热烈祝贺台州朗动科技的站长论坛隆重上线!(2012-05-28)    热烈庆祝伟大的祖国60周年生日 点击进来我们一起为她祝福吧(2009-09-26)    站长论坛禁止发布广告,一经发现立即删除。谢谢各位合作!.(2009-08-08)    热烈祝贺台州网址导航全面升级,全新版本上线!希望各位一如既往地支持台州网址导航的发展.(2009-03-28)    台州站长论坛恭祝各位新年快乐,牛年行大运!(2009-01-24)    台州Link正式更名为台州网址导航,专业做以台州网址为主的网址导航!(2008-05-23)    热烈祝贺台州Link资讯改名为中国站长资讯!希望在以后日子里得到大家的大力支持和帮助!(2008-04-10)    热烈祝贺台州Link论坛改名为台州站长论坛!希望大家继续支持和鼓励!(2008-04-10)    台州站长论坛原[社会琐碎]版块更名为[生活百科]版块!(2007-09-05)    特此通知:新台州站长论坛的数据信息全部升级成功!">特此通知:新台州站长论坛的数据信息全部升级成功!(2007-09-01)    台州站长论坛对未通过验证的会员进行合理的清除,请您谅解(2007-08-30)    台州网址导航|上网导航诚邀世界各地的网站友情链接和友谊联盟,共同引领网站导航、前进!(2007-08-30)    禁止发广告之类的帖,已发现立即删除!(2007-08-30)    希望各位上传与下载有用资源和最新信息(2007-08-30)    热烈祝贺台州站长论坛全面升级成功,全新上线!(2007-08-30)    
便民网址导航,轻松网上冲浪。
台州维博网络专业开发网站门户平台系统
您当前的位置: 首页 » JAVA/JSP编程 » Java 正则表达式解析 Html

Java 正则表达式解析 Html

论坛链接
  • Java 正则表达式解析 Html
  • 发布时间:2009-04-24 13:00:48    浏览数:8101    发布者:superadmin    设置字体【   
去年在 Uptech 的时候写过一个开源的 XMPP Robat ,当时有一个搜索天气信息的功能,我用了 HtmlParser 来解析网页,说实话 HtmlParser 的确不错,只是我没什么时间琢磨他,使用还不习惯,所以现在换成正责表达式来解析网页,其实是想尝试尝试一下,现在解析天气预报信息的方式已从 HtmlParser 转移到了 Java 正则表达式, 这是刚实现的一段代码,贴出来共享 ...

/**
* Copyright (C) 2006 the original author or authors.
*
* This software is published under the terms of the GNU Public License (GPL),
* a copy of which is included in this distribution.
*/

package com.boar.modules;

import java.io.BufferedReader;
import java.io.InputStreamReader;
import java.net.URL;
import java.net.URLConnection;
import java.util.Map;
import java.util.concurrent.ConcurrentHashMap;
import java.util.concurrent.ScheduledExecutorService;
import java.util.concurrent.ScheduledThreadPoolExecutor;
import java.util.concurrent.TimeUnit;

import java.util.regex.Matcher;
import java.util.regex.Pattern;

import com.boar.container.BasicModule;

/**
* @author <a href="Ben">zhuaming@gmail.com">Ben </a>
*/
public class WeatherModule extends BasicModule{

private static ScheduledExecutorService executor = null;
private static Map cache = new ConcurrentHashMap();

private static final String urlLink =
"http://weather.tq121.com.cn/mapanel/index.php?city=";

private static final String citys[] = {
"北京", "哈尔滨", "长春", "沈阳", "大连",
"天津", "呼和浩特","乌鲁木齐", "西宁", "银川",
"兰州", "西安", "拉萨", "成都","重庆", "贵阳",
"昆明", "太原", "石家庄", "济南", "青岛", "郑州",
"合肥", "南京", "徐州", "连云港", "上海", "武汉",
"长沙", "南昌", "杭州", "福州", "厦门", "台北",
"南宁", "桂林", "海口", "三亚", "广州", "香港", "澳门"
};

public WeatherModule() {
super(" Weather Module");
}

public void start(){
executor = new ScheduledThreadPoolExecutor(1);
executor.scheduleWithFixedDelay(new WeatherMonitor(), 0, 60 * 60, TimeUnit.SECONDS);
}

public void stop() {
if (executor != null) {
executor.shutdown();
}
if (cache != null){
cache.clear();
}
}

public String search(String city){

if (cache.containsKey(city.trim())){
return cache.get(city.trim()).toString();
}
return " Not support .";
}


private class WeatherMonitor implements Runnable {

public void run() {
cache.clear();
parse();
}

private String getWeather(String pattern, String match){
Pattern sp = Pattern.compile(pattern);
Matcher matcher = sp.matcher(match);
while(matcher.find()){
return matcher.group(1);
}
return "";
}

private void parse() {
for(int i=0;i<= citys.length-1;i++){
StringBuffer pageBuffer = new StringBuffer();
try {
URL url = new URL(urlLink + citys);
URLConnection ret = url.openConnection();
String input ;

BufferedReader br = new BufferedReader(new InputStreamReader(ret.getInputStream()));
while((input = br.readLine()) != null) {
pageBuffer.append(input);
}
}catch(Exception e){
System.out.println(e.getMessage());
}

StringBuffer weatherBuffer = new StringBuffer();

weatherBuffer.append(getWeather("<td width=\"163\" align=\"center\" valign=\"top\"><span class=\"big-cn\">(.*?)</span>",pageBuffer.toString()));
weatherBuffer.append(getWeather("<td width=\"160\" align=\"center\" valign=\"top\" class=\"weather\">(.*?)</td>",pageBuffer.toString()));
weatherBuffer.append(getWeather("<td width=\"153\" valign=\"top\"><span class=\"big-cn\">(.*?)</span>",pageBuffer.toString()));
weatherBuffer.append(getWeather("class=\"weatheren\">(.*?)</td>",pageBuffer.toString()));

cache.put(citys, weatherBuffer.toString());
}
}
}
}
娱乐休闲专区A 影视预告B 音乐咖啡C 英语阶梯D 生活百科
网页编程专区E AMPZF HTMLG CSSH JSI ASPJ PHPK JSPL MySQLM AJAX
Linux技术区 N 系统管理O 服务器架设P 网络/硬件Q 编程序开发R 内核/嵌入
管理中心专区S 发布网址T 版主议事U 事务处理