一品网
  • 首页

获取所有的列表


import urllib
import time
##读取指定的网址
url = []
page = 1
while page <= 11:
    url_con = urllib.urlopen('http://blog.sina.com.cn/s/articlelist_1193111400_0_'+str(page)+'.html').read()
    print 'con' ,url_con

    i = 0
    title = url_con.find(r'')

    print "title",title
    href = url_con.find(r'href=',title)
    print "href",href

    html = url_con.find(r'.html',href)
    print "html",html


    while title != -1 and href != -1 and html != -1 and i < 40:
        url.append(url_con[href+6:html+5])
        print page,url[i]
        title = url_con.find(r'',html)
        
        href = url_con.find(r'href=',title)
        
        html = url_con.find(r'.html',href)
        
        filename = url[-26:]

        i = i + 1
    else:
        print page, 'find end'
    page = page + 1
else:
    print 'all find end !'
j = 0
k = len(url)
print "url sum:",k
while j < k:
    content = urllib.urlopen(url[j]).read()
    filename = url[j][-26:]
    open(r'blog/'+ filename,'w').write(content)
    j = j + 1
    time.sleep(5)

 以上代码是获取所有博客文章列表,并读取其内容,并输出html

python爬虫Python

相关


学习《Python编程从入门到实践》PDF+代码训练

python-----面向对象简单理解

python多线程控制

Sublime 的安装、汉化、配置、Python环境和插件

python——time strftime() 函数表示当地时间

python 初识函数

python 函数对象 嵌套 闭包

Python栈溢出——设置python栈大小

python-面向对象-01课堂笔记

python爬虫

Python 之父的解析器系列之五:左递归 PEG 语法

Python 为了提升性能,竟运用了共享经济

标签

一品网 冀ICP备14022925号-6