python教程怎么抓起数据_介绍python 数据抓取三种方法
三種數據抓取的方法正則表達式(re庫)
BeautifulSoup(bs4)
lxml
*利用之前構建的下載網頁函數,獲取目標網頁的html,我們以https://guojiadiqu.bmcx.com/AFG__guojiayudiqu/為例,獲取html。
from get_html import download
url = 'https://guojiadiqu.bmcx.com/AFG__guojiayudiqu/'page_content = download(url)
*假設我們需要爬取該網頁中的國家名稱和概況,我們依次使用這三種數據抓取的方法實現數據抓取。
1.正則表達式from get_html import downloadimport re
url = 'https://guojiadiqu.bmcx.com/AFG__guojiayudiqu/'page_content = download(url)country = re.findall('class="h2dabiaoti">(.*?)', page_content) #注意返回的是listsurvey_data = re.findall('
(.*?)', page_content)survey_info_list = re.findall('(.*?)
', survey_data[0])survey_info = ''.join(survey_info_list)print(country[0],survey_info)2.BeautifulSoup(bs4)from get_html import downloadfrom bs4 import BeautifulSoup
url = 'https://guojiadiqu.bmcx.com/AFG__guojiayudiqu/'html = download(url)#創建 beautifulsoup 對象soup = BeautifulSoup(html,"html.parser")#搜索country = soup.find(attrs={'class':'h2dabiaoti'}).text
survey_info = soup.find(attrs={'id':'wzneirong'}).textprint(country,survey_info)
3.lxmlfrom get_html import downloadfrom lxml import etree #解析樹url = 'https://guojiadiqu.bmcx.com/AFG__guojiayudiqu/'page_content = download(url)selector = etree.HTML(page_content)#可進行xpath解析country_select = selector.xpath('//*[@id="main_content"]/h2') #返回列表for country in country_select:
print(country.text)survey_select = selector.xpath('//*[@id="wzneirong"]/p')for survey_content in survey_select:
print(survey_content.text,end='')
運行結果:
最后,引用《用python寫網絡爬蟲》中對三種方法的性能對比,如下圖:
僅供參考。相關免費學習推薦:python教程(視頻)
總結
以上是生活随笔為你收集整理的python教程怎么抓起数据_介绍python 数据抓取三种方法的全部內容,希望文章能夠幫你解決所遇到的問題。
- 上一篇: java的圆周率_java学习日记,圆周
- 下一篇: authenticationstring