词云绘制

第三方库 wordcloud pip install wordcloud 指定镜像源: pip install -i https://pypi.tuna.tsinghua.edu.cn/simple wordcloud 文档:https://amueller.github.io/word_cloud/index.html wordcloud.WordCloud() 案例 1:”政府工作报告爬取与词云绘制“ python 1 2 3 4 5 6 7 8 9 10 11 12 13 import urllib.request from bs4 import BeautifulSoup from wordcloud import WordCloud url = "https://www.gov.cn/zhuanti/2021lhzfgzbg/index.htm" response = urllib.request.urlopen(url) html = response.read().decode("utf-8") soup = BeautifulSoup(html, "html.parser") content = soup.find("div", class_="zhj-bbqw-cont").text w = WordCloud(font_path="/Fonts/simhei.ttf").generate(content) w.to_file("政府工作报告y1.png") ...

2026年1月28日 · ☕☕ 6 min · 📄 2.6k 字 · Python爬虫

动态数据爬取

单个城市天气数据爬取 确定目标网页 https://www.weather.com.cn/ 分析网页数据 python 1 2 3 4 5 6 7 8 9 10 11 12 import requests from bs4 import BeautifulSoup myHeader = {"User-Agent": "Mozilla/5.0 (Windows NT 10.0; Win64; x64)"} url = "https://www.weather.com.cn/weather1d/101010100.shtml" r = requests.get(url) html = r.content.decode('utf-8') soup = BeautifulSoup(html, "html.parser") print(soup.find('div', class_='tem')) ...

2026年1月28日 · ☕☕☕☕☕ 75 min · 📄 3.7 万字 · Python爬虫

Scrapy 爬虫框架

Scrapy 爬虫框架介绍 官网:https://www.scrapy.org/ 文档:https://docs.scrapy.net.cn/en/latest/ 快速功能强大的网络爬虫框架 Scrapy 的安装 pip install scrapy scrapy -h Scrapy 爬虫框架结构 Scrapy不是一个函数功能库,而是一个爬虫框架。 ...

2026年1月27日 · ☕☕ 4 min · 📄 1.8k 字 · Python爬虫

Re 库入门

正则表达式 regular expression, regex, RE 正则表达式是用来简洁表达一组字符串的表达式 正则表达式是一种针对字符串表达“简洁”和“特征”思想的工具 正则表达式可以用来判断某字符串的特征归属 ...

2026年1月26日 · ☕☕ 7 min · 📄 3.3k 字 · Python爬虫

信息标记与提取方法

信息标记的三种形式 信息的标记 标记后的信息可形成信息组织结构,增加了信息维度 标记的结构与信息一样具有重要价值 标记后的信息可用于通信、存储或展示 标记后的信息更利于程序理解和运用 ...

2026年1月26日 · ☕☕ 5 min · 📄 2.0k 字 · Python爬虫

Beautiful Soup 库入门

Beautiful Soup 库入门 官网:https://www.crummy.com/software/BeautifulSoup/ You didn’t write that awful page. You’re just trying to get some data out of it. Beautiful Soup is here to help. Since 2004, it’s been saving programmers hours or days of work on quick-turnaround screen scraping projects. Beautiful Soup is a Python library designed for quick turnaround projects like screen-scraping. Three features make it powerful: Beautiful Soup provides a few simple methods and Pythonic idioms for navigating, searching, and modifying a parse tree: a toolkit for dissecting a document and extracting what you need. It doesn’t take much code to write an application Beautiful Soup automatically converts incoming documents to Unicode and outgoing documents to UTF-8. You don’t have to think about encodings, unless the document doesn’t specify an encoding and Beautiful Soup can’t detect one. Then you just have to specify the original encoding. Beautiful Soup sits on top of popular Python parsers like lxml and html5lib, allowing you to try out different parsing strategies or trade speed for flexibility. Beautiful Soup parses anything you give it, and does the tree traversal stuff for you. You can tell it “Find all the links”, or “Find all the links of class externalLink”, or “Find all the links whose urls match “foo.com”, or “Find the table heading that’s got bold text, then give me that text.” ...

2026年1月25日 · ☕☕ 5 min · 📄 2.3k 字 · Python爬虫

Requests 库入门

https://python-requests.org/ Requests 库入门 安装:pip install requests 基本使用 python 1 2 3 4 5 6 import requests r = requests.get("http://www.baidu.com") r.status_code 200 r.encoding = 'utf-8' r.text ...

2026年1月25日 · ☕☕☕☕ 11 min · 📄 5.1k 字 · Python爬虫

前言-Python数据爬取与可视化

本部分是 MOOC中的《Python数据爬取与可视化》笔记 课程链接:https://www.icourse163.org/course/NHDX-1463126169?tid=1476402447

2026年1月25日 · ☕ 1 min · 📄 97 字 · Python爬虫

前言-Python网络爬虫与信息提取

本部分是 MOOC中的《Python网络爬虫与信息提取》笔记 课程链接:https://www.icourse163.org/learn/BIT-1001870001?tid=1475660446#/learn/content ...

2026年1月25日 · ☕ 1 min · 📄 159 字 · Python爬虫

语雀导出MD图片自动上传工具

前言 事情的起因是语雀导出的 md 文件的图片自带防盗链,所以想把文章发往博客还得需要人工手动复制图片再上传一次。当文章的图片很多时,就会十分吃力。 所以本人借助 AI 做了这样的一个脚本。 ...