Files
blog/write/.recycle/posts/202606041002/index.md
T
2026-06-05 13:58:26 +08:00

379 lines
12 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
---
title: 用Python给博客字体"减肥",加载速度快了一半
date: 2026-06-04T10:00:00.000Z
draft: false
categories:
- 教程
- 折腾
tags:
- Hugo
- Python
- 字体优化
- 性能优化
author: 小赵同学
layout: post
slug: '202606041002'
status: public
---
## 前言
折腾博客这么久了,性能优化一直是我比较在意的事情。前阵子给博客做了一波JS按需加载优化,效果还不错,首页JS从800KB直接砍到了350KB。
但是,我用Chrome DevTools的Network面板一看,好家伙,字体文件居然有1.2MB!比JS还大!
说起来,之前也写过一篇用fontspider压缩字体的文章,但那个工具年久失修,对Hugo这种静态博客支持不太好。最近正好在学Python,发现了一个更好用的工具——**fonttools**,今天就来分享一下,怎么用Python给字体"减肥"。
## 为什么中文字体这么大?
先给大家科普一下,为什么我们用的中文字体动不动就几MB。
中文字和英文字不一样。英文字母就26个大小写,加上数字和符号,撑死几百个字符。但中文呢?《通用规范汉字表》收录了8105个汉字,如果算上生僻字,得有好几万。
而字体文件里面,是把**所有字符的字形**都打包进去的。也就是说,不管你用没用到"龘"这个字,它都在你的字体文件里躺着,白白占空间。
但实际上,一个博客能用到多少字呢?我统计了一下我的博客,也就**2500个字符**左右。
**1.2MB的字体,实际只需要700多KB,剩下的都是浪费。**
## 解决方案:Python + fonttools
### 什么是fonttools?
fonttools是一个Python库,专门用来处理字体文件。它能做很多事情,比如读取字体信息、转换格式、提取字符等等。
我们要用到的功能是**字体子集化**(Font Subsetting)——简单来说,就是只保留我们用到的字符,把其他的都扔掉。
### 安装环境
首先,你需要安装Python。这个就不多说了,去官网下载安装就行:
https://www.python.org/downloads/
安装的时候记得勾选 **"Add Python to PATH"**,这样就能在命令行里直接用了。
装好之后,打开cmd,输入:
```bash
python --version
```
能看到版本号就说明装好了。
然后安装fonttools和brotli(用于压缩):
```bash
pip install fonttools brotli
```
如果提示权限不够,加上 `--break-system-packages`:
```bash
pip install fonttools brotli --break-system-packages
```
## 编写优化脚本
接下来就是重头戏了,我写了一个Python脚本,可以自动扫描博客的所有内容,提取用到的字符,然后生成精简版字体。
在项目根目录创建一个文件:`scripts/subset-font-safe.py`
```python
#!/usr/bin/env python3
"""
字体子集化脚本
从HTML和CSS文件中提取实际使用的字符,生成优化的子集字体
"""
import os
import re
from fontTools.ttLib import TTFont
from fontTools.subset import Subsetter, Options
def extract_chars_from_files(directories):
"""从文件中提取使用的字符"""
chars = set()
for directory in directories:
if not os.path.exists(directory):
print(f"⚠️ 目录不存在,跳过: {directory}")
continue
print(f"🔍 扫描目录: {directory}")
for root, dirs, files in os.walk(directory):
for file in files:
if file.endswith(('.html', '.md', '.css')):
filepath = os.path.join(root, file)
try:
with open(filepath, 'r', encoding='utf-8') as f:
content = f.read()
# 提取中文字符
chinese_chars = re.findall(r'[一-鿿]', content)
chars.update(chinese_chars)
# 提取英文和数字
ascii_chars = re.findall(r'[a-zA-Z0-9]', content)
chars.update(ascii_chars)
# 提取常用标点
en_punct = set('!@#$%^&*()_+-=[]{}|;:,.<>?/\\`~"\'-')
chars.update(en_punct)
except Exception as e:
print(f"⚠️ 读取失败 {filepath}: {e}")
return chars
def main():
print("🔤 字体子集化工具")
print("=" * 50)
# 配置路径
font_dir = "themes/Ying/static/font"
input_font = os.path.join(font_dir, "zql-v2.woff2")
chars_file = os.path.join(font_dir, "used_chars.txt")
# 检查输入文件
if not os.path.exists(input_font):
print(f"❌ 字体文件不存在: {input_font}")
return
# 读取现有的字符列表(如果有)
existing_chars = set()
if os.path.exists(chars_file):
with open(chars_file, 'r', encoding='utf-8') as f:
existing_chars = set(f.read())
print(f"📖 读取现有字符: {len(existing_chars)} 个")
# 扫描目录
scan_dirs = ["content", "layouts"]
if os.path.exists("public"):
scan_dirs.append("public")
# 提取字符
new_chars = extract_chars_from_files(scan_dirs)
print(f"📝 从文件提取了 {len(new_chars)} 个字符")
# 合并字符(保留手动添加的字符)
all_chars = existing_chars | new_chars
print(f"📊 合并后: {len(all_chars)} 个字符")
# 保存字符列表
with open(chars_file, 'w', encoding='utf-8') as f:
f.write(''.join(sorted(all_chars)))
# 生成woff2子集
output_woff2 = os.path.join(font_dir, "zql-v2-subset.woff2")
print(f"\n🎯 生成 woff2 子集字体...")
font = TTFont(input_font)
options = Options()
options.flavor = 'woff2'
subsetter = Subsetter(options=options)
subsetter.populate(text=''.join(all_chars))
subsetter.subset(font)
font.save(output_woff2)
# 计算优化效果
original_size = os.path.getsize(input_font)
subset_size = os.path.getsize(output_woff2)
reduction = original_size - subset_size
percentage = (reduction / original_size) * 100
print(f"\n✅ 优化完成!")
print(f"📊 原始大小: {original_size / 1024:.1f} KB")
print(f"📊 子集大小: {subset_size / 1024:.1f} KB")
print(f"📊 减少: {reduction / 1024:.1f} KB ({percentage:.1f}%)")
if __name__ == "__main__":
main()
```
### 脚本做了什么?
这个脚本的逻辑很简单,分三步:
1. **扫描**:遍历 `content`、`layouts`、`public` 目录,找出所有用到的字符
2. **提取**:把中文、英文、数字、标点符号都提取出来
3. **生成**:用fonttools生成只包含这些字符的字体
这里有个细节:脚本会读取 `used_chars.txt` 里手动添加的字符。这样如果你有一些特殊字符(比如我博客里的古风官职表),可以手动加进去,不会被覆盖。
## 运行效果
首先,确保你已经构建过Hugo:
```bash
hugo --destination=public
```
然后运行脚本:
```bash
python scripts/subset-font-safe.py
```
我这边的运行结果:
```
🔤 字体子集化工具
==================================================
📖 读取现有字符: 2489 个
🔍 扫描目录: content
🔍 扫描目录: layouts
🔍 扫描目录: public
📝 从文件提取了 2492 个字符
📊 合并后: 2500 个字符
🎯 生成 woff2 子集字体...
✅ 优化完成!
📊 原始大小: 1226.6 KB
📊 子集大小: 740.7 KB
📊 减少: 485.9 KB (39.6%)
```
**1.2MB的字体,优化后只剩740KB,减少了40%!**
## 更新CSS
字体文件生成了,还需要告诉浏览器用新的字体。打开 `themes/Ying/assets/css/main.css`,找到字体声明:
```css
@font-face {
font-family: 'zql';
src: url('../font/zql-v2.woff2') format('woff2'),
url('../font/zql-v2.woff') format('woff');
font-display: swap;
unicode-range: U+0000-007F, U+4E00-9FFF, U+2000-206F, U+3000-303F;
}
```
改成:
```css
@font-face {
font-family: 'zql';
src: url('../font/zql-v2-subset.woff2') format('woff2'),
url('../font/zql-v2-subset.woff') format('woff');
font-display: swap;
}
```
注意两个变化:
- 文件名加了 `-subset` 后缀
- 删掉了 `unicode-range`(因为子集字体已经只包含需要的字符了)
## 自动化:GitHub Actions
每次写完文章都要手动跑一遍脚本,太麻烦了。不如让GitHub Actions帮你自动做。
在 `.github/workflows/` 下创建一个 `subset-fonts.yml`:
```yaml
name: Font Subset Optimization
on:
push:
branches:
- main
paths:
- 'content/**'
- 'layouts/**'
schedule:
- cron: '0 2 * * 1' # 每周一凌晨2点
jobs:
subset-fonts:
name: Subset Fonts
runs-on: ubuntu-latest
timeout-minutes: 30
permissions:
contents: write
if: github.repository == '你的用户名/你的仓库名'
steps:
- name: Checkout repository
uses: actions/checkout@v4
- name: Set up Python
uses: actions/setup-python@v5
with:
python-version: '3.11'
- name: Install dependencies
run: |
pip install fonttools brotli
- name: Build Hugo site
uses: peaceiris/actions-hugo@v2
with:
hugo-version: '0.128.2'
extended: true
- name: Build
run: hugo --destination=public --minify
- name: Subset fonts
run: python scripts/subset-font-safe.py
- name: Commit and push
run: |
git config --local user.email "github-actions[bot]@users.noreply.github.com"
git config --local user.name "github-actions[bot]"
git add themes/Ying/static/font/zql-v2-subset.*
git add themes/Ying/static/font/used_chars.txt
if ! git diff --staged --quiet; then
git commit -m "chore: update font subset [skip ci]"
git push origin main
fi
```
注意那个 `[skip ci]`,这是为了避免触发循环——字体更新了又触发一次构建,构建完了又更新字体……
## 最终效果
优化前后的数据对比:
| 指标 | 优化前 | 优化后 | 提升 |
|------|--------|--------|------|
| 字体大小 | 1.2 MB | 740 KB | -40% |
| 首页加载 | ~3秒 | ~1.8秒 | -40% |
| PageSpeed得分 | 60 | 85 | +42% |
说实话,39.6%的减少没有达到我预期的80%+。原因是我博客里的内容比较杂,各种字符都有用到,古风官职表、数学符号、特殊标点……能优化到这个程度已经很不错了。
如果你的博客内容比较单一(比如纯技术博客),优化效果会更好,可能能到60%-70%。
## 常见问题
### Q:优化后有些字显示不出来了?
这说明有些字符没有被提取到。解决方法:
1. 检查 `used_chars.txt`,看看缺失的字符在不在里面
2. 如果不在,手动添加进去,然后重新运行脚本
3. 如果是API返回的动态内容,需要把API返回的字符也加入字符集
### Q:脚本报错 `function "try" not defined`?
这是Hugo版本问题。这个脚本和Hugo版本无关,直接用Python运行就行。
### Q:能不能把优化效果做得更好?
可以试试:
- 扫描更多目录(比如API返回的数据)
- 使用更小的字符集(比如只包含GB2312常用字6763个)
- 用在线工具手动选择字符
## 写在最后
博客优化就像减肥,找对方法,效果立竿见影。
字体优化这个事情,说大不大,说小不小。单独看可能只减少了400KB,但配合之前的JS优化,整体加载时间直接砍了一半。
而且最爽的是,这些优化都是一劳永逸的。写一次脚本,配好GitHub Actions,以后写文章完全不用管,自动就给你优化好了。
如果你也被博客加载速度困扰,不妨试试这个方案。有问题欢迎评论区交流~