Files
blog/write/.recycle/posts/202606041002/index.md
T
2026-06-05 13:58:26 +08:00

12 KiB
Raw Blame History

title, date, draft, categories, tags, author, layout, slug, status
title date draft categories tags author layout slug status
用Python给博客字体"减肥",加载速度快了一半 2026-06-04T10:00:00.000Z false
教程
折腾
Hugo
Python
字体优化
性能优化
小赵同学 post 202606041002 public

前言

折腾博客这么久了,性能优化一直是我比较在意的事情。前阵子给博客做了一波JS按需加载优化,效果还不错,首页JS从800KB直接砍到了350KB。

但是,我用Chrome DevTools的Network面板一看,好家伙,字体文件居然有1.2MB!比JS还大!

说起来,之前也写过一篇用fontspider压缩字体的文章,但那个工具年久失修,对Hugo这种静态博客支持不太好。最近正好在学Python,发现了一个更好用的工具——fonttools,今天就来分享一下,怎么用Python给字体"减肥"。

为什么中文字体这么大?

先给大家科普一下,为什么我们用的中文字体动不动就几MB。

中文字和英文字不一样。英文字母就26个大小写,加上数字和符号,撑死几百个字符。但中文呢?《通用规范汉字表》收录了8105个汉字,如果算上生僻字,得有好几万。

而字体文件里面,是把所有字符的字形都打包进去的。也就是说,不管你用没用到"龘"这个字,它都在你的字体文件里躺着,白白占空间。

但实际上,一个博客能用到多少字呢?我统计了一下我的博客,也就2500个字符左右。

1.2MB的字体,实际只需要700多KB,剩下的都是浪费。

解决方案:Python + fonttools

什么是fonttools?

fonttools是一个Python库,专门用来处理字体文件。它能做很多事情,比如读取字体信息、转换格式、提取字符等等。

我们要用到的功能是字体子集化(Font Subsetting)——简单来说,就是只保留我们用到的字符,把其他的都扔掉。

安装环境

首先,你需要安装Python。这个就不多说了,去官网下载安装就行:

https://www.python.org/downloads/

安装的时候记得勾选 "Add Python to PATH",这样就能在命令行里直接用了。

装好之后,打开cmd,输入:

python --version

能看到版本号就说明装好了。

然后安装fonttools和brotli(用于压缩):

pip install fonttools brotli

如果提示权限不够,加上 --break-system-packages:

pip install fonttools brotli --break-system-packages

编写优化脚本

接下来就是重头戏了,我写了一个Python脚本,可以自动扫描博客的所有内容,提取用到的字符,然后生成精简版字体。

在项目根目录创建一个文件:scripts/subset-font-safe.py

#!/usr/bin/env python3
"""
字体子集化脚本
从HTML和CSS文件中提取实际使用的字符,生成优化的子集字体
"""

import os
import re
from fontTools.ttLib import TTFont
from fontTools.subset import Subsetter, Options

def extract_chars_from_files(directories):
    """从文件中提取使用的字符"""
    chars = set()

    for directory in directories:
        if not os.path.exists(directory):
            print(f"⚠️  目录不存在,跳过: {directory}")
            continue

        print(f"🔍 扫描目录: {directory}")

        for root, dirs, files in os.walk(directory):
            for file in files:
                if file.endswith(('.html', '.md', '.css')):
                    filepath = os.path.join(root, file)
                    try:
                        with open(filepath, 'r', encoding='utf-8') as f:
                            content = f.read()
                            # 提取中文字符
                            chinese_chars = re.findall(r'[一-鿿]', content)
                            chars.update(chinese_chars)
                            # 提取英文和数字
                            ascii_chars = re.findall(r'[a-zA-Z0-9]', content)
                            chars.update(ascii_chars)
                            # 提取常用标点
                            en_punct = set('!@#$%^&*()_+-=[]{}|;:,.<>?/\\`~"\'-')
                            chars.update(en_punct)
                    except Exception as e:
                        print(f"⚠️  读取失败 {filepath}: {e}")

    return chars

def main():
    print("🔤 字体子集化工具")
    print("=" * 50)

    # 配置路径
    font_dir = "themes/Ying/static/font"
    input_font = os.path.join(font_dir, "zql-v2.woff2")
    chars_file = os.path.join(font_dir, "used_chars.txt")

    # 检查输入文件
    if not os.path.exists(input_font):
        print(f"❌ 字体文件不存在: {input_font}")
        return

    # 读取现有的字符列表(如果有)
    existing_chars = set()
    if os.path.exists(chars_file):
        with open(chars_file, 'r', encoding='utf-8') as f:
            existing_chars = set(f.read())
        print(f"📖 读取现有字符: {len(existing_chars)} 个")

    # 扫描目录
    scan_dirs = ["content", "layouts"]
    if os.path.exists("public"):
        scan_dirs.append("public")

    # 提取字符
    new_chars = extract_chars_from_files(scan_dirs)
    print(f"📝 从文件提取了 {len(new_chars)} 个字符")

    # 合并字符(保留手动添加的字符)
    all_chars = existing_chars | new_chars
    print(f"📊 合并后: {len(all_chars)} 个字符")

    # 保存字符列表
    with open(chars_file, 'w', encoding='utf-8') as f:
        f.write(''.join(sorted(all_chars)))

    # 生成woff2子集
    output_woff2 = os.path.join(font_dir, "zql-v2-subset.woff2")
    print(f"\n🎯 生成 woff2 子集字体...")
    
    font = TTFont(input_font)
    options = Options()
    options.flavor = 'woff2'
    
    subsetter = Subsetter(options=options)
    subsetter.populate(text=''.join(all_chars))
    subsetter.subset(font)
    font.save(output_woff2)

    # 计算优化效果
    original_size = os.path.getsize(input_font)
    subset_size = os.path.getsize(output_woff2)
    reduction = original_size - subset_size
    percentage = (reduction / original_size) * 100

    print(f"\n✅ 优化完成!")
    print(f"📊 原始大小: {original_size / 1024:.1f} KB")
    print(f"📊 子集大小: {subset_size / 1024:.1f} KB")
    print(f"📊 减少: {reduction / 1024:.1f} KB ({percentage:.1f}%)")

if __name__ == "__main__":
    main()

脚本做了什么?

这个脚本的逻辑很简单,分三步:

  1. 扫描:遍历 content、layouts、public 目录,找出所有用到的字符
  2. 提取:把中文、英文、数字、标点符号都提取出来
  3. 生成:用fonttools生成只包含这些字符的字体

这里有个细节:脚本会读取 used_chars.txt 里手动添加的字符。这样如果你有一些特殊字符(比如我博客里的古风官职表),可以手动加进去,不会被覆盖。

运行效果

首先,确保你已经构建过Hugo:

hugo --destination=public

然后运行脚本:

python scripts/subset-font-safe.py

我这边的运行结果:

🔤 字体子集化工具
==================================================
📖 读取现有字符: 2489 个
🔍 扫描目录: content
🔍 扫描目录: layouts
🔍 扫描目录: public
📝 从文件提取了 2492 个字符
📊 合并后: 2500 个字符

🎯 生成 woff2 子集字体...

✅ 优化完成!
📊 原始大小: 1226.6 KB
📊 子集大小: 740.7 KB
📊 减少: 485.9 KB (39.6%)

1.2MB的字体,优化后只剩740KB,减少了40%!

更新CSS

字体文件生成了,还需要告诉浏览器用新的字体。打开 themes/Ying/assets/css/main.css,找到字体声明:

@font-face {
    font-family: 'zql';
    src: url('../font/zql-v2.woff2') format('woff2'),
        url('../font/zql-v2.woff') format('woff');
    font-display: swap;
    unicode-range: U+0000-007F, U+4E00-9FFF, U+2000-206F, U+3000-303F;
}

改成:

@font-face {
    font-family: 'zql';
    src: url('../font/zql-v2-subset.woff2') format('woff2'),
        url('../font/zql-v2-subset.woff') format('woff');
    font-display: swap;
}

注意两个变化:

  • 文件名加了 -subset 后缀
  • 删掉了 unicode-range(因为子集字体已经只包含需要的字符了)

自动化:GitHub Actions

每次写完文章都要手动跑一遍脚本,太麻烦了。不如让GitHub Actions帮你自动做。

在 .github/workflows/ 下创建一个 subset-fonts.yml:

name: Font Subset Optimization

on:
  push:
    branches:
      - main
    paths:
      - 'content/**'
      - 'layouts/**'
  schedule:
    - cron: '0 2 * * 1'  # 每周一凌晨2点

jobs:
  subset-fonts:
    name: Subset Fonts
    runs-on: ubuntu-latest
    timeout-minutes: 30
    permissions:
      contents: write

    if: github.repository == '你的用户名/你的仓库名'

    steps:
      - name: Checkout repository
        uses: actions/checkout@v4

      - name: Set up Python
        uses: actions/setup-python@v5
        with:
          python-version: '3.11'

      - name: Install dependencies
        run: |
          pip install fonttools brotli

      - name: Build Hugo site
        uses: peaceiris/actions-hugo@v2
        with:
          hugo-version: '0.128.2'
          extended: true

      - name: Build
        run: hugo --destination=public --minify

      - name: Subset fonts
        run: python scripts/subset-font-safe.py

      - name: Commit and push
        run: |
          git config --local user.email "github-actions[bot]@users.noreply.github.com"
          git config --local user.name "github-actions[bot]"
          git add themes/Ying/static/font/zql-v2-subset.*
          git add themes/Ying/static/font/used_chars.txt
          if ! git diff --staged --quiet; then
            git commit -m "chore: update font subset [skip ci]"
            git push origin main
          fi

注意那个 [skip ci],这是为了避免触发循环——字体更新了又触发一次构建,构建完了又更新字体……

最终效果

优化前后的数据对比:

指标 优化前 优化后 提升
字体大小 1.2 MB 740 KB -40%
首页加载 ~3秒 ~1.8秒 -40%
PageSpeed得分 60 85 +42%

说实话,39.6%的减少没有达到我预期的80%+。原因是我博客里的内容比较杂,各种字符都有用到,古风官职表、数学符号、特殊标点……能优化到这个程度已经很不错了。

如果你的博客内容比较单一(比如纯技术博客),优化效果会更好,可能能到60%-70%。

常见问题

Q:优化后有些字显示不出来了?

这说明有些字符没有被提取到。解决方法:

  1. 检查 used_chars.txt,看看缺失的字符在不在里面
  2. 如果不在,手动添加进去,然后重新运行脚本
  3. 如果是API返回的动态内容,需要把API返回的字符也加入字符集

Q:脚本报错 function "try" not defined?

这是Hugo版本问题。这个脚本和Hugo版本无关,直接用Python运行就行。

Q:能不能把优化效果做得更好?

可以试试:

  • 扫描更多目录(比如API返回的数据)
  • 使用更小的字符集(比如只包含GB2312常用字6763个)
  • 用在线工具手动选择字符

写在最后

博客优化就像减肥,找对方法,效果立竿见影。

字体优化这个事情,说大不大,说小不小。单独看可能只减少了400KB,但配合之前的JS优化,整体加载时间直接砍了一半。

而且最爽的是,这些优化都是一劳永逸的。写一次脚本,配好GitHub Actions,以后写文章完全不用管,自动就给你优化好了。

如果你也被博客加载速度困扰,不妨试试这个方案。有问题欢迎评论区交流~