Skip to content

推荐系统 · 用户画像从埋点到标签 ​

日期: 2026-07-16 | 作者: sunxin(+ Cursor AI 结对讲解) 关联路线图: 推荐系统建设路线图.md §五 Phase 3 关联产品需求: 推荐系统-使用场景清单.md §二 场景 A + §十二 依赖矩阵 关联笔记: ./05-时间衰减怎么调参.md · ./04-冷启动fallback.md 触发的 ADR: ADR-019「用户兴趣画像 T+1 全量聚合 + JSON 宽表 + 加权求和公式」

TL;DR ​

  • 用户画像(User Profile)= 从一堆散乱的行为日志里,提炼出对用户的结构化描述("用户 A 是个 25 岁北京姑娘、爱美食徒步、周末活跃、7 日活跃 45 次...")。
  • 从 0 到 1 的完整链路:埋点上报(P2 已完成)→ 加权求和 → 时间衰减 → 归一化 → 打标签 → 存画像表 → 推荐时消费。
  • 本项目 P3 采用「T+1 全量重算 + JSON 画像表 + 加权求和公式」:跟头条早期一样的朴素做法,无需 Spark / Flink / 图数据库;未来数据量上来再升级到"事件驱动增量 + 特征库"。
  • 对标:头条早期(2012-2014)= T+1 全量 + Redis 存画像;小红书早期 = MySQL JSON 画像表 + 每日 Job;抖音成熟版 = Flink 实时特征 + FeatureStore(远期)。你项目做到"头条早期"就已经够用。
  • 最容易踩的坑:以为画像越"精细"越好,堆一堆没验证的字段(如"用户是否喜欢深色主题")—— 每个画像字段必须被至少 1 个推荐场景消费,否则就是死代码。

目录 ​


1. 场景:为什么先有画像才能个性化 ​

1.1 一个非常具体的例子 ​

用户 A 的行为日志(user_action_log 表 30 天累积):

LIKE      post_p1     category=food     topic="美食探店,成都"     2026-07-01
FAVORITE  post_p2     category=food     topic="美食探店"          2026-07-02
LIKE      post_p3     category=sports   topic="徒步"              2026-07-03
COMMENT   post_p4     category=food     topic="日料"              2026-07-05
VIEW      post_p5     category=travel   topic="日本旅行"          2026-07-06
LIKE      post_p6     category=food     topic="美食探店"          2026-07-08
...(共 45 条)

问题:推荐系统怎么用这堆数据?

朴素思路 1:每次推荐都扫这 45 条 → 太慢,10 万用户同时刷首页会打死数据库

朴素思路 2:每次推荐直接 SQL SELECT * FROM posts WHERE category IN (SELECT DISTINCT category FROM user_action_log WHERE user_id=A LIMIT 10) → 还是慢,且不知道 category 之间谁更重要

正确思路:预先聚合一次,产生一份结构化画像存好:

json
{
  "user_id": "A",
  "top_categories": [
    {"category": "food",   "score": 85.5},
    {"category": "travel", "score": 42.3},
    {"category": "sports", "score": 28.1}
  ],
  "top_topics": [
    {"topic_id": "t_meishi",  "name": "美食探店", "score": 68.0},
    {"topic_id": "t_tubu",    "name": "徒步",     "score": 25.0},
    {"topic_id": "t_riben",   "name": "日本旅行", "score": 15.0}
  ],
  "active_hours": [12, 20, 22],
  "geo_city": "beijing",
  "action_count_30d": 45,
  "updated_at": "2026-07-16 02:00:00"
}

这份画像 + 兴趣召回 SQL 就能立刻用:

sql
SELECT * FROM posts 
WHERE category IN ('food','travel','sports')  -- 来自画像 top_categories
ORDER BY hot_score DESC 
LIMIT 30

1.2 用户画像 vs 推荐算法的关系 ​

很多新手会把两者混为一谈,其实分工完全不同:

组件干什么类比
用户画像描述"这个用户是谁"(静态或半静态特征)电商的"用户档案"
多路召回根据画像"给这个用户找候选内容"商品搜索
排序模型给候选内容"从最像到最不像地排序"商品排序

画像不做召回、不做排序,只做"描述"。它是所有推荐逻辑的共享输入。


2. 完整链路 · 从一行 user_action_log 到一条画像标签 ​

2.1 全景图 ​

┌──────────────────────────────────────────────┐
│  P2 已完成: user_action_log 表               │
│  用户 A 的一行: LIKE + food + 美食探店       │
└──────────────────────────────────────────────┘
                  ↓
    ┌───────────────────────────────┐
    │  P3.1  T+1 定时任务(凌晨 02:00) │
    │  遍历所有活跃用户(近 30 天有行为) │
    └───────────────────────────────┘
                  ↓
    ┌──────────────────────────────────────┐
    │  P3.2  对每个用户聚合                 │
    │  ① 查询近 30 天行为                    │
    │  ② 分组 by category + topic           │
    │  ③ 每类目/话题按公式打分:              │
    │     score = Σ (action_weight × time_decay) │
    │  ④ 归一化 → 取 top N                   │
    └──────────────────────────────────────┘
                  ↓
    ┌──────────────────────────────────────┐
    │  P3.3  写入 user_profile 表(UPSERT)   │
    │  JSON 字段: top_categories/top_topics │
    └──────────────────────────────────────┘
                  ↓
    ┌──────────────────────────────────────┐
    │  P3.4  IUserProfilePort.getProfile    │
    │  P4 兴趣召回消费画像                    │
    └──────────────────────────────────────┘

2.2 一条日志到一个分数的具体计算 ​

从一行日志到画像里 category="food" 的一个分数增量:

输入:user_action_log 单行
   user_id      = "A"
   action_type  = "LIKE"
   target_category = "food"
   action_time  = 2026-07-01 12:00:00

聚合任务 2026-07-16 凌晨跑,计算 category="food" 对用户 A 的贡献:
   action_weight("LIKE")     = 3   (来自 ActionType.defaultWeight)
   days_since(action_time)   = 15
   time_decay(15)            = 1 / (15+1)^0.5 ≈ 0.25
   
   本行贡献 = 3 × 0.25 = 0.75

对 A 的所有 category="food" 行为累加:
   food_raw_score = Σ 0.75 + 1.2 + 0.8 + ... = 42.3

归一化(除以最大类目分):
   food_normalized_score = 42.3 / max_score × 100 = 85.5

最终写入:
   top_categories = [{"category":"food", "score": 85.5}, ...]

这就是"用户画像"的全部数学。没什么神秘的,就是加权求和 + 时间衰减 + 归一化。


3. 画像结构设计 ​

3.1 画像 = 「用户属性」的集合 ​

画像由多个"维度"组成,每个维度回答一个"用户是谁"的问题:

维度回答的问题来源更新频率优先级
top_categories用户喜欢什么类目user_action_log.target_category 加权求和T+1P0
top_topics用户关注什么话题user_action_log.target_topic_ids 加权求和T+1P0
top_authors用户常互动的作者user_action_log.target_author_id 分组计数T+1P1
active_hours用户什么时段活跃user_action_log.action_time 按小时分组T+1P1
geo_city用户主要在哪个城市user_action_log 关联 posts 的 city,众数T+7P1
action_count_30d用户 30 日行为总数(冷启动降级用)count(*)T+1P0
preferred_gender用户偏好搭子性别(如有)matching 领域的历史选择T+7P2
age_range用户年龄段注册信息 + 行为反推静态P2
spending_level消费档次(贵/中/便宜)报名过的活动价格分布T+7P3

3.2 P3 起步版本只做前 4 个字段 ​

P0 优先级(P3 必须做):

  • top_categories (top 5)
  • top_topics (top 10)
  • top_authors (top 20)
  • action_count_30d

P1 优先级(P3 有余力再做):

  • active_hours
  • geo_city

P2/P3 优先级:等有明确产品需求时再加,不做没有消费方的字段。

3.3 为什么不做"性别/年龄"这种传统标签 ​

传统"CRM 用户画像"包含很多人口学标签(性别 / 年龄 / 学历 / 收入)。但推荐系统的画像不需要这些,原因:

  1. 推荐关心行为,不关心身份:一个 30 岁女性可能爱看电竞视频,一个 20 岁男性可能爱看美妆 —— 用人口学标签推等于胡推
  2. 行为标签自然覆盖了人口学信号:如果一个用户长期看"婴儿用品",大概率是新手妈妈 —— 不需要单独记"性别=女"
  3. 人口学标签合规风险大:GDPR / 中国《个保法》对性别年龄有严格约束

结论:画像里只放行为衍生标签,不放人口学标签。


4. 加权求和公式 ​

4.1 完整公式 ​

score(user, category) = Σ over action in user_action_log where target_category == category {
    weight(action.action_type) × decay(action.action_time)
}

4.2 各因子说明 ​

weight(action_type) —— 已在 P2 定义(ActionType.defaultWeight):

VIEW     : 1   (曝光信号弱,稀疏 view 已去重过)
COMMENT  : 2   (主动打字,中等强度)
LIKE     : 3   (一键表态,标准强度)
SHARE    : 4   (有社交传播意愿)
FAVORITE : 5   (最强意图 - 主动保存)

decay(action_time) —— 时间衰减,用幂次衰减公式(详见 ./05-时间衰减怎么调参.md):

decay(t) = 1 / (days_since(t) + 1) ^ 0.5

具体值:
  今天(0 天前)    = 1.0
  1 天前          = 0.71
  7 天前          = 0.35
  15 天前         = 0.25
  30 天前         = 0.18

归一化(Normalization):

normalized_score = raw_score / max(所有类目的 raw_score) × 100

归一化后所有分数在 [0, 100] 区间,最高的那个类目 = 100,其它按比例。

4.3 一个完整例子 ​

用户 A 过去 30 天行为(简化到 5 条):

时间行为categoryweightdays_sincedecay贡献
07-16LIKEfood301.003.00
07-14FAVORITEfood520.582.90
07-10LIKEsports360.381.14
07-05COMMENTfood2110.290.58
06-20LIKEtravel3260.190.57

分组求和:

  • food: 3.00 + 2.90 + 0.58 = 6.48
  • sports: 1.14
  • travel: 0.57

归一化(max = food = 6.48):

  • food: 100
  • sports: 17.6
  • travel: 8.8

写入画像:

json
"top_categories": [
  {"category": "food",   "score": 100.0},
  {"category": "sports", "score": 17.6},
  {"category": "travel", "score": 8.8}
]

这份画像的解读:用户 A 主要爱 food,sports 是弱兴趣,travel 是可有可无。P4 兴趣召回时:

  • food 分类的帖子占推荐位 60%
  • sports 占 25%
  • travel 占 10%
  • 剩余 5% 探索性内容

4.4 边界情况 ​

  • 新用户无行为:action_count_30d < 30 → 走冷启动 fallback(见 §7)
  • 单一 category 独大:例如 food = 95, others 都 < 5 → 需要强制探索机制(否则一辈子只看 food)
  • 画像抖动:某天用户偶然 like 了几条 travel → 别让画像跳变,用指数移动平均 (EMA) 平滑:
    new_score = 0.7 × old_score + 0.3 × today_score

5. 存储方案(JSON 宽表 vs 独立标签表 vs Redis Hash) ​

5.1 三种方案对比 ​

方案结构优点缺点适合规模
MySQL JSON 宽表user_profile(user_id, top_categories JSON, ...)一次读全部字段;schema 灵活;MySQL 5.7+ 有 JSON 索引单字段查询慢(全表扫);不支持"找所有 food 分数 > 80 的用户"这类反查DAU < 100w
独立标签表user_tag(user_id, tag_type, tag_value, score)支持反查;标签维度可无限扩一次拉画像要 join / group_by;SQL 复杂DAU > 100w 且有反查需求
Redis HashHGETALL profile:user:A → {cat_food: 85.5, cat_travel: 42.3, ...}微秒级读;天然支持热更新数据丢失风险;需要落地方案;结构不灵活超高频访问场景

5.2 P3 推荐用「MySQL JSON 宽表」 ​

理由(跟你项目现状匹配):

  • DAU 远低于 100w —— 用不上反查
  • MySQL 已有 —— 不引入新中间件(Redis 已用于其他场景)
  • JSON 字段可自由扩展:P3 起步只写 4 个字段,P4/P5 加字段不改表结构
  • 未来演进路径清晰:数据量大了迁 Redis Hash 或 ElasticSearch,都是独立的重构,不阻塞当前 P3

5.3 建议的表结构 ​

sql
CREATE TABLE user_profile (
    user_id VARCHAR(32) PRIMARY KEY,
    top_categories JSON COMMENT '[{"category":"food","score":85.5}, ...] 长度 <= 5',
    top_topics JSON COMMENT '[{"topic_id":"t123","name":"美食探店","score":42.0}, ...] 长度 <= 10',
    top_authors JSON COMMENT '[{"author_id":"u1","score":30.0}, ...] 长度 <= 20',
    active_hours JSON COMMENT '[9, 12, 20, 22] 长度 <= 8',
    geo_city VARCHAR(32) COMMENT '基于地理埋点聚合出的主要活跃城市',
    action_count_30d INT NOT NULL DEFAULT 0 COMMENT '近 30 天行为总数(冷启动降级用)',
    computed_at DATETIME NOT NULL COMMENT '本条画像计算时间',
    KEY idx_computed_at (computed_at)  -- 按更新时间批量刷新用
) ENGINE = InnoDB COMMENT = '用户画像 v1(JSON 宽表)';

为什么加 idx_computed_at:未来做"7 天没更新的画像重刷"批量任务时能走索引。

5.4 未来演进(不阻塞当前 P3) ​

  • DAU 5w+:加 Redis 缓存层(profile:user:{userId}),MySQL 作为持久化底座
  • DAU 50w+:user_profile 单表拆按 user_id % 8 分表
  • DAU 500w+:迁到独立的"特征库"服务(TFServing 或自建)

6. 聚合方式(T+1 全量 vs 事件驱动增量) ​

6.1 两种模式对比 ​

模式何时算优点缺点
T+1 全量每天凌晨 02:00 定时任务重算所有活跃用户逻辑简单;容易 debug;今天错了明天补数据滞后 1 天;用户"刚点赞立刻刷不到相关内容"
事件驱动增量每次用户行为 → 立刻更新画像相关字段数据实时;用户感知快逻辑复杂;容易出错;需要幂等设计
混合模式T+1 全量 + 事件驱动"局部热点补丁"两个优点都占双写复杂度高

6.2 P3 选 T+1 全量(起步) ​

理由:

  1. 朴素、够用:日活 < 1w 场景下 T+1 完全够,用户感知不出 24h 延迟
  2. 可复用已有基础设施:StatAggregationTask 每天凌晨 02:00 已在跑 stat_daily_*,画像聚合挂上去就行(不新建 @Scheduled)
  3. 符合 AGENTS.md 摸底约定:已有定时任务优先扩展而非新建

6.3 P5 再上事件驱动增量 ​

触发信号:用户反馈"我刚点赞了 A 但推荐流还是全 B"

P5 增量方案(初步设想):

  • 用户 LIKE food → 增量把 food 的 score 临时 +1
  • 临时增量只影响未来 30 分钟的推荐(用 Redis TTL 存"临时画像")
  • 不影响主 user_profile 表(避免污染 T+1 结果)
  • 30 分钟后自然过期,回归 T+1 画像

为什么不直接改主表:

  • 并发写有锁竞争
  • 若增量逻辑有 bug,会污染主画像;bug 修好后要"倒回去重算"很麻烦

这是 P5 的事。P3 只做 T+1。


7. 冷启动降级 ​

冷启动:用户画像不足时的推荐策略。详见 ./04-冷启动fallback.md。

7.1 三档策略 ​

用户状态action_count_30d推荐策略
完全冷启动0编辑精选 + 6 大分类爆款平均展示
弱画像1 ~ 3020% 兴趣召回 + 80% 全站热榜
正常画像> 30完整 4 路召回按标准权重

7.2 画像里必须有 action_count_30d 字段 ​

用途:P4 兴趣召回时判断:

java
public List<Candidate> recall(String userId) {
    UserProfile profile = userProfileRepository.get(userId);
    if (profile == null || profile.getActionCount30d() < 30) {
        // 冷启动 → 走 fallback
        return hotRecallService.recall(30);
    }
    // 正常画像 → 走兴趣召回
    return doInterestRecall(profile);
}

没有这个字段就没法判断"该走冷启动还是正常召回",这是 P3 必做的核心字段之一。


8. 对标:头条 / 小红书 / 抖音怎么做 ​

8.1 头条早期(2012-2014) ​

架构:

  • MySQL 存画像 (JSON 字段)
  • 每日 T+1 定时任务重算
  • 画像维度:兴趣分类 + 关键词 + 活跃时段
  • 无 Kafka、无 Flink、无独立特征库

给你的启发:头条 5000 万 DAU 之前都用这套。你 500 DAU 用同样的架构完全没问题。

8.2 小红书早期(2014-2017) ​

架构:

  • MySQL JSON 画像表(跟本文推荐方案一样)
  • 每日凌晨聚合 Job
  • 画像维度:category + topic + author + geo
  • 阶段性引入 Redis 缓存热点用户

给你的启发:这就是 P3 的目标。照着抄就行。

8.3 抖音成熟版(2020+) ​

架构:

  • Flink 实时计算画像特征
  • 独立特征库服务(Feature Store)
  • 用户 embedding + 视频 embedding 双塔模型
  • 画像维度:数千个

给你的启发:这是 P6 之后的远景。5 年内不用考虑。

8.4 对标总结表 ​

阶段对标画像结构更新方式存储
P3 起步(本文)头条 2012 / 小红书 20154 个 JSON 字段T+1 全量MySQL
P5 增量头条 2015 / 小红书 2017上面 + 事件驱动补丁T+1 + 增量MySQL + Redis
P6+ 触发式远景抖音 2020+数千维 embedding实时 FlinkFeature Store

9. 项目里的具体落地(P3 未来架构) ​

9.1 类目结构 ​

tour-mate-platform-domain/src/main/java/com/alisunxin/api/domain/profile/
├── model/
│   ├── aggregate/
│   │   └── UserProfileAggregate.java          // 聚合根:一个用户的完整画像
│   ├── entity/
│   │   ├── CategoryPreferenceEntity.java      // {category, score}
│   │   ├── TopicPreferenceEntity.java         // {topicId, name, score}
│   │   └── AuthorPreferenceEntity.java        // {authorId, score}
│   └── valobj/
│       └── ProfileFreshness.java              // COLD / WEAK / NORMAL
├── port/
│   ├── in/
│   │   ├── IUserProfileQueryApplicationService.java  // 查画像 (P4 消费)
│   │   └── IUserProfileAggregationApplicationService.java // 聚合入口 (StatAggregationTask 调用)
│   └── out/
│       └── IUserProfileRepository.java        // 存/取画像
├── service/
│   ├── UserProfileQueryApplicationService.java
│   ├── UserProfileAggregationApplicationService.java
│   └── ScoringStrategy.java                    // 加权求和 + 时间衰减公式

tour-mate-platform-infrastructure/src/main/java/com/alisunxin/api/infrastructure/
├── dao/po/UserProfilePO.java                  // JPA 实体
├── dao/repository/UserProfileJpaRepository.java
├── dao/converter/UserProfileConverter.java
└── adapter/repository/UserProfileRepository.java   // Port 实现

9.2 聚合逻辑(伪代码) ​

java
@Service
public class UserProfileAggregationApplicationService {
    
    private final IUserActionLogRepository actionLogRepo;
    private final IUserProfileRepository profileRepo;
    private final ScoringStrategy scoring;
    
    /**
     * StatAggregationTask 每天凌晨 02:00 调用此方法
     */
    public void aggregateAll(LocalDate statDate) {
        LocalDateTime windowEnd = statDate.plusDays(1).atStartOfDay();
        LocalDateTime windowStart = windowEnd.minusDays(30);
        
        // 1. 找出近 30 天有行为的活跃用户列表
        List<String> activeUsers = actionLogRepo.findActiveUserIds(windowStart, windowEnd);
        
        log.info("[UserProfile] 开始聚合 - 活跃用户数: {}", activeUsers.size());
        
        int successCount = 0, failCount = 0;
        for (String userId : activeUsers) {
            try {
                aggregateOne(userId, windowStart, windowEnd);
                successCount++;
            } catch (Exception e) {
                log.warn("[UserProfile] 用户画像聚合失败 - userId: {}", userId, e);
                failCount++;
            }
        }
        log.info("[UserProfile] 完成 - 成功: {}, 失败: {}", successCount, failCount);
    }
    
    private void aggregateOne(String userId, LocalDateTime from, LocalDateTime to) {
        // 2. 查询该用户近 30 天所有行为
        List<UserActionLogEntity> actions = actionLogRepo.findByUserInWindow(userId, from, to);
        
        // 3. 按 category 聚合
        Map<String, Double> categoryScores = new HashMap<>();
        Map<String, Double> topicScores = new HashMap<>();
        Map<String, Double> authorScores = new HashMap<>();
        
        for (UserActionLogEntity action : actions) {
            double contribution = scoring.calculateContribution(
                action.getActionType(), action.getActionTime());
            
            // category
            if (action.getTargetCategory() != null) {
                categoryScores.merge(action.getTargetCategory(), contribution, Double::sum);
            }
            
            // topics(逗号分隔展开)
            if (action.getTargetTopicIds() != null) {
                for (String topicId : action.getTargetTopicIds().split(",")) {
                    topicScores.merge(topicId, contribution, Double::sum);
                }
            }
            
            // author
            if (action.getTargetAuthorId() != null) {
                authorScores.merge(action.getTargetAuthorId(), contribution, Double::sum);
            }
        }
        
        // 4. 归一化 + 取 top N
        List<CategoryPreference> topCategories = scoring.normalize(categoryScores, 5);
        List<TopicPreference> topTopics = scoring.normalize(topicScores, 10);
        List<AuthorPreference> topAuthors = scoring.normalize(authorScores, 20);
        
        // 5. UPSERT 画像表
        UserProfileAggregate profile = UserProfileAggregate.builder()
            .userId(userId)
            .topCategories(topCategories)
            .topTopics(topTopics)
            .topAuthors(topAuthors)
            .actionCount30d(actions.size())
            .computedAt(LocalDateTime.now())
            .build();
        profileRepo.upsert(profile);
    }
}

9.3 与 P2 已有基础设施的融合点 ​

P2 已有P3 复用方式
user_action_log 表聚合任务的数据源
IUserActionLogRepository.findByUserInWindow已提供,P3 直接调用
ActionType.defaultWeight加权求和公式直接用
StatAggregationTask @Scheduled(02:00)挂 aggregateAll 上去,不新建 @Scheduled

特别注意:StatAggregationTask 已经在跑 statAggregationApplicationService.aggregateDaily(yesterday),P3 加一句 userProfileAggregationApplicationService.aggregateAll(yesterday) 就行 —— 严格遵守 AGENTS.md 摸底约定。

9.4 逐步上线路径 ​

Milestone 1:新建 user_profile 表 + JPA 骨架(1h) Milestone 2:ScoringStrategy 加权求和 + 归一化(1h) Milestone 3:UserProfileAggregationApplicationService.aggregateAll(2h) Milestone 4:接入 StatAggregationTask(20min) Milestone 5:手工触发一次,验证 4 个 JSON 字段格式(30min) Milestone 6:IUserProfileQueryApplicationService 提供 getProfile(userId)(30min) Milestone 7:ADR-019 起草(1h) Milestone 8:sync/current.md + 路线图更新(20min)

总计约 6-7 小时,跟 P1/P2 的节奏一致。


10. 教学收获 ​

10.1 「画像」本质是"降维" ​

用户 45 条散乱行为 → 4 个 JSON 字段几百字节 = 信息压缩 100 倍。

推荐系统的所有性能优化,本质都是"降维":

  • 全量行为日志 → 用户画像(P3)
  • 全站帖子 → 每类目热榜 top 500(P4 热门召回)
  • 全网关注关系 → 每人 top 100 关注对象的信箱(P1 写扩散)

掌握"降维"思路 = 掌握推荐系统的技术直觉。

10.2 「T+1 全量重算」的价值远超"实时" ​

新手常问:为什么不实时?

答:因为 T+1 全量能重来。如果昨天聚合逻辑有 bug,今天改完重跑一遍就好。事件驱动增量出了 bug,得倒着找到底哪些数据被污染了、怎么修 —— 复杂度上一个量级。

T+1 全量是可靠性、可 debug 性的巨大红利。除非有强产品需求,都优先选 T+1。

10.3 「不加没有消费方的字段」的克制 ​

新手做画像会有"越全越好"的冲动,加一堆字段:是否喜欢深色主题 / 是否浏览过视频类内容 / 平均评论字数...

这些都是死代码。每个字段必须能回答"P4/P5 的哪个场景要用它",否则不加。

画像 schema 是产品需求的映射,不是数据科学家的自嗨。

10.4 「归一化」的必要性 ​

不做归一化,各类目的原始分数(raw_score)随用户行为量线性增长 —— 一个 3 年老用户的 food 分是 1500,一个 3 天新用户的 food 分是 30。跨用户不可比 = 无法做协同过滤。

归一化后所有分数都在 [0, 100] 区间,跨用户可比 —— 这是很多下游算法的基础前提。


11. 附录 ​

11.1 术语对照 ​

术语中文说明
User Profile用户画像结构化描述用户偏好
Feature特征画像里的单个维度(如 top_categories)
Tag标签特征的单个值(如 food)
Weight权重每种行为对画像的贡献强度
Time Decay时间衰减越老的行为权重越低
Normalization归一化把不同量级的分数拉到统一区间
Cold Start冷启动无画像用户的降级策略
Feature Store特征库存 embedding 的独立服务(远期)

11.2 延伸阅读 ​

  • 《大数据日知录》第 12 章「数据分析算法」—— 特征工程和画像基础
  • 《推荐系统实战》第 3 章「用户行为数据的洞察」
  • 《精益数据分析》—— 什么样的数据值得跟踪
  • 头条早期算法架构分享(InfoQ 2017 有一场)

11.3 项目内关联文档 ​

Powered by VitePress