User portraying method based on deep learning
By leveraging deep learning's multimodal process understanding and modular tiered quantization capability model, this approach addresses the shortcomings of existing user profiling methods, such as insufficient real-time perception and fragmented multi-source evidence. It enables real-time updates of user profiles and robust fusion of multi-source evidence, enhancing the stability and interpretability of user profiles and supporting continuous optimization and management decision-making.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-25
- Publication Date
- 2026-04-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing user profiling methods lack the ability to perceive the actual activity process in real time. The fragmentation of multi-source evidence and the lack of constraints on evidence quality and conflict make it difficult to form a systematic model for user profiling. This results in insufficient stability and weak interpretability of the profiling results, making it difficult to support continuous optimization and refined management.
By adopting a deep learning-based multimodal process understanding and modular tiered quantification capability model, a process evidence set is generated by collecting a multimodal input set. The modular tiered quantification capability model is then used for evidence cleaning and merging to achieve the evidence-based, traceable, and closed-loop updating of user profiles.
It has enabled the transformation of user profiling from post-event statistics to process-driven and evidence-traceable methods, improving the objectivity, timeliness, and scenario adaptability of the profiling, ensuring the hierarchical quantification and robust integration of multi-source evidence, and providing reliable management decision support.
Smart Images

Figure CN121786734A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of lean management of work teams, and in particular to a user profiling method based on deep learning. Background Technology
[0002] In existing technologies, user profiling methods are mostly applied to scenarios such as human resource management, safety production management, or training assessment. They are typically based on statistical analysis and rule modeling using structured business data or manually entered information. For example, user competency tags are generated from training scores, assessment records, inspection results, or manual scoring sheets, or post-event summaries are made from documents such as meeting minutes and inspection records. The data sources for these methods are mainly offline data and manual processing, and profile updates rely on periodic manual maintenance. They lack the ability to perceive the real-time process of actual activities and are difficult to objectively reflect the actual behavior and status changes of personnel in dynamic scenarios such as meetings, discussions, and collaborations.
[0003] Meanwhile, some existing technologies have begun to incorporate speech recognition, video analysis, or multimodal learning techniques, but these mostly remain at the single-task level, such as meeting transcription, emotion recognition, or behavior recognition, failing to form a systematic modeling mechanism for user profiling. Existing solutions generally suffer from problems such as fragmented multi-source evidence, lack of constraints on evidence quality and conflict, and loosely structured capability assessment models. They struggle to uniformly quantify and traceably integrate multimodal process evidence with historical capability data, and lack an effective closed-loop mechanism during profile updates, resulting in insufficient stability and weak interpretability of profile results, making it difficult to support continuous optimization and refined management needs.
[0004] Therefore, how to provide a user profiling method based on deep learning is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0005] One objective of this invention is to propose a user profiling method based on deep learning. This invention is based on the multimodal process understanding and modular tiered quantification capability model of deep learning to realize the evidence-based, traceable and closed-loop update of user profiles, and has strong objectivity, high stability and continuous optimization capabilities.
[0006] A user profiling method based on deep learning according to an embodiment of the present invention includes the following steps: Collect video streams, audio streams, participant identity information, and participation time sequence information of team activities, align execution time and unify format, and generate a multimodal input set for team activities; Input the multimodal input set of team activities into the multimodal real-time understanding model of the meeting process to generate a set of process evidence; Collect training results, task quality, inspection records and meeting contributions corresponding to the participant identity information, unify the execution criteria and remove missing data, and generate a set of personal capability source data. Perform evidence cleaning and merging on the process evidence set to generate a profile evidence feature set; Construct a modular, tiered quantitative capability model, and obtain a four-dimensional profile of personal capabilities based on the profile evidence feature set and the personal capability source data set. The feature set of profile evidence for subsequent team activities and the set of personal ability source data are input again into the modular tiered quantitative ability model to update the four-dimensional profile set of personal ability and output the closed-loop profile update result.
[0007] Optionally, the generation of the multimodal input set for team activities specifically includes: collecting the video and audio streams of team activities and writing them with timestamps; collecting participant identity information and participation timing information; performing integrity checks and quality rejections; generating a set of valid video streams and a set of valid audio streams; extracting accompanying audio tracks from the set of valid video streams and calculating time offsets with the set of valid audio streams, aligning them with the time of writing back; and performing format unification and encapsulation of participant identity information and participation timing information to form the multimodal input set for team activities.
[0008] Optionally, the generation of the process evidence set specifically includes: The multimodal input set of team activities is input into the real-time multimodal understanding model of the meeting process. Multimodal segmentation and index binding are performed to obtain a set of video frame sequences, a set of audio segments, and a set of multimodal alignment indexes. Visual and acoustic features are extracted from video frame sequence sets and audio segment sets respectively based on the multimodal alignment index set, generating visual feature sequence sets and acoustic feature sequence sets. Speaker embedding vector sets and speaker similarity matrices are generated based on the acoustic feature sequence sets. Clustering and temporal backfilling are performed to obtain speaker differentiation results consistent with the meeting time sequence information. The acoustic feature sequence set and the speaker discrimination result are input into the speech transcription link to generate speech transcription results. Based on the speaker discrimination result, the speech transcription results are segmented and merged to obtain the speaker-aligned transcription set. Input the speaker alignment transcription set, acoustic feature sequence set and visual feature sequence set into the sentiment tendency recognition link and behavior state recognition link to generate sentiment tendency results and behavior state recognition results consistent with the meeting time sequence information; The speaker differentiation results, speech transcription results, sentiment tendency results, and behavioral state recognition results are input into the structured meeting summary generation link to generate a structured meeting summary. The summary is then encapsulated according to the multimodal alignment index set to obtain the process evidence set.
[0009] Optionally, the generation of the personal capability source dataset specifically includes: Establish cross-system identity mapping relationships based on meeting participant identity information, and uniformly map personnel identifiers in the training system, task system and inspection system to meeting participant identity information, and jointly collect corresponding training results, task quality, inspection records and meeting contributions. Extract assessment records and performance records from training results and ensure consistent application of standards to form a standardized collection of training results; The standardization of the criteria involves field standardization, time window alignment, and evaluation scale normalization. Extract task completion records, review records, and rework records for task quality and implement unified standards, then map them into a task quality score set according to a unified scoring rule. The unified scoring rule maps heterogeneous capability data into capability scores with unified dimensions based on calculable indicators such as completion rate, pass rate, proportion, and frequency. Extract the inspection point arrival records, hidden danger reporting records and review conclusions from the inspection records and implement unified standards, and map them into a set of inspection record scores according to unified scoring rules; Based on the speaker differentiation results, speech transcription results, behavioral state recognition results, and structured meeting summaries in the process evidence set, meeting contribution features are extracted and standardized to form a meeting contribution score set; The standardized training results set, task quality score set, inspection record score set, and meeting contribution score set are used to limit the effective time window based on the meeting time sequence information and perform missing data removal to generate a personal capability source data set.
[0010] Optionally, the generation of the portrait evidence feature set specifically includes: The process evidence set is indexed and expanded based on the multimodal aligned index set to form a set of evidence items bound by the participant identity information and the participation time sequence information; Perform integrity checks and null value removal on the evidence item set, perform field standardization and time window alignment on each field, and generate an aligned evidence item set; A quality assessment is performed on the aligned evidence item set. Based on the proportion of effective segments in the speech transcription results, the segment continuity of the speaker differentiation results, the temporal stability of the emotion tendency results, and the duration constraints of the behavioral state recognition results, an evidence quality label set is generated and written back to the aligned evidence item set to form a labeled evidence item set. Noise suppression is performed on the set of labeled evidence items. Items that do not meet the validity constraints in the evidence quality label set are removed. Temporal smoothing is performed on the sentiment tendency results and behavioral state recognition results that meet the validity constraints but have isolated spikes, and a cleaned set of evidence items is generated. Evidence merging is performed on the cleaned evidence item set. The speech transcription results and structured meeting summaries of the same participant identity information in the same participant time sequence information window are semantically deduplicated and merged. The sentiment tendency results and behavioral state recognition results in the same participant time sequence information window are aggregated and statistically analyzed to generate merged statistical fields, thus obtaining the merged evidence item set. Perform cross-entry consistency checks on the merged evidence entry set, generate a consistency check result set based on the segment alignment relationship between the speaker differentiation results and the speech transcription results, and the key point coverage relationship between the behavior state recognition results and the structured meeting summary, and write it back to the merged evidence entry set to form a consistent merged evidence entry set; Based on a consistent set of merged evidence items, the speech transcription results and structured meeting summaries are encoded as text evidence features, the sentiment tendency results and behavioral state recognition results are encoded as behavioral evidence features, and the text evidence features and behavioral evidence features are indexed and encapsulated according to the participant identity information and the participation time sequence information to generate a set of profile evidence features.
[0011] Optionally, the generation of the four-dimensional profile of an individual's capabilities specifically includes: A modular, tiered quantitative capability model is constructed. The set of profile evidence features and the set of personal capability source data are input into the cross-source alignment layer. Based on the participant identity information and the participant time sequence information, cross-source index alignment, scale alignment and semantic alignment are performed to generate a joint representation set. The modular tiered quantification capability model consists of a cross-source alignment layer, an evidence block library, a tiered capability ladder, a ladder gating device, a ladder recursion layer, a dimension assembly layer, an evidence conflict detector, a conflict suppression mechanism, and a profile output head. An evidence block library is constructed. The joint representation set is decomposed according to the evidence source and evidence form and written into the evidence block type set. Each evidence block encapsulates the participant identity information, participation time sequence information, source tag and quality tag to form a searchable evidence block index. A set of evidence contributions is generated based on the evidence block library, which includes evidence block indexes, evidence contribution weights, and contribution type labels. Input the evidence contribution set into the hierarchical ability ladder, establish multi-level ability slots of the hierarchical ability ladder based on the dimension set corresponding to the four-dimensional profile of personal ability, and route the evidence contribution set to the corresponding ability slot according to the contribution type label to form the hierarchical input set. The tiered input set is input into the tiered gating system. Based on the temporal adjacency relationship between the evidence contribution weight and the meeting time sequence information, the gating selection result is generated. Gating suppression is performed on evidence blocks below the contribution weight threshold, and gating retention is performed on evidence blocks above the contribution weight threshold. The gating tiered input set is input into the tiered recursive layer to perform cross-window recursive fusion to generate a tiered capability set, which includes the hierarchical capability representation and hierarchical contribution index of each dimension. The tiered capability set is input into the dimensional assembly layer. Based on the participant identity information, the hierarchical capability representations of each dimension are aligned and assembled to generate a four-dimensional assembled representation set. A four-dimensional profile of personal capabilities is generated through the profile output head. The evidence contribution set, gating selection result and tiered capability set are input into the evidence conflict detector in parallel. Conflict detection is triggered based on the consistency relationship of evidence blocks within the same meeting time sequence information window under the same meeting identity information, and the conflict evidence index set and conflict type set are obtained. When the evidence conflict detector triggers conflict detection, it calls the conflict suppression mechanism to perform a set of suppression actions on the evidence blocks corresponding to the conflict evidence index set. The suppressed evidence contribution set is then fed back to the ladder gating and ladder recursion layer to update the gating step input set and ladder capability set. After passing through the dimension assembly layer and the portrait output head, the personal capability four-dimensional portrait set is re-output.
[0012] Optionally, the generation of the closed-loop profile update result specifically includes: taking the profile evidence feature set generated by subsequent team activities and the personal ability source dataset as incremental input, and inputting it again into the building block tiered quantitative ability model. Under the premise of keeping the model structure and parameter configuration unchanged, performing cross-source alignment and time positioning on the new evidence based on the participant identity information and participation time sequence information, writing the new evidence into the evidence block library and participating in the update calculation of the evidence contribution set, performing time-series recursive fusion of the ability representation of the new evidence and the existing evidence through the tiered ability ladder, ladder gating and ladder recursive layer, generating the updated tiered ability set, and outputting the updated personal ability four-dimensional profile set through the dimension assembly layer and profile output head. At the same time, the evidence conflict detector detects the consistency between the new evidence and the existing evidence. When a conflict is triggered, the conflict suppression mechanism is called to suppress the conflicting evidence and update it back, forming a closed-loop profile update result that includes the ability profile result, the evidence contribution tracing relationship and the conflict annotation information.
[0013] The beneficial effects of this invention are: This invention transforms user profiling from "post-event statistics" to "process-driven, evidence-traceable" by introducing multimodal real-time understanding of meeting processes and a modular, tiered quantification capability model. Through real-time analysis of multimodal activities such as meeting videos and audio, process information such as speaker identification, speech transcription, emotional tendencies, and behavioral states is uniformly encapsulated into structured process evidence. This evidence is then cross-sourced and integrated with historical capability data from training, tasks, and inspections. This allows the profiling results to directly reflect an individual's participation, performance, and capability contributions in real-world business scenarios, significantly improving the objectivity, timeliness, and scenario adaptability of user profiling.
[0014] Furthermore, this invention achieves hierarchical quantification and robust fusion of multi-source evidence through an evidence block library, a tiered capability ladder, and a conflict detection and suppression mechanism. It also constructs a closed-loop profile update chain, enabling newly added activity evidence to continuously participate in capability assessment without disrupting the existing model structure. This mechanism not only improves the stability and consistency of profile results over long-term evolution but also achieves traceability of evidence contributions and controllable backflow of conflict resolution, providing reliable, interpretable, and sustainably optimizable technical support for management decision-making, training recommendations, and security control. Attached Figure Description
[0015] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0016] Figure 1 This is a flowchart of a deep learning-based user profiling method proposed in this invention; Figure 2 This is a schematic diagram illustrating the relationship between multi-source evidence generation and encapsulation in a meeting process using a deep learning-based user profiling method proposed in this invention. Figure 3 This is a schematic diagram of the structure of a modular, tiered quantization capability model for a user profiling method based on deep learning proposed in this invention. Detailed Implementation
[0017] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0018] refer to Figures 1-3 A user profiling method based on deep learning includes the following steps: Collect video streams, audio streams, participant identity information, and participation time sequence information of team activities, align execution time and unify format, and generate a multimodal input set for team activities; Input the multimodal input set of team activities into the multimodal real-time understanding model of the meeting process to generate a set of process evidence; Collect training results, task quality, inspection records and meeting contributions corresponding to the participant identity information, unify the execution criteria and remove missing data, and generate a set of personal capability source data. Perform evidence cleaning and merging on the process evidence set to generate a profile evidence feature set; Construct a modular, tiered quantitative capability model, and obtain a four-dimensional profile of personal capabilities based on the profile evidence feature set and the personal capability source data set. The feature set of profile evidence for subsequent team activities and the set of personal ability source data are input again into the modular tiered quantitative ability model to update the four-dimensional profile set of personal ability and output the closed-loop profile update result.
[0019] In this embodiment, the generation of the multimodal input set for team activities specifically includes: collecting the video and audio streams of team activities and writing them with timestamps; collecting participant identity information and participation timing information; performing integrity checks and quality rejections; generating a set of valid video streams and a set of valid audio streams; extracting accompanying audio tracks from the set of valid video streams and calculating time offsets with the set of valid audio streams, aligning them with the completed time; and performing format unification and encapsulating the participant identity information and participation timing information to form the multimodal input set for team activities.
[0020] In this embodiment, the generation of the process evidence set specifically includes: The multimodal input set of team activities is input into the real-time multimodal understanding model of the meeting process. Multimodal segmentation and index binding are performed to obtain a set of video frame sequences, a set of audio segments, and a set of multimodal alignment indexes. Visual and acoustic features are extracted from video frame sequence sets and audio segment sets respectively based on the multimodal alignment index set, generating visual feature sequence sets and acoustic feature sequence sets. Speaker embedding vector sets and speaker similarity matrices are generated based on the acoustic feature sequence sets. Clustering and temporal backfilling are performed to obtain speaker differentiation results consistent with the meeting time sequence information. The speaker embedding vector set is selected based on the acoustic feature sequence set to select a set of speech activity segments, and the set of speech activity segments is merged into a speaker discrimination segment set according to the participation time information. Through the embedding mapping network, the vector representation corresponding to each speaker discrimination segment is output to form an initial embedding set. Vector normalization is performed and index backfilling is performed according to the participation time information. At the same time, multiple vector representations belonging to the same participant identity information are time-series aggregated to obtain the speaker embedding vector set corresponding to the participant identity information. The speaker similarity matrix performs pairwise similarity calculations on any two speaker embedding vectors in the speaker embedding vector set. The similarity is determined by the ratio of the inner product of the two speaker embedding vectors to the norm product of the two speaker embedding vectors. All pairwise similarity results are matrix-filled according to the index order of the speaker embedding vector set to generate the speaker similarity matrix. The acoustic feature sequence set and the speaker discrimination result are input into the speech transcription link to generate speech transcription results. Based on the speaker discrimination result, the speech transcription results are segmented and merged to obtain the speaker-aligned transcription set. Input the speaker alignment transcription set, acoustic feature sequence set and visual feature sequence set into the sentiment tendency recognition link and behavior state recognition link to generate sentiment tendency results and behavior state recognition results consistent with the meeting time sequence information; The sentiment identification link consists of an input alignment unit, a semantic representation encoding unit, an acoustic sentiment encoding unit, a visual sentiment encoding unit, a cross-modal fusion unit, a temporal smoothing unit, and a sentiment output head. The input alignment unit uses the meeting's temporal information as an index to segment the speaker alignment transcription set, the acoustic feature sequence set, and the visual feature sequence set into the same time window and generate an alignment segment set. The semantic representation encoding unit performs word segmentation, denoising, and semantic encoding on the speaker alignment transcription set to obtain a semantic sentiment feature set. The acoustic sentiment encoding unit performs prosodic feature aggregation and sentiment discrimination representation extraction on the acoustic feature sequence set to obtain an acoustic sentiment feature set. The visual emotion encoding unit extracts the expression intensity and posture tension related representations from the face region sequence set and human key point sequence set in the visual feature sequence set to obtain a visual emotion feature set. The cross-modal fusion unit performs gated fusion on the semantic emotion feature set, acoustic emotion feature set and visual emotion feature set and outputs a fused emotion representation set. The temporal smoothing unit applies consistency constraints to the fused emotion representation set in adjacent time windows and eliminates isolated spikes to obtain a smooth emotion representation set. The emotion tendency output head classifies and maps the smooth emotion representation set and backfills it according to the meeting time sequence information to generate an emotion tendency result consistent with the meeting time sequence information. The behavior state recognition link consists of an input alignment unit, a visual motion representation encoding unit, a keypoint sequence encoding unit, a scene context encoding unit, a temporal behavior decoding unit, a state consistency verification unit, and a behavior state output header. The input alignment unit uses the meeting timing information as an index to segment the visual feature sequence set into fixed time windows and generate a behavior segment set. The visual motion representation encoding unit extracts gaze direction, head posture, and facial movement changes related to the face region sequence set in the behavior segment set to obtain a face motion feature set. The keypoint sequence encoding unit performs skeleton topology encoding and motion amplitude aggregation on the human keypoint sequence set to obtain a skeleton motion feature set. The scene context encoding unit... The unit extracts representations of the seating area, standing area, and interaction area from the meeting scene sequence set to obtain the scene context feature set. The temporal behavior decoding unit performs temporal fusion and sequence decoding on the face action feature set, skeleton action feature set, and scene context feature set to obtain the candidate behavior state sequence set. The state consistency verification unit applies duration constraints and state transition constraints to the candidate behavior state sequence set and eliminates instantaneous states that do not meet the constraints to obtain the stable behavior state sequence set. The behavior state output head backfills the stable behavior state sequence set according to the meeting time sequence information and aggregates it to the time segment corresponding to the speaker differentiation result to generate a behavior state recognition result consistent with the meeting time sequence information. The speaker differentiation results, speech transcription results, sentiment tendency results, and behavioral state recognition results are input into the structured meeting summary generation link to generate a structured meeting summary. The summary is then encapsulated according to the multimodal alignment index set to obtain the process evidence set.
[0021] In this embodiment, the generation of the personal capability source data set specifically includes: Establish cross-system identity mapping relationships based on meeting participant identity information, and uniformly map personnel identifiers in the training system, task system and inspection system to meeting participant identity information, and jointly collect corresponding training results, task quality, inspection records and meeting contributions. Extract assessment records and performance records from training results and ensure consistent application of standards to form a standardized collection of training results; The standardization of the criteria involves field standardization, time window alignment, and evaluation scale normalization. Extract task completion records, review records, and rework records for task quality and implement unified standards, then map them into a task quality score set according to a unified scoring rule. The unified scoring rule maps heterogeneous capability data into capability scores with unified dimensions based on calculable indicators such as completion rate, pass rate, proportion, and frequency. Extract the inspection point arrival records, hidden danger reporting records and review conclusions from the inspection records and implement unified standards, and map them into a set of inspection record scores according to unified scoring rules; Based on the speaker differentiation results, speech transcription results, behavioral state recognition results, and structured meeting summaries in the process evidence set, meeting contribution features are extracted and standardized to form a meeting contribution score set; The standardized training results set, task quality score set, inspection record score set, and meeting contribution score set are used to limit the effective time window based on the meeting time sequence information and perform missing data removal to generate a personal capability source data set.
[0022] In this embodiment, the generation of the portrait evidence feature set specifically includes: The process evidence set is indexed and expanded based on the multimodal aligned index set to form a set of evidence items bound by the participant identity information and the participation time sequence information; The evidence entry set includes the following fields: speaker differentiation results, speech transcription results, sentiment tendency results, behavioral state recognition results, and index pointers for structured meeting summaries. Perform integrity checks and null value removal on the evidence item set, perform field standardization and time window alignment on each field, and generate an aligned evidence item set; A quality assessment is performed on the aligned evidence item set. Based on the proportion of effective segments in the speech transcription results, the segment continuity of the speaker differentiation results, the temporal stability of the emotion tendency results, and the duration constraints of the behavioral state recognition results, an evidence quality label set is generated and written back to the aligned evidence item set to form a labeled evidence item set. Noise suppression is performed on the set of labeled evidence items. Items that do not meet the validity constraints in the evidence quality label set are removed. Temporal smoothing is performed on the sentiment tendency results and behavioral state recognition results that meet the validity constraints but have isolated spikes, and a cleaned set of evidence items is generated. Evidence merging is performed on the cleaned evidence item set. The speech transcription results and structured meeting summaries of the same participant identity information in the same participant time sequence information window are semantically deduplicated and merged. The sentiment tendency results and behavioral state recognition results in the same participant time sequence information window are aggregated and statistically analyzed to generate merged statistical fields, thus obtaining the merged evidence item set. Perform cross-entry consistency checks on the merged evidence entry set, generate a consistency check result set based on the segment alignment relationship between the speaker differentiation results and the speech transcription results, and the key point coverage relationship between the behavior state recognition results and the structured meeting summary, and write it back to the merged evidence entry set to form a consistent merged evidence entry set; Based on a consistent set of merged evidence items, the speech transcription results and structured meeting summaries are encoded as text evidence features, the sentiment tendency results and behavioral state recognition results are encoded as behavioral evidence features, and the text evidence features and behavioral evidence features are indexed and encapsulated according to the participant identity information and the participation time sequence information to generate a set of profile evidence features.
[0023] In this embodiment, the generation of the four-dimensional profile of personal capabilities specifically includes: A modular, tiered quantitative capability model is constructed. The set of profile evidence features and the set of personal capability source data are input into the cross-source alignment layer. Based on the participant identity information and the participant time sequence information, cross-source index alignment, scale alignment and semantic alignment are performed to generate a joint representation set. The modular tiered quantification capability model consists of a cross-source alignment layer, an evidence block library, a tiered capability ladder, a ladder gating device, a ladder recursion layer, a dimension assembly layer, an evidence conflict detector, a conflict suppression mechanism, and a profile output head. An evidence block library is constructed. The joint representation set is decomposed according to the evidence source and evidence form and written into the evidence block type set. Each evidence block encapsulates the participant identity information, participation time sequence information, source tag and quality tag to form a searchable evidence block index. The set of evidence block types includes transcription evidence blocks, summary evidence blocks, sentiment evidence blocks, behavioral state evidence blocks, and capability source evidence blocks. Transcription evidence blocks are formed from textual evidence features corresponding to speech transcription results, summary evidence blocks are formed from textual evidence features corresponding to structured meeting summaries, sentiment evidence blocks are formed from behavioral evidence features corresponding to sentiment tendency results, behavioral state evidence blocks are formed from behavioral evidence features corresponding to behavioral state recognition results, and capability source evidence blocks are formed from capability score items corresponding to personal capability source data sets. A set of evidence contributions is generated based on the evidence block library, which includes evidence block indexes, evidence contribution weights, and contribution type labels. The evidence contribution set includes evidence block indexes, evidence contribution weights, and contribution type labels. It is generated by recalling a set of candidate evidence blocks corresponding to the same participant identity information from the evidence block library using the participant identity information as the key. The candidate evidence block set is then windowed and aggregated according to the participant time sequence information to obtain a set of windowed evidence groups. Relevance evaluation and consistency evaluation are performed on each evidence block in the set of windowed evidence groups. The relevance evaluation is determined based on the matching score between the joint representation set and the evidence block representation. The consistency evaluation is determined based on the temporal alignment relationship and semantic coverage relationship between evidence blocks within the same participant time sequence information window. The relevance evaluation results, consistency evaluation results, and evidence block quality labels are jointly mapped to evidence contribution weights. Input the evidence contribution set into the hierarchical ability ladder, establish multi-level ability slots of the hierarchical ability ladder based on the dimension set corresponding to the four-dimensional profile of personal ability, and route the evidence contribution set to the corresponding ability slot according to the contribution type label to form the hierarchical input set. The tiered input set is input into the tiered gating system. Based on the temporal adjacency relationship between the evidence contribution weight and the meeting time sequence information, the gating selection result is generated. Gating suppression is performed on evidence blocks below the contribution weight threshold, and gating retention is performed on evidence blocks above the contribution weight threshold. The gating tiered input set is input into the tiered recursive layer to perform cross-window recursive fusion to generate a tiered capability set, which includes the hierarchical capability representation and hierarchical contribution index of each dimension. The tiered capability set is input into the dimensional assembly layer. Based on the participant identity information, the hierarchical capability representations of each dimension are aligned and assembled to generate a four-dimensional assembled representation set. A four-dimensional profile of personal capabilities is generated through the profile output head. The evidence contribution set, gating selection result and tiered capability set are input into the evidence conflict detector in parallel. Conflict detection is triggered based on the consistency relationship of evidence blocks within the same meeting time sequence information window under the same meeting identity information, and the conflict evidence index set and conflict type set are obtained. The triggering conditions include segment alignment conflicts between speaker differentiation results and speech transcription results, temporal co-occurrence conflicts between emotion tendency results and behavioral state recognition results, and key point coverage conflicts between structured meeting summaries and speech transcription results. Any one of these conditions must satisfy a preset conflict determination rule to trigger conflict detection. When the evidence conflict detector triggers conflict detection, it calls the conflict suppression mechanism to perform a set of suppression actions on the evidence blocks corresponding to the conflict evidence index set. The suppressed evidence contribution set is then fed back to the ladder gating and ladder recursion layer to update the gating step input set and ladder capability set. The set is then re-output as a four-dimensional profile of personal capabilities through the dimension assembly layer and the profile output head. The set of suppression actions includes determining the suppression order according to the conflict type, performing contribution weight reduction and gating forced suppression on conflicting evidence blocks, performing aggregate representation to replace conflicting evidence blocks with non-conflicting evidence blocks in the same window, writing conflict tags to conflicting evidence blocks and writing back to the evidence block library to form a conflict-annotated evidence block index.
[0024] In this embodiment, the generation of the closed-loop profile update result specifically includes: taking the profile evidence feature set generated by subsequent team activities and the individual ability source dataset as incremental input, and inputting it again into the building block tiered quantitative ability model. Under the premise of keeping the model structure and parameter configuration unchanged, cross-source alignment and time positioning are performed on the new evidence based on the participant identity information and participation time sequence information. The new evidence is written into the evidence block library and participates in the update calculation of the evidence contribution set. The ability representation of the new evidence and the existing evidence is fused in a time sequence through the tiered ability ladder, ladder gating and ladder recursion layer to generate the updated tiered ability set. The updated four-dimensional profile set of individual ability is output through the dimension assembly layer and the profile output head. At the same time, the consistency between the new evidence and the existing evidence is detected by the evidence conflict detector. When a conflict is triggered, the conflict suppression mechanism is called to suppress the conflicting evidence and update it back, forming a closed-loop profile update result that includes the ability profile result, the evidence contribution traceability relationship and the conflict annotation information.
[0025] Example 1: To verify the feasibility of this invention in practice, it was applied to a multi-shift collaborative operation management scenario in a large manufacturing enterprise. In this scenario, different shifts frequently hold production coordination meetings daily, including both on-site and remote meetings. These meetings involve multiple roles and complex discussions, and individual competency assessments have long relied on manual experience and static performance records, making it difficult to reflect changes in employee performance during actual work processes. This leads to problems such as delayed evaluations, fragmented evidence, and subjective bias. This invention addresses these issues by integrating meeting behavior evidence with long-term competency data to dynamically and continuously update individual competency profiles.
[0026] In practical applications, the system first simultaneously collects meeting video and audio streams during the meeting, and then aligns and unifies the formats by combining participant identity information and meeting timing information to form a directly processable multimodal input set. This input set is fed into a real-time multimodal understanding model for the meeting process, which automatically parses the meeting content and continuously generates speaker differentiation results, speech-to-text results, sentiment analysis results, behavioral state recognition results, and structured meeting summaries. Simultaneously, by combining training results, task completion status, inspection behavior, and meeting participation contributions recorded in the enterprise's existing management platform, data from different sources are standardized and missing data is removed to form a stable set of individual capability source data. Subsequently, the process evidence generated during the meeting is cleaned and merged, removing noisy evidence and merging valid information within the same time window to obtain a set of profile evidence features that can be used for capability modeling.
[0027] In the competency calculation phase, the feature set of the profile evidence and the set of personal competency source data are jointly input into a modular, tiered quantification competency model. Through cross-source alignment, evidence block construction, tiered recursion, and conflict suppression, a four-dimensional profile of personal competency is generated. As subsequent meetings continue, new meeting evidence is continuously incorporated into the model, incrementally updating the competency profile without altering the model structure. This allows the personal competency profile to evolve in real time with work behavior. In actual operation, comparative analysis of competency profile changes over different time periods reveals that this method significantly improves the continuity and consistency of competency assessment, reduces the frequency of human intervention, and effectively avoids evaluation fluctuations caused by single abnormal behaviors.
[0028] From the perspective of application results, after the company has been running the system for a period of time, managers are able to understand the changing trends of employees in terms of collaborative awareness, execution status and stability more intuitively based on the four-dimensional profile of individual capabilities. The relevant decision feedback cycle is significantly shortened, the team collaboration efficiency is continuously improved, and the objectivity and traceability of employee capability assessment are significantly enhanced, which fully demonstrates the practical value and promotional significance of the invention in real application scenarios.
[0029] Table 1. Performance comparison between deep learning-based user profiling methods and traditional methods.
[0030] As shown in Table 1, this invention has a significant advantage in profile update efficiency. Traditional manual assessment methods typically use a monthly update cycle, resulting in a profile update cycle of 30 days. In contrast, this invention shortens the profile update cycle to 1 day by using meeting process evidence and capability source data as continuous incremental inputs. This change is directly reflected in a significant reduction in the average time spent generating a single profile, from 18.6 minutes to 2.3 minutes. This indicates that the modular, tiered quantitative capability model significantly improves the automation level of evidence organization and capability calculation, reducing the time consumed by manual processing and repetitive calculations.
[0031] In terms of profiling quality, the improvement in profiling evidence completeness and cross-meeting evidence consistency is particularly significant. Traditional methods suffer from prominent issues of missing evidence and semantic breaks due to the dispersion of meeting minutes, behavioral performance, and historical capability data across different media, resulting in an evidence completeness rate of only 61.4% and a consistency rate of 58.9%. This invention, through evidence cleaning, merging, and cross-source alignment of process evidence, enables multimodal evidence to be encapsulated and utilized under a unified temporal sequence and identity index, increasing the evidence completeness rate to 93.8% and the cross-meeting evidence consistency rate to 91.2%, indicating a more comprehensive and coherent foundation for capability profiling.
[0032] From the perspective of stability and robustness indicators, this invention effectively suppresses abnormal fluctuations in capability profiles. Traditional manual assessments are significantly influenced by performance in a single meeting or subjective judgment, resulting in a capability profile stability volatility of 24.7%. This invention introduces evidence contribution weights and a temporal recursive mechanism into the tiered capability ladder and ladder gating system, and combines an evidence conflict detector and a conflict suppression mechanism to suppress low-quality or conflicting evidence, reducing the volatility to 7.6%. Simultaneously, the proportion of conflicting evidence decreases from 19.3% to 4.1%, indicating that the model possesses stronger self-correcting capabilities when facing multi-source heterogeneous evidence.
[0033] In terms of management efficiency, the frequency of manual intervention by managers has decreased from 12 times per week to 3 times, reflecting a significant improvement in the reliability and usability of the competency profile results. In particular, regarding the lead time indicator for identifying competency changes, this invention can identify individual competency change trends 14 days in advance, while traditional methods can only passively discover such changes in post-event statistics. This is mainly due to the closed-loop profile update mechanism, which allows newly added meeting evidence to continuously participate in the recursive calculation of competency, reflecting the impact of behavioral changes on the competency profile in advance.
[0034] In summary, the above data fully demonstrate that the present invention, by constructing a modular, tiered quantitative capability model, structurally integrates and updates meeting process evidence with long-term capability data, achieving substantial improvements in key indicators such as efficiency, completeness, stability, and traceability. The performance improvement stems from the systematic design of evidence block organization, tiered recursive fusion, and the synergistic effect of conflict detection and suppression.
[0035] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A user profiling method based on deep learning, characterized in that, Includes the following steps: Collect video streams, audio streams, participant identity information, and participation time sequence information of team activities, align execution time and unify format, and generate a multimodal input set for team activities; Input the multimodal input set of team activities into the multimodal real-time understanding model of the meeting process to generate a set of process evidence; Collect training results, task quality, inspection records and meeting contributions corresponding to the participant identity information, unify the execution criteria and remove missing data, and generate a set of personal capability source data. Perform evidence cleaning and merging on the process evidence set to generate a profile evidence feature set; Construct a modular, tiered quantitative capability model, and obtain a four-dimensional profile of personal capabilities based on the profile evidence feature set and the personal capability source data set. The feature set of profile evidence for subsequent team activities and the set of personal ability source data are input again into the modular tiered quantitative ability model to update the four-dimensional profile set of personal ability and output the closed-loop profile update result.
2. The user profiling method based on deep learning according to claim 1, characterized in that, The generation of the multimodal input set for the work group activities specifically includes: collecting the video and audio streams of the work group activities and writing them to the collection timestamp; collecting the participant identity information and participation time sequence information; performing integrity verification and quality rejection; generating a set of valid video streams and a set of valid audio streams; extracting the accompanying audio track based on the set of valid video streams and calculating the time offset with the set of valid audio streams and writing it back to align with the completion time; and performing format unification and encapsulating the participant identity information and participation time sequence information to form the multimodal input set for the work group activities.
3. The user profiling method based on deep learning according to claim 1, characterized in that, The generation of the process evidence set specifically includes: The multimodal input set of team activities is input into the real-time multimodal understanding model of the meeting process. Multimodal segmentation and index binding are performed to obtain a set of video frame sequences, a set of audio segments, and a set of multimodal alignment indexes. Visual and acoustic features are extracted from video frame sequence sets and audio segment sets respectively based on the multimodal alignment index set, generating visual feature sequence sets and acoustic feature sequence sets. Speaker embedding vector sets and speaker similarity matrices are generated based on the acoustic feature sequence sets. Clustering and temporal backfilling are performed to obtain speaker differentiation results consistent with the meeting time sequence information. The acoustic feature sequence set and the speaker discrimination result are input into the speech transcription link to generate speech transcription results. Based on the speaker discrimination result, the speech transcription results are segmented and merged to obtain the speaker-aligned transcription set. Input the speaker alignment transcription set, acoustic feature sequence set and visual feature sequence set into the sentiment tendency recognition link and behavior state recognition link to generate sentiment tendency results and behavior state recognition results consistent with the meeting time sequence information; The speaker differentiation results, speech transcription results, sentiment tendency results, and behavioral state recognition results are input into the structured meeting summary generation link to generate a structured meeting summary. The summary is then encapsulated according to the multimodal alignment index set to obtain the process evidence set.
4. The user profiling method based on deep learning according to claim 1, characterized in that, The generation of the personal capability source data set specifically includes: Establish cross-system identity mapping relationships based on meeting participant identity information, and uniformly map personnel identifiers in the training system, task system and inspection system to meeting participant identity information, and jointly collect corresponding training results, task quality, inspection records and meeting contributions. Extract assessment records and performance records from training results and ensure consistent application of standards to form a standardized collection of training results; The standardization of the criteria involves field standardization, time window alignment, and evaluation scale normalization. Extract task completion records, review records, and rework records for task quality and implement unified standards, then map them into a task quality score set according to a unified scoring rule. The unified scoring rule maps heterogeneous capability data into capability scores with unified dimensions based on calculable indicators such as completion rate, pass rate, proportion, and frequency. Extract the inspection point arrival records, hidden danger reporting records and review conclusions from the inspection records and implement unified standards, and map them into a set of inspection record scores according to unified scoring rules; Based on the speaker differentiation results, speech transcription results, behavioral state recognition results, and structured meeting summaries in the process evidence set, meeting contribution features are extracted and standardized to form a meeting contribution score set; The standardized training results set, task quality score set, inspection record score set, and meeting contribution score set are used to limit the effective time window based on the meeting time sequence information and perform missing data removal to generate a personal capability source data set.
5. The user profiling method based on deep learning according to claim 1, characterized in that, The generation of the portrait evidence feature set specifically includes: The process evidence set is indexed and expanded based on the multimodal aligned index set to form a set of evidence items bound by the participant identity information and the participation time sequence information; Perform integrity checks and null value removal on the evidence item set, perform field standardization and time window alignment on each field, and generate an aligned evidence item set; A quality assessment is performed on the aligned evidence item set. Based on the proportion of effective segments in the speech transcription results, the segment continuity of the speaker differentiation results, the temporal stability of the emotion tendency results, and the duration constraints of the behavioral state recognition results, an evidence quality label set is generated and written back to the aligned evidence item set to form a labeled evidence item set. Noise suppression is performed on the set of labeled evidence items. Items that do not meet the validity constraints in the evidence quality label set are removed. Temporal smoothing is performed on the sentiment tendency results and behavioral state recognition results that meet the validity constraints but have isolated spikes, and a cleaned set of evidence items is generated. Evidence merging is performed on the cleaned evidence item set. The speech transcription results and structured meeting summaries of the same participant identity information in the same participant time sequence information window are semantically deduplicated and merged. The sentiment tendency results and behavioral state recognition results in the same participant time sequence information window are aggregated and statistically analyzed to generate merged statistical fields, thus obtaining the merged evidence item set. Perform cross-entry consistency checks on the merged evidence entry set, generate a consistency check result set based on the segment alignment relationship between the speaker differentiation results and the speech transcription results, and the key point coverage relationship between the behavior state recognition results and the structured meeting summary, and write it back to the merged evidence entry set to form a consistent merged evidence entry set; Based on a consistent set of merged evidence items, the speech transcription results and structured meeting summaries are encoded as text evidence features, the sentiment tendency results and behavioral state recognition results are encoded as behavioral evidence features, and the text evidence features and behavioral evidence features are indexed and encapsulated according to the participant identity information and the participation time sequence information to generate a set of profile evidence features.
6. The user profiling method based on deep learning according to claim 1, characterized in that, The generation of the aforementioned four-dimensional profile of individual capabilities specifically includes: A modular, tiered quantitative capability model is constructed. The set of profile evidence features and the set of personal capability source data are input into the cross-source alignment layer. Based on the participant identity information and the participant time sequence information, cross-source index alignment, scale alignment and semantic alignment are performed to generate a joint representation set. The modular tiered quantification capability model consists of a cross-source alignment layer, an evidence block library, a tiered capability ladder, a ladder gating device, a ladder recursion layer, a dimension assembly layer, an evidence conflict detector, a conflict suppression mechanism, and a profile output head. An evidence block library is constructed. The joint representation set is decomposed according to the evidence source and evidence form and written into the evidence block type set. Each evidence block encapsulates the participant identity information, participation time sequence information, source tag and quality tag to form a searchable evidence block index. A set of evidence contributions is generated based on the evidence block library, which includes evidence block indexes, evidence contribution weights, and contribution type labels. Input the evidence contribution set into the hierarchical ability ladder, establish multi-level ability slots of the hierarchical ability ladder based on the dimension set corresponding to the four-dimensional profile of personal ability, and route the evidence contribution set to the corresponding ability slot according to the contribution type label to form the hierarchical input set. The tiered input set is input into the tiered gating system. Based on the temporal adjacency relationship between the evidence contribution weight and the meeting time sequence information, the gating selection result is generated. Gating suppression is performed on evidence blocks below the contribution weight threshold, and gating retention is performed on evidence blocks above the contribution weight threshold. The gating tiered input set is input into the tiered recursive layer to perform cross-window recursive fusion to generate a tiered capability set, which includes the hierarchical capability representation and hierarchical contribution index of each dimension. The tiered capability set is input into the dimensional assembly layer. Based on the participant identity information, the hierarchical capability representations of each dimension are aligned and assembled to generate a four-dimensional assembled representation set. A four-dimensional profile of personal capabilities is generated through the profile output head. The evidence contribution set, gating selection result and tiered capability set are input into the evidence conflict detector in parallel. Conflict detection is triggered based on the consistency relationship of evidence blocks within the same meeting time sequence information window under the same meeting identity information, and the conflict evidence index set and conflict type set are obtained. When the evidence conflict detector triggers conflict detection, it calls the conflict suppression mechanism to perform a set of suppression actions on the evidence blocks corresponding to the conflict evidence index set. The suppressed evidence contribution set is then fed back to the ladder gating and ladder recursion layer to update the gating step input set and ladder capability set. After passing through the dimension assembly layer and the portrait output head, the personal capability four-dimensional portrait set is re-output.
7. The user profiling method based on deep learning according to claim 1, characterized in that, The generation of the closed-loop profile update result specifically includes: taking the profile evidence feature set generated from subsequent team activities and the individual ability source dataset as incremental input, and inputting it again into the modular tiered quantitative ability model. While keeping the model structure and parameter configuration unchanged, cross-source alignment and time positioning are performed on the new evidence based on the participant identity information and participation time sequence information. The new evidence is written into the evidence block library and participates in the update calculation of the evidence contribution set. The ability representation of the new evidence and the existing evidence is fused in a time sequence through the tiered ability ladder, ladder gating, and ladder recursion layer to generate the updated tiered ability set. The updated four-dimensional profile set of individual abilities is output through the dimension assembly layer and the profile output head. At the same time, the consistency between the new evidence and the existing evidence is detected by the evidence conflict detector. When a conflict is triggered, the conflict suppression mechanism is called to suppress the conflicting evidence and update it back, forming a closed-loop profile update result that includes the ability profile result, the evidence contribution tracing relationship, and the conflict annotation information.