Information matching method and device based on multiple modes, computer equipment and medium
Through multimodal processing technology and timestamp matching algorithm, the problems of low accuracy and poor timeliness of traditional character portraits are solved, and efficient and accurate portrait construction is achieved, which is suitable for employee performance evaluation and customer identification.
Patent Information
- Application Number
- CN202510364807.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-08-15
AI Technical Summary
Traditional methods have low accuracy, strong subjectivity and poor timeliness when building portraits of characters, making it difficult to meet complex business needs.
Object detection technology, voiceprint recognition technology and natural language processing technology are used to perform multi-modal processing on audio and video files, multi-dimensional feature sets are extracted, and feature comparison algorithms are used to match character portraits based on pre-recorded timestamps.
It improves the efficiency and accuracy of feature extraction, ensures the accuracy and timeliness of character portraits, reduces labor and material costs, and is suitable for scenarios such as employee performance evaluation and customer identity identification.
Smart Images

Figure CN120492941A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a multimodal information matching method, device, computer equipment and medium. Background Art
[0002] Traditional user profiling methods mainly include rule-based methods, statistical analysis methods, and questionnaire survey methods. Rule-based methods use predefined business rules to classify users into different groups based on their attributes (such as age, gender, region, etc.) and behavioral data (such as purchase history, browsing history, etc.). However, rules need to be set manually, making it difficult to adapt to complex and changing user behavior patterns. Updating rules is costly and can easily miss important features. Statistical analysis methods use statistical methods (such as cluster analysis and regression analysis) to model user data and identify user groups with similar characteristics. However, they rely on large amounts of high-quality data and are ineffective when the data is insufficient or of low quality. The model results are also difficult to interpret and lack interpretability. Questionnaire survey methods collect basic user information and preferences through designed questionnaires to construct user profiles. However, the sample is not representative enough and it is difficult to cover all user groups. The information source is single and is inevitably affected by subjective factors of the question setter and the respondent, resulting in the possibility of subjective bias when users fill out the questionnaire. Traditional methods of drawing user portraits involve either data analysts analyzing common information in large amounts of data, which is inevitably influenced by prior information and makes it difficult to capture the potential information contained therein; or using machine learning methods for mathematical modeling, which requires certain server costs.
[0003] Traditional work level assessment methods mainly include performance appraisal, 360-degree assessment, and critical incident assessment. The performance appraisal method sets specific performance indicators based on employees' work results and task completion, and conducts assessments regularly. However, the indicator settings may be too simplistic. Traditional level assessment methods often set a few indicators and use the only means of assessment based on whether the person being assessed can achieve the indicators. This cannot fully reflect the employee's actual work performance and can easily lead to short-term behavior and neglect long-term development. The 360-degree assessment method comprehensively evaluates employees from multiple perspectives (superiors, colleagues, subordinates, customers, etc.), but the evaluator may have subjective biases, which affect the objectivity of the assessment results. The feedback information is too much and difficult to process. The critical incident method records particularly good or bad events that occurred in the employee's work as the basis for evaluation. However, incomplete event records may lead to inaccurate evaluation results, making it difficult to quantify and not conducive to horizontal comparison.
[0004] The existing methods of drawing user portraits and evaluating work levels will inevitably lead to a lag in information extraction due to the time cost of collecting data. This will cause inevitable timeliness errors in the extracted information during the utilization stage. At the same time, the labor cost of data analysis and portrait construction cannot be ignored.
[0005] To sum up, traditional methods have shortcomings in practice, such as low accuracy, strong subjectivity, and poor timeliness. They have many limitations and are difficult to meet increasingly complex business needs. Summary of the Invention
[0006] In view of this, the present invention provides a multimodal information matching method, apparatus, computer equipment and medium to solve the problems of low accuracy, strong subjectivity and poor timeliness in the actual construction of character portraits by traditional methods.
[0007] In a first aspect, the present invention provides a multimodal information matching method, the method comprising:
[0008] Obtain the audio and video files to be extracted for different task objects in the target scene;
[0009] Use object detection technology, voiceprint recognition technology, and natural language processing technology to perform multimodal processing on audio and video files, and extract multi-dimensional feature sets corresponding to different task objects from the multimodal processed audio and video files;
[0010] Construct character portraits based on multi-dimensional feature sets corresponding to different task objects;
[0011] A feature matching algorithm is used to match the character profile to a specific person based on pre-recorded timestamps.
[0012] The present invention provides a multimodal information matching method for acquiring audio and video files to be extracted from different task objects within a target scene, capable of collecting information from multiple dimensions. Compared to single-source data, this multi-source data collection method can more comprehensively reflect the behavior and status of the task objects in the scene, providing a rich data foundation for subsequent analysis and avoiding inaccurate analysis results due to missing information. Multimodal processing of audio and video files using object detection, voiceprint recognition, and natural language processing technologies, and extracting multidimensional feature sets, offers significant advantages. Object detection technology can quickly and accurately identify people, objects, and their actions in the video. Voiceprint recognition technology can precisely distinguish the voices of different people, ensuring accurate attribution of speech content and avoiding information errors caused by voice confusion. Natural language processing technology analyzes the text after speech-to-text conversion to extract deep-level features such as semantics and emotion. The integrated use of multimodal technologies greatly improves the efficiency and accuracy of feature extraction, enabling the rapid and accurate extraction of valuable information from complex audio and video data. Constructing character portraits based on the multidimensional feature sets corresponding to different task objects can present a richer and more three-dimensional character image. The multi-dimensional feature set covers a wide range of information, including a person's appearance, behavioral habits, language style, and emotional tendencies. A feature matching algorithm is used based on pre-recorded timestamps to match the person's portrait to a specific person, with high accuracy and efficiency. Timestamps provide a temporal dimension identifier for the data, making data from different modalities consistent in time and facilitating synchronous analysis and matching. The feature matching algorithm can quickly calculate the similarity between the features of the person's portrait and the features of the specific person. With the assistance of timestamps, it can further narrow the matching range and improve the accuracy of the match. This matching method can quickly and accurately match newly collected person portraits with specific persons in the database. Whether in employee performance evaluation, customer identification, or other scenarios requiring person matching, it can save companies a significant amount of time and labor costs, while improving the accuracy of decision-making. It solves the problems of low accuracy, strong subjectivity, and poor timeliness in the actual construction of person portraits with traditional methods.
[0013] In an optional embodiment, the audio and video files include video files and audio files; obtaining the audio and video files to be extracted for different task objects in the target scene includes:
[0014] Obtain a video file and an audio file with a timestamp for a preset time period within the target scene;
[0015] The synchronous alignment technology is used to align the timestamps of the video files and the audio files with timestamps to obtain video files and audio files with consistent timestamps.
[0016] The present invention provides a multimodal information matching method, which can capture visual information such as a person's body movements, expressions, and position movements through video files; and can record acoustic information such as voice content, intonation, and speaking speed through audio files. By aligning the timestamps of video files and audio files with timestamps through synchronous alignment technology, accurate matching of video and audio data in the time dimension can be achieved. This means that the visual information such as the person's movements and expressions can accurately correspond to the audio information such as the words and intonation they say. Video and audio files with consistent timestamps can fully present the information in the target scene, ensuring that the information collected from multiple dimensions is complete, and providing comprehensive data support for in-depth exploration of potential laws and behavioral patterns in the scene. For analysis tasks based on audio and video data, timestamp alignment significantly improves the accuracy of the analysis. When analyzing a person's emotional state, combined with synchronized facial expression and voice intonation analysis, their true emotions can be more accurately judged.
[0017] In an optional embodiment, the natural language processing technology includes speech recognition technology and a large language model;
[0018] Use object detection technology, voiceprint recognition technology, and natural language processing technology to perform multimodal processing on audio and video files, including:
[0019] Use object detection technology to process the video modality of the video file and determine the target time period for separate conversations between different task objects in the video file;
[0020] Use voiceprint recognition technology to process the audio modality of the audio file, obtain the audio of each task object in the audio file within the target time period, and annotate the audio of each task object;
[0021] Speech recognition technology is used to perform text modal processing on the audio of each labeled task object to obtain the text content corresponding to each labeled task object audio. A large language model is used to perform contextual analysis on the text content corresponding to each labeled task object audio to verify and correct the text content that is incorrectly labeled due to the similar timbre of different task objects.
[0022] The present invention provides a multimodal information matching method that uses object detection technology to process video files in the video mode and determine the target time period for individual conversations between different task subjects. This operation can accurately locate valuable information fragments. Voiceprint recognition technology is used to process audio files in the audio mode, obtaining and annotating the audio of each task subject within the target time period, effectively solving the problem of audio confusion between multiple people. Voiceprint characteristics of different people are unique, just like fingerprints. Voiceprint recognition technology can accurately distinguish the voices of different task subjects. Speech recognition technology is used to convert the annotated task subject audio into text content. A large language model is used for contextual analysis to verify and correct mislabeled text content due to similar timbre, significantly improving the accuracy and usability of the text. By understanding the text context and performing semantic analysis, the large language model can determine the rationality of text attribution and promptly detect and correct errors. It can accurately verify and correct text, providing high-quality data for subsequent text-based analysis, and enhancing the accuracy and robustness of the audio-to-text conversion process.
[0023] In an optional embodiment, different task objects include persons with fixed identities and persons with random identities;
[0024] Extract multi-dimensional feature sets corresponding to different task objects from the multimodal processed audio and video files, including:
[0025] A large language model is used to extract the first appearance feature, first body movement feature, and first emotional expression feature of a person with a fixed identity from the video file of the target time period, and to extract the second appearance feature, second body movement feature, and second emotional expression feature of a person with a random identity;
[0026] Based on preset prompt words, a large language model is used to extract the first language expression features and chat skill features of fixed-identity people from the text content corresponding to the audio of each labeled task object, as well as the second language expression features, feedback features, and satisfaction scores of random-identity people matching the preset prompt words.
[0027] The first appearance feature, the first body movement feature, the first emotion expression feature, the first language expression feature, the chat skill feature, the feedback feature and the satisfaction score are used as the first multidimensional feature set;
[0028] The second appearance feature, the second body movement feature, the second emotion expression feature, the second language expression feature, the feedback feature and the satisfaction score are used as the second multidimensional feature set.
[0029] The present invention provides a multimodal information matching method that utilizes a large language model to extract multidimensional features from the textual content corresponding to video files and audio, enabling a comprehensive and detailed characterization of the task subject. Appearance, body language, and emotional expression features extracted from the video provide a visual representation of the task subject's outward appearance and immediate status. Language expression and conversational skills features extracted from the audio text provide a deeper understanding of the task subject's communication skills, such as a salesperson's language fluency, use of professional terminology, and skill in facilitating customer conversations. For customers with random identities, multifaceted features are similarly extracted, providing rich data for accurate analysis of different types of individuals and avoiding the limitations of single-feature analysis. Different features are extracted for both fixed-identity and random-identity individuals, providing highly targeted analysis. Fixed-identity individuals typically play specific roles in business processes, and extracting their first language expression and conversational skills features helps assess their work ability and professionalism. For random-identity individuals, second language expression features, feedback features, and satisfaction scores are extracted to focus on their reactions to products or services. By integrating these multidimensional feature sets, a highly accurate persona profile can be constructed. For fixed-identity employees, the first multidimensional feature set comprehensively reflects all aspects of their work performance, including appearance, body language, emotional management, and communication skills, assisting companies in talent assessment and training needs analysis. For random-identity employees, the second multidimensional feature set focuses on characteristics during their interactions with the business, helping companies gain a deeper understanding of customer needs, preferences, and satisfaction, providing strong support for precision marketing and product optimization. The extracted multidimensional feature set provides a rich data basis for corporate decision-making.
[0030] In an optional embodiment, constructing a character portrait based on a multi-dimensional feature set corresponding to different task objects includes:
[0031] Obtain the appearance information of each random person and each fixed person from the video files within the target time period;
[0032] Constructing a character portrait of a random person based on a preset timestamp and a second multi-dimensional feature set;
[0033] A character portrait of a person with a fixed identity is constructed based on a preset timestamp and a first multi-dimensional feature set.
[0034] The present invention provides a multimodal information matching method that obtains the appearance information of each random-identity person and fixed-identity person from video files within a target time period, accurately capturing key information. Constructing a person portrait based on a preset timestamp and a multidimensional feature set enhances the timeliness and relevance of the portrait. The timestamp provides a clear time identifier for the data, allowing the appearance information to be closely temporally correlated with other multidimensional features (such as body movements and language expressions). Constructing person portraits for fixed-identity persons and random-identity persons based on the preset timestamp and first and second multidimensional feature sets, respectively, significantly improves the completeness and accuracy of the portraits. For fixed-identity persons, the first multidimensional feature set encompasses multiple aspects such as appearance, body movements, emotional expression, language expression, and conversation skills. Combined with the timestamp, the constructed portrait comprehensively and accurately reflects their status and capabilities in work scenarios. For random-identity persons, the second multidimensional feature set, combined with appearance information, can deeply characterize their characteristics and needs during business interactions. The resulting accurate and complete person portraits provide strong support for enterprise business decision-making and optimization.
[0035] In an optional embodiment, a feature matching algorithm is used to match the character portrait to a predetermined person based on a pre-recorded timestamp, including:
[0036] Tracing the source is performed based on a preset timestamp, and a feature comparison algorithm is used to compare the portrait of the random person with the appearance information of each random person to obtain the matching random person with the given identity;
[0037] Obtain the prior information of the fixed-identity person, trace the source based on the preset timestamp, and use the feature comparison algorithm to compare the character portrait of the fixed-identity person with the prior information of each fixed-identity person to obtain the matching fixed-identity person.
[0038] The present invention provides a multimodal information matching method that performs traceability based on preset timestamps, and can accurately track the activity trajectories of people with random identities and people with fixed identities at different time points. The timestamp is like a time coordinate, which connects various behavioral data of people within the target time period. For people with random identities, at the same time, a feature comparison algorithm is used to compare the character portrait with the appearance information, which can accurately identify specific customers. For people with fixed identities, such as corporate employees, the timestamp can trace each link in their work process. Combined with the comparison of the portrait and prior information, it can confirm whether the employee's work status at a specific time meets the requirements, accurately identify the employee's identity and work performance, and help the company to carry out refined management.
[0039] In an optional implementation, the multimodal information matching method further includes:
[0040] The work level of the person with a fixed identity is evaluated based on the first multi-dimensional feature set to obtain a work evaluation result.
[0041] The present invention provides a multimodal information matching method, in which the first multidimensional feature set includes the first appearance feature, first body movement feature, first emotional expression feature, first language expression feature and chat skill feature of a person with a fixed identity. This makes the evaluation of their work level no longer limited to a single dimension, but a comprehensive consideration from multiple angles. Evaluation based on a multidimensional feature set can produce more accurate and objective work evaluation results. Each feature reflects the work status of employees from different aspects. These features are quantitatively analyzed through scientific evaluation methods, reducing the interference of subjective factors in the evaluation. Accurate work evaluation results are based on the first multidimensional feature set, which helps companies improve overall performance and strengthen cultural construction. After employees clarify their own development direction and improve their abilities through training, their work performance and results will be improved, directly promoting the improvement of corporate performance.
[0042] In a second aspect, the present invention provides a multimodal information matching device, comprising:
[0043] The module for obtaining audio and video files to be extracted is used to obtain audio and video files to be extracted for different task objects in the target scene;
[0044] A multimodal processing and multidimensional feature set extraction module, which uses object detection technology, voiceprint recognition technology, and natural language processing technology to perform multimodal processing on audio and video files, and extract multidimensional feature sets corresponding to different task objects from the multimodal processed audio and video files;
[0045] A character portrait construction module is used to construct character portraits based on multi-dimensional feature sets corresponding to different task objects;
[0046] The information matching module is used to match the character portrait to a predetermined person using a feature comparison algorithm based on a pre-recorded timestamp.
[0047] In a third aspect, the present invention provides a computer device comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the multimodal information matching method of the first aspect or any corresponding embodiment thereof by executing the computer instructions.
[0048] In a fourth aspect, the present invention provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the multimodal information matching method of the first aspect or any corresponding embodiment thereof.
[0049] In a fifth aspect, the present invention provides a computer program product comprising computer instructions for enabling a computer to execute the multimodal information matching method of the first aspect or any corresponding embodiment thereof. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the specific embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0051] Figure 1 is a flowchart of a multimodal information matching method according to an embodiment of the present invention;
[0052] Figure 2 is a flowchart of another multimodal information matching method according to an embodiment of the present invention;
[0053] Figure 3 is a flowchart of another multimodal information matching method according to an embodiment of the present invention;
[0054] FIG4( a ) is a schematic flow chart of yet another multimodal information matching method according to an embodiment of the present invention;
[0055] FIG4( b ) is a schematic flow chart of another multimodal information matching method according to an embodiment of the present invention;
[0056] Figure 5 is a structural block diagram of a multimodal information matching device according to an embodiment of the present invention;
[0057] Figure 6 Schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0058] To make the purpose, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of the present invention.
[0059] Existing methods for creating user profiles rely on the collection and analysis of large amounts of user data, and their technical background primarily relies on data mining, machine learning algorithms, and statistical analysis. These methods aim to extract features from multiple sources of information, such as user behavior data and demographic data, to construct models that reflect user interests, preferences, and behavioral patterns. For example, e-commerce websites record user browsing history, purchase history, and other information, and use cluster analysis to group users with similar shopping habits into categories, thereby customizing personalized recommendations for each category. Existing implementations for creating user profiles include:
[0060] First, data mining is used to extract valuable information from multiple data sources (such as user behavior logs, demographic data, and transaction records). After cleaning and preprocessing, this data is transformed into structured or semi-structured datasets, laying the foundation for subsequent analysis. Next, machine learning algorithms are used to classify and cluster users. For example, the K-means clustering algorithm can be used to divide users into different groups based on characteristics such as purchase frequency and preferred product categories. Alternatively, supervised learning algorithms such as decision trees and random forests can be used to predict users' potential needs and behavior patterns. Finally, statistical analysis methods are used to gain a deeper understanding of the characteristics of user groups. Descriptive statistics are used to analyze the distribution of basic user attributes, such as age, gender, and location. Correlation and regression analysis are also used to explore the relationships between different variables and reveal the key factors influencing user behavior.
[0061] Traditional performance evaluation methods are based more on performance management theory and human resource management practices. Their technical background encompasses psychological measurement, KPI (Key Performance Indicator) setting, and assessment system design. These evaluations typically focus on employee performance, efficiency, and teamwork.
[0062] Current implementations of performance evaluations include: First, using a performance management framework to collect data from multiple sources (such as work results, colleague feedback, and customer reviews). After organization and preprocessing, this data forms a structured or semi-structured dataset, laying the foundation for subsequent evaluations. Common data sources include KPIs (key performance indicators), project completion status, and 360-degree feedback surveys. Next, employ psychometric tools and data analysis methods to quantify and categorize employee performance. For example, a scale can be used to assess employees' soft skills (such as communication and teamwork), and regression analysis can be used to determine which factors significantly influence work outcomes. Alternatively, machine learning algorithms (such as random forests and support vector machines) can be used to predict employees' future performance and development potential. Finally, statistical analysis methods are used to gain a deeper understanding of the evaluation results. Descriptive statistics are used to analyze the distribution of employee performance across different dimensions, such as work efficiency, innovation, and customer satisfaction. Correlation and regression analysis can also be used to explore the relationships between different performance indicators and uncover the key factors influencing performance.
[0063] Taking the sales department as an example, companies will comprehensively assess employee performance based on quantitative and qualitative indicators such as sales performance and customer satisfaction survey results, combined with a 360-degree feedback system. This evaluation then provides targeted training and development recommendations. Many companies currently use the OKR (Objectives and Key Results) methodology to track employee progress and measure contributions. This methodology uses clear objectives and quantifiable key results to ensure focus on the most important tasks and measure progress. Objectives are inspiring, qualitative descriptions, while key results are specific, measurable indicators. Each objective typically has three to five key results. OKRs are typically implemented on a quarterly or annual basis, offering high transparency, flexibility, and self-motivation. For example, an internet company might set a goal of "increasing user engagement" and quantify and track progress against this goal through key results such as "increasing daily active users by 20%" and "launching at least two new features."
[0064] In summary, the existing methods of drawing user portraits and evaluating work levels will inevitably lead to a lag in extracting information due to the time cost of collecting data, which will inevitably cause the extracted information to have inevitable timeliness errors in the utilization stage. At the same time, the labor cost of data analysis and portrait construction cannot be ignored. The present invention has a high degree of automation, which saves a lot of manpower and material resources while dynamically analyzing portrait information to ensure its timeliness and accuracy. It can automatically extract information from a large amount of unstructured conversation text, covering more interactive scenarios and details, and avoiding errors caused by subjective factors in the data acquisition and analysis process of traditional methods.
[0065] The embodiment of the present invention provides a multimodal information matching method, which constructs character portraits in different scenarios through steps such as data processing, information extraction, information output and matching, thereby achieving the effect of accurately constructing character portraits and solving the problems of low accuracy, strong subjectivity and poor timeliness in the actual construction of character portraits by traditional methods.
[0066] According to an embodiment of the present invention, an embodiment of a multimodal information matching method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0067] In this embodiment, a multimodal information matching method is provided, which can be used in the above-mentioned computer device. Figure 1 is a flow chart of a multimodal information matching method according to an embodiment of the present invention. Figure 1 As shown, the process includes the following steps:
[0068] Step S101: obtaining audio and video files to be extracted of different task objects in a target scene.
[0069] Specifically, for various target scenarios, statistically significant task objects can be roughly divided into two categories: the first category of task objects are people with fixed identities, who have been in the scenario for a long time and often have relatively complete identity prior information, such as sales personnel in the sales field, teachers in the education field, etc. For this type of task objects, the focus of information extraction is mostly on the performance of fixed people facing different objects or events in fixed scenarios. Therefore, through monitoring, recording and other equipment in the target scenario, video files and audio files of such people during their working hours can be intercepted as effective information sources, and the data can be concentrated for comprehensive processing. The extracted information is often summarized from many video files and audio files containing original information, and the final information feature matching is done. Assigned to a specific person; the second type of task objects are people with random identities, short-term stays in the scene, and there is no way to collect all prior information. They are highly mobile and random, but have common characteristics, such as customers in the sales field. Such people are not statistically significant as individuals because their identities are random and unknown; since the means of information collection rely on scenes (monitoring and recording fixed in the target scene, etc.), compared with the first type of task objects, this type of task object will have fewer information sources and limited information that can be extracted due to its short-term stay in the scene; however, as a group, this type of character objects will show some common characteristics. This common characteristic is effective information with statistical significance and is also the main target of information collection.
[0070] Taking the sales field as an example, the focus of information extraction for the second type of task objects is mostly on the different performances of specific people facing the same salesperson or product. The purpose is to subdivide the broad concept of "customer" to promote adapting to local conditions and teaching students in accordance with their aptitude. Therefore, the amount of data is not blindly pursued to be huge, but more accurate. The extracted information is refined to achieve the goal of seeing the big picture from the details, and finally matched to the customer. The features are often abstracted into one together with the customer's basic information as the extraction result.
[0071] It's important to note that the audio and video files to be extracted for different task objects within the target scenario must be consistent with the scenario and time-matched. This means they contain the data to be extracted, ensuring the necessary information is present. For example, in the sales field, the input can be video and audio of a store during business hours, while footage recorded when the store is closed would not meet the extraction requirements.
[0072] In step S102 , object detection technology, voiceprint recognition technology, and natural language processing technology are used to perform multimodal processing on the audio and video files, and multi-dimensional feature sets corresponding to different task objects are extracted from the multimodal processed audio and video files.
[0073] Specifically, multimodality includes visual, auditory, textual, and other modalities. Visual modality primarily derives from images and video files, encompassing information such as an object's appearance, shape, color, position, and motion. For example, processing video files using object detection technology to visually identify different task objects, determine their motion and position, and thus define the target time period for the conversation is an application of the visual modality. Auditory modality uses audio as its data source, encompassing information such as speech and ambient sounds. For example, processing audio files using voiceprint recognition technology to distinguish and label the voices of different task objects and converting audio into text all involve the auditory modality. Textual modality consists of data consisting of text, including conversational text, documents, and webpages. Using large language models to perform contextual analysis on speech-to-text text, verify and correct the textual content, and mine its semantic information is considered textual modality processing. In natural language processing tasks, news article classification, sentiment analysis, and intent recognition in chat logs are all applications of textual modality data. In addition to the common modalities mentioned above, other modalities may also include sensor data (such as temperature and humidity sensor data, used for environmental monitoring and analysis), biometric data (such as fingerprint and iris recognition data, used for identity authentication), etc.
[0074] The multidimensional feature set may include features such as body movement features, language expression features, appearance features, feedback features, and satisfaction scores extracted from visual modalities, auditory modalities, and textual modalities.
[0075] Step S103: constructing a character portrait based on the multi-dimensional feature sets corresponding to different task objects.
[0076] For example, based on the multi-dimensional feature set of people with fixed identities and the multi-dimensional feature set of people with random identities, a machine learning algorithm is used to construct character portraits of people with fixed identities and task portraits of people with random identities, respectively.
[0077] Furthermore, machine learning algorithms, such as clustering algorithms, are used to analyze the collated and integrated feature data, dividing customers with similar characteristics into different groups and generating profiles for each group. Each group profile includes a description of its key characteristics, such as "young female customers are highly interested in fashion products, pay attention to product appearance, frequently mention the novelty of product styles in their feedback, and generally have high satisfaction ratings."
[0078] Step S104 , matching the character portrait to a predetermined person using a feature comparison algorithm based on a pre-recorded timestamp.
[0079] Specifically, for individuals with fixed identities, the constructed persona data was further organized to ensure that the data for each characteristic dimension was clearly identifiable. For individuals with random identities, the group portrait data was similarly organized. Key features within each group portrait, such as the age range, gender, interest in fashion products, attention to product appearance, frequency of feedback on style novelty, and average satisfaction ratings for the "young female customer group" profile, were organized into structured data, clarifying the value range and description method for each feature.
[0080] Extract relevant data about established personnel from the company's personnel database. For individuals with fixed identities, extract basic information from their files, such as name, employee number, and job title, and also extract historical work data related to the person's profile characteristics. For example, for sales personnel, extract their verbal expression records from past sales activities (sales scripts that can be obtained through audio-to-text conversion), body movement observation records (such as behavioral observation reports during training), appearance requirements and actual performance records, etc., and organize this data in a manner corresponding to the person's profile characteristics to form a feature data set for the established personnel.
[0081] For random individuals, if a historical customer information database exists, data relevant to the current person's profile is extracted. For example, information such as the customer's age, gender, past purchase history, and product reviews is extracted and organized into structured data that can be compared with the person's profile.
[0082] The character portrait data and the established personnel data are separately associated and organized with pre-recorded timestamps. For individuals with fixed identities, the time interval corresponding to the character portrait data is determined. For example, the portrait data of a salesperson is based on audio and video analysis of their sales activities during a certain quarter, and the start and end times of the quarter are recorded as the timestamp range. In the historical work data of established personnel, the time corresponding to each data item is also marked, such as the specific time of a sales speech recording, the time of a body movement observation, etc., to ensure that the data of both parties can correspond in the time dimension.
[0083] For random individuals, clearly define the time range represented by the group profile data. For example, the profile data for a "young female customer group" is based on customer behavior analysis during a recent promotion, recording the time period of the promotion. Within the historical customer information database, select customer data that transacted or interacted within that time period and annotate the specific time of each transaction or interaction to facilitate subsequent accurate timestamp-based comparison.
[0084] For individuals with fixed identities, given that their feature data includes both numerical (such as body movement frequency and language fluency indicators) and categorical (such as clothing style categories), a comparison algorithm that comprehensively considers different data types can be used, such as an algorithm that combines Euclidean distance and cosine similarity. Euclidean distance can be used to calculate the differences between numerical features, while cosine similarity is used to measure the similarity between categorical features in vector space. For individuals with random identities, since the feature data is mostly descriptive and statistical, algorithms based on fuzzy matching and correlation analysis can be used. For example, for numerical data such as age range and satisfaction ratings, the degree of similarity is measured by calculating relative error; for textual descriptive feedback features, text similarity algorithms (such as the edit distance algorithm) are used to determine similarity.
[0085] The multimodal information matching method provided in this embodiment acquires audio and video files to be extracted for different task objects within the target scene, gathering information from multiple dimensions. Compared to a single data source, this multi-source data collection method can more comprehensively reflect the behavior and status of the task objects in the scene, providing a rich data foundation for subsequent analysis and avoiding inaccurate analysis results due to missing information. The multimodal processing of audio and video files using object detection, voiceprint recognition, and natural language processing technologies, and the extraction of multidimensional feature sets, offers significant advantages. The combined application of multimodal technologies greatly improves the efficiency and accuracy of feature extraction, enabling the rapid and accurate extraction of valuable information from complex audio and video data. Constructing character portraits based on multidimensional feature sets corresponding to different task objects presents a richer and more three-dimensional character image. The multidimensional feature sets encompass a wide range of information, including a person's appearance, behavioral habits, language style, and emotional tendencies. Using a feature comparison algorithm based on pre-recorded timestamps, the character portraits are matched to specific individuals with high accuracy and efficiency. Timestamps provide a temporal dimension for data, making data from different modalities consistent in time and facilitating synchronous analysis and matching. Feature matching algorithms can quickly calculate the similarity between character portrait features and established person features. With the aid of timestamps, they can further narrow the matching range and improve matching accuracy. This matching method can quickly and accurately match newly collected character portraits with established persons in the database. Whether in employee performance evaluations, customer identification, or other scenarios requiring character matching, it can save companies significant time and labor costs while improving decision-making accuracy. It also addresses the low accuracy, subjectivity, and timeliness of traditional character portrait construction methods.
[0086] In this embodiment, a multimodal information matching method is provided, which can be used in the above-mentioned computer device. Figure 2 is a flow chart of a multimodal information matching method according to an embodiment of the present invention. Figure 2 As shown, the process includes the following steps:
[0087] Step S201 : obtaining audio and video files to be extracted of different task objects in a target scene.
[0088] Specifically, the audio and video files include video files and audio files; the above step S201 includes:
[0089] Step S2011: Acquire a video file and an audio file with a timestamp within a preset time period in the target scene.
[0090] In an optional implementation, the above step S2011 includes:
[0091] Step a1, video file acquisition includes:
[0092] Determine the data source: First, clearly identify the location and related information of the video surveillance equipment in the target scene. In a retail store scenario, you need to understand the installation location, coverage area, and corresponding storage paths of each surveillance camera within the store. These cameras may be connected to a local storage server or uploaded to cloud storage via the network.
[0093] Time Range Setting: Based on a preset time range, enter the start and end times in the monitoring device's management system or storage server's search interface. For example, to retrieve video files from a specific store between 10:00 AM and 12:00 PM on October 1, 2024, precisely set this time range in the system. The system will search and locate relevant video clips within the storage media based on this time range.
[0094] File Extraction: Once you've located video clips that fit the time range, extract them using the download or export functionality provided by the management system. If the video files are stored on a local server, they can be downloaded directly to a designated work computer using a file transfer protocol (e.g., FTP). If stored in the cloud, use the appropriate cloud storage client and follow the system prompts to download the files. Ensure that the obtained video files contain the original timestamp information, which is typically automatically generated by the surveillance device during recording and embedded in the file attributes.
[0095] Step a2, obtaining the audio file includes:
[0096] Audio Capture Device Location: Determine the audio capture devices in the target scenario, such as the location and connection methods of microphones installed within a store. These microphones may be connected to audio recording devices or integrated into video surveillance equipment to also record audio. Understand the audio capture device parameters, such as sampling rate and number of channels, for subsequent processing.
[0097] Time range matching: Similar to video files, audio recording devices or related storage systems can be filtered based on a preset time period. The audio storage system might display recorded audio clips in a timeline format. By dragging the time slider or directly entering a time range, you can locate audio files between 10:00 AM and 12:00 PM on October 1, 2024.
[0098] Audio File Export: After finding an audio file that fits the time range, export it using the audio recording device's built-in export function or related audio processing software. If the audio file format needs to be converted (for example, from the original audio format to the commonly used MP3 format), perform the format conversion during the export process. Ensure that the exported audio file retains the original timestamp information, which is consistent with the video file timestamp on the time basis, providing a basis for subsequent synchronization alignment.
[0099] Step S2012: align the timestamps of the video file and the audio file with the timestamp using a synchronization alignment technology to obtain a video file and an audio file with the same timestamps.
[0100] In some optional implementations, the above step S2012 includes:
[0101] Step b1, timestamp information parsing, includes:
[0102] Video file timestamp parsing: Use video processing software or programming libraries (such as the OpenCV library in Python) to read acquired video files. These tools can parse the video file metadata and extract the timestamp information. Timestamps may be stored in various formats, such as time-coded formats (e.g., Unix timestamps, which represent the number of seconds from January 1, 1970, 00:00:00 UTC) to the current time) or custom time formats. Using appropriate parsing algorithms, convert the timestamps into recognizable time objects for subsequent calculations and comparisons.
[0103] Audio file timestamp parsing: For audio files, use audio processing tools (such as FFmpeg or Python's pydub library) to parse the audio file's metadata and extract the timestamp information. Similar to video files, convert the timestamps in the audio file to a unified time format to ensure they are on the same time scale as the video file's timestamps for easy alignment.
[0104] Step b2, synchronous alignment algorithm application.
[0105] Time offset-based alignment: Calculate the time offset between the video and audio file timestamps. Compare the difference in the start timestamps of the video and audio files. For example, if the start timestamp of the video file corresponds to 10:00:01 on October 1, 2024, and the start timestamp of the audio file corresponds to 10:00:03 on October 1, 2024, then the audio file has a 2-second delay relative to the video file. This time offset calculation determines the required time adjustment for the audio file.
[0106] Timestamp Correction: Correct the timestamps of the audio file based on the calculated time offset. Within the audio processing tool, use the appropriate timestamp modification function or algorithm to add or subtract the corresponding offset from each timestamp in the audio file, aligning the timestamps of the audio file with those of the video file. For example, subtract 2 seconds from all timestamps in the audio file to align them with the time base of the video file.
[0107] The multimodal information matching method provided in this embodiment uses synchronous alignment technology to align the timestamps of video files with timestamps and audio files with timestamps to obtain video files and audio files with consistent timestamps, providing a reliable data foundation for subsequent analysis and processing based on multimodal data.
[0108] In step S202 , object detection technology, voiceprint recognition technology, and natural language processing technology are used to perform multimodal processing on the audio and video files, and multi-dimensional feature sets corresponding to different task objects are extracted from the multimodal processed audio and video files.
[0109] Specifically, natural language processing technology includes speech recognition technology and large language models; different task objects include fixed-identity personnel and random-identity personnel; the above step S202 includes:
[0110] Step S2021: Perform video modality processing on the video file using object detection technology to determine target time periods for individual conversations between different task objects in the video file.
[0111] In an optional implementation, the above step S2021 includes:
[0112] Step c1, object detection model selection and initialization: Select a suitable object detection model, such as the YOLO (You Only Look Once) series model or the Faster R-CNN model. These models have been trained on a large number of image datasets and can recognize a variety of object categories. According to the target scene and task requirements, determine the object category that the model needs to recognize. In the dialogue scene, the main task object (such as people) is to be recognized. Initialize the selected object detection model, including loading the pre-trained model weights and configuring the model's operating parameters, such as the size of the input image, the detection threshold, etc. For example, the input image size of the YOLO model is set to match the video frame resolution to ensure that the model can accurately process the video data, and the detection threshold is set to 0.5, that is, only when the model's recognition confidence in the object is higher than 0.5, is the object considered to be detected.
[0113] Step c2, video frame processing and person detection: Read video frames frame by frame from a video file with consistent timestamps. Use a video processing library (such as OpenCV) to decompose the video file into a continuous sequence of image frames. For each frame of the image, input it into the initialized object detection model for processing. The object detection model analyzes the input video frame, identifies the human target, and returns the location information of the person (such as bounding box coordinates). For example, the model detects that there are two people in the video frame, and returns the bounding box coordinates of person A as (x1, y1, x2, y2) and the bounding box coordinates of person B as (x3, y3, x4, y4), indicating the position range of the person in the image.
[0114] Step c3, Character Interaction Analysis and Target Time Determination: Based on the character position information in consecutive video frames, character interactions are analyzed. By tracking the characters' motion trajectories, it is determined whether the characters are in close proximity and engaging in interactive actions, such as standing face-to-face or exchanging body language. For example, if the bounding box distance between characters A and B is detected within a certain threshold range across multiple consecutive frames, and character A's body movements tend to be directed toward character B, it can be preliminarily determined that the two are interacting. Combined with the audio file's time information (since the video and audio are timestamp-aligned), the target time period for private conversations between the characters is determined. When an interaction is detected between the characters, the corresponding video frame timestamp is recorded. This time period is then associated with the corresponding time period in the audio file to determine if the two individuals are likely engaged in a private conversation within that time period. For example, if character A and character B are detected interacting from frame 100 to frame 200, and the corresponding video timestamps are 0:05-0:10, then the period 0:05-0:10 in the audio file is the potential target time period for private conversations. The entire video file is analyzed to determine the target time periods for private conversations between all different task objects.
[0115] For example, object detection technology, such as YOLOV5, can be used to capture the time points of private conversations between sales staff and customers to ensure the purity of the data source. Object detection technology can be used to detect single-person conversations between a salesperson and a customer in a store, and the start timestamp is returned. When the customer leaves the store or a second customer enters the store, the end timestamp is returned. The time between the start and end timestamps is the pure time period data for the customer.
[0116] Step S2022: Use voiceprint recognition technology to perform audio modality processing on the audio file to obtain the audio of each task object in the audio file within the target time period, and mark each task object audio.
[0117] In an optional implementation, the above step S2022 includes:
[0118] Step d1: Voiceprint recognition model preparation:
[0119] Choose a professional voiceprint recognition model, such as a deep learning-based DNN (deep neural network) or CNN (convolutional neural network) voiceprint recognition model. These models need to be trained on a large-scale voiceprint database to learn the voiceprint characteristics of different individuals. Obtain and organize a voiceprint database for training and testing. The database contains voice samples of different task subjects and their corresponding identities.
[0120] Train and optimize the voiceprint recognition model. Iteratively train the model using training data and adjust its parameters to accurately extract and identify different voiceprint features. During training, use cross-validation and other techniques to evaluate model performance. Optimize the model by adjusting its structure and training parameters to improve voiceprint recognition accuracy.
[0121] Step d2, audio data preprocessing:
[0122] Extract audio data from the target time period from audio files with consistent timestamps. Use an audio processing library (such as pydub) to segment the audio files according to the previously determined target time period, generating audio segments containing only the target conversation. Preprocess the extracted audio segments to improve voiceprint recognition accuracy. This preprocessing step includes noise removal, using filtering algorithms (such as Wiener filtering) to remove interference such as ambient noise and current noise. The audio signal amplitude is normalized to maintain consistent volume levels across different audio samples, preventing volume differences from affecting voiceprint feature extraction.
[0123] Step d3: Voiceprint recognition and audio annotation:
[0124] The preprocessed audio clip is input into the trained voiceprint recognition model. The model analyzes the audio, extracts the voiceprint features, and compares them with the known voiceprint features in the voiceprint database. For example, the model extracts the voiceprint feature vector from the audio clip and calculates the similarity with the voiceprint feature vectors of task objects A, B, C, etc. in the database one by one. Based on the similarity matching results, the identity of each task object in the audio clip is determined, and the audio is annotated. If the model determines that the voiceprint features of a certain audio segment are most similar to the voiceprint features of task object A and exceed the set similarity threshold (such as 0.8), then the audio segment is annotated as the audio of task object A. All audio clips within the target time period are identified and annotated one by one to obtain the audio annotation results of each task object within the target time period.
[0125] Exemplarily, the audio data corresponding to the time nodes confirmed by the above-mentioned video data are segmented for subsequent speaker separation and speech recognition. The audio data is then cut according to the silent points (i.e., breath points) using Voice Activity Detection (VAD) technology. The content between each breath point must be a continuous speaking audio segment of a person. Acoustic feature extraction is performed on each continuous speaking audio segment (calculating its Mel-frequency cepstral coefficients), and the acoustic features of each segment are represented by a fixed-dimensional vector generated by a deep learning model. Speaker separation technology calculates feature vectors for the voiceprints of sales staff and customers, calculates cosine similarity for the embedded vectors, and regards the two audio segments with higher similarity as being spoken by one person, and vice versa, as being spoken by two people, i.e., speaker embedding technology. After processing by speaker embedding technology, it can be ensured that the vectors of different segments of the same speaker are closer, and the segments of different speakers are farther apart, so that clustering (k-means) can be performed accordingly to distinguish the voices of sales staff and customers and label them.
[0126] The formula for segmentation according to the silence point (i.e., breath opening) in the above-mentioned voice activity detection is as follows:
[0127]
[0128] Where E represents the energy threshold, ZCR represents the zero-crossing rate, x[n] represents the audio signal, n represents the time series index of the audio signal, and represents the nth sampling point in the signal. N is the length of the audio signal, that is, the value of n is from 0 to N-1. Formulas (1) and (2) are used to determine which areas are human voice areas (the energy domain and zero-crossing rate of human voice audio are different from those of noise and silence) and retain them; which areas are non-human voice areas and delete them.
[0129] The calculation formula of Mel cepstral coefficient is as follows:
[0130] y[n]=x[n]-αx[n-1] (3);
[0131] N=f s ·T (4);
[0132]
[0133] Among them, f s represents the sampling rate, T represents the frame length, y[n] represents the pre-emphasized signal, α represents the pre-emphasis coefficient, w[n] represents the Hamming window function, X[k] represents the frequency domain representation after discrete Fourier transform, x i[n] represents the signal after pre-emphasis and windowing, j represents an imaginary number, k represents the frequency index, which represents the kth frequency component, M(f) represents the Mel frequency scale formula, which is used to map the linear frequency to the Mel frequency to make it more consistent with the human auditory system and perception characteristics, m k Indicates the Mel frequency corresponding to the center frequency of the kth Mel filter, h m (k) represents the transfer function of the mth Mel filter, which defines the response of the filter at different frequencies, m l Indicates the left boundary frequency of the Mel filter, m c Indicates the center frequency of the Mel filter, m r Indicates the right boundary frequency of the Mel filter, E m Represents the energy output of the mth Mel filter, which represents the energy retained after passing through the filter.
[0134] In step S2023, speech recognition technology is used to perform text modal processing on the audio of each marked task object to obtain the text content corresponding to each marked task object audio, and a large language model is used to perform context analysis on the text content corresponding to each marked task object audio to verify and correct the text content that is incorrectly marked due to the similar timbre of different task objects.
[0135] In an optional implementation, the above step S2023 includes:
[0136] Step e1, speech-to-text (speech recognition):
[0137] Select a mature speech-to-text tool or engine, such as Baidu Speech Recognition or iFlytek Speech Recognition, or use an open-source speech recognition framework such as Kaldi. These tools and frameworks are trained on large amounts of speech data and can convert audio into text. Perform speech-to-text conversion on each labeled audio clip for each task object. Input the audio clip into the selected speech recognition tool. The tool analyzes and processes the audio signal, identifies the speech content, and outputs the corresponding text. For example, input an audio clip from task object A into the speech recognition engine, and the engine outputs the text "I think the color of this product is very attractive." Batch process all labeled audio clips for the task objects to obtain the text content corresponding to each audio clip.
[0138] Step e2: Load and configure the large language model:
[0139] Select an appropriate large language model, such as the GPT series or LLaMA, and ensure that the model is correctly loaded into your local environment or cloud platform. Configure the large language model according to task requirements, setting parameters such as the model's input and output formats and the maximum generated text length. For example, set the model's input format to match the text format generated by speech-to-text conversion and set the maximum generated text length to 200 words to control the length of the model's output content.
[0140] Step e3: Context analysis and text verification and correction:
[0141] The text generated by speech-to-text conversion is input into the configured large language model. The large language model uses its powerful natural language understanding capabilities to perform contextual analysis on the text. The model determines the proper attribution of different sentences in the text based on the semantic and grammatical relationships between the surrounding context. For example, in a conversation, based on conversational logic and common language patterns, it can be determined that a sentence asking about the price is more likely to be uttered by a customer, while a sentence introducing product features is more likely to be uttered by a store clerk.
[0142] For text content that may be mislabeled due to similar timbre between different task subjects, the large language model verifies and corrects it through contextual analysis. If the model finds that the attribution of a certain text sentence does not match the contextual logic, for example, according to conversational logic, a sentence expressing purchase intent should be spoken by the customer, but is currently labeled as spoken by the clerk, the model will correct the annotation based on the analysis results and re-label the text sentence as spoken by the customer. All text content generated by speech-to-text conversion is subjected to contextual analysis and verification and correction to improve the accuracy of text content annotation and provide a reliable data foundation for subsequent text-based analysis and processing.
[0143] For example, in the sales field, the text content of speech recognition is rewritten into a conversational format. This is then combined with the contextual understanding of a large language model to verify and correct the information based on the conversational content and character logic. This ensures that there are no mismatches between the extracted information and the customer's identity. The conversation's words are accurately distinguished from those of the customer, preventing the salesperson's personal characteristics from contaminating the customer's, thus lowering the analysis threshold of the large language model. However, for the "salesperson performance evaluation" requirement, there is no such requirement to match customer identity with information. Instead, a comprehensive and inductive evaluation is required based on data from a large number of staff members on their performance with different customers. For this task, the daily audio data is retained in its entirety and speech recognition is performed on the entire data. The contextual analysis capabilities of the large language model are then leveraged to directly analyze the salesperson's service level.
[0144] After performing speech recognition on the processed segmented and overall audio data, the recognized text becomes the primary source of information extraction. To address the diverse needs of different fields, some words may be commonly used in certain fields but unfamiliar to the general public. To prevent these words from being mistakenly recognized as homophones during speech recognition, a specific hotword library is constructed. This library includes various proper nouns. When the speech recognition model recognizes a syllable identical to a hotword, it outputs the hotword as a whole, rather than guessing at a homophone. This process ensures speech recognition accuracy.
[0145] After speech recognition, the text content will be polished again. Since this step utilizes the function of large language models to analyze context, the polished data will not only filter out meaningless modal particles, delete duplications and fill in omissions, but also correct the same noun in the speech recognition step that is recognized as another noun with the same pronunciation but different meaning in the context, maintain semantic consistency, and further improve recognition accuracy.
[0146] Step S2024, using a large language model to extract the first appearance features, first body movement features, and first emotional expression features of a person with a fixed identity from the video file of the target time period, and extract the second appearance features, second body movement features, and second emotional expression features of a person with a random identity.
[0147] In an optional implementation, the above step S2024 includes:
[0148] Step f1, parsing and input of video files:
[0149] Video frame extraction: Use a video processing library (such as OpenCV) to open the video file for the target time period. Based on the video's frame rate information, determine the video frames to be extracted. To comprehensively and accurately extract features, video frames are usually extracted at regular intervals, such as every 10 frames. This ensures that key information in the video is covered while reducing the amount of computation. The extracted video frames are stored as image files for subsequent input into the large language model. For example, for a video file with a duration of 1 minute and a frame rate of 30 frames per second, 18 frames of image can be extracted by extracting every 10 frames.
[0150] Image preprocessing: Extracted video frames are preprocessed to improve the model's processing performance. This preprocessing step includes image enhancement, such as using histogram equalization to enhance image contrast and make features more visible; image resizing to uniformly resize all images to the model's input dimensions, such as scaling images to 224×224 pixels; and normalization to map image pixel values to a range of 0–1 to meet the model's input specifications. These preprocessing operations ensure the quality and consistency of the image data fed into the model.
[0151] Step f2, feature extraction:
[0152] Appearance Feature Extraction: Preprocessed video frames are fed into a large language model, which uses its visual understanding capabilities to identify the appearance of both fixed-identity individuals and random-identity individuals. For the first appearance feature of a fixed-identity individual, the model can identify their clothing style, such as formal suit, casual wear, or workwear, and describe the color and style of the clothing, such as "wearing a dark blue slim-fit suit, a white shirt, and a red tie." For hairstyle, the model can identify types such as short, long, or curly hair, and further describe the length and characteristics of the hairstyle, such as "straight black hair that reaches shoulder length." For facial features, the model can identify the general shape and characteristics of the eyes, nose, and mouth, such as "big eyes, double eyelids, and a high nose bridge." For the second appearance feature of a random-identity individual, the model uses the same method for identification and description, while also focusing on appearance features that may influence consumer behavior or context analysis, such as accessories (hats, scarves, glasses, etc.), such as "wearing a black baseball cap, a colorful scarf around the neck, and a pair of round glasses."
[0153] Body movement feature extraction: The large language model extracts body movement features by analyzing the joint positions and motion trajectories of people in video frames. For the first-level body movement features of fixed-identity individuals, in sales scenarios, the model can identify pointing gestures used when introducing products, such as "right hand extended, index finger pointing at the displayed product," and calculate quantitative features such as the frequency and duration of pointing gestures. For body posture, the model can determine whether the individual is standing, sitting, or walking, as well as the postural characteristics of standing or sitting, such as "slightly leaning forward, maintaining an active communication posture." For the second-level body movement features of random-identity individuals, in shopping mall scenarios, the model can identify the actions of customers picking up items, described as "picking up an item from the shelf with the left hand and carefully observing it," as well as information such as the walking route and the location of the stop, such as "staying in the electronics area for a long time, walking around the mobile phone display cabinet."
[0154] Emotional expression feature extraction: Leveraging the large language model's ability to analyze facial expressions, we extract emotional expression features. For individuals with fixed identities, the model identifies their emotional state by identifying facial muscle movement patterns, such as a smile indicating happiness and a frown indicating thought or dissatisfaction. For example, "the face appears smiling, the corners of the mouth are raised, and the eyes are slightly narrowed, indicating a positive emotion." For individuals with random identities, the model similarly determines their emotions based on facial expressions, while also combining body movements and scene information for a comprehensive analysis. For example, if a customer crosses their arms and frowns slightly while communicating with a store clerk, the model can determine that they may be expressing doubt or dissatisfaction, describing it as "arms crossed in front of the chest, brows slightly furrowed, expressing some doubt about the product the clerk is introducing."
[0155] Step f3, feature organization and structuring: The appearance, body movement, and emotional expression features extracted by the large language model from different video frames are organized. For people with fixed identities, the first appearance feature, first body movement feature, and first emotional expression feature corresponding to each video frame are arranged in chronological order to form a feature sequence. For example, in a 5-minute video, 30 frames of images are extracted. The model organizes the features extracted from each frame into a list, and each element in the list contains a description of the appearance, body movement, and emotional expression corresponding to the frame. For people with random identities, the second appearance feature, second body movement feature, and second emotional expression feature are also organized to form a corresponding feature sequence.
[0156] The organized features are structured and converted into a format that facilitates subsequent analysis and storage. Using JSON format, the feature information can be organized into a structure that includes fields such as person identification (such as employee number for fixed-identity personnel, or anonymous number for random-identity personnel), timestamp (corresponding to the time of the video frame), description of appearance features, description of body movements, and description of emotional expression.
[0157] Step S2025, based on the preset prompt words, a large language model is used to extract the first language expression features and chat skill features of the fixed-identity person from the text content corresponding to the audio of each labeled task object, as well as the second language expression features, feedback features and satisfaction scores of the random-identity person who matches the preset prompt words.
[0158] Specifically, choose an appropriate large language model with strong natural language processing capabilities, such as GPT-3.5-Turbo or Claude. Ensure that the model is initialized and that parameters related to text processing are configured, such as the maximum number of generated characters and the temperature coefficient (used to control the randomness of generated text).
[0159] Pre-set prompts are carefully designed to meet the different needs of fixed-identity and random-identity personnel. For fixed-identity personnel, prompts such as "Please analyze the speaker's language fluency, vocabulary richness, use of professional terms, and conversation-guiding skills in this text" are set to extract first language expression features and chat skills. For random-identity personnel, prompts such as "Extract customer feedback about products or services from the text, determine the positive or negative nature of the feedback, and, if mentioned, extract the satisfaction score" and "Analyze the conciseness of the customer's language expression and whether there are specific needs expressed" are set to extract second language expression features, feedback features, and satisfaction scores. These prompts are designed to guide the large language model to accurately extract the required features from the text content corresponding to the audio of the labeled task object.
[0160] Fixed-identity feature extraction: The text corresponding to the annotated fixed-identity audio is fed into a large language model, segment by segment. The model analyzes the audio based on pre-set prompts and extracts first-language expression features. For example, language fluency is assessed by counting pauses and repeated words in the text. The types and frequencies of different words in the text are counted, and vocabulary richness and the use of specialized terminology are determined by combining domain-specific vocabulary lists. Regarding conversational skill features, the model analyzes the questioning style (e.g., the ratio of open-ended to closed-ended questions), topic-initiating statements (e.g., "Next, let's talk about the advantages of this product"), and response strategies to the other party's points (e.g., whether the speaker actively affirms and expands on the topic). It extracts relevant features and provides quantitative or qualitative descriptions. For example, the output might be, "High language fluency, few pauses; moderate vocabulary richness, using three product-related professional terms; in terms of conversational skills, the speaker is adept at using open-ended questions to guide the conversation, with questions accounting for 40% of the time; and actively responds to the other party's points, affirming and expanding on the topic 60% of the time."
[0161] Random Person Feature Extraction: The text corresponding to the annotated random person audio is also input into the large language model. The model extracts second language expression features based on preset cues, analyzing the text for grammatical correctness, sentence structure complexity, and directness of expression. For example, it determines whether the customer expresses their needs concisely and clearly or with more subtlety and complexity. Regarding feedback features, the model identifies evaluative statements about products or services in the text and uses sentiment analysis algorithms (such as a bag-of-words model combined with a sentiment lexicon, or a pre-trained sentiment analysis neural network model) to determine whether the feedback is positive (e.g., "This product is great, very easy to use"), negative (e.g., "This service is terrible, I waited a long time and no one responded"), or neutral. If the text explicitly mentions a satisfaction rating, the model directly extracts that rating, such as "I give this product an 8." If a rating is not directly mentioned, but there is sufficient feedback content, the model estimates a relative satisfaction rating based on the sentiment analysis results and internally defined rating mapping rules. For example, positive feedback is mapped to a score of 7-10, negative feedback to a score of 1-3, and neutral feedback to a score of 4-6. The final output is such as "the language expression is relatively direct and the grammar is basically correct; the feedback is positive, mentioning that the product has powerful functions and a good experience; the estimated satisfaction score is 8 points."
[0162] Step S2026, taking the first appearance feature, the first body movement feature, the first emotion expression feature, the first language expression feature, the chat skill feature, the feedback feature and the satisfaction score as the first multidimensional feature set; taking the second appearance feature, the second body movement feature, the second emotion expression feature, the second language expression feature, the feedback feature and the satisfaction score as the second multidimensional feature set.
[0163] Specifically, the first appearance feature, first body movement feature, and first emotional expression feature of the fixed-identity person previously extracted from the video file are combined with the first language expression feature and chat skill features extracted from the audio text. For appearance features, the previously described information, such as clothing style, hairstyle, and facial features, is retained; for body movement features, quantitative or qualitative descriptions, such as the frequency of pointing movements and body posture, are retained; and for emotional expression features, descriptions of emotional states, such as smiling and frowning, are maintained. These features are then fused with the newly extracted language expression and chat skill features. For example, features such as "wearing a dark blue slim-fitting suit, neatly styled hair, and a smiling facial expression" and "frequently pointing with the right hand and leaning forward to communicate actively" from the previous video analysis are combined with features such as "high language fluency and good at guiding conversations" from the audio text analysis. Furthermore, if any customer feedback features and satisfaction ratings regarding the fixed-identity person's service are obtained from the audio text, these are also incorporated into the first multidimensional feature set. Ultimately, a multidimensional feature set is formed that comprehensively describes the fixed-identity person's performance in business scenarios. Each feature has a clear definition and description, facilitating subsequent analysis and application.
[0164] The second appearance features, second body movement features, and second emotional expression features of a random person extracted from the video file are integrated with the second language expression features, feedback features, and satisfaction scores extracted from the audio text. Appearance features include information such as age range, gender, and accessories; body movement features include descriptions of picking up products and walking routes; and emotional expression features include whether interest or doubt is expressed. These features are combined with language expression features (such as conciseness of expression and how the request is expressed), feedback features (positive, negative, or neutral), and satisfaction scores obtained from audio text analysis. For example, by combining features such as "young woman wearing a hat, spending a long time in the product display area, showing some interest" from video analysis with "direct language expression, positive feedback on the product color, and a satisfaction score of 9" from audio text analysis, a second multidimensional feature set is constructed that comprehensively reflects the random person's behavior and attitude in the scene. This feature set provides rich data support for companies to understand customer needs and optimize products and services.
[0165] Step S203: Construct a character portrait based on the multi-dimensional feature set corresponding to different task objects. Figure 1 Step S103 of the illustrated embodiment will not be described in detail here.
[0166] Step S204: Match the character portrait to a specific person using a feature comparison algorithm based on the pre-recorded timestamp. Figure 1 Step S104 of the illustrated embodiment will not be described in detail here.
[0167] The multimodal information matching method provided in this embodiment can capture visual information such as a person's body movements, expressions, and positional movements through video files; and acoustic information such as voice content, intonation, and speaking speed through audio files. Using synchronous alignment technology to align the timestamps of time-stamped video and audio files, it is possible to achieve precise temporal matching of video and audio data. This means that visual information such as a person's movements and expressions accurately correspond to audio information such as their spoken words and intonation. Video and audio files with consistent timestamps can fully present the information within the target scene, ensuring the integrity of information collected from multiple dimensions and providing comprehensive data support for in-depth exploration of potential patterns and behavioral patterns within the scene. Voiceprint recognition technology is used to perform audio modality processing on audio files to obtain and annotate the audio of each task subject within the target time period, effectively solving the problem of audio confusion between multiple people. Voiceprint characteristics of different people are as unique as fingerprints, and voiceprint recognition technology can accurately distinguish the voices of different task subjects. Speech recognition technology is used to convert the annotated task audio into text. A large language model is then used for contextual analysis to verify and correct mislabeled text due to similar timbre, significantly improving the accuracy and usability of the text. By understanding the text context and conducting semantic analysis, the large language model can determine the legitimacy of text attribution and promptly identify and correct errors. This ability to accurately verify and correct text provides high-quality data for subsequent text-based analysis, enhancing the accuracy and robustness of the audio-to-text conversion process.
[0168] In this embodiment, a multimodal information matching method is provided, which can be used in the above-mentioned computer device. Figure 3 is a flow chart of a multimodal information matching method according to an embodiment of the present invention. Figure 3 As shown, the process includes the following steps:
[0169] Step S301: Obtain the audio and video files to be extracted for different task objects in the target scene. Figure 2 Step S201 of the illustrated embodiment will not be described in detail here.
[0170] Step S302: Use object detection technology, voiceprint recognition technology and natural language processing technology to perform multimodal processing on the audio and video files, and extract multi-dimensional feature sets corresponding to different task objects from the multimodal processed audio and video files. Figure 2 Step S202 of the illustrated embodiment will not be described in detail here.
[0171] Step S303: constructing a character portrait based on the multi-dimensional feature sets corresponding to different task objects.
[0172] Specifically, the above step S303 includes:
[0173] Step S3031: Acquire the appearance information of each person with a random identity and the appearance information of each person with a fixed identity from the video files within the target time period.
[0174] Specifically, for the extraction of appearance information of individuals with fixed identities, image recognition technology and algorithms are used to analyze the appearance of individuals with fixed identities within the extracted video frames. First, the individual's clothing style is identified by matching the image's color distribution and texture features with pre-set clothing style templates. For example, if the image shows predominantly dark clothing with regular stripes or plaids, and the style conforms to typical business attire, the clothing style is determined to be formal business attire. For hairstyles, edge detection algorithms and shape recognition technologies are used to identify the hair's outline shape and length. For example, if the hair's outline exhibits neat, short edges and the length is within a specific range, the hairstyle can be determined to be short. For facial features, a facial landmark detection algorithm is used to locate the position and shape of key features such as the eyes, nose, mouth, and eyebrows. For example, the aspect ratio of the eyes, the shape of the corners of the eyes, and the height and width of the nose are detected to describe facial features, such as "big eyes, double eyelids, and a high and straight nose bridge." This appearance information extracted from different video frames is integrated and comprehensively judged to obtain more accurate and comprehensive appearance information of individuals with fixed identities.
[0175] Appearance Information Extraction for Random Identity Individuals: Also based on extracted video frames, the appearance of random identity individuals is identified. Due to the large number of random identity individuals and their diverse features, a more versatile and adaptable image analysis method is employed. To determine age ranges, features such as facial wrinkles, skin texture, and facial feature proportions are combined with machine learning models for prediction. For example, a trained convolutional neural network model is fed a facial image. Based on the learned facial feature patterns for different age groups, the model outputs an age range for the individual, such as "20-30 years old." For gender identification, a deep learning-based gender classification model analyzes facial contours, facial features, and hairstyle to determine the individual's gender. For accessories, image segmentation technology is used to identify the presence and type of accessories such as hats, glasses, and necklaces. For example, if a round lens and frame are detected in the head region of an image, it can be determined that the individual is wearing glasses. This appearance information is organized and recorded to form a collection of appearance information for random identity individuals.
[0176] Step S3032: construct a character portrait of the random person based on the preset timestamp and the second multi-dimensional feature set.
[0177] In an optional implementation, the above step S3032 includes:
[0178] Step g1: Timestamp association and data screening:
[0179] First, define the preset timestamp range, which corresponds to the target time period. Within the second multidimensional feature set, filter out data related to that timestamp range. For example, if the timestamp range is 10:00-12:00 on October 1, 2024, find the various feature data of random individuals recorded within that time period from the feature set, including secondary appearance features, secondary body movement features, secondary language expression features, feedback features, and satisfaction scores. Ensure that the data used is closely aligned with the target time period to avoid inaccurate profile construction due to time mismatches.
[0180] Step g2, feature integration and weight allocation:
[0181] The selected features are integrated. For physical features, the age range, gender, and accessories previously extracted from the video files are combined. Physical features include walking routes, resting locations, and interactions with products. Language features include conciseness of expression and how needs are expressed. Feedback features are determined to be positive, negative, or neutral. Satisfaction scores, if any, are directly incorporated. Weights are assigned to each feature based on business needs and the importance of each feature to the persona. For example, in a retail scenario, feedback features and satisfaction scores are crucial for understanding customer needs and purchase intentions and can be assigned higher weights, such as 0.3 and 0.25. Age range and gender, among physical features, play a role in market segmentation and can be assigned weights of 0.1 and 0.05, respectively. Physical features and language features are assigned weights of 0.15 and 0.15, respectively, based on their impact on analyzing customer behavior patterns. These weights are not fixed and can be adjusted based on actual data analysis and business experience.
[0182] Step g3, character portrait construction and presentation:
[0183] Based on the integrated and weighted feature data, a character profile of a random person is constructed. Vector representation can be used to combine each feature and its corresponding weight into a multidimensional vector. For example, the character profile vector of a random person might be [age range (0.1), gender (0.05), accessories (0.05), body movement frequency (0.15), language conciseness (0.15), feedback positivity (0.3), satisfaction score (0.25)], where the weight of each feature is in parentheses. To present the character profile more intuitively, this vector data can be converted into a visual chart, such as a radar chart. The feature of each dimension serves as an axis of the radar chart, and the vector value corresponds to the scale on the axis. The radar chart clearly shows the performance of the random person in each feature dimension, facilitating market analysis and customer segmentation for companies.
[0184] Step S3033: construct a character portrait of the person with a fixed identity based on the preset timestamp and the first multi-dimensional feature set.
[0185] In an optional implementation, the above step S3033 includes:
[0186] Step h1, timestamp matching and data extraction:
[0187] After determining the preset timestamp range, accurately match and extract the data of the fixed-identity personnel related to the timestamp in the first multi-dimensional feature set. For example, the timestamp corresponds to a certain working time period, and the first appearance features (dressing style, hairstyle, facial features, etc.), first body movement features (indicating movement frequency, body posture, etc.), first emotional expression features (emotional states such as smiling and frowning), first language expression features (language fluency, vocabulary richness, use of professional terms, etc.), chat skills features (questioning methods, topic guidance ability, etc.) and possible feedback features and satisfaction scores (if customers have feedback on their services) of the fixed-identity personnel during that time period are extracted from the feature set. Ensure that the extracted data accurately reflects the work performance and status of the fixed-identity personnel during that time period.
[0188] Step h2, feature fusion and comprehensive evaluation:
[0189] The extracted features are integrated and comprehensively evaluated based on business objectives. For appearance features, their suitability for the work scenario is considered, such as whether the salesperson's attire meets the company's image requirements. Body language features assess their effectiveness in work communication, such as whether their instructions are clear and accurate. Emotional expression features analyze their impact on the work atmosphere and customer relationships, such as whether positive emotions contribute to improved customer satisfaction. For language expression and conversation skills features, their ability to communicate and promote products or services is assessed based on the work tasks and objectives. Different weights are assigned to each feature based on its importance. For example, in sales work, language expression and conversation skills have higher weights, such as 0.3 and 0.25, respectively. Appearance features are weighted at 0.15 based on industry and job requirements. Body language and emotional expression features are weighted at 0.15 and 0.1, respectively. Through comprehensive evaluation and weighting, the performance of fixed-status individuals is comprehensively measured.
[0190] Use the weighted feature data to generate a character portrait of a fixed-identity person. A table can be used to list each feature, its corresponding description, and its weight score. The character portrait is presented in Table 1 below:
[0191] Table 1 Character portrait display
[0192] Feature Category Specific features describe Weighted score Physical characteristics Dress style Dark blue business suit 0.15 Body movement characteristics Indicates action frequency 3 times per minute 0.15
[0193] Step S304 , matching the character portrait to a predetermined person using a feature matching algorithm based on a pre-recorded timestamp.
[0194] Specifically, the above step S304 includes:
[0195] Step S3041, tracing the source based on the preset timestamp, and using a feature matching algorithm to compare the character portrait of the random identity person with the appearance information of each random identity person to obtain a matching random identity person.
[0196] In an optional implementation, taking the sales field as an example, customer information comes from several sources. One is the basic information that customers register at the store, which is known by sales staff; another is the customer's appearance information, such as gender and age, extracted through video recordings; and another is basic information that may be revealed in conversations between customers and sales staff, which is extracted through large language model analysis. The main source of customer features is only conversation. By fine-tuning the large language model, it can analyze the conversations between customers and sales staff and extract customer features such as dietary preferences, interests, and preferences for various products. Since we retain the timestamp corresponding to the data when processing data in each modality (the timestamp of the recording data is also inherited when converting speech to text), we only need to match the same timestamp to achieve a one-to-one match between features and appearance information. The sales staff who receive this customer can then fill in the basic information they know based on the customer's appearance information in the database to achieve a perfect match between customer information and customer features.
[0197] The above step S3041 includes:
[0198] Step i1, character portrait arrangement: further sort out the character portrait data of random people with previously constructed identities. Convert the character portrait data represented in vector form, such as [age range (0.1), gender (0.05), accessories (0.05), body movement frequency (0.15), language expression simplicity (0.15), feedback positivity (0.3), satisfaction score (0.25)], into structured data that is easy to compare. Create a table containing the character portrait ID (used to uniquely identify each portrait), descriptions of various features and corresponding weight values. At the same time, clarify the preset timestamp range corresponding to each portrait. For example, for a portrait with a character portrait ID of 001, the corresponding timestamp range is 10:00-11:00 on October 1, 2024.
[0199] Appearance data organization: Organize the appearance information of each randomly identified person obtained from the video file. Categorize the appearance information by person, creating a table containing detailed information such as the person's appearance ID (corresponding to the person in the video frame), age range, gender, and accessories. Ensure that the appearance information is associated with the video frame timestamp. For example, for a person with appearance ID A001, the appearance information corresponds to the video frame timestamp range of October 1, 2024, 10:10-10:20.
[0200] Timestamp Association: Carefully check the timestamps of the person portrait data and appearance information data to ensure consistency in timescale. For overlapping timestamp ranges, establish associations between person portrait and appearance information. For example, if the timestamp range of person portrait ID 001 partially overlaps with the timestamp range of person appearance ID A001, record this association to prepare for subsequent feature comparison.
[0201] Step i2, feature comparison algorithm execution phase:
[0202] Given the data characteristics of the portraits and appearance information of random-identity individuals, an algorithm combining similarity calculation and fuzzy matching was selected. For numerical features such as age range and satisfaction rating, the Euclidean distance algorithm was used to calculate the degree of difference between the two. For example, if the age range in the portrait is 20-30 years old, and the age range of the person corresponding to a certain appearance information is 25-35 years old, the difference in age between the two is calculated using Euclidean distance. For categorical features such as gender and accessories, a rule-based fuzzy matching algorithm is used. For example, if the gender in the portrait is female and the gender in the appearance information is also female, then the categorical feature has a high degree of match; for accessories, if the portrait mentions wearing glasses and glasses are also detected in the appearance information, then the accessory features are considered a match.
[0203] Set algorithm parameters based on data distribution and business needs. For the Euclidean distance algorithm, set a reasonable distance threshold. When the calculated Euclidean distance is less than the threshold, the numerical features are considered matched. For example, set the Euclidean distance threshold for the age range to 5 (indicating that the age range difference within 5 years is considered a match). For the fuzzy matching algorithm, set the matching rules and weights for the category features. For example, the gender matching weight is 0.6, and the accessory matching weight is 0.4. The comprehensive category feature matching degree is calculated through weighted calculation.
[0204] The character portrait of each random person is compared with the corresponding appearance information one by one. For the character portrait ID 001, the Euclidean distance between it and the appearance information ID A001, such as the age range and satisfaction score, is first calculated, and then the fuzzy matching degree of the categorical features such as gender and accessories is calculated. The matching results of the numerical and categorical features are combined to obtain an overall matching score. For example, after calculation, the numerical feature matching score is 0.7, and the categorical feature matching score is 0.8. According to the preset weights (assuming the numerical feature weight is 0.4 and the categorical feature weight is 0.6), the overall matching score is calculated to be 0.76. The same calculation is performed on all character portraits and appearance information to obtain a series of matching scores.
[0205] Sort the calculated matching scores from high to low. Set a matching score threshold based on business needs to filter out matching results with scores above the threshold. For example, setting a matching score threshold of 0.7 will filter out combinations of person portraits and appearance information with matching scores greater than 0.7. These combinations are considered possible matches.
[0206] Re-verify the timestamp consistency of the selected potential matches. Ensure that the timestamp range of the person portrait and the timestamp range of the video frames corresponding to the appearance information highly overlap within the critical time period. For example, the timestamp range of the person portrait ID 001 is 10:00-11:00 on October 1, 2024, and the timestamp range of the appearance information ID A001 is 10:05-10:55 on October 1, 2024. The two timestamps have a high degree of overlap and meet the matching requirements. If a timestamp inconsistency is found, the match result will be excluded even if the match score is high.
[0207] After matching score screening and timestamp consistency verification, the final match is confirmed for the specified random person. The matching person's profile ID, appearance ID, and matching score are recorded. For example, a successful match between profile ID 001 and appearance ID A001 is recorded, with a matching score of 0.85. These matching results can be used in subsequent business scenarios such as customer analysis and market research, providing data support for companies to understand customer behavior and needs.
[0208] Step S3042, obtain the prior information of the fixed-identity person, trace the source based on the preset timestamp, and use the feature matching algorithm to compare the character portrait of the fixed-identity person with the prior information of each fixed-identity person to obtain the matching fixed-identity person.
[0209] In an optional implementation, the above step S3042 includes:
[0210] Step j1, character portrait data collation: The character portrait data of the person with a fixed identity constructed based on the first multi-dimensional feature set is collated. The character portrait data is presented in a tabular form, including the character portrait ID, the first appearance feature (dressing style, hairstyle, facial features, etc.), the first body movement feature (indicating action frequency, body posture, etc.), the first emotional expression feature (emotional state such as smiling and frowning), the first language expression feature (language fluency, vocabulary richness, use of professional terms, etc.), chat skill features (questioning methods, topic guiding ability, etc.) and possible feedback features and satisfaction scores (if any), and the preset timestamp range corresponding to each portrait is clarified. For example, the portrait with the character portrait ID P001 corresponds to a timestamp range of 9:00-11:00 on October 5, 2024, the dress style is dark blue business suit, the frequency of indicating action is 4 times per minute, and other detailed feature information.
[0211] Prior information data organization: Obtain prior information about individuals with fixed identities, such as basic information from employee files (name, employee number, position information, etc.), past work performance data (sales performance, customer reviews, etc.), and training records. Associate this prior information with the characteristics of the persona. For example, associate an employee's sales performance with the language expression and conversation skills characteristics in the persona. Create a table that contains the employee ID, various prior information, and the relationship with the persona characteristics. Ensure that the prior information is associated with a timestamp, for example, the employee's work performance data on October 5, 2024, corresponds to that timestamp.
[0212] Timestamp Association: Similar to identifying random individuals, verify the timestamps of the persona data and prior information data to establish a temporal association between the two. For persona and prior information with consistent timestamp ranges or significant overlap, clarify the association. For example, if the timestamp range of persona ID P001 coincides with the timestamp range of prior work-related information for employee ID E001 on October 5, 2024, record this association to provide a basis for subsequent feature comparison.
[0213] In step j2, considering the specialized nature and complexity of data on fixed-identity individuals, a method combining the Analytic Hierarchy Process (AHP) and the cosine similarity algorithm was used. First, the AHP was used to determine the weights of different features in the matching process. For example, in sales positions, language expression and conversational skills are crucial for job performance. AHP analysis determined their weights to be 0.4; appearance, based on company image requirements, was weighted 0.2; body language and emotional expression were weighted 0.15 and 0.1, respectively; and feedback and satisfaction scores were weighted 0.15. Then, for each feature dimension, the cosine similarity algorithm was used to calculate the similarity between the persona's feature vector and the prior information feature vector. For example, for language expression, quantitative data such as language fluency and vocabulary richness in the persona's feature vector were combined. The cosine similarity algorithm was then used to calculate the similarity between the feature vector and the prior information feature vector. A threshold for cosine similarity was set; when the calculated cosine similarity exceeded the threshold, the feature dimension was considered a match. For example, the cosine similarity threshold was set to 0.6.
[0214] A feature comparison calculation is performed on the portrait of each fixed-identity individual against the corresponding prior information. For the portrait with ID P001, a weighted calculation is performed on the cosine similarity results for each feature dimension based on the weights determined by the AHP. For example, a cosine similarity of 0.8 for language expression features yields a weighted score of 0.32 based on a weight of 0.4; a cosine similarity of 0.7 for appearance features yields a weighted score of 0.14 based on a weight of 0.2, and so on. The weighted scores of all feature dimensions are combined to produce an overall matching score. The same calculation is performed on the portraits and prior information of all fixed-identity individuals, resulting in a series of matching scores.
[0215] Step j3, Match Score Sorting and Filtering: Sort the calculated match scores from high to low. Based on business needs and historical data experience, set a match score threshold to filter out matches with scores above the threshold. For example, set the match score threshold to 0.7 to filter out combinations of person profiles and prior information with a match score greater than 0.7. These combinations are considered potential matches.
[0216] Timestamp consistency and business logic verification: Screened potential matches are double-checked for timestamp consistency and business logic. Ensure that the timestamps of the persona and prior information are consistent within key business time periods. Simultaneously, the legitimacy of the matching results is determined based on business logic. For example, if the persona's profile shows that an employee has excellent sales skills during a specific time period, and prior information also indicates that the employee has high sales performance during the same period, then the business logic is met. If the persona and prior information are logically inconsistent, the match will be excluded, even if the match score is high.
[0217] Confirmation and Recording of Match Results: After matching score screening, timestamp consistency, and business logic verification, the final match is confirmed for the specific individual. Information such as the matching persona ID, employee ID, and match score are recorded. For example, a successful match between persona ID P001 and employee ID E001 is recorded, with a match score of 0.8. These matching results can be used for human resources management tasks such as employee performance evaluation, talent selection, and training plan development, providing data support for optimizing human resource allocation within the enterprise.
[0218] Step S305 : evaluating the work level of the person with a fixed identity based on the first multi-dimensional feature set to obtain a work evaluation result.
[0219] Specifically, the above step S305 includes:
[0220] Step S3051, data preparation stage:
[0221] Arrangement of the first multi-dimensional feature set: Comprehensively sort out the constructed first multi-dimensional feature set. This feature set includes the first appearance features, first body movement features, first emotional expression features, first language expression features, chat skills features, feedback features, and satisfaction scores of people with established identities. Present these features in a structured table to ensure that each feature has a clear definition and value range. For example, create a table with headers for feature categories (such as appearance features, body movement features, etc.), specific feature descriptions (such as dress style, frequency of indicating actions, etc.), feature values (such as business formals, 5 times per minute, etc.), and corresponding weights (pre-set according to business importance).
[0222] Clearly identify each person with a specific identity within the feature set, such as their employee number. Ensure data integrity and address any missing values based on data characteristics and business needs. If the number of missing values is small and their impact on the overall assessment is minimal, delete the records containing them. If the number of missing values is high, use methods such as mean filling and regression prediction filling to fill in the missing values. For example, if the lexical richness indicator in the language expression feature is missing for an individual employee, fill in the missing data based on the mean lexical richness of other employees in the same position.
[0223] Setting Performance Assessment Standards: Develop detailed performance assessment standards based on the company's business objectives and job requirements. For each characteristic dimension, determine performance levels corresponding to different value ranges. For sales positions, for example, in terms of fluency within the first language expression characteristic, fewer than three pauses per minute is considered excellent; 3-5 pauses is good; 5-8 pauses is moderate; and more than eight pauses is poor. Regarding conversational skills, a success rate of over 80% in guiding customers toward purchase intent is considered excellent; a success rate between 60% and 80% is good; a success rate between 40% and 60% is moderate; and a success rate below 40% is poor. Regarding appearance, compliance with the company's dress code and overall presentation are considered satisfactory. Furthermore, higher evaluations may be given if the dress style enhances the company's image or meets the needs of specific business scenarios.
[0224] For example, when analyzing a salesperson's performance, their basic information and shift schedule are readily available. We segment the recordings of the employee during their shifts based on their shift schedule, creating the original audio data. Similarly, speech recognition and text polishing are performed to obtain text data about the employee during their work hours. This text data is fed into a large language model that has been fine-tuned using prompts. This model is informed that this is a conversation between a salesperson and a customer during work. The model automatically distinguishes between the salesperson and the customer, and based on the conversation analysis results, outputs an assessment of the salesperson's work process, including information on topics discussed, conversational techniques used, and any instances of customer abuse. This assessment is then combined with the basic information provided by the sales company to create the final salesperson performance assessment report.
[0225] Conduct an in-depth analysis of the work level assessment results. For employees rated as excellent, summarize their strengths in various characteristic dimensions and create case studies to share with other employees. For example, excellent employees excel in verbal expression and conversation skills; experience exchange meetings can be organized to allow them to share communication skills and sales techniques. For employees rated as medium or below, analyze which characteristic dimensions they are deficient in and develop targeted training plans. For example, if an employee scores low in body movement characteristics and emotional expression characteristics, relevant business etiquette and emotion management training courses can be arranged. At the same time, link the work level assessment results with employees' performance bonuses, promotion opportunities, etc. to motivate employees to improve their work level and provide strong support for the company's human resources management and business development.
[0226] The multimodal information matching method provided in this embodiment obtains the appearance information of each random-identity person and fixed-identity person from video files within a target time period, accurately capturing key information. Constructing a person portrait based on a preset timestamp and a multidimensional feature set enhances the timeliness and relevance of the portrait. The timestamp provides a clear time identifier for the data, allowing the appearance information to be closely temporally correlated with other multidimensional features (such as body movements and language expressions). Constructing person portraits for the fixed-identity person and the random-identity person based on the preset timestamp and the first and second multidimensional feature sets, respectively, significantly improves the completeness and accuracy of the portraits. For the fixed-identity person, the first multidimensional feature set covers multiple aspects such as appearance, body movements, emotional expression, language expression, and conversation skills. Combined with the timestamp, the constructed portrait comprehensively and accurately reflects their status and capabilities in the work scenario. For the random-identity person, the second multidimensional feature set, combined with appearance information, can deeply characterize their characteristics and needs during business interactions. The resulting accurate and complete person portraits provide strong support for enterprise business decision-making and optimization. Traceability based on preset timestamps can accurately track the activities of both random and fixed-identity individuals at different points in time. Timestamps act as time coordinates, linking together various behavioral data of individuals within the target time period. For random-identity individuals, a feature comparison algorithm compares person portraits with appearance information to accurately identify specific customers. For fixed-identity individuals, such as corporate employees, timestamps can trace each step of their workflow. By combining portraits with prior information, it can confirm whether the employee's work status met requirements at a specific time. This allows for precise identification of employee identities and performance, facilitating refined management within the enterprise. The first multidimensional feature set includes the first appearance feature, first body movement feature, first emotional expression feature, first language expression feature, and conversational skill feature of a given fixed-identity individual. This allows assessment of their performance to move beyond a single dimension and encompass a comprehensive assessment from multiple perspectives. Evaluation based on this multidimensional feature set results in more accurate and objective performance evaluations. Each feature reflects an employee's work status from a different perspective. Quantitative analysis of these features through scientific evaluation methods reduces subjective influence on the evaluation. Accurate job evaluations, based on the first multidimensional feature set, help companies improve overall performance and strengthen their culture. When employees clearly identify their developmental goals and enhance their capabilities through training, their performance and results improve, directly driving improved company performance.
[0227] As one or more specific application examples of the embodiments of the present invention, the multimodal information matching method provided by the present invention is further described in detail in conjunction with FIG. 4( a ) and FIG. 4( b ), as follows:
[0228] The multimodal information matching method of the present invention includes three parts: data processing, information extraction, information output and matching. For different task objects, the data processing, information extraction and matching methods are different. For various scenarios, statistically significant task objects can be roughly divided into two categories: one is people with fixed identities, who have been in the scene for a long time and have relatively complete identity prior information, such as sales personnel in the sales field, teachers in the education field, etc. For such objects, the focus of information extraction is mostly on the performance of fixed people facing different objects or events in fixed scenes. Therefore, through monitoring, recording and other equipment in the scene, the recordings and videos of such people during their working hours can be intercepted as effective information sources, and the data can be concentrated for comprehensive processing. The extracted information is often summarized from many data containing original information, and the final information features are matched to a given The other category is people with random identities, short stays in the scene, and no way to collect all prior information. They are highly mobile and random, but they share common characteristics, such as customers in the sales field. As individuals, these people are not statistically significant because their identities are random and unknown. Because information collection methods rely on the scene (surveillance and recordings fixed to the target scene), this group of people's short stays in the scene result in fewer information sources and limited information that can be extracted compared to the first category. However, as a group, these subjects exhibit some common characteristics, which are statistically valid information and the main target. Therefore, for this type of subject, the focus of information extraction is often on the different behaviors of specific people when faced with the same salesperson or product. The goal is to segment the broad concept of "customer" and promote tailored teaching methods. Therefore, the data volume is not blindly pursued to be huge, but rather to be precise and refined, so that the extracted information can be seen from the details. The final matching result is often the characteristics that are abstracted into a whole together with the customer's basic information.
[0229] Taking the three applications of extracting and drawing user portraits in the sales field, analyzing and generating sales staff's work level, and capturing customers' evaluations of products as examples, the multimodal information matching method of the present invention is elaborated in detail, including:
[0230] First, data processing: Our input consists of video and audio files that match the scenario and timeframe of the information to be extracted. This means that the data contains the information we need. For example, in the sales field, the input could be video or audio recordings of a store during business hours. Video recordings of a store closed would not meet our input requirements.
[0231] As shown in Figure 4(a), the input is video data and audio data containing raw information. Here, we use the example of using audio and video recordings of a store to create user profiles and sales staff sales level assessments. After obtaining the data, we perform different preprocessing on the data according to different tasks.
[0232] After the data is prepared, different data processing methods are required for different needs. For example, when it comes to extracting user profiles, the key is accuracy, meaning that customer information must be accurately matched to customer features. A multimodal approach is used to verify the accuracy of the matching. For video, surveillance footage is used to confirm that there is no interference from other customers during the time period of the customer and clerk's conversation. Voiceprint recognition is used in the recording to annotate the customer and clerk's voices, and this annotation is carried over to the speech recognition task, resulting in a textual record of who said what and when. The contextual analysis capabilities of the large language model are then used to further verify the text based on character logic (e.g., statements promoting a product must be spoken by the clerk, and questions about the price of a product must be spoken by the customer). This allows for further self-verification of whether any customer or salesperson characters have been mislabeled due to voice similarities. Using three different modalities (video, audio, and text), the extracted features are guaranteed to be the characteristics of the specific customer, without contamination by information from other customers or clerks. Among them, customer information comes from several aspects. On the one hand, it is the basic information of the customer registered in the store, which is mastered by the sales staff; on the other hand, it is the customer's appearance information such as gender and age group extracted through video recording; on the other hand, it is the basic information that may be revealed in the conversation between the customer and the sales staff, which is extracted by large language model analysis.
[0233] The primary source of customer characteristics is conversation. By fine-tuning the large language model, it analyzes conversations between customers and salespeople, extracting customer characteristics such as dietary preferences, hobbies, and product preferences. Since the corresponding timestamps are retained when processing data in each modality (the timestamps of the recorded data are also inherited during speech-to-text conversion), a one-to-one match between characteristics and appearance information can be achieved by simply matching the same timestamps. The salesperson who receives the customer then fills in the database with basic information based on the customer's appearance, achieving a perfect match between customer information and characteristics. Similarly, since customers from different backgrounds may have different evaluations of different products, matching the "user feedback" reviews with the customers is particularly important to better utilize these evaluations. Therefore, for these two tasks, the "customer" is the core extraction object.
[0234] In the data processing step, the acquired video data needs to be pre-processed. Object detection technology, such as YOLOV5, is used to capture the time points when the salesperson and the customer have a private conversation to ensure the purity of the data source. Object detection technology is used to detect the single-person conversation scene in the store where there is only a clerk and a customer, and the start timestamp is returned. When the customer leaves the store or a second customer enters the store, the end timestamp is returned. The time between the start timestamp and the end timestamp is the pure data of the customer. The audio data corresponding to the time nodes confirmed by the above video data is segmented for subsequent speaker separation and speech recognition. The audio data is then cut according to the silent points (that is, the breath points) using voice activity detection technology (VAD). In this way, the content between each breath point must be a continuous speaking audio segment of one person. The acoustic features of each continuous speech audio segment are extracted (its Mel-frequency cepstral coefficients are calculated), and the acoustic features of each segment are represented by a fixed-dimensional vector using a deep learning model. Speaker separation technology is used to calculate the feature vectors of the voiceprints of sales staff and customers, and the cosine similarity of the embedded vectors is calculated. Two audio segments with higher similarity are regarded as spoken by one person, and vice versa. This is speaker embedding technology. Speaker embedding ensures that vectors for different segments of the same speaker are closer, while those for different speakers are farther apart. This allows for clustering (k-means), distinguishing and annotating the voices of salespeople and customers. Combined with the contextual understanding of a large language model, the speech recognition text is rewritten into a conversational format. Further contextual understanding of the large language model allows for verification and correction based on the conversational content and character logic. This ensures that the extracted information and customer identities are not mismatched, accurately distinguishing between the salesperson's and the customer's words in the conversation, and preventing the salesperson's personal characteristics from contaminating the customer's characteristics, thus lowering the analysis threshold of the large language model. However, the need for "salesperson performance evaluation" does not require matching customer identities with information. Instead, a comprehensive and inductive evaluation requires data from a large number of staff members on different customer experiences. Therefore, for this task, the daily audio data is retained and speech recognition is performed on the entire data set, leveraging the contextual analysis capabilities of the large language model to directly analyze the salesperson's service level.
[0235] After performing speech recognition on the processed segmented and overall audio data, the recognized text becomes the primary source of information extraction. To address the diverse needs of different fields, some words may be commonly used in certain fields but unfamiliar to the general public. To prevent these words from being mistakenly recognized as homophones during speech recognition, a specific hotword library is constructed. This library includes various proper nouns. When the speech recognition model recognizes a syllable identical to a hotword, it outputs the hotword as a whole, rather than guessing at a homophone. This process ensures speech recognition accuracy.
[0236] After speech recognition, the text data content will undergo another step of polishing. Since this step utilizes the function of large language models to analyze context, the polished data will not only filter out meaningless modal particles, delete duplications and fill in omissions, but also correct the same noun in the speech recognition step that is recognized as another noun with the same pronunciation but different meaning in the context, maintain semantic consistency, and further improve recognition accuracy.
[0237] Second, information extraction: After data processing, information extraction is performed on the processed data. Depending on different needs, information can be extracted from the text data after speech recognition based on several aspects, such as user profiles, staff performance evaluations, and user feedback on the product. User needs are provided by the users of the present invention. According to the present invention, users (e.g., product companies or practitioners in the education field) input keywords (such as user profiles, performance evaluations, user feedback, etc.) and specific feature requirements (using user profiles as an example, specific feature requirements can be interests, hobbies, living habits, dietary structure, etc.). Utilizing the powerful analysis and generation capabilities of the large prediction model, corresponding prompt words are generated based on the user's input, which are used to fine-tune the large model in the information extraction step. The large model fine-tuned with different prompt words will analyze the conversation text from different angles and extract information with different focuses. Because prompt words are created based on the user's specific needs and ultimately extract the information the user wants, the direction of information extracted varies depending on the user's identity, field of work, and focus. This is why the present invention can extract diverse information.
[0238] Information extraction can focus on different areas, as exemplified by three directions: User profiling can extract basic information about a user's personality, psychology, interests, preferences, lifestyle, and other aspects of the conversation, as well as information about products of interest to the user. Furthermore, it can capture the topics frequently used by staff members when serving customers, their preferred chat styles, and any instances of insults or verbal abuse during interactions with customers, to assess their performance. Furthermore, it can also extract customer reviews of products, providing feedback to promote product development and improvement. This information is embedded in conversations between staff and customers. After fine-tuning with specific prompts, the large language model will pay special attention to conversations in which this information is mentioned, analyze and integrate it, and generate output.
[0239] After clarifying the requirements, different data processing methods are selected based on the specific needs to obtain data containing the original information. By leveraging the large language model's powerful analytical capabilities, excellent generalization performance, and convenient fine-tuning methods, different prompt words are designed to fine-tune the large language model, allowing it to analyze the data containing the original information and obtain the most relevant feature information. Prompt words for different needs have different focuses. For example, prompt words for user portraits focus on depicting the user's image through conversation analysis; prompt words for product feedback focus on capturing product-related information; and as shown in Figure 4(b), evaluating staff performance focuses on analyzing staff chat skills and assessing their emotions.
[0240] Furthermore, during the user profile creation process, object detection is first performed on the video data to identify time periods where the customer and the employee have private conversations. This ensures that all information features within these time periods are attributed to the customer. Basic customer information (such as age and gender) is also recorded. The audio is segmented according to the time periods provided by the video processing, creating the segmented clean data. Speech recognition and speaker separation are then performed on the clean data. Speech recognition transforms the data into a form that can be analyzed by the large language model, while speaker separation prevents confusion between the employee's information and the customer's information. After text polishing to filter out meaningless modal particles, remove duplicates, and fill in gaps, and ensure consistency in context, the resulting text data contains the information to be extracted. This data then serves as input to the large language model, which has been fine-tuned using prompts specifically for character profile analysis. The large language model's powerful analytical capabilities enable it to output not only concrete information about the conversational content, such as interests, preferences, and lifestyle habits, but also more abstract information, such as personality and psychology, derived from the conversational process. This information is then combined with the previously recorded basic information to form the final customer user profile.
[0241] Third, information output and matching: Finally, information output involves extracting the characteristic information. Feedback tracing is required to identify the time period during which the customer appeared, enabling a one-to-one matching of the customer's characteristics with their basic information. The purpose of information matching is to ensure that the characteristic information extracted through conversation analysis corresponds to the appearance information captured through video capture, ensuring that the characteristic information was indeed provided by the customer. Since timestamps are inherited at each modality transition from video data to audio data to text data, despite the different modalities, timestamps are simply matched at the code level between the timestamps of different modalities, eliminating the need for manual comparison. To achieve this, timestamps are retained at each data processing step, representing the real-time information corresponding to the data period. After extracting the required information features during the information extraction step, identity tracing is performed based on the timestamps retained in the data, matching the basic information of the customer who provided the characteristic information, and establishing a one-to-one matching. This can be combined with basic customer information such as age and gender to create a statistical record of customer preferences, facilitating subsequent product recommendations for customers with similar basic information. Similarly, for the need for "staff work evaluation", the store information and duty information in the original data are also retained, so that the final output work evaluation results can be matched one-to-one with the corresponding staff, forming a complete, grassroots work evaluation report.
[0242] As shown in Figure 4(b), when analyzing a salesperson's performance, basic information and their shift schedule are readily available. Recordings of the salesperson during their shifts are segmented based on their shift schedule, creating the original audio data. Similarly, speech recognition and text polishing are performed to obtain text data about the salesperson's work time. This text data is fed into a large language model, previously fine-tuned with prompts. This model is informed that this is a conversation between a salesperson and a customer during work. The model automatically distinguishes between the salesperson and the customer, and based on the conversation analysis results, outputs an assessment of the salesperson's work process, including information on topics discussed, conversational techniques used, and whether any verbal abuse of customers occurs. This assessment is then combined with the basic information provided by the sales company to create the final salesperson performance assessment report.
[0243] The multimodal information matching method provided in this embodiment captures information from everyday events within the target scene by acquiring video and audio from the target scene, thereby ensuring the objectivity and authenticity of the information source. Furthermore, unlike traditional machine learning methods, this method does not rely on structured data; even information in natural language can be accurately captured. This method saves manpower and material resources while leveraging the powerful analytical and inductive capabilities, efficient processing speed, and lack of subjective influence of large language models to efficiently and objectively capture common features within massive amounts of unstructured information. This method uses a multi-faceted assessment of the person's performance based on their daily work. For example, for salespeople, the assessment is based not only on sales volume but also on the effectiveness of their sales pitch; for teachers, not only on their enrollment rate but also on the effectiveness of their teaching methods. This allows the incorporation of indicators that are difficult to quantify in previous methods but are crucial for assessing performance into the evaluation criteria. Furthermore, based on the evaluation results, it can promote superior pitches and methods, ensuring that the evaluation is value-added. This method can be flexibly applied to fields such as sales, education, and healthcare. For example, in the field of online education, this method can be used to draw student portraits, analyze their learning habits, knowledge mastery, interests, etc. It can also be used to evaluate the teaching quality of teachers, whether their teaching methods are appropriate, whether the interaction is positive, and whether they can effectively answer students' questions.
[0244] In this embodiment, a multimodal information matching device is also provided, which is used to implement the above-mentioned embodiments and preferred embodiments. The details already described will not be repeated here. As used below, the term "module" can refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.
[0245] This embodiment provides a multi-modal information matching device, such as Figure 5 Shown, including:
[0246] The to-be-extracted audio and video file acquisition module 501 is used to acquire the to-be-extracted audio and video files of different task objects in the target scene.
[0247] The multimodal processing and multidimensional feature set extraction module 502 is used to perform multimodal processing on the audio and video files using object detection technology, voiceprint recognition technology, and natural language processing technology, and to extract multidimensional feature sets corresponding to different task objects from the multimodal processed audio and video files;
[0248] A character portrait construction module 503 is used to construct a character portrait based on a multi-dimensional feature set corresponding to different task objects;
[0249] The information matching module 504 is used to match the character portrait to a predetermined person using a feature matching algorithm based on a pre-recorded timestamp.
[0250] In some optional implementations, the audio and video files include video files and audio files; the audio and video file acquisition module 501 to be extracted includes:
[0251] The video and audio acquisition unit is used to acquire a video file with a time stamp and an audio file with a time stamp within a preset time period in the target scene.
[0252] The time synchronization alignment unit is used to align the time stamps of the video file with the time stamp and the audio file with the time stamp by using the synchronization alignment technology to obtain the video file and the audio file with the consistent time stamps.
[0253] In some optional implementations, the natural language processing technology includes speech recognition technology and a large language model; the multimodal processing and multi-dimensional feature set extraction module 502 includes:
[0254] The video modality processing unit is used to perform video modality processing on the video file using object detection technology to determine the target time period for separate conversations between different task objects in the video file.
[0255] The audio modality processing unit is used to perform audio modality processing on the audio file using voiceprint recognition technology, obtain the audio of each task object in the audio file within the target time period, and mark the audio of each task object.
[0256] The text modality processing unit is used to perform text modality processing on the audio of each labeled task object using speech recognition technology to obtain the text content corresponding to each labeled task object audio, and use a large language model to perform context analysis on the text content corresponding to each labeled task object audio, to verify and correct the text content that is incorrectly labeled due to the similar timbre of different task objects.
[0257] In some optional implementations, different task objects include persons with fixed identities and persons with random identities; the multimodal processing and multidimensional feature set extraction module 502 further includes:
[0258] The video file feature extraction unit is used to extract the first appearance feature, first body movement feature and first emotional expression feature of a person with a fixed identity from the video file of the target time period using a large language model, and to extract the second appearance feature, second body movement feature and second emotional expression feature of a person with a random identity.
[0259] The audio file feature extraction unit is used to extract the first language expression features and chat skill features of a fixed-identity person from the text content corresponding to the audio of each labeled task object based on preset prompt words using a large language model, and to extract the second language expression features, feedback features and satisfaction scores of a random-identity person who matches the preset prompt words.
[0260] The multidimensional feature set construction unit is used to use the first appearance feature, the first body movement feature, the first emotion expression feature, the first language expression feature, the chat skill feature, the feedback feature and the satisfaction score as the first multidimensional feature set; and use the second appearance feature, the second body movement feature, the second emotion expression feature, the second language expression feature, the feedback feature and the satisfaction score as the second multidimensional feature set.
[0261] In some optional implementations, the character portrait construction module 503 includes:
[0262] The appearance extraction unit is used to obtain the appearance information of each person with a random identity and the appearance information of each person with a fixed identity from the video files within the target time period.
[0263] The character portrait construction unit of the random identity person is used to construct the character portrait of the random identity person based on a preset timestamp and a second multi-dimensional feature set.
[0264] The character portrait construction unit of the person with a fixed identity is used to construct the character portrait of the person with a fixed identity based on a preset timestamp and a first multi-dimensional feature set.
[0265] In some optional implementations, the information matching module 504 includes:
[0266] The random person matching unit with established identity is used to trace the source based on a preset timestamp, and at the same time uses a feature comparison algorithm to compare the character portrait of the random person with the appearance information of each random person with established identity to obtain a matched random person with established identity.
[0267] The fixed-identity person matching unit is used to obtain the prior information of the fixed-identity person, trace the source based on the preset timestamp, and use the feature matching algorithm to compare the character portrait of the fixed-identity person with the prior information of each fixed-identity person to obtain the matched fixed-identity person.
[0268] In some optional implementations, the multimodal information matching device further includes:
[0269] The work evaluation module is used to evaluate the work level of a person with a fixed identity based on the first multi-dimensional feature set to obtain a work evaluation result.
[0270] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.
[0271] The multimodal information matching device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.
[0272] The embodiment of the present invention also provides a computer device having the above Figure 5 The multimodal information matching device shown.
[0273] See also Figure 6 , Figure 6 is a structural diagram of a computer device provided by an optional embodiment of the present invention, such as Figure 6 As shown, the computer device includes: one or more processors 10, memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components utilize different buses to communicate with each other and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in the memory or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Equally, multiple computer devices can be connected, and each device provides part of the necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 6 A processor 10 is taken as an example.
[0274] The processor 10 may be a central processing unit, a network processor, or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic, or any combination thereof.
[0275] The memory 20 stores instructions that can be executed by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.
[0276] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created based on the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0277] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid-state drive; the memory 20 may also include a combination of the above types of memory.
[0278] The computer device further includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30 and the output device 40 may be connected via a bus or other means. Figure 6 The bus connection is taken as an example.
[0279] The input device 30 can receive input digital or character information and generate key signal input related to user settings and function control of the computer device, such as a touch screen, a keypad, a mouse, a trackpad, a touch pad, an indicator stick, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 can include a display device, an auxiliary lighting device (e.g., an LED), and a tactile feedback device (e.g., a vibration motor). The above-mentioned display device includes but is not limited to a liquid crystal display, a light emitting diode, a display, and a plasma display. In some optional embodiments, the display device can be a touch screen.
[0280] The embodiment of the present invention also provides a computer-readable storage medium. The above-mentioned method according to the embodiment of the present invention can be implemented in hardware, firmware, or implemented as a computer code that can be recorded in a storage medium, or implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the method shown in the above embodiment is implemented.
[0281] A portion of the present invention may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. Those skilled in the art should understand that the form in which the computer program instruction exists in a computer-readable medium includes, but is not limited to, a source file, an executable file, an installation package file, etc. Accordingly, the way in which the computer program instruction is executed by the computer includes, but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium may be any available computer-readable storage medium or communication medium that can be accessed by the computer.
[0282] Although the embodiments of the present invention have been described with reference to the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present invention. Such modifications and variations are all within the scope defined by the appended claims.
Claims
1. A multimodal information matching method, characterized in that: The method comprises: Obtain the audio and video files to be extracted for different task objects in the target scene; Performing multimodal processing on the audio and video files using object detection technology, voiceprint recognition technology, and natural language processing technology, and extracting multidimensional feature sets corresponding to different task objects from the multimodal processed audio and video files; Constructing a character portrait based on the multi-dimensional feature set corresponding to the different task objects; The character portrait is matched to a predetermined person using a feature comparison algorithm based on a pre-recorded timestamp.
2. The method according to claim 1, characterized in that The audio and video files include video files and audio files; the step of obtaining the audio and video files to be extracted for different task objects in the target scene includes: Obtain a video file and an audio file with a timestamp for a preset time period within the target scene; The synchronous alignment technology is used to align the timestamps of the video files and the audio files with timestamps to obtain video files and audio files with consistent timestamps.
3. The method according to claim 2, characterized in that The natural language processing technology includes speech recognition technology and large language models; The multimodal processing of the audio and video files using object detection technology, voiceprint recognition technology, and natural language processing technology includes: Performing video modality processing on the video file using object detection technology to determine target time periods for individual conversations between different task objects in the video file; Using voiceprint recognition technology to perform audio modal processing on the audio file, obtain each task object audio in the audio file within the target time period, and mark each task object audio; Speech recognition technology is used to perform text modal processing on the audio of each labeled task object to obtain the text content corresponding to each labeled task object audio. A large language model is used to perform contextual analysis on the text content corresponding to each labeled task object audio to verify and correct the text content that is incorrectly labeled due to the similar timbre of different task objects.
4. The method according to claim 3, characterized in that The different task objects include personnel with fixed identities and personnel with random identities; The step of extracting multi-dimensional feature sets corresponding to different task objects from the multimodally processed audio and video files includes: A large language model is used to extract the first appearance feature, first body movement feature, and first emotional expression feature of a person with a fixed identity from the video file of the target time period, and to extract the second appearance feature, second body movement feature, and second emotional expression feature of a person with a random identity; Based on preset prompt words, a large language model is used to extract the first language expression features and chat skill features of fixed-identity people from the text content corresponding to the audio of each labeled task object, as well as the second language expression features, feedback features, and satisfaction scores of random-identity people matching the preset prompt words. The first appearance feature, the first body movement feature, the first emotion expression feature, the first language expression feature, the chat skill feature, the feedback feature, and the satisfaction score are used as a first multidimensional feature set; The second appearance feature, the second body movement feature, the second emotion expression feature, the second language expression feature, the feedback feature and the satisfaction score are used as a second multidimensional feature set.
5. The method according to claim 4, characterized in that Constructing a character portrait based on the multi-dimensional feature sets corresponding to the different task objects includes: Obtain the appearance information of each random person and each fixed person from the video files within the target time period; Constructing a character portrait of a random person based on a preset timestamp and a second multi-dimensional feature set; A character portrait of a person with a fixed identity is constructed based on a preset timestamp and a first multi-dimensional feature set.
6. The method according to claim 4, characterized in that The method of matching the character portrait to a predetermined person using a feature comparison algorithm based on a pre-recorded timestamp includes: Tracing the source is performed based on a preset timestamp, and a feature comparison algorithm is used to compare the portrait of the random person with the appearance information of each random person to obtain the matching random person with the given identity; Obtain the prior information of the fixed-identity person, trace the source based on the preset timestamp, and use the feature comparison algorithm to compare the character portrait of the fixed-identity person with the prior information of each fixed-identity person to obtain the matching fixed-identity person.
7. The method according to claim 6, characterized in that The method further comprises: The work level of a person with a fixed identity is evaluated based on the first multi-dimensional feature set to obtain a work evaluation result.
8. A multimodal information matching device, characterized in that: The device comprises: The module for obtaining audio and video files to be extracted is used to obtain audio and video files to be extracted for different task objects in the target scene; A multimodal processing and multidimensional feature set extraction module, configured to perform multimodal processing on the audio and video files using object detection technology, voiceprint recognition technology, and natural language processing technology, and to extract multidimensional feature sets corresponding to different task objects from the multimodal processed audio and video files; A character portrait construction module, configured to construct a character portrait based on the multi-dimensional feature sets corresponding to the different task objects; The information matching module is used to match the character portrait to a predetermined person using a feature comparison algorithm based on a pre-recorded timestamp.
9. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the multimodal information matching method according to any one of claims 1 to 7 by executing the computer instructions.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the multimodal information matching method according to any one of claims 1 to 7.
Citation Information
Cited By
Multi-channel call recording recognition method and device based on single-channel artificial intelligence model
CN120895028A
Enterprise culture hot word analysis method based on content tracing
CN121435967A
Enterprise culture hot word analysis method based on content traceability
CN121435967B
Teacher ability diagnosis and evaluation method and system based on multi-modal data fusion
CN122066312A
A teacher ability diagnosis and evaluation method and system based on multi-modal data fusion
CN122066312B