Conversation-based multi-modal feature analysis method and electronic device

By using multimodal feature analysis, speech, facial expression, and text data are collected and labeled. Combined with historical data, language state influencing factors are calculated, which solves the subjectivity and inconsistency problems of traditional assessment methods and realizes scientific assessment of language state and quantitative guidance for intervention strategies.

CN121144772BActive Publication Date: 2026-02-27浙江连信数字有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511698131.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-19
Publication Date
2026-02-27
Estimated Expiration
2045-11-19

AI Technical Summary

Technical Problem

Traditional language assessment methods rely on human judgment, are easily influenced by subjective factors, lack unified quantitative standards, are difficult to fully capture language state characteristics, and cannot effectively explore the distribution patterns of language states in historical assessment data, resulting in poor reliability and consistency of assessment results.

Method used

This study employs a conversation-based multimodal feature analysis method. By collecting voice, facial expression, and text response data, marking state feature points, statistically analyzing feature frequency and intensity, and combining historical data to calculate language state influence factors, a scientific assessment of the current language state is achieved.

Benefits of technology

It enables a comprehensive and detailed capture of language status, improves the scientific rigor and reliability of the assessment, provides a quantitative assessment of the state of health, and offers clear guidance for intervention strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121144772B_ABST
    Figure CN121144772B_ABST
Patent Text Reader

Abstract

The application discloses a talk-based multi-modal feature analysis method and electronic equipment, and relates to the technical field of language analysis, and the technical solution points thereof comprise the following steps: collecting multi-modal evaluation data of a talk process of an object to be evaluated, performing feature marking on the multi-modal evaluation data to obtain state feature points; counting the occurrence frequency and feature intensity of the same features in the state feature points to obtain feature correlation parameters; extracting historical language state feature distribution rules in different evaluation scenes from a historical evaluation data set, calculating a language state influence factor between a language state tendency value in the historical evaluation data and the feature correlation parameters according to the historical language state feature distribution rules; and screening a target influence factor from the language state influence factor according to the real-time language state feature distribution characteristics of the current talk scene, so that clear guidance is provided for the subsequent development of intervention strategies.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of language analysis technology, and more specifically, to a conversation-based multimodal feature analysis method and electronic device. Background Technology

[0002] Traditional language assessment methods have many limitations. Past assessments relied heavily on human interviews, judging the subject's state based on the content of their speech and observed facial expressions and gestures. However, human assessment is susceptible to subjective factors; different assessors may interpret the same performance differently, and the dimensions of information collected and analyzed manually are limited, making it difficult to comprehensively capture the subject's language state characteristics. Traditional methods are time-consuming and labor-intensive when handling large-scale assessment tasks, and the lack of standardized quantitative assessment criteria makes it difficult to guarantee the consistency and reliability of assessment results. Furthermore, the complex relationships between language state characteristics make it difficult to effectively uncover patterns in the distribution of language states in historical assessment data, thus failing to provide valuable guidance for current assessments. Summary of the Invention

[0003] In view of the shortcomings of the existing technology, the purpose of this invention is to provide a method and electronic device for multimodal feature analysis based on conversation.

[0004] To achieve the above objectives, the present invention provides the following technical solution:

[0005] A conversation-based multimodal feature analysis method, comprising the following steps:

[0006] Collect multimodal evaluation data of the conversation process of the subject to be evaluated, and perform feature labeling on the multimodal evaluation data to obtain state feature points;

[0007] Feature association parameters are obtained by analyzing the frequency and intensity of similar features among statistical state feature points.

[0008] Extract the distribution patterns of historical language state features under different assessment scenarios from the historical assessment dataset, and calculate the language state influence factor between the language state tendency value and the feature association parameter in the historical assessment data based on the distribution patterns of historical language state features.

[0009] Based on the real-time language state feature distribution characteristics of the current conversation scenario, target influencing factors are selected from the language state influencing factors. Based on the target influencing factors, the core language state tendency value and auxiliary language state tendency value of the subject to be evaluated in the conversation process are determined, and the first language state evaluation value and the second language state evaluation value are output.

[0010] The current status assessment value is obtained by judging the health status value of the assessed object based on the first language status assessment value and the second language status assessment value;

[0011] determining an intervention strategy for the object to be evaluated according to the current situation evaluation value.

[0012] Preferably, the multi-modal evaluation data comprises voice data, facial expression data and text response data.

[0013] Preferably, the multi-modal evaluation data is feature-labeled to obtain state feature points, specifically comprising the following steps:

[0014] The voice data is tonal and speech rate feature-labeled to obtain language state features;

[0015] The facial expression data is expression action feature-labeled to obtain expression language state features;

[0016] The text response data is semantic language state feature-labeled to obtain text response language state features,

[0017] The language state features, expression language state features and text response language state features constitute the state feature points.

[0018] Preferably, the occurrence frequency and feature intensity of the same type of features in the state feature points are counted to obtain feature correlation parameters, specifically comprising the following steps:

[0019] The occurrence frequency of the same type of features in the state feature points is counted within a preset time window to obtain feature occurrence frequency;

[0020] The language state expression degree of the state feature points is calculated to obtain feature intensity values;

[0021] The feature occurrence frequency and the feature intensity values are associated and mapped to obtain feature correlation parameters.

[0022] Preferably, a language state influence factor between the language state tendency values in the historical evaluation data and the feature correlation parameters is calculated according to the historical language state feature distribution law, specifically comprising the following steps:

[0023] According to the historical language state feature distribution law, the historical language state tendency values corresponding to the language state tendency labels are extracted from the historical evaluation data;

[0024] According to the feature correlation parameters, the language state tendency differences of the same type of features in the historical language state tendency values are calculated to obtain language state difference values;

[0025] The language state difference values and the feature correlation parameters are proportionally calculated to obtain feature proportion values;

[0026] All feature proportion values are weighted and summed to obtain the language state influence factor.

[0027] Preferably, the target influence factor is screened from the language state influence factor according to the real-time language state feature distribution characteristics of the current conversation scene, and specifically includes the following steps:

[0028] The real-time language state feature distribution characteristics of the current conversation scene are obtained.

[0029] The target influence factor is screened from the language state influence factor after the real-time language state feature distribution characteristics are matched with the historical language state feature distribution characteristics in terms of similarity.

[0030] Preferably, the first language state evaluation value and the second language state evaluation value are output after the core language state tendency value and the auxiliary language state tendency value of the to-be-evaluated object in the conversation process are judged according to the target influence factor, and specifically includes the following steps:

[0031] The core language state expression dimension and the auxiliary language state expression dimension of the to-be-evaluated object in the conversation process are determined.

[0032] The first feature set corresponding to the core language state and the second feature set corresponding to the auxiliary language state are determined.

[0033] The first language state evaluation value is output after the core language state tendency value is judged according to the feature association parameter corresponding to the first feature set, the real-time language state intensity value and the target influence factor.

[0034] The second language state evaluation value is output after the auxiliary language state tendency value is judged according to the feature association parameter corresponding to the second feature set, the real-time language state intensity value and the target influence factor.

[0035] Preferably, the state health degree value of the evaluation object is judged to obtain the current condition evaluation value according to the first language state evaluation value and the second language state evaluation value, and specifically includes the following steps:

[0036] The first language state evaluation value is given a core weight value.

[0037] The second language state evaluation value is given an auxiliary weight value.

[0038] The first evaluation factor is obtained by multiplying the first language state evaluation value by the core weight value.

[0039] The second evaluation factor is obtained by multiplying the second language state evaluation value by the auxiliary weight value.

[0040] The comprehensive state evaluation factor is obtained by accumulating the first evaluation factor and the second evaluation factor.

[0041] A reference evaluation database of corresponding state health values under different language state tendency combinations and feature intensities is obtained.

[0042] The current status assessment value is obtained by matching the comprehensive status assessment factors with the reference assessment database.

[0043] Preferably, determining the intervention strategy for the object to be evaluated based on the current status assessment value specifically includes the following steps:

[0044] If the current status assessment value is lower than the status warning threshold, an intervention prompt message will be output;

[0045] After calling the preset intervention plan library based on the intervention prompt information, candidate intervention plans that match the current language state feature distribution features are extracted;

[0046] Evaluation data is obtained by assessing the implementation complexity and expected outcome of candidate intervention programs.

[0047] The best-fit intervention strategy is one that extracts evaluation data from candidate intervention programs and finds that outperforms the pre-defined intervention standards.

[0048] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the talk-based multimodal feature analysis method.

[0049] Compared with the prior art, the present invention has the following beneficial effects:

[0050] This invention, by collecting multimodal assessment data and marking state feature points, can comprehensively and meticulously capture the language state expressions of the subject being assessed during conversation. Compared to single-dimensional data collection, multimodal data fusion allows for a richer and more comprehensive mining of language state features, avoiding assessment biases caused by missing information and laying a solid foundation for subsequent accurate analysis. In the feature association and pattern mining stage, the frequency and intensity of similar features are statistically analyzed to obtain feature association parameters. These parameters are then combined with the language state feature distribution patterns extracted from historical data to calculate the language state influence factor, achieving a deep integration of historical experience and current assessment. This integration allows the assessment to move beyond the immediate characteristics of the current conversation, instead leveraging the language state and feature association patterns accumulated in historical data to provide a more valuable basis for judging the current language state tendency, improving the scientific rigor and reliability of the assessment. This application can quantify the health status of the subject being assessed, providing clear guidance for the formulation of subsequent intervention strategies. Attached Figure Description

[0051] Fig. 1 This is a schematic diagram illustrating the steps of the conversation-based multimodal feature analysis method proposed in this invention;

[0052] Fig. 2 This is a schematic diagram illustrating the steps for calculating feature correlation parameters in the conversation-based multimodal feature analysis method proposed in this invention;

[0053] Fig. 3 Figure 1 is a structural schematic diagram of an electronic device provided by an embodiment of the present application.

[0054] 610, processor; 620, communication interface; 630, memory; 640, communication bus. DETAILED DESCRIPTION

[0055] In order to make the above objectives, features and advantages of the present application more apparent, a detailed description of the specific embodiments of the present application will be given below with reference to the accompanying drawings.

[0056] In the following description, a large number of specific details are set forth in order to facilitate a thorough understanding of the present application, but the present application can also be implemented in other ways different from the description, and those skilled in the art can make similar generalizations without departing from the concept of the present application, so the present application is not limited to the specific embodiments disclosed below.

[0057] Secondly, the "one embodiment" or "embodiment" referred to herein means that the specific features, structures or characteristics can be included in at least one implementation of the present application. "In one embodiment" appearing in different places in the specification does not mean the same embodiment, nor is it an embodiment that is independent of or mutually exclusive with other embodiments.

[0058] Referring to Figs. 1-3 as shown.

[0059] The embodiments further illustrate the talk-based multi-modal feature analysis method and the electronic device proposed by the present application.

[0060] The talk-based multi-modal feature analysis method comprises the following steps:

[0061] Collecting multi-modal evaluation data of a talk process of an object to be evaluated, and performing feature labeling on the multi-modal evaluation data to obtain state feature points;

[0062] Counting the occurrence frequency and feature intensity of the same type of features in the state feature points to obtain feature correlation parameters;

[0063] Extracting historical language state feature distribution rules in different evaluation scenarios from a historical evaluation data set, and calculating a language state influence factor between a language state tendency value and the feature correlation parameters in the historical evaluation data according to the historical language state feature distribution rules;

[0064] According to the real-time language state feature distribution characteristics of the current talk scenario, the target influence factor is selected from the language state influence factor, and after judging the core language state tendency value and the auxiliary language state tendency value of the object to be evaluated in the talk process according to the target influence factor, a first language state evaluation value and a second language state evaluation value are output;

[0065] determining a current situation evaluation value according to the first language state evaluation value and the second language state evaluation value;

[0066] determining an intervention strategy for the to-be-evaluated object according to the current situation evaluation value.

[0067] First, multi-modal evaluation data of the to-be-evaluated object in the conversation process is collected, such as speech data, facial expression data and text response data, and state feature points are obtained by feature marking on the modal evaluation data. Specifically, the language state features are obtained by marking the tone and speed features of the speech data, the expression language state features are obtained by marking the expression action features of the facial expression data, and the text response language state features are obtained by marking the semantic language state features of the text response data. The three types of features together constitute the state feature points.

[0068] The occurrence frequency and feature intensity of the same type of features in the state feature points are counted to obtain the feature correlation parameter. In specific operation, the occurrence frequency of the same type of features is counted in a preset time window, the feature intensity value is calculated by calculating the language state expression degree, and then the feature correlation parameter is obtained by correlating and mapping the two.

[0069] The historical language state feature distribution rules of different evaluation scenarios are extracted from the historical evaluation data set, and the language state influence factor between the language state tendency value and the feature correlation parameter in the historical evaluation data is calculated according to these rules. That is, the historical language state tendency value corresponding to the language state tendency label is extracted from the historical data, the language state tendency difference of the same type of features in the historical language state tendency value is calculated by using the feature correlation parameter to obtain the language state difference value, the feature proportion value is obtained by proportionally calculating the language state difference value and the feature correlation parameter, and finally the language state influence factor is obtained by weighted sum of all feature proportion values.

[0070] A real-time scene recognition module is introduced in the data collection layer. The keywords of the conversation content and the environmental audio are extracted, and then the evaluation stage is automatically divided. The speech collection in the initial stage adopts a high sampling rate of 16 kHz to capture subtle tremors, and only the eye movement trajectory is extracted for facial expression; In the core evaluation stage, facial micro-expression capture is activated, and the associated muscle groups of the zygomatic major muscle and the corrugator supercilii muscle are monitored, and the weight of the fundamental frequency standard deviation in the speech feature is increased by 40%.

[0071] Dynamic weight adjustment is designed in the feature processing layer. The information gain of each modality under the current label is calculated: when the language tendency is recognized, the speech feature weight is automatically adjusted to 50%, the text feature accounts for 30%, and the facial feature accounts for 20%.

[0072] A feature fusion confidence score mechanism is constructed in the result output layer. When the credibility is lower than the threshold, the supplementary collection process is automatically started to ensure that the evaluation result remains high in dynamic adjustment.

[0073] The target influence factor is selected from the language state influence factor according to the real-time language state feature distribution characteristics of the current conversation scene. First, the real-time language state feature distribution characteristics of the current scene are obtained, which are matched with the historical language state feature distribution characteristics, and then the target influence factor is selected. The core language state tendency value and the auxiliary language state tendency value in the conversation process of the to-be-evaluated object are determined by using the target influence factor, and the core and auxiliary language state expression dimensions, and the corresponding first feature set and second feature set are determined. The core language state tendency value and the auxiliary language state tendency value are calculated by combining the feature association parameter, the real-time language state intensity value and the target influence factor, and then the first language state evaluation value and the second language state evaluation value are obtained.

[0074] The conversation scene is subdivided into core topics, and each type of scene corresponds to a dedicated feature association model. For example, in the family conflict scene, through the training of more than 100,000 samples, it is found that a specific feature combination has a high correlation with the real language tendency; in the professional stress scene, the recognition accuracy of the specific feature combination to the language tendency is improved.

[0075] A personalized feature time series library is constructed by longitudinal tracking to record the changes in feature association from the first evaluation to the Nth follow-up. For example, different stages of features are automatically marked for language tendency evaluation; the language inertia index is introduced to calculate the decay rate of the same feature association in continuous evaluation, providing a quantitative basis for evaluation.

[0076] The current condition evaluation value is obtained by judging the state health degree value of the evaluation object according to the first language state evaluation value and the second language state evaluation value. The first language state evaluation value and the second language state evaluation value are respectively given core and auxiliary weight values, and finally a comprehensive state evaluation factor is obtained by calculation, which is matched with the reference evaluation database to obtain the current condition evaluation value.

[0077] The intervention strategy is determined according to the current condition evaluation value. If the evaluation value is lower than the condition warning threshold, a candidate intervention strategy is extracted from the preset intervention scheme library, and the implementation complexity, expected effect value and other scheme evaluation data of the candidate intervention strategy are evaluated to select an adaptive intervention strategy whose scheme evaluation data is better than the preset standard, which is used to intervene the evaluation object.

[0078] The multi-modal evaluation data includes speech data, facial expression data and text response data.

[0079] The state feature points are obtained by marking the features of the multi-modal evaluation data, including the following steps:

[0080] The language state features are obtained by marking the tone and speed features of the speech data;

[0081] The facial expression data is marked with expression action features to obtain expression language state features;

[0082] The text response data is marked with semantic language state features to obtain text response language state features,

[0083] The language state features, the expression language state features and the text response language state features constitute state feature points.

[0084] The voice data records speaking tone and speaking speed, such as excited speaking speed and tone. The facial expression data captures frown and smile expression actions, such as easy frown. The text response data extracts language content, such as the text response "I am very sad" with language inclination. By marking the three types of data with features, the language state features contained in the tone and speed are mined from the voice, the expression action corresponding expression language state features are extracted from the facial expression, and the semantic associated text response language state features are analyzed from the text response, which together constitute the state feature points.

[0085] The feature association parameters are obtained by counting the occurrence frequency and feature intensity of the same type of features in the state feature points, including the following steps:

[0086] The feature occurrence frequency is counted by counting the occurrence frequency of the same type of features in the state feature points in a preset time window;

[0087] The feature intensity value is calculated by calculating the language state expression degree of the state feature points;

[0088] The feature association parameters are obtained by associating and mapping the feature occurrence frequency and the feature intensity value.

[0089] The application first focuses on the state feature points, and the same type of features in the state feature points is counted in a pre-set time window. The pre-set time window can be a selected 5-minute continuous conversation process of the to-be-evaluated object as the time window. For example, if "fast speaking speed and high tone" is classified as the same type of excited feature, the number of times of occurrence of the same type of feature in the 5-minute conversation is counted to obtain the feature occurrence frequency, which directly reflects the activity degree of the same type of language state feature in a specific period.

[0090] The language state expression degree of the state feature points is calculated. Not only the number of occurrences is concerned, but also the expression intensity is considered. For example, the same "fast speaking speed and high tone" can be slightly faster and slightly higher, or sharply faster and significantly higher. The feature intensity value is obtained by quantifying the expression degree according to the voice waveform change and expression action amplitude index, which reflects the inherent intensity of the language state feature.

[0091] The feature occurrence frequency is associated with the feature intensity value. It can be understood that a corresponding relationship is established, and the two work together to generate a feature association parameter.

[0092] According to the historical language state feature distribution law, the language state influence factor between the language state tendency value in the historical evaluation data and the feature association parameter is calculated, specifically including the following steps:

[0093] According to the historical language state feature distribution law, the historical language state tendency value corresponding to the language state tendency label is extracted from the historical evaluation data;

[0094] According to the feature association parameter, the language state tendency difference of the same type of feature corresponding to the historical language state tendency value is calculated to obtain a language state difference value;

[0095] The language state difference value is proportionally calculated with the feature association parameter to obtain a feature proportion value;

[0096] All feature proportion values are weighted and summed to obtain a language state influence factor.

[0097] First, the historical language state tendency value corresponding to the language state tendency label is extracted from the historical evaluation data by means of the historical language state feature distribution law. For example, in the related evaluation data in the tendency evaluation scene, the historical language state tendency values corresponding to the state labels in different conversation scenes are filtered out. Different conversation scenes can be job stress conversation and social conversation.

[0098] The language state tendency difference of the same type of feature in the historical language state tendency value is determined by using the feature association parameter. Taking the same type of feature "fast speech" as an example, the language state tendency corresponding to "fast speech" may have slight differences in different historical evaluation data, and some are more biased towards excitement. The language state difference value is calculated by combining the feature association parameter, which contains the occurrence frequency and intensity, and presents the language state tendency fluctuation of the same type of feature in the historical data.

[0099] The language state difference value is proportionally calculated with the feature association parameter to obtain a feature proportion value. It is assumed that the language state difference value reflects the tendency dispersion degree of the same type of feature, and the feature association parameter reflects the comprehensive performance of the feature. The proportion of the two excavates the internal correlation between the feature and the language state tendency difference. For example, a certain type of feature "pitch rise", the language state difference value is 0.6, the feature association parameter is 1.2, and the feature proportion value is 0.5 after proportional calculation, which quantifies the contribution degree of the feature in the language state tendency difference.

[0100] The language state influence factors are generated by weighted sum of all feature proportion values. Different features have different influences on language state tendency, and the feature proportion values of various types are integrated by assigning reasonable weights. For example, the weight of speech speed feature is 0.3, and the corresponding proportion value is 0.5; the weight of expression feature is 0.4, and the corresponding proportion value is 0.6. The language state influence factor is obtained by weighted sum, which can effectively associate the historical data features with the language state tendency relationship and provide historical reference basis for language state tendency judgment.

[0101] The target influence factor is selected from the language state influence factor according to the real-time language state feature distribution characteristics of the current conversation scene, specifically including the following steps:

[0102] The real-time language state feature distribution characteristics of the current conversation scene are obtained.

[0103] The target influence factor is selected from the language state influence factor after similarity matching of the real-time language state feature distribution characteristics and the historical language state feature distribution characteristics.

[0104] First, the real-time language state feature distribution characteristics of the current conversation scene are obtained. Multimodal data is collected during the conversation of the to-be-evaluated object and processed to obtain state feature points. After statistical association operation, the distribution of various language state features in the current scene is presented, including the proportion and combination of different language state features.

[0105] The real-time language state feature distribution characteristics are similarity matched with the historical language state feature distribution characteristics. The historical evaluation data set stores the language state feature distribution rules in different past scenes, such as the feature distribution of "fast speech speed, negative vocabulary, and high frequency of frowning" in a job stress scene. The language state feature distribution of the current scene is compared with the feature distribution of these historical scenes one by one to calculate the similarity degree, and the scene or data segment in history that best fits the current language state feature mode is found.

[0106] The language state influence factor is calculated based on historical data, which reflects the correlation between language state features and language state tendency. After finding the historical feature distribution highly similar to the current scene, the language state influence factor in the corresponding historical scene is extracted as the target influence factor, which allows the evaluation process to fully draw on historical experience and improve the accuracy and scientificity of language state analysis in the current conversation.

[0107] After determining the core language state tendency value and the auxiliary language state tendency value of the to-be-evaluated object in the conversation process according to the target influence factor, the first language state evaluation value and the second language state evaluation value are output, specifically including the following steps:

[0108] The core language state expression dimension and the auxiliary language state expression dimension of the to-be-evaluated object in the conversation process are determined.

[0109] determining a first feature set corresponding to the core language state and a second feature set corresponding to the auxiliary language state;

[0110] judging a core language state tendency value according to a feature correlation parameter corresponding to the first feature set, a real-time language state intensity value and a target influence factor, and outputting a first language state evaluation value;

[0111] judging an auxiliary language state tendency value according to a feature correlation parameter corresponding to the second feature set, a real-time language state intensity value and a target influence factor, and outputting a second language state evaluation value.

[0112] The application first determines the core and auxiliary language state expression dimensions. In the conversation process of the to-be-evaluated object, the core language state expression dimension that has a more critical impact on the state and the auxiliary language state expression dimension that plays a supplementary role are determined.

[0113] The first feature set is composed of features closely related to the core language state. Taking the core language state with a negative language tendency as an example, the first feature set corresponding thereto includes features such as “fast speech speed, rising tone, frown expression, negative semantic vocabulary”, so as to establish a corresponding relationship between the language state and the specific features.

[0114] The core language state tendency value is comprehensively judged according to the feature correlation parameter corresponding to the first feature set, the real-time language state intensity value and the target influence factor. The first language state evaluation value is output by quantitatively calculating the interaction of these factors, so as to present the tendency degree of the core language state. Similarly, the auxiliary language state tendency value is judged according to the feature correlation parameter corresponding to the second feature set, the real-time language state intensity value and the target influence factor, and the second language state evaluation value is output, thereby providing a key language state quantitative basis for subsequent state health degree judgment.

[0115] judging the state health degree value of the to-be-evaluated object according to the first language state evaluation value and the second language state evaluation value to obtain a current condition evaluation value, specifically including the following steps:

[0116] assigning a core weight value to the first language state evaluation value;

[0117] assigning an auxiliary weight value to the second language state evaluation value;

[0118] multiplying the first language state evaluation value by the core weight value to obtain a first evaluation factor;

[0119] multiplying the second language state evaluation value by the auxiliary weight value to obtain a second evaluation factor;

[0120] The first evaluation factor and the second evaluation factor are accumulated to obtain a comprehensive state evaluation factor;

[0121] A reference evaluation database of state health values corresponding to different language state tendency combinations and feature intensities is obtained;

[0122] The comprehensive state evaluation factor is matched with the reference evaluation database to obtain a current condition evaluation value.

[0123] The first language state evaluation value reflects the greater influence of the core language state tendency on the state judgment, so a core weight value is assigned; the second language state evaluation value corresponds to the relatively secondary influence of the auxiliary language state tendency, so an auxiliary weight value is assigned. For example, the core weight is set to 0.7, and the auxiliary weight is set to 0.3.

[0124] The first language state evaluation value is multiplied by the core weight value to obtain the first evaluation factor, which quantifies the influence of the core language state on the state; the second language state evaluation value is multiplied by the auxiliary weight value to obtain the second evaluation factor, which reflects the effect of the auxiliary language state. Assuming that the first language state evaluation value is 8 points, the core weight is 0.7, and the first evaluation factor is 8*0.7=5.6; the second language state evaluation value is 6 points, the auxiliary weight is 0.3, and the second evaluation factor is 6*0.3=1.8.

[0125] The first evaluation factor and the second evaluation factor are accumulated to obtain a comprehensive state evaluation factor, i.e. 5.6+1.8=7.4, which integrates the influence of the core and auxiliary language states and comprehensively reflects the effect of the current language state tendency on the state.

[0126] First, a reference evaluation database is constructed, and state health values corresponding to different language state tendency combinations and feature intensities are collected.

[0127] The calculated comprehensive state evaluation factor is matched with the data in the reference evaluation database. The most suitable language state combination and health value corresponding record in the database is searched to determine the state health degree of the current evaluation object, and then a current condition evaluation value is output, which provides a key quantitative basis for subsequent judgment and intervention strategy.

[0128] According to the current condition evaluation value, an intervention strategy for the evaluation object is determined, which specifically includes the following steps:

[0129] If the current condition evaluation value is lower than the condition warning threshold, an intervention prompt information is output;

[0130] According to the intervention prompt information, a preset intervention scheme library is called to extract a candidate intervention scheme matching the current language state feature distribution feature;

[0131] The implementation complexity and expected effect value of the candidate intervention scheme are evaluated to obtain scheme evaluation data;

[0132] An adaptive intervention strategy is extracted from the candidate intervention scheme whose scheme evaluation data is superior to the preset scheme standard.

[0133] The current condition evaluation value is compared with the preset condition warning threshold. If the current condition evaluation value is lower than the threshold, it means that the evaluation object state has a more serious problem and intervention is needed. At this time, an intervention prompt information is output. For example, the condition warning threshold is set to 60 points (assuming a full score of 100 points), and the current evaluation object evaluation value is 55 points, that is, the current condition evaluation value is lower than the condition warning threshold, which triggers the generation of prompt information after the intervention process.

[0134] According to the intervention prompt information, a preset intervention scheme library is called. The scheme library stores a variety of intervention schemes, including content adapted to different language state feature distribution scenarios. According to the current language state feature distribution characteristics, candidate intervention schemes matching the characteristics are extracted from the scheme library.

[0135] Multi-dimensional evaluation is carried out on the candidate intervention scheme. The evaluation process focuses on two core indicators: implementation complexity and expected effect value. Implementation complexity evaluation needs to consider the implementation scene of intervention, such as online remote execution or offline face-to-face execution, and the required resources, such as the qualification requirements of the execution personnel, the type of equipment support, and the length of the implementation cycle. For example, a scheme that requires professional personnel to participate 3 times a week has a higher implementation complexity than a scheme that is completed independently online 1 time a week. The expected effect value is quantified according to historical implementation data to quantify the state improvement rate of evaluation objects with similar language state characteristics. For example, the evaluation value of a similar characteristic object is improved by 30% by a certain scheme, and the improvement rate of another scheme is 15%. Through the quantitative scoring of these two indicators, complete scheme evaluation data is formed to provide an objective basis for subsequent screening.

[0136] First, a preset scheme standard is set, which includes specific quantitative indicators such as an upper limit of implementation complexity, such as no need for professional personnel to follow throughout, and a lower limit of expected effect value, such as a state evaluation value improvement rate of not less than 20%. The evaluation data of each candidate intervention scheme is compared with the preset scheme standard one by one, and the schemes whose evaluation data do not meet the standard are excluded. Finally, the schemes whose evaluation data are superior to the standard are extracted as adaptive intervention strategies for the current evaluation object, ensuring that the selected intervention scheme has operable conditions and can effectively promote the state of the evaluation object to change in a healthy direction.

[0137] An electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor implements a talk-based multi-modal feature analysis method when executing the program.

[0138] For example, Fig. 3As shown, the electronic device can include a processor 610, a communication interface 620, a memory 630, and a communication bus 640, wherein the processor 610, the communication interface 620, and the memory 630 can communicate with each other through the communication bus 640. The processor 610 can invoke the logic instructions in the memory 630 to execute the talk-based multi-modal feature analysis method.

[0139] In addition, the logic instructions in the memory 630 described above can be implemented in the form of a software functional unit and sold or used as an independent product, which can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory, a random access memory, a magnetic disk or an optical disk, and various media that can store program codes.

[0140] On the other hand, the present application also provides a computer program product, which includes a computer program, the computer program can be stored on a non-transitory computer readable storage medium, and the computer program can be executed by a processor to enable a computer to execute the talk-based multi-modal feature analysis method.

[0141] In yet another aspect, the present application also provides a non-transitory computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the talk-based multi-modal feature analysis method.

[0142] The device embodiments described above are only schematic, wherein the units shown as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, i.e., they can be located in one place, or distributed on multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the present embodiment. Those skilled in the art can understand and implement it without creative labor.

[0143] Those skilled in the art can clearly understand the technical solutions of the various embodiments from the above description of the embodiments, and the various embodiments can be implemented by means of software with the necessary general hardware platforms, and of course, can also be implemented by hardware. Based on such understanding, the above technical solutions, essentially or in other words, the part of the prior art that makes a contribution, can be embodied in the form of a software product, which can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, and the like, and includes a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0144] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for some technical features therein; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A conversation-based multimodal feature analysis method, characterized in that, The method includes the following steps: Collect multimodal assessment data of the conversation process of the subject to be assessed, and perform feature labeling on the multimodal assessment data to obtain state feature points. The specific steps include: Language state features are obtained by labeling speech data with intonation and speech rate features; Facial expression data is labeled with facial expression action features to obtain facial expression language state features; Semantic and linguistic state features are used to label text response data to obtain text response linguistic state features; Among them, the language state features, facial expression language state features, and text response language state features constitute state feature points; The feature association parameters are obtained by analyzing the frequency and intensity of similar features among statistical state feature points. This process includes the following steps: The frequency of occurrence of a feature is obtained by counting the frequency of occurrence of the same type of feature among state feature points within a preset time window; The feature strength value is obtained by calculating the degree of linguistic state expression of the state feature points; Feature association parameters are obtained by mapping the frequency of feature occurrence with feature intensity values; Extracting the distribution patterns of historical language state features under different assessment scenarios from historical assessment datasets, and calculating the language state influence factor between language state tendency values ​​and feature correlation parameters in historical assessment data based on the distribution patterns of historical language state features, specifically including the following steps: Based on the distribution pattern of historical language state characteristics, historical language state tendency values ​​corresponding to language state tendency labels are extracted from historical assessment data. The language state difference value is obtained by calculating the language state tendency difference corresponding to the same type of feature in the historical language state tendency value based on the feature association parameter; The feature ratio value is obtained by proportionally calculating the language state difference value and the feature association parameter. The language state influence factor is obtained by weighted summation of all feature proportions. Based on the real-time language state feature distribution characteristics of the current conversation scenario, target influencing factors are selected from the language state influencing factors. The specific steps include: Obtain the real-time language state feature distribution characteristics of the current conversation scene; After matching the real-time language state feature distribution with the historical language state feature distribution, the target influence factor is selected from the language state influence factors. After determining the core language state tendency and auxiliary language state tendency of the subject to be evaluated during the conversation based on the target impact factor, the first language state evaluation value and the second language state evaluation value are output. The specific steps include: Determine the core and auxiliary language state expression dimensions of the subject being evaluated during the conversation; Determine the first feature set corresponding to the core language state and the second feature set corresponding to the auxiliary language state; After determining the core language state tendency value based on the feature association parameters corresponding to the first feature set, the real-time language state strength value, and the target influence factor, the first language state evaluation value is output. After determining the auxiliary language state tendency value based on the feature association parameters corresponding to the second feature set, the real-time language state strength value, and the target influence factor, the second language state evaluation value is output. The current status assessment value is obtained by judging the health status value of the assessed object based on the first language status assessment value and the second language status assessment value; The intervention strategy for the subject to be evaluated is determined based on the current status assessment value.

2. The conversation-based multimodal feature analysis method according to claim 1, characterized in that, The multimodal evaluation data includes voice data, facial expression data, and text response data.

3. The conversation-based multimodal feature analysis method according to claim 1, characterized in that, The current status assessment value is obtained by judging the health status of the assessed object based on the first language status assessment value and the second language status assessment value. The specific steps include: Assign core weight values ​​to the first language status assessment values; Assign auxiliary weight values ​​to the second language status assessment values; The first evaluation factor is obtained by multiplying the first language status evaluation value by the core weight value; The second evaluation factor is obtained by multiplying the second language status assessment value by the auxiliary weight value; The comprehensive status assessment factor is obtained by summing the first assessment factor and the second assessment factor. Obtain a reference assessment database of state health values ​​corresponding to different combinations of language state tendencies and feature intensities; The current status assessment value is obtained by matching the comprehensive status assessment factors with the reference assessment database.

4. The conversation-based multimodal feature analysis method according to claim 1, characterized in that, Determine the intervention strategy for the subject to be evaluated based on the current status assessment value, specifically including the following steps: If the current status assessment value is lower than the status warning threshold, an intervention prompt message will be output; After calling the preset intervention plan library based on the intervention prompt information, candidate intervention plans that match the current language state feature distribution features are extracted; Evaluation data is obtained by assessing the implementation complexity and expected outcome of candidate intervention programs. The best-fit intervention strategy is one that extracts evaluation data from candidate intervention programs and finds that outperforms the pre-defined intervention standards.

5. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the conversation-based multimodal feature analysis method as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Psychological analysis method based on certifying psychology and mapping knowledge domain

    CN118121198A

  • Psychological counseling interaction method and device based on autonomous psychological planning architecture

    CN120636701A