Digital human posture adjustment method and device combined with image recognition

By combining image recognition technology to identify visitors' emotions and analytical demands in real time, digital people can adjust their adaptive posture and voice, solving the problem that existing digital people find it difficult to accurately perceive visitors' emotions and demands, and improving the quality of interactive experience.

CN119473209BActive Publication Date: 2025-05-06ANCIENT CONTINENTAL ARTIFICIAL INTELLIGENCE TECHNOLOGY (SUZHOU) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411558827.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-04
Publication Date
2025-05-06
Estimated Expiration
2044-11-04

AI Technical Summary

Technical Problem

It is difficult for existing digital people to accurately perceive visitors' emotions and demands, and make adaptive postures and voice adjustments, affecting the quality of the interactive experience.

Method used

By combining image recognition technology, visitors' facial image data and text feature data are collected, emotions are identified in real time and appeal summary are analyzed, adaptive response text is generated, and digital human poses and speech are dynamically adjusted through emotional pose sequences and rhythm labeling sequences.

Benefits of technology

It improves the accuracy of digital people's recognition of visitors' emotions and demands, enhances the adaptability of digital people's posture and voice, and improves the quality of interactive experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119473209B_ABST
    Figure CN119473209B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and device for adjusting the posture of a digital human combined with image recognition, and relates to the field of digital human control technology. The method includes: starting the data collection behavior of the interactive object; synchronously performing cumulative emotion recognition in the process of facial image data collection to obtain real-time interactive emotions; performing interactive demand analysis to obtain an interactive demand summary; generating an adaptive response text according to the interactive demand summary; integrating and analyzing the adaptive response text and the real-time interactive emotions, outputting an emotional posture sequence and a rhythmic annotation sequence; and synchronously and dynamically synthesizing and playing the adaptive response voice according to the rhythmic annotation sequence and the adaptive response text. The present invention solves the technical problem in the prior art that it is difficult for digital humans to accurately perceive the emotions and demands of visitors and to perform adaptive posture and voice adjustments, and achieves the technical effect of improving the recognition accuracy of digital humans to the emotions and demands of visitors and enhancing the adaptability of digital humans' postures and voices to the emotions of visitors.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of digital human control, and in particular to a method and device for adjusting the posture of a digital human combined with image recognition. Background Art

[0002] With the development of artificial intelligence technology, digital humans have been widely used in many fields. However, existing digital human technology has some limitations. In the process of interacting with visitors, digital humans often find it difficult to accurately perceive visitors' emotions and demands. Relying only on single-modal information, such as analyzing only voice or text, it is impossible to fully capture the emotional state of visitors. Moreover, the recognition of emotions is not accurate and dynamic enough, and the changing trend of emotions cannot be tracked in real time. At the same time, there are also deficiencies in adjusting posture and voice according to visitors' emotions and demands. Digital humans may react stiffly, lacking natural posture and voice expression that match visitors' emotions. For example, it is impossible to adjust their body language, facial expressions, and voice intonation according to the different emotional states of visitors, which affects the quality of the interactive experience.

[0003] The existing technology has the technical problem that it is difficult for digital humans to accurately perceive the emotions and demands of visitors and make adaptive posture and voice adjustments. Summary of the invention

[0004] The present application provides a method and device for adjusting the posture of a digital human combined with image recognition, which is used to solve the technical problem in the prior art that it is difficult for a digital human to accurately perceive the emotions and demands of visitors and to make adaptive posture and voice adjustments.

[0005] In view of the above problems, the present application provides a method and device for adjusting the posture of a digital human in combination with image recognition.

[0006] In a first aspect of the present application, a method for adjusting the posture of a digital human combined with image recognition is provided, the method comprising:

[0007] After the interactive object is allowed to interact with the target digital human, data collection of the interactive object is started; in the process of collecting facial image data of the interactive object, cumulative emotion recognition is synchronously performed to obtain real-time interactive emotions; after the text feature data collection of the interactive object is completed, the interactive demand is parsed based on the text feature data to obtain an interactive demand summary; an adaptive response text is generated according to the interactive demand summary; the adaptive response text and the real-time interactive emotions are integrated and analyzed to output an emotional posture sequence and a rhythmic annotation sequence; in the process of using the emotional posture sequence to adjust the posture of the target digital human, the adaptive response voice is synchronously and dynamically synthesized and played according to the rhythmic annotation sequence and the adaptive response text.

[0008] In a second aspect of the present application, a digital human posture adjustment device combined with image recognition is provided, the device comprising:

[0009] A data collection module, which is used to start data collection on the interactive object after allowing the interactive object to interact with the target digital person; a real-time interactive emotion acquisition module, which is used to synchronously perform cumulative emotion recognition in the process of collecting facial image data of the interactive object to obtain real-time interactive emotions; an interactive demand summary acquisition module, which is used to perform interactive demand analysis based on the text feature data after completing the text feature data collection on the interactive object to obtain an interactive demand summary; an adaptive response text generation module, which is used to generate an adaptive response text according to the interactive demand summary; a sequence output module, which is used to fuse and analyze the adaptive response text and real-time interactive emotions, and output an emotional posture sequence and a rhythmic annotation sequence; a dynamic synthesis playback module, which is used to perform synchronous dynamic synthesis playback of the adaptive response voice according to the rhythmic annotation sequence and the adaptive response text in the process of using the emotional posture sequence to adjust the posture of the target digital person.

[0010] One or more technical solutions provided in this application have at least the following technical effects or advantages:

[0011] After the interactive object is allowed to interact with the target digital human, the data collection behavior of the interactive object is started; in the process of collecting facial image data of the interactive object, cumulative emotion recognition is synchronously performed to obtain real-time interactive emotions; after the text feature data collection of the interactive object is completed, the interactive demand analysis is performed to obtain the interactive demand summary; the adaptive response text is generated according to the interactive demand summary; the adaptive response text and the real-time interactive emotions are integrated and analyzed to output the emotional posture sequence and the rhythm annotation sequence; the adaptive response voice is synchronously and dynamically synthesized and played according to the rhythm annotation sequence and the adaptive response text. The technical effect of improving the recognition accuracy of the digital human to the visitor's emotions and demands and enhancing the adaptability of the digital human's posture and voice to the visitor's emotions is achieved. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0013] Figure 1A schematic diagram of a flow chart of a method for adjusting a digital human's posture in combination with image recognition provided in an embodiment of the present application;

[0014] Figure 2 A schematic diagram of the structure of a digital human posture adjustment device combined with image recognition provided in an embodiment of the present application.

[0015] Explanation of the accompanying drawings: data collection module 10, real-time interactive emotion acquisition module 20, interactive demand summary acquisition module 30, adapted response text generation module 40, sequence output module 50, dynamic synthesis playback module 60. DETAILED DESCRIPTION

[0016] The present application provides a method and device for adjusting the posture of a digital human combined with image recognition, which is used to solve the technical problem in the prior art that it is difficult for a digital human to accurately perceive the emotions and demands of visitors and to make adaptive posture and voice adjustments.

[0017] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0018] Embodiment 1, as Figure 1 As shown, the present application provides a digital human posture adjustment method combined with image recognition, the method comprising:

[0019] Step S100: After the interactive object is allowed to interact with the target digital human, data collection on the interactive object is started.

[0020] Specifically, when the interactive scene is set to allow the interactive object to interact with the target digital human, this permission mechanism triggers the start of the entire data collection process. Once the interactive object is identified to start interacting with the target digital human, such as interacting with the digital human through voice commands, touch operations or specific gestures, the data collection behavior will be started. The data collection behavior covers multiple dimensions, aiming to fully obtain relevant information about the interactive object, including but not limited to the collection of voice data of the interactive object, recording the voice content of the interactive object and the digital human through microphones and other devices; the collection of facial image data of the interactive object, using cameras to capture the facial expressions, eyes and other image information of the interactive object; and the collection of other body movements, postures and other data of the interactive object, such as sensing the hand movements and body posture changes of the interactive object through sensors. These collected data will provide a rich data foundation for the subsequent intention analysis of the interactive object, emotion recognition and reasonable response of the digital human.

[0021] Step S200: while collecting facial image data of the interactive object, cumulative emotion recognition is synchronously performed to obtain real-time interactive emotions.

[0022] Specifically, in the key link of facial image data collection of the interactive object, the cumulative emotion recognition operation is carried out simultaneously, the purpose of which is to accurately obtain real-time interactive emotions. This emotional result is crucial for driving the reasonable feedback of the digital human, and can ensure the accuracy and real-time nature of the interactive emotions. First, an emotion recognition update interval is preset. When collecting facial image data, the camera continuously collects the facial images of the interactive object based on this preset emotion recognition update interval as the image data update interval, thereby obtaining a multi-frame image stream. These multi-frame image streams will be temporarily stored and synchronized to the cumulative emotion recognition model at intervals according to the transmission cycle set by the emotion recognition update interval. The cumulative emotion recognition model is a complex and critical part. It undertakes the important task of capturing the trend of emotional changes from the image stream. The model contains multiple hierarchical structures, including a facial feature recognition layer, an emotional state update layer, and an emotional trend extraction layer. When the multi-frame image stream is passed into the model, it first passes through the facial feature recognition layer, which will perform emotional feature recognition on multiple image frames in the image stream, thereby obtaining the first emotional feature sequence. Next, after the emotional state update layer receives the first emotional feature sequence, it will filter out the background emotions through the emotional feature selection gate to obtain the main emotional feature sequence. Then, the emotional state integration gate extracts and connects the time-series emotions of the main emotional feature sequence based on the emotional type to obtain the first set of emotional change features. Thereafter, the emotional trend extraction layer receives the first set of emotional change features, and uses the emotional mutation detection unit to traverse these features to identify emotional mutation nodes to obtain the first set of emotional mutation nodes. Finally, the emotional output regulation unit smoothly optimizes the first set of emotional change features according to the first set of emotional mutation nodes to obtain the first emotional trend vector. Through the above process, the cumulative emotion recognition model continuously captures the emotional change trend based on the multi-frame image stream uploaded at intervals, and obtains multiple emotional trend vectors. Then, these multiple emotional trend vectors are fused and analyzed to finally obtain real-time interactive emotions. This real-time interactive emotion will serve as the visitor's interactive emotion and provide a key basis for the digital human's feedback, thereby ensuring that the interaction between the digital human and the interactive object is more natural and smooth.

[0023] Step S300: After the text feature data collection of the interactive object is completed, the interactive demand is analyzed based on the text feature data to obtain an interactive demand summary.

[0024] Specifically, after the collection of text feature data of the interactive object is completed, it enters the important stage of analyzing the interactive demand based on the collected text feature data. The ultimate goal is to obtain a summary of the interactive demand. Before this, a series of data collection operations are performed on the interactive object. In the process of collecting text feature data, the voice information of the interactive object is converted into text form through speech recognition technology, or the text content input by the interactive object is directly obtained to complete the collection of text feature data. The collected text feature data contains key information in the process of interaction between the interactive object and the digital human. Based on this data, natural language processing technology and related algorithms are used to analyze the interactive demand. First, the text will be segmented and decomposed into meaningful word units. Then, through part-of-speech tagging, syntactic analysis and other means, the grammatical structure and semantic information of the text are further understood. In the analysis process, the key elements such as needs, intentions, and problems in the text will be identified. For example, if the text contains the phrase "I want to know about health", the system will identify the key verb "I want to know" that expresses the demand and the demand object "knowledge about health". By extracting and integrating these key elements, it will eventually form a concise and clear summary of the interaction demand. This summary can accurately summarize the core demands of the interaction object during the interaction process, and provide an important basis for the digital human to subsequently generate adaptive response texts.

[0025] Step S400: Generate an adapted response text according to the interaction request summary.

[0026] Specifically, after obtaining the interactive demand summary, the adapted response text is generated based on this summary. The interactive demand summary summarizes the core demands of the interactive object in the process of interacting with the digital human. Based on this summary, it will first search and match in its pre-built knowledge graph or database. This knowledge graph or database contains a large amount of knowledge information, covering professional knowledge in various fields, answers to common problems, and coping strategies in various situations. For example, if the interactive demand summary involves a certain issue in the health field, it will search for matching content in the health-related knowledge section. During the search process, multiple dimensions such as the accuracy, relevance, and depth of the problem will be considered to ensure that the most appropriate information is found to respond to the interactive object. Then, the text is organized and generated based on the matching information found, involving operations such as extraction, rewriting, and integration of the original information. If the information found is a more complex paragraph, the key sentences are extracted and rearranged in a reasonable logical order to make it more in line with the context and requirements of the response. At the same time, if some concepts need to be further explained or supplemented, relevant content will be added accordingly to ensure that the response text can fully and accurately meet the demands of the interactive object. The adaptive response text finally generated will be able to answer the questions of the interactive object in a targeted manner and meet its needs during the interaction process, thus making the interaction between the digital human and the interactive object smoother and more effective.

[0027] Step S500: Fusion analysis of the adapted response text and the real-time interactive emotion, and output of an emotion gesture sequence and a rhythm annotation sequence.

[0028] Specifically, first, we need to obtain multiple sample interaction emotion nodes and multiple sample feedback emotion nodes. These nodes come from a large number of interaction instances and cover a wealth of emotions and feedback information. Then, we construct a mapping relationship between sample interaction emotion nodes and sample feedback emotion nodes, and build a feedback emotion map based on this. The map connects different emotional states and their corresponding feedback performances. Multiple groups of sample emotion postures are stored in the feedback emotion map. These sample emotion postures include action postures and voice postures, which describe in detail the posture performance of the digital human under different emotions. After that, the real-time interaction emotion is traversed in the feedback emotion map. Once the real-time interaction emotion matches an emotion node in the map, the corresponding emotion posture sequence can be called out. This sequence consists of action posture emotions and voice posture emotions. It clarifies the posture and voice emotion characteristics that the digital human should show when responding. Finally, the voice posture emotions are used to perform rhythmic annotation on the adapted response text to obtain a rhythmic annotation sequence. Prosodic annotation covers elements such as the text's speech rhythm, intonation, and stress, ensuring that the adapted response voice can naturally convey the corresponding emotions when synthesized and played, giving the digital human's response an accurate and rich basis for emotional and postural expression.

[0029] Step S600: in the process of adjusting the posture of the target digital human by using the emotional posture sequence, synchronous dynamic synthesis and playback of the adaptive response voice are performed according to the rhythmic annotation sequence and the adaptive response text.

[0030] Specifically, in terms of posture adjustment, the emotional posture sequence is used as a guide. The emotional posture sequence includes action posture emotions and voice posture emotions, which clarifies the posture characteristics that the digital human should present in the current interactive situation. According to this sequence, the postures of various parts of the target digital human's body are adjusted, such as the angle of the head, the movement of the arms, the posture of the body, etc., so that it meets the emotional expression and action specifications required by the emotional posture sequence. At the same time, in terms of speech synthesis and playback, operations are performed based on the rhythm annotation sequence and the adaptive response text. The rhythm annotation sequence provides the adaptive response voice with specifications in terms of speech rhythm, intonation, stress, etc., and the adaptive response text is the basis of the speech content. Through specific speech synthesis technology, the adaptive response text is dynamically synthesized according to the requirements of the rhythm annotation sequence. During the synthesis process, it is considered how the intonation of the voice should change to reflect the emotion, which words should be stressed to highlight the key points, and how the rhythm of the voice should be grasped to maintain natural fluency. Finally, the adaptive response voice and the digital human posture adjustment are synchronized. While adjusting the posture, the digital human emits the adaptive response voice that meets the current interactive situation, making the interaction between the digital human and the interactive object more natural and smooth.

[0031] In a possible implementation, step S200 further includes:

[0032] Step S210: Preset the emotion recognition update interval.

[0033] Step S220: In the process of using a camera to collect facial image data of the interactive object, the camera temporarily stores the multi-frame image stream obtained by collecting with the emotion recognition update interval as the image data update interval, and synchronizes the multi-frame image stream interval to the cumulative emotion recognition model with the emotion recognition update interval as the transmission period.

[0034] Step S230: the cumulative emotion recognition model captures and updates the emotion change trend based on the multi-frame image stream uploaded at intervals to obtain the real-time interactive emotion.

[0035] Specifically, we first need to set an emotion recognition update interval. This interval is a key time parameter in the entire facial image data collection and emotion recognition process. It determines the frequency of data collection and the rhythm of model updates, laying the foundation for the subsequent accurate acquisition of real-time interactive emotions.

[0036] The facial image data of the interactive object is collected by the camera, and the collected data is processed and transmitted according to specific rules. When the facial image data of the interactive object is collected by the camera, the update interval of the image data is determined according to the preset emotion recognition update interval. This update interval is very critical. It ensures that the collected image can reflect the changes in the facial expression of the interactive object at an appropriate time interval. With this interval as the standard, the camera continuously collects facial images to obtain a multi-frame image stream. Then, the collected multi-frame image stream will be temporarily stored. This is because these image streams need to be further analyzed and transmitted in the subsequent processing process. Temporary storage can ensure the integrity and availability of the data. Finally, with the same emotion recognition update interval as the transmission period, the temporarily stored multi-frame image stream interval is synchronized to the cumulative emotion recognition model. This synchronous transmission method enables the cumulative emotion recognition model to receive the latest image data in time, so as to better identify and analyze the emotions of the interactive object. In this way, it can be ensured that the data used by the model is the latest and best reflects the current emotional state of the interactive object, thereby improving the accuracy and real-time performance of emotion recognition.

[0037] The cumulative emotion recognition model plays a key role. It captures and updates the trend of emotion changes based on the multi-frame image stream uploaded at intervals, thereby obtaining real-time interactive emotions. When the multi-frame image stream uploaded at intervals enters the cumulative emotion recognition model, the emotion features will be extracted at a specific level of the model. For example, in the facial feature recognition layer, facial feature analysis will be performed on each image frame in the image stream to identify facial feature information related to emotions, thereby obtaining a series of emotion feature sequences. Then, these emotion feature sequences will be further processed in the subsequent levels of the model. For example, in the emotional state update layer, the emotion feature sequences will be screened and integrated through some specific rules and algorithms to remove some interference information, highlight the main emotional features, and obtain a more refined emotional information representation. Then, in the emotional trend extraction layer, the processed emotional information will be trend analyzed. By observing and analyzing the emotional information on a series of time series, the model can capture trend information such as the direction and magnitude of emotional changes. In the whole process, the model will be continuously updated and adjusted according to the newly uploaded image stream data. As more image stream data is processed, the model will capture the trend of emotional changes more accurately and finely. Finally, through comprehensive analysis and integration of all processed emotional information, the model obtains real-time interactive emotions, which can accurately reflect the emotional state of the interactive object during the interaction with the digital human, and provide an important basis for the reasonable feedback of the digital human.

[0038] In a possible implementation, step S300 further includes:

[0039] Step S310: During the interactive voice data collection process for the interactive object, the audio segments collected are temporarily stored with the emotion recognition update interval as the voice data update interval, and the emotion recognition update interval is used as the analysis period to analyze the interval emotion change trend of the audio segments to obtain the verified interactive emotions.

[0040] Step S320: using the verification interaction emotion to perform credible verification of the real-time interaction emotion.

[0041] Step S330: After completing the collection of the interactive voice data of the interactive object, the text feature data is obtained by recognizing and converting the interactive voice data.

[0042] Step S340: performing interaction demand analysis based on the text feature data to obtain the interaction demand summary.

[0043] Specifically, first of all, it is very important to choose a suitable audio acquisition device. A microphone with high sensitivity and low noise is selected to ensure that the voice signal of the interactive object can be captured clearly and accurately. Then, the microphone is configured according to the emotion recognition update interval, and the sampling frequency and sampling time interval of the microphone are set through the software so that it can collect audio according to the set emotion recognition update interval as the voice data update interval. For example, if the emotion recognition update interval is set to 3 seconds, the microphone will collect audio clips at intervals of 3 seconds. The collected audio clips need to be temporarily stored for subsequent analysis, and efficient data storage structures and algorithms are used to achieve this purpose. A ring buffer is used as a container for temporarily storing audio clips. The ring buffer has a fixed size. When a new audio clip is collected, if the buffer is full, it will automatically overwrite the earliest stored audio clip, thereby ensuring real-time data update and effective use of storage space. At the same time, in order to ensure the integrity and accessibility of the data, the audio clips will be marked and indexed during the temporary storage process. These marks and indexes help to quickly locate and extract audio clips of a specific time period for subsequent analysis. For the temporarily stored audio clips, the first step is to extract speech features, which is the basis for analyzing the trend of emotional changes. Various speech signal processing technologies are used to extract different types of speech features. For example, the fundamental frequency features of speech can be extracted through the linear predictive coding (LPC) technology. The fundamental frequency is an important parameter reflecting the pitch of speech. The formant features of speech can be extracted through the Mel-frequency cepstral coefficient (MFCC) technology. The formant is related to the timbre of speech. In addition, the rhythmic features of speech such as speech rate, intonation, and pauses are also extracted. These features can be obtained by analyzing the audio clips in the time domain and frequency domain. After extracting the speech features, it is necessary to associate these features with emotional features, which can be achieved by machine learning or deep learning algorithms. First, an emotional feature association model is constructed and the model is trained using a large number of speech samples labeled with emotions. Support vector machines (SVM) are used for training. During the training process, the model learns the association between speech features and emotions. Then, the speech features extracted from the temporarily stored audio clips are input into the trained emotional feature association model, and the model outputs the corresponding emotional prediction results based on the input speech features. These emotion prediction results are further processed and analyzed to finally obtain the validation interaction emotion. For example, if the model predicts that the emotion corresponding to a certain audio clip is "happy", and this emotion has a certain change trend in consecutive audio clips, then the validation interaction emotion and its change trend can be determined.

[0044] After obtaining the validation interaction emotion, it is necessary to use it to verify the credibility of the real-time interaction emotion. Since the real-time interaction emotion is obtained through the cumulative emotion recognition in the process of facial image data collection, and the validation interaction emotion is obtained from the analysis of interactive voice data collection, the two reflect the emotions of the interactive object from different angles. By comparing the validation interaction emotion and the real-time interaction emotion, if the two are consistent or within a reasonable error range, for example, the emotion categories represented by the two are the same and the emotion intensity is similar, then the credibility of the real-time interaction emotion is improved. If the difference between the two is large, the cause needs to be further analyzed. It may be that there is a problem with the collection equipment, or the emotion recognition algorithm has limitations in some aspects, and the relevant equipment or algorithm needs to be adjusted and optimized.

[0045] After completing the key stage of collecting interactive voice data of the interactive object, the interactive voice data is then recognized and converted to obtain text feature data. This process is achieved with the help of speech recognition technology. First, the speech recognition system will pre-process the collected interactive voice data, including noise reduction processing of the audio signal to improve the clarity and recognizability of the voice. At the same time, the audio signal will be framed to divide the continuous voice signal into smaller time periods to better analyze the characteristics of the voice. The speech recognition system will extract a series of features from the pre-processed audio signal, including acoustic features and prosodic features. Acoustic features such as fundamental frequency and formant reflect the physical characteristics of speech. Prosodic features such as speech rate and intonation reflect the rhythm and intonation changes of speech. After extracting the features, the speech recognition system will perform recognition conversion based on a pre-trained speech model. This speech model is obtained through training with a large amount of speech data. It contains the correspondence between different speech features and corresponding texts. The speech recognition system will match the extracted features with the features in the speech model and finally determine the text content corresponding to the interactive voice data. The text content obtained through recognition and conversion is the required text feature data. It contains various information expressed by the interactive object during the interaction with the digital human, and provides an important data basis for subsequent analysis of interactive needs.

[0046] The text feature data is segmented. Segmentation is to divide the continuous text into meaningful word units according to semantic and grammatical rules. For example, for the text "I want to know about health and wellness", after segmentation, the words "I", "I want", "I know", "about", "health and wellness", "of", "knowledge" and so on are obtained. The accuracy of segmentation is crucial for subsequent analysis, which lays the foundation for further understanding the semantics and structure of the text. Next, part-of-speech tagging is performed. Part-of-speech tagging is to determine the grammatical category of each word in the sentence, such as noun, verb, adjective, pronoun, etc. In the above example, "I" is a pronoun, "I want" and "I know" are verbs, "about" is a preposition, "health and wellness" is a noun phrase, "of" is an auxiliary word, and "knowledge" is a noun. Part-of-speech tagging helps to understand the grammatical relationship and semantic role between words. Then, syntactic analysis is performed. Syntactic analysis aims to determine the grammatical structure of the sentence, including subject, predicate, object, attributive, adverbial and other components. In the sentence "I want to know about health and wellness", the subject is "I", the predicate is "want to know", and the object is "knowledge about health and wellness". Syntactic analysis can help us better understand the semantics and intention of the sentence. After completing the above basic natural language processing steps, we start to identify key elements in the text that express needs, intentions, questions, etc. In this example, "want to know" expresses the need, and "knowledge about health and wellness" is the object of the need. Finally, these key elements are extracted and integrated to form a concise and clear summary of the interaction demand. For the above example, the interaction demand summary may be "knowledge about health and wellness". This summary can accurately summarize the core demands of the interaction object in the interaction process, and provide an important basis for the subsequent generation of adaptive response texts based on the demands.

[0047] In a possible implementation, step S220 further includes:

[0048] Step S221: The cumulative emotion recognition model includes a cascaded facial feature recognition layer, an emotional state update layer, and an emotional trend extraction layer.

[0049] Step S222: interactively obtain a sample emotion feature set, and use the sample emotion feature set as training data to train a facial feature recognition model pre-built based on a convolutional neural network.

[0050] Step S223: The facial feature recognition layer is constructed by synchronizing the trained and optimized facial feature recognition model to the facial feature recognition layer.

[0051] Step S224: constructing an emotion feature selection gate using predefined emotion feature screening rules, and completing the construction of the emotion state update layer by cascading the emotion feature selection gate and the emotion state integration gate.

[0052] Step S225: introducing a multi-dimensional emotion fluctuation threshold to construct an emotion mutation detection unit, and completing the construction of the emotion trend extraction layer by cascading the emotion mutation detection unit and the emotion output regulation unit.

[0053] Specifically, the cumulative emotion recognition model has a hierarchical structural design, which is composed of three key layers in cascade. Among them, the facial feature recognition layer, as the basic layer, is mainly responsible for identifying various facial features from the input facial image data, providing the original data basis for subsequent emotion analysis. On the basis of the facial feature recognition layer, the emotional state update layer further processes the acquired facial feature-related information, updates and maintains the emotional state-related information to make it more consistent with the actual emotional situation. The emotional trend extraction layer is at a high level of the model. Based on the results of the first two layers, it focuses on extracting and analyzing the changing trends of emotions, so as to grasp the dynamic changes of emotions more comprehensively and accurately. The three layers work together to achieve effective recognition of cumulative emotions.

[0054] The sample emotion feature set is obtained in an interactive way. This process involves collecting a large amount of facial image data containing different emotional states and extracting relevant emotion feature information from these images. Then, the obtained sample emotion feature set is used as training data and applied to the facial feature recognition model pre-built based on the convolutional neural network. During the training process, the model continuously adjusts its parameters based on the feature information in the sample data to learn the intrinsic connection between facial features and emotions, thereby improving the accuracy and effectiveness of facial feature recognition and laying the foundation for subsequent emotion recognition.

[0055] After completing the training and optimization of the pre-built facial feature recognition model based on the convolutional neural network, the next key step is to synchronize the trained model to the facial feature recognition layer. This process is to pass the parameters, structure and related algorithm logic of the model obtained after training and optimization to the facial feature recognition layer. Through this synchronization operation, the facial feature recognition layer can inherit the capabilities of the trained model, so that it has the function of accurately identifying facial features, and finally complete the construction of the facial feature recognition layer, laying a solid foundation for subsequent emotion recognition in the entire cumulative emotion recognition model.

[0056] The emotion feature selection gate is constructed based on the predefined emotion feature screening rules such as the weight parameter range set for each emotion as the background emotion. This process requires in-depth analysis and quantification of the relevant features of different emotions. For example, for the happy emotion, the weight parameter range of certain facial muscle movement features (such as the degree of the mouth corners, the shape of the wrinkles at the corners of the eyes, etc.) and voice features (such as the pitch, the speed of speech, etc.) is set as the background emotion. Through a large number of experiments and data statistics, the relative importance of these features in different emotional states is determined, so as to construct a selection gate that can effectively screen emotion features. Then, the emotion feature selection gate and the emotion state integration gate are cascaded. When the emotion feature selection gate screens out the relevant emotion features, these features will be passed to the emotion state integration gate. The emotion state integration gate will further integrate these screened emotion features according to the emotion type. For example, for happiness and excitement, which belong to the same positive emotion category, the emotion state integration gate will reasonably integrate them according to their respective characteristics and the preset integration rules, so that these features can more accurately reflect the current emotional state. Through this cascading method, the emotional state update layer is constructed, providing a more accurate emotional state information basis for subsequent emotional trend extraction.

[0057] The concept of multidimensional emotion fluctuation threshold is introduced to construct an emotion mutation detection unit. This unit monitors emotion changes according to the set multidimensional threshold and can keenly capture emotion mutations. Then, the emotion mutation detection unit is cascaded with the emotion output regulation unit. The emotion output regulation unit will regulate and optimize the detected emotion mutations to make the emotion change trend smoother and more reasonable. Through this cascade method, the construction of the emotion trend extraction layer is finally completed, so that it has the ability to accurately extract emotion trends from relevant emotion information, providing key trend information support for the entire emotion recognition and analysis process.

[0058] In a possible implementation, step S230 further includes:

[0059] Step S231: after receiving the first multi-frame image stream, the cumulative emotion recognition model performs emotion feature recognition on a plurality of image frames in the first multi-frame image stream via the facial feature recognition layer to obtain a first emotion feature sequence.

[0060] Step S232: after receiving the first emotion feature sequence, the emotion state update layer filters out background emotions through the emotion feature selection gate to obtain a main emotion feature sequence.

[0061] Step S233: the emotional state integration gate extracts and connects the time-series emotions of the main emotional feature sequence based on the emotional type to obtain a first set of emotional change features.

[0062] Step S234: After receiving the first group of emotion change features, the emotion trend extraction layer uses the emotion mutation detection unit to traverse the first group of emotion change features to identify emotion mutation nodes, and obtain a first group of emotion mutation nodes.

[0063] Step S235: the emotion output adjustment unit performs smoothing optimization on the first group of emotion change features according to the first group of emotion mutation nodes to obtain a first emotion trend vector.

[0064] Step S236: Similarly, the cumulative emotion recognition model captures the emotion change trend based on the multiple frame image streams uploaded at intervals to obtain multiple emotion trend vectors.

[0065] Step S237: Fusion analysis of the multiple emotion trend vectors to obtain the real-time interactive emotion.

[0066] Specifically, when the cumulative emotion recognition model receives the first multi-frame image stream, the facial feature recognition layer therein will perform emotion feature recognition on each image frame in these image streams. For each image, its emotional feature composition is reflected by the weights of multiple emotion types. For example, for a certain image, it may present a weight distribution such as joy: 0.6, sadness: 0.1, anger: 0.05, surprise: 0.15, calmness: 0.1. Here, the weight of joy is 0.6, which means that the model judges that the emotion displayed by the image is more inclined to joy, and its possibility or intensity of appearing in the image is the highest; while the weight of anger is only 0.05, which indicates that the expression of anger in the image is very weak. These weights comprehensively reflect the distribution of multiple emotions in the current image. The model can calculate the overall emotional state based on these weight information, and use this state for subsequent accumulation and update operations of emotional trends, and finally obtain the first emotional feature sequence, which contains comprehensive information about emotions extracted from these image frames.

[0067] When the emotional state update layer receives the first emotional feature sequence, it uses the emotional feature selection gate to screen out background emotions. The emotional feature selection gate determines the importance of each emotional feature based on pre-set rules and standards, and screens out those features that are not heavy enough to be the main emotion and are similar to background emotions. For example, if an emotional feature has a low weight in the sequence, it indicates that its contribution to the overall emotion is small and is judged as background emotion. Through this screening process, those unimportant background emotional features are eliminated, and the main emotional features that can have a key impact on the overall emotional state are retained. These main emotional features constitute the main emotional feature sequence, which provides a basis for more accurate analysis and grasp of the emotional state in the future.

[0068] After receiving the main emotional feature sequence, the emotional state integration gate processes it according to the emotional type. It carefully analyzes the emotional type corresponding to each main emotional feature in the sequence, such as happiness, sadness, or anger, and for each emotional type, it extracts the expression of the emotion in the time series, that is, the temporal emotion. For example, if the main emotional feature sequence contains multiple features related to happy emotions, and these features appear at different time points, the emotional state integration gate will extract these happy emotional features at different time points and connect them in chronological order. Through this operation, the originally scattered main emotional features are converted into information with a time dimension, thereby obtaining the first set of emotional change features, which can more clearly reflect the changing trend of emotions over a period of time.

[0069] After obtaining the first set of emotion change features, the emotion trend extraction layer uses the emotion mutation detection unit to conduct a comprehensive analysis of them. The emotion mutation detection unit will traverse each element in the first set of emotion change features according to pre-set rules and standards. It will look for points where there are sudden changes in emotion expression, that is, emotion mutation nodes. For example, when the emotion suddenly changes from relative calmness to excitement, or from happiness to sadness, the emotion mutation detection unit can identify this drastic change in emotion and mark it as an emotion mutation node. By carefully traversing the first set of emotion change features, the first set of emotion mutation nodes are finally determined, which are crucial for the subsequent understanding of the dynamic change process of emotions.

[0070] For each first group of emotion mutation nodes, a specific range centered on the node is determined as the focus area. This area contains the emotion change features before and after the node. Its size can be set according to the actual situation, such as the emotion features corresponding to several frames of images before and after the node. For each emotion change feature in the focus area, a coefficient related to the mutation node needs to be calculated. This coefficient is used to determine the influence of the feature in the smooth optimization process. The weight is determined according to the distance between the feature and the mutation node. The closer the feature is to the mutation node, the higher its weight is, while the farther the feature is, the lower its weight is. Then, according to the calculated weight coefficient, the weight of the emotion change feature in the focus area is adjusted. For the emotion type with more drastic weight changes near the mutation node, it is moderately eased according to its corresponding weight coefficient. For example, if the weight of a certain emotion suddenly increases at the mutation node, the weight of this sudden increase is adjusted according to the weight of the surrounding features and the weight of the distance from the mutation node to make its change smoother and avoid excessive jumps. After performing the above-mentioned smoothing adjustment on the weights of all emotion types in the first group of emotion change features, we finally obtain a first emotion trend vector that comprehensively considers the emotion mutation node and its surrounding features. This vector can more accurately reflect the changing trend of emotions. It is composed of the weights of multiple emotions after smoothing and optimization, presenting a relatively smooth change curve, which provides a more reasonable description of emotion changes for subsequent emotion analysis and digital human posture adjustment.

[0071] After completing the processing of the first set of image streams and obtaining the first emotion trend vector, the cumulative emotion recognition model will process the subsequent multiple-frame image streams uploaded at intervals according to the same process. For each new set of image streams, the model will first perform emotion feature recognition by the facial feature recognition layer to obtain the corresponding emotion feature sequence. Next, the emotion state update layer uses the emotion feature selection gate to filter out background emotions and obtain the main emotion feature sequence. Then, the emotion state integration gate extracts and connects the temporal emotions based on the emotion type to obtain a set of emotion change features. The emotion trend extraction layer then uses the emotion mutation detection unit to identify the emotion mutation nodes, and the emotion output regulation unit smoothly optimizes the emotion change features based on these nodes to obtain a new emotion trend vector. This cycle repeats itself. As image streams of different time periods are continuously uploaded and processed, the cumulative emotion recognition model eventually obtains multiple emotion trend vectors, each of which represents the emotion change trend in the corresponding time period.

[0072] After obtaining multiple emotion trend vectors, fusion analysis is required to obtain real-time interactive emotions. Based on a certain time weighting mechanism, different weights are assigned to emotion trend vectors in different time periods. For example, vectors in more recent time periods are assigned higher weights because they better reflect the current emotional state. Then, these weighted vectors are comprehensively analyzed. The weighted average of the calculation vectors is used to calculate the elements of each vector. In this process, the categories of emotions and the relationship between different categories also need to be considered. For example, if happy emotions dominate in some vectors and sad emotions dominate in other vectors, these different emotional information need to be balanced and integrated according to specific algorithms and rules. Through such a fusion analysis process, a real-time interactive emotion that can accurately reflect the real-time emotional state of visitors throughout the interaction process is finally obtained.

[0073] In a possible implementation, step S237 further includes:

[0074] Step S2371: After aggregating the multiple emotion trend vectors based on the sliding time window, capture the dominant emotion of the aggregation results to obtain K dominant emotion trends.

[0075] Step S2372: performing attenuation weighted assignment of the K dominant emotion trends according to the emotion capture time to obtain K emotion association weights.

[0076] Step S2373: Using the K emotion association weights to perform weighted summation on the K dominant emotion trends to obtain a comprehensive emotion trend.

[0077] Step S2374: Use the preset sample emotion weight distribution to perform weighted summation of the comprehensive emotion trend to obtain the real-time interactive emotion.

[0078] Specifically, the size of the sliding time window is determined, and here 3-5 trend vectors are set as the frame selection interval of a window. Starting from the first window, the emotional trend vectors in the window are aggregated. Aggregation is performed by vector addition. After the aggregation is completed, the aggregation results are analyzed to capture the dominant emotions, which requires in-depth research on the emotional components represented by the aggregated vectors. The dominant emotion is determined by comparing the weights or intensities of different emotional components. For example, if the relevant indicators of happy emotions (such as weight, intensity, etc.) are significantly higher than other emotions in the aggregation results, then happy emotions may be determined as the dominant emotion in the window. In this way, as the sliding time window slides on multiple emotional trend vectors in turn, the vectors in each window are aggregated and the dominant emotions are captured, and finally K dominant emotional trends are obtained. Each dominant emotional trend represents the most significant direction of emotional change in the corresponding window interval.

[0079] K dominant emotional trends are assigned attenuated weights according to the emotional capture time. The emotional capture time refers to the information related to the time point or time period when each dominant emotional trend is captured during the entire interaction process. Generally speaking, the dominant emotional trend closer to the current time may be assigned a higher weight because it better reflects the current emotional state. Through this attenuated weighted assignment method, a suitable weight value is determined for each dominant emotional trend, thereby obtaining K emotional association weights. These weight values ​​reflect the relative importance of different dominant emotional trends after comprehensively considering the time factor.

[0080] The K obtained emotion association weights are used to perform a weighted summation operation on the corresponding K dominant emotion trends. For each dominant emotion trend, it is multiplied by the corresponding emotion association weight, and then all the products are added. Such a weighted summation process can comprehensively consider each dominant emotion trend and their weights to obtain a comprehensive emotion trend. This comprehensive emotion trend is obtained after considering the dominant emotions and their weights in different time periods, and can more comprehensively reflect the overall change trend of emotions.

[0081] The comprehensive emotional trend is weighted and summed again using the preset sample emotional weight allocation. The preset sample emotional weight allocation is based on the pre-setting of the importance of different emotional types in the overall emotional expression. For example, some emotional types are considered more important in specific application scenarios and are therefore given higher weights. By combining the comprehensive emotional trend with the preset sample emotional weight allocation, a weighted summation operation is performed to finally obtain the real-time interactive emotion. This real-time interactive emotion can accurately reflect the overall emotional state of the visitor during the interaction with the digital human, and comprehensively considers multiple factors such as emotional changes in different time periods, dominant emotional trends, and the importance of different emotional types.

[0082] In a possible implementation, step S500 further includes:

[0083] Step S510: interactively obtaining a plurality of sample interaction emotion nodes and a plurality of sample feedback emotion nodes.

[0084] Step S520: creating a mapping relationship between the plurality of sample interaction emotion nodes and the plurality of sample feedback emotion nodes, and constructing a feedback emotion graph based on the mapping relationship.

[0085] Step S530: calling multiple groups of sample emotion gestures of the multiple sample feedback emotion nodes, and storing the multiple groups of sample emotion gestures into the feedback emotion graph.

[0086] Step S540: using the real-time interactive emotion to traverse the feedback emotion map, and calling to obtain the emotion gesture sequence, wherein the emotion gesture sequence is composed of action gesture emotions and voice gesture emotions.

[0087] Step S550: Prosody-annotate the adapted response text using the voice gesture emotion to obtain the prosody-annotated sequence.

[0088] Specifically, through a large number of interaction processes, multiple sample interaction emotion nodes and multiple sample feedback emotion nodes are obtained. These nodes are quantitative representations of user emotions and system feedback emotions in different interaction scenarios. For example, in an interaction, the happy emotions expressed by the user and the emotional state when the system gives corresponding positive feedback can be recorded as a sample interaction emotion node and a corresponding sample feedback emotion node. Through multiple such interaction collections, a sufficient number of nodes are accumulated for subsequent analysis and processing.

[0089] A mapping relationship is created for the obtained multiple sample interaction emotion nodes and multiple sample feedback emotion nodes. This mapping relationship aims to establish a connection between user emotions and system feedback emotions. For example, if a sample interaction emotion node represents the user's anger, then the corresponding sample feedback emotion node may be the emotional state when the system adopts a soothing strategy. Based on this one-to-one mapping relationship, a feedback emotion map is constructed. The feedback emotion map is a structured data representation that graphically displays the relationship between different emotion nodes, providing a basic framework for subsequent emotional posture calling and analysis.

[0090] After constructing the feedback emotion map, for multiple sample feedback emotion nodes in the map, call the corresponding multiple groups of sample emotion postures. These sample emotion postures are the expressions of body posture and voice intonation associated with each sample feedback emotion node. For example, for a sample feedback emotion node representing happiness, its corresponding sample emotion posture may include smiling facial expression, cheerful voice intonation, and relaxed body posture. These multiple groups of sample emotion postures are stored in the feedback emotion map, so that the map contains not only the mapping relationship between emotion nodes, but also the specific emotion expression information related to each node.

[0091] When the real-time interactive emotion is used to traverse the feedback emotion map, the real-time interactive emotion is compared and matched with the emotion represented by each node in the map. For each node in the map, there is a corresponding emotion feature and related information. In the traversal process, the node that is most similar or related to the real-time interactive emotion is found. Once a node with a high matching degree is found, the emotion posture information associated with the node will be called. These emotion posture information include action posture emotion and voice posture emotion. Action posture emotion involves the performance of body posture, amplitude and frequency of movement, such as the tilt angle of the head, the swinging mode of the arm, etc. Voice posture emotion is related to the characteristics of the voice tone, rhythm, volume, such as whether it is a cheerful tone or a low tone. In this way, the action posture emotion and voice posture emotion obtained from the feedback emotion map together constitute the emotion posture sequence, which provides a basis for the subsequent posture adjustment and speech synthesis of the digital human.

[0092] When using speech gesture emotions to rhythmically annotate the adaptive response text, it is first necessary to clarify the various characteristic information contained in the speech gesture emotions. The speech gesture emotions cover multiple aspects of characteristics such as intonation, rhythm, volume, and stress. In terms of intonation, if the speech gesture emotions are positive and cheerful, then when the adaptive response text is rhythmically annotated, the corresponding rising tone will be annotated to make the synthesized speech sound more pleasant. For example, in some sentences expressing happiness, the tone will be appropriately raised at the end of the sentence. In terms of rhythm, it will also vary depending on the speech gesture emotions. For example, when the speech gesture emotions are more urgent, the rhythm annotation of the adaptive response text may reflect a faster rhythm and a relatively short pause between words; while if it is a soothing speech gesture emotion, the rhythm will be correspondingly slower and the pause time will be longer. Volume is also an important factor. If the speech gesture emotions emphasize strong emotional expression, a larger volume will be annotated; conversely, for milder emotions, the volume annotation will be relatively small. The annotation of stress is also based on the speech gesture emotions. If a certain emotion is emphasized in the speech gesture emotion, the corresponding words in the adapted response text will be marked as the stress position. By marking the adapted response text in terms of intonation, rhythm, volume and stress in accordance with the speech gesture emotion, a prosodic annotation sequence is finally obtained. This sequence will guide the subsequent speech synthesis process so that the synthesized adapted response speech can accurately convey the corresponding emotion.

[0093] Embodiment 2 is based on the same inventive concept as the digital human posture adjustment method combined with image recognition in the above embodiment. Figure 2 As shown, the present application provides a digital human posture adjustment device combined with image recognition, and the device and method embodiments in the present application are based on the same inventive concept. The device includes:

[0094] The data collection module 10 is used to start data collection on the interactive object after the interactive object is allowed to interact with the target digital human.

[0095] The real-time interactive emotion acquisition module 20 is used to synchronously perform cumulative emotion recognition during the process of collecting facial image data of the interactive object to obtain real-time interactive emotions.

[0096] The interaction demand summary acquisition module 30 is used to perform interaction demand analysis based on the text feature data after completing the text feature data collection of the interaction object to obtain an interaction demand summary.

[0097] The adapted response text generation module 40 is used to generate an adapted response text according to the interactive demand summary.

[0098] The sequence output module 50 is used to fuse and analyze the adapted response text and the real-time interactive emotion, and output an emotion gesture sequence and a rhythm annotation sequence.

[0099] The dynamic synthesis playing module 60 is used to perform synchronous dynamic synthesis playing of the adaptive response voice according to the rhythmic annotation sequence and the adaptive response text during the process of adjusting the posture of the target digital human using the emotional posture sequence.

[0100] Furthermore, the real-time interactive emotion acquisition module 20 also includes:

[0101] An update interval preset unit, wherein the update interval preset unit is used to preset an emotion recognition update interval.

[0102] An image stream interval synchronization unit is used for temporarily storing a multi-frame image stream obtained by collecting facial image data of the interactive object using a camera, and for synchronizing the multi-frame image stream interval to a cumulative emotion recognition model using the emotion recognition update interval as the image data update interval, and using the emotion recognition update interval as the transmission period.

[0103] A real-time interactive emotion acquisition unit is used for the cumulative emotion recognition model to capture and update the trend of emotion changes based on the multi-frame image stream uploaded at intervals to obtain the real-time interactive emotion.

[0104] Furthermore, the interactive demand summary acquisition module 30 further includes:

[0105] The verification interaction emotion acquisition unit is used for temporarily storing the audio segments obtained by the interactive voice data collection process of the interactive object, taking the emotion recognition update interval as the voice data update interval, and taking the emotion recognition update interval as the analysis period to analyze the interval emotion change trend of the audio segments to obtain the verification interaction emotion.

[0106] A credible verification unit, wherein the credible verification unit is used to use the verification interactive emotion to perform credible verification of the real-time interactive emotion.

[0107] A text feature data acquisition unit is used to obtain the text feature data by recognizing and converting the interactive voice data after completing the collection of the interactive voice data of the interactive object.

[0108] An interactive demand summary unit is configured to perform interactive demand analysis based on the text feature data to obtain the interactive demand summary.

[0109] Furthermore, the image stream interval synchronization unit also includes:

[0110] The cumulative emotion recognition model construction unit is used for the cumulative emotion recognition model including a cascaded facial feature recognition layer, an emotional state updating layer and an emotional trend extraction layer.

[0111] A training model unit is used to interactively obtain a sample emotion feature set, and use the sample emotion feature set as training data to train a facial feature recognition model pre-built based on a convolutional neural network.

[0112] A facial feature recognition layer construction unit is used to complete the construction of the facial feature recognition layer by synchronizing the trained and optimized facial feature recognition model to the facial feature recognition layer.

[0113] The emotional state update layer construction unit is used to construct an emotional feature selection gate by adopting a predefined emotional feature screening rule, and complete the construction of the emotional state update layer by cascading the emotional feature selection gate and the emotional state integration gate.

[0114] The emotion trend extraction layer construction unit is used to introduce a multi-dimensional emotion fluctuation threshold to construct an emotion mutation detection unit, and complete the construction of the emotion trend extraction layer by cascading the emotion mutation detection unit and the emotion output regulation unit.

[0115] Furthermore, the real-time interactive emotion acquisition unit further includes:

[0116] The first emotion feature sequence acquisition unit is used for the cumulative emotion recognition model to perform emotion feature recognition of multiple image frames in the first multi-frame image stream via the facial feature recognition layer after receiving the first multi-frame image stream to obtain a first emotion feature sequence.

[0117] The main emotion feature sequence acquisition unit is used for the emotion state update layer to filter out background emotions through the emotion feature selection gate after receiving the first emotion feature sequence to obtain the main emotion feature sequence.

[0118] The first group of emotion change feature acquisition units is used for the emotion state integration gate to extract and connect the time-series emotions of the main emotion feature sequence based on the emotion type to obtain the first group of emotion change features.

[0119] The first group of emotion mutation node acquisition units is used for the emotion trend extraction layer to traverse the first group of emotion change features to identify emotion mutation nodes using the emotion mutation detection unit after receiving the first group of emotion change features, so as to obtain the first group of emotion mutation nodes.

[0120] The first emotion trend vector acquisition unit is used by the emotion output adjustment unit to smoothly optimize the first group of emotion change features according to the first group of emotion mutation nodes to obtain the first emotion trend vector.

[0121] The emotion trend vector acquisition unit is used for, by analogy, capturing the emotion change trend based on the multiple frame image streams uploaded at intervals by the cumulative emotion recognition model to obtain multiple emotion trend vectors.

[0122] An emotion trend vector analysis unit is used to fuse and analyze the multiple emotion trend vectors to obtain the real-time interactive emotion.

[0123] Furthermore, the emotion trend vector analysis unit also includes:

[0124] A dominant emotion trend acquisition unit is configured to aggregate the plurality of emotion trend vectors based on a sliding time window, and then capture the dominant emotion of the aggregation result to obtain K dominant emotion trends.

[0125] The emotion association weight acquisition unit is used to perform attenuation weighted assignment of the K dominant emotion trends according to the emotion capture time to obtain K emotion association weights.

[0126] A comprehensive emotion trend acquisition unit is used to perform weighted summation on the K dominant emotion trends using the K emotion association weights to obtain a comprehensive emotion trend.

[0127] The emotion trend weighted summation unit is used to perform weighted summation of the comprehensive emotion trend by using preset sample emotion weight distribution to obtain the real-time interactive emotion.

[0128] Furthermore, the emotion trend vector analysis unit also includes:

[0129] The interaction obtains multiple sample interaction emotion nodes and multiple sample feedback emotion nodes.

[0130] A mapping relationship between the multiple sample interaction emotion nodes and the multiple sample feedback emotion nodes is created, and a feedback emotion graph is constructed based on the mapping relationship.

[0131] A plurality of groups of sample emotion gestures of the plurality of sample feedback emotion nodes are called, and the plurality of groups of sample emotion gestures are stored in the feedback emotion graph.

[0132] The real-time interactive emotion is used to traverse the feedback emotion map, and the emotion gesture sequence is called to obtain the emotion gesture sequence, wherein the emotion gesture sequence is composed of action gesture emotion and voice gesture emotion.

[0133] The adapted response text is rhythmically annotated using the voice gesture emotion to obtain the rhythmic annotation sequence.

[0134] Furthermore, the sequence output module 50 further includes:

[0135] An emotion node acquisition unit is used to interactively obtain a plurality of sample interaction emotion nodes and a plurality of sample feedback emotion nodes.

[0136] A feedback emotion map construction unit is used to create a mapping relationship between the multiple sample interaction emotion nodes and the multiple sample feedback emotion nodes, and to construct a feedback emotion map based on the mapping relationship.

[0137] A sample emotion gesture calling unit is used to call multiple groups of sample emotion gestures of the multiple sample feedback emotion nodes, and store the multiple groups of sample emotion gestures in the feedback emotion graph.

[0138] An emotion gesture sequence calling unit is used to traverse the feedback emotion map using the real-time interactive emotion, and call to obtain the emotion gesture sequence, wherein the emotion gesture sequence is composed of action gesture emotions and voice gesture emotions.

[0139] A prosody annotation unit is used to perform prosody annotation on the adapted response text using the voice gesture emotion to obtain the prosody annotation sequence.

[0140] It should be noted that the above-mentioned sequence of the embodiments of the present application is only for description and does not represent the advantages and disadvantages of the embodiments. And the above-mentioned specific embodiments of this specification are described. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0141] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application should be included in the protection scope of the present application.

[0142] This specification and the drawings are merely exemplary illustrations of the present application and are deemed to cover any and all modifications, variations, combinations or equivalents within the scope of the present application. Obviously, a person skilled in the art may make various modifications and variations to the present application without departing from the scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the present application and its equivalents, the present application intends to include these modifications and variations.

Claims

1. A digital human posture adjustment method combined with image recognition, characterized in that: The method comprises: After the interactive object is allowed to interact with the target digital person, the data collection behavior of the interactive object is started; In the process of collecting facial image data of the interactive object, cumulative emotion recognition is synchronously performed to obtain real-time interactive emotions; After the text feature data of the interactive object is collected, the interactive demand is analyzed based on the text feature data to obtain an interactive demand summary; Generate an adapted response text according to the interactive appeal summary; Fusion analysis of the adapted response text and the real-time interactive emotions, outputting an emotion posture sequence and a rhythm annotation sequence; In the process of adjusting the posture of the target digital human by using the emotional posture sequence, synchronous dynamic synthesis and playback of the adaptive response voice are performed according to the rhythmic annotation sequence and the adaptive response text; During the process of collecting facial image data of the interactive object, cumulative emotion recognition is synchronously performed to obtain real-time interactive emotions, and the method includes: Preset emotion recognition update interval; In the process of using a camera to collect facial image data of the interactive object, the camera temporarily stores the multi-frame image stream obtained by collecting with the emotion recognition update interval as the image data update interval, and synchronizes the multi-frame image stream interval to the cumulative emotion recognition model with the emotion recognition update interval as the transmission period; The cumulative emotion recognition model captures and updates the emotion change trend based on the multi-frame image stream uploaded at intervals to obtain the real-time interactive emotion; The method comprises: The cumulative emotion recognition model includes a cascaded facial feature recognition layer, an emotion state update layer, and an emotion trend extraction layer; Interactively obtain a sample emotion feature set, and use the sample emotion feature set as training data to train a facial feature recognition model pre-built based on a convolutional neural network; The facial feature recognition layer is constructed by synchronizing the trained and optimized facial feature recognition model to the facial feature recognition layer; Using predefined emotion feature screening rules to construct an emotion feature selection gate, and completing the construction of the emotion state update layer by cascading the emotion feature selection gate and the emotion state integration gate; A multi-dimensional emotion fluctuation threshold is introduced to construct an emotion mutation detection unit, and the emotion mutation detection unit and the emotion output regulation unit are cascaded to complete the construction of the emotion trend extraction layer; The adapted response text and the real-time interactive emotion are integrated and analyzed, and an emotion posture sequence and a rhythm annotation sequence are output. The method includes: Interactively obtain multiple sample interaction emotion nodes and multiple sample feedback emotion nodes; Creating a mapping relationship between the plurality of sample interaction emotion nodes and the plurality of sample feedback emotion nodes, and constructing a feedback emotion graph based on the mapping relationship; Calling multiple groups of sample emotion gestures of the multiple sample feedback emotion nodes, and storing the multiple groups of sample emotion gestures into the feedback emotion graph; The real-time interactive emotion is used to traverse the feedback emotion map, and the emotion posture sequence is called to obtain the emotion posture sequence, wherein the emotion posture sequence is composed of action posture emotion and voice posture emotion; The adapted response text is rhythmically annotated using the voice gesture emotion to obtain the rhythmic annotation sequence.

2. The digital human posture adjustment method combined with image recognition as claimed in claim 1, characterized in that: After the text feature data of the interactive object is collected, the interactive demand is parsed based on the text feature data to obtain an interactive demand summary. Before that, the method includes: During the interactive voice data collection process for the interactive object, the audio segments obtained by the collection are temporarily stored with the emotion recognition update interval as the voice data update interval, and the emotion recognition update interval is used as the analysis period to analyze the interval emotion change trend of the audio segments to obtain the effective interactive emotion; Using the verification interactive emotion to perform credible verification of the real-time interactive emotion; After completing the collection of the interactive voice data of the interactive object, obtaining the text feature data by recognizing and converting the interactive voice data; Interaction demand analysis is performed based on the text feature data to obtain the interaction demand summary.

3. The digital human posture adjustment method combined with image recognition as claimed in claim 1, characterized in that: The cumulative emotion recognition model captures and updates the emotion change trend based on the multi-frame image stream uploaded at intervals to obtain the real-time interactive emotion. The method includes: After receiving the first multi-frame image stream, the cumulative emotion recognition model performs emotion feature recognition of multiple image frames in the first multi-frame image stream via the facial feature recognition layer to obtain a first emotion feature sequence; After receiving the first emotion feature sequence, the emotion state update layer filters out background emotions through the emotion feature selection gate to obtain a main emotion feature sequence; The emotional state integration gate extracts and connects the time-series emotions of the main emotional feature sequence based on the emotional type to obtain a first group of emotional change features; After receiving the first group of emotion change features, the emotion trend extraction layer uses the emotion mutation detection unit to traverse the first group of emotion change features to identify emotion mutation nodes, and obtain a first group of emotion mutation nodes; The emotion output adjustment unit performs smooth optimization on the first group of emotion change features according to the first group of emotion mutation nodes to obtain a first emotion trend vector; By analogy, the cumulative emotion recognition model captures the emotion change trend based on the multiple frame image streams uploaded at intervals to obtain multiple emotion trend vectors; The multiple emotion trend vectors are integrated and analyzed to obtain the real-time interactive emotion.

4. The method for adjusting the digital human posture in combination with image recognition as claimed in claim 3, characterized in that: The method comprises: fusing and analyzing the multiple emotion trend vectors to obtain the real-time interactive emotion. After aggregating the multiple emotion trend vectors based on the sliding time window, capturing the dominant emotion of the aggregation results to obtain K dominant emotion trends; Perform attenuation weighted assignment of the K dominant emotion trends according to the emotion capture time to obtain K emotion association weights; Using the K emotion association weights to perform weighted summation on the K dominant emotion trends to obtain a comprehensive emotion trend; The preset sample emotion weight distribution is used to perform weighted summation of the comprehensive emotion trend to obtain the real-time interactive emotion.

5. A digital human posture adjustment device combined with image recognition, used to implement a digital human posture adjustment method combined with image recognition as claimed in any one of claims 1 to 4, characterized in that: The device comprises: A data collection module, wherein the data collection module is used to start data collection on the interactive object after the interactive object is allowed to interact with the target digital human; A real-time interactive emotion acquisition module, wherein the real-time interactive emotion acquisition module is used to synchronously perform cumulative emotion recognition during the process of collecting facial image data of the interactive object to obtain real-time interactive emotions; An interaction demand summary acquisition module, wherein after completing the collection of text feature data of the interaction object, the interaction demand is analyzed based on the text feature data to obtain an interaction demand summary; An adaptation response text generation module, the adaptation response text generation module is used to generate an adaptation response text according to the interaction demand summary; A sequence output module, the sequence output module is used to fuse and analyze the adapted response text and the real-time interactive emotion, and output an emotion posture sequence and a rhythm annotation sequence; A dynamic synthesis and playback module is used to perform synchronous dynamic synthesis and playback of the adaptive response voice according to the rhythmic annotation sequence and the adaptive response text during the process of using the emotional posture sequence to adjust the posture of the target digital human.

Citation Information

Patent Citations

  • Emotion recognition method and device, computer readable storage medium and equipment

    CN113408503A

  • AI digital human interaction method and system based on emotion recognition

    CN118519538A