A multi-mode anthropomorphic ecological reconstruction system and method thereof
Through the combination of multi-head attention model and multi-modal emotional neural network, the synchronization problem and data loss problem of multi-modal personified ecological reconstruction system during cross-modal data integration is solved, the emotion recognition and response ability of the digital personality system is improved, and the immersion and satisfaction of users are enhanced.
Patent Information
- Application Number
- CN202411284305.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-13
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2044-09-13
AI Technical Summary
The existing multimodal personified ecological reconstruction system has synchronization problems, data loss and interpretation difficulties when integrating cross-modal data, which affects the naturalness and interaction quality of digital personality.
Through the combination of data source recognition, multi-head attention model and multi-modal emotional neural network, visual, auditory and text data are preprocessed and integrated to generate emotional responses, and dynamically adjust digital personality to adapt to different situations through the generation of adversarial networks.
It improves the emotion recognition and response capabilities of the digital personality system, enhances the immersion and satisfaction of users, realizes more comprehensive data utilization and analysis, and optimizes the comprehensiveness and depth of information processing.
Smart Images

Figure CN119150911B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of anthropomorphic ecological management, and more specifically, to a multi-mode anthropomorphic ecological reconstruction system and a method thereof. Background Art
[0002] Anthropic Ecosystem is an integrated system that covers technologies such as artificial intelligence, machine learning, natural language processing, and computer vision to create and maintain digital entities that can simulate human behavior, emotions, thinking, and interactions. This concept not only includes a single technology application, but refers to a broad, multi-technology integrated environment with the goal of creating a digital life form that dynamically interacts with human society.
[0003] This is the core part of the anthropomorphic ecology, which refers to virtual human images with personality and emotions created through artificial intelligence technology. These digital personalities can simulate the thinking, behavior, language and emotional responses of real humans to varying degrees.
[0004] Deficiencies of existing technologies: In multimodal anthropomorphic ecological reconstruction, cross-modal data integration involves how to effectively merge data from different perceptual channels (such as vision, hearing, language, etc.) and generate consistent and coherent digital personality responses on this basis. Existing systems often have limitations in processing and fusing data from different modalities, such as synchronization problems, data loss, and difficulty in interpretation. These problems affect the naturalness and interaction quality of the digital personality, and thus affect the quality of anthropomorphic ecological reconstruction. Summary of the invention
[0005] In order to overcome the above-mentioned defects of the prior art, the present invention provides a multi-mode anthropomorphic ecological reconstruction system and method thereof to solve the problem of poor interaction quality in the management process of the multi-mode anthropomorphic ecology proposed in the above-mentioned background technology.
[0006] To achieve the above object, the present invention provides the following technical solutions:
[0007] A multi-mode anthropomorphic ecological reconstruction method comprises the following steps:
[0008] Identify data sources and determine the data types and sources of multimodal data of anthropomorphic ecology. Multimodal data includes visual data, auditory data, and text data.
[0009] Preprocess visual data, auditory data, and text data, and input the preprocessed multimodal data into the multi-head attention model for data integration;
[0010] Use a multimodal emotional neural network to analyze the emotional state of the integrated data and generate corresponding emotional responses. Generate digital personalities in different situations based on the generative adversarial network, and analyze the matching between the generated digital personality's response and the user's current situation.
[0011] Acquire the matching adaptation information generated in the process of matching the digital personality response with the user's current situation, determine the matching situation and generate a signal based on the matching adaptation information, and perform corresponding mode strategy management based on the generated signal.
[0012] In a preferred embodiment, the visual data, the auditory data and the text data are preprocessed, and the specific process is as follows:
[0013] The images collected in the visual data are cropped and scaled, and the noise in the images is removed using Gaussian blur technology;
[0014] Use the Sobel algorithm to extract image edges and perform shape analysis and object recognition on subsequent images;
[0015] After image edge extraction, a scale-invariant feature transformation algorithm is used to detect key points and descriptors in the image as features of visual data;
[0016] Using Mel-frequency cepstrum coefficients on auditory data to determine the characteristics of sound signals in the auditory data, and calculating the power spectrum or logarithmic power spectrum of the sound signal to analyze the frequency components;
[0017] Process text data to remove stop words, punctuation marks, and non-standard characters, and perform text standardization;
[0018] The bag-of-words model is used to extract features from text data, identify all words from the text data, and build a vocabulary index. For each document, the number of times the words in the vocabulary appear is counted, and these counts are used as elements of the feature vector.
[0019] In a preferred embodiment, after the image edge is extracted, a scale-invariant feature transformation algorithm is used to detect key points and descriptors in the image and use them as features of visual data. The specific process is as follows:
[0020] Construct a Gaussian difference scale space, apply Gaussian blur of different scales to the original image, and calculate the difference between adjacent Gaussian blurred images;
[0021] For each pixel in the Gaussian difference scale space, compare the size of the pixel with that of the adjacent points. If it is a local maximum or minimum value, it is regarded as a key point.
[0022] Based on the direction and magnitude of the local image gradient, one or more directions are assigned to each key point;
[0023] In the neighborhood around each key point, a window of corresponding size is selected according to the scale of the key point, the window is divided into small sub-blocks, and the gradient histograms in 8 directions are calculated in each sub-block to form a descriptor.
[0024] In a preferred embodiment, a multimodal emotional neural network is used to analyze the emotional state of the integrated data and generate corresponding emotional responses. A digital personality in different situations is generated according to a generative adversarial network. The specific process is as follows:
[0025] The data features from different modalities are used to train a multimodal emotion neural network to identify the user's overall emotional state. The supervised learning method is used for training based on the labeled emotion dataset.
[0026] Conduct emotional database construction, i.e. create a database containing emotional responses, including predefined text, sound or facial expressions;
[0027] Design and classify emotional responses according to emotion categories, each category includes text response, sound, expression, select or generate corresponding emotional responses from the emotion database based on real-time emotion recognition results, and use generative adversarial networks to create user-specific emotional responses in real time.
[0028] In a preferred embodiment, the matching adaptation information generated in the process of matching the digital personality response with the user's current situation is obtained, and the specific process is as follows:
[0029] The matching adaptation information includes multimodal context information and context adaptation information;
[0030] The multimodal scenario information includes a multimodal data adaptation index, and the scenario adaptation information includes a multi-level scenario adaptation index;
[0031] The obtained multi-mode data adaptation index and multi-level scenario adaptation index are jointly calculated to obtain the pattern matching stability coefficient.
[0032] In a preferred embodiment, the matching situation is determined according to the matching adaptation information and a signal is generated, and the corresponding mode strategy management is performed according to the generated signal. The specific process is as follows:
[0033] Obtaining a pattern matching stability coefficient obtained by the joint calculation, and comparing the pattern matching stability coefficient with a reconstruction adjustment threshold;
[0034] If the pattern matching stability coefficient is greater than or equal to the reconstruction adjustment threshold, an anthropomorphic ecological stability signal is generated and no additional management measures are taken;
[0035] If the pattern matching stability coefficient is less than the reconstruction adjustment threshold, a pattern reconstruction signal is generated to perform reconstruction adjustment of the anthropomorphic ecology.
[0036] A multi-mode anthropomorphic ecological reconstruction system, used to implement the above-mentioned multi-mode anthropomorphic ecological reconstruction method, comprising:
[0037] The data acquisition and processing module is used to identify the data source, determine the data type and source of the multimodal data of the anthropomorphic ecology, and pre-process the visual data, auditory data and text data, and input the pre-processed multimodal data into the multi-head attention model for data integration;
[0038] The scene matching analysis module is used to analyze the emotional state of the integrated data using a multimodal emotional neural network and generate corresponding emotional responses. It generates digital personalities in different situations based on the generative adversarial network and analyzes the matching between the generated digital personality's response and the user's current situation.
[0039] The mode control module is used to obtain the matching adaptation information generated in the process of matching the digital personality response with the user's current situation, determine the matching situation and generate a signal based on the matching adaptation information, and perform corresponding mode strategy management based on the generated signal.
[0040] Technical effects and advantages of the present invention:
[0041] The present invention improves the emotion recognition and response capabilities of the digital personality system by efficiently integrating visual, auditory and text data, and applying a multi-head attention model and a multimodal emotion neural network. Then, a generative adversarial network is used to dynamically generate and adjust the digital personality to adapt to different interactive scenarios, thereby enhancing the user's immersion and satisfaction. This cross-modal data fusion achieves more comprehensive data utilization and analysis, and also optimizes the comprehensiveness and depth of information processing, so that complex user inputs can be understood more accurately when constructing an anthropomorphic ecosystem. In addition, the reaction strategy is automatically adjusted according to the matching adaptation information generated during the real-time matching analysis process, and the operation efficiency and intelligence level are further improved by automatically generating signals and automatically managing the mode strategy accordingly, and the need for manual intervention is significantly reduced. In general, the present invention provides a highly integrated and intelligent solution that can effectively improve the adaptability of the system, the personalized user experience and the overall user satisfaction. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 The figure is a flow chart of a multi-mode anthropomorphic ecological reconstruction method of the present invention.
[0043] Figure 2 This is a schematic diagram of the structure of a multi-mode anthropomorphic ecological reconstruction system of the present invention. DETAILED DESCRIPTION
[0044] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0045] Example 1: Figure 1 As shown, a multi-model anthropomorphic ecological reconstruction method includes the following steps:
[0046] Identify data sources and determine the data types and sources of multimodal data of anthropomorphic ecology. Multimodal data includes visual data, auditory data, and text data.
[0047] Preprocess visual data, auditory data, and text data, and input the preprocessed multimodal data into the multi-head attention model for data integration;
[0048] Use a multimodal emotional neural network to analyze the emotional state of the integrated data and generate corresponding emotional responses. Generate digital personalities in different situations based on the generative adversarial network, and analyze the matching between the generated digital personality's response and the user's current situation.
[0049] Acquire the matching adaptation information generated in the process of matching the digital personality response with the user's current situation, determine the matching situation and generate a signal based on the matching adaptation information, and perform corresponding mode strategy management based on the generated signal.
[0050] Conduct data source identification, determine the data type and source of the multimodal data of the anthropomorphic ecology, and perform data processing on the acquired multimodal data to collect multimodal data from multiple sensor sources;
[0051] Visual data acquisition is to deploy multiple high-resolution cameras in the user interaction environment to capture the user's facial expressions, body language and overall movements from different angles. For example, cameras can be set up in front, on the side and on the top of the interactive space to obtain a full range of views; edge computing devices are deployed for real-time image analysis to quickly identify and respond to changes in user behavior. For example, GPU-accelerated devices can be used to execute facial expression recognition and action recognition algorithms;
[0052] Auditory data acquisition uses an environmental microphone array to collect the user's voice information, analyze the background sound in the environment, such as other people's conversations and music, understand the scene context, process sound data in real time, quickly extract voice features, and use voice activity detection (VAD) technology to determine the beginning and end of the voice, providing accurate input for speech recognition;
[0053] Text data acquisition is to synchronize users' historical and real-time text data from multiple digital channels (such as email, social media, chat applications, etc.). These text data provide users' language habits, interest preferences and social network information, and provide real-time text input interfaces, such as chat robot windows or smart assistant input boxes, allowing users to communicate directly with the system. At the same time, ensure that text input is seamless and natural, add voice-to-text conversion functions, allow users to choose voice input or direct text input, and perform immediate pre-processing after collecting text data, including language detection, spelling correction and grammar analysis, to ensure data quality and consistency.
[0054] In the reconstruction of anthropomorphic ecology, visual data processing directly affects the system's ability to interpret the user's visual information such as expressions and actions. The specific steps of visual data processing are:
[0055] The images collected in the visual data are cropped and scaled, and the image size and area are adjusted as needed for subsequent processing. Color adjustment is performed to convert the color image into a grayscale image, which is effective enough for certain types of feature extraction. Gaussian blur technology is used to remove noise from the image and improve image quality. The expression of Gaussian blur is: , where (x, y) represents the pixel position currently being processed, and σ is the standard deviation of the Gaussian distribution (or Gaussian function), which determines the smoothness of the filter;
[0056] Use the Sobel or Canny algorithm to extract the image edge, perform shape analysis and object recognition on the subsequent images, and use the Sobel operator to calculate the gradient of the image: the gradient of the image in the horizontal direction (x direction) is: ,in, is the image value at the pixel coordinate (x, y); the gradient in the vertical direction (y direction) is: , extract the image edge according to the horizontal and vertical gradients of the image;
[0057] After image edge extraction, the scale-invariant feature transform (SIFT) algorithm is used to detect key points and descriptors in the image. The specific steps are as follows:
[0058] Construct a Gaussian difference scale space (DoG, Difference of Gaussian) to identify key points at multiple scales. That is, apply Gaussian blur of different scales to the original image, and then calculate the difference between adjacent Gaussian blurred images;
[0059] For each pixel in the DoG space, compare its size with that of adjacent points (including points in space and scale). If it is a local maximum or minimum, it is considered a key point. The DoG function expression is: , where k is a constant representing the ratio of adjacent scales in the scale space. This ratio is used to determine the degree of scale separation between the two Gaussian blur layers used to generate the DoG image, and k is the scale factor of two consecutive Gaussian kernels, which is used to calculate the standard deviation of the two Gaussian kernels;
[0060] The key point positions are accurately determined by fitting a three-dimensional quadratic function, excluding key points with low contrast and key points with strong edge effects;
[0061] Based on the direction and magnitude of the local image gradient, each key point is assigned one or more directions, so that the SIFT descriptor remains invariant to image rotation;
[0062] In the neighborhood around each key point, a window of corresponding size is selected according to the scale of the key point, the window is divided into small sub-blocks, and the gradient histograms in 8 directions are calculated in each sub-block. The gradients are summarized into a descriptor and the descriptors of the key points are used as features of the visual data for subsequent computer vision analysis.
[0063] Processing of auditory data such as speech and audio for subsequent speech recognition, speech emotion analysis, and sound event detection. The following are the specific steps for auditory data processing:
[0064] Use bandpass filters and noise reduction algorithms (such as spectral subtraction) to reduce background noise and unwanted frequency components in the recording, and perform normalization to adjust the amplitude of the sound signal to a consistent volume level. The normalization expression is: , where x(t) is the original sound signal, is the normalized signal;
[0065] Mel-frequency cepstral coefficients (MFCC) are used to determine the common features of sound signals in speech recognition. They can capture the short-term energy and spectral shape of the sound and calculate the power spectrum or logarithmic power spectrum of the sound signal for analyzing the frequency components.
[0066] The calculation steps of Mel frequency cepstral coefficients are:
[0067] First, compute the Fast Fourier Transform (FFT) to obtain the spectrum of the signal and convert the frequency scale to the Mel scale: , where f is the frequency (in Hertz) and m is the Mel scale;
[0068] A Mel filter bank is applied to extract the frequency band energy. The logarithm of the output of each Mel filter is taken, and all logarithmic energy values are subjected to discrete cosine transform (DCT) to obtain the Mel frequency cepstrum coefficients. The mathematical expression of discrete cosine transform is: ,in, is the kth cepstral coefficient, is the logarithmic energy of the nth Mel filter output, N is the number of Mel filters (usually the same as the length of the DCT), K is the number of selected cepstral coefficients, usually K is less than N to keep only the most important cepstral coefficients;
[0069] Use machine learning or deep learning models (such as CNN, RNN) to identify and classify sound events, such as identifying different sound activities or environmental sounds, and use automatic speech recognition technology to convert sound signals into corresponding text information. This usually involves acoustic models (for phoneme recognition) and language models (for building and understanding sentences);
[0070] Mel-frequency cepstral coefficients convert the spectral features of sound signals into the cepstral domain, which is helpful for sound processing and speech recognition. The application of discrete cosine transform not only improves the robustness of the features, but also helps to improve the efficiency and accuracy of subsequent speech recognition models.
[0071] Process text data to remove stop words, punctuation marks, and non-standard characters, and perform text standardization, such as lowercase and stem extraction;
[0072] Use the Bag of Words (BOW) model to extract text data features, identify all unique words (text) from the entire document collection, and build a vocabulary index. For each document, count the number of times the words in the vocabulary appear, and use these counts as elements of the feature vector;
[0073] Use a pre-trained sentiment analysis model, such as a deep learning based classifier (e.g. using LSTM, CNN, etc.), or a rule-based system to input the vectors generated by BOW into the sentiment analysis model;
[0074] If you are using a pre-trained model, you can directly apply the model for prediction. If you are training the model yourself, you can do it in two steps: training and validation. You can use methods such as cross-validation to optimize the model parameters.
[0075] This process converts text into digital signals (through BOW), which are then used for in-depth analysis of sentiment and intent. This approach is suitable for extracting valuable insights from user-generated text content for a variety of applications, such as improving products and services and enhancing user experience.
[0076] Integrate visual, auditory and text data, convert the data of each modality into a unified feature representation, such as vectorized representation, to ensure the consistency and accuracy of information, and use a multi-head attention mechanism to enhance data integration between different modalities. The core formula of multi-head attention is as follows: , , , where Q, K, and V are query, key, and value matrices, respectively, which are derived from the same modality or cross-modality features. , , , is a learnable weight matrix, Is the dimension of the key vector, used to scale the dot product to prevent excessively large internal dot product values. They are all independent attention functions, and the user output is finally passed through merge;
[0077] Apply the output of the multi-head attention model to fuse information from visual, auditory, and textual data, and use appropriate classifiers or regression models to process the fused features according to specific applications (such as sentiment analysis, intent recognition, etc.);
[0078] Through the above steps and formulas, multimodal data integration can effectively integrate information from vision, hearing and text. The multi-head attention mechanism handles this type of task because it can focus on multiple different types of input data at the same time, thereby improving the accuracy and efficiency of processing.
[0079] Use deep sentiment analysis models to analyze the user's emotional state and generate corresponding emotional responses. Combined with emotion recognition algorithms (such as emotional neural networks) and generative adversarial networks (GANs), generate digital personalities that can express complex emotions in different situations. The specific steps are as follows:
[0080] Collect data features from different modalities, such as facial expressions (visual data), voice pitch and rhythm (auditory data), and language vocabulary and style (text data), combine features from all modalities, and train a deep neural network model, such as a multimodal emotion neural network, to identify the user's overall emotional state, using supervised learning methods and training based on annotated emotion datasets;
[0081] Analyze users’ expressions, voice, and text input in real time, and use trained models to identify their emotional states (such as happiness, sadness, anger, surprise, etc.);
[0082] Conduct emotional database construction, that is, create a database containing a variety of emotional responses, which can be predefined text, sound or expression responses;
[0083] Design and classify emotional responses according to common emotion categories (such as happiness, sadness, etc.). Each category includes text responses (such as comforting sentences, encouraging sentences), sounds (such as voice prompts of different emotions, laughter, crying, etc.) and expressions (animated or real-time generated facial expressions). According to the real-time emotion recognition results, select or generate corresponding emotional responses from the emotion database and assign them to the digital personality. Generative adversarial networks (GANs) can be further used to create user-specific emotional responses in real time to update the digital personality.
[0084] Generate new emotional expressions using generative adversarial networks (GANs) to enhance the diversity and authenticity of digital personality responses, optimize the selection of emotional responses through logic or machine learning algorithms, ensure that the digital personality responses match the user's current context, and after matching, integrate the selected emotional responses into the digital personality's behavior to upgrade and reconstruct the digital personality, including adjusting the pitch and speed of the voice, changing facial expressions, and modifying language responses, etc., to fully express the corresponding emotions;
[0085] Acquire matching adaptation information generated in the process of matching the digital personality response with the user's current situation, the matching adaptation information determines whether the digital personality response matches the user's current situation, and can meet the needs of the multi-modal anthropomorphic ecology, the matching adaptation information includes multi-modal situation information and situation adaptation information;
[0086] The multimodal scenario information includes the multimodal data adaptation index and is calibrated as DMS, and the scenario adaptation information includes the multi-level scenario adaptation index and is calibrated as DCC;
[0087] The multimodal data adaptation index in multimodal contextual information is used to evaluate the degree of match between the emotional response generated by the digital personality and the user's current emotional state and interactive context during the multimodal data (multimodal data) interaction process. It is used to integrate data from different sensory inputs, such as vision (facial expressions), hearing (voice pitch and rhythm) and text (verbal emotion), and evaluate the consistency and adaptability of these data in expressing user emotions;
[0088] The multimodal data fit index measures the degree of match between the recognized user emotions and the emotions actually expressed by the user. This includes judging whether the system accurately captures the user's basic emotions (such as happiness, sadness, anger, etc.), evaluating whether the generated emotional response takes into account the specific context of the interaction, such as the context of the conversation, environmental factors, etc., to ensure that the response is consistent with the current situation, and analyzing the effectiveness of data from different sensory channels in the integration process to ensure that these data can support each other during analysis and improve the overall accuracy of sentiment analysis;
[0089] The multimodal data adaptation index can improve the user experience. By ensuring the accuracy and situational adaptability of emotional responses, the multimodal data adaptation index helps the anthropomorphic ecosystem provide a more natural and attractive interactive experience, thereby increasing user satisfaction and engagement. It also provides a quantitative indicator that can be used to evaluate and optimize the performance of emotion recognition and generation algorithms, especially in multimodal data fusion.
[0090] In summary, the multimodal data adaptation index is used to evaluate and optimize the quality and effect of emotional interaction in a multimodal anthropomorphic ecosystem. Through precise emotion and context matching, the multimodal data adaptation index helps to build a more intelligent, sensitive and user-friendly interaction system.
[0091] The multi-mode data adaptation index is obtained as follows:
[0092] Obtain the user emotional data analyzed by the digital personality during the interaction and generate corresponding facial expressions, and extract the generated expression features from the generated expression images , extract facial expression features from user’s facial videos , calculate the facial expression matching degree, the calculation formula is: , where the dot product ⋅ represents the scalar product of two vectors, represents the Euclidean norm of the generated expression feature vector (i.e. the length of the vector), represents the Euclidean norm of the facial expression feature vector;
[0093] Get the time interval between two adjacent sound wave peaks in the interaction, and calculate the energy of the audio signal of all sampling times in the time interval. The calculation expression is: , where x(n) represents the amplitude of the nth sample and N is the total number of samples sampled;
[0094] Collect text inputs from user interactions and the digital personality's text responses to these inputs to obtain sentiment scores for user texts Emotional scores with digital personality responses , calculate the sentiment conformity value: , calculate the multi-mode data adaptation index, the calculation expression is: .
[0095] It should be noted that the generated expression features may include data features such as the position, shape, and dynamic changes of key facial points obtained through visual data analysis; sampling refers to selecting sample points at fixed time intervals in a continuous audio waveform. Assume that there is an audio file with a sampling rate of 44.1kHz and a sampling bit depth of 16 bits, and read the audio signal, which contains 1000 samples. The amplitude value of each sample can be read directly from the audio file (such as through audio editing software or programming interface), and then the above formula is applied to calculate the total energy of the signal; the sentiment score can be obtained through sentiment analysis tools (such as VADER), and the score provided by the sentiment analysis tool ranges from [-1, 1], where -1 represents negative, 1 represents positive, and 0 represents neutral. The complement of the absolute difference between the two sentiment scores is calculated, and TEM is close to 1, indicating high consistency; conversely, if the two differ greatly, TEM is close to 0, indicating low consistency.
[0096] The multi-level situational adaptability index in the situational adaptability information measures the digital personality's ability to adapt to multi-dimensional situational factors in interactions with users. This index provides an in-depth situational adaptability assessment of the digital personality's behavior and responses by integrating multiple levels of analysis. The multi-level situational adaptability index represents how the digital personality understands and adapts to the user's needs and emotional conditions at different times, in different environments, and in various interaction dynamics. This includes the digital personality's memory and adaptation to long-term user behavior, sensitivity to the user's environment, and appropriateness of responses in real-time interactions.
[0097] Time series behavioral adaptability measures how the digital personality adapts to the user's changing behaviors and emotions over time, focuses on the digital personality's long-term learning and memory capabilities, and reflects its response to changes in user behavior patterns; interactive dynamic adaptability measures how the digital personality responds to the user's specific needs in real-time interactions, especially in terms of speed and efficiency in responding to emergencies or specific requests.
[0098] The multi-level contextual adaptability index helps provide a more personalized and compassionate interactive experience by precisely adjusting the behavior of the digital personality to better match the user's actual context and needs. It also helps developers and researchers identify the shortcomings of the digital personality in contextual adaptability, thereby improving product design more effectively.
[0099] By evaluating and optimizing the multi-level situational adaptability of digital personalities, the Multi-Level Situational Adaptability Index promotes the development of more advanced automation and intelligent solutions, especially in service robots, virtual assistants and other artificial intelligence systems, and is used to support decision-making processes, such as optimizing strategy selection and behavior adjustment in complex interaction scenarios.
[0100] The multi-level scenario adaptation index is obtained as follows:
[0101] Collect historical interaction data between the digital personality and the user from the interaction log, obtain the total number of interactions ZT and the number of response interactions T, calculate the response adaptation value: AS=T / ZT, calculate the time series value, and the calculation expression is: , where e is a constant with a value of 2.71, λ is the attenuation coefficient, and AS(t) is the fitness value of the tth interaction;
[0102] Obtain the speed at which the digital personality responds to user input during the real-time interactive response process, determine the response time SJ for each response to user input, and obtain the number of interactions CI with an acceptable response time, calculate the response efficiency value: XL=CI / ZT, obtain the average response efficiency XLavg, and calculate the interactive dynamic adaptation value: , where is the response efficiency value of the jth interaction, and the multi-level scenario adaptation index is calculated. The calculation expression is: .
[0103] It should be noted that λ is the decay coefficient, which is used to adjust the impact of early interactions on the current evaluation. As a positive value, λ is used to adjust the weight of each interaction in the total score. The larger the value, the faster the weight of early interactions decays, which means that the weight of newer interactions in the evaluation will be relatively high. When processing user data, the most recent behavior or reaction is usually considered to better reflect the current user needs and situations than the earlier behavior. Using λ can effectively simulate this time sensitivity.
[0104] The obtained multi-mode data adaptation index and multi-level scenario adaptation index are calculated jointly to obtain the mode matching stability coefficient, which is expressed as: , where is the mode matching stability coefficient, , is the preset proportional coefficient of the multi-mode data adaptation index DMS and the multi-level scenario adaptation index DCC, and , Both are greater than 0.
[0105] It should be noted that the size of the preset proportional coefficient is a specific value obtained by quantifying each parameter. In order to facilitate subsequent comparison, the size of the coefficient depends on the amount of sample data and the preliminary setting of the corresponding preset proportional coefficient for each group of sample data by technical personnel in this field. It is not unique, as long as it does not affect the proportional relationship between the parameter and the quantized value, just like the multi-level scenario adaptation index is proportional to the pattern matching stability coefficient.
[0106] The smaller the multimodal data adaptation index and the smaller the multi-level scenario adaptation index, that is, the smaller the performance value of the pattern matching stability coefficient, the lower the DMS, indicating that the digital personality fails to accurately match the user's actual emotions and behavioral intentions when processing and responding to multimodal input data (such as visual, auditory, and text data), which may cause users to feel that the digital personality's response is unnatural, irrelevant, or inaccurate, thereby reducing user satisfaction and trust;
[0107] A low multimodal data adaptation index means that the digital personality performs poorly in understanding and adapting to the user's specific context (such as time, place, and social environment). This may make users feel that the digital personality cannot understand or ignores their environment, resulting in unsatisfactory interaction or inappropriate response;
[0108] The low value of the multi-level situational adaptation index indicates that, overall, the digital personality shows low adaptability and stability in its interaction with users. Users may experience inconsistent or broken interaction experiences, which may affect user engagement and brand loyalty in the long run.
[0109] Low values of these indicators suggest that developers and designers need to conduct in-depth analysis and optimization of the design and implementation of digital personalities, which may include enhancing the algorithm's emotion recognition capabilities, improving contextual awareness logic, or providing more customized user experience design. Through these improvements, the value of PMSC can be increased, thereby ensuring that digital personalities can better interact with users in a high-quality manner.
[0110] By comprehensively considering the multi-mode data adaptation index and the multi-level scenario adaptation index, the mode matching stability coefficient can more comprehensively evaluate the status of the simulation mode switching process.
[0111] A high multimodal data adaptation index indicates that the digital personality can effectively integrate and respond to input data from different modalities (such as vision, hearing, and text), and accurately understand the user's emotions and behavioral intentions, making the digital personality's response more natural and humane, better meeting the needs of users and improving user satisfaction and engagement;
[0112] A high multi-level situational adaptation index means that the digital personality performs well in understanding and adapting to the user's specific situation (such as time, place, and social environment), which enhances the interaction quality of the digital personality, enabling it to provide appropriate responses in the specific situation of the user, and enhancing the consistency and relevance of the user experience.
[0113] A high pattern matching stability coefficient indicates that the digital personality shows a high degree of adaptability and stability in its interaction with users. Users will experience a consistent and smooth interaction experience, which helps to increase users' trust and reliance on the digital personality, thereby improving overall user loyalty and satisfaction.
[0114] Comparing the generated mode matching stability coefficient with the reconstruction adjustment threshold, generating different control signals, and adjusting the corresponding mode switching control strategy according to the generated control signals;
[0115] After obtaining the pattern matching stability coefficient, the pattern matching stability coefficient is compared with the reconstruction adjustment threshold;
[0116] If the pattern matching stability coefficient is greater than or equal to the reconstruction adjustment threshold, an anthropomorphic ecological stability signal is generated, which indicates that the current digital personality and user interaction mode is effective and stable, and can meet the needs and expectations of users. The digital personality performs well in processing multimodal data and adapting to complex situations. The user experience is coherent and the satisfaction is high. There is no need to immediately adjust the current configuration, algorithm or behavior mode, and it is suitable to maintain the status quo.
[0117] If the mode matching stability coefficient is less than the reconstruction adjustment threshold, a mode reconstruction signal is generated, indicating that the existing interaction mode does not match the user's needs and context, and reconstruction adjustment or optimization of the multi-mode anthropomorphic ecology is required. This may be because the digital personality fails to accurately understand the user's emotions, behavior patterns or environmental factors, or because the user's needs have changed;
[0118] It is necessary to reconstruct, adjust or optimize the multi-mode anthropomorphic ecology. The steps are as follows:
[0119] If the existing model fails to keep up with changes in emotions or anthropomorphic ecology, certain aspects of the digital personality will need to be redesigned or adjusted, such as improving the sentiment analysis model, adjusting the contextual awareness algorithm, or updating the user interface. This signal can trigger the system's self-adjustment program, such as initiating the retraining process of the machine learning model, or introducing human intervention to evaluate and improve the configuration in detail.
[0120] Through such a mechanism, the multi-modal anthropomorphic ecology can be self-monitored and adaptively optimized and rebuilt to ensure that the quality of interaction always meets the actual needs of users, while also helping to maintain and improve the overall user experience. This dynamic adjustment strategy is an indispensable part of modern intelligent systems, especially in application scenarios where user interactions are highly personalized and contextually changeable.
[0121] It should be noted that the setting of the reconstruction adjustment threshold can be determined according to specific scenarios and needs, and is usually adjusted and optimized based on factors such as historical data and real-time data.
[0122] The present invention improves the emotion recognition and response capabilities of the digital personality system by efficiently integrating visual, auditory and text data, and applying a multi-head attention model and a multimodal emotion neural network. Then, a generative adversarial network is used to dynamically generate and adjust the digital personality to adapt to different interactive scenarios, thereby enhancing the user's immersion and satisfaction. This cross-modal data fusion achieves more comprehensive data utilization and analysis, and also optimizes the comprehensiveness and depth of information processing, so that complex user inputs can be understood more accurately when constructing an anthropomorphic ecosystem. In addition, the reaction strategy is automatically adjusted according to the matching adaptation information generated during the real-time matching analysis process, and the operation efficiency and intelligence level are further improved by automatically generating signals and automatically managing the mode strategy accordingly, and the need for manual intervention is significantly reduced. In general, the present invention provides a highly integrated and intelligent solution that can effectively improve the adaptability of the system, the personalized user experience and the overall user satisfaction.
[0123] Embodiment 2: This embodiment is a system embodiment of embodiment 1, which is used to implement a multi-mode anthropomorphic ecological reconstruction method introduced in embodiment 1, such as Figure 2 As shown, specifically including:
[0124] A multi-mode anthropomorphic ecological reconstruction system, used to implement the above-mentioned multi-mode anthropomorphic ecological reconstruction method, comprising:
[0125] The data acquisition and processing module is used to identify the data source, determine the data type and source of the multimodal data of the anthropomorphic ecology, and pre-process the visual data, auditory data and text data, and input the pre-processed multimodal data into the multi-head attention model for data integration;
[0126] The scene matching analysis module is used to analyze the emotional state of the integrated data using a multimodal emotional neural network and generate corresponding emotional responses. It generates digital personalities in different situations based on the generative adversarial network and analyzes the matching between the generated digital personality's response and the user's current situation.
[0127] The mode control module is used to obtain the matching adaptation information generated in the process of matching the digital personality response with the user's current situation, determine the matching situation and generate a signal based on the matching adaptation information, and perform corresponding mode strategy management based on the generated signal.
[0128] The collection, storage, use, processing, transmission, provision, disclosure and handling of user personal information involved in the technical solution of the present invention are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0129] The above formulas are all dimensionless and calculated numerically. Specific dimension removal can be achieved by various means such as standardization, which will not be elaborated here. The formula is a formula for the most recent real situation obtained by collecting a large amount of data and performing software simulation. The preset parameters in the formula are set by technicians in this field according to actual conditions.
[0130] The above embodiments may be implemented in whole or in part by software, hardware, firmware or any other combination thereof. When implemented by software, the above embodiments may be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server or data center to another website, computer, server or data center by wired or wireless (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that contains one or more available media sets. The available medium may be a magnetic medium (e.g., a floppy disk, an ATA hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium. The semiconductor medium may be a solid-state ATA hard disk.
[0131] It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0132] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0133] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0134] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, and may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0135] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0136] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art who is familiar with the present technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A multi-mode anthropomorphic ecological reconstruction method, characterized in that: The steps include: Identify data sources and determine the data types and sources of multimodal data of anthropomorphic ecology. Multimodal data includes visual data, auditory data, and text data. Preprocess visual data, auditory data, and text data, and input the preprocessed multimodal data into the multi-head attention model for data integration; Use a multimodal emotional neural network to analyze the emotional state of the integrated data and generate corresponding emotional responses. Generate digital personalities in different situations based on the generative adversarial network, and analyze the matching between the generated digital personality's response and the user's current situation. Acquire matching adaptation information generated in the process of matching the digital personality response with the user's current situation, determine the matching situation and generate a signal according to the matching adaptation information, and perform corresponding mode strategy management according to the generated signal; Obtain the matching adaptation information generated in the process of matching the digital personality's response with the user's current situation. The specific process is as follows: The matching adaptation information includes multimodal context information and context adaptation information; The multimodal scenario information includes a multimodal data adaptation index, and the scenario adaptation information includes a multi-level scenario adaptation index; The obtained multi-mode data adaptation index and multi-level scenario adaptation index are calculated jointly to obtain the mode matching stability coefficient; The matching stability coefficient is used to select the generation of anthropomorphic ecological stability signals or to make reconstruction adjustments of anthropomorphic ecology; The multi-mode data adaptation index is obtained as follows: Extracting facial expression features from user's face video , calculate the facial expression matching: , where the dot product ⋅ represents the scalar product of two vectors, represents the Euclidean norm of the generated expression feature vector, represents the Euclidean norm of the facial expression feature vector; Get the time interval between two adjacent sound wave peaks in the interaction, and calculate the energy of the audio signal of all sampling times in the time interval. The calculation expression is: , where x(n) represents the amplitude of the nth sample and N is the total number of samples sampled; Collect text inputs from user interactions and the digital personality's text responses to these inputs to obtain sentiment scores for user texts Emotional scores with digital personality responses , calculate the sentiment conformity value: , calculate the multi-mode data adaptation index, the calculation expression is: ; The multi-level scenario adaptation index is obtained as follows: Collect historical interaction data between the digital personality and the user, obtain the total number of interactions ZT and the number of response interactions T, calculate the response fitness value: AS=T / ZT, calculate the time series value, and the calculation expression is: , where e is a constant with a value of 2.71, λ is the attenuation coefficient, and AS(t) is the fitness value of the tth interaction; Obtain the speed at which the digital personality responds to user input during the real-time interactive response process, determine the response time SJ for each response to user input, and obtain the number of interactions CI with a response time within the preset acceptance range, calculate the response efficiency value: XL=CI / ZT, obtain the average response efficiency XLavg, and calculate the interactive dynamic adaptation value: , where is the response efficiency value of the jth interaction, and the multi-level scenario adaptation index is calculated. The calculation expression is: .
2. A multi-mode anthropomorphic ecological reconstruction method according to claim 1, characterized in that: Preprocess the visual data, auditory data, and text data. The specific process is as follows: The images collected in the visual data are cropped and scaled, and the noise in the images is removed using Gaussian blur technology; Use the Sobel algorithm to extract image edges and perform shape analysis and object recognition on subsequent images; After image edge extraction, a scale-invariant feature transformation algorithm is used to detect key points and descriptors in the image as features of visual data; Using Mel-frequency cepstrum coefficients on auditory data to determine the characteristics of sound signals in the auditory data, and calculating the power spectrum or logarithmic power spectrum of the sound signal to analyze the frequency components; Process text data to remove stop words, punctuation marks, and non-standard characters, and perform text standardization; The bag-of-words model is used to extract text data features, identify words from text data, and build a vocabulary index. For each document in the text data, the number of times the words in the vocabulary appear is counted, and these counts are used as elements of the feature vector.
3. A multi-mode anthropomorphic ecological reconstruction method according to claim 2, characterized in that: After image edge extraction, a scale-invariant feature transformation algorithm is used to detect key points and descriptors in the image and use them as features of visual data. The specific process is as follows: Construct a Gaussian difference scale space, apply Gaussian blur of different scales to the original image, and calculate the difference between adjacent Gaussian blurred images; For each pixel in the Gaussian difference scale space, compare the size of the pixel with that of the adjacent points. If it is a local maximum or minimum value, it is regarded as a key point. Based on the direction and magnitude of the local image gradient, one or more directions are assigned to each key point; In the neighborhood around each key point, a window of corresponding size is selected according to the scale of the key point, the window is divided into small sub-blocks, and the gradient histograms in 8 directions are calculated in each sub-block to form a descriptor.
4. A multi-mode anthropomorphic ecological reconstruction method according to claim 3, characterized in that: Use a multimodal emotional neural network to analyze the emotional state of the integrated data and generate corresponding emotional responses. Generate digital personalities in different situations based on the generative adversarial network. The specific process is as follows: The data features from different modalities are used to train a multimodal emotion neural network to identify the user's overall emotional state. The supervised learning method is used for training based on the labeled emotion dataset. Conduct emotional database construction, i.e. create a database containing emotional responses, including predefined text, sound or facial expressions; Design and classify emotional responses according to emotion categories. Each category includes text response, sound, and expression. According to the real-time emotion recognition results, select or generate corresponding emotional responses from the emotion database, and use the generative adversarial network to create the user's emotional response in real time.
5. The multi-mode anthropomorphic ecological reconstruction method according to claim 1, characterized in that: Determine the matching situation and generate a signal based on the matching adaptation information, and perform corresponding mode strategy management based on the generated signal. The specific process is as follows: Obtaining a pattern matching stability coefficient obtained by the joint calculation, and comparing the pattern matching stability coefficient with a reconstruction adjustment threshold; If the pattern matching stability coefficient is greater than or equal to the reconstruction adjustment threshold, an anthropomorphic ecological stability signal is generated and no additional management measures are taken; If the pattern matching stability coefficient is less than the reconstruction adjustment threshold, a pattern reconstruction signal is generated to perform reconstruction adjustment of the anthropomorphic ecology.
6. A multi-mode anthropomorphic ecological reconstruction system, used to implement a multi-mode anthropomorphic ecological reconstruction method according to any one of claims 1 to 5, characterized in that: include: The data acquisition and processing module is used to identify the data source, determine the data type and source of the multimodal data of the anthropomorphic ecology, and pre-process the visual data, auditory data and text data, and input the pre-processed multimodal data into the multi-head attention model for data integration; The scene matching analysis module is used to analyze the emotional state of the integrated data using a multimodal emotional neural network and generate corresponding emotional responses. It generates digital personalities in different situations based on the generative adversarial network and analyzes the matching between the generated digital personality's response and the user's current situation. The mode control module is used to obtain the matching adaptation information generated in the process of matching the digital personality response with the user's current situation, determine the matching situation and generate a signal based on the matching adaptation information, and perform corresponding mode strategy management based on the generated signal.
Citation Information
Patent Citations
Mental health assessment method and system based on digital human
CN117462130A
Emotion recognition method and system based on multiple modes
CN118312857A