Digitalized interaction method and system based on emotion recognition

Through the digital interaction method based on emotional recognition, users' emotional tags are identified in real time and corresponding interactive scenarios are implemented, the problem of lack of emotional interaction in existing digital cemeteries and memorial methods is solved, and a more humane remote memorial experience is achieved.

CN120067769AActive Publication Date: 2025-05-30北京十三陵绿都憩园管理有限公司 +1

Patent Information

Application Number
CN202510525268.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-05-30
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

The existing digital cemeteries and memorial methods lack emotional interaction and cannot meet people's deep emotional needs.

Method used

Using a digital interaction method based on emotion recognition, we use the collection and preprocessing of historical multimodal data, train deep learning models, identify users' emotional tags in real time, and execute corresponding preset digital interaction scenarios.

Benefits of technology

It realizes emotional interaction on digital cemeteries and memorial platforms, provides a more humane remote memorial experience, and meets the deep emotional needs of users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067769A_ABST
    Figure CN120067769A_ABST
Patent Text Reader

Abstract

The invention discloses a digital interaction method and system based on emotion recognition, and belongs to the technical field of digital interaction.Preprocessed historical multi-modal data and corresponding emotion labels serve as training data, a deep learning model is trained, an emotion recognition model is obtained, and in the process of interaction with a user, the emotion recognition model is obtained. And the emotion recognition model is adopted to recognize the preprocessed real-time multi-modal data, a real-time emotion label corresponding to the user is determined, and finally a preset digital interaction scene corresponding to the real-time emotion label is executed, so that interaction between a digital cemetery and a digital sacrifice can be effectively realized, and the user experience is improved. And an interactive commemorative scene is automatically recommended, so that more humanized remote reviewing experience is brought.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of digital interaction, and particularly relates to a digital interaction method and system based on emotion recognition. Background Art

[0002] Digital cemeteries are virtual cemeteries created using modern technology. Through an online platform, information about the deceased, memorial items, etc. are displayed to achieve remote remembrance. Online mourning platforms provide various virtual memorial ways, such as laying flowers, leaving messages, playing memorial videos, etc., allowing users to express their grief anytime and anywhere. The combination of the two breaks through the limitations of time and space, facilitating relatives and friends to cherish the memory of the deceased. At the same time, it saves land resources and conforms to the concept of green environmental protection. The integration of emotion recognition technology makes the interaction more user-friendly, meets the deep emotional needs of users, and leads a new trend in modern funerals. With the development of technology, digital cemeteries and digital memorials have gradually become new ways for people to remember the deceased and express their grief. However, existing digital cemeteries and memorial methods often lack emotional interaction and cannot meet people's deep emotional needs. Summary of the Invention

[0003] The present invention provides a digital interaction method and system based on emotion recognition to solve the problem that existing digital cemeteries and memorial methods often lack interaction.

[0004] On the one hand, the present invention provides a digital interaction method based on emotion recognition, including: Collecting historical multi-modal data, and after preprocessing the historical multi-modal data, obtaining the preprocessed historical multi-modal data; wherein, the historical multi-modal data includes video data and audio data; Showing the historical multi-modal data to a staff member so that the staff member inputs an emotion label corresponding to the historical multi-modal data; Using the preprocessed historical multi-modal data and its corresponding emotion label as training data to train a deep learning model to obtain an emotion recognition model; During the interaction with a user, collecting the user's real-time multi-modal data in real time, and preprocessing the real-time multi-modal data to obtain the preprocessed real-time multi-modal data; Using the emotion recognition model to recognize the preprocessed real-time multi-modal data, determining a real-time emotion label corresponding to the user, and executing a preset digital interaction scenario corresponding to the real-time emotion label to achieve digital interaction based on emotion recognition.

[0005] Further, collecting historical multi-modal data, and after preprocessing the historical multi-modal data, obtaining the preprocessed historical multi-modal data, includes: Based on a preset time period, collect the video data and audio data of the user within any one time period to obtain the historical multi-modal data corresponding to a single time period; Collect the historical multi-modal data within multiple time periods, and after preprocessing the historical multi-modal data, obtain the preprocessed historical multi-modal data.

[0006] Further, after preprocessing the historical multi-modal data, obtain the preprocessed historical multi-modal data, including: Perform sampling processing on the video data in the historical multi-modal data based on a preset data sampling frequency to obtain multiple video frames corresponding to the historical multi-modal data; Extract the speech features in the audio data of the historical multi-modal data, and construct a speech feature data matrix based on the speech features; Take the multiple video frames corresponding to the historical multi-modal data and the speech feature data matrix together as the preprocessed historical multi-modal data.

[0007] Further, extract the speech features in the audio data of the historical multi-modal data, and construct a speech feature data matrix based on the speech features, including: Extract the speech features in the audio data of the historical multi-modal data, and form a feature vector with the speech features; Obtain the importance of each feature in the feature vector based on the trained LightGBM model, and sort the features in descending order according to the importance; Obtain the average value of the importance corresponding to the speech features, filter out the speech features with importance lower than the average value of the importance, use the sequential forward algorithm to select the optimal feature subset, obtain the target speech features, and form a feature vector with the target speech features.

[0008] Further, use the preprocessed historical multi-modal data and its corresponding emotion labels as training data to train a deep learning model to obtain an emotion recognition model, including: Construct a first deep learning model and a second deep learning model; Initialize the hyperparameters of the deep learning model to be trained to obtain multiple hyperparameter individuals; wherein, the deep learning model to be trained is the first deep learning model or the second deep learning model; According to the preprocessed historical multi-modal data and its corresponding emotion labels, obtain the fitness of each hyperparameter individual, and based on the fitness of each hyperparameter individual, determine the current optimal hyperparameter individual in the current training process; Based on the current optimal hyperparameter individual, successively train multiple hyperparameter individuals using a time - adaptive local search strategy, a space - adaptive local search strategy, and a global jump search strategy; After the total number of training times reaches the preset maximum number of training times, determine the target optimal hyperparameter individual, and use the target optimal hyperparameter individual as the final hyperparameter of the deep learning model to be trained, obtaining the trained deep learning model to be trained; After training both the first deep learning model and the second deep learning model, obtain the trained first deep learning model and the trained second deep learning model; Remove the classification output layer of the trained first deep learning model to obtain a first feature extraction model; remove the classification output layer of the trained second deep learning model to obtain a second feature extraction model; Construct a third deep learning model, and based on the pre - processed historical multi - modal data, extract first feature data through the first feature extraction model and extract second feature data through the second feature extraction model; Based on the first feature data, the second feature data, and the corresponding sentiment labels, use the third deep learning model as the deep learning model to be trained for training to obtain a feature recognition model; Based on the first feature extraction model, the second feature extraction model, and the feature recognition model, obtain a sentiment recognition model.

[0009] Furthermore, the time - adaptive local search strategy includes: Based on the current number of training times, generate a time - adaptive control factor as:

[0010] where, represents the time - adaptive control factor, represents pi, represents the current number of training times, represents the expectation, represents the variance, represents the variance control parameter, and exp represents the natural constant with base e; Based on the current optimal hyperparameter individual, perform time - adaptive local search on the hyperparameter individual using the time - adaptive control factor to obtain the first target hyperparameter individual as:

[0011] where, represents the j - th hyperparameter individual in the t - th training process, and j = 1, 2, …, K, where K represents the total number of hyperparameter individuals, Denote the j-th first target hyperparameter individual, Denote the first learning factor, Denote the current optimal hyperparameter individual.

[0012] Furthermore, the spatial adaptive local search strategy includes: For the i -th first target hyperparameter individual, determine the i -1 first target hyperparameter individuals as the first adjacent individuals, and determine the i +1 first target hyperparameter individuals as the second adjacent individuals; among them, for the 1st first target hyperparameter individual, its first adjacent individual is set to other random first target hyperparameter individuals; for the K-th first target hyperparameter individual, its second adjacent individual is set to other random first target hyperparameter individuals; According to the first adjacent individuals and the second adjacent individuals, obtain the spatial adaptive control factor corresponding to the i -th first target hyperparameter individual as:

[0013]

[0014]

[0015] where Denote the spatial adaptive control factor corresponding to the i -th first target hyperparameter individual, Denote the first proportionality factor, Denote the second proportionality factor, Denote the current number of training times, T denotes the preset maximum number of training times, Denote the i -th first target hyperparameter individual and the Euclidean distance between it and its corresponding first adjacent individual, Denote the i -th first target hyperparameter individual and the Euclidean distance between it and its corresponding second adjacent individual; According to the spatial adaptive control factor corresponding to the i -th first target hyperparameter individual, perform spatial adaptive local search on the i -th first target hyperparameter individual, and obtain the second target parameter individual as:

[0016] where Denote the t -th first target hyperparameter individual during the i -th training process, Denote thei The first neighboring individual of the first target hyperparameter individual, denotes the i second neighboring individual of the first target hyperparameter individual, represents a first random number between (0, 1), denotes the i th second target parameter individual.

[0017] Furthermore, the global jump search strategy includes: According to the current number of training times, obtain the adaptive global jump probability as:

[0018] where, represents the adaptive global jump probability, sin represents the sine function, denotes the preset maximum number of training times, t represents the current number of training times, represents pi; Obtain the fitness values corresponding to all second target parameter individuals, and sort the second target parameter individuals in descending order of fitness. According to the sorted second target parameter individuals, obtain the information fusion positions corresponding to all second target parameter individuals as:

[0019]

[0020] where, denotes the t th second target parameter individual in the n th training process after sorting, represents the weighting coefficient of the second target parameter individual , represents the information fusion position corresponding to all second target parameter individuals, denotes the total number corresponding to the second target parameter individuals; For any second target parameter individual, determine the jump action corresponding to the second target parameter individual according to the adaptive global jump probability; among them, the jump action includes the need to jump or not to jump; When the jump action corresponding to the second target parameter individual is the need to jump, then perform a global jump search on the second target parameter individual according to the information fusion position, and obtain the global jump search position as:

[0021] where, denotes the t th second target parameter individual in the m th training process, Represents the second target parameter individual The corresponding global jump search position Represents the second learning factor Represents the third learning factor Represents the second random number between (0, 1) Represents the third random number between (0, 1) Represents except for the second target parameter individual The random second target parameter individual other than that Judge whether the fitness of the global jump search position is greater than the fitness of the second target parameter individual. If so, use the global jump search position as the third target parameter individual; otherwise, use the original second target parameter individual as the third target parameter individual. Among them, the third target parameter individual is the second target parameter individual after the global jump search

[0022] Furthermore, use the emotion recognition model to recognize the preprocessed real-time multimodal data, determine the real-time emotion label corresponding to the user, and execute the preset digital interaction scenario corresponding to the real-time emotion label to realize digital interaction based on emotion recognition, including: Use the emotion recognition model to recognize the preprocessed real-time multimodal data, and determine the real-time emotion label corresponding to the user Query the preset digital interaction scenario corresponding to the real-time emotion label. Among them, the preset digital interaction scenario includes but is not limited to customized music, generating interactive guidance videos or generating interactive guidance voices Execute the preset digital interaction scenario corresponding to the real-time emotion label to realize digital interaction based on emotion recognition

[0023] On the other hand, the present invention provides a digital interaction system based on emotion recognition, including: a historical data collection module, a label acquisition module, a data relationship learning module, a real-time data collection module, and a digital interaction module The historical data collection module is used to collect historical multimodal data, and after preprocessing the historical multimodal data, obtain the preprocessed historical multimodal data. Among them, the historical multimodal data includes video data and audio data The label acquisition module is used to display the historical multimodal data to the staff, so that the staff can input the emotion label corresponding to the historical multimodal data The data relationship learning module is used to use the preprocessed historical multimodal data and its corresponding emotion label as training data to train a deep learning model to obtain an emotion recognition model The real-time data acquisition module is used to collect the real-time multi-modal data of the user in real time during the interaction with the user, and preprocess the real-time multi-modal data to obtain the preprocessed real-time multi-modal data; The digital interaction module is used to identify the preprocessed real-time multi-modal data by using the emotion recognition model, determine the real-time emotion label corresponding to the user, and execute the preset digital interaction scenario corresponding to the real-time emotion label to realize digital interaction based on emotion recognition.

[0024] A digital interaction method and system based on emotion recognition provided by the present invention use the preprocessed historical multi-modal data and their corresponding emotion labels as training data to train a deep learning model to obtain an emotion recognition model, and during the interaction with the user, use the emotion recognition model to identify the preprocessed real-time multi-modal data, determine the real-time emotion label corresponding to the user, and finally execute the preset digital interaction scenario corresponding to the real-time emotion label, which can effectively realize the interaction in digital cemeteries and digital memorial ceremonies, automatically recommend interactive memorial scenarios, and thus bring a more user-friendly remote remembrance experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The accompanying drawings here are incorporated into the specification and form a part of the specification, showing the embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention.

[0026] Figure 1 It is a flowchart of a digital interaction method based on emotion recognition provided by an embodiment of the present invention.

[0027] Figure 2 It is a schematic structural diagram of a digital interaction system based on emotion recognition provided by an embodiment of the present invention.

[0028] Through the above-mentioned accompanying drawings, the clear embodiments of the present invention have been shown, and there will be more detailed descriptions hereinafter. These drawings and text descriptions are not intended to limit the scope of the inventive concept in any way, but to illustrate the concept of the present invention to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0029] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.

[0030] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0031] As Figure 1 shown, the present invention provides a digital interaction method based on emotion recognition, including: S101. Collect historical multimodal data, and after preprocessing the historical multimodal data, obtain the preprocessed historical multimodal data; wherein, the historical multimodal data includes video data and audio data; The core idea of multimodal emotion recognition is to use the learned modal feature representations and a heuristic fusion strategy to analyze the overall emotional tendency of heterogeneous modal sources. By integrating emotional information from different types of modalities, it is possible to predict and classify emotions more comprehensively and accurately, thereby improving the accuracy and robustness of the entire emotion recognition system. Therefore, the embodiments of the present invention use historical multimodal data and the corresponding emotional labels of the historical multimodal data for model training, so as to effectively improve emotion recognition in the digital memorial process and ultimately enhance the interactive experience.

[0032] It should be noted that the video data needs to include the user's face. During the preprocessing process, the face can also be tracked and cropped to obtain the face region data of the user during the digital interaction process, improving the accuracy of data recognition.

[0033] S102. Display the historical multimodal data to the staff so that the staff can input the emotional labels corresponding to the historical multimodal data; Before emotion recognition, first display the historical multimodal data to the staff. Professional staff can give emotional labels, so in the subsequent process, the emotional labels of users can be automatically recognized, and the preset digital interaction scenarios corresponding to the emotional labels can be executed to assist the staff in improving work efficiency and achieving work automation. At the same time, it can also enhance the more user-friendly remote memorial experience.

[0034] S103. Use the preprocessed historical multimodal data and its corresponding emotional labels as training data to train a deep learning model to obtain an emotion recognition model; Optionally, the deep learning model provided by the embodiments of the present invention can be a single model or a combination of multiple models, which is set according to actual needs. After training the deep learning model with historical multimodal data and its corresponding emotional labels as training data, an emotion recognition model for assisting the staff in recognizing the emotions of users is obtained.

[0035] S104. During the interaction with the user (such as when the user is conducting digital memorial ceremonies), collect the real-time multi-modal data of the user in real time, and preprocess the real-time multi-modal data to obtain the preprocessed real-time multi-modal data; To ensure accurate data recognition, the processing process of real-time multi-modal data is the same as that of historical multi-modal data, so that the emotion recognition model can achieve data recognition.

[0036] S105. Use the emotion recognition model to recognize the preprocessed real-time multi-modal data, determine the real-time emotion label corresponding to the user, and execute the preset digital interaction scenario corresponding to the real-time emotion label to achieve digital interaction based on emotion recognition.

[0037] The present invention combines emotion recognition technology with digital cemeteries and memorial platforms, providing a brand-new interaction method, meeting the diverse needs of modern people for funeral activities, and having broad application prospects and social value.

[0038] In the embodiment of the present invention, after collecting historical multi-modal data and preprocessing the historical multi-modal data, the preprocessed historical multi-modal data is obtained, including: Based on a preset time period, collect the video data and audio data of the user within any one time period to obtain the historical multi-modal data corresponding to a single time period; Collect the historical multi-modal data within multiple time periods, and after preprocessing the historical multi-modal data, obtain the preprocessed historical multi-modal data.

[0039] It should be noted that the above two modal data are only preferred examples of the embodiment of the present invention, and other modal data can also be used for recognition to improve the accuracy of interaction.

[0040] In the embodiment of the present invention, after preprocessing the historical multi-modal data, the preprocessed historical multi-modal data is obtained, including: Perform sampling processing on the video data in the historical multi-modal data based on a preset data sampling frequency to obtain multiple video frames corresponding to the historical multi-modal data; Extract the speech features in the audio data of the historical multi-modal data, and construct a speech feature data matrix based on the speech features; Use the multiple video frames corresponding to the historical multi-modal data and the speech feature data matrix together as the preprocessed historical multi-modal data.

[0041] Optionally, the video frames can also be enhanced to improve the data recognition accuracy.

[0042] In an embodiment of the present invention, extracting speech features from the audio data in the historical multi-modal data, and constructing a speech feature data matrix based on the speech features, includes: Extracting speech features from the audio data in the historical multi-modal data, and forming a feature vector with the speech features; Obtaining the importance of each feature in the feature vector based on a trained LightGBM model, and sorting the features in descending order according to the importance; Obtaining the average value of the importance corresponding to the speech features, filtering out the speech features with importance lower than the average value of the importance, selecting an optimal feature subset by using a sequential forward algorithm to obtain target speech features, and forming a feature vector with the target speech features.

[0043] Optionally, the speech features may include Feature 1 to Feature 809; wherein, the specific features of Feature 1-8 are: the mean, variance, maximum value, and minimum value of the short-time energy and its first-order difference; the specific features of Feature 9-14 are: the mean, variance, and maximum value of the sound intensity and its first-order difference; Feature 15 is the average speech rate; the specific features of Feature 16-23 are: the mean, variance, maximum value, and minimum value of the fundamental frequency and its first-order difference; the specific features of Feature 24-53 are: the mean, variance, maximum value, minimum value, and median of the first, second, and third formant frequencies and their first-order differences; the specific features of Feature 54-137 are: the mean, variance, maximum value, minimum value, median, range, and sum of the 1st-12th order Mel-frequency cepstral coefficients MFCC; the specific features of Feature 138-221 are: the mean, variance, maximum value, minimum value, median, range, and sum of the 1st-12th order gamma-frequency cepstral coefficients GFCC; the specific features of Feature 222-305 are: the mean, variance, maximum value, minimum value, median, range, and sum of the 1st-12th order Bark-frequency cepstral coefficients BFCC; the specific features of Feature 306-389 are: the mean, variance, maximum value, minimum value, median, range, and sum of the 1st-12th order linear prediction coefficients LPC; the specific features of Feature 390-473 are: the mean, variance, maximum value, minimum value, median, range, and sum of the 1st-12th order linear prediction cepstral coefficients LPCC; the specific features of Feature 474-557 are: the mean, variance, maximum value, minimum value, median, range, and sum of the 1st-12th order normalized gamma chirp cepstral coefficients NGCC; the specific features of Feature 558-641 are: the mean, variance, maximum value, minimum value, median, range, and sum of the 1st-12th order magnitude-based spectral root cepstral coefficients MSRCC; the specific features of Feature 642-725 are: the mean, variance, maximum value, minimum value, median, range, and sum of the 1st-12th order phase-based spectral root cepstral coefficients PSRCC; the specific features of Feature 726-809 are: the mean, variance, maximum value, minimum value, median, range, and sum of the 1st-12th order linear frequency cepstral coefficients LFCC.

[0044] It should be noted that the voice features in the historical multi-modal data can also be directly extracted, and the voice features can be directly constructed into a voice feature data matrix, which can also achieve the recognition of voice data.

[0045] In the embodiment of the present invention, the preprocessed historical multi-modal data and its corresponding emotion labels are used as training data to train a deep learning model to obtain an emotion recognition model, including: Construct a first deep learning model and a second deep learning model; Optionally, both the first deep learning model and the second deep learning model can be set as convolutional neural networks.

[0046] Initialize the hyperparameters of the deep learning model to be trained to obtain multiple hyperparameter individuals; wherein, the deep learning model to be trained is the first deep learning model or the second deep learning model; Optionally, a random initialization or a chaotic mapping initialization strategy can be used to initialize the hyperparameters of the deep learning model to be trained to obtain multiple hyperparameter individuals; wherein, a hyperparameter individual includes all the hyperparameters to be trained of the deep learning model to be trained, and all the hyperparameters to be trained of the deep learning model to be trained can be all its hyperparameters or some of its hyperparameters.

[0047] According to the preprocessed historical multi-modal data and its corresponding emotion labels, obtain the fitness of each hyperparameter individual, and based on the fitness of each hyperparameter individual, determine the current optimal hyperparameter individual (i.e., the hyperparameter individual with the largest fitness) in the current training process; For example, for the first deep learning model, after applying the hyperparameters in the hyperparameter individual to the first deep learning model, the video frame can be used as the input, the corresponding emotion label can be used as the expected output, obtain the loss function value corresponding to the hyperparameter individual, and after taking the negative of the loss function value, obtain the fitness corresponding to the hyperparameter individual.

[0048] For the second deep learning model, after applying the hyperparameters in the hyperparameter individual to the second deep learning model, the voice feature data matrix can be used as the input, the corresponding emotion label can be used as the expected output, obtain the loss function value corresponding to the hyperparameter individual, and after taking the negative of the loss function value, obtain the fitness corresponding to the hyperparameter individual.

[0049] Based on the current optimal hyperparameter individual, train multiple hyperparameter individuals in turn using a time-adaptive local search strategy, a space-adaptive local search strategy, and a global jump search strategy; After the total number of training times reaches the preset maximum number of training times, determine the target optimal hyperparameter individual (that is, determine the individual with the largest fitness from the second target parameter individuals in the last training process), and use the target optimal hyperparameter individual as the final hyperparameters of the deep learning model to be trained, so as to obtain the deep learning model to be trained after training; After training both the first deep learning model and the second deep learning model, obtain the trained first deep learning model and the trained second deep learning model; Remove the classification output layer of the trained first deep learning model to obtain a first feature extraction model (for a convolutional neural network, the Softmax classification output layer needs to be removed); remove the classification output layer of the trained second deep learning model to obtain a second feature extraction model; Construct a third deep learning model (such as constructing a third deep learning model using a long short-term memory network or a BP neural network), and according to the preprocessed historical multimodal data, extract first feature data through the first feature extraction model and extract second feature data through the second feature extraction model; According to the first feature data, the second feature data and the corresponding sentiment labels, use the third deep learning model as the deep learning model to be trained for training to obtain a feature recognition model; During the process of training the third deep learning model, splice the first feature data and the second feature data into a vector, and use the spliced vector as the input, and use the corresponding sentiment label as the expected output to obtain the fitness. It should be noted that it is necessary to splice the first feature data and the second feature data corresponding to all video frames into a vector.

[0050] Optionally, in order to simplify the recognition process, the first deep learning model may not remove the classification output layer. After splicing the classification labels corresponding to all video frames in the single historical multimodal data and the second feature data into a vector, then train the third deep learning model, and the feature recognition model can also be obtained.

[0051] According to the first feature extraction model, the second feature extraction model and the feature recognition model, obtain a sentiment recognition model.

[0052] In the prior art, during the hyperparameter training process, it is easy to fall into local optimality, and the training accuracy and effect are both poor. Therefore, the present invention provides a time-adaptive local search strategy, a space-adaptive local search strategy and a global jump search strategy for joint search, so as to solve the problems existing in the prior art, and finally be able to better realize the recognition of sentiment labels and accurately complete the data recognition tasks specified by the staff.

[0053] In the embodiment of the present invention, the time - adaptive local search strategy includes: Generate a time - adaptive control factor based on the current number of training times as:

[0054] wherein, represents the time - adaptive control factor, represents pi, represents the current number of training times, represents the expectation (which can be set to 0, for example), represents the variance (which can be set to 1, for example), represents the variance control parameter, and exp represents the natural constant with the natural constant e as the base; Based on the current optimal hyperparameter individual, use the time - adaptive control factor to perform time - adaptive local search on the hyperparameter individual, and obtain the first target hyperparameter individual as:

[0055] wherein, represents the j - th hyperparameter individual in the t - th training process, and j = 1, 2, …, K, where K represents the total number of hyperparameter individuals, represents the j - th first target hyperparameter individual, represents the first learning factor, represents the current optimal hyperparameter individual.

[0056] The time - adaptive local search strategy provided by the embodiment of the present invention can enable the hyperparameter individual to generate a certain fluctuation while learning the optimal position information, can increase the ability to jump out of the local optimum to a certain extent, and will not affect the convergence of the algorithm in the later stage of the algorithm.

[0057] In the embodiment of the present invention, the space - adaptive local search strategy includes: For the i -th first target hyperparameter individual, determine the i -1 - th first target hyperparameter individual as the first adjacent individual, and determine the i +1 - th first target hyperparameter individual as the second adjacent individual; wherein, for the 1 - st first target hyperparameter individual, its first adjacent individual is set to other random first target hyperparameter individuals; for the K - th first target hyperparameter individual, its second adjacent individual is set to other random first target hyperparameter individuals; According to the first adjacent individual and the second adjacent individual, obtain the space - adaptive control factor corresponding to the i -th first target hyperparameter individual as:

[0058]

[0059]

[0060] Among them, represents the spatial adaptive control factor corresponding to the i th first target hyperparameter individual, represents the first scale factor, represents the second scale factor, represents the current number of training times, and T represents the preset maximum number of training times, represents the i th Euclidean distance between the th first target hyperparameter individual and its corresponding first adjacent individual, i represents the th Euclidean distance between the i th first target hyperparameter individual and its corresponding second adjacent individual; i Perform spatial adaptive local search on the

[0061] According to the spatial adaptive control factor corresponding to the th first target hyperparameter individual, and obtain the second target parameter individual as: t where i represents the th first target hyperparameter individual in the i th training process, represents the first adjacent individual of the i th first target hyperparameter individual, represents the second adjacent individual of the th first target hyperparameter individual, i represents the

[0062] The spatial adaptive local search strategy provided by the embodiments of the present invention enables hyperparameter individuals to search based on their positions in space, try to search unfamiliar areas, avoid search collisions, and at the same time improve the ability to find the global optimum. As the algorithm converges, the positions of all hyperparameter individuals gradually gather in the solution space, and at this time, the convergence accuracy can be increased.

[0063] In the embodiments of the present invention, the global jump search strategy includes: Obtain the adaptive global jump probability according to the current number of training times as:

[0064] Among them, represents the adaptive global jump probability, sin represents the sine function, represents the preset maximum number of training times, t represents the current training times, represents pi; Obtain the fitness values corresponding to all the second target parameter individuals, and arrange the second target parameter individuals in descending order of fitness. According to the arranged second target parameter individuals, the information fusion positions corresponding to all the second target parameter individuals are:

[0065]

[0066] Among them, represents the t th second target parameter individual after arrangement in the n th training process, represents the second target parameter individual 's weighting coefficient, represents the information fusion position corresponding to all the second target parameter individuals, represents the total number corresponding to the second target parameter individuals; For any one of the second target parameter individuals, determine the jump action corresponding to the second target parameter individual according to the adaptive global jump probability; among them, the jump action includes the need to jump or not to jump; For example, based on the adaptive global jump probability, a random number between (0, 1) can be generated. When this random number is less than the adaptive global jump probability, then it is necessary to jump, otherwise there is no need to jump.

[0067] When the jump action corresponding to the second target parameter individual is the need to jump, then perform a global jump search on the second target parameter individual according to the information fusion position, and the global jump search position is obtained as:

[0068] Among them, represents the t th second target parameter individual in the m th training process, represents the second target parameter individual 's corresponding global jump search position, represents the second learning factor, represents the third learning factor, represents the second random number between (0, 1), represents the third random number between (0, 1), represents except for the second target parameter individual a random second target parameter individual other than Determine whether the fitness of the global jump search position is greater than the fitness of the second target parameter individual. If so, use the global jump search position as the third target parameter individual; otherwise, use the original second target parameter individual as the third target parameter individual. Here, the third target parameter individual is the second target parameter individual after the global jump search.

[0069] The global jump search strategy provided by the embodiments of the present invention can effectively improve the global search ability of the algorithm. Throughout the early and middle stages of the algorithm, it has a strong global search ability, and can also provide a certain global search ability in the later stage, enabling the hyperparameter individuals to search in areas far from the aggregation of hyperparameter individuals (i.e., far from the information fusion position) while performing global search. Combining the above two strategies ensures that the algorithm can effectively jump out of the local optimum, and a greedy algorithm is introduced for search control to ensure the convergence speed of the algorithm.

[0070] Optionally, instead of using a greedy algorithm for search control, an annealing simulation algorithm can be used for search control, which can also achieve the control of the algorithm convergence speed. After each search, the individuals can be processed to prevent them from going out of bounds, ensuring that the hyperparameters are within the valid range.

[0071] By providing the above hyperparameter optimization algorithm, the embodiments of the present invention can effectively improve the data learning relationship, more accurately achieve the established data recognition tasks of the staff, and ultimately improve the experience effect and accuracy of digital interaction.

[0072] In the embodiments of the present invention, the emotion recognition model is used to recognize the preprocessed real-time multimodal data, determine the real-time emotion label corresponding to the user, and execute the preset digital interaction scenario corresponding to the real-time emotion label to implement digital interaction based on emotion recognition, including: Use the emotion recognition model to recognize the preprocessed real-time multimodal data and determine the real-time emotion label corresponding to the user; Query the preset digital interaction scenario corresponding to the real-time emotion label. Here, the preset digital interaction scenario includes but is not limited to customized music, generating interactive guidance videos, or generating interactive guidance voices; For example, the preset digital interaction scenario can be playing customized music, generating virtual memorial ceremonies (videos or voices), thus bringing a more user-friendly remote remembrance experience. However, it should be noted that the above several preset digital interaction scenarios are only examples of the embodiments of the present invention, and other preset digital interaction scenarios can also be used to enrich the digital interaction experience.

[0073] Execute the preset digital interaction scenario corresponding to the real-time emotion tag to achieve digital interaction based on emotion recognition.

[0074] A digital interaction method based on emotion recognition provided by the present invention trains a deep learning model with the preprocessed historical multimodal data and its corresponding emotion tags as training data to obtain an emotion recognition model. During the interaction with the user, the emotion recognition model is used to recognize the preprocessed real-time multimodal data to determine the real-time emotion tag corresponding to the user. Finally, the preset digital interaction scenario corresponding to the real-time emotion tag is executed, which can effectively realize the interaction in digital cemeteries and digital memorial ceremonies, automatically recommend interactive memorial scenarios, and thus bring a more user-friendly remote remembrance experience.

[0075] As Figure 2 shown, an embodiment of the present invention provides a digital interaction system based on emotion recognition, including: a historical data acquisition module 201, a tag acquisition module 202, a data relationship learning module 203, a real-time data acquisition module 204, and a digital interaction module 205; The historical data acquisition module 201 is used to acquire historical multimodal data and, after preprocessing the historical multimodal data, obtain the preprocessed historical multimodal data; wherein, the historical multimodal data includes video data and audio data; The tag acquisition module 202 is used to display the historical multimodal data to the staff so that the staff can input the emotion tag corresponding to the historical multimodal data; The data relationship learning module 203 is used to train a deep learning model with the preprocessed historical multimodal data and its corresponding emotion tags as training data to obtain an emotion recognition model; The real-time data acquisition module 204 is used to, during the interaction with the user, acquire the real-time multimodal data of the user in real time and preprocess the real-time multimodal data to obtain the preprocessed real-time multimodal data; The digital interaction module 205 is used to use the emotion recognition model to recognize the preprocessed real-time multimodal data, determine the real-time emotion tag corresponding to the user, and execute the preset digital interaction scenario corresponding to the real-time emotion tag to achieve digital interaction based on emotion recognition.

[0076] A digital interaction system based on emotion recognition provided by an embodiment of the present invention can execute the above method technical solution, and its principle and beneficial effects are similar, so they will not be elaborated here.

[0077] Other embodiments of the present invention will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. The present invention is intended to cover any variations, uses, or adaptations of the invention following the general principles of the invention and including known common general knowledge or conventional technical means in the technical field not disclosed by the present invention. It should be understood that the present invention is not limited to the precise structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.

Claims

1. A digital interaction method based on emotion recognition, characterized in that: include: Collecting historical multimodal data, and preprocessing the historical multimodal data to obtain preprocessed historical multimodal data; wherein the historical multimodal data includes video data and audio data; Displaying the historical multimodal data to a staff member so that the staff member inputs an emotion label corresponding to the historical multimodal data; Using the pre-processed historical multimodal data and its corresponding emotion labels as training data, training the deep learning model to obtain an emotion recognition model; In the process of interacting with the user, real-time multimodal data of the user is collected in real time, and the real-time multimodal data is preprocessed to obtain the real-time multimodal data after preprocessing; The emotion recognition model is used to recognize the real-time multimodal data after the preprocessing, determine the real-time emotion tag corresponding to the user, and execute the preset digital interaction scene corresponding to the real-time emotion tag to realize digital interaction based on emotion recognition.

2. The digital interaction method based on emotion recognition according to claim 1, characterized in that: Collecting historical multimodal data and preprocessing the historical multimodal data to obtain preprocessed historical multimodal data includes: Based on a preset time period, the video data and audio data of the user in any time period are collected to obtain historical multimodal data corresponding to a single time period; Historical multimodal data within a plurality of time periods are collected, and the historical multimodal data are preprocessed to obtain preprocessed historical multimodal data.

3. The digital interaction method based on emotion recognition according to claim 2, characterized in that: After preprocessing the historical multimodal data, the preprocessed historical multimodal data is obtained, including: Sampling the video data in the historical multimodal data based on a preset data sampling frequency to obtain a plurality of video frames corresponding to the historical multimodal data; Extracting speech features from the audio data in the historical multimodal data, and constructing a speech feature data matrix based on the speech features; The multiple video frames and the speech feature data matrix corresponding to the historical multimodal data are taken together as the historical multimodal data after preprocessing.

4. The digital interaction method based on emotion recognition according to claim 3 is characterized in that: Extracting speech features from the audio data in the historical multimodal data, and constructing a speech feature data matrix based on the speech features, including: Extracting speech features from the audio data in the historical multimodal data, and composing the speech features into feature vectors; Based on the trained LightGBM model, the importance of each feature in the feature vector is obtained, and the features are sorted in descending order by importance; The importance average value corresponding to the speech features is obtained, and the speech features whose importance is lower than the importance average value are filtered out, and the optimal feature subset is selected by using a sequence forward algorithm to obtain the target speech features, and the target speech features are combined into a feature vector.

5. The digital interaction method based on emotion recognition according to claim 3 or 4, characterized in that: The deep learning model is trained using the pre-processed historical multimodal data and its corresponding emotion labels as training data to obtain an emotion recognition model, including: Constructing a first deep learning model and a second deep learning model; Initializing the hyperparameters of the deep learning model to be trained to obtain multiple hyperparameter individuals; wherein the deep learning model to be trained is the first deep learning model or the second deep learning model; According to the pre-processed historical multimodal data and its corresponding emotion label, the fitness of each of the hyperparameter individuals is obtained, and based on the fitness of each of the hyperparameter individuals, the current optimal hyperparameter individual in the current training process is determined; Based on the current optimal hyperparameter individual, a time-adaptive local search strategy, a space-adaptive local search strategy, and a global jump search strategy are sequentially adopted to train multiple hyperparameter individuals; After the total number of training times reaches the preset maximum number of training times, the target optimal hyperparameter individual is determined, and the target optimal hyperparameter individual is used as the final hyperparameter of the deep learning model to be trained, so as to obtain the deep learning model to be trained after training; After training the first deep learning model and the second deep learning model, a trained first deep learning model and a trained second deep learning model are obtained; The classification output layer of the first deep learning model after training is removed to obtain a first feature extraction model; the classification output layer of the second deep learning model after training is removed to obtain a second feature extraction model; Constructing a third deep learning model, extracting first feature data through the first feature extraction model and extracting second feature data through the second feature extraction model according to the historical multimodal data after the preprocessing; According to the first feature data, the second feature data and the corresponding emotion label, the third deep learning model is trained as the deep learning model to be trained to obtain a feature recognition model; An emotion recognition model is obtained according to the first feature extraction model, the second feature extraction model and the feature recognition model.

6. The digital interaction method based on emotion recognition according to claim 5, characterized in that: The time-adaptive local search strategy includes: Based on the current number of training times, the adaptive control factor of the generation time is: in, represents the time adaptive control factor, represents pi, Indicates the current number of training times. Express expectations, represents the variance, represents the variance control parameter, exp represents the natural constant with the natural constant e as the base; Based on the current optimal hyperparameter individual, a time-adaptive control factor is used to perform a time-adaptive local search on the hyperparameter individual, and the first target hyperparameter individual is obtained as follows: in, represents the jth hyperparameter individual in the tth training process, and j=1,2,…,K, K represents the total number of hyperparameter individuals, represents the jth first target hyperparameter individual, represents the first learning factor, Represents the current optimal hyperparameter individual.

7. The digital interaction method based on emotion recognition according to claim 6, characterized in that: The spatially adaptive local search strategy comprises: For the i The first target hyperparameter individuals are determined i -1 first target hyperparameter individual is the first adjacent individual, determine the i +1 first target hyperparameter individual is the second adjacent individual; among them, for the first first target hyperparameter individual, its first adjacent individual is set to other random first target hyperparameter individuals; for the Kth first target hyperparameter individual, its second adjacent individual is set to other random first target hyperparameter individuals; According to the first adjacent individual and the second adjacent individual, obtain the i The spatial adaptive control factor corresponding to the first target hyperparameter individual is: in, Indicates i The spatial adaptive control factor corresponding to the first target hyperparameter individual, represents the first scale factor, represents the second scale factor, represents the current number of training times, T represents the preset maximum number of training times, Indicates i The Euclidean distance between the first target hyperparameter individual and its corresponding first adjacent individual, Indicates i The Euclidean distance between the first target hyperparameter individual and its corresponding second adjacent individual; According to the said i The spatial adaptive control factor corresponding to the first target hyperparameter individual is i The first target hyperparameter individuals are spatially adaptively searched locally, and the second target parameter individuals are obtained as follows: in, Indicates t During the training process i The first target hyperparameter individuals, Indicates i The first neighboring individuals of the first target hyperparameter individuals, Indicates i The second neighboring individuals of the first target hyperparameter individual, represents the first random number between (0,1), Indicates i The second target parameter individual.

8. The digital interaction method based on emotion recognition according to claim 7, characterized in that: The global jump search strategy includes: According to the current training times, the adaptive global jump probability is obtained as: in, represents the adaptive global jump probability, sin represents the sine function, represents the preset maximum number of training times, t represents the current number of training times, represents pi; Obtain the fitness corresponding to all second target parameter individuals, and arrange the second target parameter individuals in descending order of fitness. According to the arranged second target parameter individuals, obtain the information fusion position corresponding to all second target parameter individuals as follows: in, Indicates t During the training n The second target parameter individuals after permutation, Represents the second target parameter individual The weighting coefficient of Indicates the information fusion position corresponding to all second target parameter individuals, Indicates the total number of individuals corresponding to the second target parameter; For any second target parameter individual, a jump action corresponding to the second target parameter individual is determined according to the adaptive global jump probability; wherein the jump action includes whether jumping is required or not; When the jump action corresponding to the second target parameter individual is that a jump is required, a global jump search is performed on the second target parameter individual according to the information fusion position, and the global jump search position is obtained as: in, Indicates t During the training m The second target parameter individuals, Represents the second target parameter individual The corresponding global jump search position, represents the second learning factor, represents the third learning factor, represents the second random number between (0,1), represents the third random number between (0,1), Indicates that except for the second target parameter individual A random second target parameter individual other than ; Determine whether the fitness of the global jump search position is greater than the fitness of the second target parameter individual. If so, use the global jump search position as the third target parameter individual, otherwise use the original second target parameter individual as the third target parameter individual; wherein the third target parameter individual is the second target parameter individual after the global jump search.

9. The digital interaction method based on emotion recognition according to claim 8, characterized in that: The emotion recognition model is used to recognize the real-time multimodal data after the preprocessing, determine the real-time emotion tag corresponding to the user, and execute the preset digital interaction scene corresponding to the real-time emotion tag to realize the digital interaction based on emotion recognition, including: Using the emotion recognition model to recognize the real-time multimodal data after the preprocessing, and determining the real-time emotion label corresponding to the user; Querying a preset digital interactive scene corresponding to a real-time emotion tag; wherein the preset digital interactive scene includes but is not limited to customized music, generating an interactive guidance video, or generating an interactive guidance voice; Execute the preset digital interactive scene corresponding to the real-time emotion tag to realize digital interaction based on emotion recognition.

10. A digital interactive system based on emotion recognition, characterized in that: include: Historical data collection module, label acquisition module, data relationship learning module, real-time data collection module and digital interaction module; The historical data acquisition module is used to collect historical multimodal data and preprocess the historical multimodal data to obtain preprocessed historical multimodal data; wherein the historical multimodal data includes video data and audio data; The label acquisition module is used to display the historical multimodal data to the staff so that the staff can input the emotion label corresponding to the historical multimodal data; The data relationship learning module is used to train the deep learning model using the pre-processed historical multimodal data and its corresponding emotion labels as training data to obtain an emotion recognition model; The real-time data acquisition module is used to collect the real-time multimodal data of the user in real time during the interaction with the user, and pre-process the real-time multimodal data to obtain the real-time multimodal data after pre-processing; The digital interaction module is used to use the emotion recognition model to identify the real-time multimodal data after preprocessing, determine the real-time emotion tag corresponding to the user, and execute the preset digital interaction scene corresponding to the real-time emotion tag to realize digital interaction based on emotion recognition.

Citation Information

Patent Citations

  • Multi-modal emotion recognition method and system based on context awareness

    CN113947702A

  • Multi-modal emotion recognition method and device

    CN116935277A

  • Multi-modal perception fusion emotion recognition method and robot emotion interaction method

    CN117994622A

  • Multi-modal emotion recognition model training method, multi-modal emotion recognition method and equipment

    CN118430043A

  • Emotion recognition method, pre-training method for emotion recognition model, and electronic device

    WO2025044647A1

Cited By

  • Psychological state assessment method and system based on deep learning

    CN120340872A