Training methods for behavioral decision-making models and adaptive interaction methods for digital humans

Through iterative training and multi-objective optimization strategies of the digital human behavior decision-making model, the problem of insufficient digital human behavior decision-making in MR scenarios was solved, effective integration and dynamic adaptation of multimodal information were achieved, and the user experience was improved.

CN120354176BActive Publication Date: 2025-09-16SHIYOU (BEIJING) TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510833804.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-09-16
Estimated Expiration
2045-06-20

AI Technical Summary

Technical Problem

Existing digital human behavior decision-making models lack the ability to effectively integrate multimodal information and dynamically adapt in MR scenarios, making it difficult to accurately perceive scene changes or reasonably respond to changes in user behavior, resulting in a lack of flexibility in behavioral decision-making and naturalness in human-computer interaction.

Method used

By obtaining sample training sets to iteratively train the behavior decision model, using multi-branch neural network architecture design and cross-modal sensor data, combined with a multi-objective optimization strategy of contrast loss and distribution alignment loss, digital human performance data adapted to the current scenario is generated.

Benefits of technology

It improves the behavioral decision-making flexibility and naturalness of human-computer interaction of digital humans in MR scenarios, enhances the adaptability to changing environments and user behaviors, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120354176B_ABST
    Figure CN120354176B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for training a behavioral decision model and an adaptive interaction method for a digital human. The training method includes: obtaining a sample training set; iteratively training the behavioral decision model to be trained using the sample training set until the objective function loss value of the behavioral decision model to be trained is less than a preset loss threshold; the objective function loss value is obtained by: generating a contrast loss based on the relative entropy between a first predicted behavior distribution output by the behavioral decision model to be trained for positive samples and a second predicted behavior distribution output for negative samples; generating a distribution alignment loss based on the mean squared error between the first predicted behavior distribution and the first target behavior distribution, and the mean squared error between the second predicted behavior distribution and the second target behavior distribution; and generating an objective function loss value based on the contrast loss and the distribution alignment loss. This invention solves the technical problem of insufficient behavioral decision-making for digital humans in MR scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of digital humans, and in particular to a training method for a behavior decision-making model and an adaptive interaction method for digital humans. Background Art

[0002] With the rapid development of virtual reality (VR), augmented reality (AR), and mixed reality (MR) technologies, MR scenarios have been widely applied in fields such as education, entertainment, healthcare, and industry. By deeply integrating virtual information with the real world, MR technology can provide users with a more immersive and interactive experience. Digital humans, as the core interactive subjects in MR scenarios, possess the ability to communicate, collaborate, and interact naturally with users. The intelligence and adaptability of their behavioral performance directly impact the user's overall interactive experience.

[0003] However, the complexity and dynamism of MR scenes in existing technologies place higher demands on the interactive capabilities of digital humans. Specifically, MR scenes can present variable environmental factors, such as changes in scene structure, lighting, and the dynamic distribution of objects. Furthermore, user behaviors, emotions, and interaction needs often exhibit significant diversity and uncertainty. Existing digital human behavioral decision-making models generally lack the ability to effectively integrate and dynamically adapt to multimodal information, making it difficult to accurately perceive scene changes or appropriately respond to changes in user behavior. This results in a lack of flexibility in behavioral decision-making and a lack of natural human-computer interaction, impacting the user experience.

[0004] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention

[0005] The embodiments of the present invention provide a method for training a behavior decision model and a method for adaptive interaction of a digital human, so as to at least solve the technical problem of insufficient behavior decision-making of a digital human in an MR scene.

[0006] According to one aspect of an embodiment of the present invention, a method for training a behavior decision model is provided, comprising: obtaining a sample training set, wherein the sample training set includes positive samples and negative samples; using the sample training set to iteratively train the behavior decision model to be trained until the objective function loss value of the behavior decision model to be trained is less than a preset loss threshold, so as to obtain the trained behavior decision model; wherein the objective function loss value is obtained by: generating a contrast loss based on the relative entropy between a first predicted behavior distribution output by the behavior decision model to be trained for the positive sample and a second predicted behavior distribution output for the negative sample; generating a contrast loss based on the relative entropy between the first predicted behavior distribution and the first predicted behavior distribution; The mean square error between the target behavior distributions, and the mean square error between the second predicted behavior distribution and the second target behavior distribution, generate a distribution alignment loss, wherein the first target behavior distribution is the target behavior distribution obtained by weighted aggregation of the first predicted behavior distribution of each training round based on the first confidence weight corresponding to each training round and the distribution entropy of the first predicted behavior distribution, and the second target behavior distribution is the target behavior distribution obtained by weighted aggregation of the second predicted behavior distribution of each training round based on the second confidence weight corresponding to each training round and the distribution entropy of the second predicted behavior distribution; based on the contrast loss and the distribution alignment loss, generate the objective function loss value.

[0007] According to another aspect of an embodiment of the present invention, a method for adaptive interaction of a digital human in an MR scene is provided, comprising: collecting real-time information of the MR scene through a cross-modal sensor to obtain cross-modal sensor data, wherein the cross-modal sensor comprises at least one of the following: a depth sensor, a visual sensor, and an environmental sensor; pre-processing the cross-modal sensor data to obtain scene perception data, and extracting scene feature information and user behavior information from the scene perception data; generating a behavior decision result of the digital human using a trained behavior decision model based on the scene feature information and the user behavior information; adaptively adjusting the digital human based on the behavior decision result to generate digital human performance data adapted to the current scene, and driving the digital human to interact in the MR scene based on the digital human performance data; wherein the trained behavior decision model is trained using the above-mentioned training method.

[0008] According to another aspect of an embodiment of the present invention, a training device for a behavior decision model is also provided, comprising: an acquisition module configured to acquire a sample training set, wherein the sample training set includes positive samples and negative samples; a training module configured to use the sample training set to iteratively train the behavior decision model to be trained until the objective function loss value of the behavior decision model to be trained is less than a preset loss threshold, so as to obtain the trained behavior decision model; wherein the objective function loss value is obtained by: generating a contrast loss based on the relative entropy between a first predicted behavior distribution output by the behavior decision model to be trained for the positive sample and a second predicted behavior distribution output for the negative sample; generating a contrast loss based on the first predicted behavior distribution; generating a contrast loss based on the first predicted behavior distribution; generating a contrast loss based on the first predicted behavior distribution; generating a contrast loss based on the first predicted behavior distribution; generating a contrast loss based on the first predicted behavior distribution; generating a contrast loss based on the first predicted behavior distribution; generating a contrast loss based on the first predicted behavior distribution; generating a contrast loss based on the The mean square error between the predicted behavior distribution and the first target behavior distribution, and the mean square error between the second predicted behavior distribution and the second target behavior distribution, generate a distribution alignment loss, wherein the first target behavior distribution is a target behavior distribution obtained by weighted aggregation of the first predicted behavior distribution of each training round based on the first confidence weight corresponding to each training round and the distribution entropy of the first predicted behavior distribution, and the second target behavior distribution is a target behavior distribution obtained by weighted aggregation of the second predicted behavior distribution of each training round based on the second confidence weight corresponding to each training round and the distribution entropy of the second predicted behavior distribution; based on the contrast loss and the distribution alignment loss, generate the objective function loss value.

[0009] According to another aspect of an embodiment of the present invention, there is provided an adaptive interaction device for a digital human in an MR scene, comprising: an acquisition module configured to acquire real-time information of the MR scene through a cross-modal sensor to obtain cross-modal sensor data, wherein the cross-modal sensor comprises at least one of the following: a depth sensor, a visual sensor, and an environmental sensor; a processing module configured to pre-process the cross-modal sensor data to obtain scene perception data, and extract scene feature information and user behavior information from the scene perception data; a decision generation module configured to generate a behavior decision result of the digital human using a trained behavior decision model based on the scene feature information and the user behavior information; a driving module configured to adaptively adjust the digital human based on the behavior decision result, generate digital human performance data adapted to the current scene, and drive the digital human to interact in the MR scene based on the digital human performance data; wherein the trained behavior decision model is trained using the above-mentioned training method.

[0010] In an embodiment of the present invention, a sample training set is obtained, wherein the sample training set includes positive samples and negative samples; the behavior decision model to be trained is iteratively trained using the sample training set until the objective function loss value of the behavior decision model to be trained is less than a preset loss threshold, so as to obtain the trained behavior decision model; wherein the objective function loss value is obtained by: generating a contrast loss based on the relative entropy between the first predicted behavior distribution output by the behavior decision model to be trained for the positive sample and the second predicted behavior distribution output for the negative sample; generating a contrast loss based on the mean square error between the first predicted behavior distribution and the first target behavior distribution , and the mean squared error between the second predicted behavior distribution and the second target behavior distribution to generate a distribution alignment loss, wherein the first target behavior distribution is a target behavior distribution obtained by weighted aggregation of the first predicted behavior distribution for each training round based on the first confidence weight corresponding to each training round and the distribution entropy of the first predicted behavior distribution, and the second target behavior distribution is a target behavior distribution obtained by weighted aggregation of the second predicted behavior distribution for each training round based on the second confidence weight corresponding to each training round and the distribution entropy of the second predicted behavior distribution; the objective function loss value is generated based on the contrast loss and the distribution alignment loss. The above solution solves the technical problem of insufficient behavioral decision-making of digital humans in MR scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0012] Figure 1 is a flowchart of an optional method for training a behavior decision model according to an embodiment of the present invention;

[0013] Figure 2 is a flowchart of another optional behavior decision model training method according to an embodiment of the present invention;

[0014] Figure 3 is a flowchart of an optional method for initializing a training objective function and hyperparameters according to an embodiment of the present invention;

[0015] Figure 4 is a flow chart of an optional adaptive interaction method of a digital human in an MR scene according to an embodiment of the present invention;

[0016] Figure 5 1 is a schematic structural diagram of an optional behavioral decision model training device according to an embodiment of the present invention;

[0017] Figure 62 is a schematic structural diagram of an optional adaptive interaction device for a digital human in an MR scene according to an embodiment of the present invention;

[0018] Figure 7 A schematic structural diagram of an electronic device suitable for implementing the embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0019] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0020] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0021] According to an embodiment of the present invention, a method embodiment of a method for training a behavioral decision model is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0022] Figure 1 is a training method for a behavior decision model according to an embodiment of the present invention, such as Figure 1 As shown, the method includes the following steps:

[0023] Step S102: Obtain a sample training set, wherein the sample training set includes positive samples and negative samples.

[0024] Step S104 , iteratively training the behavior decision model to be trained using the sample training set until the objective function loss value of the behavior decision model to be trained is less than a preset loss threshold, so as to obtain the trained behavior decision model.

[0025] The objective function loss value is obtained by:

[0026] First, a contrast loss is generated based on the relative entropy between the first predicted behavior distribution output by the behavior decision model to be trained for the positive sample and the second predicted behavior distribution output for the negative sample. For example, the first predicted behavior distribution is used as the first reference distribution, the second predicted behavior distribution is used as the first target distribution, the first divergence between the first predicted behavior distribution and the second predicted behavior distribution is calculated, and the first divergence is used as the relative entropy in the first direction; the second predicted behavior distribution is used as the second reference distribution, the first predicted behavior distribution is used as the second target distribution, the second divergence between the first predicted behavior distribution and the second predicted behavior distribution is calculated, and the second divergence is used as the relative entropy in the second direction; the relative entropy in the first direction and the relative entropy in the second direction are weighted and fused to generate the symmetric contrast loss.

[0027] Secondly, based on the mean square error between the first predicted behavior distribution and the first target behavior distribution, and the mean square error between the second predicted behavior distribution and the second target behavior distribution, a distribution alignment loss is generated. For example, the mean square error between the first predicted behavior distribution and the first target behavior distribution, and the mean square error between the second predicted behavior distribution and the second target behavior distribution are interval discretized, and the mean square error between the first predicted behavior distribution and the first target behavior distribution, and the mean square error between the second predicted behavior distribution and the second target behavior distribution are divided into multiple different numerical intervals, and the weighted residuals of the different numerical intervals are calculated; based on the weighted residuals of the different numerical intervals, the nonlinear response error change between the first predicted behavior distribution and the second predicted behavior distribution is calculated, and the nonlinear response error change is used as the distribution alignment loss. For example, a nonlinear response function is applied to the weighted residuals of the different numerical intervals to generate a dynamic interval adjustment factor; based on the dynamic interval adjustment factor and the weighted residuals of the different numerical intervals, the nonlinear response error change is calculated.

[0028] In some embodiments, the first target behavior distribution is a target behavior distribution obtained by weighted aggregation of the first predicted behavior distributions of each training round based on the first confidence weight corresponding to each training round and the distribution entropy of the first predicted behavior distribution. Specifically, first, the first confidence weight of the first predicted behavior distribution of each training round in each training round is calculated; for example, in each training round, based on the modal label to which the positive sample belongs, the first predicted behavior distribution is modally grouped to obtain multiple modal groups, and based on the multiple modal groups, the modal consistency similarity between each training round and the historical training round is calculated; based on the modal consistency similarity, the first confidence weight of the first predicted behavior distribution of each training round is determined. Then, the distribution entropy of the first predicted behavior distribution of each training round is calculated, wherein the distribution entropy of the first predicted behavior distribution is used to measure the uncertainty of the first predicted behavior distribution of each training round; finally, based on the first confidence weight and the distribution entropy of the first predicted behavior distribution, a weighted aggregation algorithm based on the training stage sensitive control factor is used to calculate the first target behavior distribution, wherein the training stage sensitive control factor is used to dynamically adjust the first confidence weight and the distribution entropy of the first predicted behavior distribution according to the loss reduction trend of the stage in which the current training round is located.

[0029] In some embodiments, the second target behavior distribution is a target behavior distribution obtained by weighted aggregation of the second predicted behavior distributions of each training round based on the second confidence weight corresponding to each training round and the distribution entropy of the second predicted behavior distribution. Specifically, first, the second confidence weight of the second predicted behavior distribution of each training round in each training round is calculated; for example, in each training round, based on the modal label to which the negative sample belongs, the second predicted behavior distribution is modally grouped to obtain multiple modal groups, and based on the multiple modal groups, the modal consistency similarity between each training round and the historical training round is calculated; based on the modal consistency similarity, the second confidence weight of the second predicted behavior distribution of each training round is determined. Then, the distribution entropy of the second predicted behavior distribution of each training round is calculated, wherein the distribution entropy of the second predicted behavior distribution is used to measure the uncertainty of the second predicted behavior distribution of each training round; finally, based on the second confidence weight and the distribution entropy of the second predicted behavior distribution, a weighted aggregation algorithm based on the training stage sensitive control factor is used to calculate the second target behavior distribution, wherein the training stage sensitive control factor is used to dynamically adjust the second confidence weight and the distribution entropy of the second predicted behavior distribution according to the loss reduction trend of the stage in which the current training round is located.

[0030] Finally, the objective function loss value is generated based on the contrast loss and the distribution alignment loss.

[0031] Figure 2 FIG. 1 is a flow chart of another method for training a behavior decision model according to an embodiment of the present invention. Figure 2 As shown, the method includes the following steps:

[0032] Step S202: Obtain a sample training set.

[0033] In this embodiment, the sample training set includes both positive and negative samples, each composed of multimodal data, including but not limited to visual image sequences, speech intonation, skeletal motion data, and environmental state parameters. To ensure sample quality, positive samples are annotated based on real MR interaction scenarios, clearly identifying the ideal behavioral strategies that the digital human should adopt in specific contexts. Negative samples are constructed from misoperation behaviors or simulated abnormal scenarios to provide data reference for training the model's discriminative capabilities.

[0034] To improve the coverage and robustness of training data, a data augmentation mechanism was introduced during the sample construction process. Visual data was transformed through rotation, cropping, and color perturbation to generate view-transformed versions. Voice data was simulated through speed changes, noise addition, and editing to simulate real-world interaction noise. Motion data was perturbed through time series to simulate differences in motion rhythms at different time scales. Environmental state data was generated using a sensor simulator to generate sensor datasets under varying lighting, temperature, and humidity conditions.

[0035] In addition, the sample training set also adds modality missing scenarios to enhance the model's robustness in behavior prediction when some modalities are unavailable. For example, when voice commands are lost, images are partially occluded, or action signals are delayed, the model can still maintain a high level of behavior prediction accuracy.

[0036] Step S204: constructing the initial structure of the behavior decision model.

[0037] The behavioral decision-making model utilizes a multi-branch neural network architecture, capable of simultaneously receiving input from multiple modalities and analyzing the input data at multiple semantic levels. The overall model structure includes a perceptual encoding module, a modal synchronization and alignment module, a multimodal fusion module, and a behavioral prediction module.

[0038] The perceptual coding module extracts features from different modal data. Visual modality input is first processed through a lightweight convolutional neural network (such as MobileNetV3) to extract image feature sequences. Speech modality input is processed through a speech encoder (such as a Transformer-based speech understanding network) to obtain semantic vectors. Skeletal motion data is processed through a graph convolutional network (GCN) to obtain temporal motion features. Environmental state data is modeled as an environmental vector representation using a fully connected neural network.

[0039] The modality synchronization module introduces an attention mechanism based on temporal alignment to dynamically align and match modalities in the temporal dimension. For example, by using timestamp playback technology and the Dynamic Time Warping (DTW) algorithm, key information from different modalities can be semantically integrated at the same time step.

[0040] The multimodal fusion module uses a multi-head self-attention mechanism to achieve weighted fusion of features. To improve the contextual adaptability of the fusion, this embodiment introduces a modality credibility gating mechanism. This dynamically adjusts the weight of each modality in the fusion process based on its current quality indicators (such as signal-to-noise ratio, occlusion rate, and keypoint missingness), thereby improving the semantic integrity of the fused vector.

[0041] The behavior prediction module, based on a multi-layer perceptron (MLP) design, maps fused features into a behavior space and outputs a predicted behavior probability distribution. The behavior space can be preset as a fixed-dimensional classification space, representing the set of actions that the digital human can perform, or it can be expanded to a parameterized policy distribution (e.g., a continuous behavior output with action parameters).

[0042] Step S206: Initialize the training objective function and hyperparameters.

[0043] In order to achieve efficient training and improve the discrimination ability, this embodiment provides a multi-objective optimization strategy that integrates contrast loss and distribution alignment loss. The method of initializing the training objective function and hyperparameters is as follows: Figure 3 As shown, the following steps are included:

[0044] Step S2062: Calculate contrast loss.

[0045] First, temperature scaling is performed. The first predicted behavior distribution is denoted as , which represents the behavior probability distribution predicted by the model under positive sample input; the second predicted behavior distribution is recorded as , represents the behavior probability distribution output by the model under negative sample input. Before calculating the KL divergence of the first predicted behavior distribution and the second predicted behavior distribution, the model first applies a temperature scaling transformation to each behavior distribution to control the smoothness of the softmax output distribution. The formula is:

[0046]

[0047] in 、 is the original logit (logistic regression) value of the i-th type of behavior, τ∈(0,1] is the temperature parameter. The smaller the temperature parameter, the sharper the behavior distribution will be, which will increase the model's sensitivity to category differences. n represents the number of samples.

[0048] In this embodiment, τ is preferably set to 0.05, and is adjusted in stages at different training stages. For example, a higher temperature is used in the early stage to ensure learning stability, and then gradually reduced in the later stage to improve the resolution of behavior discrimination.

[0049] In some other embodiments, in order to enhance the sensitivity to the comparison of the tail of the distribution (low-probability behavior), a tail emphasis factor λ may be introduced, and a nonlinear weighting function may be introduced for the low-probability term to obtain a modified first predicted behavior distribution:

[0050]

[0051] The modified first predicted behavior distribution is taken as In the same way, the modified second predicted behavior distribution can be calculated.

[0052] Next, the relative entropy in the first direction is calculated.

[0053] The relative entropy in the first direction (first divergence) is defined as follows:

[0054]

[0055] in, A very small constant added to prevent division by zero errors. This metric indicates whether high-probability events in positive sample predictions are significantly underestimated in negative sample predictions. If negative samples are insensitive to key behavior predictions, the first divergence in this direction will increase significantly.

[0056] Next, calculate the relative entropy in the second direction. In the calculation of the relative entropy in the second direction, swap the roles of reference and target, use the second predicted behavior distribution as the reference distribution, and the first predicted behavior distribution as the target distribution, and obtain the second divergence as the relative entropy in the second direction:

[0057]

[0058] The second direction mainly characterizes the self-consistency under abnormal behavior. That is, if the negative sample model outputs a high probability of certain abnormal behaviors, and the positive sample output is highly repulsive to this (that is, gives a low probability), then the second divergence in this direction will also be amplified, thereby providing an optimization signal for model training.

[0059] To prevent the model from focusing too much on low-probability behaviors in the early stages of training, which can lead to oscillatory learning, a warm-up strategy is set up in the first few training rounds, which introduces linear weighting for the second divergence:

[0060]

[0061] Among them, t represents the current training round, T warmup represents the upper limit of the warm-up rounds (e.g., 10 rounds), and r(t) represents the scaling factor. The above method can help improve the training stability of contrastive learning.

[0062] Finally, the contrast loss is generated by symmetric fusion. To overcome the asymmetric nature of the divergence itself, the relative entropy of the first direction and the second direction are weighted and fused to construct a symmetric contrast loss term:

[0063]

[0064] Here, α is the fusion coefficient, ranging from 0 to 1. The default setting is 0.5 to achieve complete symmetry, and its value can be adjusted dynamically based on the performance of the validation set. For example, if the model is found to be more likely to misclassify negative samples during training, the fusion coefficient weight can be increased to increase the importance of the relative entropy of negative samples.

[0065] Step S2064: Calculate the distribution alignment loss.

[0066] 1) Dynamically construct target behavior distribution.

[0067] In this embodiment, the first target behavior distribution and the second target behavior distribution are not composed of static labels, but are generated based on the predicted output of historical training rounds using a dynamic confidence aggregation strategy.

[0068] Taking the first target behavior distribution as an example, suppose that in the tth round of training, the first predicted behavior distribution is , and process it as follows:

[0069] First, the confidence weights are calculated.

[0070] In each round, the predicted behavior distribution is grouped by modality based on the modal label of the positive samples. For example, image modality, speech modality, action modality, etc. For each modality, the consistency of the predicted distribution of the modality in the current round and the previous rounds is calculated (for example, using cosine similarity or KL divergence antisimilarity). Finally, based on the modality consistency, the weighted average confidence is calculated as the first confidence weight:

[0071]

[0072] in, is the modality consistency similarity of the training round, is the modality consistency similarity of historical training rounds, and M represents the number of historical reference samples.

[0073] Next, calculate the distribution entropy :

[0074]

[0075] in, The first predicted behavior distribution of the i-th sample in the t-th training round is represented by the calculated distribution entropy, which is used to reflect the degree of certainty of the current predicted distribution. The distribution entropy of the second predicted behavior distribution can be calculated in the same way.

[0076] Then, the sensitive regulatory factors in the training phase are used for aggregation.

[0077] This embodiment introduces a training phase sensitivity factor to dynamically weight the confidence of different rounds according to the loss reduction rate. For example, the training phase sensitivity factor is obtained by the following formula:

[0078]

[0079] Where Lt is the loss value of the tth round, and k is the sensitivity adjustment coefficient. The first target behavior distribution is constructed as follows:

[0080]

[0081] The construction logic of the second target behavior distribution is exactly the same as above and will not be repeated here.

[0082] 2) Perform interval discretization of the mean square error and calculate the weighted residual.

[0083] First, for each training sample, the mean square error (MSE) between its predicted behavior distribution and the target behavior distribution is calculated. Let the first predicted behavior distribution be , corresponding to the first target behavior distribution is , then the mean square error is recorded as:

[0084]

[0085] Similarly, the error between the second predicted behavior distribution and the second target behavior distribution is recorded as:

[0086]

[0087] Next, the above error results are discretized. Preset several discrete value intervals {I1, I2, ..., I k Each interval represents an MSE range. The system traverses all training samples, identifies the interval to which their mean square error belongs, and records the number of samples and the average error value in each interval.

[0088] For each interval Assign a residual weight w j , the weight can be determined based on the variance of sample errors within the interval, the number of samples, or the stability index of historical training. For example, the weight is defined as:

[0089]

[0090] in is an interval The average error within, λ is the regularization coefficient, δ is a small constant to prevent division by zero, and x is a value in the interval.

[0091] Finally, the sum of the weighted residuals of all intervals is the original residual term:

[0092]

[0093] in is an interval The average error value of , k is the number of intervals.

[0094] 3) Perform nonlinear response function mapping and generate dynamic adjustment factors.

[0095] In order to more precisely determine the response relationship between the residual and the error change, the weighted residual result is input into the nonlinear response function to generate a dynamic adjustment factor. The dynamic adjustment factor can be selected in the following form:

[0096]

[0097] in, The hyperbolic tangent function is used in this embodiment to provide a stronger response in a small error range and to saturate the control error dominant weight in a large error range.

[0098] Then, the nonlinear response error change term of each interval is calculated based on the dynamic adjustment factor and the weighted residual. :

[0099] 4) Combine the response error changes of all intervals to obtain the overall distribution alignment loss:

[0100]

[0101] Compared with simple mean square error accumulation, this method more effectively integrates the amount of information and dynamic response intensity of different error intervals, so that the model can be stably optimized at different accuracy stages.

[0102] Step S2066, determine the objective function loss value.

[0103] The objective function loss value is generated by weighted summing of the contrast loss and distribution alignment loss.

[0104] Step S208: Execute the model training process.

[0105] This step uses an iterative training mechanism based on batch gradient descent, combined with a multi-sample dynamic sampling strategy to ensure a balance between training efficiency and sample distribution.

[0106] First, a batch of positive and negative samples are extracted from the training set, processed by the modal synchronization module, and then input into the model; the model outputs the behavior probability distribution; the corresponding target distribution is calculated based on the historical prediction distribution of the current round, the modal label, and the modal quality estimation; the contrast loss is calculated in sequence t Alignment loss with distribution; the weighted sum of the two losses is used as the target loss.

[0107] Model parameters are updated using the backpropagation algorithm, and the AdamW optimizer is used to enhance regularization. A cosine annealing scheduling strategy is introduced for learning rate settings, dynamically adjusting the learning rate based on the current training round position and historical loss trends to avoid local optimality.

[0108] In order to control the overfitting problem during training, this embodiment additionally introduces a modal dropout mechanism, which randomly blocks certain modal inputs in each training batch, forcing the model to learn to rely on information sources of multiple modalities for decision-making, thereby improving generalization capabilities.

[0109] This embodiment also provides a validation set cycle detection mechanism. After each fixed round of training, an independent validation set is evaluated. Metrics include behavior prediction accuracy, interaction rationality scores, and modality sensitivity indicators (such as the degree of performance degradation in the absence of a modality). If the verification performance improvement is less than a set threshold over several consecutive cycles, training is considered to be saturated, and the system automatically enters the convergence detection state. The training phase sensitivity control mechanism is activated to fine-tune the model.

[0110] Model training is considered complete when the objective function loss value of the training process converges and reaches the minimum threshold (the preset loss threshold), or when the performance of the validation set meets the design requirements. At this point, the model weight and structure configuration files are automatically exported.

[0111] In summary, the behavioral decision-making model training method for digital humans in MR scenarios provided in this embodiment has an accurate objective function and a dynamically adjustable training process. In particular, the introduction of a sensitive control mechanism and a nonlinear error response module in the training phase makes model training more stable, converges faster, and has better performance.

[0112] This application also provides an adaptive interaction method for digital humans in MR scenes, such as Figure 4 As shown, the following steps are included:

[0113] Step S402: collecting real-time information of the MR scene using a cross-modal sensor to obtain cross-modal sensor data, wherein the cross-modal sensor includes at least one of the following: a depth sensor, a visual sensor, and an environmental sensor;

[0114] Step S404: preprocessing the cross-modal sensor data to obtain scene perception data, and extracting scene feature information and user behavior information from the scene perception data;

[0115] Step S406, generating a behavior decision result of the digital human using the trained behavior decision model based on the scene feature information and the user behavior information;

[0116] Step S408: Based on the behavior decision result, the digital human is adaptively adjusted to generate digital human performance data adapted to the current scene, and the digital human is driven to interact in the MR scene based on the digital human performance data;

[0117] The trained behavior decision model is obtained by using the above training method.

[0118] This application also provides a training device for a behavior decision model, such as Figure 5As shown, it includes: an acquisition module 52, which is configured to acquire a sample training set, wherein the sample training set includes positive samples and negative samples; a training module 54, which is configured to use the sample training set to iteratively train the behavior decision model to be trained until the objective function loss value of the behavior decision model to be trained is less than a preset loss threshold, so as to obtain the trained behavior decision model; wherein the objective function loss value is obtained by: generating a contrast loss based on the relative entropy between the first predicted behavior distribution output by the behavior decision model to be trained for the positive sample and the second predicted behavior distribution output for the negative sample; based on the The mean square error between the first predicted behavior distribution and the first target behavior distribution, and the mean square error between the second predicted behavior distribution and the second target behavior distribution, generate a distribution alignment loss, wherein the first target behavior distribution is a target behavior distribution obtained by weighted aggregation of the first predicted behavior distribution of each training round based on the first confidence weight corresponding to each training round and the distribution entropy of the first predicted behavior distribution, and the second target behavior distribution is a target behavior distribution obtained by weighted aggregation of the second predicted behavior distribution of each training round based on the second confidence weight corresponding to each training round and the distribution entropy of the second predicted behavior distribution; based on the contrast loss and the distribution alignment loss, generate the objective function loss value.

[0119] This application also provides an adaptive interaction device for digital humans in MR scenes, such as Figure 6 As shown, the device includes: an acquisition module 62, configured to collect real-time information of the MR scene through a cross-modal sensor to obtain cross-modal sensor data, wherein the cross-modal sensor includes at least one of the following: a depth sensor, a visual sensor, and an environmental sensor; a processing module 64, configured to pre-process the cross-modal sensor data to obtain scene perception data, and extract scene feature information and user behavior information from the scene perception data; a decision generation module 66, configured to generate a behavior decision result of a digital human based on the scene feature information and the user behavior information using a trained behavior decision model; a driving module 68, configured to adaptively adjust the digital human based on the behavior decision result, generate digital human performance data adapted to the current scene, and drive the digital human to interact in the MR scene based on the digital human performance data; wherein the trained behavior decision model is trained using the above-mentioned training method.

[0120] It should be noted that the apparatus provided in the above embodiments is merely exemplified by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments are based on the same concept. The specific implementation process is detailed in the method embodiments and will not be repeated here.

[0121] Figure 7 Schematic diagram of the structure of an electronic device suitable for implementing the embodiment of the present disclosure is shown. Figure 7 The electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.

[0122] like Figure 7 As shown, the electronic device includes a central processing unit (CPU) 1001, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1002 or the program loaded from the storage part 1008 into the random access memory (RAM) 1003. Various programs and data required for system operation are also stored in the RAM 1003. The CPU 1001, ROM 1002 and RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.

[0123] The following components are connected to the I / O interface 1005: an input section 1006 including a keyboard, a mouse, and the like; an output section 1007 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section 1008 including a hard disk; and a communication section 1009 including a network interface card such as a LAN card or a modem. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the I / O interface 1005 as needed. A removable medium 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 1010 as needed, so that computer programs read therefrom can be installed into the storage section 1008 as needed.

[0124] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A training method for a behavioral decision model, characterized in that: include: Acquire a sample training set, wherein the sample training set includes positive samples and negative samples, and each sample in the sample training set is composed of multimodal data, and the multimodal data includes visual image sequences, speech intonation, skeletal motion data, and environmental state parameters; Iteratively training the behavior decision model to be trained using the sample training set until the objective function loss value of the behavior decision model to be trained is less than a preset loss threshold, so as to obtain the trained behavior decision model; The objective function loss value is obtained as follows: generating a contrast loss based on a relative entropy between a first predicted behavior distribution output by the behavior decision model to be trained for the positive sample and a second predicted behavior distribution output for the negative sample; Generate a distribution alignment loss based on the mean square error between the first predicted behavior distribution and the first target behavior distribution, and the mean square error between the second predicted behavior distribution and the second target behavior distribution, wherein the first target behavior distribution is a target behavior distribution obtained by weighted aggregation of the first predicted behavior distribution of each training round based on the first confidence weight corresponding to each training round and the distribution entropy of the first predicted behavior distribution, and the second target behavior distribution is a target behavior distribution obtained by weighted aggregation of the second predicted behavior distribution of each training round based on the second confidence weight corresponding to each training round and the distribution entropy of the second predicted behavior distribution; Generating the objective function loss value based on the contrast loss and the distribution alignment loss; Wherein, based on the mean square error between the first predicted behavior distribution and the first target behavior distribution, and the mean square error between the second predicted behavior distribution and the second target behavior distribution, a distribution alignment loss is generated, including: performing interval discretization processing on the mean square error between the first predicted behavior distribution and the first target behavior distribution, and the mean square error between the second predicted behavior distribution and the second target behavior distribution, dividing the mean square error between the first predicted behavior distribution and the first target behavior distribution, and the mean square error between the second predicted behavior distribution and the second target behavior distribution into multiple different numerical intervals, and calculating the weighted residuals of the different numerical intervals; based on the weighted residuals of the different numerical intervals, calculating the nonlinear response error change between the first predicted behavior distribution and the second predicted behavior distribution, and using the nonlinear response error change as the distribution alignment loss.

2. The method according to claim 1, characterized in that Generating a contrast loss based on a relative entropy between a first predicted behavior distribution output by the behavior decision model to be trained for the positive sample and a second predicted behavior distribution output for the negative sample, including: Taking the first predicted behavior distribution as a first reference distribution and the second predicted behavior distribution as a first target distribution, calculating a first divergence between the first predicted behavior distribution and the second predicted behavior distribution, and taking the first divergence as the relative entropy in a first direction; Using the second predicted behavior distribution as a second reference distribution and the first predicted behavior distribution as a second target distribution, calculating a second divergence between the first predicted behavior distribution and the second predicted behavior distribution, and using the second divergence as the relative entropy in a second direction; The relative entropy in the first direction and the relative entropy in the second direction are weightedly fused to generate the symmetric contrast loss.

3. The method according to claim 2, characterized in that Calculating a nonlinear response error change between the first predicted behavior distribution and the second predicted behavior distribution based on the weighted residuals in the different numerical intervals includes: Applying a nonlinear response function to the weighted residuals of the different numerical intervals to generate a dynamic interval adjustment factor; The nonlinear response error change is calculated based on the dynamic interval adjustment factor and the weighted residuals of the different numerical intervals.

4. The method according to claim 2, characterized in that The first target behavior distribution is obtained by: Calculating the first confidence weight of the first predicted behavior distribution of each training round in the training rounds; Calculating the distribution entropy of the first predicted behavior distribution of each training round, wherein the distribution entropy of the first predicted behavior distribution is used to measure the uncertainty of the first predicted behavior distribution of each training round; Based on the first confidence weight and the distribution entropy of the first predicted behavior distribution, a weighted aggregation algorithm based on the training stage sensitive control factor is used to calculate the first target behavior distribution, wherein the training stage sensitive control factor is used to dynamically adjust the first confidence weight and the distribution entropy of the first predicted behavior distribution according to the loss downward trend of the stage in which the current training round is located.

5. The method according to claim 3, characterized in that Calculating the first confidence weight of the first predicted behavior distribution in each of the training rounds includes: In each of the training rounds, based on the modal labels to which the positive samples belong, the first predicted behavior distribution is modally grouped to obtain a plurality of modal groups, and based on the plurality of modal groups, a modality consistency similarity between each of the training rounds and a historical training round is calculated; Based on the modality consistency similarity, the first confidence weight of the first predicted behavior distribution of each training round is determined.

6. An adaptive interaction method for digital humans in MR scenes, characterized in that: include: Real-time information collection of the MR scene is performed using a cross-modal sensor to obtain cross-modal sensor data, wherein the cross-modal sensor includes at least one of the following: a depth sensor, a visual sensor, and an environmental sensor; Preprocessing the cross-modal sensor data to obtain scene perception data, and extracting scene feature information and user behavior information from the scene perception data; Based on the scene feature information and the user behavior information, generating a behavior decision result of the digital human using a trained behavior decision model; Based on the behavioral decision result, the digital human is adaptively adjusted to generate digital human performance data adapted to the current scene, and based on the digital human performance data, the digital human is driven to interact in the MR scene; The trained behavior decision model is obtained by training using the method described in any one of claims 1 to 5.

7. A training device for a behavioral decision model, characterized in that: include: an acquisition module configured to acquire a sample training set, wherein the sample training set includes positive samples and negative samples, and each sample in the sample training set is composed of multimodal data, and the multimodal data includes visual image sequences, speech intonation, skeletal motion data, and environmental state parameters; a training module configured to iteratively train the behavior decision model to be trained using the sample training set until the objective function loss value of the behavior decision model to be trained is less than a preset loss threshold, so as to obtain the trained behavior decision model; The objective function loss value is obtained as follows: generating a contrast loss based on a relative entropy between a first predicted behavior distribution output by the behavior decision model to be trained for the positive sample and a second predicted behavior distribution output for the negative sample; Generate a distribution alignment loss based on the mean square error between the first predicted behavior distribution and the first target behavior distribution, and the mean square error between the second predicted behavior distribution and the second target behavior distribution, wherein the first target behavior distribution is a target behavior distribution obtained by weighted aggregation of the first predicted behavior distribution of each training round based on the first confidence weight corresponding to each training round and the distribution entropy of the first predicted behavior distribution, and the second target behavior distribution is a target behavior distribution obtained by weighted aggregation of the second predicted behavior distribution of each training round based on the second confidence weight corresponding to each training round and the distribution entropy of the second predicted behavior distribution; Generating the objective function loss value based on the contrast loss and the distribution alignment loss; Wherein, based on the mean square error between the first predicted behavior distribution and the first target behavior distribution, and the mean square error between the second predicted behavior distribution and the second target behavior distribution, a distribution alignment loss is generated, including: performing interval discretization processing on the mean square error between the first predicted behavior distribution and the first target behavior distribution, and the mean square error between the second predicted behavior distribution and the second target behavior distribution, dividing the mean square error between the first predicted behavior distribution and the first target behavior distribution, and the mean square error between the second predicted behavior distribution and the second target behavior distribution into multiple different numerical intervals, and calculating the weighted residuals of the different numerical intervals; based on the weighted residuals of the different numerical intervals, calculating the nonlinear response error change between the first predicted behavior distribution and the second predicted behavior distribution, and using the nonlinear response error change as the distribution alignment loss.

8. An adaptive interactive device for digital humans in MR scenes, characterized in that: include: an acquisition module configured to acquire real-time information of the MR scene through a cross-modal sensor to obtain cross-modal sensor data, wherein the cross-modal sensor includes at least one of the following: a depth sensor, a visual sensor, and an environmental sensor; a processing module configured to preprocess the cross-modal sensor data to obtain scene perception data, and extract scene feature information and user behavior information from the scene perception data; A decision generation module is configured to generate a behavior decision result of a digital human using a trained behavior decision model based on the scene feature information and the user behavior information; a driving module configured to adaptively adjust the digital human based on the behavioral decision result, generate digital human performance data adapted to the current scene, and drive the digital human to interact in the MR scene based on the digital human performance data; The trained behavior decision model is obtained by training using the method described in any one of claims 1 to 5.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the method according to any one of claims 1 to 5.

10. A computer device, characterized in that: include: memory and processor, The memory stores a computer program; The processor is configured to execute a computer program stored in the memory, wherein the computer program enables the processor to execute the method according to any one of claims 1 to 5 when the program is executed.

11. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Virtual-real fusion simulation medical teaching method and system based on AI and MR technologies

    CN117789563A