A personalized digital device training generation method based on multimodal information fusion
Through the method of multimodal data fusion, a personalized digital device training generation method is constructed, which solves the problems of rigid digital human behavior and lack of personalization, and improves training efficiency and interactive experience.
Patent Information
- Application Number
- CN202510884809.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-30
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2045-06-30
AI Technical Summary
Existing digital human generation technology relies on single-modal data input, resulting in rigid behavior, lack of personalization, poor real-time performance, low efficiency in collaborative processing of multi-modal data, and difficulty in supporting high-concurrency interactions.
By collecting multimodal data to generate training data packets, constructing initial intelligent agents and multiple user demand categories, establishing corresponding demand scenarios, and setting training sub-strategies, iterative training is carried out in combination with user feature data packets to generate personalized intelligent agents.
It improves the training efficiency and user interaction experience of personalized digital devices, realizes the personalized training needs of different types of users, and enhances the interactive adaptability between digital humans and users.
Smart Images

Figure CN120408203B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a personalized digital device training and generation method based on multimodal information fusion. Background Art
[0002] In recent years, with the rapid development of technologies such as artificial intelligence (AI), computer vision (CG), natural language processing (NLP), and text-to-speech (TTS), personalized digital humans (HDHs) have become a research hotspot in fields such as human-computer interaction, virtual reality (VR / AR), intelligent customer service, and digital entertainment. A HDH is a computer-generated avatar that simulates human appearance, voice, expressions, and behavior, allowing for natural interaction with users.
[0003] Current digital human generation and activation technologies primarily rely on single-modal data input (such as text-based conversations or voice commands) and generate corresponding speech, expressions, or actions using pre-trained models (such as GPT and VITS). However, these approaches present numerous challenges, including: Single-modal dependency: Traditional solutions rely on a single data source, such as text or voice, resulting in rigid digital human behavior; insufficient personalization, and a lack of dynamic modeling of users' long-term preferences and emotional states; poor real-time performance; and inefficient multimodal data collaborative processing, making it difficult to support high-concurrency interactions. Summary of the Invention
[0004] The purpose of this application is: to solve the above technical problems, this application provides a personalized digital device training generation method based on multimodal information fusion, aiming to improve the training efficiency of personalized digital devices and enhance the user's interactive experience.
[0005] In some embodiments of the present application, a training data packet is generated by collecting multimodal data, an initial intelligent agent and multiple user demand categories are constructed based on the training data packet, corresponding demand scenarios are established based on different user demand categories, and training sub-strategies corresponding to each demand scenario are set, thereby realizing personalized training needs for different types of users.
[0006] In some embodiments of the present application, by analyzing the user's feature data packets, a corresponding first-level training strategy is quickly constructed to improve the training efficiency of personalized digital devices, and by collecting the user's long-term preference data, the initial intelligent agent is iteratively trained to construct a first-level intelligent agent that adapts to the user and improves the user's interactive experience.
[0007] In some embodiments of the present application, a method for generating personalized digital device training based on multimodal information fusion is provided, comprising:
[0008] Generate initial intelligent agents and multiple demand scenarios based on training data packets, and build training models based on all demand scenarios;
[0009] Obtain the user's feature data package and set the first-level training strategy based on the feature data package and training model;
[0010] Generate a first-level agent based on the first-level training strategy;
[0011] Among them, when generating multiple demand scenarios, including:
[0012] Establish a demand scenario sequence A, A=(a1, a2…a i …a n ), where a i is the i-th demand scenario; n is the number of demand scenarios.
[0013] In some embodiments of the present application, the construction of a training model based on all demand scenarios includes:
[0014] Set multiple data modes according to the training data package;
[0015] Establish the data modal sequence P, P=(p1, p2…p i …p m ), where pi is the i-th data modality; m is the number of data modalities;
[0016] Set a in sequence according to the required scenario sequence A i For the target scene;
[0017] Obtain the evaluation data package of the target scene;
[0018] Set the control sub-strategy for the target scenario based on the evaluation data package;
[0019] Generate dependency evaluation values between the target scenario and each data modality;
[0020] Establish a dependent evaluation value sequence B, B=(b1,b2…b i …b n ), where b i is the dependency evaluation value between the target scene and the i-th data modality, and n is the number of dependency evaluation values;
[0021] Set the processing sub-strategy of the target scenario according to the dependent evaluation value sequence B;
[0022] Generate a training sub-strategy for the target scenario based on the control sub-strategy and the processing sub-strategy for the target scenario;
[0023] Generate training sub-strategies for each demand scenario in turn;
[0024] Establish a training sub-strategy sequence W, W=(w1,w2…wi …w n ), where w i is the training sub-strategy for the i-th demand scenario; n is the number of demand scenarios;
[0025] The training model is constructed according to the training sub-strategy sequence W.
[0026] In some embodiments of the present application, when setting a control sub-strategy for a target scenario based on an evaluation data packet, the following steps are included:
[0027] Generate acquisition device parameters for the target scenario based on the evaluation data packet;
[0028] Set acquisition sub-strategies based on acquisition device parameters;
[0029] Generate a demand evaluation value c for the target scenario based on the evaluation data package;
[0030] Set the monitoring cycle duration according to the demand evaluation value c;
[0031] Generate a control sub-strategy for the target scenario based on the acquisition sub-strategy and the monitoring cycle duration.
[0032] In some embodiments of the present application, generating the demand evaluation value c includes:
[0033]
[0034] Among them, e1 is the preset first weight coefficient; e2 is the preset second weight coefficient; Q1 is the preset first fixed coefficient; Q2 is the preset second fixed coefficient; θ1 is the number of data evaluation indicators; β i is the impact factor of the i-th data evaluation index; j i is the reference value of the i-th data evaluation index generated based on the evaluation data packet; θ2 is the number of scene evaluation indicators; η i is the impact factor of the evaluation index of the i-th scene; k i Generates the reference value of the evaluation index of the i-th scene based on the evaluation data package.
[0035] In some embodiments of the present application, when setting a first-level training strategy, it includes:
[0036] Set a in sequence according to the required scenario sequence A i The scene to be compared;
[0037] Generate an associated evaluation value v between the user and the scene to be compared;
[0038] Generate the associated evaluation values of users and each demand scenario in sequence;
[0039] Establish the associated evaluation value sequence V, V=(v1, v2…v i …vn ), where v i is the evaluation value associated with the user and the i-th demand scenario, and n is the number of demand scenarios;
[0040] Set the maximum value v in the associated evaluation value sequence V max The training sub-strategy corresponding to the demand scenario is the first-level training strategy.
[0041] In some embodiments of the present application, generating the association evaluation value v includes:
[0042]
[0043] Among them, θ3 is the number of characteristic evaluation indicators; g i is the influencing factor of the i-th characteristic evaluation index; s i is the similarity evaluation value of the i-th feature evaluation index between the user and the scene to be compared.
[0044] In some embodiments of the present application, when generating a first-level agent according to a first-level training strategy, the process includes:
[0045] Generate multiple iterative data packets based on the first-level training strategy;
[0046] Iteratively train the initial agent according to the iterative data packet;
[0047] Generate a demand evaluation value c' according to the first-level training strategy, and set the initial monitoring cycle duration t1 according to the demand evaluation value c';
[0048] According to the maximum value v in the associated evaluation value sequence V max Set the compensation coefficient α1;
[0049] Set the first-level monitoring cycle duration t, t=α1×t1;
[0050] Set multiple monitoring time nodes according to the first-level monitoring cycle length t;
[0051] Generate an iterative evaluation value for each monitoring time node, and determine whether to output a first-level intelligent agent based on the iterative evaluation value.
[0052] In some embodiments of the present application, generating the iterative evaluation value of each monitoring point includes:
[0053] Get the iteration results and monitoring data packets of the current monitoring time node;
[0054] Generate the iterative evaluation value f of the current monitoring time node;
[0055]
[0056] Among them, e3 is a preset third weight coefficient; e4 is a preset fourth weight coefficient; Q3 is a preset third fixed coefficient; Q4 is a preset fourth fixed coefficient; θ4 is the number of operation evaluation indicators; r 1i is the influence factor of the i-th operation evaluation indicator; d i is the reference value of the i-th operation evaluation indicator generated based on the iterative result; θ5 is the number of monitoring evaluation indicators; r 2i is the influence factor of the i-th monitoring evaluation indicator; h i is the reference value of the i-th monitoring evaluation indicator generated based on the monitoring data packet; U is a conversion coefficient;
[0057] Generate the output result of the current monitoring time node according to the iterative evaluation value f.
[0058] In some embodiments of the present application, when generating the output result of the current monitoring time node according to the iterative evaluation value f, it includes:
[0059] Preset the first iterative evaluation value threshold F1 and the second iterative evaluation value threshold F2, and F1 < F2;
[0060] [[ID=2*]]If f < F1, generate a first-level correction instruction at the current monitoring time node;
[0061] If F1 < f < F2, generate a first-level iterative instruction at the current monitoring time node;
[0062] If f > F2, output a first-level intelligent agent.
[0063] Compared with the prior art, the beneficial effects of a personalized digital device training generation method based on multi-modal information fusion in the embodiments of the present application are as follows:
[0064] By collecting multi-modal data to generate a training data packet, constructing an initial intelligent agent and various user demand categories according to the training data packet, establishing corresponding demand scenarios based on different user demand categories, and setting training sub-strategies corresponding to each demand scenario, the personalized training needs of different types of users can be realized.
[0065] By analyzing the feature data packet of the user, quickly constructing the corresponding first-level training strategy, improving the training efficiency of the personalized digital device, and iteratively training the initial intelligent agent by collecting the long-term preference data of the user, so as to construct a first-level intelligent agent suitable for the user and improve the user's interaction experience. Brief Description of the Drawings
[0066] Figure 1 is a schematic flowchart of a personalized digital device training generation method based on multi-modal information fusion in a preferred embodiment of the embodiments of the present application. Detailed Embodiments
[0067] The following embodiments are used to illustrate the present invention, but are not intended to limit the scope of the present invention.
[0068] In the description of this application, it should be understood that the terms "center", "up", "down", "front", "back", "left", "right", "vertical", "horizontal", "top", "bottom", "inside", "outside", etc., indicating the orientation or position relationship, are based on the orientation or position relationship shown in the accompanying drawings, and are only for the convenience of describing this application and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on this application.
[0069] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature specified as "first" or "second" may explicitly or implicitly include one or more of such features. Throughout this application, unless otherwise specified, "plurality" means two or more.
[0070] In the description of this application, it should be noted that, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they can refer to fixed connections, detachable connections, or integral connections; mechanical connections or electrical connections; direct connections or indirect connections through an intermediate medium; and internal connections between two components. Those skilled in the art will understand the specific meanings of the above terms in this application based on the specific circumstances.
[0071] like Figure 1 As shown, a preferred embodiment of the present application provides a personalized digital device training generation method based on multimodal information fusion, including:
[0072] S101: Generate an initial agent and multiple demand scenarios based on the training data package, and build a training model based on all demand scenarios;
[0073] S102: Obtain a user's feature data packet and set a first-level training strategy based on the feature data packet and the training model;
[0074] S103: Generate a first-level agent according to the first-level training strategy;
[0075] Among them, when generating multiple demand scenarios, including:
[0076] Establish a demand scenario sequence A, A=(a1, a2…a i …a n ), where a iis the i-th demand scenario; n is the number of demand scenarios.
[0077] Specifically, training data packages are generated by collecting multimodal data. Initial agents are constructed based on these training data packages. These initial agents can complete basic, general interactive tasks, such as basic chat Q&A, basic action simulation, and basic image simulation. Different user need categories are categorized based on the training data packages, and corresponding demand scenarios are established based on these different user need categories.
[0078] Specifically, a single demand scenario represents a category of user demand.
[0079] Specifically, when establishing different user demand categories, they can be divided according to multiple parameters such as the type of user modal data that can be collected, the user's interaction needs, the user's data collection device type, etc.
[0080] Specifically, when building a training model based on all required scenarios, it includes:
[0081] Set multiple data modes according to the training data package;
[0082] Establish the data modal sequence P, P=(p1, p2…p i …p m ), where pi is the i-th data modality; m is the number of data modalities;
[0083] Set a in sequence according to the required scenario sequence A i For the target scene;
[0084] Obtain the evaluation data package of the target scene;
[0085] Set the control sub-strategy for the target scenario based on the evaluation data package;
[0086] Generate dependency evaluation values between the target scenario and each data modality;
[0087] Establish a dependent evaluation value sequence B, B=(b1,b2…b i …b n ), where b i is the dependency evaluation value between the target scene and the i-th data modality, and n is the number of dependency evaluation values;
[0088] Set the processing sub-strategy of the target scenario according to the dependent evaluation value sequence B;
[0089] Generate a training sub-strategy for the target scenario based on the control sub-strategy and the processing sub-strategy for the target scenario;
[0090] Generate training sub-strategies for each demand scenario in turn;
[0091] Establish a training sub-strategy sequence W, W=(w1,w2…w i …w n ), where w i is the training sub-strategy for the i-th demand scenario; n is the number of demand scenarios;
[0092] The training model is constructed according to the training sub-strategy sequence W.
[0093] Specifically, data modalities include but are not limited to visual emotion data, text data, voice data, etc.
[0094] Specifically, the evaluation data package includes the historical interaction data of users corresponding to the target scenario, demand feedback data and collection parameters of each modal data, the types of data collection equipment that can be provided, etc.
[0095] Specifically, the larger the dependency evaluation value, the easier it is to collect data of this type in the target scenario, and the more personalized information is included in the data modality, and the more important it is for personalized training of initial intelligence.
[0096] Specifically, the processing sub-strategy includes feature extraction and fusion strategies for data of different data modalities. By establishing processing sub-strategies for each demand scenario, the efficiency of fusion processing of multimodal data is improved, thereby improving the iterative effect of the initial intelligent agent.
[0097] In a preferred embodiment of the present application, when setting a control sub-strategy for a target scenario based on an evaluation data packet, the following steps are included:
[0098] Generate acquisition device parameters for the target scenario based on the evaluation data packet;
[0099] Set acquisition sub-strategies based on acquisition device parameters;
[0100] Generate a demand evaluation value c for the target scenario based on the evaluation data package;
[0101] Set the monitoring cycle duration according to the demand evaluation value c;
[0102] Generate a control sub-strategy for the target scenario based on the acquisition sub-strategy and the monitoring cycle duration.
[0103] Specifically, the larger the demand evaluation value, the higher the data complexity of the users in the target scenario, the more difficult their personalized training is, and the shorter the corresponding monitoring cycle.
[0104] Specifically, the working parameters of each device are set according to the type of acquisition device in the target scene, thereby generating a corresponding acquisition sub-strategy.
[0105] Specifically, the collection devices include but are not limited to computers, tablets, mobile phones, wearable devices and other types of data terminals.
[0106] Specifically, when generating the demand evaluation value c, it includes:
[0107]
[0108] Among them, e1 is the preset first weight coefficient; e2 is the preset second weight coefficient; Q1 is the preset first fixed coefficient; Q2 is the preset second fixed coefficient; θ1 is the number of data evaluation indicators; β i is the impact factor of the i-th data evaluation index; j i is the reference value of the i-th data evaluation index generated based on the evaluation data packet; θ2 is the number of scene evaluation indicators; η i is the impact factor of the evaluation index of the i-th scene; k i Generates the reference value of the evaluation index of the i-th scene based on the evaluation data package.
[0109] Specifically, data evaluation indicators include but are not limited to multiple parameters such as data volume, difficulty in extracting data features, and data complexity. The influencing factors of each data evaluation indicator can be set according to the influence of the data evaluation indicator on personalized training. The greater the influence on personalized training, the greater the corresponding influence factor.
[0110] Specifically, the scenario evaluation indicators include but are not limited to multiple parameters such as user interaction needs and user interaction frequency. The influencing factors of each scenario evaluation indicator can be set according to the impact on personalized training. The greater the impact on personalized training, the greater the corresponding impact factor.
[0111] Specifically, all parameters in the model are normalized by presetting a first fixed coefficient and a second fixed coefficient, so that each parameter in the model is in the same value range.
[0112] It can be understood that in the above embodiments, a training data package is generated by collecting multimodal data, an initial intelligent agent and multiple user demand categories are constructed based on the training data package, corresponding demand scenarios are established based on different user demand categories, and training sub-strategies corresponding to each demand scenario are set, thereby realizing personalized training needs for different types of users.
[0113] In a preferred embodiment of the present application, when setting the first-level training strategy, it includes:
[0114] Set a in sequence according to the required scenario sequence A i The scene to be compared;
[0115] Generate an associated evaluation value v between the user and the scene to be compared;
[0116] Generate the associated evaluation values of users and each demand scenario in sequence;
[0117] Establish the associated evaluation value sequence V, V=(v1, v2…v i …v n ), where v i is the evaluation value associated with the user and the i-th demand scenario, and n is the number of demand scenarios;
[0118] Set the maximum value v in the associated evaluation value sequence V max The training sub-strategy corresponding to the demand scenario is the first-level training strategy.
[0119] Specifically, the larger the association evaluation value is, the higher the matching degree between the current user and the corresponding demand scenario is.
[0120] Specifically, when generating the associated evaluation value v, it includes:
[0121]
[0122] Among them, θ3 is the number of characteristic evaluation indicators; g i is the influencing factor of the i-th characteristic evaluation index; s i is the similarity evaluation value of the i-th feature evaluation index between the user and the scene to be compared.
[0123] Specifically, the feature evaluation indicators include but are not limited to the type of acquisition device, the content of personalized information in each data mode of the user, and other parameters. The influencing factors of each feature evaluation indicator can be set according to the influence on personalized training. The greater the influence on personalized training, the greater the corresponding influence factor.
[0124] In a preferred embodiment of the present application, when generating a first-level agent according to a first-level training strategy, the process includes:
[0125] Generate multiple iterative data packets based on the first-level training strategy;
[0126] Iteratively train the initial agent according to the iterative data packet;
[0127] Generate a demand evaluation value c' according to the first-level training strategy, and set the initial monitoring cycle duration t1 according to the demand evaluation value c';
[0128] According to the maximum value v in the associated evaluation value sequence V max Set the compensation coefficient α1;
[0129] Set the first-level monitoring cycle duration t, t=α1×t1;
[0130] Set multiple monitoring time nodes according to the first-level monitoring cycle length t;
[0131] Generate an iterative evaluation value for each monitoring time node, and determine whether to output a first-level intelligent agent based on the iterative evaluation value.
[0132] Specifically, the larger the maximum correlation evaluation value, the larger the corresponding compensation coefficient, which ranges from zero to one. By setting the compensation coefficient, the monitoring period can be dynamically adjusted to provide timely warnings of training deviations of the initial intelligent agent, thereby improving the training efficiency of the initial intelligent agent.
[0133] Specifically, when generating the iterative evaluation value of each monitoring point, it includes:
[0134] Get the iteration results and monitoring data packets of the current monitoring time node;
[0135] Generate the iterative evaluation value f of the current monitoring time node;
[0136]
[0137] Among them, e3 is the preset third weight coefficient; e4 is the preset fourth weight coefficient; Q3 is the preset third fixed coefficient; Q4 is the preset fourth fixed coefficient; θ4 is the number of operation evaluation indicators; r 1i is the influencing factor of the i-th operation evaluation index; d i is the reference value of the i-th operation evaluation index generated based on the iteration result; θ5 is the number of monitoring evaluation indicators; r 2i is the influencing factor of the i-th monitoring and evaluation indicator; h i is the reference value of the ith monitoring evaluation indicator generated based on the monitoring data packet; U is the conversion coefficient;
[0138] Generate the output result of the current monitoring time node according to the iterative evaluation value f.
[0139] Specifically, by presetting the conversion coefficient, the larger the reference value of the monitoring and evaluation indicator, the smaller the corresponding iterative evaluation value.
[0140] Specifically, the operation evaluation indicators include but are not limited to the interactive ability of the initial intelligent agent in the current iteration result, user satisfaction and other parameters. The larger the reference value of each operation evaluation indicator, the higher the degree of personalization of the initial intelligent agent. The influencing factors of each operation evaluation indicator can be set according to the influence on the iteration result. The greater the influence, the greater the corresponding influence factor.
[0141] Specifically, the monitoring and evaluation indicators include, but are not limited to, multiple parameters such as the deviation degree of data types, the deviation degree of user requirements, etc. The larger the reference value of each monitoring and evaluation indicator, the worse the effect of the current personalized training. The influence factors of each monitoring and evaluation indicator can be set according to their influence degree on personalized training. The greater the influence degree on personalized training, the greater the corresponding influence factor.
[0142] Specifically, all parameters in the model are normalized by presetting a third fixed coefficient and a fourth fixed coefficient, so that each parameter in the model is within the same value range.
[0143] It can be understood that in the above embodiments, by analyzing the feature data packet of the user, the corresponding first-level training strategy is quickly constructed to improve the training efficiency of the personalized digital device, and the initial agent is iteratively trained by collecting the long-term preference data of the user, so as to construct a first-level agent adapted to the user and improve the user's interaction experience.
[0144] In the preferred embodiment of the embodiment of the present application, when generating the output result of the current monitoring time node according to the iterative evaluation value f, it includes:
[0145] Preset the first iterative evaluation value threshold F1 and the second iterative evaluation value threshold F2, and F1 < F2;
[0146] If f < F1, a first-level correction instruction is generated at the current monitoring time node;
[0147] If F1 < f < F2, a first-level iterative instruction is generated at the current monitoring time node;
[0148] If f > F2, output the first-level agent.
[0149] Specifically, the first-level correction instruction means that the current first-level training strategy cannot complete the personalized training of the initial agent, and it is necessary to timely correct the acquisition sub-strategy and processing sub-strategy in the first-level training strategy, so as to improve the personalized training efficiency of the initial agent.
[0150] Specifically, the first-level iterative instruction means to continue to collect the multi-modal data of the user and iterate the current initial agent.
[0151] Specifically, when the training evaluation value is greater than the preset second training evaluation value threshold, it means that the current initial agent has completed personalized training, and the iteration result is output as the first-level agent.
[0152] According to the first concept of this application, a training data package is generated by collecting multimodal data, an initial intelligent agent and multiple user demand categories are constructed based on the training data package, corresponding demand scenarios are established based on different user demand categories, and training sub-strategies corresponding to each demand scenario are set, thereby realizing personalized training needs for different types of users.
[0153] According to the second concept of this application, by analyzing the user's feature data packets, the corresponding first-level training strategy is quickly constructed to improve the training efficiency of personalized digital devices, and by collecting the user's long-term preference data, the initial intelligent agent is iteratively trained to construct a first-level intelligent agent that adapts to the user and improve the user's interactive experience.
[0154] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and replacements can be made without departing from the technical principles of the present application. These improvements and replacements should also be regarded as the scope of protection of the present application.
Claims
1. A personalized digital device training generation method based on multimodal information fusion, characterized in that: It includes: Generate an initial agent and multiple demand scenarios based on the training data packet, and construct a training model based on all demand scenarios; Obtain the feature data packet of the user, and set the first-level training strategy according to the feature data packet and the training model; Generate a first-level agent according to the first-level training strategy; When generating multiple demand scenarios, it includes: Establish a demand scenario sequence A, A=(a1, a2…a i …a n ), where a i is the i-th demand scenario; n is the number of demand scenarios; When constructing a training model based on all demand scenarios, it includes: Set multiple data modalities according to the training data packet; Establish the data modal sequence P, P=(p1, p2…p i …p m ), where pi is the i-th data modality; m is the number of data modalities; Set a in sequence according to the required scenario sequence A i For the target scene; Obtain the evaluation data packet of the target scenario; Set the control sub-strategy of the target scenario according to the evaluation data packet; Generate the dependency evaluation value between the target scenario and each data modality; Establish a dependent evaluation value sequence B, B=(b1,b2…b i …b n ), where b i is the dependency evaluation value between the target scene and the i-th data modality, and n is the number of dependency evaluation values; Set the processing sub-strategy of the target scenario according to the dependency evaluation value sequence B; Generate the training sub-strategy of the target scenario according to the control sub-strategy and the processing sub-strategy of the target scenario; [[ID= Establish a training sub-strategy sequence W, W=(w1,w2…w i …w n ), where w i is the training sub-strategy for the i-th demand scenario; n is the number of demand scenarios; 2. The personalized digital device training generation method based on multimodal information fusion according to claim 1 is characterized in that: 3. The personalized digital device training generation method based on multimodal information fusion according to claim 2, characterized in that: Among them, e1 is the preset first weight coefficient; e2 is the preset second weight coefficient; Q1 is the preset first fixed coefficient; Q2 is the preset second fixed coefficient; θ1 is the number of data evaluation indicators; β i is the impact factor of the i-th data evaluation index; j i is the reference value of the i-th data evaluation index generated based on the evaluation data packet; θ2 is the number of scene evaluation indicators; η i is the impact factor of the evaluation index of the i-th scene; k i Generates the reference value of the evaluation index of the i-th scene based on the evaluation data package.
4. The personalized digital device training generation method based on multimodal information fusion according to claim 2, characterized in that: Set a in sequence according to the required scenario sequence A i The scene to be compared; Establish the associated evaluation value sequence V, V=(v1, v2…v i …v n ), where v i is the evaluation value associated with the user and the i-th demand scenario, and n is the number of demand scenarios; Set the maximum value v in the associated evaluation value sequence V max The training sub-strategy corresponding to the demand scenario is the first-level training strategy.
5. The personalized digital device training generation method based on multimodal information fusion according to claim 4 is characterized in that: Among them, θ3 is the number of characteristic evaluation indicators; g i is the influencing factor of the i-th characteristic evaluation index; s i is the similarity evaluation value of the i-th feature evaluation index between the user and the scene to be compared.
6. The personalized digital device training generation method based on multimodal information fusion according to claim 4 is characterized in that: According to the maximum value v in the associated evaluation value sequence V max Set the compensation coefficient α1; 7. The personalized digital device training generation method based on multimodal information fusion according to claim 6, characterized in that: Among them, e3 is the preset third weight coefficient; e4 is the preset fourth weight coefficient; Q3 is the preset third fixed coefficient; Q4 is the preset fourth fixed coefficient; θ4 is the number of operation evaluation indicators; r 1i is the influencing factor of the i-th operation evaluation index; d i is the reference value of the i-th operation evaluation index generated based on the iteration result; θ5 is the number of monitoring evaluation indicators; r 2i is the influencing factor of the i-th monitoring and evaluation indicator; h i is the reference value of the ith monitoring evaluation indicator generated based on the monitoring data packet; U is the conversion coefficient; 8. The personalized digital device training generation method based on multimodal information fusion according to claim 7 is characterized in that:
Citation Information
Patent Citations
Large model-based training method and related device
CN119905206A