Interaction method for driving digital human language understanding and corresponding reaction through artificial intelligence algorithm
By constructing environment-enhanced text vectors and a multi-module collaborative system, the digital human system improves response accuracy and naturalness in noisy and mobile scenarios, solves the problem of inaccurate response in existing technologies, and realizes modular expansion and multi-channel output.
Patent Information
- Application Number
- CN202511014941.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2025-11-21
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing digital human systems have poor accuracy and naturalness in response to noisy environments, mobile scenarios, or specific vertical fields, and are difficult to support modular expansion and multi-channel feedback.
By collecting users' historical language data, text data, action commands, and physical environment information, an environment-enhanced text vector is constructed. This vector is then combined with a pre-trained language model for secondary pre-training to generate optimized language model parameters. A named entity recognition model and an intent recognition model are then constructed. Finally, a structured knowledge graph and a response decision model are combined to achieve multimodal response.
It enhances the understanding and response capabilities of digital humans in complex scenarios, improves the accuracy and naturalness of responses, supports task switching and multi-channel output in different scenarios, and enhances the system's language understanding depth and decision-making reasoning capabilities.
Smart Images

Figure CN120995379A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of electric digital data processing, and in particular to an interaction method for language understanding and corresponding reaction of digital human driven by artificial intelligence algorithm. BACKGROUND
[0002] As an important carrier of human-computer interaction, digital human has gradually evolved from a traditional script-driven model to an interactive system with certain intelligent response capabilities. In recent years, pre-trained language models such as BERT and GPT have shown excellent performance in understanding language semantics, driving the precision of tasks such as multi-turn dialogue generation, intent recognition, and entity extraction. At the same time, the integration of speech recognition (ASR), environmental perception (such as location, time, and noise recognition), and multi-modal learning provides a technical foundation for building context-sensitive and scene-adaptive interactive systems.
[0003] Existing technologies mostly focus on modeling language itself, lacking deep integration with the physical environment (such as time, location, and situational state) in which the user is located, resulting in poor response accuracy and naturalness of digital human in noisy environments, mobile scenarios, or specific vertical fields. In addition, most systems use a single model to output end-to-end, making it difficult to support modular expansion, scene switching, and multi-channel feedback (such as speech, action, and text) collaborative output. SUMMARY
[0004] Therefore, it is necessary to provide an interaction method for language understanding and corresponding reaction of digital human driven by artificial intelligence algorithm to solve at least one of the above technical problems.
[0005] To achieve the above purpose, an interaction method for language understanding and corresponding reaction of digital human driven by artificial intelligence algorithm includes the following steps:
[0006] Step S1: Collecting historical language data, text data, action instruction data, and physical environment information of the user to obtain structured training data; performing vectorization processing on the text data to obtain a text semantic vector; extracting an environmental feature vector from the physical environment information; and concatenating the text semantic vector and the environmental feature vector to obtain an environment-enhanced text vector;
[0007] Step S2: Constructing a pre-trained language model based on the environment-enhanced text vector and performing secondary pre-training to generate optimized language model parameters; and constructing and training a dialogue generation model based on the optimized language model parameters and the historical language data;
[0008] Step S3: Performing BIO labeling on the text data based on the structured training data, and training a named entity recognition model through the pre-trained language model to generate an entity recognition result; and constructing a structured knowledge graph based on the entity recognition result and a pre-set external knowledge base;
[0009] Step S4: based on the optimized language model parameters and the environment enhanced text vector, a multi-classification intention of the recognized user input text is identified, and an intention recognition model is constructed, obtaining a user intention label; an intention-action mapping table is constructed based on historical language data and action instruction data, and a pre-training rule-learning hybrid reaction decision model is constructed based on a structured knowledge graph;
[0010] Step S5: based on the user intention label, the entity recognition result and the environment feature vector, the user input intention is explained, and a structured intention representation is obtained;
[0011] Step S6: based on the structured intention representation, the dialogue generation model and the reaction decision model, a multi-modal response action is designed and visualized, and a digital human interaction feedback result is obtained.
[0012] The present application effectively improves the understanding and response ability of digital human in complex and variable scenes by introducing the way of environment enhanced text vector, which deeply integrates user language and physical environment. Compared with the traditional interactive system which takes language modeling as the core, this scheme fully utilizes the time, location, noise and other environmental information of the user, realizes intelligent response of context perception, and helps to improve the accuracy and naturalness of digital human in noisy, mobile or specific vertical fields. At the same time, by constructing a multi-module cooperative system architecture, the modular expansion and flexible adaptation of the model are realized, which provides technical support for supporting task switching, multi-channel output (such as voice, action, text, etc.) and multi-modal fusion in different scenes. In addition, combined with pre-training language model and knowledge graph, intention recognition, entity extraction, rule learning and other intelligent components, the language understanding depth and decision reasoning ability of the system are effectively enhanced, realizing more targeted and coherent interactive feedback. Overall, this method improves the intelligence, scalability and user experience of digital human in human-computer interaction. BRIEF DESCRIPTION OF DRAWINGS
[0013] Other features, objects and advantages of the present application will become more apparent from the following detailed description of non-limiting embodiments made with reference to the accompanying drawings:
[0014] Figure 1 The step flowchart of the interactive method of the artificial intelligence algorithm driven digital human language understanding and corresponding reaction of the present application;
[0015] Figure 2 The detailed step flowchart of step S1 in the present application; Figure 1 The detailed step flowchart of step S2 in the present application;
[0016] Figure 3 The detailed step flowchart of step S3 in the present application; Figure 1 The detailed step flowchart of step S2 in the present application. DETAILED DESCRIPTION
[0017] The technical method of the present application will be described clearly and completely below in conjunction with the drawings. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0018] In addition, the drawings are only schematic illustrations of the present application and are not necessarily drawn to scale. Identical reference numerals in the drawings represent identical or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities, which do not necessarily have to correspond to physically or logically independent entities. The functional entities can be implemented in the form of software, or in one or more hardware modules or integrated circuits, or in different network and / or processor methods and / or microcontroller methods.
[0019] It should be understood that although the terms "first", "second", etc. can be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, without departing from the scope of the exemplary embodiments, a first element can be called a second element, and similarly a second element can be called a first element. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0020] To achieve the above-mentioned purpose, please refer to Figures 1 to 3 The present application provides an interactive method of artificial intelligence algorithm driven digital human language understanding and corresponding reaction, comprising the following steps:
[0021] Step S1: Collecting historical language data, text data, action instruction data and physical environment information of the user to obtain structured training data; performing vectorization processing on the text data to obtain a text semantic vector; extracting an environment feature vector from the physical environment information; splicing the text semantic vector and the environment feature vector to obtain an environment enhanced text vector;
[0022] Step S2: Constructing a pre-training language model based on the environment enhanced text vector, and performing secondary pre-training to generate optimized language model parameters; constructing and training a dialogue generation model based on the optimized language model parameters and the historical language data;
[0023] Step S3: Performing BIO labeling on the text data based on the structured training data, and training a named entity recognition model through the pre-training language model to generate an entity recognition result; constructing a structured knowledge graph based on the entity recognition result and a preset external knowledge base;
[0024] Step S4: based on the optimized language model parameters and the environment enhanced text vector, a multi-classification intention of the recognized user input text is identified, and an intention recognition model is constructed to obtain a user intention label; an intention-action mapping table is constructed based on historical language data and action instruction data, and a pre-training rule-learning hybrid reaction decision model is constructed based on a structured knowledge graph;
[0025] Step S5: based on the user intention label, the entity recognition result and the environment feature vector, the user input intention is explained to obtain a structured intention representation;
[0026] Step S6: based on the structured intention representation, a dialogue generation model and a reaction decision model, a multi-modal response action is designed and visualized to obtain a digital human interaction feedback result.
[0027] In the embodiment of the present application, reference is made to Figure 1 The flowchart of the steps of the artificial intelligence algorithm driven digital human language understanding and corresponding reaction interaction method of the present application, in the present example, the artificial intelligence algorithm driven digital human language understanding and corresponding reaction interaction method comprises the following steps:
[0028] Step S1: collect the user's historical language data, text data, action instruction data and physical environment information to obtain structured training data; the text data is vectorized to obtain a text semantic vector; the physical environment information is extracted to obtain an environment feature vector; the text semantic vector and the environment feature vector are spliced to obtain an environment enhanced text vector;
[0029] In the embodiment of the present application, first, the original voice data input by the user through voice in the past 7 consecutive days is collected by the microphone module, the sampling frequency is fixed at 16 kHz, the collection time is not less than the complete cycle of each voice instruction, and the front and rear mute is trimmed using the endpoint detection method (VAD); the text data input by the user through touch screen or voice in the application interface is collected synchronously, each text data needs to meet the condition that the number of characters is not less than 8 Chinese characters, and contains 3 types of instruction, query and casual chat; in addition, the action instruction data triggered by the user in the interface is extracted through the operation log recording module embedded in the terminal device, the recording dimensions include instruction type (such as click, slide), response component ID and operation timestamp, and the physical environment information in the corresponding scene is collected, including latitude and longitude data output by the GPS module (update frequency 1 Hz, accuracy within 5 meters), local timestamp recorded by the system clock (accuracy to minute), light intensity value read by the light sensor (unit Lux, sampling range 0-10000), environmental noise level measured by the microphone (unit dB, dynamic range 30-100 dB) and movement state information obtained by the acceleration sensor (whether in motion is judged by three-axis synthesis); then, the collected text data is subjected to Unicode standard decoding and word segmentation processing, the maximum forward matching method based on dictionary is called to divide the Chinese word boundary, the word frequency matrix is constructed after deleting the stop words, and the weight vector of each word item is calculated by using the word frequency-inverse document frequency product, and the text semantic vector with a dimension of 300 is constructed; next, the collected physical environment information is normalized according to the field respectively, the GPS coordinates are converted into three classification vectors of city, street and place type, the light intensity value and the noise level are mapped into low, medium and high three levels according to the set interval, the timestamp is converted into four time period encodings of morning, day, evening and night, and the movement state is represented by 0 / 1 binary whether in motion, and the above encoding results are combined to form an environment feature vector with a length of 50; finally, the text semantic vector with a length of 300 and the environment feature vector with a length of 50 are spliced in the direction of vector dimension to generate an environment enhanced text vector with a length of 350, and stored in the structured data set respectively with the user unique identifier as the primary key.
[0030] Step S2: constructing a pre-training language model based on the environment enhanced text vector, and performing secondary pre-training to generate optimized language model parameters; constructing and training a dialogue generation model based on the optimized language model parameters and historical language data;
[0031] In the embodiment of the application, first, the environment enhanced text vector generated in step S1 is called and input format construction processing is performed thereon, wherein each sample contains a set of text semantic vectors (dimension 300) and a set of environment feature vectors (dimension 50), a joint representation sequence with a dimension of 350 is obtained by vector splicing, each sequence is limited to a maximum length of 128 tokens, and the part exceeding the length is truncated, and the part with insufficient length is padded with zero vectors; then, based on the structured training data containing industry terms, query sentences and user casual corpus, double-task processing is performed on the joint representation sequence, the first task is a mask language modeling task, wherein 15% of the tokens in each sample are randomly selected, 80% of which are replaced with [MASK] symbols, 10% of which are replaced with random tokens, and 10% of which remain unchanged, the goal being to restore the original tokens under the context condition; the second task is a sentence continuity prediction task, the training data is organized into a sentence pair format, the first sentence comes from the user input text, and the second sentence comes from the historical dialogue corpus, the positive sample is a context coherent sentence pair, and the negative sample is a context incoherent or environment label inconsistent sentence pair, the goal being to predict whether the sentence pair is coherent and environment consistent according to the overall vector of the sentence pair; the prediction results of the above two types of tasks are calculated by a softmax probability distribution to obtain a loss value, the mask prediction loss adopts a cross-entropy function, the sentence pair prediction loss adopts a binary classification logarithmic loss function, and the two types of losses are added according to a weight ratio of 0.5:0.5 to form a total loss value, all Transformer encoding layer parameters are updated through gradient back propagation, the number of iterations is set to 30 rounds, each round contains 2000 batches, and the number of samples in each batch is 64; after the loss value decreases and tends to be stable, the parameters of all encoding layers and embedding layers are extracted to form an optimized language model parameter set; then, the optimized language model parameter set is used to construct dialogue training corpus from historical language data, wherein each corpus contains user input text, historical context, environment label and standard reply, the maximum length of the input sequence is limited to 256 tokens, the input data is vector enhanced by using position encoding and paragraph encoding, and a training sample pair (input→reply) is constructed; then, the decoder module in the Transformer structure is used to model the reply sequence, the decoder contains 6 layers of self-attention units, each layer contains 8 attention heads, each attention head has a dimension of 64, the input vector is encoded-decoded cross-operated, the difference between the predicted reply sequence and the real reply is minimized, all decoder parameters are optimized, the training process is performed for a total of 20 rounds, each round contains 1000 sample batches, and after completion, a dialogue generation model with environment perception capability is obtained, the encoder parameters in the model come from the optimized language model parameter set, the decoder parameters are the current training result, and the model is finally stored in a tensor format.
[0032] Step S3: BIO label the text data based on structured training data, and train a named entity recognition model through a pre-trained language model to generate an entity recognition result; and construct a structured knowledge graph based on the entity recognition result and a preset external knowledge base;
[0033] In the embodiment of the present application, first, the text data in the structured training data constructed in step S1 is selected, each text sentence is segmented according to the word granularity, and manual correction is performed in combination with the corresponding intent label, entity label and environment label to ensure that all entity boundaries are accurate and entity types are unique. Subsequently, the BIO tagging method is used for coding, wherein B represents the start of the entity, I represents the internal entity, and O represents the non-entity, and the entity categories are limited to five categories of time, place, organization, product and person. The BIO tagging result of each text is stored in the form of an equal-length label sequence; after the labeling is completed, the optimized language model parameters generated in step S2 are called to perform vectorization processing on the labeled text. The input text is segmented into a token sequence by a tokenizer, the total number of tokens is limited to no more than 128, the embedding layer and the 6-layer Transformer coding layer are used to generate the context semantic representation of the corresponding position, each token is finally output as a vector with a dimension of 768 to form a semantic representation enhanced sequence vector; then, the above vector and the BIO label sequence are aligned one by one to construct a training sample pair, and a conditional judgment mechanism is used to predict the entity category according to the token order, wherein a probability vector is output at each position, the category corresponding to the maximum probability is selected as the prediction result, the label cross-validation method is used to compare the predicted label and the real BIO label, the accuracy is calculated, and the parameters are updated in reverse. The training batch is set to 1000 groups, each group contains 32 samples, and the training period is set to 20 rounds; after the prediction accuracy is stable, the embedding layer and the output prediction weight are saved as the named entity recognition model parameter set, which is used for subsequent entity extraction; subsequently, the named entity recognition model is used to process the real-time input text of the user of the user terminal. The input text is uniformly converted into a token sequence, the BIO label prediction sequence is generated through the encoder and the output module, the entity recognition result list is constructed by combining the continuous B / I tagging positions, and the start and end positions, the original text and the entity type corresponding to each entity are recorded; finally, the established external knowledge base is called, the knowledge base is organized in the form of a triple, contains entity name, belonging category and semantic relationship information, each entity name in the recognition result is matched with the standard entity node in the knowledge base through string matching and semantic association judgment, the Jaccard text similarity calculation method is used, the matching threshold is set to 0.75, and the matched entity is included in the target node set; for each target node, its upper and lower concept relationship and semantic adjacent node are further searched, a "entity-relation-entity" structure triple is constructed, and the entity node is additionally attached with its corresponding geographical position label, environment state label and time period label. Finally, all triples are organized in a graph structure to construct a structured knowledge graph.
[0034] Step S4: based on the optimized language model parameters and the environment enhanced text vector, a multi-classification intention of the recognized user input text is identified, and an intention recognition model is constructed to obtain a user intention label; an intention-action mapping table is constructed based on historical language data and action instruction data, and a pre-training rule-learning hybrid reaction decision model is constructed based on a structured knowledge graph;
[0035] In the embodiment of the application, first, the optimized language model parameters generated in step S2 are applied to the semantic representation construction of the input text, and the environment enhanced text vector obtained in step S1 is combined to generate a semantic vector with a dimension of 768 after each user input text is segmented and encoded, then the corresponding environment feature vector (dimension 50) is spliced into a joint input vector with a length of 818; subsequently, the samples labeled as explicit intent categories in the historical language data are subjected to multi-classification label arrangement, the intent types are limited to five categories of "query information", "request operation", "express emotion", "idle chat response" and "control instruction", all training samples are subjected to intent labeling by using a label mapping table, a training set is constructed, an input-label pair is generated after each sample is subjected to vector encoding, a softmax multi-classification mechanism is used for training, a five-dimensional vector is output, the category corresponding to the maximum probability position is taken as the prediction result, the accuracy is evaluated by using a cross-validation method, the training batch is set to 2000 groups, the number of samples in each group is 64, the training period is set to 15 rounds, after the training is completed, all parameters are packaged into an intent recognition module for analyzing the intent label corresponding to the current input text of the user; then, based on the historical language data and action instruction data collected in step S1, the language input and interface action in the same time period are subjected to timestamp matching, "intent expression-action execution" sample pairs are extracted, the dimensions include intent label, operation component ID, action type and response delay, an intent-action reference table is established, and a multi-version mapping table is generated by grouping according to the user geographic location, time period and device type; then, the structured knowledge graph constructed in step S3 is called, the associated relationship edges and upper and lower hierarchical relationships of each target entity node therein are searched, an entity semantic path is constructed, an explicit relationship rule set is constructed through three dimensions of path length, relationship type and context semantic distance, each rule is defined as "if the intent is X, the entity is Y, and the environment is Z, then the recommended action is A", and all rules are organized into a rule set table in a unified format; finally, the user feedback records collected in step S1 are counted, the positive feedback threshold is set to 0.8, the negative feedback threshold is set to 0.3, the association between the user behavior and the action sequence is extracted, a state-action-feedback triple is constructed, the rule set and the feedback behavior record are jointly analyzed, the action selection weight is set for the scene with multiple candidate actions, and a reward update strategy is set, the weight is dynamically adjusted according to the feedback in the continuous interaction process, a reaction decision structure containing an explicit mapping rule and a dynamic weight update mechanism is constructed, a pre-trained rule-learning hybrid reaction decision model is generated, the structure is composed of a rule engine, an intent mapping table, an environment adaptation rule and a behavior feedback module, and a recommended action sequence under the user intent and environment is uniformly output.
[0036] Step S5: interpreting the user input intent based on the user intent label, the entity recognition result and the environment feature vector to obtain a structured intent representation;
[0037] In the embodiment of the application, first, the user intent label identified in step S4 is called, and the label is converted into a one-hot encoding vector with a fixed length of 5 according to a preset five-category intent label coding rule, as the first part of the structured input; second, the entity recognition result list output in step S3 is obtained, each entity is classified according to its label type (time, place, organization, product, and person), and numbered according to the order of appearance in the text, an entity information table is constructed using the position index and entity type label, and the entity original text, belonging category, position index in the sentence, and corresponding knowledge graph node ID are extracted as fields to construct an entity structure unit, and each input is limited to contain at most 3 entities, if the entity exceeds 3, only the first 3 in the sentence are retained; then, the environment feature vector generated in step S1 is called, the vector is split into 5 categories of sub-features according to the fields, which are geographic location label, time period label, illumination level, noise level, and user motion state, each category of sub-feature is encoded and reorganized to form an environment description structure unit with a fixed length of 50; next, the above three parts of information, i.e., the one-hot encoding intent vector, the entity information table, and the environment description structure unit, are aligned at the field level, a unified structure is generated through field name mapping and sequential numbering, and the fields include intent type, main entity, entity category, context position, entity node, time label, place label, noise level, illumination level, and motion state, a total of 10 fields; then, the structure fields are logically consistent, and the rules include: if the intent type is “query information”, the entity category must exist and cannot be O label; if the entity node has no corresponding relationship path in the knowledge graph, mark the field as invalid entity, and set it to null value later; if the time label is “night” and the intent is “request operation”, add a “need to mute output” field; after all the checks are passed, the final structured intent representation is generated, which is encoded in JSON format, the field names and values are organized in standard key-value pairs, and a unique number and timestamp are recorded.
[0038] Step S6: Based on the structured intent representation, the dialogue generation model, and the reaction decision model, multi-modal response actions are designed and visualized to obtain the digital human interaction feedback result.
[0039] In the embodiment of the application, first, the structured intention representation generated in step S5 is called, the intention type field therein is matched and searched, and it is input as an instruction target type to a dialogue content construction module. The main entity field corresponding to the intention and the environment label such as time and place are combined to form a context input sequence, the maximum length of the input sequence is limited to 256 characters, and after being uniformly encoded, it is transmitted to the encoding end of the dialogue generation model generated in step S2. The context fusion semantic representation vector is obtained through the encoder output, and then the decoder performs conditional text generation operation to generate a corresponding text reply. The text length is limited to within 80 words, and the output result is a pure text string. Then, the reaction decision model constructed in step S4 is called, all fields in the structured intention representation are taken as input conditions, the intention type and the environment state are rule searched, the corresponding action option list is matched, and the action selection weight in the feedback score record under the similar historical situation is combined to give a priority score to each candidate action. The highest scoring item is taken as the action instruction output, and the action type is limited to one of interface operation, voice playback, device control and visual feedback. Then, the text reply and the selected action instruction are integrated into a multi-modal response task package, and the response package field is constructed, including: reply text, response action type, target component number, execution parameter and output channel type, which are uniformly encoded in JSON format. The channel type is determined according to the noise level and motion state fields recorded in step S1. If the noise level is higher than 70 dB or the motion state is in motion, the output channel is set to “text + interface feedback”, otherwise it is set to “voice + animation feedback”. Subsequently, the response task package is transmitted into the interaction rendering engine, and the branch processing is performed according to the channel type. If the channel type is “text + interface feedback”, the interface rendering module is called to load the reply text into the dialogue window, and the graphical control instruction is called to drive the digital human animation module to execute facial expression and action gesture. If the channel type is “voice + animation feedback”, the text reply is converted into an audio stream through the TTS speech synthesis module, and the natural speech is synthesized at a sampling rate of 22050 Hz. The audio stream is transmitted to the loudspeaker module, and the synchronous animation is loaded on the digital human screen at the same time. The rendering action is controlled by the execution parameter in the response package, and the corresponding animation execution script is limited to an XML structure file, and the frame rate is set to 30 frames per second. Finally, the user response record, output content log and feedback behavior mark of the above interaction process are collected and stored in the user behavior database, and the fields include intention identification, response content, action type, execution channel, interaction time consumption and satisfaction mark.
[0040] The application effectively improves the understanding and response ability of the digital person in complex and variable scenes by deeply integrating the user language with the physical environment through the introduction of the environment enhanced text vector. Compared with the traditional interaction system which takes language modeling as the core, this scheme fully utilizes the time, location, noise and other environmental information of the user, realizes intelligent response with context awareness, and helps to improve the accuracy and naturalness of the digital person in noisy, mobile or specific vertical fields. At the same time, by constructing a multi-module collaborative system architecture, the modular expansion and flexible adaptation of the model are realized, which provides technical support for supporting task switching, multi-channel output (such as voice, action, text, etc.) and multi-modal fusion in different scenes. In addition, combined with pre-training language models and knowledge graphs, intent recognition, entity extraction, rule learning and other intelligent components, the language understanding depth and decision reasoning ability of the system are effectively enhanced, realizing more targeted and coherent interaction feedback. Overall, this method improves the intelligence, scalability and user experience of the digital person in human-computer interaction.
[0041] Preferably, step S1 comprises the following steps:
[0042] Step S11: Collecting historical language data, text data, action instruction data and physical environment information of the user end user to obtain original multi-modal historical interaction data;
[0043] Step S12: Data cleaning, corpus classification and multi-label artificial annotation are performed on the original multi-modal historical interaction data to obtain structured training data;
[0044] Step S13: The text data in the structured training data is subjected to word segmentation, tokenization and context encoding processing to obtain a text semantic vector;
[0045] Step S14: Feature engineering processing is performed based on the physical environment information to obtain an environmental feature vector;
[0046] Step S15: The text semantic vector and the environmental feature vector are spliced in the vector dimension to obtain an environment enhanced text vector.
[0047] In the embodiment of the application, firstly, the voice collection module embedded in the mobile terminal records all voice input data of the user terminal user in the continuous use period, the sampling frequency is set to 16 kHz, the audio format is PCM, single channel, 16-bit depth, the endpoint detection method based on energy and short-time zero-crossing rate joint detection is used for voice segment division, then the text data input by the user in the application interface, the action instruction data executed and the physical environment information obtained by the terminal multi-sensor system are synchronously collected, wherein the action instruction data includes component ID, trigger mode (click, slide, long press) and system response code, the physical environment information includes GPS coordinates (WGS-84 format, refresh frequency 1 Hz), illumination intensity (unit: Lux, sensor range: 0-10000), noise intensity (unit: dB, range: 30-100), timestamp (format: yyyy-MM-dd HH:mm) and three-axis acceleration value; then, in step S12, the original multi-modal historical interaction data collected is subjected to data cleaning operation, the abnormal samples with noise greater than 90 dB, illumination value of 0 and no operation behavior are deleted, the text is merged according to the dialogue round, the corpus data is manually annotated according to the intent type (query type, control type, emotion type), entity type (time, place, organization) and environment type (quiet / noisy, day / night), the multi-label annotation process is cross-checked by 3 people, and finally the structured training data containing text content, label index, voice path and environment data field are generated; in step S13, the text field in the structured training data is processed, firstly, the text is subjected to word segmentation processing by using the dictionary maximum forward matching word segmentation method, tokenization coding is performed after removing stop words, the maximum length of each text is limited to 128 tokens, the insufficient part is filled with [PAD] mark, the excess part is truncated, then position coding (using linear incremental method) and segment coding are added to each token, and the text semantic vector with a dimension of 300 is obtained; in step S14, the feature engineering operation is performed on the environment data, the GPS coordinates are mapped to 8-bit geographic code by GeoHash method, the illumination intensity is divided into five levels of 0-100, 101-500, 501-2000, 2001-5000 and 5001-10000, each level is coded as a 1-bit one-hot vector, the noise intensity is divided into four levels of 30-50, 51-70, 71-90 and 91-100, the timestamp is converted into four time interval (morning, day, evening, night) encodings, and the acceleration synthesis value is greater than the set threshold 2.0 is defined as "moving", otherwise defined as "still", the above encoding results are merged to form an environment feature vector with a length of 50; finally, in step S15, the text semantic vector with a length of 300 generated in step S13 is spliced with the environment feature vector with a length of 50 generated in step S14 in the vector dimension direction, and an environment-enhanced text vector with a length of 350 is obtained after splicing, and is output to the sample data buffer area in the JSON structured format.
[0048] The present application significantly improves the richness and representation ability of data through multi-modal collection and processing of user historical interaction data, and constructs a deep semantic representation with language, behavior and environment characteristics, laying a solid foundation for subsequent intelligent model training. Through systematic data cleaning and labeling process, the quality and label accuracy of training data are effectively improved, the influence of noise interference on model performance is reduced, and the usability and generalization ability of training data are enhanced. At the same time, the representation method of text and environment information fusion makes the generated environment-enhanced text vector can more comprehensively reflect the semantic intention and interaction scene of the user, which helps to improve the sensitivity and accuracy of the model to the context. In addition, this method realizes the unified encoding of multi-modal information, provides context-consistent and clear-structured input expression for subsequent intelligent tasks (such as intent recognition, dialogue generation, behavior prediction, etc.), and is beneficial to build more robust, efficient and scene-adaptive human-computer interaction systems.
[0049] Preferably, step S2 comprises the following steps:
[0050] Step S21: constructing an input representation based on the environment-enhanced text vector to obtain a joint representation sequence;
[0051] Step S22: processing the joint representation sequence based on the structured training data to perform a masked language model and next sentence prediction training task to obtain a pre-trained language model;
[0052] Step S23: performing secondary pre-training optimization based on the pre-trained language model to obtain optimized language model parameters;
[0053] Step S24: performing vector encoding and training sample construction based on the optimized language model parameters and historical language data to obtain a context-aware dialogue training corpus;
[0054] Step S25: performing sequence modeling and Transformer decoder optimization on the context-aware dialogue training corpus, and training a dialogue generation model to obtain the dialogue generation model.
[0055] In the embodiment of the application, first, in step S21, the environment-enhanced text vector generated in the previous step S15 is called, each vector has a length of 350, of which the first 300 dimensions are text semantic vectors and the last 50 dimensions are environment feature vectors, the vector is subjected to fixed position coding and segment coding operation according to the maximum token number 128, the position coding adopts a trigonometric function mode, the segment coding distinguishes the current sentence and the context sentence position, and after the construction is completed, the joint representation sequence is obtained, each input unified format is a two-dimensional tensor with a length of 128 and a dimension of 350; in step S22, the joint representation sequence is combined with the structured training data generated in step S12 to form a training sample, the training sample contains sentence pair structure and label pair structure, and two task operations are performed: the first task is a mask word prediction task, 15% of the tokens in each sentence are randomly selected, of which 80% are replaced with [MASK], 10% are replaced with random tokens, and 10% remain the same, the target is to predict the original token, a softmax output structure is adopted, and the label is the original token index; the second task is a next sentence relationship prediction task, a sentence pair sample is constructed from the structured data, wherein the positive sample satisfies the continuous context semantics and consistent environment label, and the negative sample is a sentence pair that is not continuous or has conflicting environment labels, the target is to predict continuity and environment consistency, the output is a binary probability distribution, and a binary classification logarithmic loss function is adopted; the loss of the two task outputs is weighted and added in proportion of 1:1 to form a total loss function, 2000 batches are set for each round of training, the number of samples per batch is 64, the total number of training rounds is 30, the gradient is calculated through back propagation and the joint representation coding parameters and output prediction parameters are updated, and after the training is completed, a pre-trained language model containing a 6-layer Transformer encoder, 12 attention units per layer and 64 dimensions per attention head is obtained; in step S23, based on the pre-trained language model obtained in step S22, all coding parameters are loaded, sentence input is re-inferred and verified, a perturbation sample set including newly added synonymous sentences, structure disturbance sentences and label replacement sentences is constructed for stability training, 15 rounds of training are performed again, 1000 batches are set for each round, the number of samples is maintained at 64, the mask proportion in the perturbation sentence is still 15%, the output target is consistent with the original task, and finally the parameters of all coding layers and embedding layers are evaluated through gradient convergence to form optimized language model parameters; in step S24, the optimized language model parameters are called, and historical language data (including user input text, historical context text, environment label and standard reply) is subjected to vector coding, the coding mode of each sample is to splice input text tokens, historical dialogue tokens and environment coding tokens, a context-enhanced sequence with a maximum input sequence length of 256 is constructed, and the output is a multi-round context semantic representation vector in a tensor format, combined with the corresponding standard reply text, a context-reply pair sample set is formed, named as a scenario-aware dialogue training corpus, and stored in a JSON array format.Finally in step S25, the context-aware dialogue training corpus is input into the dialogue generation module of the Transformer structure, the encoding end directly uses the optimized language model parameter locking structure, the decoding end adopts a 6-layer stacked structure, each layer contains 8 attention heads, adopts residual connection and feedforward transformation, and outputs a decoding text sequence, wherein the output of each time step is selected by softmax, the character cross-entropy loss of the predicted reply and the real reply is calculated, the training process is set to 20 rounds, 1000 batches per round, and the number of samples per batch is 32, and finally the dialogue generation model containing the encoder and decoder structure is derived, and the output model is saved in tensor format and bound with the model version number and timestamp.
[0056] The application introduces an environment-enhanced text vector to construct a joint representation sequence, so that the language model has sensitivity to environmental information in the pre-training stage, thereby improving the understanding ability of the model to the real context of the user. The combination of mask language modeling and next sentence prediction task effectively enhances the model's grasp of semantic continuity and context logic, laying a firmer foundation for dialogue generation. Through secondary pre-training, the language model parameters are further optimized to better fit specific interactive scenarios and user behavior characteristics, thereby improving the generalization ability and stability of the model in actual application. Based on the optimized model parameters, context-aware training corpus is constructed, which significantly improves the adaptability of the dialogue system to changes in user intent and context transfer, and helps to generate more natural, coherent and targeted response content. In addition, by optimizing the decoder mechanism of the Transformer structure, the modeling ability of the model for long-distance dependencies is strengthened, so that the dialogue system can still maintain semantic coherence and consistent strategy in complex scenarios, and the interactive intelligence level and user experience quality of the digital person are comprehensively improved.
[0057] Preferably, step S22 comprises the following steps:
[0058] Step S221: performing slicing processing on the joint representation sequence to construct input sentence pairs suitable for language modeling tasks, and obtaining a training task sentence pair set;
[0059] Step S222: labeling and grouping the training task sentence pair set based on the structured training data, constructing positive and negative sample combinations, and obtaining a supervised training sample set;
[0060] Step S223: performing random mask operation on the supervised training sample set, and performing MASK replacement, random replacement and retention operation on token, to obtain a mask input sequence;
[0061] Step S224: performing continuity and environment consistency judgment on the supervised training sample set, constructing sentence pair labels, and obtaining NSP input samples and labels;
[0062] Step S225: perform semantic encoding processing on the mask input sequence to obtain a context fusion representation sequence;
[0063] Step S226: perform prediction output modeling on the context fusion representation sequence, and perform prediction of the masked token and sentence pair continuity judgment using the prediction output modeling result to obtain a mask prediction result and a next sentence prediction result;
[0064] Step S227: calculate a joint loss function for the mask prediction result and the next sentence prediction result to obtain a language model multi-task training target;
[0065] Step S228: perform back propagation and parameter updating on the mask language model based on the joint loss function, iteratively optimize the model structure and weight parameters, and obtain a pre-trained language model.
[0066] In the embodiment of the application, the joint representation sequence generated in step S21 is divided according to a fixed length of 128, each 128 tokens are a segment, the part exceeding the length is truncated, the part insufficient is filled with a [PAD] mark, and pairing is performed in sequence to construct a sentence pair set containing a preceding-sentence structure, each sentence pair is composed of a preceding sentence (A) and a following sentence (B), and the semantic integrity between sentences and the consistency of the environment label of each sample are ensured; based on the structured training data generated in step S12, a label is added to each sample in the sentence pair set, and the samples are grouped according to two dimensions of context logical continuity and environment label consistency, the positive sample requires logical continuity between the preceding sentence and the following sentence in the sentence pair and consistency of the time, place and environment state field, and the negative sample requires construction under the conditions of semantic rupture, place change or noise level mutation, all the sentence pairs and labels constitute a supervised training sample set, each group of samples contains an input sentence pair and a label value (1 for positive and 0 for negative); a mask operation is performed on the token sequence of the preceding sentence and the following sentence in each group of supervised training samples, 15% of the tokens in each sample are randomly selected, 80% of the positions are replaced with a [MASK] mark, 10% of the positions are replaced with a random token in the vocabulary, and 10% of the positions remain unchanged, a masked input sequence is generated after the operation, and the original token is recorded as a supervised prediction target; the label labeled in step S222 is read, context logical checking and environment feature comparison judgment are performed on each pair of sentences, the sentence pair that is continuous and consistent in environment is marked as a positive example, the rest is a negative example, and the label and the sentence pair jointly constitute an NSP input sample, the fixed length of each sample is 256, a [SEP] mark is added between the sentences, and a [CLS] mark is added at the beginning of the sentence; a 6-layer Transformer encoding structure is used to perform context semantic encoding on the masked input sequence, each token input includes a word embedding vector, a position encoding and a sentence segment encoding, and the output is a 768-dimensional context fusion representation sequence corresponding to each token, wherein the output vector of the [CLS] position is used for subsequent sentence pair continuity judgment, and the other token positions are used for mask prediction tasks; two prediction tasks are performed on the context fusion representation sequence, which are a mask token prediction task and a next sentence continuity judgment task, the former predicts the vocabulary probability distribution of each masked position by using softmax, and the latter performs linear transformation on the [CLS] position vector and outputs a binary classification result after normalization, obtaining a mask prediction result and a sentence pair prediction result; the loss values of the two task results are calculated respectively, the cross-entropy loss function is used for the mask prediction task, the deviation between the predicted distribution and the real token is calculated for each masked token, the logistic regression loss function is used for the next sentence prediction task, and the binary logarithmic loss of the matching degree of the predicted label and the actual label is calculated, and finally the losses of the two are combined with equal weights to form a joint loss function; error back propagation is performed using the joint loss function value to calculate the gradient values of all embedding layers, attention layers and output layers, and the gradient values are calculated according to a set learning rate of 0.001The parameter is updated, each round of training contains 2000 batches, each batch contains 64 samples, and the training is continued for 30 rounds until the training error converges. After completing the optimization of the entire structure and parameters, the Embedding matrix, Transformer encoding parameters and output prediction module are derived, and finally a pre-trained language model with semantic understanding and environment consistency judgment ability is formed. The model structure and parameters are saved as a standard tensor format file, and the version number and training timestamp are bound.
[0067] The application combines the training task sentence pair and the positive and negative sample group to make the pre-trained language model perceive the semantic continuity and environmental consistency between languages in the training stage, effectively enhancing the modeling ability of the model to the context logic and situation association. The mask mechanism combines the random replacement and retention strategy, which improves the generalization ability of the model while effectively avoiding overfitting, enhancing the robustness and expressiveness of the model to the semantic of the vocabulary. The introduction of the sentence pair continuity judgment task enables the model to have the ability to distinguish the semantic relationship between the previous and next sentences, further improving the processing level of semantic continuity and topic coherence in multi-round dialogue. The joint loss function design unifies the mask prediction and next sentence judgment tasks to optimize, significantly improving the training efficiency and multi-task learning ability, which helps the model to maintain consistent performance in various language understanding scenarios. At the same time, the iterative optimization strategy ensures the continuous convergence and performance improvement of the model structure and parameters in the training process, so that the pre-trained language model obtained finally has stronger semantic understanding depth, environmental adaptability and task migration ability, laying a solid foundation for subsequent high-quality dialogue generation and intelligent interaction.
[0068] Preferably, step S3 comprises the following steps:
[0069] Step S31: manual correction of intent labels, entity labels and environment labels of text data based on structured training data, and BIO labeling to obtain sequence labeling training data;
[0070] Step S32: BERT embedding coding of sequence labeling training data based on optimized language model parameters to obtain semantic representation enhanced sequence vectors;
[0071] Step S33: entity category labeling training of semantic representation enhanced sequence vectors, and sequence labeling modeling to obtain a named entity recognition model;
[0072] Step S34: reasoning and identifying entity recognition results of user input text based on the named entity recognition model;
[0073] Step S35: Based on the entity recognition result and the preset external knowledge base, entity semantic matching is performed, the matching result is subjected to concept disambiguation processing, the relationship triplets between the entity and the knowledge node are constructed, and a structured knowledge graph is obtained.
[0074] In the embodiment of the application, first, the structured training data generated in step S12 is called to manually review the text field piece by piece, and the accuracy of the intent label, entity label and environment label is checked sentence by sentence. The three-person cross-labeling method is used, which is divided into three types of operations: "confirmation", "correction" and "relabeling". The labeling object is at the Chinese character level. The intent label is limited to five categories (query, request, control, emotion and casual chat). The entity label is limited to five categories of time, place, product, organization and person. The environment label is limited to three types of information such as time period, place type and noise level. After confirmation, the entity is labeled using the BIO method, in which B represents the start of the entity, I represents the continuation of the entity, and O represents non-entity characters. The output result is a label sequence with the same length as the original text. The above BIO-labeled text is input into the BERT embedding coding module. Before coding, the text is segmented. The maximum length of each sample is limited to 128 tokens. WordPiece is used for sub-word decomposition and special markers [CLS] and [SEP] are inserted. The word vector embedding dimension is fixed at 768. Combined with position encoding and paragraph encoding, a three-layer embedding representation is formed. The input is fed into the coding end of the pre-trained language model. The context-dependent features are extracted through the self-attention mechanism. The output is a semantic representation enhanced sequence vector, which is a two-dimensional tensor structure with a shape of [batch_size, sequence_length, 768]. Based on the above vector and BIO label, a training sample is constructed one by one. The output dimension is set to the number of entity categories + 1. Linear mapping structure is used to map each token position output to a specific entity category. Viterbi decoding method is used to combine BIO label rules for path constraint and structure optimization. The accuracy and F1 value between the output label and the real label are compared. In the training process, each round contains 1000 batches, each batch contains 32 samples, the number of training rounds is set to 15, and after completion, the embedding layer and output layer weights are exported to form a named entity recognition model parameter set. The user input text from the user end is converted into a token sequence and an input vector through the same segmentation and embedding method. The corresponding BIO label is output by the named entity recognition module. The entity is extracted according to the B-I continuous combination rule. The original text, start position, end position and predicted label type of each entity are recorded. The output result is stored in the form of a JSON structure body. The fields include entity_text, start_index, end_index and entity_type. For each entity in the above entity recognition result, the standard entity index table in the preset external knowledge base is called to perform entity name string similarity calculation. The Jaccard similarity method is used to calculate the edit distance and overlap. The matching threshold is set to 0.75, the matched entity mapping is the target knowledge node, if the same entity corresponds to multiple nodes, the context entity co-occurrence probability is introduced as the basis for judgment, the concept disambiguation processing is executed; subsequently, the upper and lower nodes of the target node in the knowledge graph, the category nodes to which the target node belongs and the semantic connection relationship are extracted, a triple structure with the structure of "entity A-relation R-entity B" is constructed, each triple field contains head entity, relation, tail entity, and the environment tags (such as "day" and "hospital") extracted in the context are attached as additional annotations, finally, a structured knowledge graph is generated, the storage format is RDF / XML, each triple is encoded in a standard namespace, and is uniformly stored in a knowledge graph service engine.
[0075] The application improves the labeling accuracy of training data in the aspects of intent, entity and environment by introducing multi-label artificial correction and BIO labeling mechanism, and provides high-quality supervision information basis for downstream recognition models. Combined with optimized language model parameters for BERT embedding coding, the input sequence has better context awareness and language depth in semantic expression, thereby improving the accuracy and robustness of entity recognition. The entity recognition model is trained on the basis of semantic enhanced vectors, which can accurately identify key entities in user input, especially suitable for complex and variable dialogue scenarios. Through the combination of external knowledge base for entity semantic matching and concept disambiguation processing, the understanding ability of the model for homonymy, ambiguous concepts and other situations is significantly improved, effectively reducing misrecognition and ambiguous response. The finally constructed structured knowledge graph realizes the association modeling between entities and knowledge nodes, not only enhances the knowledge reasoning ability of the system, but also provides a solid knowledge support for subsequent intent explanation, strategy generation and personalized interaction, thereby significantly improving the intelligent level and semantic understanding depth of the entire system.
[0076] Preferably, step S35 comprises the following steps:
[0077] Step S351: based on the entity recognition result, the semantic similarity of each named entity and the standard entity node in the preset external knowledge base is matched to obtain an entity preliminary matching result;
[0078] Step S352: the entity preliminary matching result is subjected to concept level analysis, upper and lower semantic reasoning and synonymous concept merging to obtain a semantic unique target entity node;
[0079] Step S353: based on the semantic unique target entity node and the preset external knowledge base, the semantic connection relationship structure between entities is extracted to obtain a knowledge triple set;
[0080] Step S354: the knowledge triple set is subjected to structure verification, semantic completion and environment tag association mapping to construct a structured knowledge graph.
[0081] In the embodiment of the present application, first, the entity recognition result generated in step S34 is read, and standardization processing is performed on each named entity, including full-width half-width conversion, unified simplified and traditional Chinese coding format, and removal of punctuation marks. Then, a standard entity node word table under the corresponding entity category is extracted from the preset external knowledge base. Each entry in the word table includes a unique entity identifier, a standardized name, an alias list, and a category field. The Jaccard character overlap degree is used as a similarity measure, and all nodes in the word table are pairwise matched with the input entity. The similarity calculation formula is the intersection length divided by the union length. The matching threshold is set to 0.75. The entity with the highest similarity and exceeding the threshold is retained as the preliminary matching result. The recorded fields include the input entity, the standard node ID, the similarity score, and the matching method. The concept hierarchy analysis operation is performed on the standard node in the preliminary matching result. The superior and parallel concept nodes corresponding to the node are searched based on the upper and lower hierarchy structure in the knowledge base. The category label and synonymous entries in the ontology are also queried. If multiple entities correspond to the same superior concept and the context semantics are consistent, they are merged into a unique target entity node. The priority of the entities with synonymous labels is sorted. The weight is superimposed according to the co-occurrence word frequency in the context sentence, the syntax structure position, and the environment label. The node with the highest weight is finally selected as the only target node. All the unique entity nodes are assigned a uniform knowledge graph internal identifier code, and are attached with semantic classification labels and hierarchy path information. For all semantic unique target entity nodes, the relationship edges between the entity and its associated nodes are extracted from the external knowledge base. The relationship types are limited to five categories: “affiliated organization”, “located at”, “contains component”, “production unit”, and “applicable time”. The “entity-relation-entity” triple set is generated in a structured format. The field format is {head_entity_id, relation_type, tail_entity_id}. All relationships need to have clear bidirectional semantic definitions in the knowledge base. If there is a ring structure, the hierarchy depth is set to no more than 3 hops to prevent redundancy diffusion. Each triple is attached with a data source identifier and a creation timestamp. The triple set is subjected to structure consistency verification to check for duplicate relationships, empty entities, and illegal relationship types. Hash retrieval structure is used to filter redundant entries. The tail entity semantic completion is performed on the triple with a missing tail entity based on the existing entity hierarchy information. The completion principle is to select the node with the highest frequency of occurrence in the same semantic path as the head entity as the missing item. After completion, the context label is associated with each triple according to the environmental feature vector field obtained in step S1. The label types include time period, geographic location, and noise level. The structured knowledge graph with environmental context attributes is constructed. It is organized in RDF triple format and stored in a graph database. It supports entity query and path reasoning call based on SPARQL statements.
[0082] The application effectively improves the matching accuracy between named entities and standard entities in the knowledge base by introducing semantic similarity matching and concept hierarchy analysis mechanism, and significantly reduces the matching errors caused by language diversity and ambiguous expression. Through up-down reasoning and synonymous concept merging, the uniqueness of entity semantics is realized, enhancing the consistency and standardization of knowledge representation, providing stable foundation support for subsequent reasoning and retrieval of the system. The further extracted knowledge triples not only clarify the semantic connection between entities, but also improve the expression integrity and logical coherence of the knowledge structure. On this basis, through structure checking and semantic completion operations, the robustness and knowledge coverage of the graph are enhanced, and the association mapping mechanism of environmental labels gives the knowledge graph dynamic context awareness ability, making it better adapt to the specific scene and semantic context of the user. Overall, this method significantly improves the understanding ability of digital people to complex semantic structures and scene adaptability, providing a solid knowledge support for precise response and intelligent interaction.
[0083] Preferably, the step S4 includes:
[0084] performing semantic analysis on the environment-enhanced text vector based on the optimized language model parameters to obtain a semantic representation sequence;
[0085] performing multi-classification label processing on the user input text based on the semantic representation sequence to obtain a multi-classification intent prediction dataset;
[0086] performing label cleaning and balanced sampling on the multi-classification intent prediction dataset to obtain a class distribution optimized training sample;
[0087] performing vector encoding and feature processing on the class distribution optimized training sample using a BERT network structure to construct an intent recognition model and generate a user intent label.
[0088] In the embodiment of the application, first, the optimized language model parameters generated in step S2 are called, and the environment enhanced text vector generated in step S15 is input as an input, and the vector length is uniformly adjusted to 350 dimensions. After input, semantic analysis and processing are performed through an embedding layer and a 6-layer Transformer coding structure. Each layer is provided with 12 attention heads, and the attention head dimension is 64. The output vector of the [CLS] position is taken as the semantic representation sequence of the whole sentence, and the sequence dimension is 768. Subsequently, based on the semantic representation sequence, the intent classification label of each user input text in the structured training data is labeled. The intent label is limited to five categories: "query information", "request operation", "control instruction", "emotional feedback" and "idle chat response". The labels are processed in a one-hot encoding format to generate a multi-class intent prediction dataset. Each sample contains an input text, a semantic vector and a corresponding label. During sample arrangement, label distribution statistics are performed on the number of samples of each category. If the proportion of samples of a certain category is less than 10%, the same type of sentences are extracted from the original corpus for supplementation, or the sample number is expanded using the synonymous sentence expansion method to ensure that the proportion of each sample is not less than 15%. During the cleaning process, sentences with a semantic repetition degree higher than 0.95 are deleted, and finally a training sample set with balanced labels is formed. Then, the above training sample is encoded and processed using the BERT network structure. The input data is encoded by WordPiece to construct a token sequence, and the maximum length is limited to 128. [CLS] and [SEP] markers are added. After encoding, the embedding tensor structure is formed by combining the embedding layer, position encoding and paragraph encoding, and is transmitted to the coding module for context construction. Finally, the full connection layer is attached to the output vector at the [CLS] position, and the output layer node number is 5, corresponding to five types of intent. Through softmax conversion, the intent type corresponding to the maximum probability is selected as the prediction output. During the training process, the loss function adopts cross-entropy loss, the optimization process uses a fixed learning rate of 0.0005, the training rounds are 20, each round has 2000 batches, and each batch size is 32. After training, the embedding layer and the Transformer layer parameters are frozen, the classification output layer weight is retained, the intent recognition model parameter set is composed, and the unique user intent label is output.
[0089] The application effectively improves the perception ability of the user's real context and semantic details in the intent recognition process by fusing the optimized language model parameters and the environment enhanced text vector for semantic analysis, so that the model can still accurately distinguish the user's intent under complex or ambiguous input. The multi-classification label processing combined with the semantic representation sequence realizes efficient coverage of the diversity and fine granularity of intent, enhances the adaptability and generalization ability of the model to the subdivided tasks. By introducing the label cleaning and balanced sampling strategy, the class imbalance problem in the training data is significantly alleviated, the recognition accuracy of low-frequency intent categories is improved, and the model bias is reduced. The BERT structure is used to further strengthen the context consistency and context sensitivity of semantic expression by deep vector encoding and feature extraction of optimized samples, thereby constructing an intent recognition model with strong robustness and high classification precision. Overall, the method effectively improves the understanding ability and intent judgment accuracy of the digital human to the user's diversified expression, and provides key support for personalized and intelligent human-computer interaction.
[0090] Preferably, the step S4 of constructing the intent-action mapping table based on the historical language data and the action instruction data, and constructing the pre-training rule-learning hybrid reaction decision model based on the structured knowledge graph comprises:
[0091] Performing semantic correlation analysis on the historical language data and the action instruction data to obtain intent-action historical mapping samples;
[0092] Performing label classification and action aggregation on the intent-action historical mapping samples to obtain intent-action initial mapping pairs;
[0093] Performing statistical modeling and rule induction on the intent-action initial mapping pairs to obtain the intent-action mapping table;
[0094] Performing entity node screening, semantic relationship extraction and context label aggregation on the structured knowledge graph to obtain a semantic enhanced knowledge subgraph;
[0095] Analyzing the concept association relationship between the intent label and the target action based on the structured knowledge graph to obtain an intent-action-knowledge triple set;
[0096] Constructing a candidate action generation rule set based on the intent-action mapping table and the intent-action-knowledge triple set;
[0097] Performing reinforcement learning training on the intent-action historical mapping samples and the semantic enhanced knowledge subgraph to obtain a pre-training reinforcement learning strategy model;
[0098] Performing integrated modeling on the candidate action generation rule set and the pre-training reinforcement learning strategy model to construct the pre-training rule-learning hybrid reaction decision model.
[0099] In the embodiment of the present application, firstly, the historical language data and the action instruction data collected in step S1 are subjected to timestamp alignment processing, and are subjected to sliding matching according to a 1-second window, and a data pair in which a statement issued by a user and an instruction executed by a device are continuous in time and related in content is screened out, and a semantic matching threshold value is 0.the vector cosine similarity method of claim 7 to perform language-instruction content correlation analysis, extract user input text, action type, response component ID and execution result to form intent-action history mapping samples; on this basis, group statistics are performed according to intent labels, repeated actions within each group are merged for processing, the most frequent action is extracted as the initial action of the intent, an intent-action initial mapping pair is formed, and a frequency matrix is constructed to record the joint occurrence probability of each intent and action; then, statistical modeling operations are performed according to the frequency matrix, explicit rules are induced using the maximum frequency priority principle, the action with the highest occurrence percentage for each group of intents is taken as the main mapping action, and the rest are taken as candidate actions, the mapping rule field includes intent label, main action ID, candidate action set, action trigger condition, average response time and success rate, and a structured intent-action mapping table is constructed; then, the target entity node after normalization is selected in the structured knowledge graph, all one-hop adjacent nodes and edge relationships thereof are extracted, the edge type is limited to five types of “functional association”, “use place”, “corresponding service”, “controlled device” and “trigger condition”, and the context environment label field (such as geographic location, time period, noise level) attached to each entity is merged and injected into the node attribute to form a semantic enhanced knowledge subgraph with context information; on the basis of the knowledge subgraph, the main associated entity node is matched for each intent label, the path relationship between the intent label and the action node is counted and recorded, and the path length, edge type combination and context label consistency are recorded, and an intent-action-knowledge path set in the form of triple is constructed, each triple is in the format of {intent_label, knowledge_relation_path, action_id}, the path length is limited to no more than 3 hops, and the path structure is de-duplicated; the mapping table and the knowledge triple set are jointly arranged, a candidate action generation rule set is constructed according to the intent label as the primary key, and the rule structure includes the target action ID, the trigger path, the context matching condition and the execution priority; then, the intent-action history mapping samples and the context path in the semantic enhanced knowledge subgraph are taken as state inputs, the corresponding action is taken as a response option, the reward label composed of the execution result and the user feedback evaluation is collected, the feedback satisfaction threshold is set to 4 (full score 5), the action selection and feedback update operations are performed, the state-action-reward information of each round is accumulated, the training process is iterated for 100 rounds, each round contains 500 groups of state-action pairs, the action priority distribution is updated, and a reinforcement selection strategy table is constructed; finally, the candidate action generation rule set and the pre-trained reinforcement strategy table are field-aligned and integrated modeling, a hybrid structure is generated using rule priority matching and strategy score weighting, the structure includes an explicit rule mapping module, a context filtering module, an action score updating module and an output candidate action control module, a pre-trained rule-learning hybrid reaction decision model is composed, and is exported in JSON structure.
[0100] The application fully excavates the intention mode behind the user behavior by semantically associating and analyzing historical language data and action instructions, provides the system with real and rich intention-action mapping basis, thereby enhancing the individualization and accuracy of interactive response. The label classification and action aggregation improve the abstract level of the mapping pair, facilitating the model to realize generalization matching under diverse inputs. The mapping table constructed through rule induction provides a stable prior rule basis for response strategy, which helps to improve the explainability and response consistency of the model. The semantic enhanced knowledge subgraph in the structured knowledge graph establishes a deeper conceptual association between intention and action, improving the reasoning ability of the decision model when facing complex contexts. The candidate action generation rule set is constructed through the intention-action-knowledge triple, enabling the model to have strong initial response ability under the rule driving. Combined with the reinforcement learning strategy model training process, the system adaptability and response flexibility in dynamic environment are improved. Finally, through the integration modeling of rules and learning models, a reaction decision system with both knowledge driving and learning adaptability is constructed, which not only improves the response ability of digital people to the user's diversified needs, but also enhances the robustness and intelligent level in different task scenarios.
[0101] Preferably, step S5 comprises the following steps:
[0102] Step S51: analyzing the semantics of the label and the original input text of the user based on the user intention label alignment, to obtain a preliminary intention semantic segment;
[0103] Step S52: performing entity binding processing on the preliminary intention semantic segment based on the entity recognition result, to obtain a joint representation of explicit semantics and named entities;
[0104] Step S53: decoding and classifying the environment feature vector to obtain an environment context description;
[0105] Step S54: performing slot extraction and filling operations on the user input text based on the user intention label, the joint representation of entities and the environment context description, to obtain an intention filling sequence;
[0106] Step S55: performing semantic structured conversion on the intention filling sequence to construct an intermediate representation;
[0107] Step S56: performing type normalization and standardization processing on the intermediate representation based on the structured knowledge graph, to obtain a unified entity semantic representation;
[0108] Step S57: performing context discrimination analysis on the intermediate representation to determine the physical social scenario in which the user is currently located, and constructing a structured intention representation.
[0109] In the embodiment of the present application, firstly, the user intent label output in step S4 is called, the label is correspondingly parsed with the user original input text, semantic alignment analysis is performed through label name and keyword matching in the text, sliding window method (window length is set to 5 tokens) is used to extract semantic clauses related to the meaning of the intent label in the input text, the extraction rule is based on syntactic dependency relationship and keyword position offset, a preliminary intent semantic fragment is generated, and the output structure is {input_text_id, intent_label, semantic_fragment}; the named entity recognized in step S34 is mapped to the above semantic fragment according to its start and end position label in the original text, a mapping index table is constructed, each entity is inserted into the corresponding position in the fragment according to the B / I label recognition result, a double-field mapping is established in the format of {fragment_token, entity_type}, and the output is an explicit semantic and named entity joint representation sequence, the dimension of each sequence is consistent with the length of the original fragment; reverse parsing is performed on the environment feature vector generated in step S14, firstly, the geographic location label is restored to the administrative region name through GeoHash encoding, the illumination level, noise level and time period number are mapped to Chinese description labels, for example, noise level above 80 dB is marked as “high noise scene”, time period number 3 is mapped to “evening period”, the text environment context description is formed after merging all fields, and the output structure body field includes location_text, time_label, noise_status and motion_state; the user intent label, the entity representation in step S52 and the context description generated in step S53 are called jointly, slot extraction and filling operations are performed on the user original input text, the extraction operation adopts syntactic rules based on part of speech and dependency relationship, the target is to identify slot types and slot values, the slot types are limited to “target object”, “operation action”, “condition restriction” and “feedback requirement”, the filling process binds the identified slot values to the corresponding slots in the form of a structure body, and the output is an intent filling sequence, the structure form is {intent_label, slot_type, slot_value}; the intent filling sequence is subjected to structure conversion processing, and an intermediate representation containing “intent type-entity parameter-environment context” three parts is generated, the entity parameter field is derived from step S52, the environment context field is derived from step S53, the output structure body is in three nested forms, which respectively represent the target operation, the involved object and the environmental condition constraint of the current intent, and the storage format is standard JSON;The structured knowledge graph constructed in step S35 is called to standardize the entity parameters in the intermediate representation, map the object name in the slot value to the corresponding entity node ID in the knowledge graph, and extract its category, synonym, semantic label and other auxiliary attributes. If the slot value has a matching relationship with multiple nodes, the unique entity is selected based on the shortest semantic path, and the slot value in the intermediate representation is replaced by the binary tuple of the knowledge graph entity identifier and the standard name; the context field of the intermediate representation is analyzed, the scene classification label is generated according to the combination of light value, noise level and motion state, whether it belongs to the scene categories such as "indoor static", "outdoor high noise" and "commuting" is determined by using a rule table, and is added to the structure field. Finally, the structured intent representation is output, which includes fields such as intent_type, entity_id, entity_class, geo_location, temporal_context and scene_label, encoded in JSON format and saved to the task scheduling buffer area.
[0110] The present application effectively improves the semantic capture ability of the system for the user's real intention by introducing the semantic alignment mechanism of the intent label and the original input text, and enhances the accuracy of semantic analysis. By binding the recognized named entity and the explicit intent semantic fragment, the semantic and entity are deeply fused, which helps to improve the recognition and understanding ability of the entity parameters in the complex instruction. The decoding of the environmental characteristics and the mapping of the context supplement the user behavior background information, so that the system has stronger scene perception ability and context reasoning ability. The slot extraction and filling operation performed on this basis improves the structured degree of intent understanding and the integrity of parameter extraction. By constructing the intermediate representation of the triple structure, the system can express the intent content in a standardized way, enhancing the data consistency and logical clarity across modules. Combined with the normalization processing of the structured knowledge graph, the integration ability of the system for multi-source heterogeneous semantics is further improved, and the ambiguity and information redundancy are reduced. Finally, the structured intent representation is formed through the context discrimination analysis, so that the digital person can also maintain the consistency of accurate judgment of user intention and semantic expression in the changing physical and social scene, thereby significantly improving the adaptability, reliability and naturalness of intelligent interaction.
[0111] Preferably, step S6 comprises the following steps:
[0112] Step S61: extracting semantic deconstruction and instruction intention based on the structured intent representation to obtain a task execution semantic unit;
[0113] Step S62: performing conditional text generation reasoning on the task execution semantic unit based on a dialogue generation model to obtain a text reply content;
[0114] Step S63: Contextual multi-strategy action reasoning on the task execution semantic unit based on the reaction decision model to obtain a response action sequence;
[0115] Step S64: Format specification and structure reorganization on the text reply content and the response action sequence to generate a multi-modal response content data package;
[0116] Step S65: Adaptation strategy evaluation on the multi-modal response content data package based on the environmental feature vector, and judgment of the preferred output channel under the current scenario to obtain an output channel selection result;
[0117] Step S66: Channel distribution and calling processing on the multi-modal response content data package to generate corresponding output content according to the output channel selection result;
[0118] Step S67: Digital human behavior rendering and interface visual design processing on the corresponding output content to obtain a digital human interaction feedback result.
[0119] In the embodiment of the application, firstly, based on a pre-designed structured intention representation template, the input natural language is hierarchically parsed, the input text is decomposed into semantic deconstruction elements and instruction intention information through grammar rules and semantic role labeling technology, the sentence components and their relationships are extracted by using a dependency syntax analyzer and a semantic graph construction tool to form a task execution semantic unit containing action items, target objects, constraint conditions and other elements, and each element in the unit is limited to no more than 128 characters, ensuring that the semantic information is complete and accurate; using a text generation process based on a conditional probability inference mechanism, the task execution semantic unit obtained in step S61 is input as a condition, a text template matching and a word replacement method based on context correlation are used to generate a text reply content with coherent structure and clear semantics, the specific process includes vocabulary table screening, syntax structure adjustment and semantic consistency verification, the length of the text reply content is limited to within 512 characters, ensuring that the expression of the reply is complete and clear; the task execution semantic unit is combined with the dialogue history context, a multi-strategy action inference method is used, that is, the decision tree branch rule and the context similarity calculation are combined to generate a response action sequence, the action sequence contains specific action identifiers, parameter settings and execution timing, and the number of actions is limited to within 5 to ensure effective execution of the action; the text reply content of step S62 and the response action sequence generated in step S63 are formatted, the text content and the action sequence are divided into fields using a JSON data structure, the fields include "text content", "action type", "action parameter", "action sequence", and a check code is added to the data packet and character encoding conversion is performed, ensuring that the data structure is standardized and unified, and the data packet size is controlled within 10KB for subsequent processing; according to the current environment feature vector composed of temperature, light intensity, user device type and network bandwidth and other multi-dimensional values collected by the internal sensors of the system, the multi-modal response content data packet is matched and scored through a pre-defined adaptive strategy mapping table, the adaptation score is calculated, the score interval is limited to 0 to 100, the output channel that best matches the scene is selected according to the score ranking, the output channel includes one of screen display, voice broadcast and gesture prompt, ensuring that the output channel best matches the environmental conditions; the multi-modal response content data packet is distributed according to the output channel selection result of step S65, the corresponding channel interface is called to realize content distribution, the specific calling process is as follows: if the screen display channel is selected, the text display module of the rendering engine is called for typesetting and presentation; if the voice broadcast channel is selected, the speech synthesis module is called to convert the text content into an audio signal output; if the gesture prompt channel is selected, the action control unit is called to execute the corresponding action sequence, the action execution time is not more than 3 seconds, ensuring that the response is timely and accurate.For the output content generated in step S66, in combination with the digital human interaction framework, the action capture technology is used to perform behavior rendering processing on the response action sequence, and at the same time, the interface element layout and interaction effect design are performed through the interface design tool, to ensure that the digital human behavior is smooth and natural, and the interface interaction is clear and visible, and finally the digital human interaction feedback result is output.
[0120] The present application realizes accurate semantic deconstruction and intention recognition of user instructions by extracting task execution semantic units from structured intention representation, ensuring the target definiteness and execution consistency of response behavior. Text generation and multi-strategy action reasoning are respectively completed by means of dialogue generation model and reaction decision model, so that the system has both natural and smooth language expression and intelligent flexibility of behavior decision, improving the personalization and context matching ability of response content. The structure reorganization and format specification operation of multi-modal response content guarantee the consistency and operability of data, providing high-quality input for subsequent modules. The channel adaptation strategy evaluation mechanism driven by environmental feature vector enhances the system's judgment ability of different physical scenes and adaptive output strategy, making the response output more context-adaptive. The selection and calling of output channels ensure that the content can be conveyed through the optimal channel (such as voice, text, action, etc.), improving the human-computer interaction efficiency and user experience. Finally, through digital human behavior rendering and interface visualization design, the system can express feedback content in a natural and lively way, realizing the intuitiveness and emotionalization of information transmission, and significantly enhancing the immersion, expressiveness and realism of digital human interaction.
[0121] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting, the scope of the present application is not limited by the above description, and therefore all changes falling within the meaning and scope of the equivalent elements of the application file are intended to be included in the present application.
[0122] The above description is only a specific implementation of the present application, enabling those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but will conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. An interactive method for language understanding and corresponding responses in digital humans driven by an artificial intelligence algorithm, characterized in that, Includes the following steps: Step S1: Collect the user's historical language data, text data, action command data, and physical environment information to obtain structured training data; perform vectorization processing on the text data to obtain text semantic vectors; extract environmental feature vectors from the physical environment information; concatenate the text semantic vectors with the environmental feature vectors to obtain environment-enhanced text vectors. Step S2: Construct a pre-trained language model based on the environment-enhanced text vectors, and perform secondary pre-training to generate optimized language model parameters; construct and train a dialogue generation model based on the optimized language model parameters and historical language data; Step S3: Perform BIO annotation on the text data based on the structured training data, and train the named entity recognition model through a pre-trained language model to generate entity recognition results; construct a structured knowledge graph based on the entity recognition results and a preset external knowledge base; Step S4: Based on optimized language model parameters and enhanced text vectors, identify multi-class intents of user input text and construct an intent recognition model to obtain user intent labels; construct an intent-action mapping table based on historical language data and action command data, and construct a pre-trained rule-learning hybrid reaction decision model based on structured knowledge graph; Step S5: Interpret the user's input intent based on the user intent label, entity recognition results, and environmental feature vectors to obtain a structured intent representation; Step S6: Design multimodal response actions based on structured intent representation, dialogue generation model and response decision model, and visualize them to obtain digital human interaction feedback results.
2. The interactive method for digital human language understanding and corresponding response driven by artificial intelligence algorithm according to claim 1, characterized in that, Step S1 includes the following steps: Step S11: Collect the user's historical language data, text data, action command data, and physical environment information to obtain the raw multimodal historical interaction data; Step S12: Perform data cleaning, corpus classification, and multi-label manual annotation on the original multimodal historical interaction data to obtain structured training data; Step S13: Perform word segmentation, tokenization, and context encoding on the text data in the structured training data to obtain the text semantic vector; Step S14: Perform feature engineering based on physical environment information to obtain environmental feature vectors; Step S15: Concatenate the text semantic vector and the environment feature vector along the vector dimension to obtain the environment-enhanced text vector.
3. The interactive method for digital human language understanding and corresponding response driven by artificial intelligence algorithm according to claim 1, characterized in that, Step S2 includes the following steps: Step S21: Construct the input representation based on the environment-enhanced text vector to obtain the joint representation sequence; Step S22: Based on the structured training data, perform masked language model and next sentence prediction training on the joint representation sequence to obtain a pre-trained language model; Step S23: Perform secondary pre-training optimization based on the pre-trained language model to obtain optimized language model parameters; Step S24: Based on the optimized language model parameters and historical language data, perform vector encoding and construct training samples to obtain context-aware dialogue training corpus; Step S25: Perform sequence modeling and Transformer decoder optimization on the context-aware dialogue training corpus, and train the dialogue generation model to obtain the dialogue generation model.
4. The interactive method for digital human language understanding and corresponding response driven by artificial intelligence algorithm according to claim 3, characterized in that, Step S22 includes the following steps: Step S221: Segment the joint representation sequence to construct input sentence pairs suitable for language modeling tasks, and obtain the training task sentence pair set; Step S222: Based on the structured training data, label and group the training task sentence pairs, construct positive and negative sample combinations, and obtain the supervised training sample set; Step S223: Perform a random masking operation on the supervised training sample set, and perform MASK replacement, random replacement and retention operations on the token to obtain the masked input sequence; Step S224: Perform continuity and environmental consistency judgment on the supervised training sample set, construct sentence pair labels, and obtain NSP input samples and labels; Step S225: Perform semantic encoding on the masked input sequence to obtain the context fusion representation sequence; Step S226: Model the prediction output of the context fusion representation sequence, and use the prediction output modeling result to perform the prediction of the masked token and the judgment of sentence pair continuity, so as to obtain the mask prediction result and the next sentence prediction result; Step S227: Calculate the joint loss function for the mask prediction result and the next sentence prediction result to obtain the multi-task training objective of the language model; Step S228: Perform backpropagation and parameter update on the masked language model based on the joint loss function, iteratively optimize the model structure and weight parameters, and obtain the pre-trained language model.
5. The interactive method for digital human language understanding and corresponding response driven by artificial intelligence algorithm according to claim 1, characterized in that, Step S3 includes the following steps: Step S31: Based on the structured training data, manually proofread the text data for intent labels, entity labels, and environment labels, and perform BIO annotation to obtain sequence labeling training data; Step S32: Perform BERT embedding encoding on the sequence labeling training data based on the optimized language model parameters to obtain semantically enhanced sequence vectors; Step S33: Train the semantic representation-enhanced sequence vectors with entity category annotations and perform sequence annotation modeling to obtain the named entity recognition model; Step S34: Based on the named entity recognition model, infer and recognize the user input text on the user terminal to obtain the entity recognition result; Step S35: Perform entity semantic matching based on entity recognition results and a preset external knowledge base, and perform concept disambiguation processing on the matching results to construct triplet relationships between entities and knowledge nodes, thereby obtaining a structured knowledge graph.
6. The interactive method for digital human language understanding and corresponding response driven by artificial intelligence algorithm according to claim 5, characterized in that, Step S35 includes the following steps: Step S351: Based on the entity recognition results, perform semantic similarity matching between each named entity and the standard entity nodes in the preset external knowledge base to obtain preliminary entity matching results; Step S352: Perform conceptual hierarchy analysis, hierarchical semantic reasoning, and synonym merging on the preliminary entity matching results to obtain semantically unique target entity nodes; Step S353: Extract the semantic connection structure between entities based on the semantically unique target entity nodes and the preset external knowledge base to obtain a set of knowledge triples; Step S354: Perform structural verification, semantic completion, and environment label association mapping on the knowledge triple set to construct a structured knowledge graph.
7. The interactive method for digital human language understanding and corresponding response driven by artificial intelligence algorithm according to claim 1, characterized in that, Step S4, which involves optimizing language model parameters and enhancing text vectors based on the environment to identify multi-classification intents of user input text and constructing an intent recognition model, includes: Semantic analysis of environment-enhanced text vectors is performed based on optimized language model parameters to obtain semantic representation sequences; Multi-class labeling is performed on user input text based on semantic representation sequences to obtain a multi-class intent prediction dataset; Label cleaning and balanced sampling are performed on the multi-class intent prediction dataset to obtain training samples with optimized category distribution; By using the BERT network structure to optimize the category distribution of training samples, vector encoding and feature processing are performed to construct an intent recognition model and generate user intent labels.
8. The interactive method for digital human language understanding and corresponding response driven by artificial intelligence algorithm according to claim 1, characterized in that, Step S4 involves constructing an intent-action mapping table based on historical language data and action command data, and building a pre-trained rule-learning hybrid response decision model based on a structured knowledge graph, including: Semantic association analysis was performed on historical language data and action command data to obtain intent-action history mapping samples; The intent-action history mapping samples are labeled and aggregated to obtain initial intent-action mapping pairs. Statistical modeling and rule induction are performed on the initial intent-action mapping pairs to obtain the intent-action mapping table; Entity node filtering, semantic relation extraction, and context label aggregation are performed on the structured knowledge graph to obtain a semantically enhanced knowledge subgraph; Based on the analysis of the conceptual relationship between intent tags and target actions using structured knowledge graphs, a set of intent-action-knowledge triples is obtained; A set of candidate action generation rules is constructed based on the intent-action mapping table and the intent-action-knowledge triple set. Reinforcement learning is performed on the intent-action history mapping samples and the semantically enhanced knowledge subgraph to obtain a pre-trained reinforcement learning policy model. An integrated modeling approach is developed, combining the candidate action generation rule set with the pre-trained reinforcement learning strategy model, to construct a hybrid reaction decision model based on pre-trained rules and learning.
9. The interactive method for digital human language understanding and corresponding response driven by artificial intelligence algorithm according to claim 1, characterized in that, Step S5 includes the following steps: Step S51: Based on the alignment analysis of user intent labels, analyze the semantics of the labels and the user's original input text to obtain preliminary intent semantic fragments; Step S52: Based on the entity recognition results, perform entity binding processing on the preliminary intent semantic fragment to obtain a joint representation of explicit semantics and named entities; Step S53: Decode and classify the environmental feature vectors to obtain an environmental context description; Step S54: Perform slot extraction and filling operations on the user input text based on user intent labels, entity joint representations, and environmental context descriptions to obtain the intent filling sequence; Step S55: Perform semantic structuring transformation on the intended filling sequence to construct an intermediate representation; Step S56: Based on the structured knowledge graph, perform type normalization and standardization on the intermediate representation to obtain a unified entity semantic representation; Step S57: Perform contextual discriminant analysis on the intermediate representation to determine the user's current physical and social context, and construct a structured intent representation.
10. The interactive method for digital human language understanding and corresponding response driven by artificial intelligence algorithm according to claim 1, characterized in that, Step S6 includes the following steps: Step S61: Extract semantic deconstruction and instruction intent based on structured intent representation to obtain task execution semantic units; Step S62: Based on the dialogue generation model, perform conditional text generation reasoning on the semantic units of task execution to obtain the text response content; Step S63: Based on the reaction decision model, perform multi-strategy action reasoning on the semantic units of task execution to obtain the response action sequence; Step S64: Standardize the format and restructure the text response content and response action sequence to generate a multimodal response content data packet; Step S65: Evaluate the adaptation strategy of the multimodal response content data packet based on the environmental feature vector, determine the preferred output channel in the current scenario, and obtain the output channel selection result; Step S66: Perform channel distribution and invocation processing on the multimodal response content data packet, and generate corresponding output content based on the output channel selection result; Step S67: Perform digital human behavior rendering and interface visualization design on the corresponding output content to obtain digital human interaction feedback results.
Citation Information
Cited By
Knowledge base retrieval method fused with natural language large model
CN121168677A
Intelligent dynamic tagging and implementation method for multi-source heterogeneous customer data
CN121614642A
Generative large language model-based intention recognition method and system
CN121833953A
Document structured analysis method and system based on basic discourse units
CN121920374A
Multi-round interaction emotion speech synthesis method and system based on context self-adaption
CN122024704A