Environment interaction large model construction method and device based on protocol learning and storage medium
By importing protocol learning templates and protocol databases into the large environmental interaction model, the problem of insufficient scenario expansion capability is solved, and high scalability and interactive compatibility are improved.
Patent Information
- Application Number
- CN202510651304.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-09-26
AI Technical Summary
Existing technologies lack scenario expansion capabilities in environmental interaction scenarios, resulting in low interaction compatibility.
By importing the protocol learning template into the initial model, protocol learning is performed to generate an interactive model. The model is modified according to the protocolized decision rules and content output rules, and data query and association are performed in conjunction with the protocol database to generate a target model.
The settings visible to model training converge with the training process, which promotes the generalization of protocols invisible to model training, achieves high scalability, and improves interactive compatibility.
Smart Images

Figure CN120706539A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method, device, and storage medium for constructing a large-scale environment interaction model based on protocol learning. Background Art
[0002] With the development of large-scale model technology, conversational robots based on large models have been widely applied in key areas such as sales, customer service, management, and smart cities. They have greatly promoted the intelligentization of today's society and provided convenience for residents. However, these technologies lack the ability to expand their scope in environmental interaction scenarios, resulting in low interaction compatibility.
[0003] The above content is only used to assist in understanding the technical solution of this application and does not constitute an admission that the above content is prior art. Summary of the Invention
[0004] The main purpose of this application is to provide a method, device and storage medium for constructing a large environmental interaction model based on protocol learning, aiming to solve the technical problem that related technologies lack scenario expansion capabilities in environmental interaction scenarios, resulting in low interaction compatibility.
[0005] To achieve the above objectives, this application proposes a large-scale model method for environmental interaction based on protocol learning, which includes:
[0006] Importing a protocol learning template into an initial model, performing protocol learning on the initial model, and generating an interaction model;
[0007] Modifying the interaction model to obtain a rule interaction model according to the protocolized decision rule and the content output rule;
[0008] The rule interaction model is associated with a protocol database through data query to generate a target model, wherein the protocol database includes an environment interaction protocol.
[0009] In one embodiment, based on the initial model, a protocol learning template is imported to perform protocol learning training for the dialogue instruction task, and a protocol learning training weight is generated;
[0010] Importing the protocol learning training weights into the initial model, performing a protocol learning test, and generating a test result;
[0011] Verifying the protocol learning and training weights according to the test results to obtain verified protocol learning and training weights;
[0012] The verified protocol learning training weights are imported into the initial model, the pre-training weights of the initial model are updated, and the interaction model is generated.
[0013] In one embodiment, based on the interaction model and reinforcement learning framework, the logical rules of the business scenario are simulated to generate state observations;
[0014] The interaction model generates a feedback decision according to the state observation value, and associates the state observation value with the feedback decision to generate the protocolized decision rule;
[0015] Based on the interaction model, combined with the protocolized decision rules and preset conditions, a content sample that meets the preset conditions is generated, and an alignment verification set is constructed through a stratified random sampling strategy;
[0016] Based on the alignment verification set, decision tree rule extraction is performed to generate the content output rule.
[0017] In one embodiment, the protocolized decision rule is imported into the interaction model, and the decision parameters of the protocolized decision are adjusted to generate the protocolized interaction model;
[0018] The content output rules are imported into the protocolized interaction model, and the screening parameters of the content output are adjusted to generate the rule interaction model.
[0019] In one embodiment, the data of the environment interaction protocol is stored in a database to construct an expandable protocol database;
[0020] Based on the rule interaction model, an associated query table is established with the protocol database to generate the target model.
[0021] In one embodiment, a structured protocol interface is established between the target model and the protocol database based on the association query table;
[0022] Scenario deployment of the target model is performed according to the target model, the protocol database, the association query table, and the structured protocol interface.
[0023] In one embodiment, when environmental interaction is involved, control signals are constructed through protocolized decision rules;
[0024] Parsing the control signal through hard-coded program code, making an interactive request to the associated resource, and obtaining a response feedback;
[0025] The response feedback is input into the target model to generate an interactive response to obtain an interactive result;
[0026] Update the content of the immediate response based on the interactive response.
[0027] In one embodiment, if the environmental interaction has an operation type outside the protocolized decision rule, the protocol database is queried according to the associated query table to obtain the interaction protocol corresponding to the operation type;
[0028] Based on the interaction protocol, the control signal constructed by the protocolized decision rule is updated.
[0029] In addition, to achieve the above-mentioned purpose, the present application also proposes an environmental interaction large model device, which includes: a memory, a processor, and a computer program stored on the memory and runnable on the processor, and the computer program is configured to implement the steps of the method for constructing an environmental interaction large model based on protocol learning as described above.
[0030] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by the processor, the steps of the method for constructing a large environmental interaction model based on protocol learning as described above are implemented.
[0031] The present application provides a method for constructing a large model of environmental interaction based on protocol learning, including importing a protocol learning template into an initial model, performing protocol learning on the initial model to generate an interaction model; modifying the interaction model according to the protocolized decision rules and content output rules to obtain a rule interaction model; and associating the rule interaction model with a protocol database for data query to generate a target model, wherein the protocol database includes an environmental interaction protocol. Based on a complete protocol learning setting, the present application can achieve efficient vertical scenario solution migration by simply performing protocol configuration and introducing an additional low-cost database, without the need for expensive manual secondary development and continued training or fine-tuning with high computing power costs.
[0032] To sum up, this application further expands the relevant protocol settings through the adaptive design of protocol learning templates and combines them with the protocol database, so that the settings visible to model training converge with the training process. At the same time, it promotes the generalization of protocols invisible to model training, overcomes the technical problem of low interactive compatibility in related technologies, and achieves the effect of high scalability. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0034] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0035] Figure 1This is a flow chart of the first embodiment of the method for constructing a large environmental interaction model based on protocol learning in this application;
[0036] Figure 2 This is a flow chart of the second embodiment of the method for constructing a large environmental interaction model based on protocol learning in this application;
[0037] Figure 3 Study the flowchart for this application agreement;
[0038] Figure 4 This is a flowchart of the third embodiment of the method for constructing a large environmental interaction model based on protocol learning in this application;
[0039] Figure 5 This is the interactive system diagram of the learning model for this application protocol;
[0040] Figure 6 This is an example diagram of the protocol database for this application;
[0041] Figure 7 This is a flowchart of the complete training program;
[0042] Figure 8 This is the industrial control software sub-item diagram in the application agreement database;
[0043] Figure 9 This is a structural diagram of the large model equipment for environmental interaction in this application.
[0044] The purpose, features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION
[0045] It should be understood that the specific embodiments described herein are merely used to explain the technical solutions of the present application and are not intended to limit the present application.
[0046] Related technologies lack scenario expansion capabilities in environmental interaction scenarios, resulting in technical problems such as low interactive compatibility.
[0047] The present application provides a solution: first, a protocol learning template is imported into an initial model, protocol learning is performed on the initial model to generate an interaction model, then, according to the protocolized decision rules and content output rules, the interaction model is modified to obtain a rule interaction model, and finally, the rule interaction model is associated with a protocol database for data query to generate a target model, wherein the protocol database includes an environment interaction protocol.
[0048] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, mobile phone, etc., or an electronic device or large-scale environmental interaction model device capable of implementing the above functions. The following uses the large-scale environmental interaction model device as an example to illustrate this embodiment and the following embodiments.
[0049] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.
[0050] The present invention provides a method for constructing a large-scale environment interaction model based on protocol learning. Figure 1 , Figure 1 This is a flow chart of the first embodiment of the method for constructing a large environmental interaction model based on protocol learning in this application.
[0051] In this embodiment, the method for constructing a large environmental interaction model based on protocol learning includes steps S10 to S30:
[0052] Step S10: importing the protocol learning template into the initial model, performing protocol learning on the initial model, and generating an interaction model.
[0053] In this embodiment, the initial model is a model that has completed basic training on large-scale general-purpose data and outputs a multimodal pedestal containing pre-trained weight parameters. The protocol learning template is a standardized set of constraints extracted from structured protocol data, used to define permitted / prohibited operations and association rules. The interaction model is an intelligent interaction model that integrates protocol reasoning capabilities, supporting real-time interaction and generating protocol-compliant outputs.
[0054] As an optional implementation method, in the scenario of building a large environmental interaction model, the protocol learning template is imported into the initial model, and the initial model is trained with protocol learning for dialogue instruction tasks using a test data set. The protocol learning training weights are obtained through the protocol learning training, and then the protocol learning training weights are imported into the initial model for testing. The test results are verified, and the initial model is updated with the verified protocol learning weights to obtain an interaction model.
[0055] As an implementation method for acquiring a test data set, in the test data set acquisition scenario, it is obtained by modifying the open source data set according to preset training requirements, or by simulating and synthesizing data through an existing large model system, or by direct manual input training.
[0056] Step S20: modifying the interaction model to obtain a rule interaction model according to the protocolized decision rule and the content output rule.
[0057] In this embodiment, protocolized decision rules refer to the judgment criteria that a machine can execute, as trained through the protocol. Simulated output refers to using a trained model to generate virtual operation results to test effectiveness. Content output rules refer to the specific rules that specify the model's responses or operations must comply with.
[0058] As an optional implementation method for generating protocolized rules, in the scenario of building a large environmental interaction model, the agreement terms in the agreement learning template are loaded through the pre-trained initial model, and a virtual decision-making scenario is built in the reinforcement learning environment. The model attempts to generate operation plans based on the current data and agreement terms. The system automatically checks whether these plans meet the requirements of the agreement. Compliant plans will be recorded, and non-compliant plans will be deleted. The operation results and corresponding records generated by each simulation are associated to generate protocolized decision rules. Then, by using these simulated operation results and corresponding records, mixed with correct operation cases in real business scenarios, the interaction model is trained to understand the specific boundaries of content output. Combined with adjustable preset rules, content output rules are finally generated that comply with the agreement terms and can solve practical problems.
[0059] As an implementation method for judging the quality of generated results, in the scenario of building a large environmental interaction model, a virtual decision-making scenario is built in a reinforcement learning environment, allowing the model to try to generate operation plans based on current data and agreement terms. The system automatically checks whether these plans meet the requirements of the agreement. Compliant plans will receive bonus points, while non-compliant plans will be deducted points and the error type will be recorded. These simulated operation results and corresponding scoring data are then used.
[0060] As an optional implementation method, in the scenario of building a large environmental interaction model, the rule conversion logic is configured according to the interaction model, and the protocolized decision rules and content output rules are converted into code logic that can be directly executed by the machine through the rule conversion logic. The code logic is adapted to the operation logic of the interaction model, and the adaptation is performed by judging whether the code logic and the operation logic conform to the preset execution logic. If they conform to the preset execution logic, the logic is integrated. If they do not conform to the preset execution logic, automatic fine-tuning or manual intervention is performed. After the logic fusion and adjustment, a rule interaction model is generated.
[0061] Step S30 , performing data query association on the rule interaction model and the protocol database to generate a target model, wherein the protocol database includes an environment interaction protocol.
[0062] In this embodiment, the protocol database refers to a database that stores business agreement terms and additional association rules, including environment interaction protocols (which refer to protocols involved in interactions with other devices) and interaction protocols beyond protocol learning templates. Data query association refers to the process of matching relevant protocol rules from the database based on the current request. The target model refers to a practical, operational decision model that integrates protocol rules and business logic.
[0063] As an optional implementation method, in the scenario of building a large environmental interaction model, the protocol database is connected through the rule interaction model, the database query protocol is used to obtain the associated constraint rules, a query association is established between the rule interaction model and the protocol database, the association adaptability of the rule interaction model and the protocol database is tested, fine-tuning or manual adjustment is performed according to the test adaptation structure, and finally the target model is generated and the deployment of the target model is executed.
[0064] As an implementation method of association adaptability judgment, it is determined whether there are any agreement clauses in the agreement database that conflict with the content output rules. If so, the rule interaction model is automatically fine-tuned to give priority to meeting the requirements of the content output rules. If the fine-tuning fails, manual intervention is notified to adjust.
[0065] Exemplarily, pre-training and instruction fine-tuning follow the general large-model training process, with an additional protocol learning step following instruction fine-tuning for model training. During the deployment phase, a protocol database is introduced to store additional protocol data not visible during model training, such as public or private APIs, locally deployed privatized models, and environmental interaction interfaces for non-AI model types (encapsulation protocols for heterogeneous terminals). The protocol learning template contains the following fields: <|system|> is used to identify the character settings related to the large model; <|protocol_define|> is used to identify the field definition and interpretation of the global protocol; <|user|> is used to identify the user input content; <|assistant|> is used to identify the generated content of the model; <|response|> is embedded to identify the response content related to the interaction between the model and humans; <|control|> is used to identify the control signal related to the interaction between the model and the external environment; <|feedback|> is used to identify the feedback signal related to the interaction with the external environment received by the model; <|extra_protocol|> is used to identify the protocol information queried from the protocol database and is used to standardize the control signal content in the <|control|> field; <|conversation|> is used to identify the start and end of a complete multi-round conversation; <|id|> is used to identify the temporal relationship of micro-operations. Micro-operation types include: <|user|><|response|><|control|><|feedback|>
[0066] Due to the adaptive design of the protocol learning template and the combination with the protocol database, the relevant protocol settings are further expanded, so that the settings visible to the model training converge with the training process. At the same time, the protocols invisible to the model training are generalized, achieving a high scalability effect.
[0067] Based on any of the above embodiments, in the second embodiment of the present application, refer to Figure 2 , Figure 2 This is a flow chart of the second embodiment of the method for constructing a large-scale environment interaction model based on protocol learning in this application. Step S10 includes steps A11 to A14:
[0068] Step A11: Based on the initial model, a protocol learning template is imported to perform protocol learning training for the dialogue instruction task, and a protocol learning training weight is generated.
[0069] In this example, a dialogue instruction task is one that requires the AI to answer questions according to rules in a specific scenario. Protocol learning training weights refer to the parameter files generated after training, which enable the model to complete basic interactive tasks.
[0070] As an optional implementation method, in the scenario of building a large environmental interaction model, based on the initial model, the protocol learning template is imported into the initial model, and the dialogue instruction task is trained. By inputting different dialogue instructions into the initial model, the relevant dialogue instructions are identified and responded to by combining the protocol learning template, and the protocol learning training weights are generated through the response.
[0071] Step A12: import the protocol learning training weights into the initial model, perform protocol learning test, and generate test results.
[0072] In this embodiment, protocol learning testing refers to the process of verifying whether the model complies with protocol rules through simulated dialogue. The test result is an evaluation report containing compliance rate indicators.
[0073] As an optional implementation method, in the scenario of building a large environmental interaction model, the protocol learning training weights obtained through training are imported into the initial model, and the protocol learning test is performed through dialogue instructions to test the model's compliance with the protocol, and the test results are generated through the compliance rate.
[0074] Step A13: Verify the protocol learning and training weights according to the test results to obtain verified protocol learning and training weights.
[0075] In this embodiment, verification refers to adjusting weight parameters by analyzing error cases to correct deviations in the model's understanding of the rules. The verified protocol learning training weights refer to the optimized parameter file, which ensures that the model adheres to the protocol rules more accurately.
[0076] As an optional implementation method, in the scenario of building a large environmental interaction model, the model fine-tunes itself according to the compliance rate and protocol conflict items displayed in the test results, and verifies and adjusts the protocol learning training weights to correct the model's understanding deviation of the rules and obtain the verified protocol learning training weights.
[0077] Step A14: import the verified protocol learning training weights into the initial model, update the pre-training weights of the initial model, and generate the interaction model.
[0078] In this embodiment, updating the pre-trained weights of the initial model refers to adjusting and updating the pre-trained weights through the verified protocol learning weights.
[0079] As an optional implementation method, in the scenario of building a large environmental interaction model, the optimized protocol learning weights are merged into the initial model, and the original weights are gradually replaced through the parameter fusion algorithm, and the original parameters are overwritten with the verified high-sensitivity parameters. After the update is completed, the updated model is adaptively fine-tuned using the online learning framework to obtain the interaction model.
[0080] For example, referring to Figure 3 , Figure 3 For the protocol learning flowchart of this application, by constructing the protocol learning data templates of the system segment, protocol segment and dialogue segment, the original data corpus used to train the model can be obtained. The template structure and definition refer to the previous description. The corresponding data content will be cleaned, synthesized and constructed for vertical domain tasks, such as by modifying the open source data set or synthesizing the data based on the existing large model system. After obtaining the original data corpus, it is converted into token data through the tokenizer of the large model, and then training can be carried out. The optional training methods include but are not limited to common model training methods such as pretraining, continue pretraining, instruction finetune, and post-training alignment. After the training is completed, the results can be tested according to the business scenario or experimental test set.
[0081] By defining the data structure through the protocol template and explicitly decoupling various fields, it is possible to introduce control and feedback signals of environmental interaction into the structured template, and introduce unstructured natural responses. Through supervised training to adapt to the protocol context, the adaptability of the large model is improved.
[0082] Based on any of the above embodiments, in the third embodiment of the present application, refer to Figure 4 , Figure 4This is a flow chart of the third embodiment of the method for constructing a large-scale environment interaction model based on protocol learning in this application. Before step S20, steps B11 to B14 are also included:
[0083] Step B11: Based on the interaction model and reinforcement learning framework, the logical rules of the business scenario are simulated to generate state observation values.
[0084] In this example, the reinforcement learning framework refers to an AI training system that uses a simulated environment to enable the model to learn through trial and error, optimizing its decision-making through a reward / penalty mechanism. The logical rules of a business scenario refer to the operational procedures and constraints within a specific business. State observations refer to the real-time environmental parameters that the model considers when making decisions.
[0085] As an optional implementation method, in the scenario of building a large environmental interaction model, a virtual scene is built based on the interaction model and combined with the reinforcement learning framework. The interaction model is placed in a specific virtual scene for reinforcement training, the logical rules of the business scenario are learned, and based on the logical rules, the state observation values of the scene are generated through the interaction model.
[0086] Step B12: The interaction model generates a feedback decision based on the state observation value, and associates the state observation value with the feedback decision to generate the protocolized decision rule.
[0087] In this embodiment, feedback decision refers to the response instructions generated based on the current state observation value. Protocolized decision rules refer to solidifying effective decisions into reusable business rules.
[0088] As an optional implementation method, in the scenario of building a large environmental interaction model, the interaction model analyzes the state observation values, identifies characters, scenes and tasks, generates feedback decisions based on the characters, scenes and tasks, and makes adaptability judgments on the feedback decisions. The feedback decisions are adjusted based on the judgment results, and finally the adjusted feedback decisions are associated with the corresponding state observation values to generate protocolized decision rules.
[0089] Step B13: Based on the interaction model, combined with the protocolized decision rules and preset conditions, generate content samples that meet the preset conditions, and construct an alignment verification set through a stratified random sampling strategy.
[0090] In this embodiment, the output content of the interaction model simulation refers to the dialogue responses or action suggestions generated by the trained interaction model. Pre-defined conditions refer to pre-defined compliance standards. Content samples refer to specific instances of model output. A stratified random sampling strategy refers to a method of sampling proportionally after grouping by business type or risk level. An alignment validation set refers to a set of test data used to verify that the model output matches the rule requirements.
[0091] As an optional implementation method, in the scenario of building a large environmental interaction model, simulation requests are processed through the interaction model, and the output content corresponding to the simulation requests is output according to the protocolized decision rules. The interaction model automatically filters out content samples that meet the preset conditions based on preset conditions, classifies them according to business scenarios through stratification strategies, and divides each type of business scenario into sub-layers according to risk levels. Then, a preset proportion of samples are randomly extracted from each sub-layer through a stratified sampling tool, and finally the preset proportion of samples are combined into several alignment verification sets, which include typical compliance cases and edge cases.
[0092] Step B14: extracting decision tree rules based on the alignment verification set to generate the content output rules.
[0093] In this embodiment, decision tree rule extraction refers to using a tree structure to analyze data and find a combination of conditions that determine whether the output is compliant. Content output rules refer to the explicit requirements that specify the format and conditions that must be met in the model's answer.
[0094] As an optional implementation method, in the scenario of building a large environmental interaction model, decision tree rules are extracted from the compliant samples and non-compliant samples in the verification set, each dialogue answer is broken down into structured features, and the key judgment conditions are found by information gain calculation. The effectiveness of the rules is then tested using samples that have not participated in the training. If a rule is found to be misjudged, the judgment threshold is adjusted by backtracking to the corresponding node of the decision tree, and the optimized rules are finally compiled into executable content output rules.
[0095] For example, in the scenario of building a large environmental interaction model, 500 compliant samples and 200 non-compliant samples in the verification set are input into the decision tree algorithm, and each dialogue answer is broken down into structured features (such as whether it contains the drug name, whether there are dosage instructions, the location of the disclaimer, and other 20 features). The information gain calculation is used to find the key judgment conditions. For example, it is found that "containing the drug name but lacking contraindication reminders" is a common feature of 80% of non-compliant cases. The trained decision tree will generate branch rules such as "If there is a drug name in the answer → check whether there is a contraindication reminder → if not, it is determined to be a violation". These rules are translated into sentences that business personnel can understand (for example, "IF the recommended drug THEN must contain the words 'please follow the doctor's advice'"). Then, the effectiveness of the rules is tested using 300 samples that did not participate in the training. If a rule is found to be misjudged (such as misjudging the reasonable abbreviation "vitamin C" as a violation), the judgment threshold is adjusted by backtracking to the corresponding node of the decision tree (such as allowing specific whitelist words). Finally, the optimized rules are compiled into protocolized decision rules in a standard format and linked to the protocol database. For example, in a medical scenario, when the model generates an answer, it automatically triggers the three-level rule verification of "drug name detection → contraindication reminder verification → format compliance check", and visualizes the rule hit status through the management background (such as using a heat map to show that "missing dosage instructions" is a high-incidence problem), supporting business personnel to add exception clauses with one click, and finally compiling the optimized output rules into content output rules in a standard format.
[0096] By training the decision rules of the interaction model and aligning the data, the model can interact in a corresponding way according to specific scenarios, thereby improving the interactive capabilities of the large model.
[0097] Based on any of the above embodiments, in the fourth embodiment of the present application, step S20 includes steps C11 to C12:
[0098] Step C11 : importing the protocolized decision rule into the interaction model, adjusting the decision parameters of the protocolized decision, and generating a protocolized interaction model.
[0099] In this embodiment, decision parameter adjustment refers to modifying the internal calculation parameters of the model to make it more strictly match the rules. The protocolized interaction model refers to a system that is highly compliant with the protocol after completing rule-intensive training.
[0100] As an optional implementation method, in the scenario of building a large environmental interaction model, based on the interaction model, the protocolized decision rules are converted into parameter adjustment instructions that the model can understand. The parameter adjustment instructions are matched and adjusted with the protocolized decision parameters in the model to make the model comply with the protocolized decision rules and generate a protocolized interaction model.
[0101] Step C12: importing the content output rules into the protocolized interaction model, adjusting the screening parameters of the content output, and generating the rule interaction model.
[0102] In this embodiment, adjusting the screening parameters refers to modifying the filtering threshold and inspection intensity of the model output. The rule interaction model refers to an upgraded intelligent system that integrates protocol rules and content specifications and can accurately control output.
[0103] As an optional implementation method, in the scenario of constructing a large environmental interaction model, based on the protocolized interaction model, by converting the content output rules into adjustable screening parameters that the model can understand, the adjustable screening parameters are matched and adjusted with the screening parameters and protocolized decision parameters of the protocolized interaction model, and the various parameters of the protocolized interaction model are adjusted while giving priority to ensuring that the content output rules are not violated, so as to generate a rule interaction model.
[0104] For example, in the scenario of building a large environmental interaction model, the protocolized decision rules (such as "medical answers must include dosage instructions") are converted into parameter adjustment instructions that the model can understand, and the neural network layers responsible for rule execution (such as the drug keyword detection layer) are located in the interaction model. The weight parameters of these layers are fine-tuned through the gradient descent algorithm - if the model misses the dosage instructions in the historical data, the update intensity of the corresponding parameters of the rule is increased, just like giving the robot's "rule memory area" a booster shot. Dual-channel verification is used in the adjustment process. On the one hand, 3,000 simulated data are used to test the rule hit rate (for example, check whether 90% of the drug recommendation answers have dosage instructions), and on the other hand, the quality of routine conversations is monitored (for example, the fluency of general health consultations cannot be reduced). For complex rules (such as financial products that need to match users at the same time), the rule hit rate is tested with 3,000 simulated data (for example, check whether 90% of the drug recommendation answers have dosage instructions). User risk level and investment period), develop a parameter linkage adjuster, automatically associate multiple rule parameters when the model processes complex conditions (for example, synchronously adjust the period verification parameter when the risk level parameter changes), design adjustable screening parameters for content output rules (such as "all drug recommendations must include dosage, frequency, and contraindications"), for example, adjust the sensitivity parameter of dosage instruction detection from 0.7 to 0.9 (stricter), scan the answers generated by the model in real time through a regular expression matcher (for example, detect whether "once a day" exists in the text), and trigger the automatic completion module to insert standard scripts if missing; at the same time, develop a parameter control panel, allowing operators to drag sliders to adjust the execution intensity of different rules (for example, adjust the priority of the "contraindication reminder" rule to the highest), and complete the adjustment of the rule interaction model.
[0105] By integrating and matching the protocolized decision-making and content output rules, the two rules can be coordinated with each other and the conflicts between them can be eliminated, making the interaction of the interactive model more natural and improving the interactive capabilities of the large model.
[0106] Based on any of the above embodiments, in the fifth embodiment of the present application, step S30 includes steps D11 to D12:
[0107] Step D11 , storing the data of the environment interaction protocol in a database, and constructing an expandable protocol database.
[0108] In this embodiment, the extensible protocol database refers to a protocol database that supports the expansion of more environment interaction protocols and device interaction protocols.
[0109] As an optional implementation, in the scenario of building a large environmental interaction model, additional protocol data including protocol data outside the protocol learning template, interaction protocol data of specific devices and interaction protocol data of specific environments are used to build a protocol database, and users update the protocol database according to their needs.
[0110] Step D12: Based on the rule interaction model, an associated query table is established with the protocol database to generate the target model.
[0111] In this embodiment, the association query table refers to an index table that needs to query protocol rules when defining model decisions.
[0112] As an optional implementation method, in the scenario of building a large environmental interaction model, an associated query table is established according to the protocol database and the rule interaction model, and database query statements are automatically generated. Relevant rules are pulled from the protocol database, and a cache mechanism is established for high-frequency query scenarios. Commonly used protocol rules are preloaded into memory to obtain the target model.
[0113] For example, in the scenario of building a large environmental interaction model, an intelligent query interface for the protocol database is developed in the rule interaction model. For example, when a user requests a prescription, the system automatically generates a database query statement based on the associated query table, pulls relevant rules from the protocol database (such as the contraindications and maximum dosage limit of the drug), establishes a cache mechanism for high-frequency query scenarios (such as financial product recommendations), preloads commonly used protocol rules into memory, allows the model to be called quickly, compresses the response time through the query optimization algorithm, ensures real-time compliance of the model when processing user requests, and generates a target model that records the agreement terms associated with each decision and deploys the target model.
[0114] Since efficient vertical scenario solution migration can be achieved through simple protocol configuration and the introduction of additional low-cost databases, without the need for expensive manual secondary development and high computing power cost of continued training or fine-tuning, the cost of building the large model is reduced.
[0115] Based on any of the above embodiments, in the sixth embodiment of the present application, after step S30, steps E11 to E12 are further included:
[0116] Step E11: establishing a structured protocol interface between the target model and the protocol database based on the association query table.
[0117] In this embodiment, the structured protocol interface refers to a standard communication channel that implements the interaction between the model and the database protocol rules based on a unified data format.
[0118] As an optional implementation method, in the scenario of building a large environmental interaction model, by integrating the protocol interface framework into the target model, pre-compiling dynamic query templates based on the associated query table, and converting the original rules returned by the protocol database into mathematical expressions that can be parsed by the target model through protocol mapping rules, the structured protocol interface is defined in a standard interface format to establish a structured protocol interface between the target model and the protocol database.
[0119] Step E12: Execute scenario deployment of the target model according to the target model, the protocol database, the association query table, and the structured protocol interface.
[0120] In this embodiment, scenario deployment refers to fine-tuning relevant weights of the target model according to the environmental adaptation parameters in the device or scenario to complete the adaptive deployment of the target model.
[0121] As an optional implementation method, in the construction scenario of the environmental interaction large model, the target model is tested for adaptability according to the environmental parameters of the deployment scenario, deployment conflict items are detected, and the relevant weights of the target model deployment conflict items are fine-tuned to make the adaptability of the target model to the deployment scenario reach the preset deployment conditions. Based on the target model and the protocol database, the associated deployment is performed, and the structured protocol interface between the target model and the protocol database is checked according to the associated query table. If it is detected that the structured protocol interface is inconsistent with the associated query table, it is modified according to the associated query table. If it is detected that the structured protocol interface conforms to the corresponding relationship in the associated query table, the deployment of the target model is completed.
[0122] For example, in the scenario of building a large environmental interaction model, a standardized communication channel is set up in the target model. This channel can quickly query the protocol from the protocol database. When the target model needs to make a decision, it first finds the rules to be queried according to the checklist prepared in advance (associated query table), and then sends a query request to the protocol database through the set data format (structured protocol interface). The protocol database quickly finds the corresponding rule clauses based on the problem, translates the protocol rules into calculation formulas that the model can understand, and then transmits them back. Then, the entire system is put into a movable Docker container, and tested in a simulated environment first (for example, trial run with real data from the past three months) to confirm the stability of the new system. Finally, the system of the target model is deployed to the scenario.
[0123] Due to the adaptive model deployment according to different scenarios, the normal operation of the target model can be guaranteed. By detecting the corresponding relationship of the interface, the normal interaction of the target model is guaranteed, and the stability of the large model is improved.
[0124] Based on any of the above embodiments, in the seventh embodiment of the present application, after step S40, steps F11 to F15 are further included:
[0125] Step F11 : Based on the input message and in combination with the target model, an immediate response oriented to natural human interaction is generated.
[0126] In this embodiment, input messages refer to interaction requests such as text, voice, or images sent by users. Natural interaction refers to communication methods that conform to human conversation habits. Instant responses refer to real-time feedback content generated within milliseconds.
[0127] As an optional implementation method of inputting messages, in the application scenario of the large environmental interaction model, the user inputs text and picture messages to the target model through the touch screen, or the user inputs voice messages through language.
[0128] As an optional implementation method of inputting messages, in the application scenario of the large environmental interaction model, the user inputs messages to the target model through remote communication, or the user directly interacts with the target model through language and actions.
[0129] As an optional implementation method of input messages, in the application scenario of the environmental interaction large model, the target model automatically generates input messages based on the recognition of the surrounding environment.
[0130] Step F12: If environmental interaction is involved, construct a control signal through a protocolized decision rule.
[0131] In this embodiment, the control signal refers to a physical instruction generated by the system according to the rules.
[0132] Step F13: parsing the control signal through hard-coded program code, making an interaction request to the associated resource, and obtaining a response feedback.
[0133] In this embodiment, hard-coded program code refers to control logic written directly into the program and cannot be dynamically modified. Associated resources refer to the physical or digital devices that require interaction. Interaction requests refer to the action of sending operational instructions to a device. Response feedback refers to the status data returned by the device.
[0134] As an optional implementation method, in the application scenario of the environmental interaction large model, if the target model involves environmental interaction, the surrounding environment is identified through the target model to obtain the surrounding scene information, and the control signal is constructed according to the scene information and data through the protocolized decision rules in the target model. The instruction content of the control signal is parsed through hard-coded program code, and the interaction is carried out with the associated resources through the instruction content to obtain response feedback from the interactive resources after the interaction.
[0135] Step F14: transmitting the response feedback to the target model to generate an interactive response to obtain the interactive result.
[0136] In this embodiment, the interaction result refers to the conclusion generated by the target model based on the feedback data.
[0137] Step F15: updating the content of the immediate response according to the interactive response.
[0138] In this embodiment, updating the content of the immediate response means, in the case of involving environmental interaction, updating the timely response generated by the input message through the interactive response generated by the environmental interaction.
[0139] As an optional implementation, in the application scenario of the environmental interaction large model, the response feedback generated by the environmental interaction is passed to the target model, and the target model generates an interactive response including the interaction result based on the response feedback, and updates the content of the timely response triggered by the input message according to the interactive response.
[0140] For example, referring to Figure 5 , Figure 5This is a diagram of the protocol learning model interaction system for this application. In the complete protocol learning model interaction steps, the predefined content of <|system|> and <|protocol|> is first defined. The model then accepts user input, forming an input message through <|user|>. The model's response begins with the <|assistant|> field, first generating an immediate response for natural human interaction, wrapped in <|response|>. Then, when interaction with the environment is required, the model begins constructing control signals through the <|control|> field. If an operation type not supported by <|protocol|> is encountered, the associated interaction protocol is retrieved from an external protocol database and wrapped in <|extra_protocol|>. Then, in the <|control|> message section, the control signal content is constructed according to the retrieved external protocol. Hard-coded program code parses the structured signal and requests interaction with resources such as associated devices or network environments. Upon receiving a response, the feedback message is wrapped in the <|feedback|> tag and passed to the model. Further iterations generate the model's <|response|> message, which represents the interaction result, to respond to the human.
[0141] Due to the joint effect of protocolized decision rules and protocol database, the target model can interact freely with any scenario and any device. The scalability of the target model is further improved through the setting of the protocol database.
[0142] Based on any of the above embodiments, in the eighth embodiment of the present application, after step F12, steps G11 to G12 are further included:
[0143] Step G11: If the environmental interaction has an operation type outside the protocolized decision rule, query the protocol database according to the associated query table to obtain the interaction protocol corresponding to the operation type.
[0144] In this embodiment, the operation types outside the protocolized decision rules refer to the interactive behaviors that are not covered by the existing rule base. The interactive protocol refers to the execution steps and constraints defined for a specific operation type.
[0145] As an optional implementation method, in the application scenario of the environmental interaction large model, if there are operation types outside the protocolized decision rules in the environmental interaction of the target model, the operation type is identified, the associated query table is queried, and the interaction protocol corresponding to the operation type is found in the protocol database.
[0146] Step G12: Based on the interaction protocol, update the control signal constructed by the protocolized decision rule.
[0147] In this embodiment, updating refers to a technical solution of replacing decision rules without interrupting service.
[0148] As an optional implementation method, in the application scenario of the environmental interaction large model, the control signal constructed by the protocolized decision rule is updated and supplemented according to the corresponding interaction protocol in the protocol database to obtain a complete control signal.
[0149] For example, referring to Figure 6 , Figure 6 This is an example diagram of the protocol database of this application. In the personal computing scenario that deeply integrates artificial intelligence technology, the external environment is specifically a PC terminal, and its control and response can be achieved by reorganizing the underlying source code of various third-party code libraries or service providers. This embodiment is combined with the aforementioned solution of the present invention to illustrate: the external database and the protocol structure in training are optional. Generally speaking, common fields include: unique identifier: a serial number used to identify the protocol. Index description: used to build a fast index of the protocol content. Protocol description: used to describe the complete protocol content. Input format convention: used to describe the input information passed by the model to the external environment. Output format convention: used to describe the output information of the external environment interaction. Exception message convention: used to describe the exception message of the external environment interaction. Special case description: The index description and the protocol description may share fields to simplify the engineering implementation. The implementation of the database is optional, including but not limited to traditional relational (databases, etc.), non-relational databases (GraphDB, KV database), vector databases, etc. The protocol database in the solution description is optional, which is a supplementary module used to expand and support the compatibility of the protocol learning model in multiple scenarios. The label identification of each protocol field in the solution description is optional. As long as the actual content and meaning of the response label are met, they fall within the scope of the solution. The pre-trained weights in the solution description are optional, including but not limited to the current open source and closed source pre-trained weights, and including but not limited to text modality and multimodal pre-trained weights. In the training phase of the solution description, except for the protocol learning step, the order and application of the remaining learning phases are optional. The optimization algorithm is optional, including but not limited to optimization algorithms based on optimizers such as SGD and Adam. The optimization scheme is optional, including but not limited to optimization of full model training, optimization based on LoRA (low rank adaptation) and optimization scheme based on reinforcement learning. The application of the training framework is optional, including but not limited to training frameworks based on torch, tensorflow, deepspeed, and megatron.
[0150] Due to the joint effect of protocolized decision rules and protocol database, the target model can interact freely with any scenario and any device, which improves the recognition ability and scalability of the large model.
[0151] Based on any of the above embodiments, in the ninth embodiment of the present application, before step S10, steps H11 to H13 are further included:
[0152] Step H11, based on the initial model, extracting semantic features by receiving multimodal raw data, and generating pre-trained weight parameters containing semantic features.
[0153] In this embodiment, deployment refers to the process of integrating a trained model into a business system for user use. An initial model refers to a basic neural network architecture that has not been trained for a specific task. Multimodal raw data refers to unlabeled mixed data in various formats, such as text, images, and audio. Semantic feature extraction refers to identifying meaningful information units from data. Pretrained weight parameters refer to the set of mathematical parameters stored by the model after pretraining, which are used for subsequent task migration.
[0154] As an optional implementation method, in the scenario of building a large environmental interaction model, the multimodal raw data is first cleaned. Based on the initial model, the cleaned multimodal raw data is received and semantic features are extracted. The initial model automatically associates and identifies the semantic features and generates pre-trained weight parameters covering the semantic features.
[0155] Step H12: Based on the pre-trained weight parameters, an attention mechanism is used to perform feature space mapping on different modal data, and the pre-trained weight parameters are optimized and aligned by minimizing the inter-modality difference function to generate target weight parameters.
[0156] In this embodiment, the attention mechanism refers to a technique that allows the model to dynamically focus on the importance of different parts of the data. Feature space mapping refers to the process of converting data from different modalities into a unified mathematical representation. The intermodal discrepancy function is a loss function that measures the degree of matching between features of different modalities. Optimal alignment refers to adjusting parameters to achieve consistent feature representation across modalities. The target weight parameter refers to the final model parameter that achieves cross-modal alignment after optimization.
[0157] As an optional implementation method, in the scenario of building a large environmental interaction model, pre-trained weight parameters are loaded into the multimodal model, paired image and text data are input, and the association mapping matrix between the local image features and the text word vectors in the image and text data is calculated through the cross-modal attention layer to generate a feature space mapping. Based on the feature space mapping, the initial model image and text matching is calibrated by minimizing the inter-modal difference function, the pre-trained weight parameters are optimized, and the data is associated and aligned to generate the target weight parameters.
[0158] Step H13: Based on the target weight parameters and in combination with the vertical domain annotation data, a contrastive learning algorithm is used to update the parameters of the initial model to generate the initial model.
[0159] In this embodiment, vertical domain annotated data refers to labeled data for a specific business scenario. Contrastive learning algorithms are training methods that optimize models by comparing the differences between positive and negative samples.
[0160] As an optional implementation method, in the scenario of constructing a large environmental interaction model, relevant scene recognition training is performed based on the initial model with imported target weight parameters and combined with the pre-processed vertical domain annotation data. By constructing positive sample pairs and negative sample pairs, the initial model is trained using a contrastive learning framework. Based on the matching degree of the scene and label data as a reference, positive sample pairs are constructed by training the initial model, and the matching degree of the positive sample pairs and the negative sample pairs is compared. The matching degree of the positive sample pairs constructed by the training initial model is higher than that of the negative sample pairs. By dynamically adjusting the matching degree of the negative sample pairs, the initial model is dynamically trained, the recognition parameters of the initial model are updated, and the initial model is generated.
[0161] For example, referring to Figure 7 , Figure 7 This is a flowchart for a complete training solution. Pre-training is performed based on multimodal original data sets (text, images, and sensor time series data). Feature space mapping is achieved through a cross-modal contrastive learning framework to generate pre-trained weight parameters with basic semantic understanding. In response to vertical domain task requirements, instruction fine-tuning technology is used to supervise the model training. Vertical domain annotation data sets (such as paired data of industrial equipment operation logs and corresponding control instructions) are used to optimize the task adaptability of the model. The protocol learning module is implemented simultaneously, and the business rule library (such as safety operating specifications and medical compliance clauses) is encoded as structured constraints. The regularized loss function is injected into the model parameter update process to introduce high-quality in the post-training stage. Align the data sets and use contrastive learning algorithms to optimize the matching degree between model output and protocol rules. During the deployment phase, dynamically couple the trained model with the protocol database, develop a real-time rule retrieval interface (based on database query engine and vector similarity matching), and implement the protocolized decision flow of environmental interaction requests (environmental state perception → protocol rule matching → control signal generation → execution feedback loop). Build an incremental learning mechanism for new interaction scenarios, continuously absorb uncovered protocol cases through the online learning framework, and update the protocol database version after verification by the digital twin system, forming a large-scale environmental interaction model architecture of "data-driven-rule-constrained-dynamic evolution".
[0162] By training with vertical domain annotation data and multimodal raw data, the model can recognize various scenarios and various semantic relationships, which improves the recognition ability of the large model.
[0163] Based on any of the above embodiments, in the tenth embodiment of the present application, refer to Figure 8 , Figure 8 This is a diagram of the industrial control software sub-item in the application agreement database.
[0164] Exemplarily, it is used to refine the control and feedback capabilities of the model system. The external environment is specifically an industrial control system. If the industrial control simply includes three capabilities: open, close, and run, the industrial control software protocol item will be placed in the protocol database. For "open" and "close", there is no need to explicitly input the content, and the output content is a true / false value. The former indicates successful completion, and the latter indicates failure to complete. For abnormal messages, it should include artificial hard-coded and natural language messages generated by different abnormal situations. For "run", the operating parameters are described and defined according to the actual situation to help the large model organize the corresponding input information into input parameters of the corresponding operation format.
[0165] The present application provides an environmental interaction large model device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the environmental interaction large model construction method based on protocol learning in the above-mentioned embodiment one.
[0166] Reference below Figure 9 , which shows a schematic diagram of the structure of an environmental interactive large model device suitable for implementing the embodiments of the present application. The environmental interactive large model device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, interactive large model devices, personal digital assistants (PDAs), tablet computers (PADs), portable multimedia players (PMPs), interactive devices, etc., as well as fixed terminals such as human-computer interaction devices and desktop computers. Figure 9 The large environmental interaction model device shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.
[0167] like Figure 9As shown, the large model device for environmental interaction may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. Various programs and data required for the operation of the large model device for environmental interaction are also stored in the random access memory 1004. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are connected to each other via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touch screen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 1003 including, for example, a magnetic tape, hard disk, etc.; and communication devices 1009. The communication devices 1009 can allow the environment interaction model device to communicate with other devices wirelessly or by wire to exchange data. Although the figure shows an environment interaction model device with various systems, it should be understood that it is not required to implement or have all of the systems shown. More or fewer systems can be implemented or have instead.
[0168] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are performed.
[0169] The large-scale model device for environmental interaction provided by this application utilizes the protocol learning-based large-scale model construction method for environmental interaction provided by the aforementioned embodiment, which can resolve the technical problem of related technologies lacking scenario expansion capabilities in environmental interaction scenarios, resulting in low interaction compatibility. Compared with the prior art, the beneficial effects of the large-scale model device for environmental interaction provided by this application are the same as those of the protocol learning-based large-scale model construction method for environmental interaction provided by the aforementioned embodiment, and the other technical features of the large-scale model device for environmental interaction are the same as those disclosed in the method of the aforementioned embodiment, and are not further described here.
[0170] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.
[0171] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
[0172] The present application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the method for constructing a large environmental interaction model based on protocol learning in the above-mentioned embodiment.
[0173] The computer-readable storage medium provided in this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination thereof. More specific examples of computer-readable storage media can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM, Erasable Programmable Read Only Memory or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM, CD-Read Only Memory), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in combination with an instruction execution system, system or device. The program code contained on the computer-readable storage medium can be transmitted using any appropriate medium, including but not limited to: wires, optical cables, radio frequency (RF, Radio Frequency), etc., or any suitable combination thereof.
[0174] The computer-readable storage medium may be included in the large-scale model device for environmental interaction, or may exist independently without being assembled into the large-scale model device for environmental interaction.
[0175] The above-mentioned computer-readable storage medium carries one or more programs. When the above-mentioned one or more programs are executed by the environmental interaction large model device, the environmental interaction large model device: based on the initial model, combined with the protocol learning template, performs model training for protocol learning to generate an interaction model; based on the interaction model, combined with the reinforcement learning framework, generates protocolized decision rules, and based on the output content simulated by the interaction model, performs alignment data training to generate content output rules; based on the interaction model, combines the protocolized decision rules and the content output rules to generate a rule interaction model; based on the rule interaction model, combines the protocolized decision rules and the content output rules, performs data query association to generate a target model, and executes the deployment of the target model.
[0176] Computer program code for performing the operations of the present application can be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0177] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.
[0178] The modules described in the embodiments of the present application may be implemented in software or hardware, wherein the name of a module does not necessarily limit the unit itself.
[0179] The readable storage medium provided in this application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., a computer program) for executing the above-mentioned method for constructing a large environmental interaction model based on protocol learning. This can solve the technical problem that the related technology lacks scenario expansion capabilities in environmental interaction scenarios, resulting in low interaction compatibility. Compared with the existing technology, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the method for constructing a large environmental interaction model based on protocol learning provided in the above-mentioned embodiment, and will not be repeated here.
[0180] The above description is only part of the embodiments of the present application and does not limit the patent scope of the present application. All equivalent structural transformations made by using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect application in other related technical fields are included in the patent protection scope of the present application.
Claims
1. A method for constructing a large environmental interaction model based on protocol learning, characterized in that: The method comprises: Importing a protocol learning template into an initial model, performing protocol learning on the initial model, and generating an interaction model; Modifying the interaction model to obtain a rule interaction model according to the protocolized decision rule and the content output rule; The rule interaction model is associated with a protocol database through data query to generate a target model, wherein the protocol database includes an environment interaction protocol.
2. The method for constructing a large environmental interaction model based on protocol learning according to claim 1, characterized in that: The steps of importing the protocol learning template into the initial model, performing protocol learning on the initial model, and generating an interaction model include: Based on the initial model, importing the protocol learning template, performing protocol learning training for the dialogue instruction task, and generating protocol learning training weights; Importing the protocol learning training weights into the initial model, performing a protocol learning test, and generating a test result; Verifying the protocol learning and training weights according to the test results to obtain verified protocol learning and training weights; The verified protocol learning training weights are imported into the initial model, the pre-training weights of the initial model are updated, and the interaction model is generated.
3. The method for constructing a large-scale environment interaction model based on protocol learning according to claim 1, characterized in that: Before the step of modifying the interaction model to obtain a rule interaction model according to the protocolized decision rule and the content output rule, the method further includes: Based on the interaction model and reinforcement learning framework, the logical rules of the business scenario are simulated to generate state observations; The interaction model generates a feedback decision according to the state observation value, and associates the state observation value with the feedback decision to generate the protocolized decision rule; Based on the interaction model, combined with the protocolized decision rules and preset conditions, a content sample that meets the preset conditions is generated, and an alignment verification set is constructed through a stratified random sampling strategy; Based on the alignment verification set, decision tree rule extraction is performed to generate the content output rule.
4. The method for constructing a large environmental interaction model based on protocol learning according to claim 1, characterized in that: The step of modifying the interaction model to obtain a rule interaction model according to the protocolized decision rule and the content output rule includes: Importing the protocolized decision rule into the interaction model, adjusting the decision parameters of the protocolized decision, and generating a protocolized interaction model; The content output rules are imported into the protocolized interaction model, and the screening parameters of the content output are adjusted to generate the rule interaction model.
5. The method for constructing a large environmental interaction model based on protocol learning according to claim 1, characterized in that: The step of performing data query and associating the rule interaction model with the protocol database to generate a target model, wherein the protocol database includes the environment interaction protocol, comprises: Storing the data of the environment interaction protocol in a database to construct an expandable protocol database; Based on the rule interaction model, an associated query table is established with the protocol database to generate the target model.
6. The method for constructing a large-scale environment interaction model based on protocol learning according to claim 1, characterized in that: After the step of performing data query and associating the rule interaction model with the protocol database to generate a target model, wherein the protocol database includes the environment interaction protocol, the method further includes: Establishing a structured protocol interface between the target model and the protocol database based on the association query table; Scenario deployment of the target model is performed according to the target model, the protocol database, the association query table, and the structured protocol interface.
7. The method for constructing a large-scale environment interaction model based on protocol learning according to claim 6, characterized in that: After the step of executing the scenario deployment of the target model according to the target model, the protocol database, the data query table, and the structured protocol interface, the method further includes: When it comes to environmental interaction, control signals are constructed through protocolized decision rules; Parsing the control signal through hard-coded program code, making an interactive request to the associated resource, and obtaining a response feedback; The response feedback is input into the target model to generate an interactive response to obtain an interactive result; Update the content of the immediate response based on the interactive response.
8. The method for constructing a large environmental interaction model based on protocol learning according to claim 7, characterized in that: If environmental interaction is involved, after the step of constructing a control signal through a protocolized decision rule, the following steps may also be included: If the environmental interaction has an operation type outside the protocolized decision rule, query the protocol database according to the associated query table to obtain the interaction protocol corresponding to the operation type; Based on the interaction protocol, the control signal constructed by the protocolized decision rule is updated.
9. A large model device for environmental interaction, characterized in that: The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the method for constructing a large environmental interaction model based on protocol learning as described in any one of claims 1 to 8.
10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by the processor, the steps of the method for constructing a large environmental interaction model based on protocol learning as described in any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Protocol vulnerability evaluation method and device and storage medium
CN114553533A
Decision model interpretable method based on decision tree
CN117332842A
Method and device for realizing adaptive learning based on large language model
CN117891903A
Heterogeneous robot cluster management platform and management method thereof
CN119996462A
Method of building a data integration environment
US20100070450A1