Cross-model knowledge editing and updating method and device, equipment and medium

By employing a cross-model knowledge editing and updating method, the problems of low efficiency and consistency in knowledge updates in multi-model deployment environments in the fintech and healthcare fields are solved. This enables knowledge sharing and migration among multiple models, improving the timeliness and reliability of the system.

CN121860013APending Publication Date: 2026-04-14PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-06
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

In the fields of fintech and healthcare, existing knowledge management methods lack the ability to share and transfer knowledge across models, resulting in low efficiency and difficulty in ensuring consistency of knowledge updates in multi-model deployment environments. Furthermore, traditional methods are prone to knowledge forgetting or error propagation, affecting system stability and reliability.

Method used

The cross-model knowledge editing and updating method includes receiving knowledge update instructions, generating knowledge editing vectors, writing intermediate layer parameters into the target model set using a cross-model mapping layer, generating an input sequence with control tags, and performing output alignment and fusion processing through a consistency fusion module. Combined with conflict detection and re-inference mechanisms, the timeliness and consistency of knowledge updates are ensured.

Benefits of technology

It enables the sharing and migration of knowledge-edited content across multiple models, avoiding the inefficiency caused by repeated editing, improving the timeliness, consistency, and reliability of knowledge updates, and ensuring system stability and accuracy in fintech and healthcare scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121860013A_ABST
    Figure CN121860013A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, can be applied to business scenes of financial science and technology, medical health and the like, and discloses a cross-model knowledge editing and updating method, device, equipment and medium, and the method comprises the steps: receiving a knowledge updating instruction, and generating a knowledge editing vector; writing the knowledge editing vector into an intermediate layer parameter of the target model set to obtain a parameter-updated target model set; generating an input sequence with a regulation mark when the input request and the knowledge update content meet semantic matching conditions, and inputting the input sequence into the target model set to obtain a model output set; performing output alignment and fusion through a consistency fusion module to generate a fusion output result; and when the output conflict is not eliminated, performing a re-reasoning operation based on the knowledge editing vector, and outputting a final result. According to the method, knowledge sharing and dynamic triggering among multiple models are realized through a cross-model mapping and consistency fusion mechanism, and the efficiency, timeliness and credibility of knowledge updating are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device, and medium for cross-model knowledge editing and updating. Background Technology

[0002] In the fintech sector, large language models have been widely applied in scenarios such as intelligent customer question answering, claims process guidance, compliance review, and risk assessment. However, existing technologies have significant shortcomings in knowledge management. Traditional approaches often rely on periodic full or incremental training to update model knowledge, a time-consuming and computationally resource-intensive process that struggles to adapt to timely updates in insurance terms, regulatory policy changes, and the rapid iteration of financial products. Furthermore, some existing knowledge editing methods can only modify local parameters of a single model, lacking cross-model sharing and transfer capabilities. This forces enterprises to perform separate update operations for each model when simultaneously deploying customer service, compliance, and risk control models. This not only leads to inefficiency but also inconsistencies in knowledge updates, impacting the overall system reliability. Moreover, existing methods are prone to interfering with the original knowledge system during knowledge modification, manifesting as knowledge forgetting or error propagation, reducing stability and credibility in financial applications.

[0003] In the healthcare sector, large language models are used for patient consultation assistance, treatment pathway recommendations, drug instruction parsing, and medical literature interpretation. Similar to the financial sector, medical knowledge systems are frequently updated; clinical guidelines, drug instructions, regulatory standards, and treatment plans are updated regularly. However, existing retraining-based models suffer from excessively long update cycles and computational costs, failing to meet the demands for rapid knowledge updates. Furthermore, existing single-model knowledge editing methods lack mechanisms for unified migration and sharing between different types of medical application models. This necessitates hospitals or medical institutions to repeatedly modify knowledge for each model when simultaneously deploying patient interaction models, medical knowledge retrieval models, and auxiliary diagnosis and treatment models. This not only increases operational complexity but can also lead to inconsistent knowledge updates, affecting the accuracy and consistency of treatment recommendations. More seriously, inappropriate knowledge update methods may result in the loss of existing medical knowledge or the creation of misleading information, posing potential risks to medical decisions and reducing the usability and security of large language models in healthcare scenarios. Summary of the Invention

[0004] The main objective of this invention is to provide a cross-model knowledge editing and updating method, apparatus, device, and storage medium, aiming to solve the technical problem that existing knowledge editing methods can only adjust local parameters for a single model and lack cross-model sharing and migration mechanisms, resulting in low efficiency and difficulty in ensuring consistency of knowledge updates in multi-model deployment environments.

[0005] To achieve the above objectives, this invention provides a cross-model knowledge editing and updating method, comprising: Receive a knowledge update instruction and parse the knowledge update content from the knowledge update instruction; The knowledge update content is encoded using the knowledge editing module to generate a knowledge editing vector; The knowledge editing vector is input into the cross-model mapping layer, and the knowledge editing vector is written into the intermediate layer parameters of each model in the target model set through the cross-model mapping layer to obtain the target model set with updated parameters. Context control tags are generated based on the knowledge update content. Input requests are converted into model input sequences. When the input requests and knowledge update content meet the semantic matching conditions, the context control tags are inserted into the model input sequences to obtain input sequences with control tags. The input sequence with the control label is input into the target model set after the parameter update to obtain a model output set containing the outputs of each model; The consistency fusion module performs output alignment and fusion processing on the model output set to generate a fused output result. When output conflicts occur in the model output set and cannot be eliminated through fusion processing, a second inference operation is performed based on the knowledge editing vector to output the final result.

[0006] Furthermore, to achieve the above objectives, the present invention provides a cross-model knowledge editing and updating apparatus, comprising: The instruction parsing module is used to receive knowledge update instructions and parse the knowledge update content from the knowledge update instructions; The knowledge editing module is used to encode the knowledge update content and generate a knowledge editing vector. The cross-model mapping module is used to input the knowledge editing vector into the cross-model mapping layer, and write the knowledge editing vector into the intermediate layer parameters of each model in the target model set through the cross-model mapping layer to obtain the target model set with updated parameters. The context control module is used to generate context control tags based on the knowledge update content, convert the input request into a model input sequence, and insert the context control tags into the model input sequence when the input request and the knowledge update content meet the semantic matching conditions, so as to obtain an input sequence with control tags. The model inference module is used to input the input sequence with the control label into the target model set after parameter update to obtain a model output set containing the outputs of each model; The consistency fusion module is used to perform output alignment and fusion processing on the model output set to generate a fused output result. The conflict handling module is used to perform a second inference operation based on the knowledge editing vector and output the final result when output conflicts occur in the model output set and cannot be eliminated by fusion processing.

[0007] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and a cross-model knowledge editing and updating program stored in the memory and executable on the processor, wherein when the cross-model knowledge editing and updating program is executed by the processor, it implements the steps of the cross-model knowledge editing and updating method as described above.

[0008] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing a cross-model knowledge editing and updating program, wherein the cross-model knowledge editing and updating program, when executed by a processor, implements the steps of the cross-model knowledge editing and updating method as described above.

[0009] Beneficial Effects: This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as fintech and healthcare. It discloses a cross-model knowledge editing and updating method, apparatus, device, and medium, comprising: receiving a knowledge update instruction and generating a knowledge editing vector; writing the knowledge editing vector into the intermediate layer parameters of a target model set to obtain a target model set with updated parameters; generating an input sequence with control tags when the input request and the knowledge update content satisfy semantic matching conditions, and inputting it into the target model set to obtain a model output set; performing output alignment and fusion through a consistency fusion module to generate a fused output result; and performing a re-inference operation based on the knowledge editing vector when output conflicts are not resolved, outputting the final result. This invention, by introducing a cross-model mapping layer and a consistency fusion mechanism in a multi-model environment, realizes the sharing and migration of knowledge editing content among multiple models, avoiding the inefficiency caused by repeatedly editing a single model in traditional methods. When there is a semantic relationship between the input request and the knowledge update content, dynamic triggering is achieved through context control tags, enabling effective utilization of knowledge updates during the inference process. Furthermore, by combining conflict detection and re-inference mechanisms, reliable final results are still generated even when model outputs are inconsistent, thereby improving the timeliness, consistency, and application reliability of knowledge updates. Attached Figure Description

[0010] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings: Figure 1 This is a schematic diagram of an application environment for a cross-model knowledge editing and updating method according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating an embodiment of the cross-model knowledge editing and updating method of the present invention; Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the cross-model knowledge editing and updating device of the present invention; Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention; Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation

[0011] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0012] The cross-model knowledge editing and updating method provided in this invention can be applied to, for example... Figure 1 In this application environment, the client communicates with the server via a network. The server can receive knowledge update instructions from the client and generate knowledge edit vectors; write the knowledge edit vectors into the intermediate layer parameters of the target model set to obtain the target model set with updated parameters; when the input request and the knowledge update content meet the semantic matching conditions, an input sequence with control tags is generated and input into the target model set to obtain the model output set; the output is aligned and fused through the consistency fusion module to generate the fused output result; when the output conflict is not eliminated, a second inference operation is performed based on the knowledge edit vector to output the final result. This invention, by introducing a cross-model mapping layer and a consistency fusion mechanism in a multi-model environment, realizes the sharing and migration of knowledge edit content among multiple models, avoiding the inefficiency caused by repeatedly editing a single model in traditional methods. When there is a semantic relationship between the input request and the knowledge update content, dynamic triggering is achieved through context control tags, enabling knowledge updates to be effectively utilized in the inference process. Furthermore, by combining conflict detection and second inference mechanisms, reliable final results are still generated even when the model outputs are inconsistent, thereby improving the timeliness, consistency, and application reliability of knowledge updates. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will now be described in detail through specific embodiments.

[0013] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the cross-model knowledge editing and updating method provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0014] likeFigure 2 As shown, the cross-model knowledge editing and updating method proposed in this invention includes the following steps: S10, Receive a knowledge update instruction and parse the knowledge update content from the knowledge update instruction; In this embodiment, the receiving of knowledge update instructions typically involves receiving structured or unstructured information from users or the system, which can take the form of text commands, configuration files, network requests, database entries, etc. The source of the knowledge update instructions may be operation commands input by business personnel, automatic trigger signals from the business system, or even update requests pushed by external monitoring platforms via interfaces. This receiving process requires a standardized interface mechanism to ensure that instructions from different sources can be accurately received and transmitted.

[0015] Parsing the received knowledge update instructions to obtain the updated knowledge content is the next step. This parsing process goes beyond simply processing the text word by word; it also requires understanding the structure and semantics of the input data. Parsing can be achieved through a rule-based parsing engine, such as keyword extraction based on regular expression matching, or through a natural language processing module that performs word segmentation, syntactic analysis, and semantic recognition on the input text. In some implementations, predefined instruction templates can be used to quickly determine the boundaries and key parameters of the updated content. For example, the insurance type, clause number, and revision item in an insurance clause update instruction can be mapped to structured fields in the updated content.

[0016] The parsing results need to be transformed into knowledge update content, which is typically stored in a structured format, such as key-value pairs, JSON, XML, or database records, ensuring that subsequent processing can be performed based on a well-defined data structure. In practice, the parsing process may also involve integrity checks and consistency checks to prevent missing fields or logical conflicts in the updated content. For example, if a drug instruction update instruction in the healthcare field lacks a drug code, the system can trigger a verification mechanism, indicating that the content is incomplete and preventing further processing.

[0017] In practice, receiving knowledge update instructions can be accomplished through various communication methods. Update requests from business systems can be received via HTTP interfaces, instructions from asynchronous task schedulers can be received via message queues, or update requests manually entered via a human-computer interaction interface. In high-concurrency scenarios, to prevent instruction loss, buffer queues or distributed log systems can be added to ensure that all update instructions are persistently stored and processed sequentially.

[0018] Multiple approaches can be employed when parsing knowledge update instructions. One approach is to pre-define common update instruction formats based on a rule engine, quickly extracting key update parameters. Alternatively, a natural language understanding module can be used to perform deep analysis of the input text, automatically extracting relevant business fields. The parsing method can be flexibly adjusted for different scenarios. For example, in the fintech sector, the focus is on extracting fields such as product category, interest rate range, and regulatory terms; in the healthcare sector, the focus is on extracting fields such as drug name, indications, and dosage restrictions.

[0019] The updated knowledge content generated after parsing can be directly written to the database or stored as a temporary cache in an in-memory data structure for frequently called update tasks. A version control mechanism can also be used to compare the updated content with the old version, generating difference records for auditing and backtracking.

[0020] This embodiment receives knowledge update instructions and parses them to obtain the updated knowledge content. This avoids relying on retraining the model to update knowledge, reducing dependence on computing resources and time. Because the parsing process can directly extract structured update information, it achieves an efficient and flexible knowledge update method, enabling the system to quickly respond to frequently changing business rules and external regulatory requirements, thereby improving the real-time performance and reliability of the application.

[0021] S20, the knowledge update content is encoded through the knowledge editing module to generate a knowledge editing vector; In this embodiment, the process of encoding the knowledge update content and generating the knowledge edit vector through the knowledge editing module can be broken down into several steps. First, the knowledge update content needs to be input into the knowledge editing module. The input format can be standardized structured text, a set of key-value pairs, or parsed serialized data. The focus of the input stage is to ensure the completeness and consistency of the content. For example, in insurance business, input fields include insurance type, clause number, and scope of application; in healthcare business, input fields include drug code, dosage information, and indication description.

[0022] Once inside the knowledge editing module, it typically contains a neural network-based encoder structure, which can employ the Transformer architecture. The Transformer encoder uses a multi-head attention mechanism to capture the dependencies between different parts of the text and extract semantic features, enabling high-quality semantic representations even when dealing with cross-domain or cross-format knowledge updates. Through the multi-head attention mechanism, each part of the knowledge update interacts with other parts, thus forming a global semantic context.

[0023] After obtaining the semantic representation, these representations need to be nonlinearly mapped using a feedforward neural network to further process the high-dimensional features to meet the subsequent knowledge vector generation requirements. A commonly used approach in this process is the multilayer perceptron, which can enhance the expressive power of the semantic representation and ensure that the knowledge edit vectors have good adaptability in subsequent cross-model mapping.

[0024] Next, normalization is applied in the module. Normalization can be achieved through layer-level normalization, primarily to prevent gradient vanishing or exploding problems in deep networks, ensuring numerical stability of the encoded results. Normalized high-dimensional features become smoother, facilitating unified processing.

[0025] In some implementations, dimensionality reduction is also required on the normalized high-dimensional vectors, mapping them to a lower-dimensional space to reduce storage and computational overhead. Dimensionality reduction can be achieved through linear transformations or adaptive projections. The final generated vectors then need to be processed by a normalization unit to normalize them into standard vectors with stable numerical distributions. This normalization result is the knowledge-edited vector, which can be stored in the module's vector repository for use across model mapping layers.

[0026] In different implementations, different techniques can be used for encoding. Deep neural networks based on Transformers can be used to encode the updated content, or lightweight bidirectional recurrent neural networks can be used to reduce computational overhead. Furthermore, in some edge computing scenarios, convolutional neural networks combined with attention mechanisms can be used to quickly encode short text updates.

[0027] The dimensionality reduction method can also be adjusted according to the application scenario. If the updated content is long and the dimensional space is large, linear dimensionality reduction methods such as principal component analysis can be used; if the updated content involves complex semantics, nonlinear dimensionality reduction using an autoencoder structure can be used to retain more contextual information. In healthcare scenarios, the updated knowledge often contains complex diagnostic and treatment rules, making nonlinear dimensionality reduction more suitable; while in fintech businesses, the updated content is mostly structured parameters, and linear dimensionality reduction methods are sufficient.

[0028] In the normalization stage, layer normalization or batch normalization can be selected. Layer normalization is suitable for handling sequential inputs, while batch normalization can improve the efficiency of training and inference when performing large-scale parallel processing.

[0029] This embodiment transforms complex and diverse update content into a unified numerical representation by inputting the updated knowledge content into the knowledge editing module and generating knowledge editing vectors. This approach preserves the semantic features of the updated content while reducing the differences in format and semantics between content from different sources, thus providing a solid foundation for cross-model sharing and transfer. This method effectively avoids misjudgment problems caused by inconsistent formats or information loss, improving the accuracy and stability of knowledge updates.

[0030] S30, the knowledge editing vector is input into the cross-model mapping layer, and the knowledge editing vector is written into the intermediate layer parameters of each model in the target model set through the cross-model mapping layer to obtain the target model set with updated parameters; In this embodiment, inputting the knowledge edit vector into the cross-model mapping layer means establishing a computational structure capable of accepting numerical representations. The cross-model mapping layer is a parameter adaptation component for multi-model environments, its core objective being to transform a single knowledge edit vector into parameter update signals that can be directly applied to multiple different models. Knowledge edit vectors are generally low-dimensional and standardized representations, but the internal parameter structures of different target models vary significantly. For example, in financial business, compliance models and customer service models may have completely different intermediate layer dimensions; therefore, a cross-model mapping layer is needed for structural matching and parameter transformation.

[0031] A cross-model mapping layer typically includes a parameter mapping network that analyzes the intermediate layer parameter configurations of each model in the target model set. Specifically, it needs to identify the dimensions, weight matrix size, and parameter distribution of the intermediate layers. For example, a customer service model might use bidirectional Transformer layers, while a claims model might use a combination of convolutional and attention layers. The parameter mapping network reads this information from metadata tables or model description files and dynamically generates mapping rules.

[0032] After obtaining the parameter structure of the target model, the cross-model mapping layer converts the knowledge edit vectors into parameter adjustment values ​​for each model. This process can be accomplished through linear mapping, nonlinear projection, or matrix factorization. The purpose of this transformation is to ensure that the knowledge edit vectors are injected into the intermediate layer in a form that meets the dimensional requirements of the target model, enabling different models to maintain consistent knowledge representation when sharing the same knowledge updates.

[0033] The adjusted parameter values ​​after transformation are written into the intermediate layer weight matrix of the corresponding language model. This writing operation is not a simple overwrite, but rather a fusion approach, commonly employing weighted updates or residual updates. Weighted updates distribute the proportions between the original weights and the adjusted values, allowing new knowledge to be gradually introduced; residual updates, on the other hand, add the adjusted values ​​to the original weights to reduce interference with existing knowledge.

[0034] Once the intermediate layers of all target models have completed parameter writing and fusion, the cross-model mapping layer will recombine these models to form a set of target models with updated parameters. This set is globally consistent, ensuring that different models can share the same knowledge update content when processing requests.

[0035] In one implementation, the cross-model mapping layer employs multi-head linear projection, simultaneously mapping the knowledge edit vector to different dimensional spaces before matching it with the parameter matrices of different models. This approach is suitable for multi-model environments with significant differences in parameter scale.

[0036] Another implementation approach is to use a hierarchical parameter mapping method. First, the knowledge editing vector is mapped to a high-dimensional general representation, then decomposed into several sub-vectors, each corresponding to the parameter update requirements of a class of models. This approach facilitates grouped updates in multi-task systems.

[0037] A dynamic weight adjustment mechanism can also be employed. The cross-model mapping layer dynamically determines the writing ratio of knowledge editing vectors in different models by monitoring the task usage frequency and update requirements of each model. For example, in the fintech business, if regulatory rules are updated more frequently, the compliance model is assigned a higher writing weight when parameters are updated. In the healthcare business, if drug information changes more frequently, the diagnosis and treatment model receives a higher writing weight.

[0038] This embodiment enables the synchronous application of a single knowledge update across multiple models by inputting the knowledge editing vector into the cross-model mapping layer and writing it into the intermediate layer parameters of multiple models, thus avoiding the resource waste caused by repeated updates. Simultaneously, the fusion update method reduces the risk of knowledge forgetting and semantic conflicts, ensuring knowledge consistency in a multi-model environment. This approach significantly improves the real-time performance and stability of knowledge updates.

[0039] S40, generate context control tags based on the knowledge update content, convert the input request into a model input sequence, and insert the context control tags into the model input sequence when the input request and the knowledge update content meet the semantic matching conditions, to obtain an input sequence with control tags; In this embodiment, generating context-modulation tags based on knowledge-updated content first requires using the knowledge-updated content as input. Knowledge-updated content typically exists in the form of standardized text, structured entries, or symbolic rules, such as revisions to financial regulatory provisions or new prohibitions in medical drug usage guidelines. By parsing this type of knowledge-updated content, the system generates an associated tag string, which is the context-modulation tag. The role of the context-modulation tag is to serve as an additional cue signal, helping the language model associate the knowledge-updated content with the input request during reasoning.

[0040] After generating context modulation tags, they need to be added to the vocabulary expansion tables of each language model in the target model set. The vocabulary expansion table is an internal vocabulary set used by the model to process input sequences; it typically contains the basic vocabulary and symbols used during model training. When knowledge updates bring new expressions or rules, if these tags are not added to the vocabulary expansion table, the model will not be able to recognize them during word segmentation or encoding, thus affecting subsequent modulation performance. Therefore, writing context modulation tags into the vocabulary expansion table ensures that they are recognized and processed during inference.

[0041] Input requests typically appear as natural language text and need to be converted into a sequence of input sequences for the model through word segmentation. Word segmentation methods may include byte-based sub-word segmentation or dictionary-based multi-granularity segmentation. The converted model input sequence consists of a series of encoded symbols or vectors, which can be directly processed by the language models in the target model set.

[0042] After obtaining the model input sequence of the input request, it is necessary to determine the semantic similarity score between the input request and the knowledge update content. The semantic similarity score is typically calculated using embedding representations, such as mapping the input request and the knowledge update content to the same semantic vector space, and then obtaining the score through cosine similarity or dot product operations. This score reflects the semantic relevance between the user request and the knowledge update content.

[0043] When the semantic similarity score exceeds a preset threshold, the system retrieves a contextual adjustment tag from the vocabulary expansion table of the corresponding language model in the target model set and inserts it at the beginning of the model's input sequence. This insertion operation ensures that the tag has priority in the model's sequence processing, allowing the model to incorporate knowledge update constraints in the initial attention allocation. This mechanism ensures that when the user request is highly relevant to the updated knowledge content, the model's reasoning process can be effectively guided by the updated knowledge.

[0044] Ultimately, the model input sequence after inserting context modulation markers is defined as a modified input sequence and used as input for subsequent inference.

[0045] In one implementation, the context control markers are directly extracted and symbolized from the key phrases of the knowledge update content, which is suitable for scenarios where the update content is a structured rule.

[0046] In another implementation, a hash encoding method can be used to generate a tag string for the knowledge update content, avoiding duplication with the original vocabulary set.

[0047] A dynamic threshold mechanism can also be used to dynamically adjust the semantic similarity threshold based on the importance of the knowledge update content. For example, in the fintech business, when regulatory policies involve high-risk clauses, the threshold can be lowered to ensure that related requests are more likely to trigger the insertion of context control markers; in the healthcare business, when knowledge updates involve drug contraindications, the threshold can be raised to avoid accidental triggering that could interfere with clinical decision-making.

[0048] This embodiment generates context-modulation tags based on knowledge update content and inserts these tags into the input sequence when the input request semantically matches the knowledge update content. This enables the target model set to more accurately incorporate new knowledge when processing input. This approach avoids frequent full model retraining and improves the coupling between input and knowledge updates, allowing knowledge update content to directly affect the inference process in a lightweight manner.

[0049] S50, input the input sequence with the control marker into the target model set after parameter update to obtain a model output set containing the outputs of each model; In this embodiment, the target model set updated by inputting the input sequence with contextual control markers first needs to clarify the input data format. The input sequence with contextual control markers is a sequence containing contextual control markers and the original request text. This sequence has already undergone word segmentation and encoding processing, enabling it to be directly recognized by various language models. The contextual control markers play a guiding role at the beginning of the sequence, guiding the model to prioritize semantics related to the knowledge update content during the reasoning process.

[0050] The updated target model set contains multiple language models, each with a knowledge edit vector written into its intermediate layer parameters to incorporate the latest knowledge. The input phase typically employs parallelization, simultaneously feeding the labeled input sequence to all models in the set to ensure processing efficiency and consistency. Upon receiving the input sequence, each language model generates its corresponding output based on its specific parameter weights and attention mechanism.

[0051] During processing, each language model outputs a raw output sequence, which includes the generated text and the probability distribution associated with that text. To enable comparison and fusion of the outputs, two core pieces of information need to be extracted from each model's raw output sequence: the output text content and the corresponding confidence score. The output text content is the natural language sequence generated by the language model, while the confidence score is a value obtained by normalizing the probability of the output text during inference, used to measure the model's trustworthiness of that output.

[0052] After the outputs of each language model have been extracted, the text content of each model needs to be combined with its corresponding confidence score to form a single model output, creating a structured output record. The individual model outputs from multiple models are then aggregated into a single set, known as the model output set. This set provides a consistent view across models, reflecting the responses of all updated language models to the same input request.

[0053] In one implementation, a distributed inference framework can be used to distribute the input sequence with control tags to language models on multiple computing nodes in parallel. Each node runs the inference task independently, and the results are finally integrated and output through the result collection module.

[0054] In another implementation, a pipelined parallel mechanism can be introduced. For the target model set with updated parameters, some models quickly generate an initial output sequence, and then the remaining models supplement the generation within the context of the initial results. This approach can shorten the overall response time while maintaining multi-model inference capabilities.

[0055] A hierarchical confidence extraction approach can also be used, further subdividing confidence scores into token-level local confidence and sentence-level global confidence, and then weighting and integrating them when combining individual model outputs. This makes the results in the model output set more robust and comparable.

[0056] This embodiment enables real-time knowledge retrieval across models by updating the target model set with input parameters of the input sequence marked with control and outputting the model output set. This approach not only ensures that all models consider the latest knowledge updates when processing input requests, but also provides a complete information foundation for subsequent fusion and consistency processing through the unified generation of the output set, thereby improving response consistency and knowledge utilization efficiency in a multi-model environment.

[0057] S60, The model output set is aligned and fused using the consistency fusion module to generate a fused output result; In this embodiment, the consistency fusion module aligns and fuses the model output set, requiring the initial reception of output sets from multiple language models. Each model output set contains multiple individual model outputs, each consisting of a text sequence and its corresponding confidence score. Since different language models may generate similar but not identical text content, the consistency fusion module is needed to align and fuse these outputs.

[0058] The specific process of output alignment typically involves comparing text sequences and unifying semantic levels. First, each output sequence is segmented and semantically encoded, converting it into a vector representation for semantic similarity calculation. The semantic similarity score determines the degree of closeness between the outputs of different models. If the similarity between all output sequences is higher than a preset threshold, it indicates a high degree of consistency in the model results, and the output with the highest confidence level can be selected as the fusion result.

[0059] If the similarity between output sequences is below a threshold, the fusion process begins. This fusion process weights the results based on the confidence scores of each model, giving higher-confidence results a larger weighting percentage, thus ensuring the final result integrates the judgments of different models. Weighting methods can include not only linear weighting, but also entropy weighting, hierarchical voting mechanisms, or task-specific weighting strategies.

[0060] Ultimately, the fusion output generated by the consistency fusion module is an aligned and weighted text sequence. This result can take into account the commonalities and differences of different model outputs, ensuring the consistency and reliability of results in a multi-model environment.

[0061] In one implementation, the consistency fusion module can use a semantic comparison method based on vector space to map the output sequence into a semantic vector space, determine the closeness between outputs by calculating cosine similarity, and determine whether to directly adopt or enter weighted fusion based on a threshold.

[0062] In another implementation, vocabulary-level edit distance calculation can be introduced during the alignment process to perform string-level comparison of the output text in order to identify potential deviations and give higher fusion weights to results with smaller semantic differences but different expressions during the fusion stage.

[0063] Another approach is to introduce a cross-model calibration mechanism to standardize the confidence scores output by each model, ensuring that the confidence scores of different models are within a comparable range, and then perform weighted fusion based on the standardized confidence scores. This method can reduce the bias caused by differences in confidence score distributions between models.

[0064] This embodiment uses a consistency fusion module to align and fuse the model outputs, achieving result consistency even when multiple models have different outputs, thus avoiding the impact of a single model's error on the overall result. Output alignment ensures the semantic consistency of the text results, while weighted fusion combines the advantages of multiple models, thereby improving the stability and reliability of the final result. This is particularly suitable for business scenarios with extremely high accuracy and consistency requirements, such as insurance, finance, and healthcare.

[0065] S70, when an output conflict occurs in the model output set and cannot be eliminated through fusion processing, a second inference operation is performed based on the knowledge editing vector to output the final result.

[0066] In this embodiment, when output conflicts occur in the model output set and cannot be eliminated through fusion processing, a second inference operation needs to be triggered to ensure the validity and consistency of the results. First, conflict detection needs to be performed on the model output set. Conflicts not only manifest as inconsistencies in text content but may also be semantic in nature. For example, one model might output a positive conclusion while another outputs a negative conclusion, or multiple models might give different calculation results in numerical inference. To accurately identify conflicts, each output sequence needs to be compared one by one, combining semantic similarity calculation and logical contradiction judgment to confirm whether a conflict exists.

[0067] After a conflict is detected, a judgment must be made based on preset conflict determination conditions. For example, if the similarity score is below a certain threshold, or if the differences between different outputs in key attributes exceed the tolerable range, it can be determined that the conflict cannot be eliminated through fusion. At this point, the knowledge editing vector needs to be invoked to execute a new inference process.

[0068] The knowledge edit vector is generated during the preprocessing stage and carries updated knowledge information. To function, the knowledge edit vector needs to be concatenated with the original input request to form input data containing contextual semantics and new knowledge content. This concatenation can be done by directly chaining the vectors together, or by using a mapping layer to map the two types of vectors to a unified space before combining them, ensuring semantic consistency and compatibility.

[0069] The concatenated input data is fed into a dedicated inference module. Unlike ordinary language models, this module typically includes additional decoding interfaces or weight adjustment mechanisms for knowledge edit vectors in its architecture, enabling it to better utilize newly added knowledge information. After processing by the dedicated inference module, the generated re-inference result contains both the contextual information of the original input request and the new knowledge carried by the knowledge edit vectors. Finally, this re-inference result is directly used as the output to replace conflicting model output sets, thereby ensuring the consistency and reliability of the system results.

[0070] In one implementation, conflict detection can be performed using a semantic embedding-based determination method, which maps the output sequence to a semantic vector space, calculates the distance between each output, and determines a conflict if the distance exceeds a set threshold.

[0071] In another implementation, the splicing and combination can be accomplished through a serialization operation, that is, by adding the encoded result corresponding to the knowledge editing vector before or after the encoded sequence of the original input text, so that the model can receive two types of information simultaneously during the reasoning process.

[0072] A gating mechanism can also be introduced through a dedicated inference module to dynamically adjust the weights of the knowledge editing vectors. When the conflict intensity is high, the weights of the knowledge editing vectors are increased so that they play a greater role in the final result; when the conflict is weak, the weights are kept low to avoid excessive interference with the original input.

[0073] This embodiment effectively avoids the uncertainty caused by the failure of the fusion mechanism by performing a second inference operation based on knowledge editing vectors when conflicts occur in the model output set. This method ensures that the updated knowledge is reflected in the final output and solves the problem of unresolved conflicts between different models. Through splicing and combination and processing by a dedicated inference module, the consistency and authority of the results with the latest knowledge can be ensured, thereby improving the stability and reliability of the system in multi-model deployment scenarios.

[0074] This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as fintech and healthcare. It discloses a cross-model knowledge editing and updating method, apparatus, device, and medium, comprising: receiving a knowledge update instruction and generating a knowledge editing vector; writing the knowledge editing vector into the intermediate layer parameters of a target model set to obtain a parameter-updated target model set; generating an input sequence with control tags when the input request and the knowledge update content satisfy semantic matching conditions, and inputting it into the target model set to obtain a model output set; performing output alignment and fusion through a consistency fusion module to generate a fused output result; and performing a re-inference operation based on the knowledge editing vector when output conflicts are not resolved, outputting the final result. This invention, by introducing a cross-model mapping layer and a consistency fusion mechanism in a multi-model environment, realizes the sharing and migration of knowledge editing content among multiple models, avoiding the inefficiency caused by repeatedly editing a single model in traditional methods. When there is a semantic relationship between the input request and the knowledge update content, dynamic triggering is achieved through context control tags, enabling effective utilization of knowledge updates during the inference process. Furthermore, by combining conflict detection and re-inference mechanisms, reliable final results are ensured even when model outputs are inconsistent, thereby improving the timeliness, consistency, and application reliability of knowledge updates.

[0075] In one embodiment, step S10 above includes: S101 receives the input knowledge update command through the interactive interface; S102, Perform format verification and integrity verification on the knowledge update instruction; S103, Extract structured text data from the verified knowledge update instructions; S104, the structured text data is converted into a standardized format, and the standardized structured text data is used as knowledge update content.

[0076] In this embodiment, the receiving of knowledge update instructions and the parsing and generation of knowledge update content revolve around a continuous data access and standardization link. The sources of knowledge update instructions can be interactive interface input, business service pushes, batch processing task submissions, or issuance by a regulatory platform, and their forms may include natural language text, semi-structured records, structured messages, or table fragments. To avoid semantic drift caused by differences in sources, the access layer first performs unified abstraction of the input channel, limits the encoding format and transmission container, and records the instruction identifier, source identifier, timestamp, idempotent key, and signature digest to ensure that subsequent processing is traceable and replayable. The interactive interface side includes four stages: authentication and authorization, rate limiting and backpressure, retry and deduplication, and persistent storage. Authentication and authorization constrain the access scope and modification permissions of the instruction submitter; rate limiting and backpressure absorb sudden high concurrency; retry and deduplication combined with idempotent keys ensure that the same instruction is processed only once in multiple submissions; and persistent storage stores the original message, access metadata, and verification digest on an immutable medium for auditing purposes.

[0077] Format validation revolves around the goal of "whether it can be correctly parsed," focusing on payload structure, field types, value ranges, and character security. Payload structure is validated using a schema description file to ensure that hierarchical relationships, field naming, and required paragraphs meet expectations. Field types are strongly constrained across numeric, text, time, and Boolean dimensions, with illegal types being rejected outright. Value range validation limits enumerated items and interval boundaries to prevent out-of-bounds input. Character security checks handle inconsistent encodings, control characters, injected symbols, and excessively long fragments, performing escaping and cleanup when necessary. When cross-language or multi-region input is involved, encoding is unified to the same character set, time is unified to Coordinated Universal Time (UTC), and currency symbols are uniformly mapped to standard currency codes, ensuring that different sources are processed within the same semantic coordinate system.

[0078] Integrity verification focuses on the sufficiency and consistency of information. Required field validation ensures that key information is not missing, such as clause number, revision subject, effective date, scope of application, and source link. Composite constraint validation handles cross-field dependencies; for example, the effective date must be later than the publication date, and revocation conditions and substitution relationships cannot be simultaneously empty. Cross-resource association validation verifies foreign key references such as insurance type codes, drug codes, and institution codes through dictionary tables or master data services to avoid isolated updates. Version conflict identification compares the sequence relationship between existing versions and submitted versions; if reverse submissions or concurrent overwriting are found, they are immediately sent to the manual review channel. Integrity verification outputs a structured verification report, marking the location of defects, severity level, and handling suggestions, and sending back non-automatically repairable issues to the interactive interface for rectification prompts.

[0079] The goal of extracting structured text data is to obtain machine-consumable key points and parameters from qualified instructions. The parser consists of lexical, syntactic, and semantic extraction units. The lexical unit performs segmentation, word segmentation, and regular expression slicing to locate candidate phrases; the syntactic unit identifies main predicates and constraint clauses to construct dependency relationships; the semantic extraction unit maps phrases to unified semantic tags, such as "revision topic," "limit type," "object scope," "applicable conditions," "exceptions," "effective date," and "termination conditions." For table and attachment content, a cell-field mapping strategy and a title alignment strategy are used to extract key-value pairs, and cross-page tables are reorganized to maintain row and column consistency. To improve robustness, a "dual alignment" mechanism of template alignment and entity dictionary alignment is introduced: template alignment is used to capture common sentence variations, and entity dictionary alignment is used to resolve synonyms and abbreviations, such as mapping "third-party liability insurance" and "third-party liability" to the same entity, and mapping drug brand names and generic names to the same drug identifier.

[0080] When converting structured data to a standardized format, a unified semantic model and field mapping table are introduced. The unified semantic model defines core entities, relationships, and attribute paradigms, covering clauses, rule units, applicable objects, threshold conditions, effective periods, and version chains. The field mapping table maintains the mapping from source fields to target fields, unit conversion strategies, and default strategies. Unit conversion covers dimensions such as amount, interest rate, dosage, and duration, ensuring values ​​are on the same scale. Terminology alignment uses a terminology library to map domain expressions to a unified terminology, avoiding repetitive standardization in subsequent steps. Time standardization unifies to a single time zone and format, and breaks down into fine-grained fields such as effective date, expiration date, observation period, and grace period. Currency and region are represented by standard codes, with decimal precision and rounding strategies. Conflict resolution is adjudicated based on priority, source credibility, and time freshness, retaining multiple versions and marking compatibility relationships when necessary. The output is transformed into an object for downstream processing, which includes resource identifier, version number, revision summary, structured parameter set, scope of application, effective and expiration time, substitution relationship, traceability information, digital signature and verification digest. All fields have clear data types and constraints, which facilitates direct consumption in subsequent encoding and writing processes.

[0081] The entire chain is equipped with closed-loop control and fault self-healing on the quality and operation and maintenance sides. During the access and verification phases, full processing logs and verification reports are generated, supporting retrieval by resource identifier and version number. Anomalies are reported simultaneously locally and on a centralized alarm platform, triggering automatic retries or manual review. Long transactions are ensured to produce only one standardized version per commit through transaction logs and breakpoint markers. During peak concurrency periods, buffer queues and backpressure strategies maintain stable throughput. In multi-tenant environments, storage and access are isolated by tenant identifiers to avoid cross-contamination. Security is enhanced by a combination of transmission encryption, storage encryption, and signature verification to resist tampering and replay attacks. The final standardized data is read-only and enters downstream channels. Any subsequent changes are submitted as a new version through the same chain rather than overwriting existing data, ensuring traceability and rollback capability of the update history.

[0082] This embodiment transforms diverse and inconsistent knowledge update information into consumable objects within a single semantic coordinate system through unified access, rigorous verification, precise extraction, and semantic standardization in the receiving and parsing link. Errors and omissions are blocked before entering the downstream, ambiguities are resolved through terminology and unit alignment, and duplication and concurrency are suppressed through idempotency and version control. This provides a clean, complete, and traceable data baseline for subsequent encoding and cross-model writing, shortens the update transmission path, reduces human intervention and backtracking costs, and improves update timeliness, consistency, and maintainability.

[0083] In one embodiment, step S20 above includes: S201, The knowledge update content is input into the encoder neural network based on the Transformer architecture in the knowledge editing module; S202, the semantic representation of the knowledge update content is determined through the multi-head attention mechanism of the encoder neural network; S203, The knowledge editing module uses a feedforward neural network to perform a nonlinear transformation on the semantic representation to generate a high-dimensional feature vector; S204, The high-dimensional feature vector is subjected to layer normalization processing through the normalization layer of the knowledge editing module to obtain the normalized high-dimensional feature vector; S205, the normalized high-dimensional feature vector is mapped to a low-dimensional space through the dimensionality reduction layer of the knowledge editing module to generate an initial low-dimensional vector; S206, The initial low-dimensional vector is normalized by the standardization unit of the knowledge editing module to obtain a standardized vector; S207, the standardized vector is stored as a knowledge editing vector in the vector repository of the knowledge editing module.

[0084] In this embodiment, before inputting the knowledge update content into the Transformer-based encoder neural network in the knowledge editing module, the standardized fields and text payloads need to be organized into a sequence that can be received by the encoder. The sequence consists of domain identifier fragments, entity identifier fragments, parameter key value fragments, and context fragments, which are concatenated in a fixed order and mapped to the index of the vocabulary or sub-vocabulary. For very long payloads, a sliding window and overlapping concatenation are used, and connection markers are introduced at the window boundaries to reduce semantic breaks. To suppress interference from irrelevant fields, weight masks are assigned to the sequence positions according to the importance level of the fields, and normalized placeholders are generated for time, effective interval, and quota-type fields, so that subsequent attention can focus on the key points of knowledge update. In the embedding stage, the index is converted into a dense vector, and position encoding and domain encoding are superimposed. Position encoding ensures the effective injection of sequence order information, and domain encoding pulls expressions from different sources such as finance and medicine into the same semantic coordinate system. Batch input uses fixed-length alignment, and short sequences are processed by padding vectors and padding masks to avoid invalid positions affecting attention acquisition.

[0085] When determining the semantic representation of knowledge updates using a multi-head attention mechanism in an encoder neural network, the query vector, key vector, and value vector are matched in parallel across the multi-head space. To reduce noise sources, a mask is used in the attention matrix to block padding positions and decorative phrases marked as non-critical. Domain-specific encoding is injected before the linear transformation of each head, ensuring that different heads focus on different types of constraint relationships. The financial domain often contains matching and exception entries, while the medical domain has taboo and adaptation pairs. Attention allocation aggregates the semantic center near the revision subject and constraint predicate by matching these paired structures. The aggregation of the entire representation employs two parallel paths: one uses the initial label vector as a global summary, and the other performs weighted pooling on all position vectors. The results from both paths are concatenated along the channel dimension, ensuring that both the overall theme and key details are preserved.

[0086] A feedforward neural network using a knowledge editing module performs nonlinear transformations on the semantic representation to improve expressive capacity and separability. The feedforward network employs a cascaded structure of linear transformations and nonlinear activations with extended dimensions. The first linear unit stretches the semantic channels to accommodate combined features, while the activation units compress unimportant dimensions and introduce nonlinear boundaries. The second linear unit restores the target dimension. To avoid overfitting and numerical jitter, random deactivation and residual channels are inserted between the two linear transformations. The residual channels add the input to the transformed output, ensuring gradient stability during deep propagation.

[0087] The knowledge editing module uses a normalization layer to perform layer-level normalization on high-dimensional feature vectors to eliminate distribution biases caused by inputs from different batches and sources. Layer-level normalization estimates the mean and variance in the channel dimension and linearly recalibrates the normalization results using learnable scaling parameters. For serialized representations, this can be performed independently in the position dimension to ensure independent and stable channel statistics at each position; for global summary representations, normalization can be performed once in the overall vector dimension to ensure consistent global scale. Placing normalization after residual summation helps suppress cumulative drift, while placing it before activation helps stabilize the activation interval; either placement method can be chosen to maintain consistency depending on the deployment environment.

[0088] The dimensionality reduction layer of the knowledge editing module maps normalized high-dimensional feature vectors to a low-dimensional space, aiming to reduce storage and subsequent mapping write overhead without compromising key semantics. The dimensionality reduction layer can use linear projection to achieve interpretable channel compression or employ a bottleneck structure to filter information in intermediate dimensions. A reservation list is set for common semantic channels that need to be shared across models; for example, channels related to revised subjects, applicable objects, threshold types, and effective ranges are excluded from reduction. Variable reduction ratios are set for domain-specific channels to balance generalization and accuracy. To reduce scale differences between different batches, weighted regularization and output pruning are added to both ends of the dimensionality reduction projection matrix to prevent individual components from having excessively large amplitudes.

[0089] The initial low-dimensional vectors are normalized using the standardization unit of the knowledge editing module to obtain standardized vectors with comparable scale and direction. The standardization unit first performs centering, aligning the vector centroids to zero, and then applies norm constraints to ensure consistent length, thus making vectors from different sources and versions comparable within the metric space. Outliers are truncated with thresholds and smoothly redistributed to prevent a few extreme components from dominating subsequent writing. When it is necessary to retain important channel differences, channel reweighting can be introduced after centering, followed by length normalization, so that relative importance is reflected on the unit sphere.

[0090] Standardized vectors are stored as knowledge editing vectors in the vector repository of the knowledge editing module, requiring reliable writing, auditable traceability, and high-performance retrieval. The stored entry index includes source identifier, domain identifier, version number, revision summary, effective period, applicable scope, and hash digest, supporting timeline queries based on version and effective period. To ensure concurrent consistency, writes employ append semantics; any changes are implemented by generating a new version, while old versions are retained read-only. To adapt to cross-model sharing and fast loading, the storage layer maintains both dense vector indexes and sparse tag indexes. Dense indexes are used for nearest neighbor retrieval and batch loading, while sparse indexes are used for precise filtering and compliance auditing. Key management and signature verification ensure entries are not tampered with, while write logs and rollback flags guarantee rapid recovery to the previous consistent state in failure scenarios.

[0091] This embodiment transforms heterogeneous knowledge update content into scale-consistent, semantically stable, and cross-model-shareable knowledge edit vectors through encoding. Attention convergence and feedforward transformation reinforce key elements such as the revised subject and constraint predicate. Normalization and dimensionality reduction control numerical scale and storage overhead, while standardization and versioned storage provide comparability and traceability. This forms a unified carrier for multi-model writing, reducing adaptation costs caused by format and domain differences, shortening update propagation paths, and reducing reliance on retraining. While ensuring sufficient representation, it improves vector loading and mapping efficiency, providing a stable input foundation for subsequent cross-model parameter writing and context triggering.

[0092] In one embodiment, step S30 above includes: S301, input the knowledge editing vector into the parameter mapping network of the cross-model mapping layer; S302, Analyze the parameter structure of each language model in the target model set through the parameter mapping network; S303, Based on the parameter structure analysis results, the knowledge editing vector is converted into parameter adjustment values ​​corresponding to each language model; S304, write the adjustment values ​​of each parameter into the intermediate layer weight matrix of the corresponding language model, and perform parameter fusion processing on the updated part of the intermediate layer weight matrix of each language model. S305 combines all language models that have completed parameter fusion processing into a set of target models with updated parameters.

[0093] In this embodiment, when inputting a knowledge editing vector into the cross-model mapping layer, the vector-to-mapping input channel must first be shaped and validated. The shaping process encapsulates the vector and its metadata into a mapping request unit. The metadata includes a source identifier, version number, effective range, applicable scope, and write policy flag, used for routing and permission verification in subsequent paths. The validation process checks the vector norm, numerical range, and missing components to avoid amplifying the risk of writing abnormal values. After receiving the request, the cross-model mapping layer creates a one-time transaction context for each write, binding the request unit to a list of target model sets to ensure consistency and rollbackability of cross-model updates.

[0094] When knowledge edit vectors are fed into the parameter mapping network, the network reads the target model set from the transaction context and loads a structural summary for each model in the set. The structural summary consists of elements such as intermediate layer type, channel size, weight tensor shape, activation form, residual and normalized arrangement, and cross-layer connectivity, and can originate from the model's self-description file or registry entries. Based on the structural summary, the mapping network constructs a mapping graph, which describes the path aligning the unified vector channels to the intermediate layer channels of each model, including common channels that need to be retained, domain channels that need to be remapped, and sensitive channels that are prohibited from being written to. For models with multi-branch intermediate layers, the mapping graph explicitly labels the branch weights and writing order to avoid race conditions in concurrent branches.

[0095] Parametric structure analysis is performed after the mapping graph is generated to determine the specific writable locations and constraints. The analysis process traverses each intermediate layer weight matrix, extracting information such as row and column dimensions, parameter blocks, sparse masks, shared weight ranges, and quantization scales, and matches this information with the write strategy. The write strategy specifies the scope and intensity of the update, such as applying only to attention-related weights, only to the feedforward projection layer, or covering both types of weights simultaneously; the intensity is represented by a coefficient or threshold to control the adjustment magnitude from crossing safety boundaries. If the structure analysis identifies unwritable blocks due to quantization or pruning, the mapping network removes these blocks from the writable list and marks the reason in the transaction record.

[0096] When converting knowledge edit vectors into parameter adjustment values ​​corresponding to each model, the mapping network performs channel rearrangement and scale adaptation based on the mapping graph. Channel rearrangement maps unified channels to model channels, ensuring a one-to-one correspondence between common semantic channels, while domain channels are remapped or linearly combined according to the mapping table. Scale adaptation compresses the numerical scale of the edit vectors to a range compatible with the target weights, avoiding excessive perturbation of intermediate layer behavior. To reduce noise, soft gating can be applied to the vectors before conversion. The gating weights are derived from the channel importance assessment during the structural analysis phase, with low-importance channels suppressed or set to zero. The conversion result, i.e., the parameter adjustment values, is organized into blocks of the same shape as the target weights according to the writing position, facilitating subsequent block-by-block writing.

[0097] When writing the adjusted parameter values ​​into the intermediate layer weight matrix of the corresponding language model, block writing and fine-grained fusion are performed. Block writing is performed according to the natural or logical blocks of the weight matrix. Before writing each block, the original weight snapshot is read, the adjustment estimate is calculated, and it is determined whether the protection threshold is triggered. If the condition is met, the fusion process begins, using either residual weighting or gated injection. Residual weighting outputs a linear combination of the original weights and parameter adjustment values, with the combination coefficients determined by the writing strategy and channel importance. Gated injection first calculates the consistency score between the adjustment values ​​and the original weights, and then opens and closes the gates according to the score, allowing components with high consistency to receive a larger injection ratio. The consistency score combines directional consistency and amplitude rationality to avoid reverse perturbation. After writing is completed, an update part is formed, which is marked with a version number and writing factor for backtracking and auditing.

[0098] The goal of parameter fusion processing for the updated part is to smoothly integrate the newly introduced knowledge with the existing representation. Fusion processing is performed at two levels: intra-layer and inter-layer. Intra-layer fusion balances multiple blocks within the same intermediate layer, corrects channel offsets caused by independent block writes, and adjusts the layer normalization scale through lightweight recalibration. Inter-layer fusion focuses on cross-layer propagation, matching the output statistics of the affected layer with the input statistics of the downstream layer, and fine-tuning the trainable biases or calibration parameters of the downstream layer when necessary to mitigate the cascading biases caused by local writes. After fusion, a fast consistency check is performed on the affected paths, checking metrics such as output distribution drift magnitude, residual channel saturation, and activation sparsity to ensure that updates are implemented within a controllable range.

[0099] When combining the language models that have undergone fusion processing into a set of target models with updated parameters, the cross-model mapping layer generates a set list based on the transaction context, recording the version increment, write range, update summary, and consistency check results for each model. The set list is written to the registry and exposes a single update load handle, allowing subsequent inference phases to load the entire model set with a single click, avoiding inconsistencies within the cluster due to partial loading. To improve rollback availability, a rollback record corresponding to this write is also generated, saving the index information of the original weight snapshot. If an anomaly is detected during subsequent checks, the system can quickly restore the system to its pre-write state using the handle.

[0100] This embodiment translates a unified knowledge edit vector into executable parameter adjustment values ​​in a multi-model architecture via a parameter mapping network. These values ​​are then written into the intermediate layer weight matrices of each model in a controlled, rollback-proof, and auditable manner. Through intra-layer and inter-layer fusion of the updated portions, the new knowledge is aligned with the existing representation distribution, ultimately outputting the target model set with updated parameters. This approach synchronously diffuses one-time knowledge updates across the multi-model environment, reducing the risk of repeated modifications and inconsistencies. Channel rearrangement, scale adaptation, and gating injection reduce disturbance intensity and direction errors. Versioned transactions and rollback records construct a traceable and recoverable update mechanism, providing a stable and homogeneous model baseline for subsequent unified inference and consistent fusion.

[0101] In one embodiment, step S40 above includes: S401, Generate a tag string associated with the knowledge update content as a context control tag based on the knowledge update content; S402, add the context control tag to the vocabulary expansion table of each language model in the target model set; S403, converts the input request into a model input sequence through word segmentation; S404, determine the semantic similarity score between the input request and the knowledge update content; S405, when the semantic similarity score exceeds a preset threshold, obtain the context control marker from the vocabulary expansion table of the corresponding language model in the target model set, and insert the context control marker into the starting position of the model input sequence; S406, the model input sequence after inserting the context control marker is used as the input sequence with control marker.

[0102] In this embodiment, when generating context control tags based on knowledge update content, the standardized update entries are first subjected to semantic point extraction and tag construction. Semantic point extraction generates stable feature summaries from the revised subject, applicable objects, constraint predicates, exception conditions, and effective intervals. Tag construction maps the summaries to tag strings that can be stably recognized by the word segmenter, using a combination of prefixes and invisible delimiters that do not conflict with existing vocabularies to ensure uniqueness and distinguishability across different models. To ensure cross-language and cross-domain usability, the tag strings simultaneously retain semantic hash fragments and domain-specific fragments. The hash fragments are used to resist name collisions, and the domain-specific fragments are used to infer the triggering context in financial or medical scenarios. After generation, the tag strings undergo orthogonality verification, including upper bound checks on similarity to existing vocabulary entries, encoding reversibility checks, and regular expression whitelist checks, to avoid ambiguity with common words or control characters.

[0103] When adding context-modification tags to the vocabulary expansion tables of each language model in the target model set, vocabulary supplementation and embedding initialization must be completed without compromising the stability of existing word segmentation statistics and embedding spaces. Vocabulary supplementation involves registering new entries in the word segmentation configuration of each model and assigning consecutive lexical numbers. If the model uses sub-word decomposition, to avoid further splitting, the tag string is registered as a priority merging unit in the merging rules and placed as a high-priority merging item in the BPE or Unigram table. If the model uses byte-level encoding, start and end boundary symbols are registered for the tag string to ensure it is encoded as an independent unit. Embedding initialization employs a two-pronged strategy executed simultaneously: first, an initialization vector is obtained by weighted summation of the embedding vectors of several lexicals most similar to the updated content; second, candidate vectors are obtained by linear projection from the knowledge editing vector. The two results are fused into the final initialization vector after scale alignment and orthogonal decorrelation, and written into the embedding matrix of each model. To maintain rollback capability and consistency, the vocabulary version, merging rule checksum, and embedding entry hash are recorded before and after expansion, and a versioning strategy of adding only entries and not deleting entries is set.

[0104] When converting input requests into model input sequences through word segmentation, the word segmenter segments the natural language requests according to the vocabulary and merging rules of each model, and performs conflict avoidance on segments adjacent to the tagged string. Conflict avoidance checks for potential prefix or suffix overlaps through a context window, and if necessary, injects zero-width delimiter mapping bits before segmentation to ensure that subsequent insertions do not change the stability of the original segmentation. During the encoding stage, a token index sequence and a position index sequence are generated. For excessively long requests, a sliding window and overlapping concatenation are used. To facilitate subsequent insertion operations, placeholder indexes are reserved at the beginning of the sequence and marked as writable positions in the attention mask. For requests containing sensitive entities, a desensitization mapping table is used to complete the replacement before word segmentation and record the reverse mapping so that reversible restoration can be performed after inference.

[0105] When determining the semantic similarity score between the input request and the knowledge update content, the request and the updated content need to be aligned and represented within a unified semantic space. The request side generates sentence vectors using a lightweight embedding encoder, while the update content side loads the knowledge edit vector or its summary vector bound to that version. Similarity calculation uses a normalized cosine metric supplemented with temperature scaling to avoid gradient vanishing due to overly dense distribution in high-dimensional space. To reduce the bias caused by cross-domain word literal differences, a domain correction term is introduced, centering and rescaling the vector distributions for finance and healthcare, respectively. To reduce the impact of noisy input, a stabilizer below the confidence threshold is introduced to converge the similarity scores of short texts or noisy texts to a conservative range. The final score is compared with a preset threshold, which can be dynamically adjusted according to the importance level of the updated item, the weight of the publishing institution, and the proximity of the effective time. These dynamic rules are recorded during version release.

[0106] When the semantic similarity score exceeds a preset threshold, contextual regulation markers are retrieved from the vocabulary expansion table of the corresponding language model in the target model set and inserted at the beginning of the model input sequence. The retrieval process directly locates the lexicon number using the vocabulary index and simultaneously loads the initial embedding verification hash to prevent accidental replacement of the vocabulary during runtime. During insertion, the lexicon index sequence and position index sequence are updated, and after rearrangement, relative displacement correction is performed on the position encoding to ensure that subsequent attention calculations are not offset by head insertion. The attention mask simultaneously opens the first and second attention channels, enabling the model to perceive the regulation signal in the first round of attention convergence. If the requested length is close to the maximum length limit, tail truncation is performed and the truncated summary index is saved for necessary completion and reconstruction after inference. After insertion, the updated sequence, position index, and mask are collectively defined as the input sequence with regulation markers and bound to the corresponding vocabulary version and verification information to align to the same marker definition during parallel inference of multiple models.

[0107] This embodiment constructs a unified context-controlled tag based on the semantic key points of the updated entries, and performs consistent vocabulary supplementation and embedding initialization in the target model set, ensuring that the tags have the same recognizability and semantic orientation across different models. It calculates the similarity between the request and the updated content using a unified semantic space and employs an auditable threshold strategy to trigger insertion, ensuring that the tag only takes effect on relevant requests. Synchronous correction of position and mask maintains the stability of sequence encoding and attention distribution. This forms a closed loop from tag generation, vocabulary expansion, request encoding to conditional insertion, allowing new knowledge to be explicitly guided at the inference entry point in a lightweight manner, reducing disturbances to the original vocabulary and embedding distribution, lowering the probability of false triggering and omissions, and providing aligned trigger signals and traceable version information for subsequent multi-model parallel inference and result consistency fusion.

[0108] In one embodiment, step S50 above includes: S501, the input sequence with control markers is input in parallel into each language model in the target model set with updated parameters; S502, each language model processes the input sequence with control markers to generate the corresponding original output sequence; S503 extracts the output text content and corresponding confidence scores from the original output sequence of each language model; S504 combines the output text content and confidence score of each language model into a single model output; S505 aggregates the outputs of individual models from all language models, forming a model output set.

[0109] In this embodiment, the target model set updated with the input parameters of the input sequence with control tags needs to establish a stable and consistent input channel and a parallel distribution mechanism. The input sequence with control tags is loaded into the input buffer with a uniform tensor layout, which includes three types of information: lexical index, position index, and attention mask. The context control tags retained at the head position remain unchanged so that different language models receive the same guiding signal during the first round of attention aggregation. During the distribution phase, the parallel scheduler copies the input tensor according to the target model list, binds the vocabulary version and tag verification summary one by one, and then pushes it to the inference entry of each language model. To avoid temporal drift, a consistent decoding temperature, sampling threshold, and stopping criterion are used within the same batch, and the random source and bundle width settings are fixed to form a reproducible generation trajectory. For requests close to the length limit, the tail truncation strategy and truncation summary index are transmitted with the batch for easy subsequent restoration.

[0110] Each language model processes the input sequence with regulatory markers and generates the original output sequence. The original output sequence consists of text fragments and positional probability trajectories. The text fragments originate from the lexical decoding process, and the probability trajectories are derived from the distribution snapshots given at each generation step. To mitigate the bias caused by differences in vocabulary between different models, the inference entry point loads a vocabulary version handle and embedding calibration vector consistent with the update batch. The mapping layer unifies the amplitude and direction of the embedding of regulatory markers on the encoding side, ensuring that the intermediate layers exhibit similar response patterns across different architectures. For requests with mixed cross-domain terminology, a domain mask is introduced in the attention graph to suppress noise fragments and maintain the proportion of lexical units directly related to the knowledge update content in the attention focus.

[0111] Output text content and confidence scores are extracted from the original output sequences of each language model using a text-probabilistic dual-channel extraction process. The text channel traces back along the generation path to obtain continuous word sequences, and character set and punctuation normalization is performed for cross-model alignment. The probabilistic channel derives sentence-level confidence scores from position-level distribution snapshots. Common practices include log-likelihood cumulative normalization, inter-frame entropy value reduction, and stabilization after length penalty. Multiple indicators are scaled and then synthesized into a single score. The extraction results are linked one-to-one with language model identifiers, vocabulary versions, and regulatory tag hashes, forming a traceable set of meta-information.

[0112] When combining the output text content and confidence scores into a single model output, a minimal complete record structure is introduced. This record structure includes the text body, sentence-level confidence scores, a truncated summary index, a domain coverage vector, a regulatory marker participation index, and a generation strategy summary. The text body retains the original vocabulary without semantic rewriting; the domain coverage vector reflects the distribution of semantic units related to the knowledge update content within the text body, and the regulatory marker participation index reflects the effective influence range of the regulatory markers in the generation path. The generation strategy summary records temperature, bundle width, stopping criteria, and random source fingerprints for subsequent reproduction.

[0113] To aggregate the outputs of all language models into a model output set, deduplication, alignment, and indexing are required. The deduplication stage uses text hashing and nearest-neighbor semantic comparison to eliminate completely identical or highly overlapping outputs. The alignment stage establishes a cross-model alignment table at the semantic unit granularity, annotating the coverage relationships of key points such as object range, threshold type, effective interval, and exception entries. The indexing stage constructs a two-level retrieval structure: the top level is organized by language model identifier and version number, and the bottom level is organized by semantic unit and domain coverage. The final model output set is archived with batch number and update timestamp, along with a control tag hash and vocabulary version snapshot, ensuring that subsequent consistent fusion and re-inference can be reused in the same state.

[0114] This embodiment uses a controlled input sequence to enter the inference channel in parallel within the updated target model set. The tags, vocabulary, and decoding strategies are aligned across models. The original output sequence is exported simultaneously along both text and probability paths. Individual model outputs are stored on disk with a complete record structure, and cross-model alignment and retrieval indexes are established. This results in a stable, reproducible, and traceable multi-model response view: parallel distribution improves throughput and timeliness; vocabulary and embedding alignment reduces cross-model drift; unified confidence score scaling facilitates subsequent weighted processing; semantic unit-level alignment provides fine-grained support for consistent fusion; and batch and version metadata binding provides a basis for compliance auditing and issue backtracking.

[0115] In one embodiment, step S60 above includes: S601, The semantic similarity score between individual model outputs is determined by comparing the output text content of each individual model output in the model output set through the consistency fusion module. S602, when the semantic similarity score of all individual model outputs is greater than or equal to the preset consistency threshold, select the single model output with the highest confidence as the fusion output result; S603, if the semantic similarity score output by any single model is less than the preset consistency threshold, weighted fusion is performed based on the confidence scores output by each single model, and the weighted fusion result is used as the fusion output result.

[0116] In this embodiment, after receiving the model output set, the consistency fusion module first establishes a unified alignment space. The output text content enters the standardization pipeline, completing the character set unification, punctuation style standardization, capitalization and whitespace folding, and unit mapping for numbers and time expressions, while retaining the bidirectional index from the original text to the standardized text. Subsequently, semantic unit extraction is performed, decomposing the text into key fragments such as object scope, constraints, exceptions, effective intervals, and quantization thresholds, and recording the start and end positions and weight percentages of each fragment to form a comparable structured representation. To reduce the bias caused by differences in decoding between different models, length normalization and stop fragment masking are introduced. Length normalization eliminates the inherent advantage of long and short texts in similarity, and stop fragment masking removes the interference of polite or formatted fragments on the comparison results.

[0117] Semantic similarity scores are determined within a unified representation space. Each individual model outputs a two-level representation: a whole-sentence vector and a semantic unit vector. These two levels of representations undergo domain correction and scale recalibration to ensure comparability between financial and medical corpora within the same metric coordinate system. Similarity matrices are calculated at both the whole-sentence and unit levels, and then weighted by adjustable coefficients to synthesize corresponding semantic similarity scores. To improve robustness, a confidence lower bound stabilizer is introduced to converge the scores of extremely short or noisy texts to a conservative range, preventing occasional high similarity from leading to misclassification.

[0118] Consistency thresholds are version-managed, written to the fusion log with each batch, and dynamically adjusted. A dynamic strategy integrates source weights, update levels, proximity of effective times, and business sensitivity levels to generate an effective threshold range for the current batch. Semantic similarity scores are compared with the thresholds to form two paths. Path one, for consistency scenarios, involves performing cross-model calibration on confidence scores to eliminate scale differences between sources when all individual model outputs reach or exceed the threshold. Then, the single model output corresponding to the highest calibrated score is selected as the fusion output. If ties exist, a secondary criterion is used to prioritize text with more comprehensive semantic units and a more balanced positional distribution, maintaining the original wording and structure without semantic rewriting, only performing format-level unification.

[0119] Path Two addresses scenarios with discrepancies by incorporating weighted fusion. Weighted fusion establishes a segment-level weight field, with the primary weight derived from the calibrated confidence score, and secondary weights combining semantic unit coverage and source diversity. An alignment table guides unit-by-unit aggregation. Fully aligned units are directly synthesized linearly into candidate expressions according to the weight field; partially aligned units or those with significant wording differences are first merged into synonymous expressions within their nearest neighbor set, then representative candidates are selected according to the weight field; missing units are filled by extracting the smallest sufficient fragment from the text containing the unit through contextual constraints. After the unit sequence is completed, it is concatenated into a complete result according to the alignment order and subjected to two checks. The factual consistency check verifies whether key points such as object scope, threshold class, and time interval remain consistent across sources after concatenation; the language coherence check uses lightweight constraints to perform smoothing at the connection points, eliminating concatenation traces. If a check fails, it reverts to the candidate layer to fine-tune weights or replaces candidates until the consistency threshold is met or an avoidance strategy is triggered. The final weighted fusion result, along with the source composition, weight summary, threshold snapshot, and similarity matrix fingerprint, is recorded to form a traceable product.

[0120] A global verification is performed before the fusion output is returned. The verification covers the semantic unit list, cross-model calibration summary, threshold and weight parameters, source decomposition vectors, and log fingerprints. After successful verification, the results are written to a versioned cache and a read-only handle is exposed for subsequent auditing or reproduction. The entire process maintains a reversible mapping from the fusion output to the outputs of each individual model, allowing any fragment to be traced back to its source text location and corresponding confidence score.

[0121] This embodiment achieves comparability of different model texts under the same metric through a unified alignment space and two-level representation. Version consistency thresholds ensure rapid decision-making in consistent scenarios, cross-model calibration avoids bias caused by inconsistent confidence scales, and weighted fusion aggregates information with fragment-level weights and maintains consistency of key points in divergent scenarios. A dual-verification mechanism constrains output stability in terms of both facts and language. Thus, in a multi-model deployment environment, it obtains semantically consistent, source-transparent, and re-verifiable fusion output results, reducing the risk of single-source bias, shortening decision-making time, and providing a high-quality, traceable input foundation for subsequent conflict rollback and re-inference.

[0122] In one embodiment, a cross-model knowledge editing and updating apparatus is provided, which corresponds one-to-one with the cross-model knowledge editing and updating method described in the above embodiments. (Refer to...) Figure 3 , Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the cross-model knowledge editing and updating device of the present invention. The modules include an instruction parsing module 10, a knowledge editing module 20, a cross-model mapping module 30, a context control module 40, a model reasoning module 50, a consistency fusion module 60, and a conflict handling module 70. Detailed descriptions of each functional module are as follows: Instruction parsing module 10 is used to receive knowledge update instructions and parse the knowledge update content from the knowledge update instructions; The knowledge editing module 20 is used to encode the knowledge update content and generate a knowledge editing vector. The cross-model mapping module 30 is used to input the knowledge editing vector into the cross-model mapping layer, and write the knowledge editing vector into the intermediate layer parameters of each model in the target model set through the cross-model mapping layer to obtain the target model set with updated parameters. The context control module 40 is used to generate context control tags based on the knowledge update content, convert the input request into a model input sequence, and insert the context control tags into the model input sequence when the input request and the knowledge update content meet the semantic matching conditions, so as to obtain an input sequence with control tags. The model inference module 50 is used to input the input sequence with the control label into the target model set after parameter update to obtain a model output set containing the outputs of each model. The consistency fusion module 60 is used to perform output alignment and fusion processing on the model output set to generate a fused output result. The conflict handling module 70 is used to perform a second reasoning operation based on the knowledge editing vector and output the final result when an output conflict occurs in the model output set and cannot be eliminated by fusion processing.

[0123] For specific limitations regarding the cross-model knowledge editing and updating device, please refer to the aforementioned limitations on the cross-model knowledge editing and updating method, which will not be repeated here. Each module in the aforementioned cross-model knowledge editing and updating device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0124] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When executed by the processor, the computer program implements the functions or steps of a cross-model knowledge editing and updating method on the server side.

[0125] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements client-side functions or steps of a cross-model knowledge editing and updating method.

[0126] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: Receive a knowledge update instruction and parse the knowledge update content from the knowledge update instruction; The knowledge update content is encoded using the knowledge editing module to generate a knowledge editing vector; The knowledge editing vector is input into the cross-model mapping layer, and the knowledge editing vector is written into the intermediate layer parameters of each model in the target model set through the cross-model mapping layer to obtain the target model set with updated parameters. Context control tags are generated based on the knowledge update content. Input requests are converted into model input sequences. When the input requests and knowledge update content meet the semantic matching conditions, the context control tags are inserted into the model input sequences to obtain input sequences with control tags. The input sequence with the control label is input into the target model set after the parameter update to obtain a model output set containing the outputs of each model; The consistency fusion module performs output alignment and fusion processing on the model output set to generate a fused output result. When output conflicts occur in the model output set and cannot be eliminated through fusion processing, a second inference operation is performed based on the knowledge editing vector to output the final result.

[0127] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor: Receive a knowledge update instruction and parse the knowledge update content from the knowledge update instruction; The knowledge update content is encoded using the knowledge editing module to generate a knowledge editing vector; The knowledge editing vector is input into the cross-model mapping layer, and the knowledge editing vector is written into the intermediate layer parameters of each model in the target model set through the cross-model mapping layer to obtain the target model set with updated parameters. Context control tags are generated based on the knowledge update content. Input requests are converted into model input sequences. When the input requests and knowledge update content meet the semantic matching conditions, the context control tags are inserted into the model input sequences to obtain input sequences with control tags. The input sequence with the control label is input into the target model set after the parameter update to obtain a model output set containing the outputs of each model; The consistency fusion module performs output alignment and fusion processing on the model output set to generate a fused output result. When output conflicts occur in the model output set and cannot be eliminated through fusion processing, a second inference operation is performed based on the knowledge editing vector to output the final result.

[0128] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0129] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0130] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0131] It should be noted that if any software tools or components not belonging to this company appear in the embodiments of this application, they are merely illustrative examples and do not represent actual use. The embodiments described above are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

[0132] The user personal information involved in this application embodiment is all authorized (knowing and consenting) by the relevant parties or fully authorized by all parties, and the executing entity can obtain it through various open, legal and compliant means. The collection, storage, use, processing, transmission, provision and disclosure of the information, data and signals involved all comply with the relevant laws and regulations of the relevant countries and regions, and do not violate public order and good morals.

Claims

1. A cross-model knowledge editing and updating method, characterized in that, Includes the following steps: Receive a knowledge update instruction and parse the knowledge update content from the knowledge update instruction; The knowledge update content is encoded using the knowledge editing module to generate a knowledge editing vector; The knowledge editing vector is input into the cross-model mapping layer, and the knowledge editing vector is written into the intermediate layer parameters of each model in the target model set through the cross-model mapping layer to obtain the target model set with updated parameters. Context control tags are generated based on the knowledge update content. Input requests are converted into model input sequences. When the input requests and knowledge update content meet the semantic matching conditions, the context control tags are inserted into the model input sequences to obtain input sequences with control tags. The input sequence with the control label is input into the target model set after the parameter update to obtain a model output set containing the outputs of each model; The consistency fusion module performs output alignment and fusion processing on the model output set to generate a fused output result. When output conflicts occur in the model output set and cannot be eliminated through fusion processing, a second inference operation is performed based on the knowledge editing vector to output the final result.

2. The cross-model knowledge editing and updating method as described in claim 1, characterized in that, Receive a knowledge update instruction and parse the knowledge update content from the knowledge update instruction, including: Receive knowledge update instructions through the interactive interface; The knowledge update instruction is subjected to format and integrity verification. Extract structured text data from verified knowledge update instructions; The structured text data is converted into a standardized format, and the standardized structured text data is used as knowledge update content.

3. The cross-model knowledge editing and updating method as described in claim 1, characterized in that, The knowledge update content is encoded using the knowledge editing module to generate a knowledge editing vector, including: The updated knowledge content is input into the encoder neural network based on the Transformer architecture in the knowledge editing module; The semantic representation of the knowledge update content is determined by the multi-head attention mechanism of the encoder neural network. The knowledge editing module uses a feedforward neural network to perform a nonlinear transformation on the semantic representation to generate a high-dimensional feature vector. The high-dimensional feature vector is normalized by the normalization layer of the knowledge editing module to obtain the normalized high-dimensional feature vector. The normalized high-dimensional feature vector is mapped to a low-dimensional space through the dimensionality reduction layer of the knowledge editing module to generate an initial low-dimensional vector. The initial low-dimensional vector is normalized by the standardization unit of the knowledge editing module to obtain a standardized vector; The standardized vectors are stored as knowledge editing vectors in the vector repository of the knowledge editing module.

4. The cross-model knowledge editing and updating method as described in claim 1, characterized in that, The knowledge edit vector is input into the cross-model mapping layer, and then written into the intermediate layer parameters of each model in the target model set through the cross-model mapping layer, resulting in the target model set with updated parameters, including: The knowledge editing vector is input into a parameter mapping network across model mapping layers; The parameter structure of each language model in the target model set is analyzed using the parameter mapping network. Based on the results of the parameter structure analysis, the knowledge editing vector is converted into parameter adjustment values ​​corresponding to each language model; Each parameter adjustment value is written into the intermediate layer weight matrix of the corresponding language model, and the updated part of the intermediate layer weight matrix of each language model is subjected to parameter fusion processing. All language models that have completed parameter fusion processing are combined into a set of target models with updated parameters.

5. The cross-model knowledge editing and updating method as described in claim 1, characterized in that, Based on the knowledge update content, a context control tag is generated. The input request is converted into a model input sequence. When the input request and the knowledge update content satisfy the semantic matching condition, the context control tag is inserted into the model input sequence to obtain an input sequence with control tags, including: Based on the knowledge update content, a tag string associated with the knowledge update content is generated as a context control tag; Add the context modulation tags to the vocabulary expansion table of each language model in the target model set; The input request is converted into a model input sequence through word segmentation. Determine the semantic similarity score between the input request and the knowledge update content; When the semantic similarity score exceeds a preset threshold, a context control marker is obtained from the vocabulary expansion table of the corresponding language model in the target model set, and the context control marker is inserted into the starting position of the model input sequence. The model input sequence after inserting the context control marker is used as the input sequence with control marker.

6. The cross-model knowledge editing and updating method as described in claim 1, characterized in that, The input sequence with the control label is input into the target model set after parameter update to obtain a model output set containing the outputs of each model, including: The input sequence with the control label is input in parallel into each language model in the target model set after parameter update; Each language model processes the input sequence with the control marker to generate the corresponding original output sequence; Extract the output text content and corresponding confidence score from the original output sequence of each language model; The output text content and confidence score of each language model are combined into a single model output; The individual model outputs of all language models are aggregated to form a model output set.

7. The cross-model knowledge editing and updating method as described in claim 1, characterized in that, The consistency fusion module performs output alignment and fusion processing on the model output set to generate a fused output result, including: The semantic similarity score between individual model outputs is determined by comparing the output text content of each individual model output in the model output set through the consistency fusion module. When the semantic similarity score of all individual model outputs is greater than or equal to the preset consistency threshold, the output of the single model with the highest confidence is selected as the fusion output result. If the semantic similarity score output by any single model is less than the preset consistency threshold, a weighted fusion is performed based on the confidence scores output by each single model, and the weighted fusion result is used as the fusion output result.

8. A cross-model knowledge editing and updating device, characterized in that, The cross-model knowledge editing and updating device includes: The instruction parsing module is used to receive knowledge update instructions and parse the knowledge update content from the knowledge update instructions; The knowledge editing module is used to encode the knowledge update content and generate a knowledge editing vector. The cross-model mapping module is used to input the knowledge editing vector into the cross-model mapping layer, and write the knowledge editing vector into the intermediate layer parameters of each model in the target model set through the cross-model mapping layer to obtain the target model set with updated parameters. The context control module is used to generate context control tags based on the knowledge update content, convert the input request into a model input sequence, and insert the context control tags into the model input sequence when the input request and the knowledge update content meet the semantic matching conditions, so as to obtain an input sequence with control tags. The model inference module is used to input the input sequence with the control label into the target model set after parameter update to obtain a model output set containing the outputs of each model; The consistency fusion module is used to perform output alignment and fusion processing on the model output set to generate a fused output result. The conflict handling module is used to perform a second inference operation based on the knowledge editing vector and output the final result when output conflicts occur in the model output set and cannot be eliminated by fusion processing.

9. A computer device, characterized in that, The computer device includes a memory, a processor, and a cross-model knowledge editing and updating program stored in the memory and executable on the processor. When executed by the processor, the cross-model knowledge editing and updating program implements the steps of the cross-model knowledge editing and updating method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The storage medium stores a cross-model knowledge editing and updating program, which, when executed by a processor, implements the steps of the cross-model knowledge editing and updating method as described in any one of claims 1-7.

Citation Information

Cited By

  • Large model key value pair fusion processing method, device and system oriented to discontinuous video memory architecture

    CN122065270A