Task processing method, traffic task processing method, and task processing model training method
By determining the key processing units in the artificial intelligence model and adjusting parameters, the problem of unclear model decision-making process is solved, and higher interpretability and transparency are achieved, which enhances user trust and training efficiency.
Patent Information
- Application Number
- PCT/CN2024/125896
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-27
- Filing Date
- 2024-10-18
- Publication Date
- 2025-07-03
AI Technical Summary
The decision-making process and interpretation mechanism of the artificial intelligence model are unclear, which affects users' trust and acceptance of model results, and it is urgent to improve the interpretability and transparency of the model.
By determining the key processing units in the initial processing model, parameter adjustments are made based on the differences between the model processing results and the reverse model processing results, the task processing model is obtained, and the interpretability and transparency of the model are improved.
It improves the interpretability of task processing results and the transparency of the model, enhances the user's trust in model results, and improves the training efficiency of task processing models.
Smart Images

Figure CN2024125896_03072025_PF_FP_ABST
Abstract
Description
Task processing, traffic task processing, and task processing model training methods
[0001] Cross-reference
[0002] This disclosure claims priority to the Chinese patent application filed with the China Patent Office on December 27, 2023, with application number 202311839740.X and invention name “Task processing, traffic task processing and task processing model training method”, the entire contents of which are incorporated by reference in this disclosure. Technical Field
[0003] The embodiments of the present disclosure relate to the field of computer technology, and in particular to task processing, traffic task processing, and task processing model training methods. Background Art
[0004] With the development of computer technology, artificial intelligence models have begun to shine, demonstrating extraordinary capabilities in language understanding, generation, interaction, and reasoning, and are widely used in processing fields such as dialogue, translation, and code generation.
[0005] However, the decision-making process and interpretation mechanism of artificial intelligence models are still unclear. The ability to explain artificial intelligence models is crucial to building trust and acceptance of the model processing process. Understanding the reasoning process of artificial intelligence models can make it easier for users and stakeholders to accept and trust the results of the model. Therefore, a task processing solution with high explainability is urgently needed.
[0006] Summary of the Invention
[0007] In view of this, embodiments of the present disclosure provide a task processing method. One or more embodiments of the present disclosure also relate to a traffic task processing method, a task processing model training method, a task processing device, a traffic task processing device, a task processing model training device, a computing device, a computer-readable storage medium, and a computer program to address technical deficiencies in the prior art.
[0008] According to a first aspect of an embodiment of the present disclosure, there is provided a task processing method, comprising:
[0009] Obtain data to be processed for the target task;
[0010] The data to be processed is input into the task processing model to obtain the task processing result output by the task processing model, wherein the task processing model is obtained based on parameter adjustment of the key processing unit in the initial processing model, and the key processing unit is obtained based on the difference between the model processing result of the initial processing model and the reverse model processing result.
[0011] According to a second aspect of an embodiment of the present disclosure, a method for processing a traffic task is provided, comprising:
[0012] Obtaining traffic data to be processed for a target traffic task;
[0013] The traffic data to be processed is input into the task processing model to obtain the task processing results output by the task processing model, wherein the task processing model is obtained based on parameter adjustment of the key processing units in the initial processing model, and the key processing units are obtained based on the difference between the model processing results of the initial processing model and the reverse model processing results.
[0014] According to a third aspect of an embodiment of the present disclosure, there is provided an initial processing model training method, which is applied to a cloud-side device and includes:
[0015] In response to a model training request for the task processing model, obtaining a model processing result of the initial processing model and unit processing results of a plurality of processing units in the initial processing model;
[0016] For a first processing unit, fixing unit processing results of processing units other than the first processing unit among the plurality of processing units, and reversely adjusting the unit processing result of the first processing unit to obtain a reverse model processing result output by the initial processing model, wherein the first processing unit is any one of the plurality of processing units;
[0017] Determining a unit weight of a first processing unit according to the model processing result and the reverse model processing result;
[0018] Determining key processing units in an initial processing model based on unit weights of multiple processing units;
[0019] Adjust the unit parameters of key processing units and obtain the trained task processing model.
[0020] According to a fourth aspect of an embodiment of the present disclosure, there is provided a task processing device, including:
[0021] A first acquisition component is configured to acquire data to be processed for a target task;
[0022] The first input component is configured to input the data to be processed into the task processing model and obtain the task processing results output by the task processing model, wherein the task processing model is obtained based on parameter adjustment of the key processing unit in the initial processing model, and the key processing unit is obtained based on the difference between the model processing results of the initial processing model and the reverse model processing results.
[0023] According to a fifth aspect of an embodiment of the present disclosure, a traffic task processing device is provided, comprising:
[0024] A second acquisition component is configured to acquire traffic data to be processed for a target traffic task;
[0025] The second input component is configured to input the traffic data to be processed into the task processing model and obtain the task processing results output by the task processing model, wherein the task processing model is obtained based on parameter adjustment of the key processing unit in the initial processing model, and the key processing unit is obtained based on the difference between the model processing results of the initial processing model and the reverse model processing results.
[0026] According to a sixth aspect of an embodiment of the present disclosure, there is provided an initial processing model training apparatus, which is applied to a cloud-side device and includes:
[0027] a third acquisition component configured to, in response to a model training request for the task processing model, acquire a model processing result of the initial processing model and unit processing results of a plurality of processing units in the initial processing model;
[0028] a first adjustment component configured to, for a first processing unit, fix unit processing results of processing units other than the first processing unit among the plurality of processing units, and reversely adjust the unit processing result of the first processing unit to obtain a reverse model processing result output by the initial processing model, wherein the first processing unit is any one of the plurality of processing units;
[0029] a first determining component configured to determine a unit weight of a first processing unit according to a model processing result and a reverse model processing result;
[0030] a second determining component configured to determine a key processing unit in the initial processing model based on unit weights of the plurality of processing units;
[0031] The second adjustment component is configured to adjust the unit parameters of the key processing unit and obtain the trained task processing model.
[0032] According to a seventh aspect of an embodiment of the present disclosure, there is provided a computing device, including:
[0033] memory and processor;
[0034] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method provided in the first aspect, the second aspect, or the third aspect are implemented.
[0035] According to an eighth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, which stores computer-executable instructions. When the instructions are executed by a processor, the steps of the method provided in the first aspect, the second aspect or the third aspect are implemented.
[0036] According to a ninth aspect of an embodiment of the present disclosure, a computer program is provided, wherein when the computer program is executed in a computer, the computer is caused to execute the steps of the method provided in the first aspect, the second aspect, or the third aspect above.
[0037] According to a tenth aspect of an embodiment of the present disclosure, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the steps of the method provided in the first aspect, the second aspect, or the third aspect.
[0038] According to the eleventh aspect of the embodiments of the present disclosure, a computer program product is provided, comprising a non-volatile computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the steps of the method provided in the first aspect, the second aspect, or the third aspect are implemented.
[0039] An embodiment of the present disclosure provides a task processing method, comprising: obtaining data to be processed for a target task; inputting the data to be processed into a task processing model, and obtaining a task processing result output by the task processing model, wherein the task processing model is obtained based on parameter adjustment of key processing units in an initial processing model, and the key processing units are obtained based on the difference between the model processing result of the initial processing model and the processing result of a reverse model. By determining the key processing units in the initial processing model based on the difference between the model processing result and the processing result of the reverse model, the key processing units that determine the model decision process are determined, thereby improving the interpretability and transparency of the model, further improving the interpretability of the task processing results, and since the task processing model is obtained by adjusting the parameters of the key processing units, there is no need to adjust the unit parameters of all units of the initial processing model, thereby improving the training efficiency of the task processing model. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] FIG1 is an architecture diagram of a task processing system provided by one embodiment of the present disclosure;
[0041] FIG2 is an architecture diagram of another task processing system provided by an embodiment of the present disclosure;
[0042] FIG3 is a flowchart of a task processing method provided by one embodiment of the present disclosure;
[0043] FIG4 is a flowchart of training a task processing model in a task processing method provided by one embodiment of the present disclosure;
[0044] FIG5 is a flowchart of a task processing method according to an embodiment of the present disclosure;
[0045] FIG6 is a flowchart of another task processing method provided by an embodiment of the present disclosure;
[0046] FIG7 is a flow chart of a method for processing a traffic task provided by one embodiment of the present disclosure;
[0047] FIG8 is a flowchart of a task processing model training method provided by one embodiment of the present disclosure;
[0048] FIG9 is a schematic structural diagram of a task processing device provided by one embodiment of the present disclosure;
[0049] FIG10 is a schematic structural diagram of a traffic task processing device provided by one embodiment of the present disclosure;
[0050] FIG11 is a structural diagram of a task processing model training device provided by one embodiment of the present disclosure;
[0051] FIG12 is a structural block diagram of a computing device provided by an embodiment of the present disclosure.
[0052] FIG13 is a schematic diagram of the structure of a processor provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0053] The following description sets forth many specific details to facilitate a full understanding of the present disclosure. However, the present disclosure can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of the present disclosure. Therefore, the present disclosure is not limited to the specific implementations disclosed below.
[0054] The terms used in one or more embodiments of the present disclosure are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of the present disclosure. The singular forms "a", "the", and "the" used in one or more embodiments of the present disclosure and the appended claims are also intended to include plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of the present disclosure refers to and includes any or all possible combinations of one or more associated listed items.
[0055] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of the present disclosure, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of the present disclosure, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".
[0056] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of the present disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.
[0057] In one or more embodiments of the present disclosure, a large model refers to a deep learning model with large-scale model parameters, which typically contains hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than ten trillion model parameters. A large model can also be called a cornerstone model / foundation model. The large model is pre-trained using large-scale unlabeled corpus to produce a pre-trained model with more than 100 million parameters. This model can adapt to a wide range of downstream tasks and has good generalization capabilities, such as a large-scale language model (LLM) and a multi-modal pre-training model.
[0058] When large models are used in practice, only a small number of samples are needed to fine-tune the pre-trained model and it can be applied to different tasks. Large models can be widely used in natural language processing (NLP), computer vision and other fields. Specifically, they can be applied to computer vision tasks such as visual question answering (VQA), image caption (IC), and image generation, as well as natural language processing tasks such as text-based sentiment classification, text summary generation, and machine translation. The main application scenarios of large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.
[0059] First, the terms involved in one or more embodiments of the present disclosure are explained.
[0060] Chain of Thought (CoT) refers to a series of logically related thought steps that form a complete thought process. This is often used in the prompt learning of large models. It breaks down the model's reasoning process into multiple steps and presents them intuitively, encouraging the model to first reason and then output the answer, thereby improving the model's accuracy in handling reasoning tasks.
[0061] The emergence of large models has attracted widespread attention. However, the decision-making processes and explanation mechanisms of these models are still unclear. The ability to explain large models is crucial to building trust and acceptance. Understanding the model's reasoning process can make it easier for users and stakeholders to accept and trust the model's results.
[0062] Explainability is a key factor in evaluating and improving model performance. By delving deeper into a model's prediction process, researchers and developers can uncover and address potential biases or errors. This is crucial for ensuring the fairness and accuracy of models. For example, in areas such as security and regulation, the use of AI is strictly regulated, requiring algorithms to provide rationale for their decisions. Therefore, interpretability can help meet regulatory compliance and ethical standards, ensuring that the model's decision-making process is reliable and explainable.
[0063] On this basis, the embodiments of the present disclosure aim to provide a preliminary exploration of the interpretability of the model, study the internal mechanism of the model to complete the thought chain reasoning, locate and explain the key processing units in the model that are related to the thought chain reasoning ability, and verify the behavior patterns of the key processing units when completing the reasoning tasks, laying the foundation for the development of more transparent, interpretable and reliable large models, helping to improve the trust in the model output, improve and correct the deviations and errors in the model, and meet regulatory and ethical requirements, providing guidance and inspiration for future research and development.
[0064] Specifically, the embodiment of the present disclosure proposes a task processing method, which obtains data to be processed for a target task; inputs the data to be processed into a task processing model, and obtains a task processing result output by the task processing model, wherein the task processing model is obtained based on parameter adjustment of key processing units in the initial processing model, and the key processing units are obtained based on the difference between the model processing result of the initial processing model and the processing result of the reverse model. By determining the key processing units in the initial processing model based on the difference between the model processing result and the processing result of the reverse model, the key processing units that determine the decision-making process of the model are determined, thereby improving the interpretability and transparency of the model, and further improving the interpretability of the task processing results. Moreover, since the task processing model is obtained by adjusting the parameters of the key processing units, there is no need to adjust the unit parameters of all units of the initial processing model, thereby improving the training efficiency of the task processing model.
[0065] In the present disclosure, a task processing method is provided. The present disclosure also relates to a traffic task processing method, a task processing model training method, a task processing device, a traffic task processing device, a task processing model training device, a computing device, and a computer-readable storage medium, which are described in detail one by one in the following embodiments.
[0066] Referring to FIG1 , FIG1 shows an architecture diagram of a task processing system provided by an embodiment of the present disclosure. The task processing system may include a client 100 and a server 200;
[0067] The client 100 is used to send the data to be processed for the target task to the server 200;
[0068] The server 200 is configured to input the data to be processed into the task processing model, obtain the task processing results output by the task processing model, wherein the task processing model is obtained by adjusting the parameters of the key processing units in the initial processing model, and the key processing units are obtained based on the difference between the model processing results of the initial processing model and the processing results of the reverse model; and send the task processing results to the client 100;
[0069] The client 100 is also used to receive the task processing result sent by the server 200.
[0070] By applying the solution of the embodiment of the present disclosure, the key processing units in the initial processing model are determined based on the difference between the model processing results and the reverse model processing results, thereby determining the key processing units that determine the model decision-making process, improving the interpretability and transparency of the model, and further improving the interpretability of the task processing results. Moreover, since the task processing model is obtained by adjusting the parameters of the key processing units, there is no need to adjust the unit parameters of all units of the initial processing model, thereby improving the training efficiency of the task processing model.
[0071] Referring to Figure 2, Figure 2 shows an architecture diagram of another task processing system provided by one embodiment of the present disclosure. The task processing system may include multiple clients 100 and a server 200. The clients 100 may include end-side devices, and the server 200 may include cloud-side devices. Multiple clients 100 may establish communication connections through the server 200. In a task processing scenario, the server 200 is used to provide task processing services between multiple clients 100. Multiple clients 100 may act as senders or receivers, respectively, and communicate through the server 200.
[0072] Users can interact with the server 200 through the client 100 to receive data sent by other clients 100, or send data to other clients 100. In the task processing scenario, users can publish data streams to the server 200 through the client 100. The server 200 generates task processing results based on the data stream and pushes the task processing results to other clients with which communication has been established.
[0073] The client 100 and the server 200 are connected via a network. The network provides a medium for the communication link between the client 100 and the server 200. The network can include various connection types, such as wired or wireless communication links or fiber optic cables. The data transmitted by the client 100 may need to be encoded, transcoded, compressed, or other processing before being released to the server 200.
[0074] The client 100 can be a browser, an application (APP), a web application such as a HyperText Markup Language 5 (H5) application, a light application (also known as a mini-program, a lightweight application), or a cloud application. The client 100 can be developed and obtained based on the software development kit (SDK) of the corresponding service provided by the server 200, such as a real-time communication (RTC) SDK. The client 100 can be deployed in an electronic device and needs to rely on the device to run or certain applications in the device to run. For example, the electronic device can have a display screen and support information browsing, such as a personal mobile terminal such as a mobile phone, a tablet computer, a personal computer, etc. Various other types of applications can also be configured in the electronic device, such as human-computer dialogue applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, social platform software, etc.
[0075] The server 200 may include servers that provide various services, such as servers that provide communication services to multiple clients, servers that provide background training to support models used on clients, and servers that process data sent by clients. It should be noted that the server 200 can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. The server can also be a server for a distributed system, or a server that is integrated with a blockchain. The server can also be a cloud server for basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.
[0076] It is worth noting that the task processing methods provided in the embodiments of the present disclosure are generally executed by the server. However, in other embodiments of the present disclosure, the client may also have similar functions to the server and thus execute the task processing methods provided in the embodiments of the present disclosure. In other embodiments, the task processing methods provided in the embodiments of the present disclosure may also be executed jointly by the client and the server.
[0077] 3 , which shows a flowchart of a task processing method provided by an embodiment of the present disclosure, specifically comprising the following steps:
[0078] Step 302: Obtain data to be processed for the target task.
[0079] In one or more embodiments of the present disclosure, the data to be processed of the target task can be used to perform task processing based on the data to be processed, and obtain the task processing result corresponding to the data to be processed.
[0080] Specifically, target tasks can be tasks in different scenarios, such as reasoning tasks, traffic tasks, question-answering tasks, retrieval tasks, etc. The data to be processed can be data in different formats, such as voice data, text data, video data, etc. The data to be processed can also be data in different languages, such as English data, Chinese data, etc.
[0081] In practical applications, there are multiple ways to obtain the data to be processed for the target task, and the specific method to be selected depends on the actual situation. The embodiments of the present disclosure do not impose any restrictions on this. In one possible implementation of the present disclosure, the data to be processed for the target task can be received from the user. In another possible implementation of the present disclosure, the data to be processed for the target task can be read from other data acquisition devices or databases.
[0082] Step 304: Input the data to be processed into the task processing model, and obtain the task processing results output by the task processing model, wherein the task processing model is obtained based on parameter adjustment of the key processing unit in the initial processing model, and the key processing unit is obtained based on the difference between the model processing results of the initial processing model and the reverse model processing results.
[0083] By applying the solution of the embodiment of the present disclosure, the key processing units in the initial processing model are determined based on the difference between the model processing results and the reverse model processing results, thereby determining the key processing units that determine the model decision-making process, improving the interpretability and transparency of the model, and further improving the interpretability of the task processing results. Moreover, since the task processing model is obtained by adjusting the parameters of the key processing units, there is no need to adjust the unit parameters of all units of the initial processing model, thereby improving the training efficiency of the task processing model.
[0084] In an optional embodiment of the present disclosure, referring to FIG4 , FIG4 shows a flow chart of training a task processing model in a task processing method provided by an embodiment of the present disclosure. Before inputting the data to be processed into the task processing model and obtaining the task processing result output by the task processing model, the following steps 402 to 408 may also be included:
[0085] Step 402: Obtain the model processing result of the initial processing model and the unit processing results of multiple processing units in the initial processing model.
[0086] Specifically, the initial processing model is an artificial intelligence model, which can be a large model or other pre-trained neural network models. The model processing result refers to the output result obtained by the initial processing model processing the input data. For example, the input data is the question to be answered "Where does the sun rise from", and the model processing result is the answer "The sun rises from the east of the earth". The processing unit refers to the unit that processes the model input data. The processing unit includes but is not limited to the attention unit and the linear mapping unit. The unit processing result of the processing unit is determined according to the processing unit. For example, if the processing unit is an attention head, the unit processing result is the attention vector.
[0087] In actual applications, there are many ways to obtain the model processing results of the initial processing model and the unit processing results of multiple processing units in the initial processing model. The specific selection is based on the actual situation, and the embodiments of the present disclosure do not impose any restrictions on this.
[0088] In a possible implementation of the present disclosure, the model processing result of the initial processing model and the unit processing results of multiple processing units in the initial processing model may be read from other data acquisition devices or databases.
[0089] In another possible implementation of the present disclosure, an initial processing model may be used to generate model processing results and unit processing results of multiple processing units in real time. The initial processing model includes multiple processing units, residual units, and a multi-layer neural network. The above-mentioned acquisition of the model processing results of the initial processing model and the unit processing results of the multiple processing units in the initial processing model may include the following steps:
[0090] Get factual data;
[0091] Processing the fact data through multiple processing units to obtain unit processing results of the multiple processing units;
[0092] The unit processing results are mapped through the residual unit and the multi-layer neural network to obtain the model processing results of the initial processing model.
[0093] Specifically, factual data can include the question to be inferred and the reasoning data for the question to be inferred. Based on this factual data, answers that conform to objective facts can be obtained. The format of factual data can be <question, reason, answer>. Reasoning data can be understood as forward reasoning ideas and processes. Because the model processing results and unit processing results obtained based on this factual data can be used to compare differences with the reverse model processing results to determine the criticality of the processing unit, factual data can also be called reference data.
[0094] It should be noted that there are multiple ways to obtain factual data, and the specific method to be selected depends on the actual situation. The embodiments of this disclosure do not impose any restrictions on this. In one possible implementation of this disclosure, factual data sent by a user can be received. In another possible implementation of this disclosure, factual data can be read from other data acquisition devices or databases.
[0095] By applying the solution of the embodiment of the present disclosure, fact data is obtained; the fact data is processed by multiple processing units to obtain unit processing results of the multiple processing units; the unit processing results are mapped by residual units and multi-layer neural networks to obtain model processing results of the initial processing model. Since the unit processing results and the model processing results are generated based on the fact data, the failure of locating the key processing units due to the unprepared model processing results is avoided. Therefore, it is ensured that the key processing units can be accurately determined through subsequent result comparison.
[0096] In an optional embodiment of the present disclosure, the in-context learning capability of the pre-trained language model may be utilized to construct fact data to stimulate the reasoning capability of the initial processing model. That is, the above-mentioned acquisition of fact data may include the following steps:
[0097] Obtain the question to be reasoned;
[0098] Input the question to be inferred and the reasoning prompt information into the pre-trained language model to obtain the reasoning data corresponding to the question to be inferred;
[0099] Construct factual data based on the question to be inferred and the inference data.
[0100] Specifically, the pre-trained language model can be a large model or a natural language neural network model trained based on multiple sample questions and the corresponding reasoning data labels. Reasoning prompt information is used to guide the pre-trained language model in generating positive reasoning data. The reasoning prompt information includes reasoning prompt templates that indicate the tasks to be performed by the pre-trained language model, as well as example reasoning question-answer pairs for the pre-trained language model to learn. These example reasoning question-answer pairs include both the reasoning process and the reasoning answer.
[0101] It should be noted that there are multiple ways to obtain the question to be reasoned, and the specific method to be selected depends on the actual situation. The embodiments of this disclosure do not impose any restrictions on this. In one possible implementation of this disclosure, the question to be reasoned can be received from a user. In another possible implementation of this disclosure, the question to be reasoned can be read from other data acquisition devices or databases.
[0102] For example, assuming that the question to be inferred is "What home entertainment equipment requires cables? The answer options are: (a) radio shed; (b) substation; (c) cabinet; (d) TV; (e) table", the reasoning prompt information and the question to be inferred are input into the pre-trained language model, and the reasoning data corresponding to the question to be inferred is "Among the above choices, only the TV requires cables. Therefore, the answer is".
[0103] By applying the solution of the embodiment of the present disclosure, a question to be inferred is obtained; the question to be inferred and the reasoning prompt information are input into a pre-trained language model to obtain the reasoning data corresponding to the question to be inferred; and factual data is constructed based on the question to be inferred and the reasoning data, thereby utilizing the in-text learning capability of the pre-trained language model to ensure the accuracy of the factual data.
[0104] Step 404: For the first processing unit, fix the unit processing results of the processing units other than the first processing unit among the multiple processing units, and reversely adjust the unit processing results of the first processing unit to obtain the reverse model processing results output by the initial processing model, wherein the first processing unit is any one of the multiple processing units.
[0105] It should be noted that, since the multiple processing units in the initial processing model constitute different reasoning paths, in order to determine whether each processing unit is a critical processing unit, a path patching method can be used to apply causal intervention to the reasoning path where the first processing unit is located, that is, the unit processing result is reversely adjusted. The disturbance will propagate along the reasoning path where the first processing unit is located to the output node to obtain the reverse model processing result. When reversely adjusting the unit processing result of the first processing unit, in order to ensure independent observation of the impact of the first processing unit on the model output result, the unit processing results of the processing units other than the first processing unit among the multiple processing units can be fixed, so that only the unit processing result of the first processing unit is causally disturbed.
[0106] In practical applications, there are multiple ways to reversely adjust the unit processing result of the first processing unit, and the specific method is selected based on actual circumstances. The embodiments of the present disclosure do not impose any restrictions on this. In one possible implementation of the present disclosure, noise can be added to the unit processing result to obtain a reverse unit processing result that is significantly different from the unit processing result. The reverse unit processing result is then propagated through the initial processing model to obtain a reverse model processing result.
[0107] In another possible implementation of the present disclosure, the unit processing result of the first processing unit may be reversely adjusted using counterfactual data. That is, the reverse adjustment of the unit processing result of the first processing unit to obtain the reverse model processing result output by the initial processing model may include the following steps:
[0108] Obtain counterfactual data;
[0109] The counterfactual data is processed by the first processing unit to obtain a reverse unit processing result of the first processing unit;
[0110] The unit processing result of the first processing unit is replaced by the reverse unit processing result, and the reverse unit processing result is propagated through the initial processing model to obtain the reverse model processing result.
[0111] Specifically, counterfactual data includes the problem to be inferred and the reverse reasoning data of the problem to be inferred. Based on the counterfactual data, an incorrect answer to the problem to be inferred can be obtained. Reverse reasoning data can be understood as the reverse reasoning idea and reverse reasoning process.
[0112] It should be noted that there are multiple ways to obtain counterfactual data, and the specific method to be selected depends on the actual situation. The embodiments of this disclosure do not impose any restrictions on this. In one possible implementation of this disclosure, counterfactual data sent by a user can be received. In another possible implementation of this disclosure, counterfactual data can be read from other data acquisition devices or databases.
[0113] Referring to FIG5 , FIG5 shows a flowchart of a processing process of a task processing method provided by an embodiment of the present disclosure. As shown in FIG5 , the lines represent the reasoning path of the initial processing model, which includes multiple residual units, an output unit, a multi-layer neural network (Multi-Layer Perceptron, abbreviated as MLP) 1, attention head 1.0, attention head 1.1, multi-layer neural network 0, attention head 0.0, and attention head 0.1. Assume that the current judgment is whether attention head 0.0 is a key processing unit: First, the last token in the embedding vector is input into the initial processing model, and the unit processing results of all processing units (multi-layer neural network and attention head) are collected. Then, the unit processing results of other attention heads (for example, attention head 1.0 and attention head 1.1) on the factual data are frozen to perform a causal perturbation on attention head 0.0, replacing its original unit processing result with the reverse unit processing result (obtained based on the counterfactual data). The perturbation effect will propagate to the output unit along the reasoning path shown by the solid line, obtaining the reverse model processing result. Finally, based on the model processing results (obtained based on the actual data) and the reverse model processing results, it is determined whether the attention head 0.0 is a key processing unit.
[0114] It should be noted that to ensure independent observation of the impact of attention head 0.0, the inference path shown by the solid line includes the forward path through the residual unit connection and MLP, but does not include other attention heads.
[0115] By applying the solution of the embodiment of the present disclosure, counterfactual data is obtained; the counterfactual data is processed by the first processing unit to obtain the reverse unit processing result of the first processing unit; the unit processing result of the first processing unit is replaced by the reverse unit processing result, and the reverse unit processing result is propagated through the initial processing model to obtain the reverse model processing result, thereby ensuring the accuracy of the reverse model processing result.
[0116] In an optional embodiment of the present disclosure, counterfactual data may be generated based on factual data. That is, the above-mentioned acquisition of counterfactual data may include the following steps:
[0117] Acquiring factual data, wherein the factual data includes inference data;
[0118] Adjust the inference data to replacement data that is unrelated to the inference data to obtain counterfactual data.
[0119] For example, citing the example of the above question to be inferred as "What home entertainment equipment requires cables? Answer options: (a) Radio shed (b) Substation (c) Cabinet (d) TV (e) Table", the inference data "Among the above choices, only the TV requires cables. Therefore, the answer is" is adjusted to the replacement data "Playful dolphins jump and play in the sparkling blue ocean. The cat sleeps peacefully on the comfortable cushion. The answer is" to obtain the counterfactual data.
[0120] In practical applications, the inference data is adjusted to replacement data that is unrelated to the inference data. Before obtaining the counterfactual data, the replacement data can be obtained. There are multiple ways to obtain the replacement data, and the specific method to be selected depends on the actual situation. The embodiments of the present disclosure do not impose any restrictions on this. In one possible implementation of the present disclosure, the replacement data sent by the user can be received. In another possible implementation of the present disclosure, the replacement data can be read from other data acquisition devices or databases.
[0121] By applying the solution of the embodiment of the present disclosure, fact data is obtained, wherein the fact data includes inference data; the inference data is adjusted to replacement data unrelated to the inference data to obtain counterfactual data, thereby ensuring the accuracy of the counterfactual data and further ensuring the accuracy of the inverse model processing results.
[0122] Step 406: Determine the unit weight of the first processing unit according to the model processing result and the reverse model processing result.
[0123] Specifically, the unit weight is used to characterize the importance of the first processing unit in the initial processing model reasoning process. After causal perturbation is performed on the first processing unit, the reverse model processing result is obtained, and the difference between the model processing result and the reverse model processing result is further compared. If the reverse model processing result differs significantly from the model processing result, it means that the unit weight of the first processing unit is larger, that is, the first processing unit is more important in the initial processing model reasoning process and is a key processing unit; if the reverse model processing result differs slightly from the model processing result, it means that the unit weight of the first processing unit is smaller, that is, the first processing unit is less important in the initial processing model reasoning process and is not a key processing unit.
[0124] In practical applications, there are multiple ways to determine the unit weight of the first processing unit based on the model processing results and the reverse model processing results. The specific selection is based on the actual situation, and the embodiments of the present disclosure do not impose any restrictions on this.
[0125] In a possible implementation of the present disclosure, the importance of the first processing unit can be determined by observing the causal effect in the model processing result after the causal perturbation. That is, determining the unit weight of the first processing unit based on the model processing result and the reverse model processing result can include the following steps:
[0126] Parsing the reverse model processing result to determine a first associated keyword in the reverse model processing result, and parsing the model processing result to determine a second associated keyword in the model processing result, wherein both the first associated keyword and the second associated keyword are related to the problem to be inferred;
[0127] Determining a first weight metric value of the first processing unit based on the first associated keyword and the reverse model processing result, and determining a second weight metric value of the first processing unit based on the second associated keyword and the model processing result;
[0128] A unit weight of the first processing unit is determined according to the first weight metric value and the second weight metric value.
[0129] It should be noted that since the word distribution in the model processing results involves a very large number of words, the resulting disturbance is very sparse. Therefore, the embodiment of the present disclosure measures the disturbance effect by observing and answering the probability changes of related keywords related to the question to be inferred, thereby quickly locating the key processing units. Among them, the related keywords can be understood as words of interest (WoI for short), and the unit weights determined based on the related keywords can be used to measure the changes in the model's thinking chain reasoning ability.
[0130] In practical applications, the weight metric value can be determined by the following formula (1), and the sum of the first weight metric value and the second weight metric value can be obtained. The rate of change of the first weight metric value and the second weight metric value can be used as a metric for evaluating the unit weight of the first processing unit. That is, the unit weight can be determined by the following formula (2):
[0131] Among them, t represents the weight measurement value, P gt represents the probability of associated keywords, ∑P cd represents the sum of the probabilities of each word, h i,j represents the first processing unit, represents the unit weight of the first processing unit, represents the first weighted metric value, and t0 represents the second weighted metric value. A negative value indicates that the relative confidence of the reverse model processing results decreases, indicating that the inference performance of the initial processing model is impaired. The greater the drop, the more important the first processing unit is.
[0132] By applying the solution of the embodiment of the present disclosure, the reverse model processing result is parsed to determine the first associated keyword in the reverse model processing result, and the model processing result is parsed to determine the second associated keyword in the model processing result, wherein both the first associated keyword and the second associated keyword are related to the problem to be inferred; based on the first associated keyword and the reverse model processing result, the first weight measurement value of the first processing unit is determined, and based on the second associated keyword and the model processing result, the second weight measurement value of the first processing unit is determined; based on the first weight measurement value and the second weight measurement value, the unit weight of the first processing unit is determined, thereby measuring the disturbance effect by the probability change of the associated keywords before and after the causal disturbance, thereby ensuring the accuracy of the unit weight.
[0133] In another possible implementation of the present disclosure, determining the unit weight of the first processing unit based on the model processing result and the reverse model processing result may include the following steps:
[0134] Parsing the reverse model processing result to determine a first associated keyword in the reverse model processing result, wherein the first associated keyword is related to the problem to be inferred;
[0135] Inputting the first associated keyword into the initial processing model to obtain a prediction processing result output by the initial processing model;
[0136] The unit weight of the first processing unit is determined according to the model processing result and the prediction processing result.
[0137] It should be noted that when parsing the reverse model processing results and determining the first associated keyword in the reverse model processing results, the keyword at the last token position in the result sequence corresponding to the model processing results can be directly used as the first associated keyword based on the keyword order. Alternatively, a keyword location function can be used to directly extract the first associated keyword from the reverse model processing results.
[0138] In practical applications, since the initial processing model performs the same task when generating the reverse model processing results and the model processing results, in order to avoid the impact of the task on the unit weight, after determining the first associated keyword, the first associated keyword can be input into the initial processing model, and the initial processing model can freely generate the predicted processing results.
[0139] Furthermore, when determining the unit weight of the first processing unit based on the model processing results and the predicted processing results, the model processing results and the predicted processing results can be directly compared. If the predicted processing result and the model processing result differ significantly, it indicates that the unit weight of the first processing unit is larger, that is, the first processing unit is more important in the initial processing model reasoning process and is a key processing unit; if the predicted processing result and the model processing result differ slightly, it indicates that the unit weight of the first processing unit is smaller, that is, the first processing unit is less important in the initial processing model reasoning process and is not a key processing unit. The unit weight of the first processing unit can also be generated using a pre-trained language model.
[0140] By applying the solution of the embodiment of the present disclosure, the reverse model processing result is parsed to determine the first associated keyword in the reverse model processing result, wherein the first associated keyword is related to the problem to be inferred; the first associated keyword is input into the initial processing model to obtain the predicted processing result output by the pre-trained language model; and the unit weight of the first processing unit is determined based on the model processing result and the predicted processing result. Since the predicted processing result is freely generated by the initial processing model and is not limited to a certain task, the predicted processing result is more flexible.
[0141] In a possible implementation of the present disclosure, determining the unit weight of the first processing unit based on the model processing result and the prediction processing result may include the following steps:
[0142] The weight generation prompt information, the model processing result, and the prediction processing result are input into the pre-trained language model to obtain the unit weight of the first processing unit.
[0143] Specifically, the weight generation prompt information is used to guide the pre-trained language model to generate unit weights. For example, the weight generation prompt information may be "Please score the model processing results and the degree of change in the predicted processing results to obtain the unit weights."
[0144] By applying the solution of the embodiment of the present disclosure, weight generation prompt information, model processing results and prediction processing results are input into the pre-trained language model to obtain the unit weight of the first processing unit, thereby improving the efficiency of obtaining the unit weight.
[0145] Step 408: Determine the key processing unit in the initial processing model based on the unit weights of the multiple processing units, adjust the parameters of the key processing unit, and obtain the task processing model.
[0146] Specifically, a key processing unit refers to a processing unit among multiple processing units that is relatively important or even indispensable for the initial processing model reasoning process. A key processing unit can also be understood as a processing unit that has a decisive influence on the reasoning result of the initial processing model.
[0147] In practical applications, there are many ways to determine the key processing units in the initial processing model based on the unit weights of multiple processing units. The specific selection is based on the actual situation, and the embodiments of the present disclosure do not impose any restrictions on this. In one possible implementation of the present disclosure, the unit weights of multiple processing units can be sorted, and the top N processing units can be determined as key processing units. In another possible implementation of the present disclosure, a weight threshold can be obtained, and the processing units whose unit weights are greater than or equal to the weight threshold can be determined as key processing units. Among them, N and the weight threshold are set according to the actual situation.
[0148] By applying the solution of the embodiment of the present disclosure, by reversely adjusting the unit processing results and observing the changes in the adjusted model processing results, the importance of the processing unit in the model decision-making process can be accurately determined, thereby improving the transparency and trust of the model output and ensuring that the model decision-making process is reliable and explainable.
[0149] In an optional embodiment of the present disclosure, after determining the key processing units in the initial processing model based on the unit weights of the multiple processing units, the following steps may be further included:
[0150] Perform criticality verification on key processing units and determine the verification results of key processing units.
[0151] It should be noted that the verification result indicates whether the key processing unit is the one that determines the model's decision-making process. After determining the key processing unit, its criticality (importance) can be verified, while the non-criticality of other processing units can be confirmed, thereby ensuring the accuracy of the key processing unit. Furthermore, the verification results of the key processing unit can be fed back to the user, improving the transparency of the model's decision-making process.
[0152] In practical applications, there are many ways to perform criticality verification on key processing units and determine the verification results of key processing units. The specific method is selected according to the actual situation, and the embodiments of the present disclosure do not impose any limitations on this.
[0153] In one possible implementation of the present disclosure, the unit processing result of the key processing unit can be fixed, and the unit processing results of the processing units other than the key processing unit can be reversely adjusted to obtain a first verification result output by the initial processing model. If the first verification result is less different from the model processing result before the reverse adjustment, it means that the key processing unit is an accurate key processing unit. There are many ways to determine the difference between the first verification result and the model processing result, and the specific selection is made according to the actual situation. The embodiment of the present disclosure does not impose any restrictions on this. In one possible implementation of the present disclosure, the cosine similarity between the first verification result and the model processing result can be calculated. If the cosine similarity is less than or equal to a preset verification threshold, it is determined that the difference is small and the key processing unit is an accurate key processing unit. The preset verification threshold is set according to the actual situation. In another possible implementation of the present disclosure, the first verification result and the model processing result can be input into a difference comparison model to obtain the difference between the first verification result and the model processing result.
[0154] In another possible implementation of the present disclosure, the criticality of the key processing unit may be verified by a unit knockout method. That is, the above-mentioned criticality verification of the key processing unit and determination of the verification result of the key processing unit may include the following steps:
[0155] Screening out a control treatment unit from multiple treatment units;
[0156] Fixing unit processing results of processing units other than the key processing unit among the multiple processing units, and reversely adjusting the unit processing results of the key processing unit to obtain reverse key processing results output by the initial processing model;
[0157] Fixing unit processing results of processing units other than the control processing unit among the multiple processing units, and reversely adjusting the unit processing results of the control processing unit to obtain a reverse control processing result output by the initial processing model;
[0158] The verification result of the key processing unit is determined based on the reverse key processing result and the reverse control processing result.
[0159] Specifically, the number of control processing units randomly selected from the plurality of processing units may be one or more. If there is one control processing unit, then the control processing unit is a processing unit other than the key processing unit. If there are multiple control processing units, then the control processing unit includes at least one processing unit other than the key processing unit. Preferably, all control processing units are processing units other than the key processing units, and their number is the same as the number of key processing units.
[0160] It should be noted that the implementation method of “fixing the unit processing results of the processing units other than the key processing units in the multiple processing units, and reversely adjusting the unit processing results of the key processing units to obtain the reverse key processing results output by the initial processing model; fixing the unit processing results of the processing units other than the control processing units in the multiple processing units, and reversely adjusting the unit processing results of the control processing units to obtain the reverse control processing results output by the initial processing model” can refer to the above-mentioned implementation method of “fixing the unit processing results of the processing units other than the first processing unit in the multiple processing units, and reversely adjusting the unit processing results of the first processing unit to obtain the reverse model processing results output by the initial processing model”, and the embodiments of the present disclosure will not be repeated here.
[0161] In practice, the unit processing results of the key processing unit are reversed, while the unit processing results of the control processing unit are reversed as a control group. The initial processing model is then run through forward reasoning to obtain the reverse key processing results and the reverse control processing results. If the accuracy of the reverse key processing results is lower than that of the reverse control results, it indicates that the reasoning ability of the initial processing model has declined and the key processing unit is an accurate key processing unit.
[0162] It's worth noting that eliminating key processing units through unit elimination significantly reduced the model's reasoning ability. Furthermore, some key processing units play a key role in determining the final answer, while others play a key role in synthesizing and gradually thinking to arrive at the answer. This corresponds to the two stages of the chain of thought reasoning process: first, gradually thinking to obtain intermediate ideas, and then answering the question based on these ideas.
[0163] By applying the solution of the embodiment of the present disclosure, the unit processing results of the processing units other than the key processing units among the multiple processing units are fixed, and the unit processing results of the key processing units are reversely adjusted to obtain the reverse key processing results output by the initial processing model; the unit processing results of the processing units other than the control processing units among the multiple processing units are fixed, and the unit processing results of the control processing units are reversely adjusted to obtain the reverse control processing results output by the initial processing model; the verification results of the key processing units are determined based on the reverse key processing results and the reverse control processing results, thereby ensuring the accuracy of the verification results.
[0164] Referring to Figure 6, Figure 6 shows a flowchart of the processing process of another task processing method provided by an embodiment of the present disclosure. When the unit elimination method is used to perform criticality verification on the key processing unit, it is assumed that the reasoning ability of the initial processing model is not affected after the control processing unit is eliminated, but it shows a significant deterioration when the key processing unit is eliminated. In this case, it can be verified that the key processing unit contributes to the reasoning task. As shown in Figure 5, during the unit elimination process, the unit processing result of each key processing unit in the factual data can be replaced with the reverse unit processing result corresponding to the counterfactual data. Then, observe and compare the changes in the prediction of the "answer" output by the initial processing model before and after the unit elimination.
[0165] In an optional embodiment of the present disclosure, after determining the key processing units in the initial processing model based on the unit weights of the multiple processing units, the following steps may be further included:
[0166] Adjust unit parameters of key processing units and obtain an initial processing model after parameter adjustment.
[0167] It should be noted that after determining the key processing units in the initial processing model, since the key processing units only account for a part of the multiple processing units in the initial processing model, only adjusting the unit parameters of the key processing units can improve the training speed of the initial processing model.
[0168] In practical applications, there are many ways to adjust the unit parameters of the key processing unit, and the specific selection is based on the actual situation. The embodiments of the present disclosure do not impose any restrictions on this. In one possible implementation of the present disclosure, the target unit parameters sent by the user can be obtained, and the unit parameters of the key processing unit can be directly replaced with the target unit parameters to obtain the initial processing model after parameter adjustment. In another possible implementation of the present disclosure, a training sample set can be obtained, and multiple training samples in the training sample set can be input into the initial processing model to obtain the prediction results of the initial processing model. The loss value is calculated based on the true label of the training sample and the prediction result. The unit parameters of the key processing unit are adjusted based on the loss value until the preset stop condition is reached, and the initial processing model after parameter adjustment is obtained. At this time, the initial processing model after parameter adjustment can be understood as the initial processing model after training.
[0169] By applying the solution of the embodiment of the present disclosure, the unit parameters of the key processing units are adjusted, and the initial processing model after the parameter adjustment is obtained, the training speed of the initial processing model is improved.
[0170] The following further illustrates the task processing method provided by the present disclosure in conjunction with FIG7 , taking its application in a smart traffic scenario as an example. FIG7 shows a flow chart of a traffic task processing method provided by an embodiment of the present disclosure, which specifically includes the following steps:
[0171] Step 702: Obtain traffic data to be processed for the target traffic task.
[0172] Step 704: Input the traffic data to be processed into the task processing model, and obtain the task processing results output by the task processing model, wherein the task processing model is obtained based on parameter adjustment of the key processing unit in the initial processing model, and the key processing unit is obtained based on the difference between the model processing results of the initial processing model and the reverse model processing results.
[0173] It should be noted that the implementation of steps 702 to 704 can refer to the implementation of steps 302 to 304 above, and will not be described in detail in this embodiment. Target traffic tasks include but are not limited to traffic route planning tasks, map planning tasks, and the like.
[0174] By applying the solution of the embodiment of the present disclosure, the key processing units in the initial processing model are determined based on the difference between the model processing results and the reverse model processing results, thereby determining the key processing units that determine the model decision-making process, improving the interpretability and transparency of the model, and further improving the interpretability of the task processing results. Moreover, since the task processing model is obtained by adjusting the parameters of the key processing units, there is no need to adjust the unit parameters of all units of the initial processing model, thereby improving the training efficiency of the task processing model.
[0175] In an optional embodiment of the present disclosure, after inputting the traffic data to be processed into the task processing model and obtaining the task processing result output by the task processing model, the following steps may be included:
[0176] Receive adjustment data sent by the user based on the task processing result, and adjust the model parameters of the task processing model according to the adjustment data.
[0177] It should be noted that after obtaining the task processing result, the task processing result can be sent to the client so that the client can display the task processing result to the user. There are many ways for the client to display the task processing result to the user, and the specific selection is made according to the actual situation. The embodiments of the present disclosure do not impose any restrictions on this. In one possible implementation of the present disclosure, the task processing result can be directly displayed to the user. In another possible implementation of the present disclosure, the task processing result can be displayed to the user based on the user's display requirement information. Among them, the display requirement information represents the user's demand for viewing the task processing results, and the display requirement information includes but is not limited to displaying only the task processing results, displaying the storage path of the task processing results, and displaying the data to be processed and the task processing results.
[0178] In actual applications, the user may not be satisfied with the task processing results. In this case, the system can receive adjustment data sent by the user based on the task processing results and adjust the model parameters of the task processing model according to the adjustment data. The adjustment data includes but is not limited to the adjusted task processing results. The process of adjusting the model parameters of the task processing model according to the adjustment data can be referred to the training process of the task processing model, and will not be further described in this embodiment.
[0179] By applying the solution of the embodiment of the present disclosure, adjustment data sent by the user based on the task processing results is received, and the model parameters of the task processing model are adjusted according to the adjustment data, so that the task processing model can be updated based on the adjustment data fed back by the user, thereby improving the accuracy of the task processing model and, at the same time, improving the user experience.
[0180] 8 , which shows a flow chart of a task processing model training method provided by one embodiment of the present disclosure. The initial processing model training method is applied to a cloud-side device and specifically includes the following steps:
[0181] Step 802: In response to a model training request for a task processing model, obtain the model processing result of the initial processing model and the unit processing results of multiple processing units in the initial processing model.
[0182] Step 804: For the first processing unit, fix the unit processing results of the processing units other than the first processing unit among the multiple processing units, and reversely adjust the unit processing results of the first processing unit to obtain the reverse model processing results output by the initial processing model, wherein the first processing unit is any one of the multiple processing units.
[0183] Step 806: Determine the unit weight of the first processing unit according to the model processing result and the reverse model processing result.
[0184] Step 808: Determine the key processing unit in the initial processing model according to the unit weights of the multiple processing units.
[0185] Step 810: Adjust the unit parameters of the key processing unit and obtain the trained task processing model.
[0186] It should be noted that the implementation of steps 802 to 810 may refer to the implementation of steps 402 to 408 above, and will not be described in detail in this embodiment of the present disclosure.
[0187] By applying the solution of the disclosed embodiments, by reversely adjusting the unit processing results and observing the changes in the adjusted model processing results, the importance of the processing unit in the model decision-making process can be accurately determined, thereby improving the transparency and trustworthiness of the model output and enhancing the interpretability of the model decision-making process. Furthermore, after determining the key processing unit, the unit parameters of the key processing unit can be directly adjusted, reducing the amount of parameter adjustment and improving the efficiency of task processing model training.
[0188] Corresponding to the above-mentioned task processing method embodiment, the present disclosure also provides a task processing device embodiment. FIG9 shows a schematic diagram of the structure of a task processing device provided by one embodiment of the present disclosure. As shown in FIG9 , the device includes:
[0189] A first acquisition component 902 is configured to acquire data to be processed for a target task;
[0190] The first input component 904 is configured to input the data to be processed into the task processing model and obtain the task processing results output by the task processing model, wherein the task processing model is obtained based on parameter adjustment of the key processing unit in the initial processing model, and the key processing unit is obtained based on the difference between the model processing result of the initial processing model and the reverse model processing result.
[0191] Optionally, the device also includes: a task processing model training component, configured to obtain the model processing results of the initial processing model and the unit processing results of multiple processing units in the initial processing model; for the first processing unit, fix the unit processing results of the processing units other than the first processing unit in the multiple processing units, and reversely adjust the unit processing results of the first processing unit to obtain the reverse model processing results output by the initial processing model, wherein the first processing unit is any one of the multiple processing units; determine the unit weight of the first processing unit based on the model processing results and the reverse model processing results; determine the key processing units in the initial processing model based on the unit weights of the multiple processing units, and adjust the parameters of the key processing units to obtain the task processing model.
[0192] Optionally, the task processing model training component is further configured to parse the reverse model processing results, determine the first associated keyword in the reverse model processing results, and parse the model processing results to determine the second associated keyword in the model processing results, wherein both the first associated keyword and the second associated keyword are related to the problem to be inferred; determine the first weight measurement value of the first processing unit based on the first associated keyword and the reverse model processing results, and determine the second weight measurement value of the first processing unit based on the second associated keyword and the model processing results; determine the unit weight of the first processing unit based on the first weight measurement value and the second weight measurement value.
[0193] Optionally, the task processing model training component is further configured to parse the reverse model processing results, determine the first associated keyword in the reverse model processing results, wherein the first associated keyword is related to the problem to be inferred; input the first associated keyword into the initial processing model to obtain the predicted processing results output by the initial processing model; and determine the unit weight of the first processing unit based on the model processing results and the predicted processing results.
[0194] Optionally, the task processing model training component is further configured to input weight generation prompt information, model processing results and prediction processing results into the pre-trained language model to obtain the unit weight of the first processing unit.
[0195] Optionally, the task processing model training component is further configured to filter out control processing units from multiple processing units; fix the unit processing results of the processing units other than the key processing units in the multiple processing units, and reversely adjust the unit processing results of the key processing units to obtain the reverse key processing results output by the initial processing model; fix the unit processing results of the processing units other than the control processing units in the multiple processing units, and reversely adjust the unit processing results of the control processing units to obtain the reverse control processing results output by the initial processing model; determine the verification result of the key processing unit based on the reverse key processing result and the reverse control processing result.
[0196] Optionally, the task processing model training component is further configured to obtain counterfactual data; process the counterfactual data through the first processing unit to obtain the reverse unit processing result of the first processing unit; replace the unit processing result of the first processing unit with the reverse unit processing result, and propagate the reverse unit processing result through the initial processing model to obtain the reverse model processing result.
[0197] Optionally, the task processing model training component is further configured to obtain factual data, wherein the factual data includes inference data; adjust the inference data to replacement data that is unrelated to the inference data, and obtain counterfactual data.
[0198] Optionally, the initial processing model includes multiple processing units, residual units and multi-layer neural networks; the task processing model training component is further configured to obtain fact data; the fact data is processed by multiple processing units to obtain unit processing results of multiple processing units; the unit processing results are mapped by residual units and multi-layer neural networks to obtain model processing results of the initial processing model.
[0199] Optionally, the task processing model training component is further configured to obtain the problem to be inferred; input the problem to be inferred and the reasoning prompt information into the pre-trained language model to obtain the reasoning data corresponding to the problem to be inferred; and construct fact data based on the problem to be inferred and the reasoning data.
[0200] By applying the solution of the embodiment of the present disclosure, the key processing units in the initial processing model are determined based on the difference between the model processing results and the reverse model processing results, thereby determining the key processing units that determine the model decision-making process, improving the interpretability and transparency of the model, and further improving the interpretability of the task processing results. Moreover, since the task processing model is obtained by adjusting the parameters of the key processing units, there is no need to adjust the unit parameters of all units of the initial processing model, thereby improving the training efficiency of the task processing model.
[0201] The above is a schematic scheme of a task processing device of this embodiment. It should be noted that the technical scheme of the task processing device and the technical scheme of the task processing method described above are of the same concept. For details not described in detail in the technical scheme of the task processing device, please refer to the description of the technical scheme of the task processing method described above.
[0202] Corresponding to the above-mentioned traffic task processing method embodiment, the present disclosure also provides a traffic task processing device embodiment. FIG10 shows a schematic structural diagram of a traffic task processing device provided by one embodiment of the present disclosure. As shown in FIG10 , the device includes:
[0203] The second acquisition component 1002 is configured to acquire traffic data to be processed for the target traffic task;
[0204] The second input component 1004 is configured to input the traffic data to be processed into the task processing model and obtain the task processing results output by the task processing model, wherein the task processing model is obtained based on parameter adjustment of the key processing unit in the initial processing model, and the key processing unit is obtained based on the difference between the model processing results of the initial processing model and the reverse model processing results.
[0205] Optionally, the apparatus further includes: a receiving component configured to receive adjustment data sent by a user based on the task processing result, and adjust model parameters of the task processing model according to the adjustment data.
[0206] By applying the solution of the embodiment of the present disclosure, the key processing units in the initial processing model are determined based on the difference between the model processing results and the reverse model processing results, thereby determining the key processing units that determine the model decision-making process, improving the interpretability and transparency of the model, and further improving the interpretability of the task processing results. Moreover, since the task processing model is obtained by adjusting the parameters of the key processing units, there is no need to adjust the unit parameters of all units of the initial processing model, thereby improving the training efficiency of the task processing model.
[0207] The above is a schematic diagram of a traffic task processing device according to this embodiment. It should be noted that the technical solution of the traffic task processing device and the technical solution of the aforementioned traffic task processing method are based on the same concept. For details not described in detail in the technical solution of the traffic task processing device, please refer to the description of the technical solution of the aforementioned traffic task processing method.
[0208] Corresponding to the above-mentioned task processing model training method embodiment, the present disclosure also provides a task processing model training device embodiment. Figure 11 shows a schematic diagram of the structure of a task processing model training device provided by one embodiment of the present disclosure. As shown in Figure 11, the device is applied to a cloud-side device and includes:
[0209] A third acquisition component 1102 is configured to obtain, in response to a model training request for the task processing model, a model processing result of the initial processing model and unit processing results of multiple processing units in the initial processing model;
[0210] a first adjustment component 1104 configured to, for a first processing unit, fix unit processing results of processing units other than the first processing unit among the plurality of processing units, and reversely adjust the unit processing result of the first processing unit to obtain a reverse model processing result output by the initial processing model, wherein the first processing unit is any one of the plurality of processing units;
[0211] A first determining component 1106 is configured to determine a unit weight of a first processing unit based on the model processing result and the reverse model processing result;
[0212] The second determining component 1108 is configured to determine the key processing unit in the initial processing model based on the unit weights of the plurality of processing units;
[0213] The second adjustment component 1110 is configured to adjust the unit parameters of the key processing unit and obtain the trained task processing model.
[0214] By applying the solution of the disclosed embodiments, by reversely adjusting unit processing results and observing the changes in the adjusted model processing results, the importance of processing units in the model decision-making process can be accurately determined, thereby improving the transparency and trustworthiness of the model output and ensuring that the model's decision-making process is reliable and explainable. Furthermore, after determining the key processing units, the unit parameters of the key processing units can be directly adjusted, reducing the amount of parameter adjustment and improving the efficiency of task processing model training.
[0215] The above is a schematic diagram of a task processing model training device according to this embodiment. It should be noted that the technical solution of the task processing model training device and the technical solution of the task processing model training method described above are based on the same concept. For details not described in detail in the technical solution of the task processing model training device, please refer to the description of the technical solution of the task processing model training method described above.
[0216] Figure 12 shows a block diagram of a computing device according to an embodiment of the present disclosure. Components of the computing device 1200 include, but are not limited to, a memory 1210 and a processor 1220. The processor 1220 is connected to the memory 1210 via a bus 1230, and a database 1250 is used to store data.
[0217] The computing device 1200 also includes an access device 1240 that enables the computing device 1200 to communicate via one or more networks 1260. Examples of such networks include a Public Switched Telephone Network (PSTN), a Local Area Network (LAN), a Wide Area Network (WAN), a Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 1240 may include one or more of any type of network interface (e.g., a Network Interface Card (NIC)) whether wired or wireless, such as an IEEE 802.11 Wireless Local Area Networks (WLAN) wireless interface, a World Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and the like.
[0218] In one embodiment of the present disclosure, the aforementioned components of the computing device 1200 and other components not shown in FIG12 may also be connected to each other, for example, via a bus. It should be understood that the computing device structure block diagram shown in FIG12 is for illustrative purposes only and does not limit the scope of the present disclosure. Those skilled in the art may add or replace other components as needed.
[0219] Computing device 1200 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or personal computer (PC). Computing device 1200 may also be a mobile or stationary server.
[0220] Among them, the processor 1220 is used to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the above-mentioned task processing method or traffic task processing method or task processing model training method.
[0221] The above is a schematic diagram of a computing device according to this embodiment. It should be noted that the technical solution of this computing device is based on the same concept as the technical solutions of the aforementioned task processing method, traffic task processing method, and task processing model training method. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solutions of the aforementioned task processing method, traffic task processing method, or task processing model training method.
[0222] The present disclosure also provides a processor. FIG13 is a schematic diagram of the structure of a processor provided by an embodiment of the present disclosure. As shown in FIG13 , the processor 1300 is configured to run a program, wherein the program, when run by the processor, executes the method in the above embodiment.
[0223] In the embodiment of the present disclosure, the processor 1300 may execute the operating program of the method in the embodiment.
[0224] Optionally, the processor 1300 can be configured to perform the following steps: obtaining data to be processed for the target task; inputting the data to be processed into the task processing model, and obtaining the task processing results output by the task processing model, wherein the task processing model is obtained based on parameter adjustment of the key processing units in the initial processing model, and the key processing units are obtained based on the difference between the model processing results of the initial processing model and the reverse model processing results.
[0225] Optionally, the processor 1300 can be configured to perform the following steps: obtaining traffic data to be processed for a target traffic task; inputting the traffic data to be processed into a task processing model, and obtaining a task processing result output by the task processing model, wherein the task processing model is obtained based on parameter adjustment of a key processing unit in the initial processing model, and the key processing unit is obtained based on the difference between the model processing result of the initial processing model and the reverse model processing result.
[0226] Optionally, the processor 1300 can be configured to perform the following steps: in response to a model training request for a task processing model, obtain the model processing results of the initial processing model and the unit processing results of multiple processing units in the initial processing model; for the first processing unit, fix the unit processing results of the processing units other than the first processing unit in the multiple processing units, and reversely adjust the unit processing results of the first processing unit to obtain the reverse model processing results output by the initial processing model, wherein the first processing unit is any one of the multiple processing units; determine the unit weight of the first processing unit based on the model processing results and the reverse model processing results; determine the key processing units in the initial processing model based on the unit weights of the multiple processing units; adjust the unit parameters of the key processing units to obtain a trained task processing model.
[0227] An embodiment of the present disclosure also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the above-mentioned task processing method, traffic task processing method, or task processing model training method.
[0228] The above is a schematic diagram of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium is based on the same concept as the technical solutions of the aforementioned task processing method, traffic task processing method, and task processing model training method. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solutions of the aforementioned task processing method, traffic task processing method, or task processing model training method.
[0229] The embodiments of the present disclosure further provide a computer program product, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method provided in the embodiments of the present disclosure is implemented.
[0230] An embodiment of the present disclosure further provides a computer program, wherein when the computer program is executed in a computer, the computer is caused to execute the steps of the above-mentioned task processing method or traffic task processing method or task processing model training method.
[0231] The above is an illustrative solution of a computer program according to this embodiment. It should be noted that the technical solution of this computer program is based on the same concept as the technical solutions of the aforementioned task processing method, traffic task processing method, and task processing model training method. For details not described in detail in the technical solution of the computer program, please refer to the description of the technical solutions of the aforementioned task processing method, traffic task processing method, or task processing model training method.
[0232] Optionally, the above-mentioned computer program implements the program code of the following steps when executed by the processor: obtaining the data to be processed for the target task; inputting the data to be processed into the task processing model, and obtaining the task processing result output by the task processing model, wherein the task processing model is obtained based on parameter adjustment of the key processing unit in the initial processing model, and the key processing unit is obtained based on the difference between the model processing result of the initial processing model and the reverse model processing result.
[0233] Optionally, the above-mentioned computer program implements the program code of the following steps when executed by the processor: obtaining the traffic data to be processed for the target traffic task; inputting the traffic data to be processed into the task processing model, and obtaining the task processing result output by the task processing model, wherein the task processing model is obtained based on parameter adjustment of the key processing unit in the initial processing model, and the key processing unit is obtained based on the difference between the model processing result of the initial processing model and the reverse model processing result.
[0234] Optionally, the above-mentioned computer program implements the program code of the following steps when executed by the processor: in response to a model training request for the task processing model, obtain the model processing results of the initial processing model and the unit processing results of multiple processing units in the initial processing model; for the first processing unit, fix the unit processing results of the processing units other than the first processing unit in the multiple processing units, and reversely adjust the unit processing results of the first processing unit to obtain the reverse model processing results output by the initial processing model, wherein the first processing unit is any one of the multiple processing units; determine the unit weight of the first processing unit based on the model processing results and the reverse model processing results; determine the key processing units in the initial processing model based on the unit weights of the multiple processing units; adjust the unit parameters of the key processing units to obtain a trained task processing model.
[0235] The foregoing description describes specific embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0236] Computer instructions include computer program code, which may be in source code form, object code form, executable files, or some intermediate form. Computer-readable media may include any entity or device capable of carrying computer program code, recording media, USB flash drives, mobile hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signals, telecommunications signals, and software distribution media. It should be noted that the content of computer-readable media may be appropriately increased or decreased based on the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media does not include electric carrier signals and telecommunications signals.
[0237] It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present disclosure are not limited by the order of the actions described, because according to the embodiments of the present disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and components involved are not necessarily required for the embodiments of the present disclosure.
[0238] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0239] The preferred embodiments of the present disclosure disclosed above are only used to help illustrate the present disclosure. The optional embodiments do not describe all details in detail, nor do they limit the invention to only the specific implementation methods described above. Obviously, many modifications and changes can be made based on the content of the embodiments of the present disclosure. The present disclosure selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of the present disclosure, so that those skilled in the art can better understand and utilize the present disclosure. The present disclosure is limited only by the claims and their full scope and equivalents. Industrial Applicability
[0240] The solution provided by the embodiment of the present disclosure can be applied in the task processing process to obtain the data to be processed for the target task; the data to be processed is input into the task processing model to obtain the task processing results output by the task processing model, wherein the task processing model is obtained based on parameter adjustment of the key processing units in the initial processing model, and the key processing units are obtained based on the difference between the model processing results of the initial processing model and the reverse model processing results, thereby solving the technical problem of poor interpretability of the task processing results.
Claims
1. A task processing method, comprising: Obtaining data to be processed for a target task; Inputting the data to be processed into a task processing model, and obtaining a task processing result output by the task processing model, where the task processing model is obtained by adjusting parameters of key processing units in an initial processing model, and the key processing units are obtained based on the difference between the model processing result and the reverse model processing result of the initial processing model.
2. The method according to claim 1, the method further comprising: Obtaining the model processing result of the initial processing model and the unit processing results of multiple processing units in the initial processing model; For a first processing unit, fixing the unit processing results of the processing units other than the first processing unit in the multiple processing units, and reversely adjusting the unit processing result of the first processing unit to obtain the reverse model processing result output by the initial processing model, where the first processing unit is any one of the multiple processing units; Determining the unit weight of the first processing unit according to the model processing result and the reverse model processing result; Determining the key processing unit in the initial processing model according to the unit weights of the multiple processing units, and adjusting parameters of the key processing unit to obtain the task processing model.
3. The method according to claim 2, determining the unit weight of the first processing unit according to the model processing result and the reverse model processing result, comprising: Analyzing the reverse model processing result to determine a first associated keyword in the reverse model processing result, and analyzing the model processing result to determine a second associated keyword in the model processing result, where both the first associated keyword and the second associated keyword are related to the problem to be inferred; Determining a first weight measurement value of the first processing unit according to the first associated keyword and the reverse model processing result, and determining a second weight measurement value of the first processing unit according to the second associated keyword and the model processing result; Determining the unit weight of the first processing unit according to the first weight measurement value and the second weight measurement value.
4. The method according to claim 2, determining the unit weight of the first processing unit according to the model processing result and the reverse model processing result, comprising: Analyzing the reverse model processing result to determine a first associated keyword in the reverse model processing result, where the first associated keyword is related to the problem to be inferred; Inputting the first associated keyword into the initial processing model to obtain a predicted processing result output by the initial processing model; Determining the unit weight of the first processing unit according to the model processing result and the predicted processing result.
5. The method according to claim 4, determining the unit weight of the first processing unit according to the model processing result and the predicted processing result, comprising: Input the weight generation prompt information, the model processing result, and the prediction processing result into a pre-trained language model to obtain the unit weight of the first processing unit, where the weight generation prompt information is used to guide the pre-trained language model to generate the unit weight of the first processing unit.
6. The method according to claim 2, the method further comprising: Screen out a control processing unit from the plurality of processing units; Fix the unit processing results of the processing units other than the key processing unit among the plurality of processing units, and reversely adjust the unit processing result of the key processing unit to obtain the reverse key processing result output by the initial processing model; Fix the unit processing results of the processing units other than the control processing unit among the plurality of processing units, and reversely adjust the unit processing result of the control processing unit to obtain the reverse control processing result output by the initial processing model; Determine the verification result of the key processing unit according to the reverse key processing result and the reverse control processing result.
7. The method according to claim 2, reversely adjusting the unit processing result of the first processing unit to obtain the reverse model processing result output by the initial processing model, comprising: Obtain counterfactual data; Process the counterfactual data through the first processing unit to obtain the reverse unit processing result of the first processing unit; Replace the unit processing result of the first processing unit with the reverse unit processing result, and perform propagation processing on the reverse unit processing result through the initial processing model to obtain the reverse model processing result.
8. The method according to claim 7, obtaining counterfactual data, comprising: Obtain factual data, where the factual data includes inference data; Adjust the inference data to replacement data independent of the inference data to obtain the counterfactual data.
9. The method according to claim 2, the initial processing model includes the plurality of processing units, residual units, and a multi-layer neural network; Obtaining the model processing result of the initial processing model and the unit processing results of a plurality of processing units in the initial processing model, comprising: Obtain factual data; Process the factual data through the plurality of processing units to obtain the unit processing results of the plurality of processing units; Perform mapping processing on the unit processing results through the residual unit and the multi-layer neural network to obtain the model processing result of the initial processing model.
10. The method according to claim 9, obtaining factual data, comprising: Obtain a question to be inferred; Input the question to be inferred and inference prompt information into a pre-trained language model to obtain inference data corresponding to the question to be inferred; Construct the factual data according to the question to be inferred and the inference data.
11. A traffic task processing method, comprising: Obtain traffic data to be processed for a target traffic task; Input the traffic data to be processed into a task processing model to obtain a task processing result output by the task processing model, where the task processing model is obtained by adjusting parameters of key processing units in an initial processing model, and the key processing units are obtained based on the difference between the model processing result and the reverse model processing result of the initial processing model.
12. According to the method described in claim 11, after inputting the traffic data to be processed into a task processing model to obtain a task processing result output by the task processing model, the method further includes: Receiving adjustment data sent by a user based on the task processing result, and adjusting the model parameters of the task processing model according to the adjustment data.
13. A method for training a task processing model, applied to a cloud-side device, includes: In response to a model training request for a task processing model, obtaining a model processing result of an initial processing model and unit processing results of multiple processing units in the initial processing model; For a first processing unit, fixing the unit processing results of the processing units other than the first processing unit among the multiple processing units, and reversely adjusting the unit processing result of the first processing unit to obtain a reverse model processing result output by the initial processing model, where the first processing unit is any one of the multiple processing units; Determining the unit weight of the first processing unit according to the model processing result and the reverse model processing result; Determining key processing units in the initial processing model according to the unit weights of the multiple processing units; Adjusting the unit parameters of the key processing units to obtain the trained task processing model.
14. A task processing device includes: A first acquisition component configured to acquire data to be processed for a target task; A first input component configured to input the data to be processed into a task processing model to obtain a task processing result output by the task processing model, where the task processing model is obtained by adjusting parameters of key processing units in an initial processing model, and the key processing units are obtained based on the difference between the model processing result and the reverse model processing result of the initial processing model.
15. A traffic task processing device includes: A second acquisition component configured to acquire traffic data to be processed for a target traffic task; A second input component configured to input the traffic data to be processed into a task processing model to obtain a task processing result output by the task processing model, where the task processing model is obtained by adjusting parameters of key processing units in an initial processing model, and the key processing units are obtained based on the difference between the model processing result and the reverse model processing result of the initial processing model.
16. A computer program product, wherein, Including a computer program, which when executed by a processor implements the method according to any one of claims 1 to 10 or any one of claims 11 to 12 or claim 13.
17. A computer program product, wherein, Comprising a non-volatile computer-readable storage medium that stores a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 10 or any one of claims 11 to 12 or claim 13.
18. A computer program, wherein, The computer program, when executed by a processor, implements the method according to any one of claims 1 to 10 or any one of claims 11 to 12 or claim 13.
19. A computing device, comprising: a memory and a processor; The memory is configured to store computer-executable instructions, and the processor is configured to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the method according to any one of claims 1 to 10 or any one of claims 11 to 12 or claim 13.
20. A computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the method according to any one of claims 1 to 10 or any one of claims 11 to 12 or claim 13.
Citation Information
Patent Citations
Neural network prediction method and device
CN109670566A
Task processing method and device, electronic equipment and storage medium
CN116934571A
Task processing method, traffic task processing method and task processing model training method
CN117971420A
Systems and methods for traffic flow prediction
US11238729B1
Cited By
Model training method and text generation method
CN120725092A
Transverse mixed attention mechanism model training method, medium, device and program product
CN121031665A
Road route design evaluation system based on data analysis
CN121167851A
Large language model thinking chain fracture traceability method and device based on anti-fact analysis
CN121457609A