Task information acquisition method and related device
By automatically obtaining data and processing results associated with the target task, and generating and selecting task description information, the high cost and inefficiency problems caused by manual writing in the prior art are solved, and more efficient model training is achieved.
Patent Information
- Application Number
- PCT/CN2024/134174
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-27
- Filing Date
- 2024-11-25
- Publication Date
- 2025-06-05
AI Technical Summary
In the prior art, task description information is usually written manually, resulting in high cost of model training and lack of diversity in description information, which affects the model training effect.
A task information acquisition method is provided, by automatically obtaining data and processing results associated with the target task, generating task description information, including first description information and a plurality of second description information, any second description information may include a placeholder, and finally selecting the optimal third description information as training data.
This method can save human resources, reduce model training costs, and improve model training effect through diversified description information.
Smart Images

Figure CN2024134174_05062025_PF_FP_ABST
Abstract
Description
A method for obtaining task information and related equipment
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on November 27, 2023, with application number 202311604028.1 and invention name “A method for obtaining task information and related equipment”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present invention relates to artificial intelligence (AI) technology, and more particularly to a method for obtaining task information and related equipment. Background Art
[0003] Instruction tuning involves constructing natural language instructions as training data to train a neural network model, thereby obtaining a trained neural network model. Instruction tuning helps the model better understand the content of the instructions and respond appropriately to them.
[0004] In related technologies, in order to obtain instructions as training data, it is often necessary to prepare a task template. A task template can also be understood as the description information of a task. The task description information is usually presented as text used to describe the task, and the text contains placeholders. After obtaining the task description information, the corresponding data can be filled into the placeholders contained in the task description information to form instructions. Then, these instructions can be used to complete model training.
[0005] In the above process, since the description information of the task is usually written manually, it requires a lot of human resources, resulting in excessively high costs for model training. Summary of the Invention
[0006] The embodiments of the present application provide a task information acquisition method and related equipment, which can automatically obtain task description information and use it as training data for model training, thereby saving human resources and reducing the cost of model training.
[0007] A first aspect of an embodiment of the present application provides a method for obtaining task information, the method comprising:
[0008] When training a model that can complete a target task, we first obtain the data associated with the target task and the processing results obtained after processing the data based on the target task. It should be noted that the processing results obtained after processing the data based on the target task can also be understood as the processing results of the data associated with the target task.
[0009] After obtaining the data associated with the target task and the processing result of the data associated with the target task, the data associated with the target task and the processing result of the data associated with the target task can be processed to obtain the first description information of the target task. It should be noted that the first description information of the target task generally does not include placeholders.
[0010] After obtaining the first description information of the target task, the data associated with the target task and the first description information of the target task can be processed to obtain multiple second description information of the target task. It should be noted that any second text of the multiple second texts of the target task may include a placeholder for carrying the data associated with the target task.
[0011] After obtaining multiple pieces of second description information for the target task, one or more pieces of second description information can be selected from these pieces of second description information to serve as the third description information for the target task. Data associated with the target task can then be inserted into the selected third description information, thereby obtaining third description information containing the data associated with the target task. In this way, the third description information containing the data associated with the target task can be used to complete model training, thereby obtaining a model capable of completing the target task.
[0012] The above method demonstrates that it provides a framework for automatically generating task description information. Based on the data associated with a task (i.e., the aforementioned target task) and the results of processing this data, the framework automatically generates multiple descriptions of the task (i.e., multiple second descriptions of the target task). The framework then selects the optimal description of the task (i.e., the third description of the target task) from the multiple descriptions and uses this as training data to complete model training. Because the framework's operation does not involve excessive human intervention, it saves human resources and reduces the cost of model training.
[0013] In one possible implementation, the data includes at least one sub-data and at least one data category corresponding to the at least one sub-data, and any second descriptive information includes at least one placeholder corresponding to the at least one data category, and the at least one placeholder is used to carry the at least one sub-data. In the aforementioned implementation, the data associated with the target task may include at least one sub-data and at least one data category corresponding to the at least one sub-data. Since the multiple second descriptive information of the target task is obtained based on the data associated with the target task and the first descriptive information of the target task, any second descriptive information may include at least one placeholder corresponding to the at least one data category, and the at least one placeholder can be used to respectively carry the at least one sub-data included in the data associated with the target task. It can be seen from this that the multiple second descriptive information of the target task generated in this way can be directly used as candidate templates for the target task, and these candidate templates all include placeholders, so the optimal template (i.e., the third descriptive information) selected from these candidate templates can be inserted into the data associated with the target task as training data for model training.
[0014] In one possible implementation, based on the data and the processing results, generating the first descriptive information of the target task includes: processing the data and the processing results through a first neural network model to obtain the first descriptive information of the target task. In the aforementioned implementation, after obtaining the data associated with the target task and the processing results of the data associated with the target task, the data associated with the target task can be used as the input of the target task, and the processing results of the data associated with the target task can be used as the output of the target task. Then, a first instruction can be constructed based on the input and output of the target task. Then, the first instruction can be processed by the first neural network model to obtain the first descriptive information of the target task. It can be seen from this that the automatic generation framework of the descriptive information of the task provided in the embodiment of the present application, the operation process of the framework can use the neural network model to complete the initial template of the target task (i.e., the aforementioned first descriptive information), and can more comprehensively consider various factors so that the optimal template of the target task finally selected has a certain quality. Therefore, the training data constructed based on the optimal template is conducive to improving the effect of model training.
[0015] In one possible implementation, based on the data and the first description information of the target task, generating multiple second description information of the target task includes: processing the data and the first description information of the target task through a second neural network model to obtain multiple second description information of the target task. In the aforementioned implementation, after obtaining the first description information of the target task, at least one data category contained in the data associated with the target task can be regarded as a keyword of the target task. Then, a second instruction can be constructed based on the first description information of the target task and the keyword of the target task. Then, the second instruction can be input into the second neural network model to process the second instruction through the second neural network model to obtain multiple second description information of the target task. It can be seen from this that the automatic generation framework of the description information of the task provided in the embodiment of the present application, the operation process of the framework can use the neural network model to complete the candidate template of the target task (i.e., the aforementioned second description information), and can consider various factors more comprehensively, so that the optimal template of the target task finally selected has a certain quality, so the training data constructed based on the optimal template is conducive to improving the effect of model training.
[0016] In one possible implementation, the number of third description information is multiple, and selecting the third description information from the multiple second description information includes: clustering the multiple second description information to obtain multiple information categories, one information category in the multiple information categories contains at least one second description information; and selecting multiple third description information from the multiple information categories. In the aforementioned implementation, after obtaining multiple second description information of the target task, a certain clustering algorithm can be used to cluster the multiple second description information of the target task to obtain multiple information categories. It can be understood that, among the multiple information categories, any one information category can contain at least one second description information of the target task. After obtaining multiple information categories, the optimal second description information in each information category can be determined as the third description information of the target task, so that multiple third description information of the target task can be obtained in the end. It can be seen that the automatic generation framework of the description information of the task provided in the embodiment of the present application can select the optimal template of the target task by clustering, so that the selected optimal template has diversity, so the training data constructed based on the optimal template is conducive to further improving the effect of model training.
[0017] In one possible implementation, multiple second description information is presented in text form, and clustering the multiple second description information to obtain multiple information categories includes: converting the multiple second description information presented in text form into multiple second description information presented in vector form; clustering the multiple second description information presented in vector form to obtain multiple information categories. In the aforementioned implementation, the multiple second description information of the target task output by the second neural network model is usually presented in text form, so before clustering, the multiple second description information presented in text form can be calculated to obtain multiple second description information presented in vector form. Then, the multiple second description information presented in vector form can be clustered to obtain multiple information categories. It can be seen that by converting the candidate template in text form into the candidate template in vector form, the clustering efficiency for the candidate template can be improved.
[0018] In one possible implementation, the method further includes: processing multiple pieces of second descriptive information carrying data using a third neural network model to obtain multiple probabilities of processing results; using the multiple probabilities as evaluation values for the multiple pieces of second descriptive information; and selecting multiple pieces of third descriptive information from multiple categories includes: selecting the second descriptive information with the highest evaluation value in any one of the multiple categories as the third descriptive information. In the aforementioned implementation, after obtaining the multiple pieces of second descriptive information for the target task, the data associated with the target task can be inserted into the multiple pieces of second descriptive information to obtain multiple pieces of second descriptive information carrying the data associated with the target task. The multiple pieces of second descriptive information carrying the data associated with the target task can be input into the third neural network model to predict multiple probabilities of processing results for the data associated with the target task. Since these multiple probabilities correspond one-to-one to the multiple pieces of second descriptive information carrying the data associated with the target task, these multiple probabilities can be used as evaluation values for the multiple pieces of second descriptive information for the target task. Thus, after obtaining the multiple categories of information, the second descriptive information with the highest evaluation value in each category can be selected as the third descriptive information for the target task, ultimately obtaining multiple pieces of third descriptive information for the target task. It can be seen that the automatic generation framework of task description information provided in the embodiment of the present application can select the optimal template for the target task through clustering + evaluation, so that the selected optimal template has diversity and accuracy. Therefore, the training data constructed based on the optimal template is conducive to further improving the effect of model training.
[0019] In one possible implementation, the object of model training is the fourth neural network model, and the model that can complete the target task is the fifth neural network model.
[0020] The second aspect of an embodiment of the present application provides a task information acquisition device, which includes: an acquisition module, used to acquire data associated with a target task, and a processing result obtained after processing the data based on the target task; a first generation module, used to generate first description information of the target task based on the data and the processing result; a second generation module, used to generate multiple second description information of the target task based on the data and the first description information of the target task, any one of the second description information includes a placeholder for carrying data; a selection module, used to select third description information from multiple second description information, and the third description information carrying data is used for model training to obtain a model that can complete the target task.
[0021] In one possible implementation, the data includes at least one sub-data and at least one data category corresponding to the at least one sub-data, and any second description information includes at least one placeholder corresponding to the at least one data category, and the at least one placeholder is used to carry the at least one sub-data.
[0022] In one possible implementation, the first generation module is used to process the data and the processing results through a first neural network model to obtain first description information of the target task.
[0023] In one possible implementation, the second generation module is used to process the data and the first description information of the target task through a second neural network model to obtain multiple second description information of the target task.
[0024] In one possible implementation, the selection module is configured to: cluster the plurality of second description information to obtain a plurality of information categories, wherein one information category among the plurality of information categories contains at least one second description information; and select a plurality of third description information from the plurality of information categories.
[0025] In one possible implementation, multiple second description information are presented in text form, and the selection module is used to: convert the multiple second description information presented in text form into multiple second description information presented in vector form; cluster the multiple second description information presented in vector form to obtain multiple information categories.
[0026] In one possible implementation, the device also includes: a processing module for processing multiple second description information carrying data separately through a third neural network model to obtain multiple probabilities of the processing results; an evaluation module for using the multiple probabilities as evaluation values of the multiple second description information; and a selection module for using the second description information with the highest evaluation value as the third description information in any information category of the multiple information categories.
[0027] In one possible implementation, the object of model training is the fourth neural network model, and the model that can complete the target task is the fifth neural network model.
[0028] A third aspect of an embodiment of the present application provides a task information acquisition device, which includes a memory and a processor; the memory stores code, and the processor is configured to execute the code. When the code is executed, the task information acquisition device performs the method described in the first aspect or any possible implementation method of the first aspect.
[0029] A fourth aspect of an embodiment of the present application provides a circuit system, which includes a processing circuit, and the processing circuit is configured to execute the method described in the first aspect or any possible implementation of the first aspect.
[0030] A fifth aspect of an embodiment of the present application provides a chip system, which includes a processor for calling a computer program or computer instructions stored in a memory so that the processor executes the method described in the first aspect or any possible implementation of the first aspect.
[0031] In a possible implementation, the processor is coupled to the memory through an interface.
[0032] In one possible implementation, the chip system further includes a memory, in which a computer program or computer instructions is stored.
[0033] A sixth aspect of the embodiments of the present application provides a computer storage medium storing a computer program, which, when executed by a computer, enables the computer to implement the method described in the first aspect or any possible implementation of the first aspect.
[0034] A seventh aspect of the embodiments of the present application provides a computer program product, which stores instructions. When the instructions are executed by a computer, the computer implements the method described in the first aspect or any possible implementation of the first aspect.
[0035] In the embodiment of the present application, when it is necessary to train a model that can complete the target task, the data associated with the target task and the processing result obtained after the data is processed based on the target task can be obtained first. Then, the data associated with the target task and the processing result of these data can be used to generate the first description information of the target task. Then, the data associated with the target task and the first description information of the target task can be used to generate multiple second description information of the target task, and any second description information includes a placeholder for carrying the data associated with the target task. Finally, the third description information of the target task can be selected from the multiple second description information of the target task, so the third description information carrying the data associated with the target task can be used as training data to complete model training, thereby obtaining a model that can complete the target task. Based on the above process, the embodiment of the present application provides an automatic generation framework for the description information of a task, which can generate multiple description information of the task (i.e., multiple second description information of the aforementioned target task) based on the data associated with a certain task and the processing result of these data, and select the optimal description information of the task (i.e., the third description information of the aforementioned target task) from the multiple description information of the task as training data, thereby completing model training. Since the operation process of this framework does not involve too much manual participation, it can save human resources and thus reduce the cost of model training. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Figure 1 is a schematic diagram of the structure of the artificial intelligence main framework;
[0037] FIG2a is a schematic diagram of the structure of a task information acquisition system provided in an embodiment of the present application;
[0038] FIG2 b is another structural diagram of the task information acquisition system provided in an embodiment of the present application;
[0039] FIG2c is a schematic diagram of a device related to task information acquisition according to an embodiment of the present application;
[0040] FIG3 is a schematic diagram of a system architecture provided in an embodiment of the present application;
[0041] FIG4 is a flow chart of a method for obtaining task information according to an embodiment of the present application;
[0042] FIG5 is a schematic diagram of obtaining task description information according to an embodiment of the present application;
[0043] FIG6 is another schematic diagram of obtaining task description information provided in an embodiment of the present application;
[0044] FIG7 is a schematic diagram of selecting task description information provided in an embodiment of the present application;
[0045] FIG8 is a schematic diagram of the structure of a task information acquisition device provided in an embodiment of the present application;
[0046] FIG9 is a schematic structural diagram of an execution device provided in an embodiment of the present application;
[0047] FIG10 is a schematic diagram of the structure of a training device provided in an embodiment of the present application;
[0048] FIG11 is a schematic structural diagram of a chip provided in an embodiment of the present application. DETAILED DESCRIPTION
[0049] The embodiments of the present application provide a task information acquisition method and related equipment, which can automatically obtain task description information and use it as training data for model training, thereby saving human resources and reducing the cost of model training.
[0050] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequential order. It should be understood that the terms used in this way can be interchangeable under appropriate circumstances, and this is merely a way of distinguishing the objects of the same attributes when describing them in the embodiments of the present application. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, so that the process, method, system, product or equipment comprising a series of units need not be limited to those units, but may include other units that are not clearly listed or inherent to these processes, methods, products or equipment.
[0051] Instruction fine-tuning involves constructing natural language instructions as training data to train the neural network model, thereby obtaining a trained neural network model. Instruction fine-tuning helps the model better understand the content of the instructions and respond appropriately to them.
[0052] In related technologies, in order to obtain instructions as training data, it is often necessary to prepare a task template. A task template can also be understood as the descriptive information of a task. The descriptive information of the task is usually presented as text used to describe the task, and the text contains placeholders. After obtaining the descriptive information of the task, the corresponding data can be filled into the placeholders contained in the descriptive information of the task to form instructions. Then, these instructions can be used to complete model training. For example, the descriptive information of the question-answering task is "Please answer [question] based on [article]." Since a certain data contains the article "XX Enterprise Development History" and the question "In which year was XX Enterprise established?", this data can be inserted into the descriptive information of the question-answering task to form the instruction "Please answer [In which year was XX Enterprise established] based on [XX Enterprise Development History]." Therefore, this instruction can be used to train a neural network model that can complete the question-answering task.
[0053] In the above process, since the description information of the task is usually written manually, the construction of training data requires a lot of human resources (for example, the description information of the task requires staff to spend a lot of time to write and requires a certain number of staff to participate in the writing, etc.), resulting in excessively high cost of model training.
[0054] Furthermore, since the description information of the task is usually written manually, the thinking of the writers is often limited, resulting in a lack of diversity in the description information of the task (for example, the staff cannot write a sufficient amount of description information about a certain task, that is, a sufficient number of templates for the task, etc.), which in turn leads to poor model training results.
[0055] In order to solve the above problems, an embodiment of the present application provides a method for obtaining task information, which can be implemented in combination with artificial intelligence (AI) technology. AI technology is a technical discipline that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence. AI technology obtains the best results by sensing the environment, acquiring knowledge and using knowledge. In other words, artificial intelligence technology is a branch of computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Using artificial intelligence for data processing is a common application of artificial intelligence.
[0056] First, we will describe the overall workflow of an AI system. See Figure 1, which illustrates a schematic diagram of the AI framework. This framework will be explained from two perspectives: the "intelligent information chain" (horizontal axis) and the "IT value chain" (vertical axis). The "intelligent information chain" reflects the entire process from data acquisition to processing. For example, it could be the general process of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, and intelligent execution and output. Throughout this process, data undergoes a condensed journey from "data-information-knowledge-wisdom." The "IT value chain," spanning the underlying infrastructure of human intelligence, information (provided and processed by technology), and the system's industrial ecosystem, reflects the value that AI brings to the information technology industry.
[0057] (1) Infrastructure
[0058] Infrastructure provides computing power for AI systems, enabling communication with the outside world and supporting this through a foundational platform. External communication occurs through sensors; computing power is provided by intelligent chips (CPUs, NPUs, GPUs, ASICs, FPGAs, and other hardware accelerators). The foundational platform includes a distributed computing framework and network-related platform guarantees and support, including cloud storage and computing, and interconnected networks. For example, sensors communicate with the outside world to acquire data, which is then fed into the intelligent chips within the distributed computing system provided by the foundational platform for computation.
[0059] (2) Data
[0060] Data above the infrastructure layer represents data sources for AI. This data includes graphics, images, voice, and text, as well as IoT data from traditional devices. This includes business data from existing systems and sensor data such as force, displacement, liquid level, temperature, and humidity.
[0061] (3) Data processing
[0062] Data processing generally includes data training, machine learning, deep learning, search, reasoning, decision-making, etc.
[0063] Among them, machine learning and deep learning can symbolize and formalize data for intelligent information modeling, extraction, preprocessing, and training.
[0064] Reasoning refers to the process of simulating human intelligent reasoning in computers or intelligent systems, using formalized information to perform machine thinking and solve problems based on reasoning control strategies. Typical functions are search and matching.
[0065] Decision-making refers to the process of making decisions after intelligent information is reasoned, and usually provides functions such as classification, sorting, and prediction.
[0066] (4) General ability
[0067] After the data has undergone the data processing mentioned above, some general capabilities can be further formed based on the results of the data processing, such as algorithms or a general system, for example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.
[0068] (5) Smart products and industry applications
[0069] Smart products and industry applications refer to the products and applications of artificial intelligence systems in various fields. They are the encapsulation of the overall artificial intelligence solution, which productizes intelligent information decision-making and realizes practical application. Its application areas mainly include: smart terminals, smart transportation, smart medical care, autonomous driving, smart cities, etc.
[0070] The following will introduce the hardware devices used to implement the methods provided in the embodiments of the present application.
[0071] Figure 2a is a schematic diagram of the structure of a task information acquisition system provided in an embodiment of the present application. The task information acquisition system includes a user device and a data processing device. The user device includes an intelligent terminal such as a mobile phone, a personal computer, or an information processing center. The user device is the initiator of the task information acquisition request, typically initiated by a user through the user device.
[0072] The aforementioned data processing device can be a device or server with data processing capabilities, such as a cloud server, network server, application server, or management server. The data processing device receives task information acquisition requests from smart terminals via an interactive interface and then processes the task information through a memory device that stores data and a processor for data processing. The memory in the data processing device is a general term that includes local storage and a database that stores historical data. The database can be located on the data processing device or on another network server.
[0073] In the task information acquisition system shown in FIG2a, the user device can receive the user's instructions. For example, the user device can obtain a task input / selected by the user, and then initiate a request to the data processing device so that the data processing device executes the task information acquisition application for the task obtained by the user device, thereby obtaining the description information of the task. For example, the user device can obtain the target task input by the user, and then initiate a request to the data processing device (the request usually contains data associated with the target task, and the processing results obtained after the data is processed based on the target task, etc.). Subsequently, the data processing device can perform a series of processing (for example, information generation and information selection, etc.) on the target task based on the request, thereby obtaining the description information of the target task.
[0074] In FIG. 2 a , the data processing device may execute the task information acquisition method according to an embodiment of the present application.
[0075] Figure 2b is another structural diagram of the task information acquisition system provided in an embodiment of the present application. In Figure 2b, the user device directly serves as a data processing device. The user device can directly obtain input from the user and directly process it by the hardware of the user device itself. The specific process is similar to that of Figure 2a. Please refer to the above description and will not be repeated here.
[0076] In the task information acquisition system shown in Figure 2b, the user device can receive the user's instructions. For example, the user device can obtain the target task input by the user, and then the user device can perform a series of processing (for example, information generation and information selection, etc.) on the target task based on the data associated with the target task and the processing results obtained after the data is processed based on the target task, etc., so as to obtain the description information of the target task.
[0077] In FIG2 b , the user device itself can execute the task information acquisition method of the embodiment of the present application.
[0078] FIG2c is a schematic diagram of related equipment for acquiring task information provided in an embodiment of the present application.
[0079] The user device in the above Figures 2a and 2b can specifically be the local device 301 or the local device 302 in Figure 2c, and the data processing device in Figure 2a can specifically be the execution device 210 in Figure 2c, wherein the data storage system 250 can store the data to be processed of the execution device 210, and the data storage system 250 can be integrated on the execution device 210, or it can be set on the cloud or other network servers.
[0080] The processors in Figures 2a and 2b can perform data training / machine learning / deep learning through a neural network model or other models (for example, a model based on a support vector machine), and use the model finally trained or learned from the data to perform task information acquisition applications for tasks, thereby obtaining corresponding processing results.
[0081] Figure 3 is a schematic diagram of the system 100 architecture provided in an embodiment of the present application. In Figure 3, the execution device 110 is configured with an input / output (I / O) interface 112 for data interaction with an external device. The user can input data to the I / O interface 112 through the client device 140. The input data can include: various tasks to be scheduled, callable resources and other parameters in the embodiment of the present application.
[0082] When the execution device 110 preprocesses the input data, or when the computing module 111 of the execution device 110 performs calculations and other related processing (for example, using symbolic expressions to complete the solution of mixed integer programming equations), the execution device 110 can call the data, code, etc. in the data storage system 150 for the corresponding processing, and can also store the data, instructions, etc. obtained from the corresponding processing in the data storage system 150.
[0083] Finally, the I / O interface 112 returns the processing result to the client device 140 so as to provide it to the user.
[0084] It is worth noting that the training device 120 generates corresponding target models (for example, the first neural network model, the second neural network model, and the third neural network model in the task information acquisition method provided in the embodiment of the present application) / rules based on different training data for a certain goal (for example, obtaining the description information of the target task in the task information acquisition method provided in the embodiment of the present application), and the corresponding target models / rules can be used to achieve the above-mentioned goals. The training data can be stored in the database 130 and come from the training samples collected by the data acquisition device 160.
[0085] In the scenario shown in FIG3 , the user can manually input data, which can be performed through the interface provided by I / O interface 112. Alternatively, client device 140 can automatically send input data to I / O interface 112. If user authorization is required for client device 140 to automatically send input data, the user can set the corresponding permissions in client device 140. The user can view the output of execution device 110 on client device 140, which can be presented in a display, sound, action, or other specific form. Client device 140 can also serve as a data acquisition terminal, collecting input data and output results from I / O interface 112 as new sample data and storing them in database 130. Of course, collection can also be performed without client device 140, with I / O interface 112 directly storing the input data and output results from I / O interface 112 as new sample data in database 130.
[0086] It is worth noting that FIG3 is only a schematic diagram of a system architecture provided by an embodiment of the present application, and the positional relationship between the devices, components, modules, etc. shown in the figure does not constitute any limitation. For example, in FIG3, the data storage system 150 is an external memory relative to the execution device 110. In other cases, the data storage system 150 can also be placed in the execution device 110. As shown in FIG3, the target model can be obtained by training according to the training device 120.
[0087] The present application also provides a chip including a neural network processor (NPU). The chip can be provided in the execution device 110 shown in FIG3 to complete the computational work of the computation module 111. The chip can also be provided in the training device 120 shown in FIG3 to complete the training work of the training device 120 and output the target model / rule.
[0088] The neural network processor (NPU) is a coprocessor mounted on the host central processing unit (CPU), which assigns tasks to it. The core of the NPU is the arithmetic circuit, which is controlled by a controller to extract data from memory (weight memory or input memory) and perform calculations.
[0089] In some implementations, the arithmetic circuit includes multiple processing engines (PEs). In some implementations, the arithmetic circuit is a two-dimensional systolic array. The arithmetic circuit can also be a one-dimensional systolic array or other electronic circuit capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit is a general-purpose matrix processor.
[0090] For example, consider an input matrix A, a weight matrix B, and an output matrix C. The computational circuit retrieves the corresponding data for matrix B from the weight memory and caches it on each PE within the computational circuit. The computational circuit then performs a matrix operation on the matrix A data from the input memory and matrix B, storing the partial or final matrix results in the accumulator.
[0091] The vector computation unit can further process the output of the arithmetic circuit, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. For example, the vector computation unit can be used for network calculations in non-convolutional / non-FC layers in neural networks, such as pooling, batch normalization, and local response normalization.
[0092] In some implementations, the vector computation unit can store the processed output vectors in a unified buffer. For example, the vector computation unit can apply a nonlinear function to the output of the computation circuit, such as a vector of accumulated values, to generate activation values. In some implementations, the vector computation unit generates normalized values, merged values, or both. In some implementations, the processed output vectors can be used as activation inputs to the computation circuit, such as for use in subsequent layers in a neural network.
[0093] The unified memory is used to store input data and output data.
[0094] The weight data is directly transferred from the external memory to the input memory and / or unified memory through the direct memory access controller (DMAC), the weight data in the external memory is stored in the weight memory, and the data in the unified memory is stored in the external memory.
[0095] The bus interface unit (BIU) is used to implement interaction between the main CPU, DMAC and instruction fetch memory through the bus.
[0096] An instruction fetch buffer connected to the controller, used to store instructions used by the controller;
[0097] The controller is used to call the instructions cached in the memory to control the working process of the computing accelerator.
[0098] Generally, unified memory, input memory, weight memory and instruction fetch memory are all on-chip memories, and external memory is memory outside the NPU, which can be double data rate synchronous dynamic random access memory (DDR SDRAM), high bandwidth memory (HBM) or other readable and writable memory.
[0099] Since the embodiments of the present application involve the application of a large number of neural networks, in order to facilitate understanding, the relevant terms and related concepts such as neural networks involved in the embodiments of the present application are first introduced below.
[0100] (1) Neural Network
[0101] A neural network can be composed of neural units. A neural unit can refer to an operation unit with xs and intercept 1 as input. The output of the operation unit can be:
[0102] Where s = 1, 2, ... n, n is a natural number greater than 1, Ws is the weight of xs, and b is the bias of the neural unit. f is the activation function of the neural unit, which is used to introduce nonlinear characteristics into the neural network to convert the input signal in the neural unit into the output signal. The output signal of the activation function can be used as the input of the next convolutional layer. The activation function can be a sigmoid function. A neural network is a network formed by connecting many of the above-mentioned single neural units together, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field. The local receptive field can be an area composed of several neural units.
[0103] The operation of each layer in a neural network can be described by the mathematical expression y = a(Wx + b). From a physical perspective, the operation of each layer in a neural network can be understood as transforming the input space (a set of input vectors) into the output space (i.e., from the row space to the column space of a matrix) through five operations. These operations include: 1. Dimensionality increase / decrease; 2. Scaling / reduction; 3. Rotation; 4. Translation; and 5. "Bending." Operations 1, 2, and 3 are performed by Wx, 4 by +b, and 5 by a(). The word "space" is used here because the objects being classified are not individual things, but rather a class of things, and space refers to the collection of all individuals within that class. W is the weight vector, each value in which represents the weight of a neuron in that layer of the neural network. This vector W determines the spatial transformation from input space to output space described above. That is, the weights W of each layer control how the space is transformed. The goal of training a neural network is to ultimately obtain the weight matrix for all layers of the trained neural network (a weight matrix formed by the vectors W of many layers). Therefore, the training process of a neural network is essentially about learning how to control spatial transformations, and more specifically, about learning the weight matrix.
[0104] Because we want the output of the neural network to be as close as possible to the value we really want to predict, we can compare the current network's predicted value with the desired target value, and then update the weight vector of each layer of the neural network based on the difference between the two (of course, there is usually an initialization process before the first update, which is to pre-configure the parameters for each layer in the neural network). For example, if the network's predicted value is too high, the weight vector is adjusted to make it predict a lower value, and this adjustment is continued until the neural network can predict the desired target value. Therefore, it is necessary to pre-define "how to compare the difference between the predicted value and the target value." This is the loss function or objective function, which is an important equation used to measure the difference between the predicted value and the target value. Among them, taking the loss function as an example, the higher the output value (loss) of the loss function, the greater the difference, so the training of the neural network becomes a process of minimizing this loss as much as possible.
[0105] (2) Backpropagation algorithm
[0106] Neural networks can use the back propagation (BP) algorithm to correct the size of the parameters in the initial neural network model during training, reducing the reconstruction error loss of the neural network model. Specifically, the forward propagation of the input signal to the output generates error loss. This error loss information is then backpropagated to update the parameters in the initial neural network model, thereby converging the error loss. The BP algorithm is a backward propagation movement dominated by error loss, aiming to obtain the optimal parameters of the neural network model, such as the weight matrix.
[0107] It is worth noting that the task information acquisition method provided in the embodiment of the present application is mainly used to obtain the descriptive information of the task. The application of the neural network model may be involved in the process of obtaining the descriptive information of the task (for example, one or more trained neural network models are used to obtain the descriptive information of the task). After obtaining the descriptive information of the task, the descriptive information of the task can be used in the subsequent training of the neural network model (for example, the descriptive information of the task can be used as training data for a neural network model to be trained). In order to further understand the process, the process is further introduced below in conjunction with Figure 4. Figure 4 is a flow chart of the task information acquisition method provided in the embodiment of the present application. As shown in Figure 4, the method includes:
[0108] 401. Acquire data associated with the target task, and a processing result obtained after processing the data based on the target task.
[0109] In this embodiment, when a model needs to be trained to complete a target task, data associated with the target task and the processing results obtained after processing the data based on the target task (hereinafter referred to as the processing results of the data associated with the target task) can be obtained from the data warehouse. It should be noted that the target task, the data associated with the target task, and the processing structure of the data associated with the target task are all stored in the data warehouse.
[0110] Specifically, a data warehouse can be built in the following ways:
[0111] (1) Assume that there are multiple tasks. For any one of the multiple tasks, data associated with the task and a processing result of the data associated with the task may be collected in advance. Furthermore, the data associated with the task may include at least one sub-data and at least one data category corresponding to the at least one sub-data.
[0112] It should be noted that, for the remaining tasks among the multiple tasks, the same operations as those performed on the task can also be performed, so that data associated with the multiple tasks and processing results of the data associated with the multiple tasks can be collected in the end.
[0113] For example, suppose there are N tasks T1, T2, ..., T N (N is an integer greater than or equal to 2), for task T t (t=1,...,N),the data can be collected with T t M associated data (M is an integer greater than or equal to 1), and the processing results of these M data For T t The associated i-th data In terms of It can contain C sub-data and the categories of these C sub-data, that is, (C is an integer greater than or equal to 1), Contains the j-th sub-data and the category of the j-th sub-data.
[0114] T t Can be used for a variety of tasks, accordingly, with T t The associated data is also different. For example, when T t For classification tasks, data associated with the text classification task You can write "Content: XXX donut selfie, the mysterious angle is so beautiful, beauty attracts everything", The processing results It can be "Category: Entertainment News". It contains a sub-data "XXX donut selfie, the mysterious angle is so beautiful, beauty attracts everything", and the category of this sub-data is "content".
[0115] For example, when T t For question-answering tasks, data associated with the question-answering task It can be "Article: XX Enterprise Development History: XX Enterprise is a technology-based enterprise, which was established in...; Question: In which year was XX Enterprise established?" The processing results It can be "Answer: XX company was established in 1987". It contains two sub-data, namely "XX Enterprise Development History: XX Enterprise is a technology-based enterprise, which was established in..." and "In which year was XX Enterprise established?". The categories of these two sub-data are "articles" and "questions", etc.
[0116] (2) After collecting data associated with multiple tasks and processing results of the data associated with multiple tasks, for any task, a dedicated storage area can be set up in the data warehouse, and the data associated with the task and the processing results of the data associated with the task can be stored in the storage area, and the task can be used as an index of the storage area.
[0117] It should be noted that the remaining tasks in the plurality of tasks can also perform the same operations as those performed on the task, so that the data associated with the plurality of tasks and the processing results of the data associated with the plurality of tasks can be stored in the plurality of storage areas of the data warehouse accordingly. In this way, the data warehouse is successfully constructed.
[0118] More specifically, the data associated with the target task and the processing results of the data associated with the target task can be obtained from the data warehouse in the following ways:
[0119] After building the data warehouse, since multiple tasks serve as indexes of multiple storage areas of the data warehouse, if you need to obtain data associated with a target task (which can be understood as a task among multiple tasks) and the processing results of the data associated with the target task, you can directly find the corresponding storage area from the data warehouse based on the target task, and read the data associated with the target task and the processing results of the data associated with the target task from the storage area.
[0120] 402. Generate first description information of the target task based on the data and the processing result.
[0121] After obtaining the data associated with the target task and the processing results of the data associated with the target task, the data associated with the target task and the processing results of the data associated with the target task can be processed to obtain first description information of the target task. It should be noted that the first description information of the target task can be understood as the first text used to describe the target task, and the first text does not contain any placeholders.
[0122] Specifically, the first description information of the target task can be obtained in the following ways:
[0123] After obtaining the data associated with the target task and the processing result of the data associated with the target task, the data associated with the target task can be used as the input of the target task, and the processing result of the data associated with the target task can be used as the output of the target task. Then, a first instruction can be constructed based on the input and output of the target task. The first instruction is used to instruct to generate an initial description of the target task with reference to the input and output of the target task. Then, the first instruction can be input into a first neural network model (which can also be understood as a trained neural network model) to process the first instruction through the first neural network model, thereby obtaining the first description information of the target task.
[0124] Still in the above example, as shown in FIG5 (FIG5 is a schematic diagram of obtaining task description information provided by an embodiment of the present application), it is assumed that T t Template (assuming T tFor classification tasks), read the data from the data warehouse with T t Related data as well as The processing results After that, you can "Content: XXX donut selfie, the mysterious angle is so beautiful, beauty attracts everything" is regarded as the input of the classification task, and "Classification: Entertainment News" is considered the output of the classification task. Next, based on the input and output of the classification task, an instruction can be constructed: "Based on the input and output of the classification task, generate a description of the classification task that is as rich and detailed as possible, and is fluent and natural, and conforms to human writing habits." This instruction can then be input into the trained large language model, causing the model to output the initial description of the classification task d t "Based on the given content, select the answer from multiple options to judge and identify which category of news the content belongs to."
[0125] 403. Generate multiple second description information of the target task based on the data and the first description information of the target task, where any second description information includes a placeholder for carrying data.
[0126] After obtaining the first description information of the target task, the data associated with the target task and the first description information of the target task can be processed to obtain multiple second description information of the target task (which can also be called multiple templates of the target task). It should be noted that the multiple second description information of the target task can be understood as multiple second texts used to describe the target task (the content expressed in the second text and the first text is usually similar, but the expression method is different), and any second text in the multiple second texts can contain a placeholder for carrying data associated with the target task.
[0127] Specifically, the second description information of the target task can be presented in the following ways:
[0128] Based on the foregoing, it can be seen that the data associated with the target task may include at least one sub-data and at least one data category corresponding one-to-one to the at least one sub-data. Therefore, accordingly, among the multiple second description information of the target task obtained based on the data associated with the target task and the first description information of the target task, any one of the second description information may include at least one placeholder corresponding one-to-one to the at least one data category. It is understandable that, among the multiple second description information of the target task, the at least one placeholder contained in any one of the second description information can be used to respectively carry at least one sub-data contained in the data associated with the target task.
[0129] Still as in the above example, due to The sub-data included is "XXX donut selfie, the mysterious angle is so beautiful, beauty attracts everything", the category of this sub-data is "content", and the initial description of the classification task is d t The task is to "select an answer from multiple options based on the given content, so as to determine and identify which category of news the content belongs to." Therefore, the final description of the classification task can be "Given the article [content] and the alternative category options 1, 2, and 3, determine the category to which the article belongs." Among them, [content] is a placeholder contained in the final description of the classification task. This placeholder can be used to fill in the sub-data "XXX donut selfie, the mysterious angle is so beautiful, beauty attracts everything."
[0130] More specifically, the second description information of the target task can be obtained in the following way:
[0131] After obtaining the first description information of the target task, at least one data category contained in the data associated with the target task can be regarded as a keyword of the target task. Then, a second instruction can be constructed based on the first description information of the target task and the keyword of the target task. The second instruction is used to instruct to insert the keyword of the target task into the initial description of the target task to generate a final description of the target task. Then, the second instruction can be input into the second neural network model (which can also be understood as a trained neural network model) to process the second instruction through the second neural network model, thereby obtaining multiple second description information of the target task.
[0132] Still as in the above example, as shown in FIG6 (FIG6 is another schematic diagram of obtaining task description information provided by an embodiment of the present application), d t After that, you can The category "content" of the contained sub-data is regarded as the keyword of the classification task. t And the keyword construction instruction for the classification task: "Based on the keywords of the classification task, rewrite the initial description of the classification task to make it more diverse and in line with people's speaking and writing style, while retaining the keywords of the classification task." Then, this instruction can be input into the trained large language model, so that the model outputs multiple final descriptions of the classification task. (that is, multiple templates for classification tasks), where the final description For "Based on the article [content] and the given options 1, 2, and 3, please judge which one best represents the concept or event described in the article", the final description For "Given the article [content] and the alternative categories 1, 2, and 3, please determine the category to which the article belongs." For "Based on the given article [content], please correctly judge which category the article belongs to from the alternative answer options 1, 2 and 3" and For example, "Please select the topic that is most relevant to the article [content] from the following options 1, 2, and 3."
[0133] 404. Select third description information from the plurality of second description information, and use the third description information carrying the data for model training to obtain a model that can complete the target task.
[0134] After obtaining multiple second description information of the target task, one or more second description information can be selected from the multiple second description information of the target task as the third description information of the target task. Then, the data associated with the target task can be inserted into the selected third description information, thereby obtaining the third description information carrying the data associated with the target task (which can also be understood as a third instruction, the third instruction is used to instruct the data associated with the target task to be processed based on the target task). In this way, the third description information carrying the data associated with the target task can be used to complete model training, thereby obtaining a model that can complete the target task.
[0135] Specifically, the third description information of the target task can be obtained in the following ways:
[0136] (1) After obtaining the plurality of second description information of the target task, a clustering algorithm (e.g., a K-means algorithm, etc.) may be used to cluster the plurality of second description information of the target task, thereby obtaining a plurality of information categories. It is understood that any one of the plurality of information categories may contain at least one second description information of the target task.
[0137] Generally, the multiple second description information of the target task output by the second neural network model is usually presented in text form. Therefore, in order to improve the efficiency of clustering, before clustering, the multiple second description information presented in text form can be calculated to obtain multiple second description information presented in vector (embedding) form. Then, the multiple second description information presented in vector form can be clustered to obtain multiple information categories.
[0138] As in the above example, we get multiple final descriptions of the classification task Afterwards, due to It is a multiple final description in text form. You can first convert the multiple final descriptions in text form into a final description in vector form. And use K-means algorithm to Clustering is performed to obtain P categories (P is an integer greater than or equal to 2), each of which contains at least one final description in vector form.
[0139] (2) After obtaining the multiple information categories, for any one of the multiple information categories, the optimal second description information can be determined as the third description information of the target task from the at least one second description information contained in the information category. The same operation as that performed on the remaining information categories can also be performed on the information category, so that multiple third description information of the target task can be obtained.
[0140] More specifically, multiple third description information of the target task can be obtained from multiple information categories in the following manner:
[0141] (1) After obtaining multiple second description information of the target task, the data associated with the target task can be inserted into the multiple second description information to obtain multiple second description information carrying the data associated with the target task (which can also be understood as multiple fourth instructions, and these multiple fourth instructions are all used to instruct the data associated with the target task to be processed based on the target task). For any second description information carrying the data associated with the target task, it can be input into the third neural network model (the trained neural network model) to process the second description information carrying the data associated with the target task through the third neural network model, thereby predicting a probability of obtaining a processing result of the data associated with the target task.
[0142] It should be noted that similar operations can also be performed on the remaining second descriptive information carrying data associated with the target task, so that multiple probabilities of the processing results of the data associated with the target task can be obtained in the end. Since these multiple probabilities correspond one-to-one to the multiple second descriptive information carrying data associated with the target task, these multiple probabilities can be used as evaluation values of the multiple second descriptive information of the target task.
[0143] Still as in the above example, as shown in FIG7 (FIG7 is a schematic diagram of selecting task description information provided by an embodiment of the present application), we obtain After that, you can Fill in In order to obtain the instruction set Any instruction in the instruction set can be Input into the large language model to predict Probability Can be used directly as the final description In this way, we can finally get All evaluation values
[0144] (2) After obtaining the multiple information categories, for any one of the multiple information categories, the second description information with the highest evaluation value among the at least one second description information contained in the information category can be used as the third description information of the target task. The same operation as that performed on the remaining information categories can also be performed on the information category, so that multiple third description information of the target task can be obtained.
[0145] As in the above example, after obtaining P categories, the final description with the highest evaluation value in any category can be used to determine the final description, so the classification task T can be obtained. t The P available final descriptions, that is, T t There are P available templates.
[0146] It should be understood that in this embodiment, in the examples shown in FIG. 5 to FIG. 7 , only the data corresponding to T is read from the data warehouse. t Related data This is just a schematic introduction and does not limit the amount of data to be read. In actual applications, data related to T t Some related data Etc., in the subsequent process, the processing of multiple data is similar to that of single data, which will not be repeated here.
[0147] It should also be understood that in this embodiment, the first neural network model, the second neural network model and the third neural model can be the same model. Of course, the first neural network model, the second neural network model and the third neural model can also be different models. For example, in the examples shown in Figures 5 to 7, all three are the same large language model.
[0148] It should also be understood that after the third descriptive information of the target task is selected, the third descriptive information of the target task can be used for model training. The object of the model training is the fourth neural network model (the neural network model to be trained), and the result of the model training is the fifth neural network model (the trained neural network model, derived from the fourth neural network model, that is, the model that can complete the target task). The fifth neural network model and the first neural network model are usually different models. Similarly, the fifth neural network model and the second neural network model are also usually different models, and the fifth neural network model and the third neural network model are also usually different models.
[0149] In the embodiment of the present application, when it is necessary to train a model that can complete the target task, the data associated with the target task and the processing result obtained after the data is processed based on the target task can be obtained first. Then, the data associated with the target task and the processing result of these data can be used to generate the first description information of the target task. Then, the data associated with the target task and the first description information of the target task can be used to generate multiple second description information of the target task, and any second description information includes a placeholder for carrying the data associated with the target task. Finally, the third description information of the target task can be selected from the multiple second description information of the target task, so the third description information carrying the data associated with the target task can be used as training data to complete model training, thereby obtaining a model that can complete the target task. Based on the above process, the embodiment of the present application provides an automatic generation framework for the description information of a task, which can generate multiple description information of the task (i.e., multiple second description information of the aforementioned target task) based on the data associated with a certain task and the processing result of these data, and select the optimal description information of the task (i.e., the third description information of the aforementioned target task) from the multiple description information of the task as training data, thereby completing model training. Since the operation process of this framework does not involve too much manual participation, it can save human resources and thus reduce the cost of model training.
[0150] Furthermore, the embodiment of the present application provides an automatic generation framework for task description information. During its operation, the framework can use a neural network model to complete the generation of task description information, and can use two-stage processing (clustering and evaluation) to complete the selection of task description information. It can consider various factors more comprehensively, so that the description information finally selected has accuracy and diversity, which is conducive to improving the effect of model training.
[0151] The above is a detailed description of the task information acquisition method provided by the embodiment of the present application. The following is an introduction to the task information acquisition device provided by the embodiment of the present application. Figure 8 is a structural diagram of the task information acquisition device provided by the embodiment of the present application. As shown in Figure 8, the device includes:
[0152] The acquisition module 801 is used to acquire data associated with the target task and the processing results obtained after processing the data based on the target task.
[0153] The first generating module 802 is configured to generate first description information of the target task based on the data and the processing result.
[0154] The second generating module 803 is configured to generate a plurality of second description information of the target task based on the data and the first description information of the target task, wherein any one of the second description information includes a placeholder for carrying data.
[0155] The selection module 804 is used to select third description information from multiple second description information. The third description information carrying data is used for model training to obtain a model that can complete the target task.
[0156] In the embodiment of the present application, when it is necessary to train a model that can complete the target task, the data associated with the target task and the processing result obtained after the data is processed based on the target task can be obtained first. Then, the data associated with the target task and the processing result of these data can be used to generate the first description information of the target task. Then, the data associated with the target task and the first description information of the target task can be used to generate multiple second description information of the target task, and any second description information includes a placeholder for carrying the data associated with the target task. Finally, the third description information of the target task can be selected from the multiple second description information of the target task, so the third description information carrying the data associated with the target task can be used as training data to complete model training, thereby obtaining a model that can complete the target task. Based on the above process, the embodiment of the present application provides an automatic generation framework for the description information of a task, which can generate multiple description information of the task (i.e., multiple second description information of the aforementioned target task) based on the data associated with a certain task and the processing result of these data, and select the optimal description information of the task (i.e., the third description information of the aforementioned target task) from the multiple description information of the task as training data, thereby completing model training. Since the operation process of this framework does not involve too much manual participation, it can save human resources and thus reduce the cost of model training.
[0157] In one possible implementation, the data includes at least one sub-data and at least one data category corresponding to the at least one sub-data, and any second description information includes at least one placeholder corresponding to the at least one data category, and the at least one placeholder is used to carry the at least one sub-data.
[0158] In one possible implementation, the first generation module 802 is used to process the data and the processing results through a first neural network model to obtain first description information of the target task.
[0159] In one possible implementation, the second generation module 803 is used to process the data and the first description information of the target task through a second neural network model to obtain multiple second description information of the target task.
[0160] In one possible implementation, the selection module 804 is configured to: cluster the plurality of second description information to obtain a plurality of information categories, wherein one information category of the plurality of information categories contains at least one second description information; and select a plurality of third description information from the plurality of information categories.
[0161] In one possible implementation, multiple second description information are presented in text form, and the selection module 804 is used to: convert the multiple second description information presented in text form into multiple second description information presented in vector form; cluster the multiple second description information presented in vector form to obtain multiple information categories.
[0162] In one possible implementation, the device also includes: a processing module, which is used to process multiple second description information carrying data separately through a third neural network model to obtain multiple probabilities of the processing results; an evaluation module, which is used to use the multiple probabilities as evaluation values of the multiple second description information; and a selection module 804, which is used to use the second description information with the highest evaluation value as the third description information in any information category of the multiple information categories.
[0163] In one possible implementation, the object of model training is the fourth neural network model, and the model that can complete the target task is the fifth neural network model.
[0164] It should be noted that the information interaction, execution process, etc. between the modules / units of the above-mentioned device are based on the same concept as the method embodiment of the present application, and the technical effects they bring are the same as those of the method embodiment of the present application. For specific contents, please refer to the description in the method embodiment shown above in the embodiment of the present application, and no further details will be given here.
[0165] The embodiment of the present application also relates to an execution device, and Figure 9 is a structural diagram of the execution device provided by the embodiment of the present application. As shown in Figure 9, the execution device 900 can be specifically manifested as a mobile phone, a tablet, a laptop computer, a smart wearable device, a server, etc., which are not limited here. Among them, the task information acquisition device described in the embodiment corresponding to Figure 8 can be deployed on the execution device 900 to implement the function of task information acquisition in the embodiment corresponding to Figure 4. Specifically, the execution device 900 includes: a receiver 901, a transmitter 902, a processor 903 and a memory 904 (wherein the number of processors 903 in the execution device 900 can be one or more, and Figure 9 takes one processor as an example), wherein the processor 903 may include an application processor 9031 and a communication processor 9032. In some embodiments of the present application, the receiver 901, the transmitter 902, the processor 903 and the memory 904 may be connected via a bus or other means.
[0166] The memory 904 may include a read-only memory and a random access memory, and provides instructions and data to the processor 903. A portion of the memory 904 may also include non-volatile random access memory (NVRAM). The memory 904 stores processor and operation instructions, executable modules, or data structures, or subsets or extended sets thereof. The operation instructions may include various operation instructions for implementing various operations.
[0167] Processor 903 controls the operation of the execution device. In specific applications, the various components of the execution device are coupled together via a bus system. In addition to a data bus, the bus system may also include a power bus, a control bus, and a status signal bus. However, for clarity, all bus systems are referred to as a bus system in the figure.
[0168] The methods disclosed in the above embodiments of the present application can be applied to or implemented by the processor 903. The processor 903 can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits in the processor 903 or by software instructions. The above processor 903 can be a general-purpose processor, a digital signal processor (DSP), a microprocessor, or a microcontroller, and can further include an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The processor 903 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of the present application can be directly implemented as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in memory 904, and processor 903 reads information in memory 904 and, in conjunction with its hardware, completes the steps of the above method.
[0169] Receiver 901 can be used to receive input digital or character information and generate signal input related to executing device-related settings and function control. Transmitter 902 can be used to output digital or character information through the first interface. Transmitter 902 can also be used to send instructions to the disk pack through the first interface to modify data in the disk pack. Transmitter 902 can also include a display device such as a display screen.
[0170] In an embodiment of the present application, in one case, the processor 903 is used to obtain the optimal description information of the target task (i.e., the aforementioned third description information) through the first neural network model, the second neural network model, and the third neural network model in the embodiment corresponding to Figure 4.
[0171] The present application also relates to a training device. FIG10 is a schematic diagram of the structure of the training device provided by the present application. As shown in FIG10 , the training device 1000 is implemented by one or more servers. The training device 1000 may vary significantly due to different configurations or performance. It may include one or more central processing units (CPUs) 1010 (e.g., one or more processors) and memory 1032, and one or more storage media 1030 (e.g., one or more mass storage devices) storing application programs 1042 or data 1044. The memory 1032 and storage medium 1030 may be temporary storage or permanent storage. The program stored in the storage medium 1030 may include one or more modules (not shown), each module may include a series of instruction operations in the training device. Furthermore, the CPU 1010 may be configured to communicate with the storage medium 1030 to execute the series of instruction operations in the storage medium 1030 on the training device 1000.
[0172] The training device 1000 may also include one or more power supplies 1026, one or more wired or wireless network interfaces 1050, one or more input and output interfaces 1058; or, one or more operating systems 1041, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.
[0173] Specifically, the training device can receive the optimal description information of the target task sent by the execution device shown in Figure 9, and use the optimal description information of the target task as training data to train the fourth neural network model using the training data, thereby obtaining a fifth neural network model that can complete the target task.
[0174] An embodiment of the present application also relates to a computer storage medium, which stores a program for signal processing. When the computer storage medium is run on a computer, it enables the computer to execute the steps executed by the aforementioned execution device, or enables the computer to execute the steps executed by the aforementioned training device.
[0175] An embodiment of the present application also relates to a computer program product, which stores instructions that, when executed by a computer, enable the computer to execute the steps executed by the aforementioned execution device, or enable the computer to execute the steps executed by the aforementioned training device.
[0176] The execution device, training device or terminal device provided in the embodiments of the present application can specifically be a chip, and the chip includes: a processing unit and a communication unit, the processing unit can be, for example, a processor, and the communication unit can be, for example, an input / output interface, a pin or a circuit, etc. The processing unit can execute the computer execution instructions stored in the storage unit, so that the chip in the execution device executes the data processing method described in the above embodiment, or so that the chip in the training device executes the data processing method described in the above embodiment. Optionally, the storage unit is a storage unit in the chip, such as a register, a cache, etc. The storage unit can also be a storage unit located outside the chip in the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.
[0177] Specifically, see Figure 11, which is a schematic diagram of the structure of a chip provided in an embodiment of the present application. The chip can be represented as a neural network processor NPU 1100. NPU 1100 is mounted on the host CPU as a coprocessor and assigned tasks by the host CPU. The core of the NPU is arithmetic circuit 1103, which is controlled by controller 1104 to extract matrix data from memory and perform multiplication operations.
[0178] In some implementations, the arithmetic circuit 1103 includes multiple processing units (PEs). In some implementations, the arithmetic circuit 1103 is a two-dimensional systolic array. The arithmetic circuit 1103 may also be a one-dimensional systolic array or other electronic circuitry capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 1103 is a general-purpose matrix processor.
[0179] For example, assume there are input matrix A, weight matrix B, and output matrix C. The arithmetic circuit retrieves the corresponding data of matrix B from weight memory 1102 and caches it on each PE in the arithmetic circuit. The arithmetic circuit retrieves the data of matrix A from input memory 1101 and performs a matrix operation on matrix B. The partial or final matrix result is stored in accumulator 1108.
[0180] Unified memory 1106 is used to store input and output data. Weight data is directly transferred to weight memory 1102 through the Direct Memory Access Controller (DMAC) 1105. Input data is also transferred to unified memory 1106 through the DMAC.
[0181] BIU stands for Bus Interface Unit, i.e., bus interface unit 1111 , which is used for interaction between the AXI bus, DMAC, and instruction fetch buffer (IFB) 1109 .
[0182] The bus interface unit 1111 (BIU) is used for the instruction fetch memory 1109 to obtain instructions from the external memory, and is also used for the storage unit access controller 1105 to obtain the original data of the input matrix A or the weight matrix B from the external memory.
[0183] DMAC is mainly used to move input data in the external memory DDR to the unified memory 1106 or to move weight data to the weight memory 1102 or to move input data to the input memory 1101.
[0184] The vector calculation unit 1107 includes multiple operation processing units. When necessary, it further processes the output of the operation circuit 1103, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. It is mainly used for non-convolutional / fully connected layer network calculations in neural networks, such as batch normalization, pixel-level summation, and upsampling of the predicted label plane.
[0185] In some implementations, the vector calculation unit 1107 can store the processed output vector to the unified memory 1106. For example, the vector calculation unit 1107 can apply a linear function or a nonlinear function to the output of the operation circuit 1103, such as linear interpolation of the predicted label plane extracted by the convolution layer, or another example of a vector of accumulated values to generate an activation value. In some implementations, the vector calculation unit 1107 generates a normalized value, a pixel-level summed value, or both. In some implementations, the processed output vector can be used as an activation input to the operation circuit 1103, for example, for use in subsequent layers in a neural network.
[0186] An instruction fetch buffer 1109 connected to the controller 1104 is used to store instructions used by the controller 1104;
[0187] Unified memory 1106, input memory 1101, weight memory 1102, and instruction fetch memory 1109 are all on-chip memories. External memories are private to the NPU hardware architecture.
[0188] The processor mentioned in any of the above places can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the above program.
[0189] It should also be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines.
[0190] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course can also be implemented by special hardware including application-specific integrated circuits, special CPUs, special memories, special components, etc. In general, all functions performed by computer programs can be easily implemented with corresponding hardware, and the specific hardware structures used to implement the same function can also be diverse, such as analog circuits, digital circuits or special circuits, etc. However, for the present application, software program implementation is a better implementation method in most cases. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., and includes a number of instructions to enable a computer device (which can be a personal computer, training equipment, or network equipment, etc.) to execute the methods described in each embodiment of the present application.
[0191] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.
[0192] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, a computer, a training device or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website, a computer, a training device or a data center. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. that includes one or more available media integrations. The available medium can be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).
Claims
1. A method for obtaining task information, characterized in that: The method comprises: Acquire data associated with a target task, and a processing result obtained after processing the data based on the target task; Based on the data and the processing result, generating first description information of the target task; Based on the data and the first description information of the target task, generate a plurality of second description information of the target task, any one of the second description information includes a placeholder for carrying the data; The third description information is selected from the multiple second description information, and the third description information carrying the data is used for model training to obtain a model that can complete the target task.
2. The method according to claim 1, characterized in that The data includes at least one sub-data and at least one data category corresponding to the at least one sub-data, and any one of the second description information includes at least one placeholder corresponding to the at least one data category, and the at least one placeholder is used to carry the at least one sub-data.
3. The method according to claim 1 or 2, characterized in that: The generating the first description information of the target task based on the data and the processing result includes: The data and the processing results are processed by a first neural network model to obtain first description information of the target task.
4. The method according to any one of claims 1 to 3, characterized in that: The generating a plurality of second description information of the target task based on the data and the first description information of the target task comprises: The data and the first description information of the target task are processed by a second neural network model to obtain multiple second description information of the target task.
5. The method according to any one of claims 1 to 4, characterized in that: The number of the third description information is multiple, and the selecting the third description information from the multiple second description information includes: Clustering the plurality of second description information to obtain a plurality of information categories, wherein one information category of the plurality of information categories includes at least one second description information; A plurality of third descriptive information is selected from the plurality of information categories.
6. The method according to claim 5, characterized in that The plurality of second description information are presented in text form, and the clustering of the plurality of second description information to obtain the plurality of information categories includes: Converting the plurality of second description information presented in text form into the plurality of second description information presented in vector form; Clustering is performed on the plurality of second description information presented in the form of vectors to obtain a plurality of information categories.
7. The method according to claim 5 or 6, characterized in that: The method further comprises: Processing the plurality of second description information carrying the data respectively through a third neural network model to obtain a plurality of probabilities of the processing results; Using the multiple probabilities as evaluation values of the multiple second description information; The selecting a plurality of third description information from the plurality of categories comprises: In any one of the plurality of information categories, the second description information having the highest evaluation value is used as the third description information.
8. The method according to any one of claims 1 to 7, characterized in that: The object of the model training is the fourth neural network model, and the model that can complete the target task is the fifth neural network model.
9. A task information acquisition device, characterized in that: The device comprises: An acquisition module, used for acquiring data associated with a target task, and a processing result obtained after processing the data based on the target task; A first generating module, used for generating first description information of the target task based on the data and the processing result; A second generating module, configured to generate a plurality of second description information of the target task based on the data and the first description information of the target task, wherein any one of the second description information includes a placeholder for carrying the data; A selection module is used to select third description information from the multiple second description information, and the third description information carrying the data is used for model training to obtain a model that can complete the target task.
10. The device according to claim 9, characterized in that The data includes at least one sub-data and at least one data category corresponding to the at least one sub-data, and any one of the second description information includes at least one placeholder corresponding to the at least one data category, and the at least one placeholder is used to carry the at least one sub-data.
11. The device according to claim 9 or 10, characterized in that The first generating module is used to process the data and the processing results through a first neural network model to obtain first description information of the target task.
12. The device according to any one of claims 9 to 11, characterized in that The second generating module is used to process the data and the first description information of the target task through a second neural network model to obtain multiple second description information of the target task.
13. The device according to any one of claims 9 to 12, characterized in that The selection module is used to: Clustering the plurality of second description information to obtain a plurality of information categories, wherein one information category of the plurality of information categories includes at least one second description information; A plurality of third descriptive information is selected from the plurality of information categories.
14. The device according to claim 13, characterized in that The plurality of second description information are presented in text form, and the selection module is used to: Converting the plurality of second description information presented in text form into the plurality of second description information presented in vector form; Clustering is performed on the plurality of second description information presented in the form of vectors to obtain a plurality of information categories.
15. The device according to claim 13 or 14, characterized in that The device also includes: A processing module, used for processing the plurality of second description information carrying the data respectively through a third neural network model to obtain a plurality of probabilities of the processing results; An evaluation module, configured to use the multiple probabilities as evaluation values of the multiple second description information; The selection module is used to select the second description information with the highest evaluation value in any one of the multiple information categories as the third description information.
16. The device according to any one of claims 9 to 15, characterized in that The object of the model training is the fourth neural network model, and the model that can complete the target task is the fifth neural network model.
17. A task information acquisition device, characterized in that: The device comprises a memory and a processor; the memory stores codes, and the processor is configured to execute the codes. When the codes are executed, the task information acquisition device executes the method according to any one of claims 1 to 8.
18. A computer storage medium, characterized in that: The computer storage medium stores one or more instructions, which, when executed by one or more computers, enable the one or more computers to implement the method of any one of claims 1 to 8.
19. A computer program product, characterized in that The computer program product stores instructions, which, when executed by a computer, enable the computer to implement the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Machine learning method and related device
CN111062495A
Model training method, model prediction method, and model control system
CN112799850A
Data processing method and system
CN113568735A
Task information acquisition method and related equipment
CN117852603A
Neural network construction method and apparatus
WO2022083536A1