Model training method and computing device
By embedding hidden state layer models in computing devices, the big model is trained in real time to match user portraits, solving the problem that big model training and use cannot be performed simultaneously, and improving user experience and model adaptability.
Patent Information
- Application Number
- CN202510354287.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-08-08
AI Technical Summary
In the prior art, large models cannot be performed simultaneously during training and use, resulting in the inability to use new data for training in time, affecting the user experience and the model's ability to adapt to new situations.
Embed the hidden state layer model in the computing device, predict user satisfaction through inference tasks and results, and train the hidden state layer model in real time to match the user portrait, real-time training of the inference model is achieved.
It realizes that the large model is carried out simultaneously while training and use, enhances the model's adaptability and user experience, and improves training efficiency and accuracy.
Smart Images

Figure CN120449962A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a model training method and computing device. Background Art
[0002] With the continuous development of big model technology, intelligent services based on big models, such as speech synthesis and image processing, are also gaining more and more attention. These intelligent services provide users with a convenient and efficient experience through big model technology.
[0003] Currently, it's impossible to train and use a large model for inference simultaneously. Therefore, when optimizing a large model (for example, adding functionality or updating its data), you need to pause inference and resume using the model after optimization is complete.
[0004] Based on the above approach, when using large models for inference, new data cannot be used for training in a timely manner. The model can only infer based on old data and cannot adapt to new situations. In addition, during the training process of the large model, the inference service of the large model is interrupted, which affects the user experience. Summary of the Invention
[0005] The embodiments of the present application provide a model training method and a computing device, which can perform real-time training on the inference model during the operation of the inference model in the computing device, thereby effectively enhancing the user experience.
[0006] To achieve the above objectives, the embodiments of the present application adopt the following technical solutions:
[0007] In a first aspect, a model training method is provided for use on a computing device. The method comprises: obtaining an inference task corresponding to a user profile and an inference result output by an inference model for the inference task. Based on the inference task and the inference result, an estimated satisfaction level is determined. During the execution of the inference model, a hidden state layer model is trained based on the inference task, the inference result, and the estimated satisfaction level.
[0008] Among them, the inference model runs on a computing device, and a hidden state layer model is embedded in the hidden layer of the inference model. The hidden state layer model is used to modify the parameters of the hidden layer during the operation of the inference model so that the inference results output by the inference model match the user portrait. The expected satisfaction refers to the user's satisfaction with the inference results predicted by the computing device.
[0009] The above technical solution enables real-time training of the hidden state layer model embedded in the inference model while the inference model is running on a computing device, thereby achieving real-time training of the inference model. This means that large models can be trained and used simultaneously, which not only enhances the adaptability of the inference model but also effectively improves the user experience.
[0010] In an optional embodiment, a prediction model may be deployed in the computing device. Based on this, determining the expected satisfaction level based on the reasoning task and the reasoning result may specifically include: inputting the reasoning task and the reasoning result into the prediction model to obtain the expected satisfaction level.
[0011] Through the above technical solution, the user's satisfaction with the inference results, that is, the expected satisfaction, can be directly predicted through the prediction model. In this way, the efficiency of determining the expected satisfaction can be effectively improved, thereby improving the efficiency of real-time training of the hidden state layer model.
[0012] In an optional embodiment, the method may further include obtaining a plurality of task samples, a result sample corresponding to each task sample, and a user's satisfaction with each result sample. For each task sample, the pre-trained model is trained using the task sample, the result sample corresponding to the task sample, and the user's satisfaction with the result sample to obtain a prediction model.
[0013] Among them, the user's satisfaction with each result sample is used to reflect the degree of matching between the result sample and the user portrait.
[0014] The above technical solution is a process of training the pre-trained model through task samples to obtain a prediction model. In this way, the prediction model can learn the user portrait based on the task samples, so that accurate estimated satisfaction can be obtained based on the user portrait prediction in the future, thereby effectively improving the accuracy of training the hidden state layer model.
[0015] In an optional embodiment, the above-mentioned training of the hidden state layer model based on the reasoning task, reasoning result and expected satisfaction can specifically include: feeding back the reasoning task, reasoning result and expected satisfaction to the hidden layer of the reasoning model to adjust the parameters of the hidden state layer model based on the reasoning task, reasoning result and expected satisfaction.
[0016] The above technical solution shows that the reasoning task, reasoning results, and expected satisfaction can be directly fed back to the hidden layer of the reasoning model to adjust the parameters of the hidden state layer model. This can speed up the training efficiency of the hidden state layer model.
[0017] In an optional embodiment, the above-mentioned training of the hidden state layer model based on the reasoning task, the reasoning result and the expected satisfaction may specifically include: when the expected satisfaction is greater than a preset satisfaction threshold, the hidden state layer model is trained based on the reasoning task, the reasoning result and the expected satisfaction.
[0018] Through the above technical solution, it can be seen that the hidden state layer model is only trained when the expected satisfaction is greater than the preset satisfaction threshold. In this way, the computing power resources of the computing device can be effectively saved.
[0019] In an optional embodiment, before training the hidden state layer model based on the reasoning tasks, reasoning results, and predicted satisfaction, the method may further include obtaining user satisfaction and determining a comprehensive satisfaction score based on the user satisfaction score and the predicted satisfaction score. Based on this, the training of the hidden state layer model based on the reasoning tasks, reasoning results, and predicted satisfaction scores may specifically include training the hidden state layer model based on the reasoning tasks, reasoning results, and comprehensive satisfaction scores.
[0020] Among them, user satisfaction refers to the user's feedback on the satisfaction with the inference results.
[0021] Through the above technical solution, the hidden state layer model is trained by comprehensively considering the satisfaction of user feedback (i.e., user satisfaction) and the expected satisfaction predicted by the computing device. This can enable the hidden state layer model to more accurately capture the user portrait and improve the accuracy of the hidden state layer model in the reasoning process. In this way, users can obtain results that are more in line with their expectations when using the hidden state layer model, thereby improving user experience and satisfaction.
[0022] In an optional embodiment, the above-mentioned determination of comprehensive satisfaction based on user satisfaction and expected satisfaction may specifically include: obtaining a first weight and a second weight, and determining comprehensive satisfaction based on the first weight, the second weight, user satisfaction, and expected satisfaction.
[0023] The first weight refers to the weight corresponding to user satisfaction, and the second weight refers to the weight corresponding to expected satisfaction.
[0024] The above technical solution describes a method for determining comprehensive satisfaction based on user satisfaction and expected satisfaction, which can effectively improve the feasibility of this application.
[0025] In an optional implementation, the first weight is greater than the second weight.
[0026] In the above technical solution, the first weight being greater than the second weight can result in a greater emphasis on user satisfaction when training the hidden state layer model based on user satisfaction and predicted satisfaction. This allows the hidden state layer model to more accurately capture user profiles, further improving the accuracy of the hidden state layer model during inference.
[0027] In an optional embodiment, the above-mentioned determination of comprehensive satisfaction based on user satisfaction and expected satisfaction may specifically include: obtaining a first priority and a second priority; and determining comprehensive satisfaction based on the first priority, the second priority, user satisfaction, and expected satisfaction.
[0028] The first priority refers to the priority corresponding to user satisfaction, and the second priority refers to the priority corresponding to expected satisfaction.
[0029] The above technical solution describes another way to determine comprehensive satisfaction based on user satisfaction and expected satisfaction, which can effectively improve the feasibility of this application.
[0030] In an optional implementation, the first priority is higher than the second priority.
[0031] In the above technical solution, prioritizing the first priority over the second priority allows the hidden layer model to be trained based on user satisfaction and predicted satisfaction, placing greater emphasis on user satisfaction. This allows the hidden layer model to more accurately capture user profiles, further improving the accuracy of the hidden layer model during inference.
[0032] In an optional implementation, the above method may further include: receiving a reasoning task input by a user, inputting the reasoning task into a reasoning model, obtaining a reasoning result, and outputting the reasoning result.
[0033] By embedding the hidden state layer model within the inference model through this technical solution, we can personalize the inference tasks input by users and obtain inference results that match their profiles. This can better meet user needs and provide personalized intelligent services for users.
[0034] In the second aspect, a model training device is provided, which includes: a functional unit for executing any one of the methods provided in the first aspect, and the actions performed by each functional unit are implemented by hardware or by hardware executing corresponding software. For example, the model training device may include: an acquisition unit, a processing unit, and a training unit. The acquisition unit is used to obtain the reasoning task corresponding to the user portrait, and the reasoning result output by the reasoning model for the reasoning task. The processing unit is used to determine the expected satisfaction based on the reasoning task and the reasoning result. The training unit is used to train the hidden state layer model based on the reasoning task, the reasoning result, and the expected satisfaction during the operation of the reasoning model.
[0035] In a third aspect, a computing device is provided, comprising: a processor and a memory, wherein the processor is connected to the memory, the memory is used to store computer-executable instructions, and the processor executes the computer-executable instructions stored in the memory, thereby implementing any one of the methods provided in the first aspect.
[0036] In a fourth aspect, a chip is provided, comprising: a processor and an interface circuit; the interface circuit is configured to receive code instructions and transmit them to the processor; and the processor is configured to run the code instructions to execute any one of the methods provided in the first aspect.
[0037] In a fifth aspect, a computer-readable storage medium is provided, which stores computer execution instructions. When the computer execution instructions are run on a computer, the computer executes any one of the methods provided in the first aspect.
[0038] In a sixth aspect, a computer program product is provided, comprising computer execution instructions, which, when executed on a computer, enable the computer to execute any one of the methods provided in the first aspect.
[0039] Among them, the technical effects brought about by any implementation method in the second to sixth aspects can refer to the technical effects brought about by different implementation methods in the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 A schematic diagram of the architecture of a communication system provided in an embodiment of the present application;
[0041] Figure 2 A schematic diagram of the structure of a computing device provided in an embodiment of the present application;
[0042] Figure 3 A flowchart of a model training method provided in an embodiment of the present application;
[0043] Figure 4 A schematic diagram of the structure of an inference model and a hidden state layer model provided in an embodiment of the present application;
[0044] Figure 5 A time series diagram provided in an embodiment of the present application;
[0045] Figure 6 A schematic diagram of the structure of a model training device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0046] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.
[0047] In the description of this application, unless otherwise specified, " / " indicates that the objects associated before and after are in an "or" relationship, for example, A / B can represent A or B; "and / or" in this application is merely a description of the association relationship of associated objects, indicating that three relationships may exist, for example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural.
[0048] Furthermore, in the description of this application, unless otherwise specified, "plurality" means two or more than two. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or plural.
[0049] In addition, in order to facilitate the clear description of the technical solutions of the embodiments of the present application, in the embodiments of the present application, words such as "first" and "second" are used to distinguish between identical or similar items with substantially the same functions and effects. Those skilled in the art will understand that words such as "first" and "second" do not limit the quantity and execution order, and words such as "first" and "second" do not necessarily limit differences. At the same time, in the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or design schemes. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a concrete way for easy understanding.
[0050] First, the application scenarios of the embodiments of the present application are exemplarily introduced.
[0051] With the continuous development of big model technology, intelligent services based on big models, such as speech synthesis and image processing, are also gaining more and more attention. These intelligent services provide users with a convenient and efficient experience through big model technology.
[0052] An embodiment of the present application provides a model training method applied to a computing device, wherein the computing device can obtain an inference task corresponding to a user profile and an inference result output by an inference model for the inference task. A hidden state layer model is embedded in the hidden layer of the inference model, and the hidden state layer model is used to modify the parameters of the hidden layer during the operation of the inference model so that the inference result output by the inference model matches the user profile. Thereafter, the computing device can determine an expected satisfaction level based on the inference task and the inference result, and train the hidden state layer model based on the inference task, the inference result, and the expected satisfaction level during the operation of the inference model itself.
[0053] The expected satisfaction refers to the user's satisfaction with the inference results predicted by the computing device.
[0054] Through this technical solution, a computing device can train the hidden state layer model embedded in the inference model in real time while running the inference model, thereby achieving real-time training of the inference model. In other words, training and using the large model can be performed simultaneously, effectively enhancing the user experience.
[0055] The following is an exemplary introduction to the system architecture of the embodiment of the present application.
[0056] Figure 1 This is a schematic diagram of the architecture of a communication system provided in an embodiment of the present application. Figure 1 As shown, the communication system may include a terminal device 101 and a computing device 102. The terminal device 101 and the computing device 102 are in communication connection.
[0057] The terminal device 101 may also be referred to as user equipment (UE) or terminal equipment (TE). Exemplarily, the terminal device may include a personal digital assistant (PDA), an ultra-mobile personal computer (UMPC), a laptop computer, a netbook, a desktop computer, an all-in-one computer, a mobile phone, a tablet computer (pad), an in-vehicle device, or a wearable device.
[0058] Computing device 102 may be a network device. Network devices may include servers, etc. A server may be a single physical server, or two or more physical servers sharing different responsibilities and collaborating to implement various server functions, or a virtual server (also referred to as a virtual machine) running on a physical server. For example, the server may be a blade server, a high-density server, a rack server, or a tower server.
[0059] It should be noted that the embodiment of the present application does not limit the device form of the computing device 102. The system architecture of the computing device 102 provided in the embodiment of the present application is described below using a server as an example.
[0060] Figure 2 Schematic diagram of a computing device 102 provided in an embodiment of the present application. Figure 2 As shown, the computing device 102 includes a processor 202 and a memory 204. The processor 202 is connected to the memory 204 via a double data rate (DDR) bus 203. Here, the DDR bus 203 can also be replaced with other types of buses, and the embodiment of the application does not limit the bus type. In addition, the computing device 102 also includes various input / output (I / O) devices, and the processor 202 can access these I / O devices 207 via a high-speed peripheral component interconnect express (PCIe) bus 205.
[0061] The processor 202 is the computing core and control core of the computing device 102. The processor 202 may include one or more processor cores 201. The processor 202 may be a very large-scale integrated circuit. An operating system and other software programs are installed in the processor 202, so that the processor 202 can access the memory 204 and various PCIe devices. It is understood that in the embodiment of the present invention, the core 201 in the processor 202 may be, for example, a central processing unit (CPU) or other application-specific integrated circuit (ASIC). The processor 202 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. In actual applications, the computing device 102 may also include multiple processors.
[0062] A memory controller is a bus circuit controller within computing device 102 that controls memory 204 and manages and schedules data transfers from memory 204 to core 201. The memory controller enables data exchange between memory 204 and core 201. The memory controller can be a separate chip connected to core 201 via the system bus.
[0063] Those skilled in the art will appreciate that the memory controller may be integrated into processor 202, embedded in the north bridge, or a separate memory controller chip. The embodiments of the present invention do not limit the specific location or form of the memory controller. In practical applications, the memory controller may control the necessary logic to write data to or read data from memory 204. Memory controller 204 may be a memory controller in a processor system such as a general-purpose processor, a dedicated accelerator, a GPU, an FPGA, or an embedded processor.
[0064] Memory 204 is the main memory of computing device 102. Memory 204 is typically used to store various running software in the operating system, input and output data, and information exchanged with external memory. To improve the access speed of processor 202, memory 204 needs to have a fast access speed. In traditional computer system architectures, dynamic random access memory (DRAM) is typically used as memory 204. Processor 202 can access memory 204 at high speed through a memory controller, performing read and write operations on any storage unit in memory 204. In addition to DRAM, memory 204 can also be other random access memories, such as static random access memory (SRAM). Memory 204 can also be read-only memory (ROM). For example, ROM can be programmable read-only memory (PROM) or erasable programmable read-only memory (EPROM). This embodiment does not limit the quantity or type of memory 204. In addition, the memory 204 can be configured to have a power-saving function. The power-saving function means that the data stored in the memory will not be lost when the system loses power and then powers on again. The memory 204 with the power-saving function is called a non-volatile memory.
[0065] I / O device 207 refers to hardware capable of data transmission and can also be understood as a device that interfaces with an I / O interface. Common I / O devices include network cards, printers, keyboards, mice, and the like. All external storage devices, such as hard drives, floppy disks, and optical disks, can also serve as I / O devices. Processor 202 can access each I / O device 207 via PCIe bus 205. It should be noted that PCIe bus 205 is only an example and can be replaced by other buses, such as a unified bus (UB).
[0066] The baseboard management controller (BMC) 206 can perform firmware upgrades, manage the device's operating status, and troubleshoot problems even when the computing device 102 is powered off. The processor 202 can access the BMC 206 via the PCIe bus 205. The BMC 206 can also be connected to at least one sensor. The sensor can acquire status data from the computing device 102, including temperature, current, and voltage data. The type of status data is not specifically limited in this application. The BMC 206 communicates with the processor 202 via the PCIe bus or other bus types, for example, transmitting acquired status data to the processor 202 for processing. The BMC 206 can also maintain program code in memory, including upgrading or restoring it. The BMC 206 can also control the power supply circuit or clock circuit within the computing device 102. In summary, the BMC 206 can manage the computing device 102 using the above methods. However, the BMC 206 is an optional device. In some implementations, the processor 202 can communicate directly with the sensor to directly manage and maintain the computing device 102 .
[0067] It should be noted that the system architecture and application scenarios described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Ordinary technicians in this field can know that with the evolution of the system architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0068] The following Figure 1 Taking the computing device shown as an example, the reasoning method provided in the embodiment of the present application is introduced in detail. Figure 3 A flowchart of a reasoning method provided in an embodiment of the present application is shown in FIG. Figure 3 As shown, the method includes S301-S303.
[0069] S301, obtaining the reasoning task corresponding to the user portrait and the reasoning result output by the reasoning model for the reasoning task.
[0070] The reasoning task corresponding to a user profile can be any one or more reasoning tasks previously entered by the user, that is, any reasoning tasks entered by the user before the current moment. Reasoning tasks can reflect user needs or topics of interest. User profiles are used to reflect user preferences and personal information. User preferences may include, but are not limited to, preferred communication methods and topics of interest. Personal information may include, but is not limited to, occupation and address.
[0071] The embodiments of this application do not specifically limit the content of the reasoning task. In one example, the reasoning task can be text information input by the user, such as tomorrow's weather forecast or instructions on how to use a rice cooker. In another example, the reasoning task can be image information input by the user. In another example, the reasoning task can be audio information input by the user.
[0072] The inference model runs in a computing device. The hidden state layer model is embedded in the hidden layer of the inference model. Therefore, the hidden state layer model can also be called a hidden layer model or a model within a model. The hidden state layer model is used to modify the parameters of the hidden layer during the operation of the inference model so that the inference results output by the inference model match the user profile. In one example, the inference model and the hidden state layer model can refer to Figure 4 Deploy as shown.
[0073] It can be understood that the above-mentioned modification of the parameters of the hidden layer during the operation of the inference model means: training the hidden state layer model during the operation of the inference model. The specific process can be referred to the description in S303 and will not be described here.
[0074] Hidden layers refer to the layers in the inference model other than the input layer and output layer. The main function of the hidden layer is to extract the features of the inference task input by the user so that the inference model can understand the user's intention.
[0075] The embodiments of this application do not specifically limit the inference model. For example, the inference model can be a large language model (LLM), a multimodal question-answering model, an embedding model, etc., built based on a transformer network, a recursive neural network (RNN), or a convolutional neural network (CNN).
[0076] The embodiment of the present application does not specifically limit the timing when the computing device obtains the reasoning task corresponding to the user portrait, and the reasoning result (hereinafter referred to as the first reasoning result) output by the reasoning model for the reasoning task.
[0077] In one example, the computing device may obtain the reasoning task and first reasoning result corresponding to the user profile at a preset time interval. For example, taking the preset time interval as 7 days, the computing device may obtain the reasoning task and first reasoning result corresponding to the user profile every 7 days. In another example, the computing device may obtain the reasoning task and first reasoning result corresponding to the user profile after the reasoning model outputs the first reasoning result. The following describes the contents of S301 using the example of the computing device obtaining the reasoning task and first reasoning result corresponding to the user profile after the reasoning model outputs the first reasoning result.
[0078] Specifically, taking the user's terminal device as a mobile phone as an example, an artificial intelligence (AI) question-and-answer platform can be deployed in the mobile phone, and the AI question-and-answer platform is used to provide intelligent question-and-answer services. On this basis, when the user intends to use AI for intelligent question-and-answer services, he can start the AI question-and-answer platform on his mobile phone and log in to his account on the AI question-and-answer platform. After logging in, the user can enter the corresponding reasoning task on the AI question-and-answer platform based on his own needs. After the mobile phone receives the reasoning task, it can send the reasoning task to the computing device. After receiving the reasoning task, the computing device can encode the reasoning task and obtain the vector encoding information (also called token) corresponding to the reasoning task. The computing device can input the token corresponding to the reasoning task into its own reasoning model to obtain the reasoning result corresponding to the reasoning task, that is, the first reasoning result. Afterwards, the computing device obtains the reasoning task and the first reasoning result.
[0079] In the embodiments of the present application, there is no specific limitation on the AI question-and-answer platform. In one example, the AI question-and-answer platform can be AI software installed in a terminal device that can be used to provide intelligent question-and-answer services. At this time, the process for the above-mentioned user to start the AI question-and-answer platform on his or her mobile phone is as follows: an icon of the AI software can be displayed on the mobile phone desktop, and the user can click on the icon of the AI software to start the AI software. In another example, the AI question-and-answer platform can be an AI website for providing intelligent question-and-answer services. At this time, the process for the above-mentioned user to start the AI question-and-answer platform on his or her mobile phone is as follows: the user can enter the URL of the AI website in the mobile phone browser to log in to the AI website.
[0080] For example, let's assume a user needs to know tomorrow's weather conditions, the AI Q&A platform is AI software, and the inference model is a large language model. After launching the AI software on their phone, the user can enter the corresponding inference task in the AI software's dialog box: "Tomorrow's weather conditions." After receiving the inference task, the phone can send the inference task to the computing device. After receiving the inference task, the computing device can input the token corresponding to the inference task into the large language model, obtaining the first inference result "Sunny to cloudy." The computing device then obtains the inference task "Tomorrow's weather conditions" and the first inference result "Sunny to cloudy."
[0081] In order to avoid the user inputting an inference task and waiting for the computing device to output the first inference result for too long, in an embodiment of the present application, the above-mentioned computing device receives the inference task input by the user and inputs the inference task into its own inference model. After obtaining the first inference result, it can also output the first inference result to the user.
[0082] It can be understood that the first inference result output by the above inference model is represented by a vector. Therefore, before outputting the first inference result to the user, the computing device needs to convert the first inference result into text and output it to the user.
[0083] It should be noted that in the embodiment of the present application, the computing device outputs the first inference result to the user, and the computing device uses the inference task and the first inference result to perform real-time training on the hidden state layer model. These are asynchronous operations and do not affect each other. In this way, the first inference result can be fed back to the user in a timely manner, and real-time training of the hidden state layer model can be achieved.
[0084] S302: Determine the expected satisfaction level based on the reasoning task and the reasoning result.
[0085] Estimated satisfaction refers to the user's satisfaction with the inference results predicted by the computing device.
[0086] The embodiments of this application do not specifically limit the method for representing the estimated satisfaction level. In one example, the estimated satisfaction level can be represented by a score, where a higher score indicates a higher predicted user satisfaction with the first reasoning result, and a lower score indicates a lower predicted user satisfaction with the first reasoning result. In another example, the estimated satisfaction level can be represented by text, such as "very satisfied," "relatively satisfied," "satisfied," "relatively dissatisfied," or "unsatisfied."
[0087] Specifically, after the computing device obtains the reasoning task and the first reasoning result through S301, the computing device may predict the user's satisfaction with the first reasoning result, that is, the expected satisfaction.
[0088] To increase the efficiency of the computing device in determining the estimated satisfaction level, in an optional embodiment, a prediction model (also known as a scoring model) may be deployed in the computing device, and this prediction model is used to determine the estimated satisfaction level. Based on this, the above S302 can be replaced by: inputting the reasoning task and the first reasoning result into the prediction model to obtain the estimated satisfaction level.
[0089] Specifically, after the computing device obtains the reasoning task and the first reasoning result through S301, the reasoning task and the first reasoning result can be input into the prediction model, and the prediction model can output the user's satisfaction with the first reasoning result, that is, the expected satisfaction.
[0090] Through the above method, the expected satisfaction level can be directly predicted by the prediction model, which can effectively improve the efficiency of determining the expected satisfaction level and further improve the efficiency of real-time training of the hidden state layer model.
[0091] In an optional embodiment, before using the prediction model, the computing device may further train a pre-trained model to obtain a prediction model. Specifically, the computing device may obtain multiple task samples, a result sample corresponding to each task sample, and a user's satisfaction with each result sample (hereinafter referred to as a first satisfaction level). For each task sample, the computing device may train the pre-trained model using the task sample, the result sample corresponding to the task sample, and the user's satisfaction with the result sample to obtain a prediction model.
[0092] The embodiments of this application do not specifically limit the method for representing the first satisfaction level. In one example, the first satisfaction level can be represented by a score, where a higher score indicates a higher user satisfaction with the result sample, and a lower score indicates a lower user satisfaction with the result sample. In another example, the first satisfaction level can be represented by text, such as "very satisfied," "satisfied," "unsatisfied," etc.
[0093] A task sample refers to a historical reasoning task entered by a user, whose reasoning results correspond to user satisfaction feedback. For example, if the user entered reasoning task A before the current moment, asking for tomorrow's weather conditions, and the reasoning model outputs the inference result for task A as "sunny turning cloudy," and if the user entered reasoning task B, asking for tomorrow's temperature, and the reasoning model outputs the inference result for task B as "1 degree Celsius," and the user reported satisfaction with the inference result for task A but not with the inference result for task B, the computing device may use task A as a task sample and the inference result for task A as a result sample.
[0094] User satisfaction with the result sample refers to the user's satisfaction with the result sample based on historical input.
[0095] It can be understood that when the first reasoning result described in S301 above corresponds to the user feedback satisfaction, the task sample and the reasoning task described in S301 above can be the same reasoning task. At this time, the user's satisfaction with the result sample is: the user's feedback satisfaction with the first reasoning result (i.e., user satisfaction).
[0096] The embodiments of the present application do not specifically limit the way in which users provide feedback on their satisfaction. In one example, an input box may be displayed on the AI question-and-answer platform. After the user sees the inference result on the AI question-and-answer platform, he or she may enter his or her satisfaction with the inference result in the input box, such as satisfaction in text form such as satisfied or dissatisfied, or satisfaction in the form of scores such as 4 points or 1 point. In another example, a satisfaction control (such as a like control, a step-down control, or a score control, etc.) may be displayed on the AI question-and-answer platform. After the user sees the inference result on the AI question-and-answer platform, he or she may operate different controls to provide feedback on his or her satisfaction with the inference result. For example, assuming that the user is relatively satisfied with the inference result, the user may operate the like control. For another example, assuming that the user is dissatisfied with the inference result, the user may operate the step-down control.
[0097] The embodiments of the present application do not specifically limit the pre-training model. For example, the pre-training model can be a model based on a transformer network, a recursive neural network (RNN), or a convolutional neural network (CNN).
[0098] Specifically, the computing device may store a correspondence between user identification and reasoning tasks. When the computing device uses a task sample corresponding to a certain user to train a pre-trained model, it may first obtain the user identification of the user, and based on the user identification, determine the reasoning task corresponding to the user identification from the correspondence stored in itself, and determine multiple task samples corresponding to the user identification from the reasoning tasks, as well as the result samples corresponding to each task sample and the user's satisfaction with each result sample, that is, the first satisfaction level.
[0099] Afterwards, for each task sample, the computing device may input the task sample and the result sample corresponding to the task sample into a pre-trained model. The pre-trained model outputs a satisfaction level (hereinafter referred to as the second satisfaction level). The computing device may use a preset loss function to determine the loss value between the first satisfaction level and the second satisfaction level. The computing device may use a preset optimization algorithm to adjust the parameters of the pre-trained model to minimize the aggregated loss value between the first satisfaction level and the second satisfaction level corresponding to each result sample (hereinafter referred to as the loss value corresponding to each result sample). In this way, a prediction model can be obtained.
[0100] The embodiments of the present application do not specifically limit the above-mentioned optimization algorithm. For example, the optimization algorithm can be a gradient descent method, a stochastic gradient descent method (SGD), a momentum method, etc.
[0101] The embodiment of the present application does not specifically limit the aggregation result of the loss values corresponding to each result sample. For example, the aggregation result of the loss values corresponding to each result sample can be the sum of the loss values corresponding to each result sample. For another example, the aggregation result of the loss values corresponding to each result sample can be the average value of the loss values corresponding to each result sample.
[0102] For example, the optimization algorithm is a gradient descent method, the first satisfaction is represented by a score, and the aggregation result of the loss values corresponding to each result sample is the sum of the loss values corresponding to each result sample. Assume that the multiple task samples include task A, task B, and task C, and the result sample corresponding to task A includes result A1, the result sample corresponding to task B includes result B1, and the result sample corresponding to task C includes result C1.
[0103] Taking result A1 as an example, assuming the first satisfaction level corresponding to result A1 is M1, the computing device can input task A and result A1 into the pre-trained model. The pre-trained model can output a second satisfaction level M2 for result A1. The computing device can then use a preset loss function to determine the loss value between the first satisfaction level M1 and the second satisfaction level M2 corresponding to result A1. Repeating the above steps, the computing device can obtain the loss value corresponding to result A1 (hereinafter referred to as loss1), the loss value corresponding to result B1 (hereinafter referred to as loss3), and the loss value corresponding to result C1 (hereinafter referred to as loss3).
[0104] The computing device can then use a gradient descent method to adjust the parameters of the pre-trained model to minimize the sum of loss1, loss2, and loss3. When the sum of loss1, loss2, and loss3 is minimized, it indicates that the training of the pre-trained model is complete and a prediction model is obtained.
[0105] S303: During the operation of the inference model, the hidden state layer model is trained based on the inference task, the inference result, and the expected satisfaction.
[0106] Specifically, after the computing device obtains the reasoning task, the first reasoning result and the expected satisfaction through the above method, it can feed back the reasoning task, the first reasoning result and the expected satisfaction to the hidden state layer model, so that the hidden state layer model adopts a preset training method and adjusts its own parameters based on the reasoning task, the first reasoning result and the expected satisfaction. In this way, the training of the hidden state layer model can be achieved.
[0107] The process of the above-mentioned hidden state layer model adjusting its own parameters based on the reasoning task, the first reasoning result and the expected satisfaction can specifically include: when the expected satisfaction is greater than the satisfaction threshold, the hidden state layer model can adjust its own parameters based on the reasoning task, the first reasoning result and the expected satisfaction through positive training, using a preset training method; when the expected satisfaction is less than the satisfaction threshold, the hidden state layer model can adjust its own parameters based on the reasoning task, the first reasoning result and the expected satisfaction through negative training, using a preset training method.
[0108] The embodiments of the present application do not specifically limit the preset training method. For example, the preset training method may be a self-supervised learning method, an online support vector machine (SVM), or the like.
[0109] The above technical solution enables real-time training of the hidden state layer model embedded in the inference model while the inference model is running on a computing device, thereby achieving real-time training of the inference model. This means that large models can be trained and used simultaneously, which not only enhances the adaptability of the inference model but also effectively improves the user experience.
[0110] From the above description, it can be seen that the hidden state layer model is embedded in the hidden layer of the reasoning model. Therefore, the above training of the hidden state layer model based on the reasoning task, the first reasoning result and the expected satisfaction can specifically include: feeding back the reasoning task, the first reasoning result and the expected satisfaction to the hidden layer of the reasoning model to adjust the parameters of the hidden state layer model based on the reasoning task, the first reasoning result and the expected satisfaction.
[0111] Through the above technical solution, the reasoning task, the first reasoning result and the expected satisfaction can be directly fed back to the hidden layer of the reasoning model without passing through the embedding layer before the hidden layer. In this way, the feedback path of the reasoning task, the first reasoning result and the expected satisfaction can be shortened, thereby accelerating the efficiency of training the hidden state layer model.
[0112] The inference model receives a very large number of inference tasks every day. If each inference task and the first inference result corresponding to the inference task were used to train the hidden state layer model embedded in the inference model, a large amount of computing resources of the computing device would be required. To conserve the computing resources of the computing device, in an optional embodiment, the computing device may train the hidden state layer model based on the inference task, the first inference result, and the expected satisfaction level when the expected satisfaction level is greater than a preset satisfaction threshold.
[0113] The satisfaction threshold may be preset by the user.
[0114] Specifically, a satisfaction threshold may be stored in the computing device. Accordingly, after obtaining the predicted satisfaction level, the computing device may compare the predicted satisfaction level with the satisfaction threshold. If the predicted satisfaction level is less than the satisfaction threshold, the computing device may delete the inference task and first inference result corresponding to the predicted satisfaction level, i.e., not use the predicted satisfaction level, the inference task and the first inference result corresponding to the predicted satisfaction level to train the hidden state layer model. If the predicted satisfaction level is greater than the satisfaction threshold, the computing device may use the predicted satisfaction level, the inference task and the first inference result corresponding to the predicted satisfaction level to train the hidden state layer model.
[0115] The embodiments of the present application do not specifically limit the case where the expected satisfaction level is equal to the satisfaction threshold. In one example, when the expected satisfaction level is equal to the satisfaction threshold, the computing device may delete the reasoning task and the first reasoning result corresponding to the expected satisfaction level. In another example, when the expected satisfaction level is equal to the satisfaction threshold, the computing device may use the expected satisfaction level, the reasoning task and the first reasoning result corresponding to the expected satisfaction level to train the hidden state layer model.
[0116] For example, the expected satisfaction is represented by a score, the satisfaction threshold is 3 points, and a prediction model is deployed in the computing device, such as Figure 5 As shown, assume that at time t1, the hidden state layer model embedded in the hidden layer of the inference model is in stage t0. That is, the last time the hidden state layer model was trained was time t0. The user inputs the inference task Q1, and the inference result output by the inference model for question Q1 is answer A1. The computing device can output answer A1 to the user and input question Q1 and answer A1 into the prediction model, which outputs the estimated satisfaction score corresponding to answer A1. Assuming that the estimated satisfaction score for answer A1 is 4 points, which is greater than the satisfaction threshold, the computing device can train the hidden state layer model using question Q1, answer A1, and the estimated satisfaction score of 4 points. After training, the hidden state layer model embedded in the hidden layer of the inference model is in stage t1.
[0117] Assuming that time t2 is after time t1, then at time t2, the hidden state layer model embedded in the hidden layer of the inference model is at stage t1. The inference task input by the user at time t2 is question Q2, and the inference result output by the inference model for question Q2 is answer A2. The computing device can output answer A2 to the user and input question Q2 and answer A2 into the prediction model, which outputs the estimated satisfaction score corresponding to answer A2. Assuming that the estimated satisfaction score for answer A2 is 2 points, which is less than the satisfaction threshold, the computing device can delete question Q2 and answer A2, that is, not use question Q2, answer A2, and the estimated satisfaction score of 2 to train the hidden state layer model.
[0118] Assuming that time t3 is after time t2, then at time t3, the hidden state layer model embedded in the hidden layer of the inference model is still in stage t1. The inference task input by the user at time t3 is question Q3, and the inference result output by the inference model for question Q3 is answer A3. The computing device can output answer A3 to the user and input question Q3 and answer A3 into the prediction model, which outputs the estimated satisfaction score corresponding to answer A3. Assuming that the estimated satisfaction score for answer A3 is 5 points, which is greater than the satisfaction threshold, the computing device can train the hidden state layer model using question Q3, answer A3, and the estimated satisfaction score of 5 points. After training, the hidden state layer model embedded in the hidden layer of the inference model is in stage t3.
[0119] Assuming that time t4 is after time t3, then at time t4, the hidden state layer model embedded in the hidden layer of the inference model is at stage t3. The inference task input by the user at time t4 is question Q4, and the inference result output by the inference model for question Q4 is answer A4. The computing device can output answer A4 to the user and input question Q4 and answer A4 into the prediction model, which outputs the estimated satisfaction score corresponding to answer A4. Assuming that the estimated satisfaction score for answer A4 is 1, which is less than the satisfaction threshold, the computing device can delete question Q4 and answer A4, that is, not use question Q4, answer A4, and the estimated satisfaction score of 1 to train the hidden state layer model.
[0120] Through the above technical solution, the hidden state layer model is only trained when the expected satisfaction level is greater than a preset satisfaction threshold. This not only effectively saves computing power resources of the computing device, but also prevents malicious user input from affecting the quality of the hidden state layer model, further enhancing the robustness of the hidden state layer model. In addition, through the above method, the hidden state layer model can be continuously trained to make it more closely fit the user profile, meet user needs, better serve users, and enhance the user experience.
[0121] In an optional embodiment, the computing device may also receive user feedback on satisfaction with the first reasoning result (hereinafter referred to as user satisfaction). Based on this, before training the hidden state layer model based on the reasoning task, the first reasoning result, and the expected satisfaction, the computing device may also obtain user satisfaction and determine a comprehensive satisfaction based on the user satisfaction and the expected satisfaction. Accordingly, the aforementioned training of the hidden state layer model based on the reasoning task, the first reasoning result, and the expected satisfaction can be replaced by training the hidden state layer model based on the reasoning task, the first reasoning result, and the comprehensive satisfaction.
[0122] Specifically, the process of training the hidden state layer model based on the reasoning task, the first reasoning result and the comprehensive satisfaction can refer to the description in S303 above, which will not be repeated here.
[0123] The embodiments of this application do not specifically limit the method for representing user satisfaction. In one example, user satisfaction can be represented by a score, where a higher score indicates higher user satisfaction and a lower score indicates lower user satisfaction. In another example, user satisfaction can be represented by text, such as "very satisfied," "satisfied," or "unsatisfied."
[0124] The following describes in detail the process of determining the overall satisfaction of a computing device based on user satisfaction and expected satisfaction.
[0125] The computing device can determine the overall satisfaction in two ways as described below.
[0126] Method 1: The computing device may obtain a weight corresponding to user satisfaction (hereinafter referred to as a first weight) and a weight corresponding to expected satisfaction (hereinafter referred to as a second weight). The computing device may then determine a comprehensive satisfaction score based on the first weight, the second weight, the user satisfaction score, and the expected satisfaction score.
[0127] Specifically, the computing device may store a first weight and a second weight. After obtaining the user satisfaction and the expected satisfaction, the computing device may take the product of the first weight and the user satisfaction and the sum of the product of the second weight and the expected satisfaction as the comprehensive satisfaction.
[0128] For example, assuming that the first weight is 80% and the second weight is 20%, both user satisfaction and expected satisfaction are represented by scores, and user satisfaction is 6 points and expected satisfaction is 8 points, then the comprehensive satisfaction = 80%*6+20%*8=6.4 points.
[0129] In order to better meet the needs of users and make the first inference result output by the trained inference model more consistent with the user portrait, in an optional embodiment, the above-mentioned first weight is greater than the second weight.
[0130] Second approach: The computing device may obtain a priority level corresponding to user satisfaction (hereinafter referred to as a first priority level) and a priority level corresponding to expected satisfaction (hereinafter referred to as a second priority level). The computing device may then determine a comprehensive satisfaction level based on the first priority level, the second priority level, the user satisfaction level, and the expected satisfaction level.
[0131] The smaller the priority value, the higher the priority. For example, if the first priority is 1 and the second priority is 2, the first priority is higher than the second priority.
[0132] Specifically, in some embodiments, a computing device may store a first priority and a second priority. After obtaining user satisfaction and predicted satisfaction, the computing device may obtain weights corresponding to the first and second priorities, where the higher the priority, the greater the weight. Assuming that the weight corresponding to the first priority is weight A and the weight corresponding to the second priority is weight B, the computing device may calculate the product of weight A and user satisfaction, and the sum of weight B and predicted satisfaction as the comprehensive satisfaction.
[0133] The weight corresponding to the first priority and the weight corresponding to the second priority are preset.
[0134] For example, assuming that the first priority is 1, the second priority is 2, and the weight A corresponding to the first priority is 60%, the weight B corresponding to the second priority is 40%, user satisfaction and expected satisfaction are both represented by scores, and user satisfaction is 6 points, and the expected satisfaction is 8 points, then the comprehensive satisfaction = 60%*6+40%*8=6.8 points.
[0135] In order to better meet the needs of users and make the first inference result output by the trained inference model more consistent with the user portrait, in an optional implementation, the above-mentioned first priority is greater than the second priority.
[0136] The first weight being greater than the second weight, or the first priority being greater than the second priority, can result in a greater emphasis on user satisfaction when training the hidden state layer model based on user satisfaction and predicted satisfaction. This allows the hidden state layer model to more accurately capture user profiles, further improving the accuracy of the hidden state layer model during inference.
[0137] In an optional embodiment, after the hidden state layer model is trained based on the reasoning task, the first reasoning result and the expected satisfaction, the user can input the above reasoning task into the computing device again. After the computing device receives the reasoning task, it can respond to the reasoning task input by the user and input the reasoning task into the trained reasoning model to obtain the reasoning result (hereinafter referred to as the second reasoning result).
[0138] The trained inference model refers to an inference model that has a trained hidden state layer model embedded in it. The degree of match between the second inference result and the user profile is greater than the degree of match between the first inference result and the user profile.
[0139] For example, suppose at time t1, the user inputs question q, obtains answer a, and the computing device trains the hidden state layer model based on question q, answer a, and the predicted satisfaction level corresponding to answer a. At time t2, the user inputs question q again, and obtains answer c. The degree of match between answer c and the user profile is greater than the degree of match between answer a and the user profile.
[0140] The above mainly introduces the solution provided by the embodiment of the present application from the perspective of the method. In order to realize the above functions, the model training device includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should easily realize that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0141] The embodiment of the present application can, according to the above method, exemplarily divide the model training device into functional modules. For example, the model training device can include various functional modules corresponding to the various functional divisions, or two or more functions can be integrated into one processing module. The above-mentioned integrated modules can be implemented in the form of hardware or in the form of software functional modules. It should be noted that the division of modules in the embodiment of the present application is schematic and is only a logical functional division. There may be other division methods in actual implementation.
[0142] For example, Figure 6 A possible structural diagram of the model training device involved in the above embodiment is shown. The device includes: an acquisition unit 601, a processing unit 602 and a training unit 603.
[0143] An acquisition unit 601 acquires the inference task corresponding to the user profile and the inference result output by the inference model for the inference task. A processing unit 602 is configured to determine the expected satisfaction level based on the inference task and the inference result. A training unit 603 is configured to train the hidden state layer model based on the inference task, the inference result, and the expected satisfaction level during the execution of the inference model.
[0144] Among them, the inference model runs on a computing device, and a hidden state layer model is embedded in the hidden layer of the inference model. The hidden state layer model is used to modify the parameters of the hidden layer during the operation of the inference model so that the inference results output by the inference model match the user portrait. The expected satisfaction refers to the user's satisfaction with the inference results predicted by the computing device.
[0145] Optionally, a prediction model may be deployed in the computing device. On this basis, the processing unit 602 is specifically used to input the reasoning task and the reasoning result into the prediction model to obtain the expected satisfaction.
[0146] Optionally, the training unit 603 may also be configured to obtain multiple task samples, a result sample corresponding to each task sample, and user satisfaction with each result sample. For each task sample, the pre-trained model is trained using the task sample, the result sample corresponding to the task sample, and the user's satisfaction with the result sample to obtain a prediction model.
[0147] Among them, the user's satisfaction with each result sample is used to reflect the degree of matching between the result sample and the user portrait.
[0148] Optionally, the training unit 603 is specifically used to: feed back the reasoning task, reasoning result and expected satisfaction to the hidden layer of the reasoning model, so as to adjust the parameters of the hidden state layer model based on the reasoning task, reasoning result and expected satisfaction.
[0149] Optionally, the training unit 603 is specifically used to: when the expected satisfaction is greater than a preset satisfaction threshold, train the hidden state layer model based on the reasoning task, the reasoning result and the expected satisfaction.
[0150] Optionally, the processing unit 602 may also be used to obtain user satisfaction and determine a comprehensive satisfaction based on the user satisfaction and the expected satisfaction. On this basis, the training unit 603 is specifically used to train the hidden state layer model based on the reasoning task, the reasoning result, and the comprehensive satisfaction.
[0151] Among them, user satisfaction refers to the user's feedback on the satisfaction with the inference results.
[0152] Optionally, the processing unit 602 is specifically configured to obtain a first weight and a second weight, and determine a comprehensive satisfaction level based on the first weight, the second weight, user satisfaction, and expected satisfaction level.
[0153] The first weight refers to the weight corresponding to user satisfaction, and the second weight refers to the weight corresponding to expected satisfaction.
[0154] Optionally, the first weight is greater than the second weight.
[0155] Optionally, the processing unit 602 is specifically configured to: obtain a first priority and a second priority; and determine a comprehensive satisfaction level based on the first priority, the second priority, user satisfaction, and expected satisfaction level.
[0156] The first priority refers to the priority corresponding to user satisfaction, and the second priority refers to the priority corresponding to expected satisfaction.
[0157] Optionally, the first priority is higher than the second priority.
[0158] Optionally, the computing device may be equipped with a receiving unit, an input unit, and an output unit. The receiving unit is configured to receive inference tasks input by a user. The input unit is configured to input the inference tasks into the computing device's inference model to obtain inference results. The output unit is configured to output the inference results.
[0159] For a detailed description of the above optional methods, please refer to the aforementioned method embodiments, which will not be repeated here. In addition, the explanation of any of the above-mentioned model training devices and model reasoning devices and the description of their beneficial effects can refer to the above-mentioned corresponding method embodiments, which will not be repeated here.
[0160] The present application also provides a computing device comprising a processor and a memory, wherein the processor is connected to the memory, and the memory stores computer-executable instructions. When the processor executes the computer-executable instructions, the method of the above embodiment is implemented. The present application does not impose any restrictions on the specific form of the computing device. For example, the computing device can be a terminal device or a network device.
[0161] The term "terminal device" may be referred to as a terminal, user equipment (UE), terminal device, access terminal, subscriber unit, subscriber station, mobile station, remote station, remote terminal, mobile device, user terminal, wireless communication device, user agent, or user device. Specifically, a terminal device may be a mobile phone, augmented reality (AR) device, virtual reality (VR) device, tablet computer, laptop computer, ultra-mobile personal computer (UMPC), netbook, personal digital assistant (PDA), etc. Specifically, a network device may be a server, etc.
[0162] The server may be a physical or logical server, or may be two or more physical or logical servers that share different responsibilities and work together to implement various functions of the server.
[0163] An embodiment of the present application further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is run on a computer, the computer is caused to execute the method executed by any one of the computing devices provided above.
[0164] For explanations of the relevant contents and descriptions of the beneficial effects of any of the computer-readable storage media provided above, reference may be made to the corresponding embodiments described above, and no further details will be given here.
[0165] The embodiment of the present application also provides a chip. The chip integrates a control circuit and one or more ports for implementing the functions of the above-mentioned computing device. Optionally, the functions supported by the chip can be referred to above and will not be repeated here. A person of ordinary skill in the art will understand that all or part of the steps of implementing the above-mentioned embodiment can be completed by a program to instruct the relevant hardware. The program can be stored in a computer-readable storage medium. The storage medium mentioned above can be a read-only memory, a random access memory, etc. The above-mentioned processing unit or processor can be a central processing unit, a general-purpose processor, an application specific integrated circuit (ASIC), a microprocessor (digital signal processor, DSP), a field programmable gate array (FPGA) or other programmable logic device, transistor logic device, hardware component or any combination thereof.
[0166] The present application also provides a computer program product comprising instructions, which, when executed on a computer, causes the computer to perform any of the methods described in the above embodiments. The computer program product comprises one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated therein. The available media may be magnetic media (e.g., floppy disk, hard disk, tape), optical media (e.g., DVD), or semiconductor media (e.g., SSD).
[0167] It should be noted that the above-mentioned devices for storing computer instructions or computer programs provided in the embodiments of the present application, such as but not limited to the above-mentioned memories, computer-readable storage media and communication chips, etc., are all non-transitory.
[0168] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using a software program, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function according to the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more media that can be integrated. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a DVD), or a semiconductor medium (eg, a solid state disk (SSD)).
[0169] Although the present application is described herein in conjunction with various embodiments, in the process of implementing the claimed application, those skilled in the art may understand and implement other variations of the disclosed embodiments by reviewing the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude multiple situations. A single processor or other unit may implement several functions listed in the claims. Certain measures are recorded in different dependent claims, but this does not mean that these measures cannot be combined to produce good results.
[0170] Although the present application has been described with reference to specific features and embodiments thereof, it is apparent that various modifications and combinations may be made thereto without departing from the spirit and scope of the present application. Accordingly, this specification and the drawings are merely illustrative of the present application as defined by the appended claims and are deemed to cover any and all modifications, variations, combinations or equivalents within the scope of the present application. Obviously, those skilled in the art may make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, the present application is intended to include such modifications and variations as fall within the scope of the claims of the present application and their equivalents.
Claims
1. A model training method, characterized in that: Applied to a computing device, the method includes: Obtaining an inference task corresponding to the user profile and an inference result output by an inference model for the inference task, wherein the inference model runs on the computing device, and a hidden state layer model is embedded in a hidden layer of the inference model, wherein the hidden state layer model is used to modify parameters of the hidden layer during the running of the inference model so that the inference result output by the inference model matches the user profile; Determining an expected satisfaction level based on the reasoning task and the reasoning result; the expected satisfaction level refers to a user's satisfaction level with the reasoning result predicted by the computing device; During the operation of the inference model, the hidden state layer model is trained based on the inference task, the inference result and the expected satisfaction.
2. The method according to claim 1, characterized in that The computing device is provided with a prediction model, and determining the expected satisfaction level based on the reasoning task and the reasoning result includes: The reasoning task and the reasoning result are input into the prediction model to obtain the expected satisfaction level.
3. The method according to claim 2, characterized in that The method further comprises: Obtaining multiple task samples, result samples corresponding to each task sample, and user satisfaction with each result sample; the user satisfaction with each result sample is used to reflect the degree of match between the result sample and the user profile; For each of the task samples, the pre-training model is trained using the task sample, the result sample corresponding to the task sample, and the user's satisfaction with the result sample to obtain the prediction model.
4. The method according to any one of claims 1 to 3, characterized in that The training of the hidden state layer model based on the reasoning task, the reasoning result, and the expected satisfaction level includes: The reasoning task, the reasoning result and the expected satisfaction are fed back to the hidden layer of the reasoning model to adjust the parameters of the hidden state layer model based on the reasoning task, the reasoning result and the expected satisfaction.
5. The method according to any one of claims 1 to 3, characterized in that The training of the hidden state layer model based on the reasoning task, the reasoning result, and the expected satisfaction level includes: When the expected satisfaction level is greater than a preset satisfaction level threshold, the hidden state layer model is trained based on the reasoning task, the reasoning result, and the expected satisfaction level.
6. The method according to any one of claims 1 to 3, characterized in that Before training the hidden state layer model based on the reasoning task, the reasoning result, and the expected satisfaction, the method further includes: Obtaining user satisfaction; the user satisfaction refers to the user's feedback on the satisfaction of the inference result; Determining a comprehensive satisfaction level based on the user satisfaction level and the expected satisfaction level; The training of the hidden state layer model based on the reasoning task, the reasoning result, and the expected satisfaction level includes: The hidden state layer model is trained based on the reasoning task, the reasoning result and the comprehensive satisfaction.
7. The method according to claim 6, characterized in that Determining the comprehensive satisfaction level based on the user satisfaction level and the expected satisfaction level includes: Obtain a first weight and a second weight; the first weight refers to the weight corresponding to the user satisfaction, and the second weight refers to the weight corresponding to the expected satisfaction; The comprehensive satisfaction level is determined based on the first weight, the second weight, the user satisfaction level, and the expected satisfaction level.
8. The method according to claim 7, characterized in that The first weight is greater than the second weight.
9. The method according to claim 1, characterized in that The method further comprises: receiving the reasoning task input by a user; Inputting the reasoning task into the reasoning model to obtain the reasoning result; The inference result is output.
10. A computing device, characterized in that include: processor and memory; The processor is connected to a memory, the memory is used to store computer-executable instructions, and the processor executes the computer-executable instructions stored in the memory to enable the computing device to implement the method according to any one of claims 1 to 9.