Intent analysis large model training method based on webpage search and related device
By combining self-supervised training and reinforcement learning algorithms to train the model, the problem that traditional intent analysis methods cannot obtain Internet information in real time is solved, and efficient and accurate web search intent analysis is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-21
- Publication Date
- 2026-03-17
AI Technical Summary
Traditional intent analysis methods cannot obtain information from the Internet in real time, resulting in low efficiency and poor accuracy in web search intent analysis, making it difficult to achieve real-time analysis of large-scale data.
By acquiring the original dataset and neural network model for self-supervised training, obtaining and labeling search records, generating training data, and merging the trained models using reinforcement learning algorithms, a target training model is generated.
It improves the accuracy and efficiency of web search intent analysis and prediction, enhances intent analysis capabilities, and improves the accuracy of prediction results.
Smart Images

Figure CN116663647B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of deep learning technology, and in particular to a method and related equipment for training a large-scale model of intent analysis based on web search. Background Technology
[0002] With the development of deep learning technology, various natural language models have been widely applied in the field of intent analysis. However, traditional intent analysis methods cannot instantly acquire information from the internet to analyze web search intent. Furthermore, traditional intent analysis methods often require significant manpower and time, resulting in low efficiency, low accuracy and efficiency in intent analysis and prediction, and difficulty in achieving real-time analysis of large-scale data. Summary of the Invention
[0003] In view of this, the purpose of this application is to propose a training method and related equipment for a large-scale model of intent analysis based on web search, so as to solve the problem that existing intent analysis methods cannot obtain information from the Internet in real time to perform intent analysis of web search, thereby improving the accuracy and efficiency of web search intent analysis and prediction.
[0004] To achieve the above objectives, this application provides a method for training a large-scale intent analysis model based on web search, comprising:
[0005] Obtain the original dataset and neural network model;
[0006] The neural network model is self-supervised training is performed based on the original data to obtain a pre-trained model.
[0007] Obtain search records, label them, and generate training data;
[0008] The pre-trained model is trained based on the training data and preset feedback data to obtain the target training model.
[0009] Furthermore, the search record includes multiple search action data;
[0010] The process of acquiring search records, labeling them, and generating training data includes:
[0011] Acquire multiple search action data, compare the multiple search action data with preset search action data, and obtain the comparison results;
[0012] In response to the comparison results reaching a preset threshold, the search action data is labeled and used as qualified search action data; wherein each qualified search action data includes search data and at least one action data.
[0013] Use all the qualified search action data as training data.
[0014] Further, the step of training the pre-trained model based on the training data and preset feedback data to obtain the target training model includes:
[0015] The pre-trained model is trained based on the training data to obtain a first training model;
[0016] The pre-trained model is trained based on preset feedback data to obtain a second trained model;
[0017] The first and second training models are combined and trained using a reinforcement learning algorithm.
[0018] The process of training the first training model, the second training model, and the combined training of the two models is iterated. When the number of iterations reaches the iteration threshold, the iterative training stops and the target training model is generated.
[0019] Further, the step of training the pre-trained model based on the training data and preset feedback data to obtain the target training model includes:
[0020] The pre-trained model is trained based on the training data to obtain a first training model;
[0021] The pre-trained model is trained based on preset feedback data to obtain a second trained model;
[0022] The first training model and the second training model are combined and trained using a reinforcement learning algorithm.
[0023] The process of training the first training model, the second training model, and the combined training of the two models is iterated. Upon receiving a termination instruction, the training stops and the target training model is generated.
[0024] Further, training the pre-trained model based on the training data includes:
[0025] The action data includes multiple action sub-data with time tags, and the multiple action sub-data includes a current time data, and at least one action sub-data located after the current time data is the next time data;
[0026] Input an action sub-data as the current time-time data into the pre-trained model;
[0027] The next time step data corresponding to the current time step data is used as the output of the pre-trained model, and the action sub-data is one of referencing, scrolling up, or scrolling down.
[0028] Further, training the pre-trained model based on the training data includes:
[0029] The search data of the qualified search action data is input into the pre-trained model;
[0030] At least one of the aforementioned action data is used as the output of the pre-trained model, and the action data includes at least one of referencing, scrolling up, and scrolling down.
[0031] Further, training the pre-trained model based on preset feedback data includes:
[0032] The search data or the action data of the qualified search action data are input into the pre-trained model;
[0033] The preset action data is used as the output of the pre-trained model, and the preset action data includes at least one of referencing, scrolling up, and scrolling down.
[0034] Furthermore, the method also includes: verifying the target training model based on preset test data and preset answers to determine whether the target training model is qualified.
[0035] Further, the step of validating the target training model based on preset test data and preset answers to determine whether the target training model is qualified includes:
[0036] Input the preset test data into the target training model to obtain the output result;
[0037] The output result is compared with the preset answer to obtain the comparison value;
[0038] If the comparison value reaches a preset comparison value, the target training model is determined to be qualified.
[0039] To achieve the above objectives, this application provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the method described in any of the above descriptions.
[0040] In view of the above objectives, a third aspect of this application provides a non-transitory computer-readable storage medium storing computer instructions for causing a computer to perform the method described in any of the preceding claims.
[0041] As described above, the large-scale model training method for web search intent analysis provided in this application first obtains the original dataset and a neural network model; then, it performs self-supervised training on the neural network model based on the original data to obtain a pre-trained model; next, it acquires search records, labels them, and generates training data; finally, it trains the pre-trained model based on the training data and preset feedback data to obtain the target training model. This method, by labeling qualified search action data and using it as training data, trains the model using the training data and preset feedback data, making the intent analysis process of the target training model more accurate and the prediction results more accurate, thereby improving the accuracy and efficiency of web search intent analysis and prediction. Attached Figure Description
[0042] To more clearly illustrate the technical solutions in this application or related technologies, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0043] Figure 1 This is a schematic diagram of a large-scale model training method for intent analysis based on web page search, according to an embodiment of this application.
[0044] Figure 2 This is a schematic diagram of a large-scale model training device for intent analysis based on web page search, according to an embodiment of this application.
[0045] Figure 3 This is a schematic diagram of the hardware structure of an electronic device according to an embodiment of this application. Detailed Implementation
[0046] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with specific embodiments and the accompanying drawings.
[0047] It should be noted that, unless otherwise defined, the technical or scientific terms used in the embodiments of this application should have the ordinary meaning understood by one of ordinary skill in the art to which this application pertains. The terms "first," "second," and similar terms used in the embodiments of this application do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed after the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are only used to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0048] refer to Figure 1 This application provides a method for training a large-scale model for intent analysis based on web search, including:
[0049] Step S100: Obtain the original dataset and neural network model.
[0050] In this step, the original dataset is obtained according to different scenario requirements. The original dataset can be a publicly available dataset on the Internet or a custom dataset uploaded.
[0051] The original dataset may be text or articles, such as large amounts of text or text explanations, or text-related articles (e.g., articles from Wikipedia). The neural network model is a large model, which includes at least one of GPT-2 and GPT-3.
[0052] Step S200: Perform self-supervised training on the neural network model based on the original data to obtain a pre-trained model.
[0053] In this step, based on the use case, an existing public or uploaded custom dataset is selected and input into the GPT network (at least one of GPT-2 and GPT-3) for self-supervised training to obtain a pre-trained model.
[0054] Step S300: Obtain search records, label them, and generate training data.
[0055] In this step, the search records are news, social media, forum, and other public opinion data crawled from the internet using web crawlers, or data can be obtained in real time using APIs provided by search engines such as Google and Bing. The acquired data needs to undergo data cleaning, deduplication, filtering, and labeling to determine qualified search records. These qualified search records are then used as training data to make the pre-trained model learn more accurately.
[0056] Step S400: Train the pre-trained model according to the training data and preset feedback data to obtain the target training model.
[0057] In this step, a pre-trained model is trained using training data to obtain a training model. Then, the pre-trained model is trained using preset feedback data to obtain another training model. The two training models are then merged for training to obtain a target training model. This target training model is trained using a reinforcement learning algorithm by merging the two different training models, resulting in a stronger intention analysis capability.
[0058] Specifically, the target training model generated through the above steps S100-S400 uses labeled qualified search data as training data and is trained on the pre-trained model using preset feedback data to obtain two different training models. The two training models are then merged and trained using a reinforcement learning algorithm, which makes the intention analysis capability of the target training model stronger, the intention analysis process more accurate, and the prediction results more accurate.
[0059] In some embodiments, the search record includes multiple search action data;
[0060] The process of acquiring search records, labeling them, and generating training data includes:
[0061] Acquire multiple search action data, compare the multiple search action data with preset search action data, and obtain the comparison results;
[0062] In response to the comparison results reaching a preset threshold, the search action data is labeled and used as qualified search action data; wherein each qualified search action data includes search data and at least one action data.
[0063] Use all the qualified search action data as training data.
[0064] Specifically, the preset search action data includes pre-defined search data and at least one pre-defined action data (at least one of referencing, scrolling up, and scrolling down) corresponding to the search data, with a preset threshold of 100%. For example, if the search term is entered (e.g., 4G), the action data is scrolling down, and the next action data is referencing (e.g., the referenced webpage shows the second record); if the search term for the preset search action data is 4G, the preset action data is scrolling down, and the next preset action data is referencing (e.g., the referenced webpage shows the second record). Multiple search action data are compared with the preset search action data. If the comparison result is 100% and reaches the preset threshold, the search action data is considered qualified and is then labeled. Training the pre-trained model with qualified search action data enhances the pre-trained model's intent analysis capabilities and makes the intent analysis process more accurate.
[0065] In some embodiments, the pre-trained model is trained based on the training data and preset feedback data to obtain the target training model;
[0066] The pre-trained model is trained based on the training data to obtain a first training model;
[0067] The pre-trained model is trained based on preset feedback data to obtain a second trained model;
[0068] The first and second training models are combined and trained using a reinforcement learning algorithm.
[0069] The process of training the first training model, the second training model, and the combined training of the two models is iterated. When the number of iterations reaches the iteration threshold, the iterative training stops and the target training model is generated.
[0070] Specifically, for example, the training data includes a search data (i.e., a search term) and multiple action data (scroll down, scroll up, and reference). The search term is input into the pre-trained model, the scroll down action is used as the first answer of the output result of the training model, the scroll up action is used as the second answer, and the reference is used as the third answer, thus obtaining the first training model.
[0071] For example, the preset feedback data includes a search data (i.e. a search term) and multiple action data (scroll down, scroll up, and reference). The search term is input into the pre-trained model, the reference action is used as the first answer of the output result of the training model, the scroll up action is used as the second answer, and the scroll down action is used as the third answer, thus obtaining the second training model.
[0072] The first and second training models are combined and trained using a reinforcement learning algorithm (i.e., the PPO algorithm).
[0073] Repeat the above steps iteratively until the number of iterations reaches the threshold, then stop training and generate the target training model. That is, the first and second training models are trained using multiple search datasets, and the process of training the first and second training models is iteratively performed using reinforcement learning algorithms, resulting in a target training model with more accurate intent analysis.
[0074] In some embodiments, the pre-trained model is trained based on the training data and preset feedback data to obtain the target training model;
[0075] The pre-trained model is trained based on the training data to obtain a first training model;
[0076] The pre-trained model is trained based on preset feedback data to obtain a second trained model;
[0077] The first training model and the second training model are combined and trained using a reinforcement learning algorithm.
[0078] The process of training the first training model, the second training model, and the combined training of the two models is iterated. Upon receiving a termination instruction, the training stops and the target training model is generated.
[0079] Specifically, for example, the training data includes a search data (i.e., a search term) and multiple action data (scroll down, scroll up, and reference). The search term is input into the pre-trained model, the scroll down action is used as the first answer of the output result of the training model, the scroll up action is used as the second answer, and the reference is used as the third answer, thus obtaining the first training model.
[0080] For example, the preset feedback data includes a search data (i.e. a search term) and multiple action data (scroll down, scroll up, and reference). The search term is input into the pre-trained model, the reference action is used as the first answer of the output result of the training model, the scroll up action is used as the second answer, and the scroll down action is used as the third answer, thus obtaining the second training model.
[0081] The first and second training models are combined and trained using a reinforcement learning algorithm (i.e., the PPO algorithm).
[0082] Repeat the above steps iteratively until a termination instruction is received, at which point training stops and the target training model is generated. That is, the first and second training models are trained using multiple search datasets, and the process of training the first and second training models is iteratively performed using reinforcement learning algorithms, resulting in a target training model with more accurate intent analysis.
[0083] In some embodiments, training the pre-trained model based on the training data includes:
[0084] The action data includes multiple action sub-data with time tags, and the multiple action sub-data includes a current time data, and at least one action sub-data located after the current time data is the next time data;
[0085] Input an action sub-data as the current time-time data into the pre-trained model;
[0086] The next time step data corresponding to the current time step data is used as the output of the pre-trained model, and the action sub-data is one of referencing, scrolling up, or scrolling down.
[0087] Specifically, for example, the time of the current moment data (the current action sub-data is scrolling down) is 13:04:50, the time of the next moment data (the next action sub-data is scrolling down) is 13:05:20, the time of the next moment data (the next action sub-data is scrolling up) is 13:05:50, and the time of the next moment data (the next action sub-data is referencing) is 13:06:30.
[0088] Input the action sub-data (scroll down) into the pre-trained model; use the scroll down action as the first answer of the output result of the pre-trained model, the scroll up action as the second answer, and the reference as the third answer.
[0089] In some embodiments, training the pre-trained model based on the training data includes:
[0090] The search data of the qualified search action data is input into the pre-trained model;
[0091] At least one of the aforementioned action data is used as the output of the pre-trained model, and the action data includes at least one of referencing, scrolling up, and scrolling down.
[0092] Specifically, for example, the search data (search term 4G) is input into the pre-trained model; the scroll-down action is used as the first answer of the output of the pre-trained model, the scroll-up action is used as the second answer, and the reference is used as the third answer.
[0093] In some embodiments, training the pre-trained model based on preset feedback data includes:
[0094] The search data or the action data of the qualified search action data are input into the pre-trained model;
[0095] The preset action data is used as the output of the pre-trained model, and the preset action data includes at least one of referencing, scrolling up, and scrolling down.
[0096] Specifically, for example, the search data (search term 4G) is input into the pre-trained model; and the pre-set answers (reference action is the first answer, scroll up is the second answer, and reference is the third answer) are used as the output of the trained model.
[0097] Alternatively, input the action data (scroll down) corresponding to the search data into the preset training model; and use the pre-set answers (reference action as the first answer, scroll up as the second answer, and reference as the third answer) as the output of the training model.
[0098] In some embodiments, the method further includes: verifying the target training model based on preset test data and preset answers to determine whether the target training model is qualified.
[0099] Specifically, for example, the preset test data includes the search term "4G" and the action data "scroll down" and "reference". The preset answer includes the search term "4G" and the action data includes "scroll down" and "reference". The preset test data is input into the target training model, and the output of the target training model is compared with the preset answer to obtain a comparison value. Based on the comparison value, it is determined whether the target training model is qualified.
[0100] In some embodiments, the step of validating the target training model based on preset test data and preset answers to determine whether the target training model is qualified includes:
[0101] Input the preset test data into the target training model to obtain the output result;
[0102] The output result is compared with the preset answer to obtain the comparison value;
[0103] If the comparison value reaches a preset comparison value, the target training model is determined to be qualified.
[0104] Specifically, for example, the search term 4G of the preset test data is input into the target training model, and the output results are scroll down and reference. The output results are compared with the preset answer (scroll down and reference), and the comparison value is 100%. The comparison value is equal to the preset threshold. Therefore, the output result of the target training model is consistent with the preset answer, and the target training model is determined to be qualified.
[0105] Furthermore, "scroll up" and "scroll down" refer to the action of scrolling the page window vertically. "Quote" refers to clicking on related articles displayed on the page.
[0106] It should be noted that the embodiments of this application can also be further described in the following ways:
[0107] Select the target audience for intent analysis, generate keywords, and input the keywords into the target training model to perform intent analysis:
[0108] The target training model continuously outputs results based on the input keywords (its output results are the process of a search engine, i.e., each action data, such as at least one of the following: an upper scroll window, a lower scroll window, and a cited article), and the target training model outputs the final result.
[0109] It should be noted that the method in this embodiment can be executed by a single device, such as a computer or server. The method can also be applied in a distributed scenario, where multiple devices cooperate to complete the task. In such a distributed scenario, one of these devices may execute only one or more steps of the method in this embodiment, and the multiple devices will interact with each other to complete the method described.
[0110] It should be noted that the above description describes some embodiments of this application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the above embodiments and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0111] Based on the same inventive concept, and corresponding to any of the above embodiments, this application also provides a large-scale model training device for intent analysis based on web page search.
[0112] refer to Figure 2 The aforementioned large-scale model training device for intent analysis based on web search includes:
[0113] Module 201 is used to acquire the original dataset and neural network model;
[0114] The first training module 202 is used to perform self-supervised training on the neural network model based on the original data to obtain a pre-trained model.
[0115] The generation module 203 is used to acquire search records, annotate them, and generate training data;
[0116] The second training module 204 is used to train the pre-trained model based on the training data and preset feedback data to obtain the target training model.
[0117] For ease of description, the above devices are described in terms of function, divided into various modules. Of course, in implementing this application, the functions of each module can be implemented in one or more software and / or hardware.
[0118] The apparatus described above is used to implement a large model training method for intent analysis based on web page search in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0119] Based on the same inventive concept, corresponding to any of the above embodiments, this application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements a web search-based intent analysis large model training method as described in any of the above embodiments.
[0120] Figure 3 This embodiment illustrates a more specific hardware structure of an electronic device, which may include a processor 1010, a memory 1020, an input / output interface 1030, a communication interface 1040, and a bus 1050. The processor 1010, memory 1020, input / output interface 1030, and communication interface 1040 are interconnected internally via the bus 1050.
[0121] The processor 1010 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this specification.
[0122] The memory 1020 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 1020 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented by software or firmware, the relevant program code is stored in the memory 1020 and is called and executed by the processor 1010.
[0123] The input / output interface 1030 is used to connect input / output modules to realize information input and output. Input / output modules can be configured as components within the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Input devices may include keyboards, mice, touchscreens, microphones, various sensors, etc., while output devices may include displays, speakers, vibrators, indicator lights, etc.
[0124] The communication interface 1040 is used to connect a communication module (not shown in the figure) to enable communication between this device and other devices. The communication module can communicate via wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0125] Bus 1050 includes a pathway for transmitting information between various components of the device, such as processor 1010, memory 1020, input / output interface 1030, and communication interface 1040.
[0126] It should be noted that although the above-described device only shows the processor 1010, memory 1020, input / output interface 1030, communication interface 1040, and bus 1050, in specific implementations, the device may also include other components necessary for normal operation. Furthermore, those skilled in the art will understand that the above-described device may only include the components necessary for implementing the embodiments of this specification, and not necessarily all the components shown in the figures.
[0127] The electronic device described in the above embodiments is used to implement a large model training method for intent analysis based on web page search in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0128] Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this application also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute a large-scale intent analysis model training method based on web page search as described in any of the above embodiments.
[0129] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.
[0130] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to execute a large model training method for intent analysis based on web page search as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0131] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of this application (including the claims) is limited to these examples; within the framework of this application, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the embodiments of this application as described above, which are not provided in the details for the sake of brevity.
[0132] Additionally, to simplify the description and discussion, and to avoid obscuring the embodiments of this application, the well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Furthermore, the apparatus may be shown in block diagram form to avoid obscuring the embodiments of this application, and this also takes into account the fact that the details of the implementation of these block diagram apparatuses are highly dependent on the platform on which the embodiments of this application will be implemented (i.e., these details should be fully understood by those skilled in the art). While specific details (e.g., circuits) have been set forth to describe exemplary embodiments of this application, it will be apparent to those skilled in the art that the embodiments of this application can be implemented without these specific details or with variations thereof. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0133] Although this application has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.
[0134] The embodiments of this application are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the embodiments of this application should be included within the protection scope of this application.
Claims
1. A method for training an intent analysis large model based on web search, characterized in that, include: Obtain the original dataset and the neural network model; wherein the original dataset is text data and the neural network model is a large model; The neural network model is self-supervised training is performed based on the original data to obtain a pre-trained model. The process involves acquiring and labeling search records to generate training data. The search records include multiple search action data points. Acquiring and labeling these records to generate training data includes: acquiring multiple search action data points; comparing these data points with preset search action data to obtain a comparison result; labeling the search action data points as qualified search action data points when the comparison result reaches a preset threshold; each qualified search action data point includes search data and at least one action data point; and using all qualified search action data points as training data; the action data points are user browsing actions. The pre-trained model is trained based on the training data and preset feedback data to obtain the target training model, wherein the preset feedback data includes preset search data and preset action data.
2. The method according to claim 1, characterized in that, The step of training the pre-trained model based on the training data and preset feedback data to obtain the target training model includes: The pre-trained model is iteratively trained based on the training data to obtain a first training model; The pre-trained model is iteratively trained based on preset feedback data to obtain a second training model; The first and second training models are combined and trained using a reinforcement learning algorithm. The process of training the first training model, the second training model, and the combined training of the two models is iterated. When the number of iterations reaches the iteration threshold, the iterative training stops and the target training model is generated.
3. The method according to claim 1, characterized in that, The step of training the pre-trained model based on the training data and preset feedback data to obtain the target training model includes: The pre-trained model is trained based on the training data to obtain a first training model; The pre-trained model is trained based on preset feedback data to obtain a second trained model; The first training model and the second training model are combined and trained using a reinforcement learning algorithm. The process of training the first training model, the second training model, and the combined training of the two models is iteratively trained. Upon receiving a termination command, the iterative training stops, and the target training model is generated.
4. The method according to claim 2 or 3, characterized in that, The step of training the pre-trained model based on the training data includes: The action data includes multiple action sub-data with time tags, and the multiple action sub-data includes a current time data, and at least one action sub-data located after the current time data is the next time data; Input an action sub-data as the current time-time data into the pre-trained model; The next time step data corresponding to the current time step data is used as the output of the pre-trained model, and the action sub-data is one of referencing, scrolling up, or scrolling down.
5. The method according to claim 2 or 3, characterized in that, The step of training the pre-trained model based on the training data includes: The search data of the qualified search action data is input into the pre-trained model; At least one of the aforementioned action data is used as the output of the pre-trained model, and the action data includes at least one of referencing, scrolling up, and scrolling down.
6. The method according to claim 2 or 3, characterized in that, The step of training the pre-trained model based on preset feedback data includes: The search data or the action data of the qualified search action data are input into the pre-trained model; The preset action data is used as the output of the pre-trained model, and the preset action data includes at least one of referencing, scrolling up, and scrolling down.
7. The method according to claim 1, characterized in that, The method further includes: The target training model is validated based on preset test data and preset answers to determine whether the target training model is qualified.
8. The method according to claim 7, characterized in that, The step of validating the target training model based on preset test data and preset answers to determine whether the target training model is qualified includes: Input the preset test data into the target training model to obtain the output result; The output result is compared with the preset answer to obtain the comparison value; If the comparison value reaches a preset comparison value, the target training model is determined to be qualified.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Model training method and device and intention recognition method and device
CN113076080A
Intention recognition method and device, equipment and medium
CN113360751A