Major language model-based material scheduling method and device, and medium
By using material scheduling methods based on large language models in logistics scheduling, the problems of complex resource allocation and real-time decision-making in logistics scheduling are solved, and efficient and personalized scheduling decisions and efficiency improvement are achieved.
Patent Information
- Application Number
- CN202510102818.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-01-22
AI Technical Summary
The prior art is difficult to effectively handle complex resource allocation, path optimization and real-time decision-making in logistics scheduling, especially when data complexity and real-time requirements are high.
The material scheduling method based on the large language model is adopted, and the material scheduling instruction data set is constructed, and the pre-trained large language model is supervised and fine-tuned, and a reward model is created to calculate the preference sorting loss, and the target large language model is finally obtained.
Effectively align the output probability of the large language model with human preferences, significantly reduce training costs, and form a user-friendly logistics scheduling large language model, which can provide accurate and personalized scheduling decisions, and automatically integrate multiple data sources to improve efficiency.
Smart Images

Figure CN120069400A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of logistics scheduling problems, and in particular, to a material scheduling method, device and medium based on a large language model. Background Art
[0002] In recent years, the application of large language models in logistics scheduling has attracted wide attention. Logistics scheduling involves complex challenges such as resource allocation, path optimization, and real-time decision-making. With the expansion of the logistics network scale and the increase in data complexity, traditional scheduling methods often fall short in dealing with these problems. Large language models can process multi-modal data and understand and generate complex scheduling strategies through natural language processing technology, providing a new intelligent solution for logistics scheduling.
[0003] The application of large language models in logistics scheduling is mainly reflected in the following aspects. First, large language models can process and understand a large amount of text data, such as order instructions, customer feedback, and logistics contracts, and assist scheduling decisions by converting this unstructured data into actionable information. Second, by analyzing historical data and real-time data, large language models can generate optimized scheduling strategies, thereby improving resource utilization and transportation efficiency. In addition, large language models can also quickly adjust scheduling plans according to emergencies (such as traffic jams or equipment failures) through a real-time response mechanism to ensure the stability of the logistics system.
[0004] Although the application of large language models in logistics scheduling has broad prospects, it still faces challenges in data complexity and real-time requirements. Effectively combining deep learning and reinforcement learning technologies to optimize the performance of large language models in logistics scheduling has become an important research direction. Through continuous technological innovation and practical exploration, large language models are expected to further improve the intelligent level of logistics scheduling and promote the digital transformation and efficiency improvement of the logistics industry. Summary of the Invention
[0005] To at least to some extent solve one of the technical problems existing in the prior art, an object of the present invention is to provide a material scheduling method, system, electronic device and storage medium based on a large language model and preference ranking loss.
[0006] The first technical solution adopted by the present invention is:
[0007] A material scheduling method based on a large language model, comprising the following steps:
[0008] Data collection and preprocessing;
[0009] Construct a material scheduling instruction dataset based on preset material information, environmental information, and response information;
[0010] Construct an initial large language model;
[0011] Pre-train the initial large language model on an unlabeled text dataset to obtain a pre-trained large language model;
[0012] Perform supervised fine-tuning on the pre-trained large language model on a material scheduling instruction dataset to obtain a supervised fine-tuning model;
[0013] On the basis of the supervised fine-tuning model, create a reward model to calculate the preference ranking loss;
[0014] Construct a material scheduling response dataset;
[0015] Train the supervised fine-tuning model on the material scheduling response dataset to obtain a target large prediction model;
[0016] Communicate with the target large prediction model through natural language to obtain a logistics scheduling response result.
[0017] Further, the data collection and preprocessing include:
[0018] Collect data related to logistics scheduling, such as material information, environmental information, response information, etc.;
[0019] Filter and clean the collected data.
[0020] Further, the material scheduling instruction dataset is a manually labeled dataset; the input of the model is an instruction, and the output is the expected answer of the model.
[0021] Further, the pre-training of the initial large language model on the unlabeled text dataset includes:
[0022] The training objective of the initial large language model is for the model to predict the next word based on the provided text.
[0023] Further, the supervised fine-tuning of the pre-trained large language model on the material scheduling instruction dataset includes:
[0024] Use the material scheduling instruction dataset to fine-tune the pre-trained large language model in a supervised manner;
[0025] The training objective of the pre-trained large language model is for the model to predict the expected answer based on the provided instruction.
[0026] Further, the architecture of the reward model is to splice a regression layer on the basis of the supervised fine-tuning model;
[0027] The reward model is obtained through training. Its input is the answers of the supervised fine-tuning model to a number of instructions, and its output is the reward score for each reply content.
[0028] The expected value of the reward score is obtained by the annotator ranking the answers of the supervised fine-tuning model according to personal preferences.
[0029] The training objective of the reward model is to fit the ranking of the annotator for the reply content of the supervised fine-tuning model.
[0030] Further, for the material scheduling response dataset, answers are generated for the instructions using, including but not limited to, the supervised fine-tuning model, other large models, and humans, and then the answers are scored and ranked by the reward model or humans.
[0031] Further, training the supervised fine-tuning model on the material scheduling response dataset to obtain the target large prediction model includes:
[0032] Taking the instructions in the material scheduling response dataset and the answers from various data sources as inputs, and feeding them into the supervised fine-tuning model and the reward model respectively.
[0033] The reward model outputs the score of the answer, and the supervised fine-tuning model outputs the log probability of the answer as an evaluation.
[0034] Calculate the preference ranking loss between the output of the reward model and the output of the supervised fine-tuning model, backpropagate the gradient, and train the supervised fine-tuning model until convergence to obtain the target large language model.
[0035] Further, the specific process of calculating the preference ranking loss is as follows:
[0036] To align the finally trained target large language model π(y|x) with human preferences, first calculate the score for each answer:
[0037]
[0038] In the formula, x is the input instruction, y i is the answer generated by the i-th model or human, y i,t is the t-th token of y i is the first t - 1 tokens of y i,<t is y i is the number of tokens of y i ||y i || is the number of tokens of y π t is the number of tokens, and logP
[0039] After obtaining the score for each answer, combined with the true ranking, obtain the ranking loss:
[0040]
[0041] wherein, r i is the true score of the i-th answer, and r j is the true score of the j-th answer, and p i is the predicted score of the i-th answer, and p j is the predicted score of the j-th answer;
[0042] where the true ranking is obtained from the true scores;
[0043] To improve the generation quality and answer diversity, the cross-entropy loss is calculated:
[0044] {i′ 1 , i′ 2 ,..., i′ k} = topkr i
[0045]
[0046] wherein, {i′ 1 , i′ 2 ,..., i′ k} are the top k best answers; k is a hyperparameter set according to experience, is the t-th token of the i′ j -th generated answer and is 's first t tokens; where j ∈ {1, 2,..., k};
[0047] The ranking loss and the cross-entropy loss are added together to obtain the preference ranking loss:
[0048] L = L rank + L ft .
[0049] The second technical solution adopted by the present invention is:
[0050] A material scheduling system based on a large language model, including a data acquisition module, a model construction module, a pre-training module, a supervised fine-tuning module, a preference alignment module, and a result output module;
[0051] The data acquisition module is used to acquire an unannotated text dataset, a material scheduling instruction dataset, and a material scheduling response dataset; the unannotated text dataset is any open-source text dataset; the material scheduling instruction dataset constructs the material scheduling response dataset based on preset material information, environmental information, and response information; use, including but not limited to, supervised fine-tuning models, other large models, and humans to generate answers for the instructions, and then use a reward model or humans to score and rank the answers;
[0052] The model construction module is used to construct an initial large language model and a reward model;
[0053] The pre-training module is used to pre-train the initial large language model on the unannotated text dataset to obtain a pre-trained large language model;
[0054] The supervised fine-tuning module is used to perform supervised fine-tuning on the pre-trained large language model on the material scheduling instruction dataset to obtain a supervised fine-tuning model;
[0055] The preference alignment module is used to supervise and train the supervised fine-tuning model with the reward model to obtain a target large prediction model;
[0056] The result output module is used to communicate with the target large prediction model through natural language to obtain a logistics scheduling response result.
[0057] The third technical solution adopted by the present invention is:
[0058] An electronic device, the electronic device includes a processor and a memory, and at least one instruction, at least one program, a code set, or an instruction set is stored in the memory, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement a material scheduling method based on a large language model as described above.
[0059] The fourth technical solution adopted by the present invention is:
[0060] A computer-readable storage medium, and at least one instruction, at least one program, a code set, or an instruction set is stored in the storage medium, and the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement a material scheduling method based on a large language model as described above.
[0061] The fifth technical solution adopted by the present invention is:
[0062] A computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to execute the above-mentioned material scheduling method based on a large language model.
[0063] The beneficial effects of the present invention are as follows: Firstly, by introducing preference ranking loss, the present invention can effectively align the output probability of the large language model with human preferences, while significantly reducing the training cost, and finally form a user-friendly large language model for logistics scheduling to meet various practical application requirements; Secondly, it can interact with users through natural language, understand complex context information, and thus provide accurate and personalized scheduling decisions; Finally, it can automatically integrate multiple data sources, adjust the scheduling strategy in real time, improve efficiency and reduce costs, and continuously optimize the decision-making process based on historical data. This combination enables the large language model to show great potential in enhancing the flexibility and accuracy of logistics scheduling. BRIEF DESCRIPTION OF THE DRAWINGS
[0064] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following introduces the relevant technical solution drawings in the embodiments of the present invention or the prior art. It should be understood that the drawings introduced below are only for conveniently and clearly presenting some embodiments of the technical solutions in the present invention. For those skilled in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0065] Figure 1 It is a schematic flowchart of the material scheduling method based on a large language model and preference ranking loss in the embodiments of the present invention;
[0066] Figure 2 It is a step flowchart of the material scheduling method based on a large language model and preference ranking loss in the embodiments of the present invention;
[0067] Figure 3 It is a schematic structural diagram of the material scheduling system based on a large language model and preference ranking loss in the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0068] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where like or similar reference numerals denote like or similar elements or elements having like or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary only for explaining the present invention and should not be construed as limiting the present invention. For the step numbers in the following embodiments, they are only set for the convenience of description and illustration, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.
[0069] In the description of the present invention, it should be understood that with respect to the orientation description, such as the orientations or positional relationships indicated by up, down, front, back, left, right, etc., are based on the orientations or positional relationships shown in the accompanying drawings. This is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention.
[0070] In the description of the present invention, the meaning of "several" is one or more, the meaning of "multiple" is more than two, and understandings such as "greater than", "less than", "exceeding", etc. do not include the recited number, while understandings such as "above", "below", "within", etc. include the recited number. If there is a description of "first" and "second", it is only for the purpose of distinguishing technical features and should not be construed as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features or the sequence of the indicated technical features.
[0071] In the description of the present invention, unless otherwise clearly defined, words such as "set", "installed", "connected", etc. should be understood in a broad sense, and those skilled in the art can reasonably determine the specific meanings of the above words in the present invention in combination with the specific content of the technical solution.
[0072] The main challenges faced in the process of material scheduling include complex and changeable situations, low efficiency, and inaccurate decision-making. These difficulties make it often difficult for traditional scheduling methods to cope with the dynamic changes in demand, the finiteness of resources, and complex constraints, resulting in low scheduling efficiency and increased costs. While large language models can process large-scale data and learn the potential patterns and rules from it, thus better understanding and predicting scheduling requirements in different situations. In view of the existing technical problems, the present invention proposes a material scheduling scheme based on a large language model and preference ranking loss. This scheme obtains a supervised fine-tuning model by constructing a material scheduling instruction dataset and performing supervised fine-tuning on a pre-trained large language model. On this basis, a reward model is created, and a material scheduling response dataset is constructed. On this dataset, the supervised fine-tuning model is further trained using the reward model to finally obtain the target large language model. By communicating with the target large language model in natural language, the logistics scheduling response result can be obtained. The present invention effectively aligns the output probability of the large language model with human preferences through preference ranking loss, significantly reducing the training cost, and finally forming a user-friendly large language model for logistics scheduling.
[0073] Embodiment 1
[0074] As Figure 1 and Figure 2 shown, this embodiment proposes a material scheduling method based on a large language model and preference ranking loss, including the following steps:
[0075] S1. Data collection and preprocessing.
[0076] Specifically, collect data related to logistics scheduling, such as material information, environmental information, response information, etc.; filter and clean the collected data.
[0077] S2. Construct a material scheduling instruction dataset based on preset material information, environmental information, and response information.
[0078] In this embodiment, the material scheduling instruction dataset is a dataset that needs to be manually annotated; the input of the model is an instruction, and the output is the expected answer of the model.
[0079] S3. Construct an initial large language model, and pre-train the initial large language model on an unannotated text dataset to obtain a pre-trained large language model.
[0080] In some embodiments, the unannotated text dataset contains billions to trillions of tokens. The training objective of the initial large language model is for the model to predict the next word based on the provided text.
[0081] S4. Perform supervised fine-tuning on the pre-trained large language model using the material scheduling instruction dataset to obtain a supervised fine-tuning model.
[0082] As an alternative implementation, step S4 includes:
[0083] S41. Fine-tune the pre-trained large language model in a supervised manner using the material scheduling instruction dataset;
[0084] S42. The training objective of the pre-trained large language model is for the model to predict the expected answer based on the provided instructions.
[0085] S5. Based on the supervised fine-tuning model, create a reward model to calculate the preference ranking loss.
[0086] In some embodiments, the reward model architecture concatenates a regression layer on the basis of a copy of the supervised fine-tuning language model. The reward model is obtained through training, with its input being the answers of the supervised fine-tuning model to a number of instructions, and its output being the reward scores for each reply content. The expected value of the reward scores is obtained by annotators ranking the answers of the above-mentioned supervised fine-tuning model according to their personal preferences. The training objective of the reward model is to fit the ranking of the reply content of the supervised fine-tuning model by the annotators.
[0087] S6. Construct a material scheduling response dataset, and train the supervised fine-tuning model on the material scheduling response dataset to obtain the target large prediction model.
[0088] In this embodiment, for the material scheduling response dataset, answers are generated for the instructions using, including but not limited to, the supervised fine-tuning model, other large models, and humans, and then the answers are scored and ranked by the reward model or humans.
[0089] In some embodiments, step S6 specifically includes the following steps:
[0090] S61. The instructions in the material scheduling response dataset and the answers from various data sources are used as inputs and fed into the supervised fine-tuning model and the reward model respectively;
[0091] S62. The reward model outputs the scores of the answers, and the supervised fine-tuning model outputs the log probability of the answers as an evaluation;
[0092] S63. Calculate the preference ranking loss between the output of the reward model and the output of the supervised fine-tuning model, backpropagate the gradients, and train the supervised fine-tuning model until convergence to obtain the target large language model.
[0093] Furthermore, the process of calculating the preference ranking loss is specifically as follows:
[0094]
[0095] Among them, x is the input instruction, and y i is the answer generated by the i-th model or human, and y i,t is the t-th token of y i and y i,<t is the t-th token of y i is the first t - 1 tokens of y, ||y i || is the number of tokens of y, logP i is the log probability; π is the log probability;
[0096] After obtaining the score p of each answer i and combining with the true ranking, the ranking loss is obtained:
[0097]
[0098] where the true ranking is obtained through true scoring;
[0099] To improve the generation quality and answer diversity, the cross-entropy loss is calculated:
[0100] {i′ 1 ,i′ 2 ,...,i′ k}=topk r i #(3)
[0101]
[0102] where r i is the score of the i-th answer, and {i′ 1 ,i′ 2 ,...,i′ k} are the top k best answers;
[0103] The ranking loss and the cross-entropy loss are added together to obtain the total loss:
[0104] L=L rank +L ft #(5)
[0105] S7. Communicate with the target large prediction model through natural language to obtain the logistics scheduling response result.
[0106] In summary, the present invention addresses the problems of complex and changeable situations, low efficiency, and inaccurate decision-making in the process of material scheduling. It proposes a material scheduling method based on a large language model and preference ranking loss. First, by introducing preference ranking loss, it effectively aligns the output probability of the large language model with human preferences, significantly reducing the training cost. Finally, a user-friendly large language model for logistics scheduling is formed to meet various practical application requirements. Second, it can interact with users through natural language, understand complex context information, and thus provide accurate and personalized scheduling decisions. Finally, it can automatically integrate multiple data sources, adjust scheduling strategies in real time, improve efficiency and reduce costs, and continuously optimize the decision-making process based on historical data. This combination enables the large language model to show great potential in improving the flexibility and accuracy of logistics scheduling.
[0107] Embodiment 2
[0108] As Figure 3 shown, this embodiment provides a material scheduling system based on a large language model and preference ranking loss, including a data acquisition module, a model construction module, a pre-training module, a supervised fine-tuning module, a preference alignment module, and a result output module;
[0109] The data acquisition module is used to acquire an unlabeled text data set, a material scheduling instruction data set, and a material scheduling response data set; the unlabeled text data set is any open-source text data set; the material scheduling instruction data set is constructed based on preset material information, environmental information, and response information to construct the material scheduling response data set; use, including but not limited to, a supervised fine-tuning model, other large models, and humans to generate answers for the instructions, and then use a reward model or humans to score and rank the answers;
[0110] The model construction module is used to construct an initial large language model and a reward model;
[0111] The pre-training module is used to pre-train the initial large language model on the unlabeled text data set to obtain a pre-trained large language model;
[0112] The supervised fine-tuning module is used to perform supervised fine-tuning on the pre-trained large language model on the material scheduling instruction data set to obtain a supervised fine-tuning model;
[0113] The preference alignment module is used to supervise and train the supervised fine-tuning model with the reward model to obtain a target large prediction model;
[0114] The result output module is used to communicate with the target large prediction model through natural language to obtain a logistics scheduling response result.
[0115] It should be noted that the material scheduling system based on large language model and preference ranking loss of the present invention corresponds one to one with the material scheduling method based on large language model and preference ranking loss of the present invention. The technical features and beneficial effects described in the above-mentioned embodiment of the material scheduling method based on large language model and preference ranking loss are applicable to the embodiment of the material scheduling system based on large language model and preference ranking loss. For specific contents, please refer to the description in the embodiment of the method of the present invention, which will not be repeated here. This is hereby declared.
[0116] In addition, in the implementation of the material scheduling based on large language model and preference ranking loss in the above embodiment, the logical division of each program module is only an example. In actual application, the above functions can be assigned to different program modules as needed, for example, for the convenience of corresponding hardware configuration requirements or software implementation. That is, the internal structure of the material scheduling system based on large language model and preference ranking loss is divided into different program modules to complete all or part of the functions described above.
[0117] Example 3
[0118] An embodiment of the present invention further provides an electronic device, the electronic device comprising a processor and a memory, the memory storing at least one instruction, at least one program, a code set or an instruction set, the at least one instruction, the at least one program, the code set or the instruction set being loaded and executed by the processor to implement the following Figure 1 and Figure 2 A material scheduling method based on a large language model and preference ranking loss is shown.
[0119] It is understood that the memory may include a random access memory (RAM) or a read-only memory (ROM). Optionally, the memory includes a non-transitory computer-readable storage medium. The memory may be used to store instructions, programs, codes, code sets, or instruction sets. The memory may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function, instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store data created according to the use of the server, etc.
[0120] The processor may include one or more processing cores. The processor uses various interfaces and circuits to connect various parts within the entire server. By running or executing instructions, programs, code sets, or instruction sets stored in the memory, and by invoking data stored in the memory, it performs various functions of the server and processes data. Optionally, the processor may be implemented in at least one of the hardware forms of digital signal processing (DSP), field-programmable gate array (FPGA), and programmable logic array (PLA). The processor may integrate a combination of one or more of a central processing unit (CPU) and a modem, etc. Among them, the CPU mainly processes the operating system and application programs, etc.; the modem is used for processing wireless communications. It can be understood that the above-mentioned modem may not be integrated into the processor and may be implemented separately by a single chip.
[0121] Since this electronic device is an electronic device corresponding to a material scheduling method based on a large language model and preference ranking loss in an embodiment of the present invention, and the principle of how this electronic device solves problems is similar to that of this method, the implementation of this electronic device can refer to the implementation process of the above method embodiment, and repeated parts will not be elaborated.
[0122] Embodiment 4
[0123] An embodiment of the present invention further provides a computer-readable storage medium, in which at least one instruction, at least one segment of program, code set, or instruction set is stored, and the at least one instruction, the at least one segment of program, the code set, or the instruction set is loaded and executed by a processor to implement a material scheduling method based on a large language model and preference ranking loss as Figure 1 and Figure 2 shown.
[0124] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program, and this program can be stored in a computer-readable storage medium. The storage medium includes read-only memory (ROM), random access memory (RAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), one-time programmable read-only memory (OTPROM), electrically-erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM), or other optical disc memories, magnetic disk memories, tape memories, or any other computer-readable medium capable of carrying or storing data.
[0125] Since this storage medium is the storage medium corresponding to a material scheduling method based on a large language model and preference ranking loss in the embodiments of the present invention, and the principle of solving problems by this storage medium is similar to that of this method, the implementation of this storage medium can refer to the implementation process of the above method embodiments, and the repeated parts will not be described again.
[0126] Embodiment 5
[0127] In some possible implementation manners, various aspects of the method in the embodiments of the present invention can also be implemented in the form of a program product, which includes program code. When the program product runs on a computer device, the program code is used to cause the computer device to execute the steps of a material scheduling method based on a large language model and preference ranking loss according to various exemplary implementation manners described above in this specification. Among them, the executable computer program code or "code" for executing each embodiment can be written in a high-level programming language such as C, C++, C#, Smalltalk, Java, JavaScript, Visual Basic, structured query language (e.g., Transact-SQL), Perl, or in various other programming languages.
[0128] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one of the following techniques known in the art or a combination thereof can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0129] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0130] The above embodiments are only for illustrating the technical concept and features of the present invention, and the purpose is to enable those of ordinary skill in the art to understand the content of the present invention and implement it accordingly, and it cannot be used to limit the protection scope of the present invention. Any equivalent changes or modifications made according to the essence of the content of the present invention should be covered within the protection scope of the present invention.
Claims
1. A material dispatching method based on a large language model, characterized in that: The following steps are involved: Data collection and preprocessing; Construct a material dispatch instruction dataset; Build an initial large language model; Pre-training the initial large language model on an unlabeled text dataset to obtain a pre-trained large language model; Performing supervised fine-tuning on the pre-trained large language model on the material dispatch instruction dataset to obtain a supervised fine-tuning model; Based on the supervised fine-tuning model, create a reward model to calculate the preference ranking loss; Construct material dispatch response data set; A supervised fine-tuning model is trained on the material dispatch response dataset to obtain the target large prediction model; By communicating with the target large prediction model through natural language, the logistics scheduling response result is obtained.
2. A material dispatching method based on a large language model according to claim 1, characterized in that: The data collection and preprocessing include: Collect data related to logistics scheduling; Filter and clean the collected data.
3. The material dispatching method based on a large language model according to claim 1, characterized in that: The material dispatch instruction dataset is a manually annotated dataset; the input of the model is an instruction, and the output is the expected answer of the model.
4. The material dispatching method based on a large language model according to claim 1, characterized in that: The pre-training of the initial large language model on the unlabeled text dataset includes: The training goal of the initial large language model is for the model to predict the next word based on the provided text.
5. The material dispatching method based on a large language model according to claim 1, characterized in that: The supervised fine-tuning of the pre-trained large language model on the material dispatch instruction dataset includes: Fine-tune the pre-trained large language model in a supervised manner using a material dispatch instruction dataset; The training goal of the pre-trained large language model is for the model to predict the expected answer based on the provided instructions.
6. The material dispatching method based on a large language model according to claim 1, characterized in that: The architecture of the reward model is to splice a regression layer on the basis of the supervised fine-tuning model; The reward model is obtained through training, and its input is the answers of the supervised fine-tuning model to several instructions, and its output is the reward score of each reply content; The expected value of the reward score is obtained by ranking the answers of the supervised fine-tuning model according to the annotator's personal preferences; The training objective of the reward model is to fit the ranking of the response content of the annotator to the supervised fine-tuning model.
7. The material dispatching method based on a large language model according to claim 1, characterized in that: The supervised fine-tuning model is trained on the material dispatch response data set to obtain the target large prediction model, including: The instructions in the material dispatch response dataset and the answers from various data sources are used as input to the supervised fine-tuning model and the reward model respectively; The reward model outputs the score of the answer, and the supervised fine-tuning model outputs the probability of the answer as an evaluation; Calculate the preference ranking loss of the reward model output and the supervised fine-tuning model output, back-propagate the gradient, train the supervised fine-tuning model until convergence, and obtain the target large language model.
8. The material dispatching method based on a large language model according to claim 1, characterized in that: The preference ranking loss calculation process is specifically as follows: In order to align the final trained target language model π(y|x) with human preferences, we first calculate the score of each answer: In the formula, x is the input instruction, y is i is the i-th generated answer, y i,t for y i The tth word, y i,<t for y i The first t-1 tokens of , ‖y i ‖ is y i The number of lemmas, t is the number of lemmas, logP π is the log probability; After getting the score of each answer, combined with the actual ranking, we get the ranking loss: In the formula, r i is the true score of the ith answer, r j is the true score of the jth answer, p i is the predicted score of the ith answer, p j is the predicted score of the jth answer; In order to improve the generation quality and answer diversity, the cross entropy loss is calculated: {i′1,i′2,…,i′ k }=topkri i In the formula, {i′1,i′2,…,i′ k } are the top k best answers; k is a hyperparameter, is the i′th j Generated answers The tth word of yes The first t words of ; where j∈{1,2,…,k}; The ranking loss is added to the cross entropy loss to obtain the preference ranking loss: L=L rank +L ft .
9. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the memory stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the method described in any one of claims 1 to 8.
10. A computer-readable storage medium, characterized in that: The storage medium stores at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Skillroute labour integration solution
CA2555116A1
Self-supervised visual language navigation pre-training method and device and storage medium
CN116168333A
Cloud-edge collaborative big language model intelligent customer service deployment optimization method
CN117808481A
Reinforcement learning model training method and device
CN118428442A
Systems and methods for predicting patient recruitment at clinical sites
US20240221874A1