Method and apparatus for evaluating a large logistics model
By acquiring a test set in the logistics field and conducting multiple rounds of problem-solving, and combining semantic entropy information to evaluate a large-scale logistics model, the problem of the accuracy and reliability of the logistics model in specific tasks is solved, and a more efficient evaluation method is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SF TECH CO LTD
- Filing Date
- 2024-11-27
- Publication Date
- 2026-05-29
AI Technical Summary
Large-scale vertical models in the logistics field are difficult to apply to specific tasks, and lack effective evaluation methods.
By acquiring test sets configured for each stage of the logistics process, inputting them into a large logistics model for multiple rounds of problem-solving, and using semantic entropy information for similarity comparison, entropy change parameters are determined to evaluate model performance.
This improves the accuracy of large-scale logistics model evaluation and ensures the stability and accuracy of the model when handling logistics tasks.
Smart Images

Figure CN122111804A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of computer application technology, specifically to the evaluation technology of large-scale logistics models in the field of computer application technology, and more specifically to the evaluation method and related apparatus for large-scale logistics models. Background Technology
[0002] The data in the logistics field is complex and diverse, involving multiple links such as transportation, warehousing, and distribution. Large-scale vertical models in the logistics field need to learn knowledge from each link. However, whether they can be applied in the end and help solve business problems requires targeted evaluation. Therefore, how to design a special evaluation method to ensure the accuracy and reliability of the model when handling these specific tasks has become a challenge. Summary of the Invention
[0003] This specification provides an evaluation method and related apparatus for large-scale logistics models, improving the accuracy of large-scale logistics model evaluation.
[0004] To achieve the above technical objectives, the embodiments of this specification provide the following technical solutions:
[0005] Firstly, one embodiment of this specification provides a method for evaluating a large-scale logistics model, comprising:
[0006] Obtain test sets configured for each stage of the logistics process;
[0007] The test questions in the test set are input into the logistics big model to obtain answer information through multiple rounds of answering;
[0008] The probability parameters are obtained by comparing the words in the answer information with the standard answers corresponding to the test questions, and the semantic entropy information is obtained by combining the probability parameters.
[0009] The entropy change parameter is determined based on the semantic entropy difference corresponding to each test item in the semantic entropy information, and the performance of the logistics big model is evaluated based on the entropy change parameter.
[0010] Secondly, one embodiment of this specification provides an evaluation device for a large-scale logistics model, comprising:
[0011] The acquisition unit is used to acquire test sets configured for various aspects of the logistics field;
[0012] An evaluation unit is used to input the test questions from the test set into the logistics big model to obtain answer information through multiple rounds of answering;
[0013] The evaluation unit is also used to obtain probability parameters by comparing the words in the answer information with the standard answer corresponding to the test question, and to obtain semantic entropy information by combining the probability parameters.
[0014] The evaluation unit is further configured to determine the entropy change parameter based on the semantic entropy difference corresponding to each test item in the semantic entropy information, so as to evaluate the performance of the logistics big model based on the entropy change parameter.
[0015] Optionally, in one possible implementation, the evaluation unit is configured to: determine the selection options corresponding to the test questions when the test questions in the test set are input into the logistics big model for multiple rounds of answering to obtain answer information; adjust the order of the selection options to update the test questions; input the updated test questions into the logistics big model for multiple rounds of answering to obtain answer information, wherein the order of the selection options is different for each round.
[0016] Optionally, in one possible implementation, the evaluation unit is configured to, when inputting test questions from the test set into the logistics big model for multiple rounds of answering to obtain answer information, acquire a logistics knowledge base associated with the test questions; extract logistics knowledge associated with the test questions through the logistics knowledge base; configure the test questions according to the logistics knowledge; input the configured test questions into the logistics big model for multiple rounds of answering to obtain answer information, wherein the logistics knowledge corresponding to each round is different.
[0017] Optionally, in one possible implementation, the evaluation unit is configured to, when obtaining the logistics knowledge base associated with the test question, acquire key information in the test question; perform a logistics domain retrieval based on the key information to obtain domain tags; and determine the logistics knowledge base associated with the domain tags.
[0018] Optionally, in one possible implementation, the evaluation unit is configured to, when obtaining probability parameters by comparing the characters in the answer information with the standard answer corresponding to the test question to obtain semantic entropy information, perform similarity comparisons between the characters in the answer information and the standard answer corresponding to the test question to obtain probability parameters; obtain length information corresponding to the standard answer; and weight the probability parameters based on the length information to obtain the semantic entropy information.
[0019] Optionally, in one possible implementation, the evaluation unit is configured to, when obtaining probability parameters by comparing the words in the answer information with the standard answer corresponding to the test question, acquire a logistics vocabulary configured for the logistics big data model; determine the logistics feature words in the answer information that are associated with the logistics vocabulary; and obtain probability parameters by comparing the logistics feature words in the answer information with the standard answer corresponding to the test question.
[0020] Optionally, in one possible implementation, the evaluation unit is configured to determine the answer option corresponding to the answer information when obtaining probability parameters by comparing the words in the answer information with the standard answer corresponding to the test question; obtain process information corresponding to the logistics big data model in determining the answer option; and obtain probability parameters by comparing the words in the process information with the standard answer corresponding to the test question.
[0021] Optionally, in one possible implementation, the evaluation unit is configured to, when determining the entropy change parameter based on the semantic entropy difference corresponding to each of the test items in the semantic entropy information, and performing performance evaluation on the logistics big model based on the entropy change parameter, determine the extreme semantic entropy corresponding to each of the test items in the semantic entropy information; determine the semantic entropy difference based on the extreme semantic entropy; calculate the mean of the semantic entropy differences corresponding to each of the test items to obtain the entropy change parameter; and perform performance evaluation on the logistics big model based on the entropy change parameter.
[0022] Optionally, in one possible implementation, the evaluation unit is configured to, when performing performance evaluation on the logistics big model based on the entropy change parameter, determine the question type information corresponding to the entropy change parameter; obtain the entropy change threshold corresponding to the question type information; and perform performance evaluation on the logistics big model based on the entropy change threshold.
[0023] Thirdly, one embodiment of this specification also provides a computing device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the evaluation method of the large-scale logistics model as described above.
[0024] Fourthly, one embodiment of this specification also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the evaluation method for the large-scale logistics model as described above.
[0025] Fifthly, embodiments of this specification provide a computer program product or computer program, the computer program product including a computer program that can be stored in a computer-readable storage medium or in the cloud; the processor of the computer device reads the computer program, and when the processor executes the computer program, it implements the steps of the above-described evaluation method for the large-scale logistics model.
[0026] As can be seen from the above technical solution, the logistics large-scale model evaluation method provided in this specification obtains a test set configured for each link in the logistics field; then, the test questions in the test set are input into the logistics large-scale model for multiple rounds of answering to obtain response information; the probability parameters are obtained by comparing the words in the response information with the standard answers corresponding to the test questions, and semantic entropy information is obtained by combining the probability parameters; then, entropy change parameters are determined based on the semantic entropy difference corresponding to each test question in the semantic entropy information, and the performance of the logistics large-scale model is evaluated based on the entropy change parameters. This realizes a logistics large-scale model evaluation process based on semantic entropy. Because the logistics large-scale model is used for multiple rounds of question-and-answering for each logistics link, and the stability of the logistics large-scale model's answers to the same question is evaluated by combining the semantic entropy in the response information, the accuracy of the model evaluation is improved. Attached Figure Description
[0027] To more clearly illustrate the technical solutions in the embodiments or prior art of this specification, the drawings used in the description of the embodiments or prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this specification. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0028] Figure 1 Network architecture diagram for the evaluation system of the large-scale logistics model;
[0029] Figure 2 A flowchart illustrating the evaluation process of a large-scale logistics model provided in this application embodiment;
[0030] Figure 3 A flowchart illustrating an evaluation method for a large-scale logistics model, provided as one embodiment of this specification;
[0031] Figure 4 A schematic diagram illustrating a scenario for an evaluation method of a large-scale logistics model, provided as one embodiment of this specification.
[0032] Figure 5 A schematic diagram illustrating a scenario for an evaluation method of another large-scale logistics model provided as one embodiment of this specification;
[0033] Figure 6A schematic diagram of the functional modules of an evaluation device for a large-scale logistics model provided in one embodiment of this specification;
[0034] Figure 7 This is a schematic diagram of the structure of a computing device provided for one embodiment of this specification. Detailed Implementation
[0035] Unless otherwise defined, the technical or scientific terms used in the embodiments of this specification shall have the ordinary meaning understood by one of ordinary skill in the art to which this specification pertains. The terms "first," "second," and similar terms used in the embodiments of this specification do not indicate any order, quantity, or importance, but are merely used to avoid confusion of constituent elements.
[0036] Unless the context otherwise requires, throughout this specification, "a plurality of" means "at least two," and "including" is interpreted as open-ended or encompassing, that is, "including, but not limited to." In the description of this specification, terms such as "one embodiment," "some embodiments," "exemplary embodiment," "example," "specific example," or "some examples" are intended to indicate that a particular feature, structure, material, or characteristic associated with that embodiment or example is included in at least one embodiment or example of this specification. The illustrative representations of the above terms do not necessarily refer to the same embodiment or example.
[0037] It should be understood that the logistics big data model evaluation method provided in this application can be applied to systems or programs in terminal devices that include logistics big data model evaluation functions, such as logistics management applications. Specifically, the logistics big data model evaluation system can run on systems such as... Figure 1 In the network architecture shown, such as Figure 1 The diagram shown illustrates the network architecture of the logistics big data model evaluation system. As can be seen, the system can provide evaluation processes for logistics big data models from multiple information sources. Specifically, evaluation requests are determined through acquisition operations on the terminal side, triggering the server to perform the corresponding evaluation process. This can be understood as... Figure 1 The diagram illustrates various terminal devices, which can be computer devices. In real-world scenarios, more or fewer types of terminal devices may participate in the evaluation process of the large-scale logistics model. The specific number and types depend on the actual scenario and are not limited here. Additionally, Figure 1 The example shows one server, but in real-world scenarios, multiple servers can be involved, especially in multidisciplinary output scenarios. The specific number of servers depends on the actual scenario.
[0038] In this embodiment, the server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, etc., but is not limited to these. The terminal and server can be directly or indirectly connected via wired or wireless communication, and the terminal and server can be connected to form a blockchain network; this application does not impose any restrictions.
[0039] It is understandable that the aforementioned logistics big data model evaluation system can run on personal mobile terminals, such as logistics management applications, or on servers, or on third-party devices to provide evaluations of the logistics big data model and obtain the evaluation results of the logistics big data model from the information source. Specifically, the logistics big data model evaluation system can run as a program on the aforementioned devices, or as a system component within the aforementioned devices, or as a cloud service program. The specific operating mode depends on the actual scenario and is not limited here.
[0040] The data in the logistics field is complex and diverse, involving multiple links such as transportation, warehousing, and distribution. Large-scale vertical models in the logistics field need to learn knowledge from each link. However, whether they can be applied in the end and help solve business problems requires targeted evaluation. Therefore, how to design a special evaluation method to ensure the accuracy and reliability of the model when handling these specific tasks has become a challenge.
[0041] To address the aforementioned issues, this application proposes an evaluation method for a large-scale logistics model, which is applied to... Figure 2 In the evaluation process framework of the large logistics model shown, such as Figure 2 The diagram shown is a flowchart of the evaluation process for a large logistics model provided in this application embodiment. The terminal sends a test set to the server through a test request, which enables the large logistics model configured in the server to respond. The model performance is then evaluated based on the semantic entropy in the response information.
[0042] It is understood that the logistics large-scale model evaluation method provided in this application can be a program written as processing logic in a hardware system, or it can be a logistics large-scale model evaluation device, implementing the above processing logic in an integrated or external manner. As one implementation, the logistics large-scale model evaluation device acquires a test set configured for each link in the logistics field; then inputs the test questions from the test set into the logistics large-scale model for multiple rounds of answering to obtain response information; and obtains probability parameters by comparing the similarity of words in the response information with the standard answers corresponding to the test questions, combining the probability parameters to obtain semantic entropy information; furthermore, it determines the entropy change parameter based on the semantic entropy difference corresponding to each test question in the semantic entropy information, and evaluates the performance of the logistics large-scale model based on the entropy change parameter. This realizes a semantic entropy-based logistics large-scale model evaluation process. Because the logistics large-scale model is used for multiple rounds of question-and-answering for logistics links, and the stability of the logistics large-scale model's answers to the same question is evaluated by combining the semantic entropy in the response information, the accuracy of the model evaluation is improved.
[0043] Based on the above process architecture, the evaluation method for the large-scale logistics model in this application will be introduced below. Please refer to [link / reference needed]. Figure 3 , Figure 3 A flowchart illustrating an evaluation method for a large-scale logistics model provided in this application embodiment, which includes at least the following steps:
[0044] 301. Obtain test sets for configurations of various aspects of the logistics field.
[0045] In this embodiment, the various links in the logistics field include multiple links such as transportation, warehousing, and distribution, and there are connections between these links. Therefore, the corresponding logistics big model has certain correlations between the questions and answers of each link. This embodiment uses test sets configured for each link in the logistics field to realize a fine-grained evaluation process for the logistics big model.
[0046] Specifically, the logistics big model is a large-scale model used for question answering. It is a vertical domain big model configured for the logistics field. The logistics big model is targeted at questions and answers in the logistics field. Compared with general big models, the logistics big model answers logistics questions more accurately and can be applied to various links in the logistics field for question answering.
[0047] In one possible scenario, the test set may include multiple-choice or short-answer questions, but other question types are also possible and not limited here. For example, based on the training data of the logistics big data model, the test set could be defined as follows:
[0048] A. Logistics and transportation related multiple-choice / short-answer questions: 100 correct answers / 100 correct answers.
[0049] B. Logistics product related multiple choice / short answer questions: 100 correct answers / 100 items.
[0050] The specific test set association process and question types depend on actual needs and are not limited here.
[0051] 302. Input the test questions from the test set into the logistics big data model to obtain the answer information through multiple rounds of solution.
[0052] In this embodiment, the question-and-answer process of the logistics big data model is carried out in multiple rounds. That is, each question is input into the logistics big data model separately, and each question needs to be answered multiple times. The same question is asked multiple times, and the big data model's answer each time is evaluated to determine the stability of the answer content.
[0053] Specifically, because the test questions in the test set have different question types, the composition of the answer information also differs for different question types. For multiple-choice questions, the order of the options input into the large model varies depending on the frequency of the answer. For short-answer questions, each question needs to be input into the large model using a mixture of different knowledge bases. The accuracy of the mixed answers for each multiple-choice question needs to be evaluated.
[0054] In one possible scenario, when the test question is a multiple-choice question, the corresponding options can be determined; then, the order of the options can be adjusted to update the test question; and the updated test question can be input into a large-scale logistics model for multiple rounds of answering to obtain the response information, with the order of the options differing in each round. For example, in... Figure 4 In the scene shown, Figure 4 This is a schematic diagram illustrating a scenario for an evaluation method of a large-scale logistics model, provided as one embodiment of this specification. For a target question, the order of its options is shuffled, and then input into the large-scale logistics model to obtain the answer information.
[0055] In one possible scenario, when the test question is an open-ended question, the logistics knowledge base associated with the test question can be obtained first; then, the logistics knowledge associated with the test question can be extracted from the logistics knowledge base; and the test question can be configured according to the logistics knowledge; subsequently, the configured test question is input into a large logistics model for multiple rounds of answering to obtain the answer information, with different logistics knowledge corresponding to each round. For example, in... Figure 5 In the scene shown, Figure 5 This diagram illustrates a scenario of an evaluation method for another large-scale logistics model provided as one embodiment of this specification. The diagram shows how, for a target question, different dimensions of answers are obtained by combining different knowledge bases to arrive at the corresponding answer information.
[0056] Understandably, a knowledge base can be the contextual knowledge needed to search for answers to short-answer questions, specifically logistics-related laws and regulations, or other logistics standards. Therefore, a targeted knowledge base can be determined by first acquiring key information from the test questions, such as the region and process; then, performing a logistics domain search based on this key information to obtain domain tags; and finally, identifying the logistics knowledge base associated with these domain tags, such as transportation regulations for region A, thereby improving the accuracy of the question-and-answer process.
[0057] 303. Obtain probability parameters by comparing the words in the answer information with the standard answers to the test questions, and combine the probability parameters to obtain semantic entropy information.
[0058] In this embodiment, a semantic entropy-based method is used to measure the model's capabilities. That is, if the large model can correctly understand the question, it can correctly grasp the knowledge regardless of how confusing the options or how mixed the information is. Therefore, we can evaluate the accuracy of the model's answers to objectively assess its ability to learn knowledge. Because the format of each large model's answer may be different, traditional text comparison methods, such as BLUE, cannot accurately reflect the accuracy of the answers. Therefore, this embodiment introduces "semantic entropy" to measure the degree of matching between each large model's answer and the standard answer.
[0059] Specifically, in this embodiment, semantic entropy is calculated by statistically determining the similarity probability of each character. A higher probability results in lower semantic entropy, indicating a smaller model entropy. This means the model's inferred answer is more similar to the standard answer, and vice versa. Furthermore, considering that the length of different answers may affect semantic entropy—longer answers may have higher semantic entropy—targeted adjustments are needed. First, probability parameters are obtained by comparing the characters in the answer information with the standard answer corresponding to the test question. Then, the length information of the standard answer is obtained. Finally, the probability parameters are weighted based on the length information to obtain the semantic entropy information. Therefore, based on the above description, the semantic entropy is designed as follows:
[0060]
[0061] Where wi is the i-th character in the standard answer, d is the length of the standard answer, and P(wi) is the probability of each character inferred by the large model. From the formula, we can deduce that the larger the probability of P(wi), the smaller H, indicating a smaller model entropy. This means the model's inferred answer will be more similar to the standard answer, and vice versa.
[0062] It is understandable that, since the evaluation method in this embodiment evaluates a general vertical domain model of logistics, the entropy generated by the logistics vocabulary of the vertical domain model will be more accurate than that of a general open-source model. Therefore, a logistics vocabulary list configured for a large logistics model can be obtained first; then, logistics feature words associated with the logistics vocabulary list in the answer information can be identified; and then, probability parameters can be obtained by comparing the logistics feature words in the answer information with the standard answers corresponding to the test questions. This improves the accuracy of the semantic entropy of answers in the logistics domain.
[0063] Furthermore, regarding the semantic entropy configuration for multiple-choice questions, since the large model outputs the relevant thought process when providing answers, this thought process can be used to calculate entropy, which can characterize the capabilities of the large model. Specifically, the process involves first determining the answer options corresponding to the given information; then acquiring the process information of the logistics large model in determining the answer options; and finally, comparing the words in the process information with the standard answers corresponding to the test questions to obtain probability parameters. This process enables the determination of semantic entropy for different question types.
[0064] 304. Determine the entropy variation parameter based on the semantic entropy difference of each test item in the semantic entropy information, and evaluate the performance of the logistics big model based on the entropy variation parameter.
[0065] In this embodiment, the semantic entropy difference is used to evaluate the stability of the logistics model for answering the same question. Specifically, it can first determine the extreme semantic entropy corresponding to each test question in the semantic entropy information, such as the maximum and minimum values; then determine the semantic entropy difference based on the extreme semantic entropy; furthermore, calculate the average of the semantic entropy differences corresponding to each test question to obtain the entropy change parameter; and then evaluate the performance of the logistics model based on the entropy change parameter. That is, the result of each entropy change is statistically analyzed, and the entropy change of each question, i.e., the semantic entropy difference, is:
[0066] ΔH(problem) = Hmax - Hmin
[0067] The entropy change result of the model is the mean entropy change of all questions, that is, the entropy change parameter is:
[0068] ΔH(model) = (ΔH(problem 0) + ΔH(problem 1) + ... + ΔH(problem N)) / N.
[0069] The performance evaluation of the large logistics model can be based on threshold comparison. For example, if ΔH(model) > 0.01, it indicates that the model entropy increases significantly and the model's understanding of knowledge is inaccurate; if ΔH(model) <= 0.01, it indicates that the model entropy increases less and the model's understanding of knowledge is more accurate.
[0070] Furthermore, considering the varying difficulty and complexity of different question types, to accurately reflect model performance, we can determine the question type information corresponding to the entropy change parameters; then, obtain the entropy change threshold corresponding to the question type information; and finally, evaluate the performance of the large-scale logistics model based on the entropy change threshold. For example, the entropy change threshold for short-answer questions is greater than that for multiple-choice questions, thereby improving the accuracy of model evaluation.
[0071] In summary, this embodiment acquires test sets configured for various stages of the logistics field; then inputs the test questions from the test sets into a large-scale logistics model for multiple rounds of question-and-answer processing to obtain response information; and obtains probability parameters by comparing the similarity of words in the response information with the standard answers corresponding to the test questions, and combines the probability parameters to obtain semantic entropy information; furthermore, it determines entropy change parameters based on the semantic entropy differences corresponding to each test question in the semantic entropy information, and evaluates the performance of the large-scale logistics model based on the entropy change parameters. This achieves a semantic entropy-based evaluation process for the large-scale logistics model. Because the large-scale logistics model is used for multiple rounds of question-and-answer processing for each logistics stage, and the stability of the large-scale logistics model's responses to the same question is evaluated by combining the semantic entropy in the response information, the accuracy of the model evaluation is improved.
[0072] It should be noted that the various embodiments described in this specification emphasize the parts that differ from other embodiments, and the embodiments can be explained by comparison with each other. Any combination of the various embodiments described in this specification based on general technical knowledge is covered within the scope of this specification.
[0073] In one exemplary embodiment of this specification, an evaluation device 600 for a large-scale logistics model is also provided, such as... Figure 6 As shown, Figure 6 A functional module diagram of an evaluation device for a large-scale logistics model provided in one embodiment of this specification, the interactive device 600 including:
[0074] Acquisition unit 601 is used to acquire test sets configured for various aspects of the logistics field;
[0075] Evaluation unit 602 is used to input the test questions in the test set into the logistics big model to obtain answer information through multiple rounds of answering;
[0076] The evaluation unit 602 is further configured to obtain probability parameters by comparing the words in the answer information with the standard answer corresponding to the test question, and to obtain semantic entropy information by combining the probability parameters.
[0077] The evaluation unit 602 is further configured to determine the entropy change parameter based on the semantic entropy difference corresponding to each test item in the semantic entropy information, so as to evaluate the performance of the logistics big model based on the entropy change parameter.
[0078] Optionally, in one possible implementation, the evaluation unit 602 is configured to: determine the selection options corresponding to the test questions when the test questions in the test set are input into the logistics big model for multiple rounds of answering to obtain answer information; adjust the order of the selection options to update the test questions; input the updated test questions into the logistics big model for multiple rounds of answering to obtain answer information, wherein the order of the selection options is different for each round.
[0079] Optionally, in one possible implementation, the evaluation unit 602 is configured to: obtain a logistics knowledge base associated with the test questions when the test questions in the test set are input into the logistics big model for multiple rounds of answering to obtain answer information; extract logistics knowledge associated with the test questions through the logistics knowledge base; configure the test questions according to the logistics knowledge; input the configured test questions into the logistics big model for multiple rounds of answering to obtain answer information, wherein the logistics knowledge corresponding to each round is different.
[0080] Optionally, in one possible implementation, the evaluation unit 602 is configured to, when obtaining the logistics knowledge base associated with the test question, acquire key information in the test question; perform a logistics domain retrieval based on the key information to obtain domain tags; and determine the logistics knowledge base associated with the domain tags.
[0081] Optionally, in one possible implementation, the evaluation unit 602 is configured to, when obtaining probability parameters by comparing the characters in the answer information with the standard answer corresponding to the test question to obtain semantic entropy information, perform a similarity comparison of the characters in the answer information with the standard answer corresponding to the test question to obtain probability parameters; obtain length information corresponding to the standard answer; and weight the probability parameters based on the length information to obtain the semantic entropy information.
[0082] Optionally, in one possible implementation, the evaluation unit 602 is configured to, when obtaining probability parameters by comparing the words in the answer information with the standard answer corresponding to the test question, acquire a logistics vocabulary configured for the logistics big model; determine the logistics feature words in the answer information that are associated with the logistics vocabulary; and obtain probability parameters by comparing the logistics feature words in the answer information with the standard answer corresponding to the test question.
[0083] Optionally, in one possible implementation, the evaluation unit 602 is configured to determine the answer option corresponding to the answer information when obtaining probability parameters by comparing the words in the answer information with the standard answer corresponding to the test question; obtain process information corresponding to the logistics big data model in determining the answer option; and obtain probability parameters by comparing the words in the process information with the standard answer corresponding to the test question.
[0084] Optionally, in one possible implementation, the evaluation unit 602 is configured to, when determining the entropy change parameter based on the semantic entropy difference corresponding to each of the test items in the semantic entropy information, and to perform performance evaluation on the logistics big model based on the entropy change parameter, determine the extreme semantic entropy corresponding to each of the test items in the semantic entropy information; determine the semantic entropy difference based on the extreme semantic entropy; calculate the mean of the semantic entropy difference corresponding to each of the test items to obtain the entropy change parameter; and perform performance evaluation on the logistics big model based on the entropy change parameter.
[0085] Optionally, in one possible implementation, the evaluation unit 602 is used to determine the question type information corresponding to the entropy change parameter when performing performance evaluation on the logistics big model based on the entropy change parameter; obtain the entropy change threshold corresponding to the question type information; and perform performance evaluation on the logistics big model based on the entropy change threshold.
[0086] Specifically, the acquisition unit and evaluation unit in this embodiment can correspond to physical components. For example, the processing unit can be a processing module such as a CPU, GPU, or FPGA. The specific physical component can be any component or combination of components with the above functions. The specific method depends on the actual scenario and is not limited here.
[0087] The aforementioned evaluation device acquires test sets configured for various stages of the logistics field; then, it inputs the test questions from the test sets into a large-scale logistics model for multiple rounds of question-and-answer processing to obtain response information; it then compares the words in the response information with the standard answers corresponding to the test questions to obtain probability parameters, and combines these probability parameters to obtain semantic entropy information; finally, it determines entropy change parameters based on the semantic entropy differences corresponding to each test question in the semantic entropy information, and evaluates the performance of the large-scale logistics model based on these entropy change parameters. This achieves a semantic entropy-based evaluation process for the large-scale logistics model. By using the large-scale logistics model for multiple rounds of question-and-answer processing for each logistics stage, and combining the semantic entropy in the response information to evaluate the stability of the large-scale logistics model's responses to the same question, the accuracy of the model evaluation is improved.
[0088] Specific limitations regarding the evaluation device for the large-scale logistics model can be found in the limitations of the evaluation method for the large-scale logistics model mentioned above, and will not be repeated here. Each unit module in the aforementioned evaluation device for the large-scale logistics model can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0089] Another embodiment of this application also proposes a computing device, see [link to relevant documentation] Figure 7 As shown, an exemplary embodiment of this specification also provides a computing device, including: a memory and a processor, the memory storing a computer program, the processor executing the computer program to perform the steps in the evaluation method of the large logistics model according to various embodiments of this specification described above.
[0090] The internal structure of the computing device can be as follows: Figure 7 As shown, the computing device includes a processor, memory, network interface, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it follows the steps of the large-scale logistics model evaluation method according to various embodiments of this specification as described in the above embodiments.
[0091] The processor may include the main processor, as well as baseband chips, modems, etc.
[0092] The memory stores a program that executes the technical solution of this invention, and may also store an operating system and other critical business functions. Specifically, the program may include program code, which includes computer operation instructions. More specifically, the memory may include read-only memory (ROM), other types of static storage devices capable of storing static information and instructions, random access memory (RAM), other types of dynamic storage devices capable of storing information and instructions, disk storage, flash memory, etc.
[0093] The processor can be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits used to control the execution of the program of the present invention. It can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0094] Input devices may include devices that receive data and information input by the user, such as keyboards, mice, cameras, scanners, light pens, voice input devices, touch screens, pedometers, or gravity sensors.
[0095] Output devices may include devices that allow information to be output to the user, such as displays, printers, speakers, etc.
[0096] The communication interface may include any transceiver-like device for communicating with other devices or communication networks, such as Ethernet, Radio Access Network (RAN), Wireless Local Area Network (WLAN), etc.
[0097] The processor executes the program stored in the memory and calls other devices, which can be used to implement the various steps of any of the logistics large-scale model evaluation methods provided in the above embodiments of this application.
[0098] The computing device may also include a display component and a voice component. The display component may be a liquid crystal display screen or an e-ink display screen. The input device of the computing device may be a touch layer covering the display component, or a button, trackball or touchpad set on the casing of the computing device, or an external keyboard, touchpad or mouse, etc.
[0099] Those skilled in the art will understand that Figure 7 The structures shown are merely block diagrams of some structures related to the solutions in this specification and do not constitute a limitation on the computing devices on which the solutions in this specification are applied. Specific computing devices may include more or fewer components than those shown in the figures, or combine certain components, or have different component arrangements.
[0100] In addition to the methods and devices described above, the logistics large-scale model evaluation method provided in the embodiments of this specification can also be a computer program product, which includes a computer program that, when run by a processor, causes the processor to perform the steps in the logistics large-scale model evaluation method according to various embodiments of this specification as described in the "Exemplary Methods" section above.
[0101] The computer program product described herein can be written in any combination of one or more programming languages to perform the operations of the embodiments described herein. These programming languages include object-oriented programming languages such as Java and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0102] Furthermore, embodiments of this specification also provide a computer-readable storage medium having a computer program stored thereon, the computer program being executed by a processor of the steps in the evaluation method of the large logistics model according to various embodiments of this specification as described in the "Exemplary Methods" section above.
[0103] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, storage, databases, or other media used in the embodiments provided in this specification can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.
[0104] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0105] The embodiments described above are merely illustrative of several implementation methods outlined in this specification. While the descriptions are specific and detailed, they should not be construed as limiting the scope of the solutions provided in this specification. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this specification, and these all fall within the scope of protection of this specification. Therefore, the scope of protection for this patent should be determined by the appended claims.
Claims
1. A method for evaluating a large-scale logistics model, characterized in that, include: Obtain test sets configured for each stage of the logistics process; The test questions in the test set are input into the logistics big model to obtain answer information through multiple rounds of answering; The probability parameters are obtained by comparing the words in the answer information with the standard answers corresponding to the test questions, and the semantic entropy information is obtained by combining the probability parameters. The entropy change parameter is determined based on the semantic entropy difference corresponding to each test item in the semantic entropy information, and the performance of the logistics big model is evaluated based on the entropy change parameter.
2. The method according to claim 1, characterized in that, The process of inputting the test questions from the test set into the logistics big data model for multiple rounds of answering to obtain response information includes: Determine the options corresponding to the test questions; The order of the options is adjusted to update the test questions; The updated test questions are input into the logistics model to obtain the answer information through multiple rounds of answering, with a different order of options for each round.
3. The method according to claim 1, characterized in that, The process of inputting the test questions from the test set into the logistics big data model for multiple rounds of answering to obtain response information includes: Obtain the logistics knowledge base associated with the test questions; The logistics knowledge associated with the test questions is extracted from the logistics knowledge base. Configure the test questions based on the aforementioned logistics knowledge; The configured test questions are input into the logistics model to obtain the answer information through multiple rounds of questioning, with different logistics knowledge corresponding to each round.
4. The method according to claim 3, characterized in that, The step of obtaining the logistics knowledge base associated with the test questions includes: Obtain key information from the test questions; Based on the aforementioned key information, a search is performed in the logistics field to obtain field tags; Determine the logistics knowledge base associated with the domain label.
5. The method according to claim 1, characterized in that, The step of obtaining probability parameters by comparing the words in the answer information with the standard answers corresponding to the test questions, and then combining the probability parameters to obtain semantic entropy information, includes: The probability parameter is obtained by comparing the words in the answer information with the standard answer corresponding to the test question; Obtain the length information corresponding to the standard answer; The probability parameters are weighted based on the length information to obtain the semantic entropy information.
6. The method according to claim 5, characterized in that, The step of obtaining probability parameters by comparing the words in the answer information with the standard answers corresponding to the test questions includes: Obtain a logistics vocabulary list for the configuration of the large logistics model; Identify the logistics feature words in the answer information that are associated with the logistics vocabulary; The probability parameter is obtained by comparing the logistics feature words in the answer information with the standard answer corresponding to the test question.
7. The method according to claim 5, characterized in that, The step of obtaining probability parameters by comparing the words in the answer information with the standard answers corresponding to the test questions includes: Determine the answer options corresponding to the given answer information; Obtain process information corresponding to the determination of the answer options in the logistics big data model; The probability parameter is obtained by comparing the characters in the process information with the standard answers corresponding to the test questions.
8. The method according to claim 1, characterized in that, The step of determining the entropy change parameter based on the semantic entropy difference corresponding to each test item in the semantic entropy information, and then evaluating the performance of the logistics big model based on the entropy change parameter, includes: Determine the extreme semantic entropy corresponding to each test question in the semantic entropy information; The semantic entropy difference is determined based on the extreme value semantic entropy; The semantic entropy difference corresponding to each of the test questions is averaged to obtain the entropy change parameter; The performance of the large-scale logistics model is evaluated based on the entropy change parameters.
9. The method according to claim 8, characterized in that, The performance evaluation of the large-scale logistics model based on the entropy change parameter includes: Determine the question type information corresponding to the entropy change parameter; Obtain the entropy change threshold corresponding to the question type information; The performance of the large-scale logistics model is evaluated based on the entropy change threshold.
10. An evaluation device for a large-scale logistics model, characterized in that, include: The acquisition unit is used to acquire test sets configured for various aspects of the logistics field; An evaluation unit is used to input the test questions from the test set into the logistics big model to obtain answer information through multiple rounds of answering; The evaluation unit is also used to obtain probability parameters by comparing the words in the answer information with the standard answer corresponding to the test question, and to obtain semantic entropy information by combining the probability parameters. The evaluation unit is further configured to determine the entropy change parameter based on the semantic entropy difference corresponding to each test item in the semantic entropy information, so as to evaluate the performance of the logistics big model based on the entropy change parameter.
11. A computing device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the evaluation method of the large logistics model according to any one of claims 1 to 9.
12. A computer program product, characterized in that, include: A computer program, when executed by a processor, implements the evaluation method for the large-scale logistics model according to any one of claims 1 to 9.