E-commerce logistics comment classification method and device based on large language model, and medium

By setting up multi-label data sets and quantitative configuration large language models on the e-commerce platform, the adaptability and interpretability problems of logistics comment classification are solved, and the automation and intelligent classification of e-commerce logistics comments are realized, and the classification accuracy and interpretability are improved.

CN120296172APending Publication Date: 2025-07-11ZHEJIANG GONGSHANG UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510460969.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The prior art is difficult to extract key information from logistics comments on e-commerce platforms efficiently and accurately and classify them, and the large language model is insufficient in adaptability and interpretability in logistics scenarios.

Method used

By setting up a multi-label data set, obtaining and marking e-commerce logistics review data, quantifying and configuring large language models, combining e-commerce logistics review data for training, optimizing large language models to adapt to logistics scenarios, and building an interpretable classification system.

Benefits of technology

It realizes the automated and intelligent classification of e-commerce logistics comments, improves the adaptability and classification accuracy of the model, and generates interpretable classification results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296172A_ABST
    Figure CN120296172A_ABST
Patent Text Reader

Abstract

The invention discloses an e-commerce logistics comment classification method and device based on a large language model, and a medium, and the method comprises the steps: setting a multi-label data set, and obtaining e-commerce logistics comment data; labeling each piece of e-commerce logistics comment data according to the multi-label data set; an initial large language model is obtained and quantized, and the quantized large language model is configured; training a big language model based on the e-commerce logistics comment data and the corresponding labels; optimizing the trained large language model based on e-commerce logistics comment feedback data; and inputting to-be-processed e-commerce logistics comment data into the optimized large language model to obtain a prediction probability value of each piece of e-commerce logistics comment data on each label, thereby obtaining a classification result of each piece of e-commerce logistics comment data. According to the method, automation, intelligentization and high efficiency of logistics comment classification are realized, and the problem of insufficient classification accuracy and scene adaptability in a traditional method is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of large language models, and particularly relates to an e-commerce logistics review classification method, device, and medium based on a large language model. Background Art

[0002] With the rapid development of e-commerce, the importance of logistics services in the supply chain has become increasingly prominent. Consumers frequently post logistics-related reviews on e-commerce platforms, and these reviews contain a large amount of information about the logistics service experience, such as delivery time, delivery location, and service quality. However, due to the large volume and complex structure of review data, traditional text analysis methods are difficult to efficiently and accurately extract key information and complete classification.

[0003] In the prior art, the classification of logistics reviews usually follows the following process: data collection and cleaning, feature extraction and representation, model training and prediction, and result application. By designing a dedicated deep neural network structure, multi-layer feature extraction is performed on the review text to improve the classification accuracy. However, such methods have problems of high dependence on data and poor generalization ability.

[0004] In recent years, with the development of deep learning technology, the application of large language models (LLMs) in the field of text processing has become a research hotspot. Large language models have excellent performance in a variety of natural language processing tasks with their powerful pre-training ability and rich context understanding ability. However, applying large language models (LLMs) to logistics review classification still faces the following difficulties:

[0005] 1. It is necessary to fine-tune the model for the logistics scenario to improve the adaptability;

[0006] 2. How to construct an interpretable e-commerce logistics review classification system in combination with the actual needs of the logistics industry. Summary of the Invention

[0007] Aiming at the deficiencies of the prior art, the present invention provides an e-commerce logistics review classification method, device, and medium based on a large language model.

[0008] In a first aspect, an embodiment of the present invention provides an e-commerce logistics review classification method based on a large language model, and the method includes the following steps:

[0009] Set up a multi-label data set, and obtain e-commerce logistics review data; label each e-commerce logistics review data according to the multi-label data set;

[0010] Obtain an initial large language model and quantize it, and configure the quantized large language model;

[0011] Train the large language model based on the e-commerce logistics review data and the corresponding multi-labels;

[0012] Input the e-commerce logistics review data to be processed into the trained large language model to obtain the predicted probability values of each e-commerce logistics review data on each label, so as to obtain the classification results of each e-commerce logistics review data.

[0013] In a second aspect, an embodiment of the present invention provides an electronic device, including:

[0014] At least one processor; and

[0015] A memory communicatively connected to the at least one processor; wherein,

[0016] The memory stores one or more computer programs executable by the at least one processor, and the one or more computer programs are executed by the at least one processor so that the at least one processor can execute the e-commerce logistics review classification method based on the large language model as described above.

[0017] In a third aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored, and the computer program realizes the e-commerce logistics review classification method based on the large language model as described above when executed by a processor.

[0018] In a fourth aspect, an embodiment of the present invention provides a computer program product, including a computer program / instructions, and the computer program / instructions realize the e-commerce logistics review classification method based on the large language model as described above when executed by a processor.

[0019] Compared with the prior art, the beneficial effects of the present invention are:

[0020] The present invention provides an e-commerce logistics review classification method based on a large language model, fine-tunes the large language model in the e-commerce logistics scenario to improve the adaptability, and the present invention combines the e-commerce logistics review data and the corresponding multi-labels to train the large language model. The multi-labels are combined with the actual e-commerce logistics industry scenario, so as to quantify the e-commerce logistics review classification and construct an interpretable e-commerce logistics review classification system. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0022] Figure 1 It is a schematic flowchart of the logistics review classification method based on the large language model provided by the embodiment of the present invention;

[0023] Figure 2 Schematic diagram of an electronic device provided by an embodiment of the present invention. Specific embodiments

[0024] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0025] It should be noted that, without conflict, the features in the following embodiments and implementation manners can be combined with each other.

[0026] As Figure 1 shown, an embodiment of the present invention provides a logistics comment classification method based on a large language model. The method includes the following steps:

[0027] Step S1, set a multi-label data set, and obtain e-commerce logistics comment data; mark labels for each e-commerce logistics comment data according to the multi-label data set.

[0028] Specifically, set a multi-label data set, denoted as Y = {y1, y2,..., y m}, where m represents the number of labels; exemplarily, the labels include: distribution logistics cost, inspection of goods, professional personnel quality, distribution range, damage of goods, loss of goods, information technology level, information transmission level, accuracy of distribution location, accuracy of distribution time, delivery flexibility, and logistics irrelevance.

[0029] The e-commerce logistics comment data is denoted as D = {d1, d2,..., d i ,..., d n}, where d i is the i-th e-commerce logistics comment data, i ∈ [1, n]; the i-th multi-label data corresponding to the i-th e-commerce logistics comment data d i is denoted as Y i = {y i1 , y i2 , ……, y ik}, where k is the number of labels corresponding to the i-th e-commerce logistics comment data d i , and k ≤ m.

[0030] Furthermore, in this example, the logistics comment data can be collected through various channels, and the logistics comment data mainly comes from e-commerce platforms (such as JD.com, Taobao, Pinduoduo, etc.).

[0031] Further, the step S1 further includes: cleaning, deduplicating, annotating, and standardizing the obtained logistics review data, denoted as X = {x1, x2, ……, x n}, so as to ensure the consistency and adaptability of the logistics review data.

[0032] Further, the process of standardizing the logistics review data includes: converting each comment into a standardized format through the formatting function formatting_prompts_func:

[0033] text{"Instruction:{(prompt)},Input:{d i},Response:{Y i}"}

[0034] where Instruction is the classification instruction, Input is the input review text, and Response is the set of target classification labels of the review.

[0035] Step S2, obtain the initial large language model and quantize it, and configure the quantized large language model.

[0036] Further, set the maximum input sequence length L of the large language model max ≤ 4096, set the data type dtype, set the quantization load_in_4bit, and quantize the large language model. The expression is as follows:

[0037] M quant =f quant (M pre ,q)

[0038] In the formula, q = 4bit, indicating the use of 4bit quantization technology to reduce memory occupancy, and f quant (.) represents the quantization operation.

[0039] Further, set the rank value, scaling factor, random dropout rate, query projection module q_proj, key projection module k_projk, and value projection module v_proj to configure the large language model. The expression is as follows:

[0040] M = f(M quant ,r,α,p,t)

[0041] In the formula, r is the rank value, α is the scaling factor, p is the random dropout rate, t is the query projection module q_proj, key projection module k_projk, and value projection module v_proj, and f(.) represents the configuration function.

[0042] Step S3: Train a large language model based on e-commerce logistics review data and corresponding multi-labels.

[0043] Further, set a loss function and train the large language model based on e-commerce logistics review data and corresponding labels, and use an optimizer (such as AdamW 8bit) to minimize the loss function; the expression of the loss function is as follows:

[0044]

[0045] In the formula, l is the loss function, h(x i ; θ) represents the set of labels predicted by the large language model, P(y j ∣x j ; θ) represents the predicted probability of the large language model for label y j , y i is the actual label of sample x i , and m is the number of samples.

[0046] In this example, the loss function uses a multi-label cross-entropy loss function, and the expression is as follows:

[0047]

[0048] In the formula, y ik represents the true value of the i-th e-commerce logistics review data on label k; p ik represents the predicted value of the i-th e-commerce logistics review data output by the large language model on label k.

[0049] The expression of the label prediction probability p ik output by the large language model is as follows:

[0050]

[0051] In the formula, σ is the Sigmoid activation function; W k , b k are the weight and bias parameters of the large language model, and h i is the feature representation of the i-th e-commerce logistics review data.

[0052] Further, the expression of the trained large language model is as follows:

[0053] M final ={W final ,T final ,C final}

[0054] In the formula, W final is the weight file, T final is the tokenizer, and C final is the configuration file.

[0055] Furthermore, the training process of the large language model further includes:

[0056] Calculating the prediction error matrix; the expression is as follows:

[0057]

[0058] In the formula, y i represents the true value corresponding to the i-th e-commerce logistics review data, represents the predicted value corresponding to the i-th e-commerce logistics review data;

[0059] When the 2-norm of the prediction error matrix is greater than the threshold, optimize the trained large language model based on the e-commerce logistics review feedback data;

[0060] Optimizing the trained large language model based on the e-commerce logistics review feedback data, the expression is as follows:

[0061] M opt = f opt (M final , F)

[0062] In the formula, f opt (.) represents the feedback optimization function, M final represents the trained large language model, F = {f1, f2,..., f t} represents the e-commerce logistics review feedback data, and t represents the total number of e-commerce logistics review feedback data;

[0063] Input the e-commerce logistics review data to be processed into the optimized large language model to obtain the predicted probability value of each e-commerce logistics review data on each label, so as to obtain the classification result of each e-commerce logistics review data.

[0064] It should be noted that the fine-tuned large language model is deployed to the actual application scenario of e-commerce logistics review classification for real-time logistics review classification. Generate service optimization suggestions according to the classification results to form an optimization target set

[0065] O = {o1, o2, ……, o n}, such as improving delivery time management, delivery location planning management, improving the attitude of professional personnel, etc.

[0066] Furthermore, the method further includes:

[0067] The classification result calculates the accuracy rate by comparing the matching degree between the actual label and the predicted label set, and the expression of the accuracy rate is as follows:

[0068]

[0069] Among them, h(x i ; θ) is the set of predicted labels of the large language model for x i , and y i is the set of actual labels.

[0070] After statistics, the total amount of experimental review data is 19,261, including 2,120 multi-label data and 17,141 single-label data. The data format is shown in Table 1. Table 2 shows the data scales of each label review in the training set, validation set, and test set. The training set contains 15,409 data, accounting for about 80% of the total data. A larger training set can provide more samples, which helps the model better learn the data features and improve its generalization ability and prediction accuracy. The data volumes of the validation set and the test set are 1,927 and 1,925 respectively, each accounting for 10% of the total data, and are mainly used for the final performance evaluation.

[0071] Table 1 Data Format

[0072]

[0073] Table 2 Scales of the Training Set, Validation Set, and Test Set

[0074]

[0075] Table 3 Scales of the Training Set, Validation Set, and Test Set

[0076]

[0077]

[0078] The experimental results show that the Llama3 model demonstrates excellent performance, with balanced and stable performance in various indicators, reflecting the powerful ability of generative large models in handling complex tasks.

[0079] To further explore the generalization ability of large models, comparative experiments were designed with data of different scales. Specifically, three sub-datasets of different scales were constructed on the basis of the original training dataset, named Group A, Group B, and Group C, with their data volumes increasing in sequence, being 192, 488, and 956 respectively. The experimental results show that when the data volume is small, the performance of the Llama3 model is significantly better than that of the other models. Especially in Group A with the smallest data volume, its classification performance advantage is particularly significant, and the F1-value is nearly about 19% higher than that of the second-ranked model, indicating that it still has good generalization ability in the case of small samples.

[0080] Table 4 F1-Value (%) Results under Data of Different Scales

[0081]

[0082] In summary, the present invention provides an e-commerce logistics review classification method based on a large language model, which fine-tunes the large language model in the e-commerce logistics scenario to improve adaptability. Moreover, the present invention combines e-commerce logistics review data and corresponding multi-labels to train the large language model. The multi-labels are combined with the actual e-commerce logistics industry scenario, realizing the automation, intelligence, and high efficiency of logistics review classification, quantifying the e-commerce logistics review classification, and constructing an interpretable e-commerce logistics review classification system.

[0083] Correspondingly, the present application also provides an electronic device, including: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the above-mentioned e-commerce logistics review classification method based on a large language model. As Figure 2 shown, it is a hardware structure diagram of any device with data processing capabilities where the e-commerce logistics review classification method based on a large language model provided by the embodiment of the present invention is located. In addition to Figure 2 the processors, memory, and network interfaces shown, any device with data processing capabilities where the device in the embodiment is located usually includes other hardware according to the actual functions of the device with data processing capabilities, which will not be elaborated here.

[0084] Correspondingly, the present application also provides a computer-readable storage medium, on which computer instructions are stored. When the instructions are executed by a processor, the above-mentioned e-commerce logistics review classification method based on a large language model is implemented. The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium may also be an external storage device, such as a plug-in hard disk, a Smart Media Card (SMC), an SD card, a Flash Card, etc. equipped on the device. Further, the computer-readable storage medium may also include both an internal storage unit of any device with data processing capabilities and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and may also be used to temporarily store data that has been output or will be output.

[0085] Those skilled in the art will readily think of other implementation schemes of the present application after considering the specification and practicing the content disclosed herein. The present application aims to cover any variations, uses, or adaptive changes of the present application, which follow the general principles of the present application and include common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and embodiments are only regarded as exemplary.

[0086] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope.

Claims

1. A method for classifying e-commerce logistics reviews based on large language models, characterized in that The method includes the following steps: Set up a multi-label dataset and obtain e-commerce logistics review data; label each piece of e-commerce logistics review data according to the multi-label dataset; Obtain an initial large language model and quantize it, and configure the quantized large language model; Train the large language model based on the e-commerce logistics review data and the corresponding multi-labels; Input the e-commerce logistics review data to be processed into the trained large language model to obtain the predicted probability value of each piece of e-commerce logistics review data on each label, so as to obtain the classification result of each piece of e-commerce logistics review data.

2. The e-commerce logistics review classification method based on a large language model according to claim 1, wherein The labels in the multi-label dataset include: distribution logistics cost, inspection, professional personnel quality, distribution range, goods damage, loss of parcels, information technology level, information transmission level, accuracy of distribution location, accuracy of distribution time, delivery flexibility, and logistics irrelevance.

3. A method for classifying e-commerce logistics reviews based on a large language model according to claim 1, characterized in that, The process of obtaining e-commerce logistics review data further includes: Standardize the e-commerce logistics review data, so as to convert the e-commerce logistics review data into a standardized text format, and the standardized text format is as follows: text{"Instruction:{(Prompt)},Input:{d i},Response:{Y i}"} Where Instruction is the classification instruction, Input is the input e-commerce logistics review data, and d i is the i-th e-commerce logistics review data, Response is the classification label corresponding to the e-commerce logistics review data, and Y i represents the i-th review text d i corresponding to the i-th multi-label data, and Y i ={y i1 , y i2 , ……, y ik}, k is the number of labels corresponding to the i-th e-commerce logistics review data d i .

4. A method for classifying e-commerce logistics reviews based on a large language model according to claim 1, characterized in that The process of obtaining an initial large language model and quantizing it, and configuring the quantized large language model includes: Set the maximum input sequence length of the large language model and quantize the large language model, and the expression is as follows: M quant = f quant (M pre , q) where q = 4bit, indicating the adoption of 4bit quantization technology to reduce memory occupancy, f quant (.) represents the quantization operation, which is used to quantize the large language model parameters M_pre into M_quant; Set the rank value, scaling factor, random dropout rate, query projection module, key projection module, and value projection module to configure the large language model, and the expression is as follows: M = f(M quant , r, α, p, t) In the formula, r is the rank value, α is the scaling factor, p is the random dropout rate, t is the query projection module, key projection module, or value projection module, and f(.) represents the configuration function, which is used to combine the quantized model with the adjustment parameters to generate the final base model M.

5. A method for classifying e-commerce logistics reviews based on a large language model according to claim 1, wherein, The process of training the large language model based on the e-commerce logistics review data and the corresponding multi-labels includes: Set the loss function, and train the large language model based on the e-commerce logistics review data and the corresponding labels until the loss function converges; the expression of the loss function is as follows: where \(l\) is the loss function, \(h(x i ;\theta)\) represents the set of labels predicted by the large language model, \(P(y j |x j ;\theta)\) represents the predicted probability of the label \(y j \) by the large language model, \(y i \) is the actual label of the sample \(x i \), and \(m\) is the number of samples.

6. A method for classifying e-commerce logistics reviews based on a large language model according to claim 5, characterized in that, The loss function adopts the multi-label cross-entropy loss function.

7. A method for classifying e-commerce logistics reviews based on a large language model according to claim 1 or 5, characterized in that, The process of training the large language model based on the e-commerce logistics review data and the corresponding multi-labels further includes: Calculate the prediction error matrix; When the 2-norm of the prediction error matrix is greater than the threshold, optimize the trained large language model based on the e-commerce logistics review feedback data; Optimize the trained large language model based on the e-commerce logistics review feedback data, and the expression is as follows: M opt = f opt (M final , F) where f opt (.) represents the feedback optimization function, M final represents the trained large language model, F = {f1, f2, …, f t} represents the e-commerce logistics comment feedback data, and t represents the total number of e-commerce logistics comment feedback data; Input the e-commerce logistics review data to be processed into the optimized large language model to obtain the predicted probability value of each piece of e-commerce logistics review data on each label, so as to obtain the classification result of each piece of e-commerce logistics review data.

8. An electronic device, characterized in that, Including: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores one or more computer programs executable by the at least one processor, and the one or more computer programs are executed by the at least one processor so that the at least one processor can execute the e-commerce logistics review classification method based on the large language model according to any one of claims 1-7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program, when executed by a processor, implements the method for classifying e-commerce logistics reviews based on a large language model according to any one of claims 1-7.

10. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by a processor, the method for classifying e-commerce logistics reviews based on a large language model according to any one of claims 1-7 is implemented.