Automatic auditing method, device and equipment for test lading bill and readable storage medium
By performing text feature extraction of test bills of lading and comprehensive use of machine learning classifiers, the subjectivity and efficiency problems of traditional manual audits are solved, and more accurate and efficient audit results are achieved.
Patent Information
- Application Number
- CN202510029563.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-23
AI Technical Summary
Traditional test bill of lading audits rely on manual labor, which has problems such as inconsistent subjectivity, time-consuming and resource-consuming and lack of automation tools.
By extracting the text content of the Bill of Lading to be audited, obtaining the text features using the BERT model, and inputting them into the first classifier and the second classifier, comprehensively performing auditing of the classification results.
It realizes the conversion of manual audit results into a unified and semantic quantitative feature space, improves audit efficiency and accuracy, and reduces human errors and results differences.
Smart Images

Figure CN120030408A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method, device, equipment and computer-readable storage medium for automated auditing of test bills of lading. Background Art
[0002] Traditional test bill of lading audits mainly rely on manual work, which has the following problems:
[0003] 1. Manual audits are easily affected by individual subjective opinions and experiences. Different auditors may come to different conclusions, resulting in inconsistent audit results.
[0004] 2. Manual audits usually require a lot of time and resources, so for large-scale projects or frequent testing, the scale of the audit may be limited;
[0005] 3. Traditional manual audits usually lack the support of automated tools and data analysis technology, and are unable to quickly analyze a large number of test bills of lading, resulting in insufficient audit efficiency. Summary of the invention
[0006] In order to solve the above technical problems, the present application provides a test bill of lading automated audit method, device, equipment and computer-readable storage medium.
[0007] In a first aspect, an embodiment of the present application provides a test bill of lading automated audit method, the test bill of lading automated audit method comprising:
[0008] Perform feature extraction on the text content of the audit test bill of lading to obtain text features;
[0009] Inputting the text features into the first classifier and the second classifier respectively;
[0010] The audit result of the test bill of lading to be audited is obtained by combining the first classification result output by the first classifier and the second classification result output by the second classifier.
[0011] In combination with the first aspect, in one implementation, the feature extraction of the text content of the test bill of lading to be audited to obtain text features includes:
[0012] Input the text content of the test bill of lading to be audited into the BERT model and select the vector output by the target layer in the BERT model;
[0013] Based on the vector output by the target layer, text features are obtained.
[0014] In combination with the first aspect, in one implementation, the target layer includes the 9th to 12th layers, and obtaining the text feature based on the vector output by the target layer includes:
[0015] The vectors output from the 9th to 12th layers are combined to obtain the initial features;
[0016] When the dimension of the initial feature is greater than the number of dimensions that can be supported by the number of training samples in the training set, the initial feature is subjected to dimensionality reduction processing to obtain text features, and the training set is used to train the first classifier and the second classifier.
[0017] In combination with the first aspect, in one implementation, the dimensionality reduction processing of the initial features adopts a principal component analysis method.
[0018] In combination with the first aspect, in one implementation, the first classification result includes a first probability value P of category i i , i∈[1,Q], Q is the number of categories; the second classification result includes the second probability value P of category i i ′, i∈[1,Q], Q is the number of categories; based on the first classification result output by the first classifier and the second classification result output by the second classifier, the audit result of the test bill of lading to be audited includes:
[0019] P corresponding to category i i and P i ′Weighted summation to obtain the overall probability value of category i;
[0020] The category corresponding to the maximum overall probability value is taken as the audit result of the test bill of lading to be audited.
[0021] In combination with the first aspect, in one implementation, the first classifier is a linear classifier, the second classifier is a nonlinear classifier, the first classifier and the second classifier are trained based on a training set, and when the number of training samples in the training set is less than or equal to a threshold, When the number of training samples in the training set is less than or equal to the threshold, Where w is P i The weight of P i ′’s weight.
[0022] In combination with the first aspect, in one implementation, before the text features are input into the first classifier and the second classifier respectively, the method further includes:
[0023] Constructing a training set, wherein the training set includes a plurality of training samples, each training sample includes a sample text feature and category labeling information corresponding to a sample test bill of lading;
[0024] Training a first classifier to be trained and a second classifier to be trained based on the training set;
[0025] When the training end condition is met, the first classifier and the second classifier are obtained.
[0026] In a second aspect, an embodiment of the present application provides a test bill of lading automated auditing device, the test bill of lading automated auditing device comprising:
[0027] A feature extraction module is used to extract features from the text content of the audit test bill of lading to obtain text features;
[0028] An input module, used for inputting the text features into the first classifier and the second classifier respectively;
[0029] The audit module is used to synthesize the first classification result output by the first classifier and the second classification result output by the second classifier to obtain the audit result of the test bill of lading to be audited.
[0030] In a third aspect, an embodiment of the present application provides a test order automated audit device, which includes a processor, a memory, and a test order automated audit program stored in the memory and executable by the processor, wherein when the test order automated audit program is executed by the processor, the steps of the test order automated audit method described in the first aspect are implemented.
[0031] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a test order automated audit program is stored, wherein when the test order automated audit program is executed by a processor, the steps of the test order automated audit method described in the first aspect are implemented.
[0032] The beneficial effects brought by the technical solution provided in the embodiments of the present application include:
[0033] In the embodiment of the present application, the text content of the test bill of lading to be audited is subjected to feature extraction to obtain text features; the text features are respectively input into the first classifier and the second classifier; the first classification result output by the first classifier and the second classification result output by the second classifier are combined to obtain the audit result of the test bill of lading to be audited. Through the embodiment of the present application, the text content of the test bill of lading to be audited is subjected to feature extraction to obtain text features based on natural language processing technology, and the differentiated bill of lading description text with inconsistent expressions generated manually is converted into a unified and semantically consistent quantitative feature space to form a numerical feature expression, which is used as the input of the first and second classifiers, and finally, the output of the first and second classifiers is combined to obtain the audit result. Among them, the method based on machine learning can quickly handle the audit work for a large number of test bills of lading, thereby improving the speed of the audit; the method based on machine learning can identify problems and anomalies, reducing the possibility of human errors; avoiding manual participation in the audit work, and reducing the differences in audit results caused by inconsistent audit conditions such as auditors and audit environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 A flowchart of an embodiment of an automated audit method for bills of lading for testing this application;
[0035] Figure 2 This is a schematic diagram of converting the text content of the test bill of lading to be audited in an embodiment of the test bill of lading automated audit method of the present application;
[0036] Figure 3 A schematic diagram of obtaining text features in an embodiment of the method for automated auditing of bills of lading for testing this application;
[0037] Figure 4 A schematic diagram of a scenario for testing an embodiment of an automated bill of lading audit method for this application;
[0038] Figure 5 This is a functional module diagram of an embodiment of an automated audit device for bills of lading for testing this application;
[0039] Figure 6 This is a schematic diagram of the hardware structure of the test bill of lading automated auditing equipment involved in the embodiment of the present application. DETAILED DESCRIPTION
[0040] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0041] In order to make the objectives, technical solutions and advantages of the present application clearer, the implementation methods of the present application will be further described in detail below with reference to the accompanying drawings.
[0042] In a first aspect, an embodiment of the present application provides a method for automated auditing of a test bill of lading.
[0043] In one embodiment, referring to Figure 1 , Figure 1 This is a flow chart of an embodiment of the method for testing the automated audit of bills of lading in this application. Figure 1 As shown, the automated audit method for testing bills of lading includes:
[0044] Step S10, extracting features from the text content of the audited test bill of lading to obtain text features;
[0045] In this embodiment, the text content is related to the quality of the test bill of lading. Specifically, data related to the quality of the test bill of lading can be mined from the perspective of auditing and organized into several elements as test bill of lading audit text attributes, taking the following six attributes as examples:
[0046] a. (Attribute 1) Feature name: Indicate the corresponding diagnostic feature. The name should be filled in based on the project diagnostic feature range specified in the project design requirements table.
[0047] b. (Attribute 2) Diagnostic conclusion: clearly define the location conclusion and point out the specific fault interface.
[0048] c. (Attribute 3) Diagnosis step: Fill in the final diagnosis and positioning results and screenshots of the last step.
[0049] d. (Attribute 4) Problem description: A concise and to-the-point summary of the error, using keywords instead of describing the location, specific conditions, and phenomenon where the bug occurred.
[0050] e. (Attribute 5) Test environment: Detailed description of the environment where the problem occurred. Highlight key elements of the environment, such as networking information, topology, single disk information, business configuration, etc.
[0051] f. (Attribute 6) Test steps: Describe in detail the steps and abnormal phenomena that caused the fault.
[0052] Feature extraction of text content is equivalent to feature encoding of text content, which is the process of converting text data into a form that computers can understand and process.
[0053] Furthermore, in one embodiment, step S10 includes:
[0054] Step S101, input the text content of the test bill of lading to be audited into the BERT model, and select the vector output by the target layer in the BERT model;
[0055] In this embodiment, the text content of the audit test bill of lading is feature extracted based on natural language processing technology to obtain text features. Specifically, the BERT model that has been pre-trained on a large corpus can be used. Specifically, Bert-base or Bert-large can be selected as the pre-training model. The main difference between the two is the size of the parameter scale. Considering the convenience of expression, this embodiment uses Bert-base as an example to describe the text feature encoding process. The Bert-base model is composed of a multi-layer bidirectional Transformer encoder, L (number of Transformer layers) = 12, H (hidden layer dimension) = 768, A (number of Attention heads) = 12. The text content containing the above six attributes is used as input for quantization encoding to form a numerical feature expression.
[0056] Furthermore, considering that the English version of the BERT model has more performance and resources, the Chinese text content description is first translated into English; then the English text content is converted into a format suitable for the English version of the BERT model input, such as [CLS] (indicating the beginning of the text) and [SEP] (indicating the end or separator of the text). Usually, the [CLS] tag is placed at the beginning of the text, and the [SEP] tag is used to separate different sentences or text segments; after the English text content is formatted, it is used as the input of the text encoder.
[0057] Reference Figure 2 , Figure 2 This is a schematic diagram of converting the text content of the test bill of lading to be audited in the first embodiment of the test bill of lading automated audit method of this application. Figure 2 As shown in the figure, the original text is first translated into English and then adjusted into a format suitable for input to the English version of the BERT model.
[0058] Step S102, obtaining text features based on the vector output by the target layer.
[0059] In this embodiment, refer to Figure 3 , Figure 3 This is a schematic diagram of obtaining text features in an embodiment of the method for automated auditing of bills of lading for testing this application. Figure 3As shown, the Input is input into the BERT model (taking Bert-base as an example). The BERT model has a total of 12 layers. L1-L5 are basic functional layers. Starting from L6 are repeated self-attention and feedforward neural network layers. That is, one layer from L6 to L12 can be selected as the extraction feature of the input text, or features of multiple layers can be selected and fused as the extraction feature of the input text.
[0060] Further, in one embodiment, the target layer includes the 9th to 12th layers, and step S102 includes:
[0061] Step S1021, synthesizing the vectors output from the 9th to the 12th layers to obtain initial features;
[0062] Step S1022, when the dimension of the initial feature is greater than the number of dimensions that can be supported by the number of training samples in the training set, the initial feature is subjected to dimensionality reduction processing to obtain text features, and the training set is used to train the first classifier and the second classifier.
[0063] In this embodiment, according to the empirical theory of pattern recognition, the number of training samples must be more than ten times the feature dimension to ensure the generalization performance of the classification model. Considering that the current number of training samples is about 2500, and the single-layer feature dimension of the hidden layer of the BERT model is 768, that is, only taking a single-layer feature dimension is already greater than the number of feature dimensions that the current number of samples can support, which is 2500*10%. However, taking only a single-layer feature is prone to loss of details, so it is necessary to select features. Continue to refer to Figure 3 In order to take into account both high semantics and detailed features, and to avoid introducing too much noise, the output vectors of the four layers L9, L10, L11, and L12 are expressed as F9, F10, F11, and F12 respectively. The output vector dimension of each layer is 768. Taking four layers will obtain features with dimension m = 768*4 = 3072
[0064] As shown above, in order to support the feature of dimension m = 3072, the minimum number of samples required is 30720. Therefore, when the number of training samples in the training set is less than 30720, (i.e., initial features) are processed by dimensionality reduction to obtain text features. The dimensionality reduction processing method is selected according to actual needs.
[0065] Furthermore, in one embodiment, the dimensionality reduction processing of the initial features adopts a principal component analysis method.
[0066] In this embodiment, the sampling principal component analysis method (PCA) finds the main variance direction in the data to reduce the dimension and reduce the redundancy of the data. In this embodiment, the dimension k after dimensionality reduction is determined by taking the eigenvalue matrix (λi, i∈[0,m]; λ1≥λ2≥…≥λm, which is the eigenvalue of the covariance matrix constructed in the PCA process) The eigenvalue λ1≥λ2≥…≥λk is used as the way to determine the dimension k. In general, PCA can usually significantly reduce the dimensionality of the data.
[0067] In this embodiment, feature extraction is performed based on the BERT-base pre-training model. The hidden layer of the BERT model contains a large amount of semantic information. The hidden states of different levels contain information of different levels. The output features of the shallow layer have strong semantic noise and insufficient expression ability. The deeper the layer, the stronger the semantics, but at the same time, some detailed information will be lost. By integrating the features of different layers of BERT, the purpose of simultaneously capturing the high-level semantics and detailed information of the text can be achieved, ensuring the completeness of the original text content capture.
[0068] Step S20, inputting the text features into the first classifier and the second classifier respectively;
[0069] In this embodiment, two classifiers are obtained by pre-training, which are recorded as a first classifier and a second classifier. Each classifier is used to predict the category to which the text feature belongs.
[0070] Furthermore, in one embodiment, before step S20, the method further includes:
[0071] A training set is constructed, wherein the training set includes a plurality of training samples, each training sample includes sample text features and category labeling information corresponding to a sample test bill of lading; a first classifier to be trained and a second classifier to be trained are trained based on the training set; and when a training end condition is met, a first classifier and a second classifier are obtained.
[0072] In this embodiment, refer to Figure 4 , Figure 4 This is a schematic diagram of a scenario for testing an embodiment of the method for automated auditing of bills of lading in this application. Figure 4As shown, the text attributes of the test bill of lading history data are defined, that is, the text content corresponding to each sample test bill of lading is determined, and the text encoder (the BERT model described above) is input to obtain the sample text features, and the expert annotation audit results (i.e., category annotation information) corresponding to the text content of the sample test bill of lading are obtained at the same time, thereby obtaining a training sample; each training sample is processed in the same way to obtain a training set including multiple training samples, that is, to construct a test bill of lading test set. Then, classifier training is performed based on the test set, thereby obtaining the trained first classifier and the second classifier, which are used for subsequent automated auditing of the test bill of lading to be audited.
[0073] Step S30, synthesizing the first classification result output by the first classifier and the second classification result output by the second classifier to obtain the audit result of the test bill of lading to be audited.
[0074] In this embodiment, the probabilities of the first classification result and the second classification result for different categories may be different, so the final audit result needs to be determined comprehensively.
[0075] Furthermore, in one embodiment, the first classification result includes a first probability value P of category i. i , i∈[1,Q], Q is the number of categories; the second classification result includes the second probability value P of category i i ′, i∈[1,Q], Q is the number of categories; step S30 includes:
[0076] Step S301: P corresponding to category i i and P i ′Weighted summation to obtain the overall probability value of category i;
[0077] Step S302: taking the category corresponding to the maximum overall probability value as the audit result of the test bill of lading to be audited.
[0078] In this embodiment, taking Q=3 as an example, the first classification result includes the first probability value P of category 1 1 , the first probability value P of category 2 2 And the first probability value P of category 3 3 The second classification result includes the second probability value P of category 1 1 ', the second probability value P of category 2 2 ' and the second probability value P of category 3 3 '.
[0079] The overall probability value of category 1
[0080] The overall probability value of category 2
[0081] Overall probability value for category 3
[0082] Wherein, w and w′ represent the weights of the first classifier and the second classifier respectively.
[0083] Furthermore, in one embodiment, the first classifier is a linear classifier, the second classifier is a nonlinear classifier, the first classifier and the second classifier are trained based on a training set, and when the number of training samples in the training set is less than or equal to a threshold, When the number of training samples in the training set is less than or equal to the threshold, Where w is P i The weight of P i ′’s weight.
[0084] In this embodiment, when the number of training samples is small, considering the characteristics of the linear classifier that the number of training samples is low and the classification generalization is good, a higher weight is given to the classification result obtained by the linear classifier. When the training sample reaches a certain scale, the nonlinear classifier has a more powerful modeling ability and can capture the complex nonlinear relationship in the data. The classification results obtained by the nonlinear classifier are given a higher weight.
[0085] This embodiment obtains a more powerful integrated classifier by weighted averaging the prediction results of the linear support vector machine (Linear SVM) and the nonlinear support vector machine (NonlinearSVM). This method can improve the linear and nonlinear classification performance and generalization of the classifier.
[0086] In this embodiment, two individual classifiers are selected, one is a linear support vector machine (Linear SVM, based on liblinear), and the other is a nonlinear support vector machine (Nonlinear SVM, based on libsvm). Linear support vector machine: has a strong classification ability for linearly separable samples, and the training cost is small. Nonlinear support vector machine: has a strong classification ability for linearly inseparable samples, and can form an effective complement to the linear support vector machine.
[0087] Classifier training: Referring to the above description of the training process of the first classifier and the second classifier, each training sample includes a sample text feature corresponding to a sample test bill of lading (hereinafter, for convenience of expression, it is represented by X) and category labeling information y = (0, 1, 2), that is, the training sample is expressed as (X, y), and the training set is expressed as: D = {(X 1 ,y 1 ), (X 2 ,y 2 ),......,(XN ,y N )}, N is the number of samples.
[0088] Among them, the linear support vector machine (Linear SVM) is a model used to solve linear classification problems. It divides samples of different categories by finding an optimal hyperplane. Its decision function can be expressed as: f linear (x) is the decision function of the linear SVM, X is the input sample feature vector, W linear is the normal vector of the hyperplane, b linear is the bias term. Through training, the weight vector W of the hyperplane can be determined linear and bias b linear .
[0089] Nonlinear support vector machine (Nonlinear SVM) uses kernel function to deal with nonlinear classification problems. Here we use the most common kernel function, radial basis function (RBF). Its decision function can be expressed as:
[0090]
[0091] f nonlinear (x) is the decision function of the nonlinear SVM, φ(x) is the mapping function, which maps the inseparable samples in the original space to the new linearly separable feature space; at the same time, in order to solve the difficulty of calculating the inner product of the high-dimensional new feature space, the kernel function is introduced during training x i , x j ∈D, so that the kernel technique and the Lagrange multiplier method can be used to obtain the dual problem of nonlinear SVM; α i is the Lagrange multiplier of the support vector, y i is the corresponding class label of the support vector, b nonlinear is the bias term. Through training, the Lagrange multiplier α of the support vector can be determined i and bias b nonlinear .
[0092] In the embodiment of the present application, the text content of the test bill of lading to be audited is subjected to feature extraction to obtain text features; the text features are respectively input into the first classifier and the second classifier; the first classification result output by the first classifier and the second classification result output by the second classifier are combined to obtain the audit result of the test bill of lading to be audited. Through the embodiment of the present application, the text content of the test bill of lading to be audited is subjected to feature extraction to obtain text features based on natural language processing technology, and the differentiated bill of lading description text with inconsistent expressions generated manually is converted into a unified and semantically consistent quantitative feature space to form a numerical feature expression, which is used as the input of the first and second classifiers, and finally, the output of the first and second classifiers is combined to obtain the audit result. Among them, the method based on machine learning can quickly handle the audit work for a large number of test bills of lading, thereby improving the speed of the audit; the method based on machine learning can identify problems and anomalies, reducing the possibility of human errors; avoiding manual participation in the audit work, and reducing the differences in audit results caused by inconsistent audit conditions such as auditors and audit environments.
[0093] In a second aspect, an embodiment of the present application also provides a test bill of lading automated auditing device.
[0094] In one embodiment, referring to Figure 5 , Figure 5 This is a functional module diagram of an embodiment of the automatic audit device for testing bills of lading in this application. Figure 5 As shown, the test bill of lading automatic audit device includes:
[0095] A feature extraction module 10 is used to extract features from the text content of the audit test bill of lading to obtain text features;
[0096] An input module 20, used to input the text features into the first classifier and the second classifier respectively;
[0097] The audit module 30 is used to synthesize the first classification result output by the first classifier and the second classification result output by the second classifier to obtain the audit result of the test bill of lading to be audited.
[0098] Furthermore, in one embodiment, the feature extraction module 10 is used to:
[0099] Input the text content of the test bill of lading to be audited into the BERT model and select the vector output by the target layer in the BERT model;
[0100] Based on the vector output by the target layer, text features are obtained.
[0101] Furthermore, in one embodiment, the target layer includes the 9th to 12th layers, and the feature extraction module 10 is used to:
[0102] The vectors output from the 9th to 12th layers are combined to obtain the initial features;
[0103] When the dimension of the initial feature is greater than the number of dimensions that can be supported by the number of training samples in the training set, the initial feature is subjected to dimensionality reduction processing to obtain text features, and the training set is used to train the first classifier and the second classifier.
[0104] Furthermore, in one embodiment, the dimensionality reduction processing of the initial features adopts a principal component analysis method.
[0105] Furthermore, in one embodiment, the first classification result includes a first probability value P of category i. i , i∈[1,Q], Q is the number of categories; the second classification result includes the second probability value P of category i i ′, i∈[1,Q], Q is the number of categories; the audit module 30 is used to:
[0106] P corresponding to category i i and P i ′Weighted summation to obtain the overall probability value of category i;
[0107] The category corresponding to the maximum overall probability value is taken as the audit result of the test bill of lading to be audited.
[0108] Furthermore, in one embodiment, the first classifier is a linear classifier, the second classifier is a nonlinear classifier, the first classifier and the second classifier are trained based on a training set, and when the number of training samples in the training set is less than or equal to a threshold, When the number of training samples in the training set is less than or equal to the threshold, Where w is P i The weight of P i ′’s weight.
[0109] Furthermore, in one embodiment, the test bill of lading automated auditing device further includes a training module for:
[0110] Constructing a training set, wherein the training set includes a plurality of training samples, each training sample includes a sample text feature and category labeling information corresponding to a sample test bill of lading;
[0111] Training a first classifier to be trained and a second classifier to be trained based on the training set;
[0112] When the training end condition is met, the first classifier and the second classifier are obtained.
[0113] Among them, the functional implementation of each module in the above-mentioned test bill of lading automated audit device corresponds to the various steps in the above-mentioned test bill of lading automated audit method embodiment, and its functions and implementation processes will not be repeated here one by one.
[0114] In a third aspect, an embodiment of the present application provides a test bill of lading automated auditing device, which may be a personal computer (PC), a laptop computer, a server, or other device with data processing capabilities.
[0115] Reference Figure 6 , Figure 6 The hardware structure diagram of the test bill automatic audit device involved in the embodiment of the present application is shown in FIG. In the embodiment of the present application, the test bill automatic audit device may include a processor, a memory, a communication interface, and a communication bus.
[0116] The communication bus may be of any type and is used to interconnect the processor, the memory, and the communication interface.
[0117] The communication interface includes input / output (I / O) interface, physical interface and logical interface, etc., which are used to realize the interconnection of devices inside the test order automation audit device, and the interface used to realize the interconnection of the test order automation audit device with other devices (such as other computing devices or user devices). The physical interface can be an Ethernet interface, a fiber optic interface, an ATM interface, etc.; the user device can be a display, a keyboard, etc.
[0118] The memory can be various types of storage media, such as random access memory (RAM), read-only memory (ROM), non-volatile RAM (NVRAM), flash memory, optical storage, hard disk, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), etc.
[0119] The processor may be a general-purpose processor, and the general-purpose processor may call the test bill of lading automated audit program stored in the memory, and execute the test bill of lading automated audit method provided by the embodiment of the present application. For example, the general-purpose processor may be a central processing unit (CPU). Among them, the method executed when the test bill of lading automated audit program is called can refer to the various embodiments of the test bill of lading automated audit method of the present application, which will not be repeated here.
[0120] Those skilled in the art will understand that Figure 6 The hardware structure shown in the figure does not constitute a limitation on the present application, and may include more or less components than shown in the figure, or combine certain components, or arrange the components differently.
[0121] In a fourth aspect, an embodiment of the present application also provides a computer-readable storage medium.
[0122] The computer-readable storage medium of the present application stores a test bill automation audit program, wherein when the test bill automation audit program is executed by a processor, the steps of the test bill automation audit method as described above are implemented.
[0123] Among them, the method implemented when the test bill of lading automated audit program is executed can refer to the various embodiments of the test bill of lading automated audit method of this application, and will not be repeated here.
[0124] It should be noted that the serial numbers of the above-mentioned embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments.
[0125] The terms "including" and "having" and any variations thereof in the specification and claims of this application and the above-mentioned drawings are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally includes steps or units that are not listed, or optionally includes other steps or units inherent to these processes, methods, products or devices. The terms "first", "second" and "third" are used to distinguish different objects, etc., and do not represent a sequence, nor do they limit "first", "second" and "third" to different types.
[0126] In the description of the embodiments of the present application, "exemplary", "for example" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary", "for example" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary", "for example" or "for example" is intended to present related concepts in a specific way.
[0127] In the description of the embodiments of the present application, unless otherwise specified, “ / ” means or, for example, A / B can mean A or B; the “and / or” in the text is merely a description of the association relationship of associated objects, indicating that three relationships may exist, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the embodiments of the present application, “multiple” refers to two or more than two.
[0128] In some processes described in the embodiments of the present application, multiple operations or steps that appear in a specific order are included, but it should be understood that these operations or steps may not be executed in the order in which they appear in the embodiments of the present application or in parallel, and the sequence number of the operation is only used to distinguish the different operations, and the sequence number itself does not represent any execution order. In addition, these processes may include more or fewer operations, and these operations or steps may be executed in sequence or in parallel, and these operations or steps may be combined.
[0129] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, disk, CD) as described above, and includes a number of instructions for a terminal device to execute the methods described in each embodiment of the present application.
[0130] The above are only preferred embodiments of the present application, and are not intended to limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A test bill of lading automated audit method, characterized in that: The test bill of lading automated audit method comprises: Perform feature extraction on the text content of the audit test bill of lading to obtain text features; Inputting the text features into the first classifier and the second classifier respectively; The audit result of the test bill of lading to be audited is obtained by combining the first classification result output by the first classifier and the second classification result output by the second classifier.
2. The test bill of lading automated audit method according to claim 1, characterized in that: The feature extraction of the text content of the audited test bill of lading obtains the following text features: Input the text content of the test bill of lading to be audited into the BERT model and select the vector output by the target layer in the BERT model; Based on the vector output by the target layer, text features are obtained.
3. The test bill of lading automated audit method according to claim 2, characterized in that: The target layer includes the 9th to 12th layers, and the text features obtained based on the vector output by the target layer include: The vectors output from the 9th to 12th layers are combined to obtain the initial features; When the dimension of the initial feature is greater than the number of dimensions that can be supported by the number of training samples in the training set, the initial feature is subjected to dimensionality reduction processing to obtain text features, and the training set is used to train the first classifier and the second classifier.
4. The test bill of lading automated audit method according to claim 3, characterized in that: The dimensionality reduction process for the initial features adopts a principal component analysis method.
5. The test bill of lading automated audit method according to claim 1, characterized in that: The first classification result includes a first probability value P of category i i , i∈[1,Q], Q is the number of categories; the second classification result includes the second probability value P of category i i ′, i∈[1,Q], Q is the number of categories; based on the first classification result output by the first classifier and the second classification result output by the second classifier, the audit result of the test bill of lading to be audited includes: P corresponding to category i i and P i ′Weighted summation to obtain the overall probability value of category i; The category corresponding to the maximum overall probability value is taken as the audit result of the test bill of lading to be audited.
6. The test bill of lading automated audit method according to claim 5, characterized in that: The first classifier is a linear classifier, the second classifier is a nonlinear classifier, the first classifier and the second classifier are trained based on a training set, and when the number of training samples in the training set is less than or equal to a threshold, When the number of training samples in the training set is less than or equal to the threshold, Where w is P i The weight of P i ′’s weight.
7. The test bill automated audit method according to any one of claims 1 to 6, characterized in that: Before the text features are input into the first classifier and the second classifier respectively, the method further includes: Constructing a training set, wherein the training set includes a plurality of training samples, each training sample includes a sample text feature and category labeling information corresponding to a sample test bill of lading; Training a first classifier to be trained and a second classifier to be trained based on the training set; When the training end condition is met, the first classifier and the second classifier are obtained.
8. A test bill of lading automated audit device, characterized in that: The test bill of lading automatic audit device comprises: A feature extraction module is used to extract features from the text content of the audit test bill of lading to obtain text features; An input module, used for inputting the text features into the first classifier and the second classifier respectively; The audit module is used to synthesize the first classification result output by the first classifier and the second classification result output by the second classifier to obtain the audit result of the test bill of lading to be audited.
9. A test bill of lading automated auditing device, characterized in that: The test order automated audit device includes a processor, a memory, and a test order automated audit program stored in the memory and executable by the processor, wherein when the test order automated audit program is executed by the processor, the steps of the test order automated audit method as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a test order automation audit program, wherein when the test order automation audit program is executed by the processor, the steps of the test order automation audit method according to any one of claims 1 to 7 are implemented.