Express business problem solving method and system based on knowledge base management system

By applying multimodal data fusion and real-time knowledge matching technology based on the knowledge base management system on the high-speed sorting line, the classification error problem caused by package label destruction and image recognition errors is solved, real-time error correction and system self-update are achieved, and sorting efficiency and customer satisfaction are improved.

CN119740941BActive Publication Date: 2025-06-06ZHEJIANG WANCUN INTERNET EXPRESS CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510247621.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-06
Estimated Expiration
2045-03-04

AI Technical Summary

Technical Problem

The prior art is difficult to achieve real-time error correction on high-speed sorting lines, resulting in classification error problems caused by package label destruction and image recognition errors seriously affect sorting efficiency and customer satisfaction.

Method used

Using a method based on the knowledge base management system, through multimodal data fusion and real-time knowledge matching, the package morphology geometric topology matrix and the face single semantic keyword vector are obtained, structured analysis and exception marking processing are carried out, many-to-many mapping relationships are established, the association matrix and graph node state matrix are generated, the cross-modal mapping function cluster is constructed, and the knowledge recommendation list is generated to solve error correction problems in express delivery services.

Benefits of technology

It realizes effective real-time error correction for classification errors on the high-speed sorting line, improves sorting efficiency and customer satisfaction, and improves the adaptability and practicality of the system through self-updating and iteration of the knowledge base.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119740941B_ABST
    Figure CN119740941B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of express business management, and discloses a method and system for solving express business problems based on a knowledge base management system. The method first obtains the package form and semantic features, obtains a graph node state matrix through structured analysis, generates an abnormal marking matrix, establishes a mapping relationship, and then obtains multi-level calibration parameters, constructs a cross-modal mapping function cluster to obtain a package feature matrix, and then designs a knowledge base retrieval and sorting mechanism to generate a knowledge recommendation list. Finally, a solution is generated based on the list and fed back to the knowledge base. The method significantly improves the accuracy and efficiency of express business processing through multimodal data processing and the use of a knowledge base, enhances the practicality and adaptability of the knowledge base, and provides strong support for the intelligent development of express business.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of express business management, and more specifically, to a method and system for solving express business problems based on a knowledge base management system. Background Art

[0002] With the rapid development of express delivery business, parcel sorting efficiency and accuracy have become the key to express delivery companies improving service quality. On high-speed sorting lines, problems such as parcel label damage and image recognition errors frequently lead to classification errors, seriously affecting sorting efficiency and customer satisfaction. When dealing with these problems, traditional sorting systems are often unable to call historical abnormal cases for error correction in real time, resulting in a high error rate. Existing technologies have obvious deficiencies in real-time knowledge matching and multimodal data fusion, and it is difficult to meet the real-time error correction needs of high-speed sorting lines.

[0003] The Chinese patent application with publication number CN105678490A discloses a system and method for processing the Internet of Things technology for express delivery services. Through components such as an identification code generation system, an identification code feedback module, an express company backend information management platform, an express document printer, and an embedded intelligent management module, the automation and informatization of the express delivery service processing process are realized. However, the system mainly focuses on the overall process optimization of the express delivery service, lacks support for the real-time error correction function on the high-speed sorting line, and cannot effectively handle the classification error problem caused by the contamination of the package label and the image recognition error. The Chinese patent application with publication number CN108335059A discloses an express delivery service processing method and an express delivery service processing device. By providing multiple first terminal identification information and its associated business object information, the synchronous processing of collection and delivery is realized, and the utilization rate of personnel and the responsiveness of the mail are improved. However, the method mainly focuses on the collection and delivery links of the express delivery service, fails to solve the real-time error correction problem on the high-speed sorting line, and lacks a processing mechanism for abnormal situations such as contamination of the package label and image recognition errors.

[0004] The existing technology has obvious deficiencies in real-time error correction on high-speed sorting lines and cannot effectively handle classification errors caused by damaged package labels and image recognition errors. Traditional systems lack the ability to call historical abnormal cases for error correction in real time, resulting in a high sorting error rate, which seriously affects sorting efficiency and customer satisfaction. Summary of the invention

[0005] In order to overcome the above-mentioned defects of the prior art, the present invention provides an express business problem solving method and system based on a knowledge base management system, which can effectively solve the classification error problem on the high-speed sorting line through multimodal data fusion and real-time knowledge matching.

[0006] To achieve the above object, the present invention provides the following technical solutions:

[0007] The express delivery business problem-solving method based on the knowledge base management system includes:

[0008] Obtain the package geometry topology matrix to form a package geometry feature set; obtain the waybill semantic keyword vector to form a waybill semantic feature set; perform structured analysis on the package geometry topology matrix to divide the packages into three package priorities, and generate an abnormal marking matrix for the waybill semantic keyword vector; establish a many-to-many mapping relationship between the package priority and the abnormal marking matrix to generate an association matrix; dynamically generate a graph node state matrix based on the association matrix;

[0009] According to the graph node state matrix, multi-level calibration parameters are obtained; based on the multi-level calibration parameters, a cross-modal mapping function cluster is constructed to obtain the package feature matrix; for the package feature matrix, a knowledge base retrieval and sorting mechanism is designed to generate a knowledge recommendation list;

[0010] Based on the knowledge recommendation list, solutions to express delivery business problems are generated, and the newly generated solutions are fed back to the knowledge base to achieve self-update and iteration of the knowledge base.

[0011] Furthermore, the obtaining of the package morphology geometric topology matrix to form a package morphology feature set includes:

[0012] The multi-angle high-definition cameras deployed at the key nodes of the sorting line collect the video image sequence of the package delivery in real time; the collected video image sequence is segmented and feature extracted frame by frame, and a lightweight CNN is used to generate a 128-dimensional package shape geometric topology matrix to form a package shape feature set;

[0013] The method of obtaining the semantic keyword vector of the delivery note and forming the semantic feature set of the delivery note includes: performing OCR analysis on the package delivery note information through a high-speed scanner equipped with a sorting line, extracting key text information from the delivery note, and converting it into a 64-dimensional delivery note semantic keyword vector to form a delivery note semantic feature set.

[0014] Furthermore, the structural analysis of the package geometry topology matrix is ​​performed to classify the packages into three package priorities, including:

[0015] Calculate the variance of each dimension feature for the package geometry topology matrix and the corresponding confidence level ;according to and , the packages are divided into three package priorities: A, B, and C; among them, is the variance of the j-th dimension feature in the 128-dimensional package morphology geometric topology matrix, for The corresponding confidence level.

[0016] Furthermore, according to and , the packages are divided into three categories: A, B, and C. The package priorities include:

[0017] Set the first variance threshold V 1 and the second variance threshold V 2 , where V 1 <V 2 ;like <V 1 , then give Weight W 1 ; If V 1 ≤ <V 2 , then give Weight W 2 ;like ≥V 2 , then give Weight W 3; Where W 1 >W 2 >W 3 ;

[0018] For each package’s 128-dimensional package geometry topology matrix, according to the confidence level and the weight W assigned 1 , W 2 , W 3 , calculate the weighted average to obtain the comprehensive confidence score CON of the package;

[0019] Set the first confidence threshold T 1 and the second confidence threshold T 2 , where T 1 >T 2 ; If CON>T 1 , the package is classified as Class A; if T 2 <CON≤T 1 , the package is classified as Class B; if CON≤T 2 , the package is classified as Class C.

[0020] Furthermore, generating an abnormal label matrix for the single semantic keyword vector includes:

[0021] For the 64-dimensional semantic keyword vector of a face, the frequency of occurrence of semantic keywords in each dimension is counted, and the probability distribution of the frequency of occurrence is calculated; the information entropy H of each dimension is calculated based on the probability distribution k ;H k represents the information entropy of the kth dimension of the semantic keyword vector of the face order, where 1≤k≤64; set the information entropy threshold E, when Hk >E, the corresponding abnormal marking matrix position is marked as 0; when H k When ≤E, the corresponding position of the abnormal marking matrix is ​​marked with 1; the number of abnormal markings for each bit in the abnormal marking matrix is ​​counted, the missing probability of each semantic keyword is calculated, and a missing probability vector corresponding to the dimension of the semantic keyword vector of the face order is generated; the missing probability vector is normalized to generate a standardized abnormal marking matrix.

[0022] Furthermore, the dynamically generating a graph node state matrix according to the association matrix includes:

[0023] Deconstruct the association matrix into two sub-matrices: priority vector and anomaly mark vector;

[0024] With the priority vector as the skeleton and the abnormal mark vector as the decoration, a cross-modal semantic association graph with a dynamic structure is generated through a random walk algorithm. Each node in the cross-modal semantic association graph represents a package priority category, and the edges between nodes represent the transition probability between package priorities.

[0025] Embedded learning is performed on the cross-modal semantic association graph to obtain the graph node state matrix.

[0026] Furthermore, obtaining the multi-level calibration parameters according to the graph node state matrix includes:

[0027] Taking the graph node state matrix as input, the hidden layer feature vector is extracted by stacking denoising autoencoders; the hidden layer feature vector is soft-thresholded, and the amplitude of the hidden layer feature vector is used as the significance score to generate the weight coefficient matrix of the morphological feature attenuation factor and the weight coefficient matrix of the semantic compensation coefficient; the weight coefficient matrix of the morphological feature attenuation factor and the weight coefficient matrix of the semantic compensation coefficient are normalized to obtain the normalized morphological feature attenuation factor weight vector and the normalized semantic compensation coefficient weight vector;

[0028] The normalized morphological feature attenuation factor weight vector is dot-producted with the hidden layer feature vector to obtain the morphological feature attenuation factor; the normalized semantic compensation coefficient weight vector is dot-producted with the hidden layer feature vector to obtain the semantic feature compensation coefficient; the morphological feature attenuation factor and the semantic feature compensation coefficient constitute a multi-level calibration parameter.

[0029] Furthermore, the construction of a cross-modal mapping function cluster to obtain a parcel feature matrix includes:

[0030] The package morphological feature set is taken as the morphological feature matrix A, and the order semantic feature set is taken as the semantic feature matrix B, and a semantic association matrix M of the morphological feature matrix A and the semantic feature matrix B is generated;

[0031] The semantic association matrix M is decomposed into three elements by matrix decomposition method to generate semantic topic matrix T, morphological semantic matching matrix P, and semantic morphological matching matrix Q.

[0032] The morphological feature matrix A is multiplied by the semantic morphological matching matrix Q to obtain a first matrix; the semantic feature matrix B is multiplied by the morphological semantic matching matrix P to obtain a second matrix; the first matrix and the second matrix have the same dimension;

[0033] The morphological feature attenuation factor is used as the weight of the first matrix, the semantic feature compensation coefficient is used as the weight of the second matrix, the first matrix and the second matrix are fused to establish a cross-modal mapping function cluster; according to the cross-modal mapping function cluster, the parcel feature matrix is ​​obtained.

[0034] Furthermore, the generation of a semantic association matrix M of the morphological feature matrix A and the semantic feature matrix B includes: subject clustering the semantic feature matrix B to obtain m semantic topics; calculating the association strength of each package on the m semantic topics to generate the semantic association matrix M; the semantic association matrix M is an m'×m matrix, where m' represents the number of packages.

[0035] Furthermore, generating a knowledge recommendation list includes:

[0036] Divide the package feature matrix into n2 hash buckets, each hash bucket corresponds to a knowledge subgraph of the knowledge base;

[0037] In each hash bucket, the structural similarity between the query matrix and the knowledge subgraph is detected, and semantic association edges are dynamically added to obtain an updated knowledge subgraph; the query matrix is ​​obtained by wrapping the feature matrix;

[0038] Perform random walks on the updated knowledge subgraph, calculate the importance score of each knowledge point, and generate a sorted knowledge recommendation list.

[0039] Furthermore, the generation of solutions to express business problems based on the knowledge recommendation list includes: clustering the express business problems according to the knowledge recommendation list to obtain problem clusters; for each problem cluster, calling relevant knowledge points in the knowledge recommendation list, combining the graph node state matrix for case reasoning, and generating a set of candidate solutions.

[0040] The express business problem solving system based on the knowledge base management system is used to implement the express business problem solving method based on the knowledge base management system. The system includes:

[0041] Package priority classification module: used to obtain the package geometry topology matrix to form a package feature set; perform structured analysis on the package geometry topology matrix to classify the packages into three package priorities;

[0042] Graph space fusion module: used to obtain the semantic keyword vector of the waybill to form the semantic feature set of the waybill; generate the abnormal marking matrix of the semantic keyword vector of the waybill; establish a many-to-many mapping relationship between the package priority and the abnormal marking matrix to generate the association matrix; dynamically generate the graph node state matrix according to the association matrix;

[0043] Knowledge recommendation module: According to the graph node state matrix, multi-level calibration parameters are obtained; based on the multi-level calibration parameters, a cross-modal mapping function cluster is constructed to obtain the package feature matrix; for the package feature matrix, a knowledge base retrieval and sorting mechanism is designed to generate a knowledge recommendation list;

[0044] Solution generation module: Generates solutions to express delivery business problems based on the knowledge recommendation list, and feeds the newly generated solutions back to the knowledge base to achieve self-update and iteration of the knowledge base.

[0045] Compared with the prior art, the present invention has the following beneficial effects:

[0046] The present invention realizes a closed-loop process from problem identification to solution generation and knowledge base update by comprehensively collecting the morphological and semantic features of the package, performing structured analysis, multimodal fusion, and knowledge base retrieval and sorting. During the package processing process, the package priority can be accurately divided, and the missing semantic information of the waybill can be effectively identified, providing an accurate basis for subsequent processing. By using cross-modal mapping and knowledge graph technology, multimodal data is deeply integrated to make the acquired knowledge recommendation list more targeted. At the same time, generating solutions based on knowledge recommendations and feeding back to the knowledge base can not only efficiently solve various problems in the express delivery business, but also continuously optimize the knowledge base and improve its adaptability and practicality. On the whole, this method comprehensively improves the accuracy and efficiency of express delivery business processing, enhances the ability of enterprises to cope with complex business scenarios, and effectively promotes the intelligent development of express delivery business. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0048] Figure 1 It is a principle flow chart of the express business problem solving method based on the knowledge base management system in the present invention;

[0049] Figure 2A flow chart of a method for forming a package morphological feature set and a delivery order semantic feature set in a solution method for express business problems based on a knowledge base management system of the present invention;

[0050] Figure 3 A flow chart of a method for generating an abnormal marking matrix in a solution method for express business problems based on a knowledge base management system of the present invention;

[0051] Figure 4 A flow chart of a method for dynamically generating a graph node state matrix according to an association matrix in a solution method for express business problems based on a knowledge base management system of the present invention;

[0052] Figure 5 A flow chart of a method for obtaining multi-level calibration parameters according to a graph node state matrix in a solution method for express business problems based on a knowledge base management system of the present invention;

[0053] Figure 6 A flow chart of a method for constructing a cross-modal mapping function cluster and obtaining a package feature matrix in a solution method for express business problems based on a knowledge base management system of the present invention;

[0054] Figure 7 A flow chart of a method for generating a knowledge recommendation list in a solution to express business problems based on a knowledge base management system of the present invention;

[0055] Figure 8 It is a functional module diagram of the express business problem solving system based on the knowledge base management system in the present invention. DETAILED DESCRIPTION

[0056] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0057] Example 1

[0058] See also Figure 1 As shown, this embodiment provides a solution to express business problems based on a knowledge base management system, including:

[0059] Step S1000, obtaining a package morphology geometric topology matrix to form a package morphology feature set; obtaining a waybill semantic keyword vector to form a waybill semantic feature set; performing a structured analysis on the package morphology geometric topology matrix, dividing the packages into three package priorities, and generating an abnormal marking matrix for the waybill semantic keyword vector; establishing a many-to-many mapping relationship between the package priorities and the abnormal marking matrix to generate an association matrix; dynamically generating a graph node state matrix according to the association matrix;

[0060] Furthermore, step S1000 includes:

[0061] Step S1100, obtaining a package geometry topology matrix to form a package geometry feature set; obtaining a waybill semantic keyword vector to form a waybill semantic feature set;

[0062] Furthermore, if Figure 2 As shown, step S1100 includes:

[0063] Step S1110, using multi-angle high-definition cameras deployed at key nodes of the sorting line to collect video image sequences of the transported packages in real time;

[0064] Step S1120, performing frame-by-frame segmentation and feature extraction on the collected video image sequence, using a lightweight CNN to generate a 128-dimensional package morphology geometric topology matrix to form a package morphology feature set;

[0065] Step S1130, perform OCR analysis on the parcel delivery note information through the high-speed scanner of the sorting line, extract the key text information in the delivery note, and convert it into a 64-dimensional delivery note semantic keyword vector to form a delivery note semantic feature set.

[0066] Specifically, in sub-step S1110, multi-angle high-definition cameras are deployed at key nodes of the sorting line. These cameras can collect video image sequences of the conveyed packages in real time. The reason for choosing to deploy at key nodes of the sorting line is that these locations can fully and accurately capture the appearance information of the packages, ensuring that the collected data is representative. For example, in a large express sorting center, the packages move on the conveyor belt, and the cameras at key nodes can clearly capture all sides of the packages, avoiding blind spots in shooting, and providing guarantees for subsequent accurate analysis of the package shape.

[0067] In sub-step S1120, the captured video image sequence is segmented and features are extracted frame by frame. Frame segmentation is to split the continuous video images into static images frame by frame, so as to analyze the package features in each image in more detail. A lightweight convolutional neural network (CNN) is used to generate a 128-dimensional package morphology geometric topology matrix, and then form a package morphology feature set. Lightweight CNN is an optimized neural network architecture. While ensuring feature extraction capabilities, it has the characteristics of small calculation amount and fast running speed. It is very suitable for use in scenarios such as express delivery business that require real-time processing of large amounts of data. The 128-dimensional package morphology geometric topology matrix is ​​used to accurately depict the appearance contour of the package. Each dimension represents a specific feature of the package shape. For example, the aspect ratio of the package, the corner radian and other information may correspond to different dimensions. Through such a matrix representation, the complex appearance of the package can be converted into digital features that are easy for the computer to process, which facilitates subsequent analysis.

[0068] In sub-step S1130, the optical character recognition (OCR) analysis of the package delivery bill information is performed by the high-speed scanner equipped with the sorting line. OCR technology can convert the text information on the delivery bill into computer-recognizable text data, extract key text information, such as the sender and recipient, origin, destination, etc., and convert it into a 64-dimensional delivery bill semantic keyword vector, thereby forming a delivery bill semantic feature set. Taking an actual package as an example, the "a city b district c street d number" filled in on the delivery bill is used as the destination information. After being extracted by OCR analysis, it is converted into the values ​​of several dimensions in the semantic keyword vector. These values ​​reflect the semantic features related to the transportation destination of the package.

[0069] Step S1100 has many beneficial effects by collecting multi-source heterogeneous features of express parcels online in real time. From the perspective of data acquisition, the collaborative work of multi-angle high-definition cameras and high-speed scanners ensures the comprehensive collection of parcel morphology and semantic information, avoiding information omissions. This provides sufficient data support for the subsequent accurate analysis of parcels and solving express business problems. In terms of improving business processing efficiency, the application of lightweight CNN and high-speed scanners enables data collection and feature extraction to be completed quickly, adapting to the characteristics of the huge number of parcels and high processing speed requirements in the express business. From the perspective of intelligent processing, the acquisition of these multi-source heterogeneous features lays the foundation for the subsequent use of deep learning and other technologies for intelligent analysis, which helps to realize the automation and intelligence of express delivery services. For example, it can automatically allocate transportation routes and predict delivery times based on parcel morphology and semantic features, thereby improving the operational efficiency and service quality of the entire express business.

[0070] Step S1200, perform structured analysis on the package geometry topology matrix, divide the packages into three package priorities, and generate an abnormal marking matrix for the single semantic keyword vector; establish a many-to-many mapping relationship between the package priority and the abnormal marking matrix, and generate an association matrix.

[0071] Further, step S1200 includes:

[0072] Step S1210, for the package morphology geometric topology matrix, calculate the variance of each dimensional feature and the corresponding confidence level ;according to and , the packages are divided into three package priorities: A, B, and C; among them, is the variance of the j-th dimension feature in the 128-dimensional package morphology geometric topology matrix, for The corresponding confidence level;

[0073] Further, step S1210 includes:

[0074] Step S1211, setting the first variance threshold V 1 and the second variance threshold V 2 , where V 1 <V 2 ;like <V 1 , then give Weight W 1; If V 1 ≤ <V 2 , then give Weight W 2; like ≥V 2 , then give Weight W 3; Where W 1 >W 2 >W 3 ;

[0075] Step S1212, for each package's 128-dimensional package geometry topology matrix, according to the confidence level and the weight W assigned 1 , W 2 , W 3 , calculate the weighted average to obtain the comprehensive confidence score CON of the package;

[0076] Step S1213, setting a first confidence threshold T 1 and the second confidence threshold T 2 , where T1 >T 2 ; If CON>T 1 , the package is classified as Class A; if T 2 <CON≤T 1 , the package is classified as Class B; if CON≤T 2 , the package is classified as Class C.

[0077] Specifically, in sub-step S1211, for the 128-dimensional package morphology geometric topology matrix, the variance of each dimensional feature is calculated. and the corresponding confidence level Variance is an important indicator used in statistics to measure the degree of data dispersion. In the context of the package geometry topology matrix, variance It reflects the variation of the j-th dimension feature between different packages. For example, if a dimension represents the length of a package, a small variance of this dimension means that most packages have little difference in length, the data is more concentrated, and the corresponding eigenvalue is more stable and reliable; conversely, a large variance indicates that the packages have obvious differences in the dimensional feature, and the stability and reliability of the eigenvalue are low.

[0078] In order to more effectively evaluate the reliability of each dimension feature, the first variance threshold V is set 1 and the second variance threshold V 2 (V 1 <V 2 ), and assign confidence levels based on the relationship between the variance and the threshold Different weights. When the variance value on the dimension is less than V 1 When , it means that the dimension feature changes very little between different packages, and its feature value confidence is high, and the weight W is given. 1 ; When the variance value is V 1 and V 2 When the eigenvalue confidence is medium, the weight W is assigned 2 ; When the variance value is greater than or equal to V 2 When , the confidence of the eigenvalue is low, and the weight W is assigned 3 , and W 1 >W 2 >W 3 This weight distribution method based on variance and threshold can quantify the confidence level of each dimensional feature, providing an important basis for the subsequent comprehensive evaluation of the reliability of package morphological features. For example, in actual operation, if the variance of the "aspect ratio" dimension of a package is less than V 1 , indicating that the package has a higher consistency in aspect ratio compared with other packages, and contributes more to the reliability of judging the package shape, so a higher weight W is given. 1 .

[0079] In sub-step S1212, the 128-dimensional package geometry topology matrix of each package is sorted according to the confidence level. The weights assigned are used to calculate the weighted average and obtain the comprehensive confidence score CON of the package. This calculation method comprehensively considers the confidence of each dimensional feature in the matrix, avoiding the influence of the deviation of a single dimensional feature on the overall evaluation. Because the package shape is a complex multi-dimensional feature set, the importance of features of different dimensions to the overall judgment of the package shape is different. The weighted average can more objectively and accurately reflect the overall reliability of the package shape features. For example, some dimensional features of a package have a small variance, high confidence, and a large weight, and their contribution to the comprehensive confidence score is greater; while other dimensions have a large variance, low confidence, and a small weight, and their influence in the comprehensive evaluation is relatively weak. This weighted calculation method can more reasonably reflect the role of each dimensional feature in the overall evaluation.

[0080] In sub-step S1213, a first confidence threshold T is set. 1 and the second confidence threshold T 2 (T 1 >T 2 ), prioritize the packages according to the comprehensive confidence score CON. If CON>T 1 , the package is classified as Class A; if T 2 <CON≤T 1 , it is classified as Class B; if CON≤T 2 , it is classified as Class C. For example, assuming T 1 The value is 90%, T 2 The value is 60%. When the calculated CON of a package is 95%, it means that the reliability of the morphological characteristics of the package is very high, and it is classified as Class A; if CON is 70%, it is classified as Class B; if CON is 50%, it is classified as Class C.

[0081] This series of steps has brought about many beneficial effects. From the perspective of express delivery business resource allocation, by dividing the parcel priorities, resources can be reasonably arranged according to the importance and processing difficulty of the parcel. For Class A parcels, due to their high reliability of morphological characteristics, they can be processed first, reducing waiting time, improving overall processing efficiency, and ensuring that important parcels can be delivered quickly and accurately; for Class C parcels, they can be processed when resources are sufficient to avoid affecting the overall business progress due to excessive attention to low-priority parcels, thereby achieving optimal resource allocation. From the perspective of risk control, the division of different priorities helps to identify potential risks. Class C parcels may have unstable morphological characteristics. In subsequent processing, they can be focused on and countermeasures can be prepared in advance, such as increasing manual verification links to prevent transportation and sorting errors, thereby reducing business risks. In terms of improving data processing efficiency and accuracy, this quantitative priority division method makes subsequent analysis and processing of parcels more targeted. For example, in an automatic sorting system, the sorting strategy can be adjusted according to the parcel priority, and different sorting processes and equipment parameters can be used for parcels of different priorities to reduce unnecessary waste of computing resources, improve sorting accuracy and efficiency, and thus improve customer satisfaction. In addition, this priority division method also provides strong support for the intelligent development of express delivery business, facilitating integration with other intelligent systems to achieve more efficient and accurate express delivery services.

[0082] Step S1220, performing missing value analysis on the face order semantic keyword vector to generate an abnormal labeling matrix corresponding to the dimension of the face order semantic keyword vector;

[0083] Furthermore, if Figure 3 As shown, step S1220 includes:

[0084] Step S1221, for the 64-dimensional semantic keyword vector of the face sheet, count the occurrence frequency of the semantic keyword in each dimension and calculate the probability distribution of the occurrence frequency;

[0085] Step S1222, calculate the information entropy H of each dimension according to the probability distribution k ;H k represents the information entropy of the k-th dimension of the semantic keyword vector of the face order, where 1≤k≤64;

[0086] Step S1223, set the information entropy threshold E, when H k >E, the corresponding abnormal marking matrix position is marked as 0; when H k When ≤E, the corresponding abnormal marking matrix position is marked with 1;

[0087] Step S1224, counting the number of abnormal markings for each bit in the abnormal marking matrix, calculating the missing probability of each semantic keyword, and generating a missing probability vector corresponding to the dimension of the semantic keyword vector of the face order;

[0088] Step S1225 , normalize the missing probability vector to generate a standardized abnormal labeling matrix.

[0089] Specifically, the semantic keyword vector of the package bill is obtained by performing OCR analysis on the package bill information and extracting key text information. Each dimension represents specific semantic features such as sender / receiver, origin, and destination. Taking "recipient name" as an example, when counting large-scale package data, the number of times it appears in all package bills is counted and then divided by the total number of packages to obtain the frequency of occurrence of the semantic keyword. The calculation of probability distribution is to normalize the frequency of occurrence of each semantic keyword so that the sum of the frequency of occurrence of all semantic keywords is 1. This can more clearly show the relative probability of occurrence of each semantic keyword in the overall data, which is helpful for the subsequent analysis of the regularity and stability of the occurrence of semantic keywords. Information entropy is an important concept in information theory, which is used to measure the uncertainty of a random variable. For a certain dimension k of the semantic keyword vector of the package bill, the calculation of its information entropy Hk is based on the probability distribution of the frequency of occurrence of the semantic keywords in this dimension. When the probability distribution of semantic keywords on a dimension in different packages is relatively uniform, it means that the semantic features represented by the dimension are varied and have high uncertainty, and the information entropy is larger; on the contrary, if the probability distribution is very unbalanced, such as a semantic keyword either appears frequently or is missing frequently, then the information entropy of the dimension is small. For example, if the semantic keyword "destination" is distributed in a large number of different regions in many packages, its probability distribution is uniform, and the information entropy is large; if it is concentrated in a few regions, the probability distribution is concentrated, and the information entropy is relatively small. By calculating the information entropy, the uncertainty of the semantic features of each dimension can be quantified, providing a numerical basis for the subsequent judgment of the stability of the semantic features.

[0090] The information entropy threshold E is set to judge the stability of the semantic features and mark the abnormal label matrix accordingly. k >E, it indicates that the semantic features of this dimension are stable, and the corresponding abnormal labeling matrix position is marked as 0; when H k ≤E, it means that the semantic features of this dimension are unstable and prone to missing, and the corresponding abnormal marking matrix position is marked as 1. Assuming that the information entropy threshold E is set to a certain value, for the semantic keyword dimension "sender contact information", if the calculated information entropy H k If H is greater than E, it means that among many packages, the sender's contact information is relatively stable and less likely to be missing, so the corresponding position in the abnormal marking matrix is ​​marked as 0; on the contrary, if Hk If it is less than or equal to E, it means that the semantic keyword is filled in very differently in different packages and is easily missing, and the corresponding position is marked as 1. This marking method based on information entropy threshold can automatically identify unstable semantic features, quickly locate missing anomalies in the processing of massive package data, and facilitate subsequent targeted processing.

[0091] Sub-step S1224 counts the number of abnormal marks for each bit in the abnormal mark matrix, calculates the missing probability of each semantic keyword, and then generates a missing probability vector corresponding to the dimension of the semantic keyword vector of the face sheet. For example, after counting the semantic keyword "recipient address" of 1,000 packages, it is found that the abnormal mark (marked as 1) appears 100 times, then the missing probability of the semantic keyword "recipient address" is 100÷1000=0.1, and the value of the corresponding dimension in the missing probability vector is 0.1. In this way, the abnormal mark situation is quantified into a specific missing probability value, so that the degree of missing semantic features can be accurately reflected, providing more quantitative data support for the subsequent more accurate evaluation of the quality of face sheet information and hierarchical fusion processing. Sub-step S1225 normalizes the missing probability vector to generate a standardized abnormal mark matrix. Normalization is a common data processing method that maps each value in the missing probability vector to the range of 0-1, so that the missing probabilities of different dimensions are comparable. For example, if the missing probabilities of some dimensions in the missing probability vector are 0.2, 0.3, and 0.1 respectively, after normalization, they may become 0.4, 0.6, and 0.2 (this is just an example, the actual calculation is based on the specific normalization formula). In this way, the value of each bit in the standardized anomaly labeling matrix obtained is the missing probability of the corresponding semantic keyword, which can more intuitively reflect the severity of the missingness of each semantic keyword.

[0092] This series of steps has brought about significant beneficial effects in many aspects. From the perspective of data quality assessment, through information entropy calculation and abnormal marking matrix generation, it is possible to accurately identify the unstable and easily missing parts of the semantic information of the waybill, help express delivery companies to timely discover data quality problems, and provide direction for subsequent data completion and correction. In terms of improving the accuracy of express delivery business processing, the standardized abnormal marking matrix provides a clear data basis for subsequent hierarchical fusion processing. For example, when automatically sorting parcels, if the probability of missing "destination" in the semantic keywords of the waybill of a parcel is high, the system can automatically adjust the sorting strategy, give priority to manual verification of the parcel or take other auxiliary measures, reduce sorting errors caused by missing information, and improve the accuracy of express delivery. From the perspective of intelligent processing, these data processing steps provide basic data support for realizing the intelligence of express delivery business. Through the analysis of the abnormal marking matrix, the system can automatically learn the missing patterns of waybill information of different types of parcels, thereby optimizing the information processing process and improving the overall business processing efficiency. At the same time, this quantitative analysis method also helps express delivery companies to manage and monitor data, discover abnormal data fluctuations in a timely manner, take corresponding improvement measures, and improve the company's operational management level. In addition, the standardized anomaly labeling matrix also plays an important role in cross-modal information fusion. It can be used as an intermediate data structure to provide a unified data format for subsequent multimodal analysis based on package morphological features, facilitating more efficient information fusion and processing, thereby further improving the accuracy and efficiency of express delivery business processing.

[0093] Step S1230, establish a many-to-many mapping relationship between the package priority and the abnormal marking matrix, and generate an association matrix reflecting the hierarchical fusion result of the package multimodal features.

[0094] Specifically, the core task of step S1230 is to establish a many-to-many mapping relationship between package priority and exception marking matrix, and generate an association matrix, so as to realize the hierarchical fusion of multimodal features of packages and provide comprehensive and targeted data support for subsequent business processing.

[0095] The parcel priority is divided into three levels: A, B, and C, based on the analysis results of the parcel morphology geometric topology matrix, which reflects the reliability of the parcel morphology features; the abnormal marking matrix is ​​obtained after the missing value analysis of the face sheet semantic keyword vector, which quantifies the missing probability of each semantic keyword. To establish a many-to-many mapping relationship between the two, it is necessary to deeply explore the potential connection between the parcel morphology and semantic information. For example, in the statistical analysis of a large amount of parcel data, it is found that the missing probability of the semantic keyword "recipient name" in A-level parcels is generally low, which indicates that not only the morphological features of A-level parcels are reliable, but also the integrity of the face sheet semantic information is relatively high; while the missing probability of this semantic keyword in C-level parcels is relatively high, which means that C-level parcels may have certain problems in both morphological features and semantic information integrity. By analyzing many such semantic keywords and associating each parcel priority category with the missing probability of each semantic keyword, a many-to-many mapping relationship can be constructed.

[0096] The association matrix generated based on this mapping relationship is a data structure that can comprehensively reflect the hierarchical fusion results of the multimodal features of the package. Each row of the association matrix represents a package priority category, and each column represents the missing probability dimension of a semantic keyword. The elements in the matrix represent the degree of association between the corresponding package priority and the missing probability of the semantic keyword. Its value can be determined according to the specific association algorithm. For example, it can be a numerical value representing the strength of association, ranging from 0 to 1. The larger the value, the higher the degree of association. Assuming that it is calculated by a certain algorithm that the strength of association between Class A packages and the lower missing probability of "recipient name" is 0.9, this indicates that there is a strong positive correlation between the two, that is, the low missing probability of "recipient name" in Class A packages is more common.

[0097] At the data fusion level, this association matrix organically integrates the morphological features (reflected by the package priority) and semantic features (reflected by the abnormal marking matrix) of the package, breaking the barriers between different modal data, so that subsequent processing can simultaneously consider multiple feature information of the package, and improve the efficiency of data utilization. For example, when planning express delivery routes, the package priority and the missing semantic keywords can be combined to give priority to the transportation of packages with high priority and high information completeness to improve the delivery timeliness; for packages with low priority and more missing information, the manual intervention process is planned in advance to ensure that the package can be delivered accurately. From the perspective of algorithm optimization, the association matrix provides a unified format and rich data foundation for subsequent algorithm processing. When performing data analysis and model training, the association matrix can be directly used as input to avoid the processing difficulties caused by inconsistent formats of different modal data. At the same time, the multimodal fusion information in the matrix helps to improve the accuracy and adaptability of the algorithm. In terms of express delivery business decision support, the association matrix can provide managers with more comprehensive and accurate data basis. By analyzing the association matrix, managers can understand the semantic information missing of packages of different priorities and formulate targeted business strategies, such as optimizing the design of the waybill and adjusting the information collection process, so as to improve the service quality and operational efficiency of the entire express delivery business. For example, if it is found that the probability of missing the semantic keyword "destination" for packages in a certain area is high, and these packages are mostly low-priority, it is possible to consider strengthening the information collection standard training in the area, or optimizing the design of the "destination" filling area on the waybill to reduce the situation of missing information. In addition, the association matrix can also be used as an intermediate data structure to provide strong support for subsequent knowledge graph construction, knowledge base retrieval and other links, making the entire express delivery business problem solving process more coherent and efficient. In the construction of the knowledge graph, the association matrix can provide data support for the establishment of relationships between nodes, so that the knowledge graph can more accurately reflect the internal connection between package information; when searching the knowledge base, the association matrix can help quickly filter out knowledge information related to a specific package, improving retrieval efficiency and accuracy.

[0098] Step S1300, dynamically generate a graph node state matrix based on the association matrix.

[0099] Furthermore, if Figure 4 As shown, step S1300 includes:

[0100] Step S1310, deconstructing the association matrix into two sub-matrices: a priority vector and an abnormality mark vector;

[0101] Step S1320, using the priority vector as a skeleton and the abnormal mark vector as a decoration, a cross-modal semantic association graph with a dynamic structure is generated through a random walk algorithm; each node in the cross-modal semantic association graph represents a package priority category, and the edges between nodes represent the transition probability between package priorities;

[0102] Step S1330, perform embedded learning on the cross-modal semantic association graph to obtain a graph node state matrix.

[0103] Specifically, sub-step S1310 aims to deconstruct the association matrix into two sub-matrices, namely, a priority vector and an abnormal marking vector. The association matrix is ​​generated in step S1230, and it comprehensively reflects the result of the hierarchical fusion of the multimodal features of the package. The deconstruction operation is to decompose this complex matrix into two relatively simple but functionally clear sub-matrices. The priority vector focuses on the priority information of the package, which reflects the reliability of the package morphological characteristics; the abnormal marking vector focuses on the missing of the semantic keywords of the face sheet, and records the missing probability information of each semantic keyword. The purpose of this deconstruction is to separate the information of different properties in the association matrix, so as to carry out more targeted processing from the two perspectives of package priority and semantic information missing. For example, when processing a batch of package data, the priority vector obtained by deconstructing the association matrix can clearly show the distribution of packages of different priorities, while the abnormal marking vector can clearly point out the missing points of the semantic information of each package face sheet, providing a clear data basis for subsequent operations. Its beneficial effect is to simplify the data structure and reduce the complexity of subsequent processing. By separating the priority and exception marking information, it is possible to use this information more efficiently in subsequent operations such as generating cross-modal semantic association graphs, thereby improving the accuracy and efficiency of data processing. At the same time, this separation also helps to analyze the multimodal characteristics of packages more deeply, providing a more detailed basis for the optimization of express delivery services.

[0104] Sub-step S1320 uses the priority vector as the skeleton and the abnormal marking vector as the decoration to generate a cross-modal semantic association graph with a dynamic structure through a random walk algorithm. The priority vector provides a macroscopic framework for the entire association graph, in which each element corresponds to a package priority category and determines the basic composition of the nodes in the graph; the abnormal marking vector gives each node and edge more detailed information, affecting the connection and relationship between nodes. The random walk algorithm is an algorithm that simulates a random process. In this step, it generates the edges between nodes and the length of the edges according to the similarity of the distribution of semantic features between package priorities. Specifically, the edge represents the transition probability between package priorities, and the length of the edge is proportional to the difficulty of calling semantic features across priorities to fill in missing morphological features. For example, assuming that in actual business, A-level packages and B-level packages have a high similarity in semantic features, and it is relatively easy to call semantic features to fill in morphological features from A-level packages to B-level packages, then the edge between them will be shorter and the transition probability will be relatively higher; while the semantic features of A-level packages and C-level packages are different, the transition difficulty is high, the edge will be longer, and the transition probability is lower. The cross-modal semantic association graph generated in this way is a directed acyclic graph, which can intuitively show the relationship between different package priorities and the difficulty of calling semantic features. The purpose of this step is to construct a network structure that can reflect the relationship between the multimodal features of the package, and provide an intuitive model for subsequent analysis. Its beneficial effects are reflected in many aspects. First, this graph structure can clearly present the complex relationship between package priorities and semantic features, helping express delivery companies to better understand package data. For example, when planning express delivery routes, this graph structure can quickly determine which packages have stronger information correlation, so as to arrange the delivery order more reasonably and improve delivery efficiency. Secondly, it provides an effective way for cross-modal information fusion. By connecting and setting the length of the edge, it is possible to obtain information from the semantic features of packages of different priorities to supplement the morphological features, improve the utilization efficiency of information, and reduce package processing errors caused by missing information. Finally, this dynamic structure graph can adapt to the changes in package data in the express business scenario. With the addition of new package data, the association graph can be updated accordingly to maintain an accurate reflection of the business situation.

[0105] Sub-step S1330 performs embedded learning on the cross-modal semantic association graph to obtain a graph node state matrix. Embedded learning is a technology that maps complex graph structure data to a low-dimensional vector space, which can retain important feature information of nodes and edges in the graph. In this step, through embedded learning, the rich morphological semantic information and priority-related information in the cross-modal semantic association graph are refined to generate a graph node state matrix. This matrix, as the input of the downstream calibration link, covers the full process information of grading, labeling, and fusion, and can provide a multi-granular reference for the calibration link. For example, when classifying packages or handling abnormal packages, the graph node state matrix can provide information from package priority to missing semantic information, helping the system make more accurate decisions. The purpose of this step is to extract valuable information from complex graph structures and convert it into a matrix form that is convenient for subsequent processing. Its beneficial effect is that, on the one hand, embedded learning can effectively reduce data dimensions and reduce the complexity of data processing, while retaining key information and improving data processing efficiency. On the other hand, as a unified information carrier, the graph node state matrix integrates various information about the package in the previous steps, providing comprehensive and orderly data support for subsequent cross-modal mapping, knowledge base retrieval and other operations, and enhancing the coherence and accuracy of the entire express business processing process. In addition, since the matrix is ​​related to priority, it can better make decisions based on the importance of the package. For example, for high-priority packages, more detailed and accurate information in the matrix can be used for rapid processing during processing, improving service quality.

[0106] In summary, step S1300 and its sub-steps achieve effective mapping and fusion of multi-granularity features in the knowledge graph space by deconstructing the association matrix, generating a cross-modal semantic association graph, and performing embedded learning, which provides key data support and analysis basis for subsequent express business processing, helps to improve the processing efficiency, accuracy and intelligence level of express business, and enhance the service quality and competitiveness of express companies.

[0107] Step S2000, obtaining multi-level calibration parameters according to the graph node state matrix; constructing a cross-modal mapping function cluster based on the multi-level calibration parameters to obtain a package feature matrix; designing a knowledge base retrieval and sorting mechanism for the package feature matrix to generate a knowledge recommendation list;

[0108] Furthermore, step S2000 includes:

[0109] Step S2100, obtaining multi-level calibration parameters according to the graph node state matrix;

[0110] Furthermore, if Figure 5 As shown, step S2100 includes:

[0111] Step S2110, taking the graph node state matrix as input, extracting the hidden layer feature vector by stacking the denoising autoencoder;

[0112] Step S2120, performing soft threshold transformation on the hidden layer feature vector, taking the amplitude of the hidden layer feature vector as the significance score, and generating a weight coefficient matrix of the morphological feature attenuation factor and a weight coefficient matrix of the semantic compensation coefficient;

[0113] Step S2130, normalizing the weight coefficient matrix of the morphological feature attenuation factor and the weight coefficient matrix of the semantic compensation coefficient to obtain a normalized morphological feature attenuation factor weight vector and a normalized semantic compensation coefficient weight vector;

[0114] Step S2140, performing a dot product operation on the normalized morphological feature attenuation factor weight vector and the hidden layer feature vector to obtain the morphological feature attenuation factor; performing a dot product operation on the normalized semantic compensation coefficient weight vector and the hidden layer feature vector to obtain the semantic feature compensation coefficient;

[0115] Step S2150, a multi-level calibration parameter is formed by the morphological feature attenuation factor and the semantic feature compensation coefficient.

[0116] Specifically, step S2100 mainly obtains multi-level calibration parameters based on the graph node state matrix, and extracts key information from the high-dimensional graph node matrix through a series of complex data processing and algorithm operations, laying the foundation for the subsequent construction of cross-modal mapping function clusters and optimizing package feature representation. Sub-step S2110 takes the graph node state matrix as input, extracts hidden layer feature vectors through stacked denoising autoencoders, and realizes dimensionality reduction compression of the state matrix. Stacked denoising autoencoders are a deep learning model, which is stacked by multiple denoising autoencoders. The denoising autoencoder adds noise to the input data during training, and then learns to remove the noise and reconstruct the original data, so that the model can learn a more robust feature representation of the data. In this step, the graph node state matrix contains a large amount of package multimodal feature information, with a high dimension, and direct processing will face problems such as high computational complexity and information redundancy. By stacking denoising autoencoders, key features can be automatically extracted, and the high-dimensional graph node state matrix can be converted into a low-dimensional hidden layer feature vector to achieve dimensionality reduction compression. For example, assuming that the state matrix of the graph node has hundreds of dimensions, after being processed by the stacked denoising autoencoder, the dimension of the hidden feature vector can be reduced to dozens, greatly reducing the data dimension while retaining key information. The purpose of this step is to simplify the data structure, extract key features, and reduce the complexity of subsequent processing. The beneficial effect is that the hidden feature vector after dimensionality reduction and compression not only reduces the overhead of data storage and calculation, but also highlights the key information in the multimodal features of the package, avoiding interference with subsequent analysis due to excessive redundant information. At the same time, this feature extraction method based on deep learning can automatically learn the potential patterns in the data, which is more adaptable and accurate than the traditional manual feature extraction method, and provides strong support for more accurate analysis of package features in the future.

[0117] Sub-step S2120 performs a soft thresholding transformation on the hidden layer feature vector, and uses the amplitude of the hidden layer feature vector as the significance score to generate a weight coefficient matrix of the morphological feature attenuation factor and a weight coefficient matrix of the semantic compensation coefficient. Soft thresholding is a signal processing technique that can suppress noise in the signal while retaining the important features of the signal. In this step, the amplitude of the hidden layer feature vector is used as a score to measure the significance of the feature. Through soft thresholding, different weights are assigned according to different amplitudes, thereby generating two weight coefficient matrices. For the weight coefficient matrix of the morphological feature attenuation factor, the matrix element (i 1 ,j 1 ) represents the i-th 1 The jth priority of the package 1 The contribution weight of the hidden layer features to the morphological feature attenuation factor; for the weight coefficient matrix of the semantic compensation coefficient, the matrix element (i 2 ,j 2 ) represents the i-th 2The jth priority of the package 2 The contribution weight of each hidden feature to the semantic compensation coefficient. For example, if the amplitude of a hidden feature vector is large, it means that the feature has a greater impact on a certain characteristic of the package. When generating the weight coefficient matrix, the element value at the corresponding position will be relatively large. The purpose of this step is to provide a weight basis for the subsequent morphological feature attenuation and semantic feature compensation according to the importance of the hidden features. Its beneficial effect is that in this way, the degree of influence of different hidden features on the morphological and semantic feature processing can be dynamically adjusted. In the process of package processing, the morphological and semantic information integrity of different packages is different. The use of these two weight coefficient matrices can be targeted according to the actual situation. Different features can be processed in a targeted manner to improve the accuracy and flexibility of the processing. For example, for packages with incomplete morphological features, the weight of the morphological feature attenuation factor can be adjusted to more reasonably use other information for compensation; for packages with more missing semantic information, the weight of the semantic compensation coefficient can be adjusted to better mine the potential semantic information in the morphological features.

[0118] Sub-step S2130 normalizes the weight coefficient matrix of the morphological feature attenuation factor and the weight coefficient matrix of the semantic compensation coefficient to obtain the normalized morphological feature attenuation factor weight vector and the normalized semantic compensation coefficient weight vector. Row normalization is a data standardization method that normalizes each row element of the matrix so that the sum of each row element is 1. Through this treatment, the elements in the weight coefficient matrix have probabilistic meaning, which is convenient for subsequent calculation and analysis. For example, in a weight coefficient matrix, after the elements of a row are normalized, the value of each element is between 0 and 1, and their sum is 1, which indicates the relative importance ratio of the hidden layer features corresponding to the row in the corresponding feature processing. The purpose of this step is to make the weight coefficients comparable and interpretable, which is convenient for subsequent calculation and analysis. Its beneficial effect is that the normalized weight vector can more intuitively reflect the relative importance of different hidden layer features in the overall feature processing, avoiding calculation errors and analysis deviations caused by differences in the magnitude of weight values. When subsequently calculating the morphological feature attenuation factor and the semantic feature compensation coefficient, the normalized weight vector can more accurately measure the contribution of each hidden layer feature and improve the reliability and stability of the calculation results.

[0119] Sub-step S2140 performs a dot product operation on the normalized morphological feature attenuation factor weight vector and the hidden layer feature vector to obtain the morphological feature attenuation factor; and performs a dot product operation on the normalized semantic compensation coefficient weight vector and the hidden layer feature vector to obtain the semantic feature compensation coefficient. The dot product operation is a vector operation. By multiplying and summing the corresponding elements of two vectors, a scalar value can be obtained. In this step, the weight vector is combined with the hidden layer feature vector by the dot product operation to obtain the compensation and attenuation factors related to the package morphology and semantic features. The purpose of this step is to calculate the key parameters for adjusting the package morphology and semantic features based on the weight vector and hidden layer feature vector generated previously. The beneficial effect is that through this calculation method, the importance of the hidden layer features can be combined with the actual features of the package to generate compensation and attenuation factors with practical significance. These factors can be used to adaptively adjust the morphological and semantic features according to the specific situation of the package, thereby improving the accuracy and effectiveness of the package feature representation. For example, when the morphological features of a package are partially missing, the morphological feature attenuation factor can reasonably adjust the weight of the morphological features according to the importance of the hidden layer features, so that the actual situation of the package can be more accurately reflected in subsequent processing.

[0120] Sub-step S2150 is composed of a morphological feature attenuation factor and a semantic feature compensation coefficient, which constitute a multi-level calibration parameter. These multi-level calibration parameters integrate the morphological and semantic feature information of the package, as well as their relative importance and completeness, and provide key parameter support for the subsequent construction of a cross-modal mapping function cluster. For example, when processing a package with partially missing information on the face sheet but relatively complete morphological features, the semantic compensation coefficient will be relatively high, and the morphological feature attenuation factor will be relatively low. The multi-level calibration parameters formed by the combination of the two can guide the subsequent cross-modal mapping function cluster to be more inclined to mine semantic information from morphological features and achieve accurate representation of package features. The purpose of this step is to integrate the key parameters calculated previously to form a unified calibration parameter set to provide support for subsequent cross-modal mapping and package feature optimization. Its beneficial effect is that the multi-level calibration parameters organically integrate the multi-modal feature information of the package, so that subsequent processing can more comprehensively consider the various characteristics of the package. When constructing a cross-modal mapping function cluster, these parameters can adaptively adjust the parameters of the mapping function according to the specific circumstances of the package, improve the accuracy and adaptability of the cross-modal mapping, and thus obtain the package feature matrix more accurately, providing more reliable data support for subsequent knowledge base retrieval and problem solving.

[0121] In summary, step S2100 and its sub-steps obtain multi-level calibration parameters from the graph node state matrix through a series of data processing and algorithm operations. These parameters can be adaptively adjusted according to the multimodal characteristics of the package, providing key support for subsequent cross-modal mapping and package feature optimization, helping to improve the accuracy and intelligent processing level of package feature analysis in express delivery services, and improving the overall efficiency and quality of express delivery business processing.

[0122] Step S2200, constructing a cross-modal mapping function cluster based on multi-level calibration parameters to obtain a parcel feature matrix;

[0123] Furthermore, if Figure 6 As shown, step S2200 includes:

[0124] Step S2210, using the package morphological feature set as the morphological feature matrix A and the sheet semantic feature set as the semantic feature matrix B, generating a semantic association matrix M between the morphological feature matrix A and the semantic feature matrix B;

[0125] Further, step S2210 includes:

[0126] Step S2211, performing topic clustering on the semantic feature matrix B to obtain m semantic topics;

[0127] Step S2212, calculate the association strength of each package on the m semantic topics, and generate a semantic association matrix M; the semantic association matrix M is an m'×m matrix, where m' represents the number of packages.

[0128] Specifically, in sub-step S2211, the semantic feature matrix B is clustered by topic to obtain m semantic topics. Here, the non-negative matrix factorization (NMF) and other methods are used to decompose the matrix B into the product of two non-negative matrices, that is, B≈W*H. Among them, W is an m'×m matrix, m' represents the number of packages, m is the preset number of topics, and each column of the matrix W represents a semantic topic; H is an m×k' matrix, k' is the dimension of the semantic features of the delivery order, and each row of the matrix H represents the representation of a semantic topic in the original feature space. The purpose of topic clustering is to summarize and refine the complex and diverse semantic information in the delivery order semantic feature matrix B to find representative semantic topics. For example, when processing a large amount of express parcel delivery order data, there may be many semantic keywords such as addresses and recipient information. Through the NMF method, these information can be clustered into a limited number of semantic topics such as "same-city express delivery", "different-city express delivery", and "business express delivery", making subsequent analysis more focused and efficient. From the perspective of data dimensionality reduction, NMF decomposes the high-dimensional semantic feature matrix B into two relatively low-dimensional matrices W and H, effectively reducing the data dimension and the complexity of data processing. At the same time, this decomposition can retain the key semantic information in the original data, laying the foundation for the subsequent accurate analysis of the semantic features of the package. In terms of improving data comprehensibility, the semantic topics obtained by topic clustering have clear semantic meanings, which facilitates express delivery companies to understand the distribution of semantic information of the waybill from a macro level and helps to formulate more reasonable business strategies. For example, if it is found that the number of packages with the theme of "same-city express delivery" has increased significantly in a certain period of time, the company can adjust the same-city distribution resource allocation in a targeted manner. In terms of improving the level of intelligence of express delivery business, these semantic topics, as intermediate results, provide convenience for the subsequent establishment of semantic association matrix and the realization of cross-modal fusion, which helps to achieve more accurate package classification, route planning and other functions, thereby improving the operational efficiency and service quality of the entire express delivery business.

[0129] Sub-step S2212 calculates the association strength of each package on the m semantic topics to generate a semantic association matrix M. The semantic association matrix M is an m'×m matrix, where the element M(i 3 ,j 3 ) represents the i-th 3 Packages and j 3 The correlation strength between semantic topics ranges from [0,1], and the larger the value, the stronger the correlation. In the specific calculation, using the matrix W and the morphological feature matrix A, let the morphological feature vector of the i*th package be A(i*,:), then its correlation strength on the j*th topic is M(i*,j*)= (A(i*,:),W(:,j*)), where Represents cosine similarity, which is used to measure the similarity between two vectors. For example, suppose there is a morphological feature vector A(1,:), and the cosine similarity is calculated with the vector W(:,1) representing the semantic theme of "same-city express delivery". If the result is close to 1, it means that the package is highly associated with the theme of "same-city express delivery" and may be a package delivered in the same city. The purpose of this step is to establish a quantitative association between the morphological features of the package and the semantic theme of the delivery bill, providing a key connection bridge for subsequent cross-modal fusion. Its beneficial effects are significant. First, by calculating the association strength, the morphological features of the package can be closely combined with the semantic theme, breaking the barriers between the modalities and realizing the preliminary fusion of multimodal information. This helps to understand the attributes of the package more comprehensively. For example, for a package with a regular shape and a high degree of association with the theme of "business express delivery", it can be inferred that it may be a business document express delivery, thereby providing a more targeted basis for the express delivery process. Secondly, the semantic association matrix M provides an important data basis for the subsequent matrix decomposition and cross-modal mapping function cluster construction, so that cross-modal fusion can more accurately reflect the actual characteristics of the package and improve the accuracy and reliability of the package feature representation. In actual applications of express delivery business, this association matrix can be used to optimize the parcel sorting process. According to the strength of association between the parcel and the semantic topic, the parcel can be automatically assigned to the corresponding processing channel, thereby improving sorting efficiency and accuracy and reducing manual intervention and error rate.

[0130] Step S2220, performing ternary decomposition of the semantic association matrix M by a matrix decomposition method to generate a semantic topic matrix T, a morphological semantic matching matrix P, and a semantic morphological matching matrix Q;

[0131] Step S2230, performing a product operation on the morphological feature matrix A and the semantic morphological matching matrix Q to obtain a first matrix; performing a product operation on the semantic feature matrix B and the morphological semantic matching matrix P to obtain a second matrix; the first matrix and the second matrix have the same dimension;

[0132] Step S2240, using the morphological feature attenuation factor as the weight of the first matrix, using the semantic feature compensation coefficient as the weight of the second matrix, fusing the first matrix and the second matrix to establish a cross-modal mapping function cluster; according to the cross-modal mapping function cluster, obtaining the package feature matrix.

[0133] Specifically, step S2220 performs ternary decomposition of the semantic association matrix M by a matrix decomposition method to generate a semantic topic matrix T, a morphological semantic matching matrix P, and a semantic morphological matching matrix Q. Matrix decomposition is a technique for decomposing a complex matrix into multiple matrices with specific meanings. In this step, by decomposing the semantic association matrix M, a deeper relationship between the package morphological features and the semantic features can be excavated. For example, the semantic topic matrix T can further refine the core features of different semantic themes, which can more clearly reflect the essential differences between different semantic themes while reducing the dimension; the morphological semantic matching matrix P and the semantic morphological matching matrix Q respectively establish the matching relationship between morphological features and semantic features from different angles. The purpose of this step is to deeply analyze the semantic association matrix M, obtain a matrix that can better reflect the relationship between the multimodal features of the package, and provide a more accurate tool for subsequent cross-modal fusion. Its beneficial effects are reflected in multiple levels. From the perspective of knowledge extraction, the semantic topic matrix T obtained by ternary decomposition realizes the secondary extraction of semantic information, further simplifies and abstracts the complex semantic associations, highlights the key semantic features, and helps express delivery companies to have a deeper understanding of the internal structure of package semantic information, thereby optimizing business processes. In terms of cross-modal fusion, the morphological semantic matching matrix P and the semantic morphological matching matrix Q provide an effective way to accurately fuse morphological features and semantic features in the future, and can more accurately capture the associations between modalities, so that the final package feature matrix can better reflect the true characteristics of the package. For example, when analyzing a package with a special shape and ambiguous semantic information, these two matching matrices can help the system match relevant semantic information more accurately, improve the accuracy of package feature representation, and thus improve the accuracy and efficiency of express delivery business processing. In addition, this matrix decomposition method also provides a more explanatory data structure for subsequent data analysis and model training, which is convenient for researchers and engineers to further optimize algorithms and models related to express delivery business.

[0134] Step S2230 performs a product operation on the morphological feature matrix A and the semantic morphological matching matrix Q to obtain a first matrix; and performs a product operation on the semantic feature matrix B and the morphological semantic matching matrix P to obtain a second matrix. In this step, the purpose of the product operation is to further fuse the morphological features and semantic features of the package based on the matching matrix obtained previously. Taking the product of the morphological feature matrix A and the semantic morphological matching matrix Q as an example, through this operation, the morphological features can be recombined and transformed according to the rules defined by the semantic morphological matching matrix Q, so that they can be better integrated with the semantic features. For example, assuming that the morphological feature matrix A contains information such as the size and shape of the package, a row in the semantic morphological matching matrix Q represents the weight distribution of morphological features under a specific semantic theme. Through the product operation, the morphological feature representation after weighted combination under the semantic theme can be obtained, that is, the first matrix. Similarly, the second matrix obtained by the product of the semantic feature matrix B and the morphological semantic matching matrix P also adjusts and fuses the semantic features according to the rules of the morphological semantic matching matrix P. The beneficial effect of this step is that the initial fusion of the package morphological features and semantic features is achieved through the product operation, laying the foundation for subsequent deeper cross-modal fusion. On the one hand, this fusion method can make full use of the inter-modal correlation information contained in the matching matrix, so that the fused matrix can better reflect the internal connection of the multimodal features of the package, and improve the accuracy of feature representation. For example, for a package whose appearance feature is displayed as a cuboid and whose semantic information is related to "electronic product transportation", the matrix obtained after the product operation can more accurately reflect the comprehensive characteristics of the package in the "electronic product transportation" scenario, which helps express delivery companies to more accurately judge the attributes and processing methods of the package. On the other hand, the first matrix and the second matrix obtained provide a suitable form for subsequent weight-based fusion, which is convenient for further processing and optimization according to different needs and scenarios, and enhances the flexibility and adaptability of the system to the package feature representation, so as to better meet the diverse processing needs in the express delivery business and improve the overall business processing capabilities.

[0135] Step S2240 uses the morphological feature attenuation factor as the weight of the first matrix, uses the semantic feature compensation coefficient as the weight of the second matrix, fuses the first matrix and the second matrix, establishes a cross-modal mapping function cluster, and obtains the package feature matrix according to the cross-modal mapping function cluster. The morphological feature attenuation factor and the semantic feature compensation coefficient are obtained through complex data processing in step S2100, and they are closely related to the morphological and semantic information integrity of the package. When the package's face sheet information is relatively complete, the semantic compensation coefficient is relatively low, and the morphological feature attenuation factor is relatively high. At this time, the fusion process focuses more on the morphological features; on the contrary, when the face sheet information is seriously missing, the semantic compensation coefficient increases, the morphological feature attenuation factor decreases, and the fusion process relies more on semantic features to supplement the morphological information. For example, assuming that there is a package whose address information on the face sheet is missing, but the morphological features are relatively obvious, the morphological feature attenuation factor will be relatively small, and the semantic feature compensation coefficient will be relatively large. When the first matrix and the second matrix are fused, the role of the semantic feature will be enhanced, and the mapping function will find more clues from the semantic information to complete the morphological contour, so that the obtained package feature matrix can more accurately reflect the actual situation of the package. The cross-modal mapping function cluster is a set of functions constructed based on these weights and matrices. They can adaptively adjust the sensitivity to different modal components according to the specific situation of the package, and realize the conversion from the original multimodal features to the unified package feature matrix. The purpose of the step is to establish a cross-modal mapping function cluster that can be adaptively adjusted by reasonably setting the weights and fusing the two matrices, so as to obtain a more accurate package feature matrix that can better reflect the true characteristics of the package. Its beneficial effects are significant. First, this weight-based fusion method can be dynamically adjusted according to the actual information of the package, making full use of the multimodal information of the package, and improving the accuracy and reliability of the package feature representation. In the express delivery business, accurate package feature representation helps to improve the accuracy of package classification, route planning and other links, reduce error processing caused by inaccurate information, and improve express delivery efficiency. Secondly, the establishment of the cross-modal mapping function cluster provides a general framework for the processing of package features, which has strong adaptability and extensibility. No matter how the morphology and semantic information of the package changes, the appropriate package feature matrix can be obtained by adjusting the weights and function clusters, which provides strong support for subsequent knowledge base retrieval and problem solving. In addition, this fusion method can also explore the deeper connection between package form and semantic features, help discover some potential package attributes and patterns, provide new ideas and methods for the optimization and innovation of express delivery business, and enhance the core competitiveness of express delivery companies.

[0136] Step S2300, designing a knowledge base search and sorting mechanism for the package feature matrix and generating a knowledge recommendation list;

[0137] Furthermore, if Figure 7 As shown, step S2300 includes:

[0138] Step S2310, dividing the package feature matrix into n2 hash buckets, each hash bucket corresponds to a knowledge subgraph of the knowledge base;

[0139] Further, step S2310 includes:

[0140] Step S2311, using a local sensitive hashing method, clustering packages according to the similarity of package feature vectors in the package feature matrix, and classifying packages with similarity higher than a preset first threshold into the same hash bucket;

[0141] Step S2312, traverse all knowledge points in the knowledge base, and divide the knowledge points into the most similar hash buckets according to the similarity between the feature vector of each knowledge point and the center of the hash bucket;

[0142] Step S2313, all knowledge points corresponding to each hash bucket are organized into a knowledge subgraph in the knowledge base.

[0143] Specifically, step S2310 is intended to process the package feature matrix, divide it into multiple hash buckets, and construct a corresponding knowledge subgraph so as to perform knowledge base retrieval more efficiently later. Sub-step S2311 adopts the locality sensitive hashing method to cluster the packages according to the similarity of the package feature vectors in the package feature matrix, and divide the packages with similarity higher than the preset first threshold into the same hash bucket. Locality-Sensitive Hashing (LSH) is an effective algorithm for similarity search in high-dimensional space. Its core idea is to map similar vectors to the same hash bucket, so that data similar to the original space also has a high probability of being mapped to the same position in the hash space, thereby accelerating the similarity search process. In this embodiment, the package feature vector contains multimodal information such as the morphological features and semantic features of the package. By calculating the similarity between these vectors, the similarity between the packages can be determined. For example, for two packages, one package is in the shape of a cuboid and the waybill information shows that it is an electronic product delivery, and the other package is also close to a cuboid in shape and also involves electronic product-related semantics, so their feature vectors may have a high similarity. The preset first threshold is an artificially set similarity standard used to determine which packages can be divided into the same hash bucket. The purpose of this step is to preliminarily cluster the packages and group similar packages together to facilitate subsequent rapid retrieval and processing. Its beneficial effect is that through the local sensitive hashing method, the efficiency of package clustering can be greatly improved, avoiding the high-complexity operation of comparing all packages two by two. When processing massive package data, this clustering method can quickly locate similar packages and reduce the amount of calculation for subsequent processing. At the same time, concentrating similar packages in the same hash bucket helps to improve the accuracy of retrieval, because similar packages may face similar processing problems. Grouping them into a group can more efficiently share processing experience and knowledge, and improve the overall efficiency of express business processing.

[0144] Sub-step S2312 traverses all knowledge points in the knowledge base, and divides the knowledge points into the most similar hash buckets according to the similarity between the feature vector of each knowledge point and the center of the hash bucket. The center of the hash bucket refers to a certain central representation of the feature vectors of all packages in the hash bucket, such as the mean vector of these vectors. The feature vector of the knowledge point is a digital representation of each knowledge point in the knowledge base, which contains various attributes and information involved in the knowledge point. By calculating the similarity between the feature vector of the knowledge point and the center of the hash bucket, the similarity between the knowledge point and the package in the hash bucket can be determined. For example, if a knowledge point is about shockproof measures during the distribution of electronic products, and most of the packages in a hash bucket are electronic product-related packages, then the similarity between this knowledge point and the center of the hash bucket may be high, and it will be divided into this hash bucket. The purpose of this step is to associate the knowledge points in the knowledge base with the package clustering results, so that each hash bucket contains knowledge points similar to the packages in the bucket, providing a basis for the subsequent rapid acquisition of relevant knowledge from the knowledge base. The beneficial effect is that, in this way, the knowledge in the knowledge base is effectively matched with the actual package. When a package in a hash bucket needs to be processed, the relevant knowledge points can be directly found from the corresponding hash bucket, which improves the pertinence and efficiency of knowledge retrieval. At the same time, this association also helps to explore the potential connection between packages and knowledge, and provides a basis for further optimizing express delivery business processes. For example, if it is found that a specific problem often occurs in a package in a hash bucket, and there is no good solution in the knowledge points associated with it, it can prompt the company to supplement and improve the relevant knowledge.

[0145] Sub-step S2313 forms a knowledge subgraph in a knowledge base with all the knowledge points corresponding to each hash bucket. The knowledge subgraph is a local subgraph divided from the entire knowledge base graph, and the features of the knowledge points inside the subgraph are highly similar to the features of the packages in the corresponding hash bucket. For example, in a hash bucket mainly composed of electronic product packages, the corresponding knowledge subgraph may contain relevant knowledge points such as electronic product packaging specifications, transportation precautions, and common fault handling. These knowledge points are interrelated to form a knowledge system for this type of package. The purpose of this step is to construct a knowledge subset with a clear structure and strong pertinence, which is convenient for rapid retrieval and application when processing packages later. Its beneficial effect is significant. The construction of the knowledge subgraph makes the structure of the knowledge base clearer and easier to manage and maintain. When processing a specific type of package, by directly accessing the corresponding knowledge subgraph, relevant knowledge can be quickly obtained, avoiding blind retrieval in a huge knowledge base, and greatly improving the speed and accuracy of knowledge retrieval. At the same time, the existence of the knowledge subgraph is also conducive to the updating and expansion of knowledge. When it is found that the knowledge in a certain knowledge subgraph is not enough to solve practical problems, the subgraph can be updated and improved in a targeted manner without affecting other parts of the entire knowledge base. In addition, the knowledge subgraph can also be analyzed and studied as an independent module, which helps to discover the intrinsic connections and laws between knowledge and provide support for the intelligent development of express delivery business.

[0146] To summarize, step S2310 and its sub-steps cluster the packages through the local sensitive hashing method, and associate the knowledge points in the knowledge base with them to construct a knowledge subgraph, which provides an efficient infrastructure for subsequent knowledge base retrieval and express business problem solving, and helps to improve the efficiency, accuracy and intelligence level of express business processing.

[0147] Step S2320, in each hash bucket, detecting the structural similarity between the query matrix and the knowledge subgraph, dynamically adding semantic association edges, and obtaining an updated knowledge subgraph; the query matrix is ​​obtained by wrapping the feature matrix;

[0148] Further, step S2320 includes:

[0149] Step S2321, extracting the package feature vector corresponding to each package in the package feature matrix to form a query matrix;

[0150] Step S2322, calculating the structural similarity between each package feature vector in the query matrix and each knowledge point feature vector in the knowledge subgraph, and if the similarity is higher than a preset second threshold, adding a semantic association edge between the package node in the query matrix and the knowledge point in the knowledge subgraph; the package node refers to the node corresponding to each package in the query matrix;

[0151] Step S2323, performing graph embedding representation on the knowledge subgraph after adding the semantic association edge.

[0152] Specifically, the main task of step S2320 is to detect the structural similarity between the query matrix and the knowledge subgraph in each hash bucket, and dynamically add semantically associated edges to obtain an updated knowledge subgraph. Substep S2321 extracts the parcel feature vector corresponding to each parcel in the parcel feature matrix to form a query matrix. The query matrix is ​​a matrix composed of the feature vectors corresponding to each parcel in the parcel feature matrix output by step S2240, representing a batch of parcel queries to be retrieved. In the express delivery business, the parcel feature matrix integrates the morphological and semantic features of the parcel, and the feature vector of each parcel contains detailed information about the parcel. For example, the feature vector of a parcel may contain morphological information such as its size and shape, as well as semantic information such as the address of the sender and the type of item. By extracting these feature vectors to form a query matrix, a query basis is provided for subsequent similarity retrieval in the knowledge subgraph. The purpose of this step is to organize the feature information of the parcel into a format suitable for retrieval so as to match it with the knowledge subgraph. Its beneficial effect is that the unified query matrix format makes the retrieval process more standardized and efficient. When performing knowledge retrieval, the query matrix can be directly operated, avoiding the tedious process of processing the feature vectors of a single package one by one, and improving the retrieval efficiency. At the same time, the construction of the query matrix also facilitates batch retrieval of multiple packages, and can obtain relevant knowledge recommendations for multiple packages at one time, further improving the efficiency of express delivery business processing.

[0153] Sub-step S2322 calculates the structural similarity between the feature vector of each package in the query matrix and the feature vector of each knowledge point in the knowledge subgraph. If the similarity is higher than the preset second threshold, a semantic association edge is added between the package node in the query matrix and the knowledge point in the knowledge subgraph. Structural similarity is used to measure the structural similarity between two vectors, which takes into account factors such as the element distribution and mutual relationship of the vectors. In this embodiment, by calculating this similarity, the potential connection between the package and the knowledge point can be determined. For example, if the feature vector of a package indicates that it is a large and fragile item, and there is a knowledge point about the transportation protection of large and fragile items in the knowledge subgraph, when their structural similarity is higher than the preset second threshold, it is considered that there is a strong association between the two. The preset second threshold is an artificially set standard for controlling the strictness of adding semantic association edges. The purpose of this step is to explore the semantic association between the package and the knowledge point, enrich the structure of the knowledge subgraph, and provide a more accurate basis for subsequent knowledge recommendation. Its beneficial effects are reflected in many aspects. First, the added semantic association edge can more accurately reflect the intrinsic connection between the package and the knowledge, making the knowledge subgraph closer to actual business needs. When processing a package, these semantic association edges can quickly find related knowledge points, improving the accuracy of knowledge recommendation. Secondly, this association mining based on structural similarity helps to discover some potential knowledge application scenarios, such as discovering new associations between certain package features and specific knowledge, thereby providing new ideas for optimizing express delivery services. Finally, rich semantic association edges are also conducive to improving the scalability of the knowledge subgraph. With the addition of new packages and new knowledge, new association edges can be continuously added based on structural similarity, allowing the knowledge subgraph to continuously adapt to business changes.

[0154] Sub-step S2323 performs graph embedding representation on the knowledge subgraph after adding the semantically associated edge. Graph embedding representation is a technique for converting graph structure data into low-dimensional vector representation, which can retain important feature information of nodes and edges in the graph. In this embodiment, through graph embedding representation, the updated knowledge subgraph can be converted into a vector form that is easier for the computer to process for subsequent calculation and analysis. For example, the knowledge subgraph containing package nodes and knowledge point nodes and the semantically associated edges between them is converted into a set of low-dimensional vectors, each vector representing the position of a node in the low-dimensional space. The purpose of this step is to reduce the dimension of complex graph structure data and extract its key features to facilitate subsequent efficient calculation and analysis. Its beneficial effect is that the graph embedding representation can greatly reduce the dimension of the data, reduce the amount of calculation, and improve the processing efficiency. When performing knowledge recommendation, calculation based on these low-dimensional vectors can obtain results faster, meeting the real-time requirements of the express delivery business. At the same time, the graph embedding representation can retain the structural information of the graph, so that the association between the package and the knowledge point can still be reflected in the low-dimensional space, ensuring the accuracy of the knowledge recommendation. In addition, this representation method also makes it easy to combine knowledge subgraphs with other machine learning algorithms to further mine the potential information in the knowledge subgraphs and provide stronger support for intelligent decision-making in express delivery services.

[0155] To summarize, step S2320 and its sub-steps enrich the semantic expression of the knowledge subgraph and improve the accuracy and efficiency of knowledge retrieval by constructing a query matrix, calculating structural similarity, adding semantic association edges, and performing graph embedding representation, laying a solid foundation for the subsequent generation of high-quality knowledge recommendation lists, and helping to improve the level of intelligence and service quality of express business processing.

[0156] Step S2330, performing random walks on the updated knowledge subgraph, calculating the importance score of each knowledge point, and generating a sorted knowledge recommendation list.

[0157] Further, step S2330 includes:

[0158] Step S2331, based on the graph embedding representation of the knowledge subgraph, calculating the semantic relevance between each package node in the query matrix and each knowledge point in the knowledge subgraph;

[0159] Step S2332, assigning corresponding priority attributes to knowledge points in the knowledge subgraph based on the package priority;

[0160] Step S2333, taking the semantic relevance of each node as the transition probability and the priority attribute of the node as the restart probability, perform a random walk on the knowledge subgraph, iteratively calculate the steady-state probability of each knowledge point as the importance score of each knowledge point, and generate a recommended ranked list of knowledge points from high to low according to the importance score.

[0161] Specifically, the core purpose of step S2330 is to perform a series of operations on the updated knowledge subgraph, calculate the importance score of each knowledge point, and then generate a sorted knowledge recommendation list to provide accurate knowledge support for solving express business problems. Sub-step S2331 calculates the semantic relevance of each package node in the query matrix and each knowledge point in the knowledge subgraph based on the graph embedding representation of the knowledge subgraph. Graph embedding representation is a technology that converts graph structure data into low-dimensional vector representation, which can retain important feature information of nodes and edges in the graph. In this embodiment, after the knowledge subgraph is processed by graph embedding, each knowledge point and package node are mapped to a low-dimensional vector space, so that their semantic relevance can be measured in this space by calculating the similarity between vectors. For example, if the feature vector of a package is close to the vector distance of a certain knowledge point in the low-dimensional space, it means that they have a high semantic relevance. This way of calculating semantic relevance aims to quantify the semantic connection between the package and the knowledge point, and provide a basis for the subsequent determination of the importance of the knowledge point. Its beneficial effect is that through graph embedding and semantic relevance calculation, the intrinsic semantic association between the package and the knowledge can be captured more accurately. When dealing with express delivery business, this helps to quickly screen out knowledge points related to specific packages and improve the accuracy and efficiency of knowledge retrieval. Compared with the traditional retrieval method based on simple keyword matching, this calculation method based on semantic relevance can take into account the semantic depth and relevance of knowledge and package features, avoiding the deviation of retrieval results caused by inaccurate keyword matching, thereby providing more valuable knowledge recommendations for express delivery business.

[0162] Sub-step S2332 assigns corresponding priority attributes to the knowledge points in the knowledge sub-graph based on the parcel priority. The parcel priority is obtained in the previous step based on the analysis of the parcel morphology geometric topology matrix, and is divided into three levels: A, B, and C, representing the urgency and importance of parcel processing. In this step, the priority attribute of the parcel is assigned to the knowledge points in the knowledge sub-graph in order to comprehensively consider the relevance of knowledge and the timeliness of parcel processing in the subsequent knowledge recommendation process. For example, for an A-level parcel, the knowledge points related to it will be given a higher priority weight in the subsequent sorting. The purpose of this is to make the knowledge recommendation results not only meet the knowledge needs of the parcel, but also take into account the timeliness requirements of the business. Its beneficial effects are reflected in many aspects. First, this method ensures that when processing urgent parcels, knowledge points related to urgent parcels can be recommended first, which improves the timeliness of express delivery business processing. For example, for some fresh express deliveries with high timeliness requirements (which may be classified as A-level parcels), knowledge points related to freshness preservation and fast delivery are given priority when recommending knowledge, which helps to ensure the quality of fresh products. Secondly, by combining the association between package priority and knowledge points, the overall strategy of knowledge recommendation is optimized, making the recommendation results more in line with the needs of actual business scenarios and improving the service quality and customer satisfaction of express delivery business.

[0163] Sub-step S2333 uses the semantic relevance of each node as the transition probability and the priority attribute of the node as the restart probability, performs random walks on the knowledge subgraph, iteratively calculates the steady-state probability of each knowledge point as the importance score of each knowledge point, and generates a recommended sorted list of knowledge points from high to low according to the importance score. Random walk is an algorithm for random movement on a graph structure. In this step, the process of finding relevant knowledge points in a knowledge subgraph is simulated by random walks. The transition probability determines the possibility of moving from one node to another. Here, the semantic relevance is used as the transition probability, which means that the higher the semantic relevance of the knowledge point to the current package, the greater the probability of being visited; and the restart probability is used to control the process of random walks, using the priority attribute of the node as the restart probability, so that during the walk, the knowledge points related to the high-priority package are more likely to be revisited and valued. For example, in a knowledge subgraph, there are multiple knowledge points related to the package. Through random walks, the knowledge points with high semantic relevance and corresponding to the high-priority package will have a higher probability of being visited multiple times, and their steady-state probability will be higher, and finally ranked higher in the recommendation list. The purpose of this step is to comprehensively consider the two key factors of semantic relevance and package priority, comprehensively evaluate and sort the knowledge points in the knowledge subgraph, and generate a high-quality knowledge recommendation list. Its beneficial effect is remarkable. On the one hand, this scoring strategy makes full use of the semantic association between the package and the knowledge point and the priority information of the package, so that the recommendation results are both accurate and timely. In the actual express delivery business, it can provide the most relevant and timely knowledge recommendations according to the specific situation of the package, helping staff to quickly find solutions to problems and improve work efficiency. On the other hand, the introduction of the random walk algorithm increases the flexibility and comprehensiveness of the recommendation process, avoids the problem of local optimal solutions, enables the recommendation results to cover a wider range of knowledge, and improves the reliability and practicality of knowledge recommendations. The knowledge recommendation list generated in this way can better meet the needs of express delivery business in complex and changing scenarios, and provide strong support for the efficient operation of express delivery business.

[0164] Step S3000, based on the knowledge recommendation list, generate a solution to the express business problem, and feed the newly generated solution back to the knowledge base to achieve self-update and iteration of the knowledge base.

[0165] Furthermore, step S3000 includes:

[0166] Step S3100, clustering the express business problems according to the knowledge recommendation list to obtain problem clusters;

[0167] Step S3200: for each problem cluster, call the relevant knowledge points in the knowledge recommendation list, perform case reasoning in combination with the graph node state matrix, and generate a candidate solution set;

[0168] Step S3300, conduct manual expert feedback learning on the candidate solution set, optimize and determine the final executable solution, and feed back the key steps and execution effects in the executable solution as new knowledge points to the knowledge base.

[0169] Specifically, step S3000 aims to generate solutions to express business problems based on the knowledge recommendation list, and feed back the newly generated solutions to the knowledge base to achieve self-update and iteration of the knowledge base, and realize the closed loop from knowledge recommendation to actual business application and knowledge accumulation through multi-stage operations. Sub-step S3100 clusters express business problems according to the knowledge recommendation list to obtain problem clusters. In this step, the problem template corresponding to each knowledge point in the knowledge recommendation list is first extracted. The problem template is an abstract representation of common express business problems, which contains the key features and structure of the problem. Then, the problems are clustered according to the template similarity to form several problem clusters. For example, there may be different ways of expressing problems involving package loss, but through the analysis of the problem template, these problems with similar expressions and the same essence can be classified into one problem cluster. Then, the number, time distribution and priority distribution of problems in each problem cluster are counted, and the problem clusters are sorted to obtain a sorted problem cluster sequence. Finally, according to the problem cluster sequence, the actual problems in the express business scenario are classified to obtain a set of express business problems corresponding to the problem clusters. The purpose of this step is to systematically sort out and classify a large number of express delivery business problems so that more targeted solutions can be found later. Its beneficial effect is that through cluster analysis, similar problems can be dealt with together to improve the efficiency of problem solving. For example, when it is found that there are a large number of problems in a certain problem cluster and the time distribution is concentrated, express delivery companies can formulate special solutions and response strategies for such problems, concentrate resources to deal with them, and avoid duplication of work. At the same time, the priority sorting of problem clusters helps companies to reasonably allocate resources, give priority to urgent and important problems, and ensure the normal operation of express delivery business. In addition, this classification method also provides a clear framework for subsequent case reasoning, so that relevant problem types and knowledge can be located more quickly when looking for solutions.

[0170] Sub-step S3200 calls the relevant knowledge points in the knowledge recommendation list for each problem cluster, performs case reasoning in combination with the graph node state matrix, and generates a set of candidate solutions. The specific operation is to retrieve the knowledge points with the highest matching degree with the attribute values ​​in the knowledge recommendation list according to the key attributes of the problems in the problem cluster as candidate knowledge points. For example, for a problem cluster about package damage compensation, the key attributes may include package type, damage degree, etc., and the knowledge points matching with them are searched in the knowledge recommendation list according to these attributes. Then, the cross-modal semantic association information contained in the graph node state matrix is ​​used to calculate the correlation between the candidate knowledge points and the problem cluster, and the Top-N most relevant knowledge points are selected. The graph node state matrix contains multimodal information such as the morphology and semantics of the package and the associations between them. Through this information, the relevance of the knowledge points and the problem cluster can be more accurately evaluated. Finally, the historical cases corresponding to the Top-N knowledge points are reasoned and migrated, and a set of candidate solutions are generated in combination with the characteristic attributes of each problem in the problem cluster. For example, a historical case similar to the current package damage problem is found, and the solution in the historical case is adjusted and optimized according to the specific characteristics of the current problem to form a candidate solution suitable for the current problem. The purpose of this step is to use the knowledge recommendation list and multimodal information to generate targeted candidate solutions, providing multiple possible ways to solve actual express delivery business problems. The beneficial effect is that through case reasoning, the company's accumulated knowledge and experience in the past are fully utilized, avoiding the inefficient way of thinking about solutions from scratch every time a problem is encountered. At the same time, combined with the multimodal information in the graph node state matrix, it can consider all aspects of the problem more comprehensively, improving the accuracy and feasibility of candidate solutions. For example, when dealing with a problem of damaged parcels with special shapes, the morphological information in the graph node state matrix can help better understand the characteristics of the parcel, so as to select more appropriate solutions from historical cases for adjustment and improve the success rate of problem solving.

[0171] Sub-step S3300 conducts artificial expert feedback learning on the candidate solution set, optimizes and determines the final executable solution, and feeds back the key steps and execution effects in the executable solution as new knowledge points to the knowledge base. Artificial expert feedback learning refers to inviting experts with rich experience in express delivery business to evaluate and guide the candidate solutions. The experts point out the advantages and disadvantages of the candidate solutions based on their professional knowledge and practical experience, and put forward improvement suggestions. According to the feedback of experts, the candidate solutions are optimized to determine the final executable solution. For example, experts may find that a candidate solution has some potential risks in actual operation, and adjust some steps in the solution to make it more feasible and safe. Then the key steps and execution effects in the executable solution are fed back to the knowledge base as new knowledge points, that is, these new knowledge and experience are added to the knowledge base for use in subsequent processing of similar problems. The purpose of this step is to ensure the reliability and effectiveness of the solution and continuously enrich and improve the knowledge base. Its beneficial effects are reflected in many aspects. First, the participation of artificial experts ensures the quality of the final solution. Their experience and professional knowledge can make up for the shortcomings of algorithms and models and improve the practicality and operability of the solution. Secondly, feeding new solutions back to the knowledge base realizes the accumulation and inheritance of knowledge, so that the knowledge base can be continuously updated and improved. As time goes by, the knowledge in the knowledge base will become richer and more comprehensive, providing stronger support for the development of express delivery business. For example, when encountering new express delivery business problems, the system can use the updated knowledge base to find solutions more quickly, improve the company's ability to deal with problems, and also help the company continuously optimize business processes and enhance overall competitiveness. In this way, a closed loop from knowledge recommendation, problem solving to knowledge updating is formed, which promotes the continuous development and optimization of express delivery business.

[0172] Example 2

[0173] This embodiment provides a solution to express delivery business problems based on the knowledge base management system on the basis of embodiment 1, including:

[0174] Step S1224, counting the number of abnormal markings for each bit in the abnormal marking matrix, calculating the missing probability of each semantic keyword, and generating a missing probability vector corresponding to the dimension of the semantic keyword vector of the face order;

[0175] The calculation of the missing probability of each semantic keyword includes:

[0176]

[0177] in:

[0178] : No. The missing probability of a semantic keyword.

[0179] : No. The number of abnormal tags in the semantic keyword dimension is calculated by The sum of the column elements is obtained.

[0180] : No. The number of abnormal tags in the semantic keyword dimension.

[0181] m' : The total number of packages, which can be determined when counting the anomaly marker matrix.

[0182] : The average number of abnormal markings for all semantic keyword dimensions.

[0183] : Smoothing factor is a preset small positive number (such as 0.01) used to avoid the situation where the denominator is 0 and ensure the stability of the formula.

[0184] : The dimension of the semantic keyword vector of the face order.

[0185] when When it increases, other conditions remain unchanged, the numerator Increase, This indicates that the more times the semantic keyword dimension has abnormal tags, the higher the probability of missing it.

[0186] when (the average number of abnormal markings of all semantic keyword dimensions) increases, if constant, The value of becomes relatively small, The value of will decrease, so that This means that when the number of overall abnormal markings increases, the relative proportion of abnormal markings in a single semantic keyword dimension decreases, and its missing probability also decreases accordingly, reflecting a comprehensive consideration of the overall situation. As the denominator increases, Increase, will decrease. However, due to Usually a small value, which is The impact is relatively small, and it mainly plays a role in stabilizing calculations.

[0187] This formula comprehensively considers the number of abnormal markings for each semantic keyword dimension and the distribution of the overall abnormal marking number. By calculating the average number of abnormal markings for all semantic keyword dimensions , and use This form associates the number of anomaly marks in each dimension with the overall average level, making the calculation of missing probability more reasonable. Smoothing factor The introduction of ensures that the formula can be calculated normally under any circumstances, enhances the stability and robustness of the formula, and avoids calculation errors caused by the denominator being 0. By accurately calculating the missing probability of each semantic keyword, the parts of the semantic information of the face order that are prone to missing can be more accurately located, providing strong support for subsequent information completion and business decision-making.

[0188] Example 3

[0189] This embodiment provides a solution to express delivery business problems based on the knowledge base management system on the basis of embodiment 1, including:

[0190] Step S2322, calculating the structural similarity between each package feature vector in the query matrix and each knowledge point feature vector in the knowledge subgraph, and if the similarity is higher than a preset second threshold, adding a semantic association edge between the package node in the query matrix and the knowledge point in the knowledge subgraph; the package node refers to the node corresponding to each package in the query matrix;

[0191] Assume that the query matrix is , among which The package feature vector is , the first The feature vector of a knowledge point is ,and and Both dimensional vector. The structural similarity formula is defined as:

[0192]

[0193] in:

[0194] : represents the first Packet feature vector and the feature vector of the g-th knowledge point in the knowledge subgraph The structural similarity of The closer the value is to 1, the higher the similarity is, the closer it is to -1, the greater the difference is, and 0 means there is no obvious similarity between the two.

[0195] : is a weight coefficient used to adjust the The importance of dimensional features in similarity calculation, and . It can be determined by analyzing the importance of different dimensional features in the express business data to the matching of packages and knowledge points. For example, for express business, dimensions such as package weight and volume may have a greater impact on the matching results and can be given a higher weight; while some relatively minor dimensions, such as package packaging color, can be given a lower weight. Machine learning algorithms, such as feature selection algorithms, can also be used to automatically determine the weight coefficient. Its significance lies in highlighting the impact of key features on similarity calculation, so that the calculation results are more in line with actual business needs.

[0196] : is the first Package feature vector No. Dimensional component.

[0197] : is the query matrix for all the wrapped feature vectors in the first The mean of the dimension, is the number of packets in the query matrix.

[0198] :It is the first Knowledge point feature vector No. Dimensional component.

[0199] : is the feature vector of all knowledge points in the knowledge subgraph in the first The mean of the dimension.

[0200] and Directly obtain the dimensional components of the corresponding vector from the query matrix and the knowledge subgraph respectively. When the covariance part of the numerator increases, other conditions remain unchanged. will increase. This means that the wrapping eigenvector and knowledge point feature vector The more similar the co-variance trends in each dimension are, the higher their structural similarity is. For example, if the weight and volume change trends of a package are consistent with the weight and volume change trends of the package type represented by a certain knowledge point, then the covariance of the two vectors in these dimensions is large, which will increase the overall similarity.

[0201] when As the numerator of the exponential increases, The value of will decrease, resulting in This indicates that the greater the absolute difference between the corresponding dimension elements of the package feature vector and the knowledge point feature vector, the lower the similarity. For example, if the destination of the package is very different from the destination in the package represented by the knowledge point, this difference will be reflected in the absolute difference of the corresponding dimension, thereby reducing the similarity between the two.

[0202] For the weight coefficient , if we increase the weight of a dimension that is important for matching (such as the package weight dimension) , when other conditions remain unchanged, when the package on this dimension is more similar to the characteristics of the knowledge point, it will make Increase; on the contrary, if the feature difference on this dimension is large, it will make This reflects the adjustment effect of the weight coefficient on the similarity calculation results, highlighting or weakening the impact of certain dimensions on similarity according to business needs.

[0203] This formula comprehensively considers the covariance and absolute difference of each dimension of the vector: the covariance part of the numerator in the formula measures the degree of coordinated change of the two vectors in each dimension, and the square root part of the denominator is used for normalization, which is similar to the traditional Pearson correlation coefficient and can reflect the degree of linear correlation between vectors. The exponential part considers the absolute difference of the elements of the corresponding dimensions of the two vectors. When the difference is larger, the value of the exponential part is smaller, and the overall similarity will also decrease, thereby more comprehensively measuring the similarity between vectors.

[0204] Weight adjustment mechanism: The introduction of can flexibly adjust the importance of each dimension feature according to business needs. In the express delivery business scenario, different features have different importance for judging the matching degree between packages and knowledge points. By setting weights reasonably, the similarity calculation can be more in line with the actual business and improve the accuracy of matching.

[0205] Adaptability to complex business scenarios: This formula is capable of processing high-dimensional vectors, and by comprehensively considering multiple factors, it can more accurately measure the structural similarity between the package feature vector and the knowledge point feature vector in complex express delivery business data, providing a more reliable basis for the subsequent addition of semantic association edges, and helping to build a more accurate knowledge graph and a more effective knowledge recommendation system.

[0206] Example 4

[0207] This embodiment provides an express business problem solving system based on the knowledge base management system on the basis of embodiment 1. Figure 8 As shown, including:

[0208] Package priority classification module: used to obtain the package geometry topology matrix to form a package feature set; perform structured analysis on the package geometry topology matrix to classify the packages into three package priorities;

[0209] Graph space fusion module: used to obtain the semantic keyword vector of the waybill to form the semantic feature set of the waybill; generate the abnormal marking matrix of the semantic keyword vector of the waybill; establish a many-to-many mapping relationship between the package priority and the abnormal marking matrix to generate the association matrix; dynamically generate the graph node state matrix according to the association matrix;

[0210] Knowledge recommendation module: According to the graph node state matrix, multi-level calibration parameters are obtained; based on the multi-level calibration parameters, a cross-modal mapping function cluster is constructed to obtain the package feature matrix; for the package feature matrix, a knowledge base retrieval and sorting mechanism is designed to generate a knowledge recommendation list;

[0211] Solution generation module: Generates solutions to express delivery business problems based on the knowledge recommendation list, and feeds the newly generated solutions back to the knowledge base to achieve self-update and iteration of the knowledge base.

[0212] In the package priority classification module, the acquisition of the package morphology geometric topology matrix and the formation of the package morphology feature set include: using multi-angle high-definition cameras deployed at key nodes of the sorting line to collect video image sequences of conveying packages in real time; performing frame-by-frame segmentation and feature extraction on the collected video image sequences, and using a lightweight CNN to generate a 128-dimensional package morphology geometric topology matrix to form a package morphology feature set.

[0213] In the package priority classification module, the classification of packages into three package priorities includes: calculating the variance of each dimensional feature for the package geometry topology matrix and the corresponding confidence level ;according to and , the packages are divided into three package priorities: A, B, and C; among them, is the variance of the j-th dimension feature in the 128-dimensional package morphology geometric topology matrix, for The corresponding confidence level;

[0214] Set the first variance threshold V 1 and the second variance threshold V 2 , where V 1 <V 2 ;like <V 1 , then give Weight W 1 ; If V 1 ≤ <V2 , then give Weight W 2 ;like ≥V 2 , then give Weight W 3; Where W 1 >W 2 >W 3 ; For each package’s 128-dimensional package geometry topology matrix, according to the confidence level and the weight W assigned 1 , W 2 , W 3 , calculate the weighted average to obtain the comprehensive confidence score CON of the package; set the first confidence threshold T 1 and the second confidence threshold T 2 , where T 1 >T 2 ; If CON>T 1 , the package is classified as Class A; if T 2 <CON≤T 1 , the package is classified as Class B; if CON≤T 2 , the package is classified as Class C.

[0215] In the graph space fusion module, generating an abnormal label matrix for a single semantic keyword vector includes:

[0216] Step S1221, for the 64-dimensional semantic keyword vector of the face sheet, count the occurrence frequency of the semantic keyword in each dimension and calculate the probability distribution of the occurrence frequency;

[0217] Step S1222, calculate the information entropy H of each dimension according to the probability distribution k ;H k represents the information entropy of the k-th dimension of the semantic keyword vector of the face order, where 1≤k≤64;

[0218] Step S1223, set the information entropy threshold E, when H k >E, the corresponding abnormal marking matrix position is marked as 0; when H k When ≤E, the corresponding abnormal marking matrix position is marked with 1;

[0219] Step S1224, counting the number of abnormal markings for each bit in the abnormal marking matrix, calculating the missing probability of each semantic keyword, and generating a missing probability vector corresponding to the dimension of the semantic keyword vector of the face order;

[0220] Step S1225 , normalize the missing probability vector to generate a standardized abnormal labeling matrix.

[0221] In the knowledge recommendation module, the multi-level calibration parameters are obtained according to the graph node state matrix, including:

[0222] Step S2110, taking the graph node state matrix as input, extracting the hidden layer feature vector by stacking the denoising autoencoder;

[0223] Step S2120, performing soft threshold transformation on the hidden layer feature vector, taking the amplitude of the hidden layer feature vector as the significance score, and generating a weight coefficient matrix of the morphological feature attenuation factor and a weight coefficient matrix of the semantic compensation coefficient;

[0224] Step S2130, normalizing the weight coefficient matrix of the morphological feature attenuation factor and the weight coefficient matrix of the semantic compensation coefficient to obtain a normalized morphological feature attenuation factor weight vector and a normalized semantic compensation coefficient weight vector;

[0225] Step S2140, performing a dot product operation on the normalized morphological feature attenuation factor weight vector and the hidden layer feature vector to obtain the morphological feature attenuation factor; performing a dot product operation on the normalized semantic compensation coefficient weight vector and the hidden layer feature vector to obtain the semantic feature compensation coefficient;

[0226] Step S2150, a multi-level calibration parameter is formed by the morphological feature attenuation factor and the semantic feature compensation coefficient.

[0227] In the knowledge recommendation module, generating a knowledge recommendation list includes:

[0228] Step S2310, dividing the package feature matrix into n2 hash buckets, each hash bucket corresponds to a knowledge subgraph of the knowledge base;

[0229] Step S2320, in each hash bucket, detecting the structural similarity between the query matrix and the knowledge subgraph, dynamically adding semantic association edges, and obtaining an updated knowledge subgraph; the query matrix is ​​obtained by wrapping the feature matrix;

[0230] Step S2330, performing random walks on the updated knowledge subgraph, calculating the importance score of each knowledge point, and generating a sorted knowledge recommendation list.

[0231] The method and system of the present application may be implemented in many ways. For example, the method and system of the present application may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of steps for the method is for illustration only, and the steps of the method of the present application are not limited to the order specifically described above, unless otherwise specifically stated.

[0232] In addition, the parts of the above-mentioned technical solutions provided in the embodiments of the present application that are consistent with the implementation principles of the corresponding technical solutions in the prior art are not described in detail to avoid excessive redundancy.

[0233] The specific implementation modes as described above further describe the purpose, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation mode of the present invention and is not intended to limit the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. A solution to express delivery business problems based on a knowledge base management system, characterized in that: The method comprises: Obtain the package geometry topology matrix to form a package geometry feature set; obtain the waybill semantic keyword vector to form a waybill semantic feature set; perform structured analysis on the package geometry topology matrix to divide the packages into three package priorities, and generate an abnormal marking matrix for the waybill semantic keyword vector; establish a many-to-many mapping relationship between the package priority and the abnormal marking matrix to generate an association matrix; dynamically generate a graph node state matrix based on the association matrix; According to the graph node state matrix, multi-level calibration parameters are obtained; based on the multi-level calibration parameters, a cross-modal mapping function cluster is constructed to obtain the package feature matrix; for the package feature matrix, a knowledge base retrieval and sorting mechanism is designed to generate a knowledge recommendation list; Generate solutions to express delivery business problems based on the knowledge recommendation list, and feed the newly generated solutions back to the knowledge base to achieve self-update and iteration of the knowledge base; The method of dynamically generating a graph node state matrix according to the association matrix includes: deconstructing the association matrix into two sub-matrices, a priority vector and an abnormal marking vector; using the priority vector as a skeleton and the abnormal marking vector as a modifier, generating a cross-modal semantic association graph with a dynamic structure through a random walk algorithm; each node in the cross-modal semantic association graph represents a package priority category, and the edges between nodes represent the transition probability between package priorities; performing embedded learning on the cross-modal semantic association graph to obtain a graph node state matrix; The method of obtaining multi-level calibration parameters according to the graph node state matrix includes: Taking the graph node state matrix as input, the hidden layer feature vector is extracted by stacking denoising autoencoders; the hidden layer feature vector is soft-thresholded, and the amplitude of the hidden layer feature vector is used as the significance score to generate the weight coefficient matrix of the morphological feature attenuation factor and the weight coefficient matrix of the semantic compensation coefficient; the weight coefficient matrix of the morphological feature attenuation factor and the weight coefficient matrix of the semantic compensation coefficient are normalized to obtain the normalized morphological feature attenuation factor weight vector and the normalized semantic compensation coefficient weight vector; Performing a dot product operation on the normalized morphological feature attenuation factor weight vector and the hidden layer feature vector to obtain the morphological feature attenuation factor; performing a dot product operation on the normalized semantic compensation coefficient weight vector and the hidden layer feature vector to obtain the semantic feature compensation coefficient; the morphological feature attenuation factor and the semantic feature compensation coefficient constitute a multi-level calibration parameter; The cross-modal mapping function cluster is constructed to obtain a package feature matrix including: The package morphological feature set is taken as the morphological feature matrix A, and the order semantic feature set is taken as the semantic feature matrix B, and a semantic association matrix M of the morphological feature matrix A and the semantic feature matrix B is generated; The semantic association matrix M is decomposed into three elements by matrix decomposition method to generate semantic topic matrix T, morphological semantic matching matrix P, and semantic morphological matching matrix Q. The morphological feature matrix A is multiplied by the semantic morphological matching matrix Q to obtain a first matrix; the semantic feature matrix B is multiplied by the morphological semantic matching matrix P to obtain a second matrix; the first matrix and the second matrix have the same dimension; The morphological feature attenuation factor is used as the weight of the first matrix, the semantic feature compensation coefficient is used as the weight of the second matrix, the first matrix and the second matrix are fused to establish a cross-modal mapping function cluster; according to the cross-modal mapping function cluster, the parcel feature matrix is ​​obtained.

2. The express delivery business problem solving method based on the knowledge base management system according to claim 1 is characterized in that: The obtaining of the package morphology geometric topology matrix to form a package morphology feature set comprises: The multi-angle high-definition cameras deployed at the key nodes of the sorting line collect the video image sequence of the package delivery in real time; the collected video image sequence is segmented and feature extracted frame by frame, and a lightweight CNN is used to generate a 128-dimensional package shape geometric topology matrix to form a package shape feature set; The method of obtaining the semantic keyword vector of the delivery note and forming the semantic feature set of the delivery note includes: performing OCR analysis on the package delivery note information through a high-speed scanner equipped with a sorting line, extracting key text information from the delivery note, and converting it into a 64-dimensional delivery note semantic keyword vector to form a delivery note semantic feature set.

3. The express delivery business problem solving method based on the knowledge base management system according to claim 2 is characterized in that: The structural analysis of the package geometry topology matrix is ​​performed to classify the packages into three package priorities, including: Calculate the variance of each dimension feature for the package geometry topology matrix and the corresponding confidence level ;according to and , the packages are divided into three package priorities: A, B, and C; among them, is the variance of the j-th dimension feature in the 128-dimensional package morphology geometric topology matrix, for The corresponding confidence level.

4. The express delivery business problem solving method based on the knowledge base management system according to claim 3 is characterized in that: The basis and , the packages are divided into three categories: A, B, and C. The package priorities include: Set the first variance threshold V1 and the second variance threshold V2, where V1<V2; if <V1, then assign Weight W1; if V1≤ <V2, then assign Weight W 2; like ≥V2, then give Weight W3; where W1>W2>W3; For each package’s 128-dimensional package geometry topology matrix, according to the confidence level and the assigned weights W1, W2, W3, calculate the weighted average to obtain the comprehensive confidence score CON of the package; A first confidence threshold T1 and a second confidence threshold T2 are set, wherein T1>T2; if CON>T1, the package is classified as Class A; if T2<CON≤T1, the package is classified as Class B; if CON≤T2, the package is classified as Class C.

5. The express delivery business problem solving method based on the knowledge base management system according to claim 2 is characterized in that: The generating of an abnormal tag matrix for a single semantic keyword vector comprises: For the 64-dimensional semantic keyword vector of a face, the frequency of occurrence of semantic keywords in each dimension is counted, and the probability distribution of the frequency of occurrence is calculated; the information entropy H of each dimension is calculated based on the probability distribution k ;H k represents the information entropy of the kth dimension of the semantic keyword vector of the face order, where 1≤k≤64; set the information entropy threshold E, when H k >E, the corresponding abnormal marking matrix position is marked as 0; when H k When ≤E, the corresponding position of the abnormal marking matrix is ​​marked with 1; the number of abnormal markings for each bit in the abnormal marking matrix is ​​counted, the missing probability of each semantic keyword is calculated, and a missing probability vector corresponding to the dimension of the semantic keyword vector of the face order is generated; the missing probability vector is normalized to generate a standardized abnormal marking matrix.

6. The express delivery business problem solving method based on the knowledge base management system according to claim 1 is characterized in that: The method of generating a semantic association matrix M of the morphological feature matrix A and the semantic feature matrix B includes: subject clustering the semantic feature matrix B to obtain m semantic topics; calculating the association strength of each package on the m semantic topics to generate the semantic association matrix M; the semantic association matrix M is an m'×m matrix, wherein m' represents the number of packages.

7. The express delivery business problem solving method based on the knowledge base management system according to claim 1 is characterized in that: The generating of the knowledge recommendation list comprises: Divide the package feature matrix into n2 hash buckets, each hash bucket corresponds to a knowledge subgraph of the knowledge base; In each hash bucket, the structural similarity between the query matrix and the knowledge subgraph is detected, and semantic association edges are dynamically added to obtain an updated knowledge subgraph; the query matrix is ​​obtained by wrapping the feature matrix; Perform random walks on the updated knowledge subgraph, calculate the importance score of each knowledge point, and generate a sorted knowledge recommendation list.

8. The express delivery business problem solving method based on the knowledge base management system according to claim 1 is characterized in that: The method of generating solutions to express business problems based on the knowledge recommendation list includes: clustering the express business problems according to the knowledge recommendation list to obtain problem clusters; for each problem cluster, calling relevant knowledge points in the knowledge recommendation list, combining the graph node state matrix for case reasoning, and generating a set of candidate solutions.

9. A system for solving express business problems based on a knowledge base management system, which is used to implement the express business problem solving method based on a knowledge base management system as claimed in any one of claims 1 to 8, characterized in that: The system comprises: Package priority classification module: used to obtain the package geometry topology matrix to form a package feature set; perform structured analysis on the package geometry topology matrix to classify the packages into three package priorities; Graph space fusion module: used to obtain the semantic keyword vector of the waybill to form the semantic feature set of the waybill; generate the abnormal marking matrix of the semantic keyword vector of the waybill; establish a many-to-many mapping relationship between the package priority and the abnormal marking matrix to generate the association matrix; dynamically generate the graph node state matrix according to the association matrix; Knowledge recommendation module: According to the graph node state matrix, multi-level calibration parameters are obtained; based on the multi-level calibration parameters, a cross-modal mapping function cluster is constructed to obtain the package feature matrix; for the package feature matrix, a knowledge base retrieval and sorting mechanism is designed to generate a knowledge recommendation list; Solution generation module: Generates solutions to express delivery business problems based on the knowledge recommendation list, and feeds the newly generated solutions back to the knowledge base to achieve self-update and iteration of the knowledge base.

Citation Information

Patent Citations

  • Express delivery service internet of things technology processing system and method

    CN105678490A

  • Express business processing method and express business processing device

    CN108335059A

  • Visual perception recommendation method and system based on cross-modal semantic reasoning and fusion

    CN114936901A

  • Comment data processing method and device, equipment, storage medium and product

    CN116955657A