A logistics intelligent sorting method and device based on a double database drive
By constructing a logistics knowledge graph and vector database, the problems of data heterogeneity and missing cargo relationships in logistics sorting were solved, achieving efficient and accurate logistics sorting, reducing the risk of cargo damage, and improving logistics operation efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- XINJIANG SHUTU TECH CO LTD
- Filing Date
- 2026-04-10
- Publication Date
- 2026-07-10
AI Technical Summary
Existing logistics sorting technologies suffer from problems such as coarse classification granularity, lack of cargo correlation, and heterogeneous data that is difficult to integrate when handling fragile or other special-attribute goods. This results in low sorting efficiency and a high risk of damage.
A dual-database-driven approach is adopted. By acquiring logistics management data, preprocessing and standardizing it, constructing a logistics knowledge graph, extracting hidden feature vectors, and using vector databases and graph databases to perform similarity calculations, the physical and chemical attributes of the goods to be sorted are classified.
It has achieved the integration of multi-source heterogeneous data, improved the processing efficiency and rationality of logistics sorting, reduced the risk of cargo damage, and enhanced logistics operation efficiency and service quality.
Smart Images

Figure CN122365064A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent logistics management technology, and in particular to an intelligent logistics sorting method and apparatus based on dual databases. Background Technology
[0002] Logistics, as a crucial support for modern socio-economic operations, directly impacts supply chain efficiency and transaction experience. Among various logistics transport objects, goods with special attributes such as fragile, perishable, flammable, explosive, and easily spoiled goods place higher demands on the transportation environment, packaging methods, and loading strategies. In actual logistics operations, these special goods are prone to damage, leakage, or failure during sorting and transportation.
[0003] Currently, most logistics nodes rely primarily on manual experience to identify and classify goods, or on surface-level information such as barcodes, RFID tags, and keywords in the goods' names for automated sorting. However, this approach has significant limitations when dealing with a large number of goods with unique attributes: on the one hand, manual sorting is time-consuming and labor-intensive, and is susceptible to fatigue and differences in experience, leading to a high error rate; on the other hand, relying solely on goods names or barcodes cannot provide a deep understanding of the goods' physical properties (such as brittleness, density, flammability, etc.) and their potential relationships with other goods. This results in existing sorting systems generally lacking the ability to model the collaborative relationships between goods, limiting the optimization of sorting, packaging, and transportation strategies. Simultaneously, the logistics data environment is becoming increasingly complex, exhibiting typical characteristics of multi-source heterogeneity, fragmentation, and data silos, preventing the effective extraction and utilization of a large amount of potential goods attribute information and loading knowledge. This further exacerbates the challenges of intelligent sorting systems and affects the efficiency of logistics classification and management.
[0004] In summary, existing logistics sorting technologies suffer from significant drawbacks when dealing with fragile or other special-attribute goods, including coarse-grained classification, lack of cargo correlation, and heterogeneous data that is difficult to integrate. How to fully utilize potential cargo attribute information and loading knowledge to form a sustainable, maintainable, and intelligent complete cargo classification and sorting system has become a key issue requiring in-depth research within the logistics industry. Summary of the Invention
[0005] This invention provides a logistics intelligent sorting method and apparatus based on dual databases, which solves the shortcomings of the limited efficiency of logistics classification and management in the prior art, realizes systematic logistics intelligent sorting, improves logistics classification and management efficiency, and reduces the risk of logistics failure.
[0006] In a first aspect, the present invention provides a logistics intelligent sorting method based on dual database driving, comprising: Acquire logistics management data and perform preprocessing and standardization to obtain standardized logistics management data; The standardized logistics management data is used to jointly extract logistics entities and relationships, and a logistics knowledge graph is constructed using a graph database. The semantic features of each logistics entity and relationship in the logistics knowledge graph are extracted by the knowledge graph embedding method, the hidden feature vectors of each logistics entity and relationship are obtained, and the hidden feature vectors are stored in the vector database. By querying the vector database using similarity calculations, the classification results of the physicochemical properties of the goods to be sorted are determined. The goods to be sorted are sorted according to the classification results of the physicochemical properties.
[0007] According to the present invention, a logistics intelligent sorting method based on dual database driving is provided, wherein the acquisition of logistics management data and the preprocessing and standardization processing to obtain standardized logistics management data include: The logistics management data is acquired and then cleaned, format-converted, and noise-reduced to obtain preprocessed logistics management data. The logistics management data originates from multiple components of the logistics management system, including at least the order system, inventory system, and transportation system. The logistics entities in the preprocessed logistics management data are standardized to obtain standardized logistics management data.
[0008] According to the present invention, a logistics intelligent sorting method based on dual databases is provided, wherein the standardized logistics management data is subjected to joint extraction of logistics entities and relationships, and a logistics knowledge graph is constructed through a graph database, including: The logistics entity design is based on the classification of the physical and chemical properties of logistics entities and the transportation and loading relationships between logistics goods; wherein, the classification of the physical and chemical properties of logistics entities includes at least one of fragile goods, toxic chemicals, refrigerated goods, frozen goods, liquid goods, granular and powder goods, easily oxidized goods, easily corrosive goods, and sensitive goods; the transportation and loading relationships include prohibited loading and cooperative loading. The standardized logistics management data is subjected to joint extraction of logistics entities and relationships using a pre-trained information extraction model, resulting in multiple logistics triples. The multiple logistics triples are instantiated and stored in a graph database to construct a basic knowledge graph; Entity alignment is performed on the basic knowledge graphs of different logistics management systems through graph structure modeling; wherein, the different logistics management systems include logistics management systems with different organizational structures and in different regions; The entity alignment results are used to fuse the basic knowledge graphs of the different logistics management systems to construct a logistics knowledge graph.
[0009] According to the present invention, a logistics intelligent sorting method based on dual databases is provided, wherein a pre-trained information extraction model is used to jointly extract logistics entities and relationships from the standardized logistics management data to obtain multiple logistics triples, including: Based on the logistics ontology design, determine the types of logistics entities and relationships to be extracted, and construct prompt words; The logistics information text and the prompt words in the standardized logistics management data are input into a pre-trained information extraction model to jointly extract logistics entities and relationships, thereby obtaining multiple logistics triples corresponding to the standardized logistics management data.
[0010] According to the present invention, a logistics intelligent sorting method based on dual databases is provided, wherein the semantic features of each logistics entity and relationship in the logistics knowledge graph are extracted by a knowledge graph embedding method to obtain hidden feature vectors of each logistics entity and relationship, and the hidden feature vectors are stored in a vector database, including: The structural semantics of the subgraph structure surrounding each logistics entity in the logistics knowledge graph are extracted using a graph neural network, and the relationships are incorporated into the calculation process of the structural semantics extraction based on the translation hypothesis to obtain the structural semantics embedding vector. The degree centrality feature vector of each logistics entity is determined using a webpage ranking algorithm; Random Gaussian noise is introduced into the degree centrality feature vector to simulate variable centrality features, and trainable variable centrality feature vectors are generated. The basic embedding vector, the degree centrality feature vector, and the variable centrality feature vector are respectively fused to obtain the hidden feature vectors corresponding to each of the logistics entities.
[0011] According to the present invention, a logistics intelligent sorting method based on dual database driving is provided, the method further includes: If querying the vector database fails, the pre-trained classifier is used to classify the physical and chemical properties of the goods to be sorted, and a preliminary classification result of the physical and chemical properties is obtained. The preliminary physicochemical property classification results are confirmed and corrected using a large language model to obtain the final physicochemical property classification results. The goods to be sorted are sorted according to the final physicochemical property classification results.
[0012] According to the present invention, a logistics intelligent sorting method based on dual database driving is provided, the method further includes: Update the logistics knowledge graph and the vector database based on the classification results of the physicochemical properties of the goods to be sorted or the final classification results of the physicochemical properties. If the preliminary physicochemical property classification results fail to predict, the pre-trained classifier is optimized.
[0013] According to the present invention, a logistics intelligent sorting method based on a dual-database driver is provided, wherein the logistics sorting of the goods to be sorted according to the classification results of the physicochemical properties includes: Logistics sorting is carried out based on the classification results of the physical and chemical properties of the goods to be sorted, as well as the transportation and loading relationship between the goods to be sorted and other logistics goods.
[0014] Secondly, the present invention also provides a logistics intelligent sorting device based on dual database driving, comprising: The data preprocessing module is used to acquire logistics management data and perform preprocessing and standardization to obtain standardized logistics management data. The knowledge graph construction module is used to jointly extract logistics entities and relationships from the standardized logistics management data, and construct a logistics knowledge graph through a graph database. The vector database construction module is used to extract the semantic features of each logistics entity and relationship in the logistics knowledge graph through the knowledge graph embedding method, obtain the hidden feature vectors of each logistics entity and relationship, and store each of the hidden feature vectors into the vector database. The logistics sorting module is used to query the vector database through similarity calculation to determine the classification results of the physical and chemical attributes of the goods to be sorted; and to perform logistics sorting on the goods to be sorted according to the classification results of the physical and chemical attributes.
[0015] Thirdly, the present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the intelligent logistics sorting method based on dual database driving as described above.
[0016] Fourthly, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the intelligent logistics sorting method based on dual database driving as described above.
[0017] Fifthly, the present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the intelligent logistics sorting method based on dual database driving as described above.
[0018] The beneficial effects of the technical solutions provided by some embodiments of the present invention include at least the following: 1) This invention provides a logistics intelligent sorting method and device based on dual databases. By constructing a unified logistics knowledge graph, it achieves the fusion of multi-source heterogeneous data and extracts hidden feature vectors of various logistics entities and relationships based on the logistics knowledge graph to construct a vector database. By querying the vector database, the physical and chemical attribute classification of goods to be sorted can be quickly determined, thereby efficiently identifying goods with similar physical and chemical attribute categories, improving the processing efficiency and rationality of logistics sorting, and minimizing the damage to goods caused by improper classification and sorting. In addition, this invention utilizes dual databases of graph database and vector database, which can greatly improve the efficiency of logistics goods classification and sorting.
[0019] 2) This invention fully utilizes knowledge graph technology to construct a complete logistics cargo classification system that is easy to expand and maintain. It can also model the collaborative loading relationship between different goods based on the logistics knowledge graph. Combined with vector database retrieval, it can improve the query and calculation speed of the relationship between goods, which is conducive to the automatic identification of special goods such as fragile items, the rapid retrieval of similar items, and the joint sorting and loading decision of logistics, ultimately reducing the risk of cargo damage and improving logistics operation efficiency and service quality.
[0020] 3) At the data governance level, this invention, on the one hand, models the transportation and loading relationships between different logistics goods by constructing a logistics knowledge graph based on a two-layer architecture of ontology concept layer and instance layer, which can provide technical support for subsequent logistics sorting and loading decisions; at the same time, it solves the problems of multi-source heterogeneity and data silos in logistics management data by integrating knowledge from different logistics management systems through entity alignment strategy; on the other hand, by introducing a dual database-driven approach of graph database and vector database, it can improve the storage and usage efficiency of logistics data. Among them, the graph database can handle complex relationships in the logistics knowledge graph, supporting chain-relational knowledge queries and graph traversal operations; the vector database not only supports efficient retrieval of vector data, but also has good data persistence, real-time insertion and query capabilities, which is particularly suitable for nearest neighbor (ANN) retrieval scenarios in the field of natural language processing.
[0021] 4) In terms of the application of scientific methods, this invention, on the one hand, incorporates strategies such as graph centrality statistical measurement and variable centrality learning based on the advanced experience of existing graph neural networks in processing graph data, in order to further improve the effect of knowledge graph embedding learning, in order to generate entity relationship hidden feature vectors containing more complete information; on the other hand, by using existing advanced large language models, and through the design of carefully constructed prompts, the large language model is used to further confirm or correct the logistics entity classification results of the classifier, thereby improving the accuracy of classification. At the same time, this can also be used as a basis for system maintenance and adaptive optimization.
[0022] 5) This invention forms a perfect closed system with "two-layer architecture knowledge graph mode → dual database drive → intelligent classification based on graph neural network and large language model → system adaptive optimization". It designs a complete intelligent logistics sorting process from bottom to top, realizing a perfect closed loop process from accurate data acquisition and standardized data processing at the source of logistics data, to in-depth mining and analysis using scientific methods, and finally realizing the efficient application of results in logistics scenarios, ensuring seamless transformation from theory to practice and maximizing value. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0024] Figure 1 This is a schematic diagram illustrating the classification of the physical and chemical properties of logistics entities and their transportation and loading relationships provided by this invention. Figure 2 This is one of the flowcharts of the intelligent logistics sorting method based on dual databases provided by the present invention; Figure 3 This is the second flowchart of the intelligent logistics sorting method based on dual databases provided by the present invention. Figure 4 This is a schematic diagram of the intelligent logistics sorting device based on dual database driving provided by the present invention. Figure 5 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0025] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0026] Logistics sorting involves various classifications, such as by destination, product category, and customer order attributes. These classifications are used to optimize loading plans, reduce empty load rates and transportation frequency, thereby lowering logistics costs. It should be noted that the logistics classification described in this invention primarily focuses on the physicochemical properties of the logistics entities; other classifications can also be implemented using the embodiments of this invention. Figure 1 The diagram illustrates the classification of logistics entities by physicochemical properties and their transportation and loading relationships, as provided by this invention. Logistics entities are classified according to their physicochemical properties into fragile goods, soft cushioning materials, granular and powdery goods, liquids, toxic chemicals, and refrigerated foods, among others. Each type of logistics cargo has its own packaging and preservation requirements during warehousing and transportation, and there are potential correlations between different logistics goods. For example, fragile goods (bottled beer, glassware, ceramic products, etc.) often share similar risk attributes. Quickly identifying and clustering similar fragile goods helps in adopting a unified sorting and packaging strategy. Furthermore, fragile goods and soft items (such as clothing, sponges, foam materials, etc.) have a natural buffering synergy; proper loading can significantly reduce physical impact damage during transportation. However, existing sorting logic does not systematically incorporate this type of correlation knowledge.
[0027] Therefore, by establishing loading relationships between each type of logistics goods, such as cooperative loading and prohibited loading relationships, a basic framework for an intelligent logistics sorting system can be constructed. This invention manages logistics sorting and transportation based on intelligent logistics classification and cooperative loading relationships between goods, which can reduce logistics transportation losses, optimize logistics resource allocation, and improve the level of intelligent logistics management.
[0028] Please see Figure 2 , Figure 2 One of the flowcharts for an intelligent logistics sorting method based on a dual-database driver, provided as an embodiment of the present invention, includes: S110. Obtain logistics management data and perform preprocessing and standardization to obtain standardized logistics management data; S120. Perform joint extraction of logistics entities and relationships from standardized logistics management data, and construct a logistics knowledge graph through a graph database; S130. Extract the semantic features of each logistics entity and relationship in the logistics knowledge graph by using the knowledge graph embedding method, obtain the hidden feature vectors of each logistics entity and relationship, and store each hidden feature vector in the vector database. S140. By querying the vector database through similarity calculation, determine the classification results of the physical and chemical properties of the goods to be sorted; S150. Based on the classification results of physical and chemical properties, perform logistics sorting of the goods to be sorted.
[0029] This invention achieves multi-source heterogeneous data fusion by constructing a unified logistics knowledge graph. Based on the logistics knowledge graph, it extracts hidden feature vectors of various logistics entities and relationships to build a vector database. By querying the vector database, the physical and chemical attribute classification of goods to be sorted can be quickly determined, thereby efficiently identifying goods with similar physical and chemical attribute categories. Furthermore, it can model the collaborative loading relationship between different goods based on the logistics knowledge graph, which is conducive to the automatic identification of special goods such as fragile items, rapid retrieval of similar items, and joint logistics sorting and loading decisions, ultimately reducing the risk of cargo damage and improving logistics operation efficiency and service quality.
[0030] like Figure 3 The diagram shown is a second schematic of a logistics intelligent sorting method based on dual databases provided by the present invention, wherein the classification of the physical and chemical properties of the goods to be sorted can be achieved through an intelligent logistics application.
[0031] The following is combined Figure 2 , Figure 3 The method of the present invention will be described in detail.
[0032] In the above implementation, step S110 mainly addresses the questions of where to obtain data and how to process it, with the aim of analyzing and mining relevant entities in the logistics management data.
[0033] Logistics management is a complex systems engineering project involving data across the entire chain, from planning, sourcing, manufacturing, delivery, and returns. Complete logistics management requires collecting data from each stage of the logistics process, such as the status of transportation vehicles (e.g., GPS tracks, sensor data), cargo attributes (e.g., value, fragility), warehousing information, customs declaration data, environmental data (e.g., weather, road conditions), and financial data (e.g., transaction records). Data preprocessing primarily handles the reprocessing of acquired knowledge, such as data noise reduction, entity standardization, and entity disambiguation, to ensure the quality of knowledge extraction.
[0034] Specifically, logistics management data can originate from multiple stages and systems within logistics management, such as order systems, inventory systems, transportation systems, customer systems, and equipment systems. This embodiment focuses primarily on logistics entities originating from the order system, inventory system, and transportation system. The order system contains basic information about the logistics entity, including quantity and amount. The inventory system provides the inventory quantity and location of the logistics entity. The transportation system provides the waybill number, quantity, weight, volume, and transportation time of the logistics entity. After identifying the data sources, data from different systems can be integrated into a single data platform using API interfaces, middleware, and other technologies. ETL (Extract, Transform, Load) tools are then used to extract data from each data source system, perform preliminary data cleaning and format conversion, and obtain the raw logistics management data.
[0035] The main challenges in acquiring logistics management data include data silos, data quality, and data security. This invention provides corresponding solutions to these problems. For data silos, inconsistent data formats and incompatible interfaces between different data source systems hinder data sharing. The solution is to establish unified data standards and interface specifications, and adopt a microservice architecture to achieve loose coupling of data. For data quality issues, data may contain missing, incorrect, or duplicate data. The solution is to establish a data quality management system and improve data quality through data cleaning and verification. For data security, logistics data involves customer privacy and corporate trade secrets, requiring strict security and confidentiality. The solution is to employ encryption technology, access control, and security auditing to ensure data security.
[0036] In some possible embodiments, logistics management data is acquired and preprocessed and standardized to obtain standardized logistics management data, including: Acquire logistics management data and perform data cleaning, format conversion and noise reduction on the logistics management data to obtain preprocessed logistics management data; The logistics entities in the preprocessed logistics management data are standardized to obtain standardized logistics management data.
[0037] Specifically, in addition to basic preprocessing such as data cleaning and format conversion, a further step of noise reduction is needed for data quality control. In the process of noise reduction for logistics data, the main solutions are determined based on the type and impact of the logistics noise data, including: Duplicate data: The same order, product, or customer information appears repeatedly in different systems, causing bias in data analysis. The main solution is deduplication based on hash tables. By calculating the hash value of the data, data with the same hash value are considered duplicates, and only one of them is retained.
[0038] Missing data: Some key fields (such as product specifications, customer addresses, etc.) are missing, affecting data integrity and analysis results. The main solution is to fill in the missing values using default values, averages, medians, etc.
[0039] Error data: Data anomalies caused by data entry errors or system malfunctions, such as incorrect order quantities or shipping times. The main solution is a combination of business rule validation and manual review. Data is validated using business rules and data dictionaries to identify and correct errors. For critical data, such as order amounts and customer addresses, manual review and correction are performed.
[0040] Abnormal data: Data outside the normal range, such as abnormally high order amounts or abnormally low inventory quantities, may be due to system errors or human error. The main solution is to use anomaly detection algorithms (such as Isolation Forest, LOF, etc.) to automatically identify abnormal data.
[0041] After completing the above preprocessing, further data standardization processing has been performed. In this embodiment, standardization processing mainly refers to unifying the attributes and behaviors of logistics entities into a standard format to facilitate data management and analysis. Through standardization processing, inconsistencies and errors in the data are eliminated, improving the accuracy and completeness of the data. Unified data standards and formats facilitate data sharing and exchange between different systems. This invention mainly focuses on the classification of logistics goods. When standardizing the names of logistics goods, methods such as removing spaces, full-width characters, and special symbols (such as " / ", "-", and "()"), unifying capitalization (such as unifying "iPhone" and "IPHONE" as "iPhone"), and removing redundant words (such as marketing terms such as "new model," "hot-selling model," and "free shipping") are often adopted. For example, Coca-Cola 330ml and Cola-330ml are uniformly named Coca-Cola 330ml canned.
[0042] In the above implementation, step S120 mainly involves ontology design, knowledge extraction, and knowledge fusion to construct a logistics knowledge graph.
[0043] To achieve intelligent sorting in intelligent logistics, the first step is to carefully design a detailed classification system for logistics goods and the relationships between each category. This involves ontology design to obtain the ontology concept layer. With the ontology concept layer, knowledge extraction can be performed based on its defined classification system, identifying relevant logistics entities and their interrelationships from the acquired structured or unstructured logistics management data. Subsequently, the instance layer can replicate the ontology concept layer model to construct a logistics instance knowledge graph. Furthermore, logistics entities may originate from different regions and e-commerce systems. For the same logistics entity, different regions and countries may have different languages and local names, and different e-commerce systems may have different descriptive names for the same product. Therefore, knowledge graph alignment technology is also needed to align the different names of the same entity, achieving knowledge fusion and ultimately constructing a unified logistics knowledge graph.
[0044] For example, in the knowledge extraction process, logistics management data may be structured or unstructured. The specific process of extracting entities and relationships from the logistics knowledge graph corresponding to different types of data is as follows: If the acquired logistics management data is a structured logistics information table, such as an Excel-like table containing fields like goods name, goods type, origin, destination, shipping date, shipping company, transportation cost, transportation status, transportation conditions, and transportation mode, knowledge extraction can be performed as follows: First, extract key entities from the table and map them to the same ontology. Key entity types include company, location, goods, order, and transportation mode. Next, use NLP tools such as HanLP for entity identification and standardize entity names. Second, define semantic relationships between entities based on the table fields to form the basic building blocks of the knowledge graph: triples (head entity - relation - tail entity).
[0045] If the acquired logistics management data is unstructured, such as emails, contracts, customer communication records, and images and videos describing the transportation process, these unstructured data share common characteristics: no fixed fields, rich semantics, and diverse formats, making them difficult to store and query directly using traditional database tables. When extracting knowledge from this unstructured logistics information, OCR technology is first used to convert image-based documents (such as contracts and handwritten documents) into readable text. Then, classic knowledge extraction methods, such as using a traditional LSTM+CRF architecture, are employed to implement text sequence labeling and output relevant entities. Finally, a relational classification model is used to establish relationships between different entities, thus forming triples.
[0046] For example, in the process of knowledge fusion, for the same entity with different names in different systems and regions, graph neural network technology and deep learning technology can be used to fully explore the graph structure and relationships of the entities by using aligned seed entity pairs in the data as training samples, in order to discover more unaligned entity pairs and ultimately achieve knowledge fusion.
[0047] In some possible embodiments, standardized logistics management data undergoes joint extraction of logistics entities and relationships, and a logistics knowledge graph is constructed using a graph database, including: S120-1. Design the logistics ontology based on the classification of the physical and chemical properties of logistics entities and the transportation and loading relationships between logistics goods. S120-2. By using a pre-trained information extraction model, logistics entities and relationships are jointly extracted from standardized logistics management data to obtain multiple logistics triples. S120-3. Instantiate multiple logistics triples and store them in a graph database to construct a two-layer basic knowledge graph. S120-4. Entity alignment of the basic knowledge graphs of different logistics management systems through graph structure modeling; S120-5. Based on the results of entity alignment, knowledge is integrated into the basic knowledge graphs of different logistics management systems to construct a logistics knowledge graph.
[0048] Specifically, in step S120-1 above, during the logistics ontology design process, this embodiment classifies logistics goods according to their physical and chemical properties, dividing entity types into several major categories such as fragile goods, soft cushioning materials, liquids, granular and powdery goods, toxic chemicals, and refrigerated foods. Furthermore, based on the transportation and loading relationships between logistics goods, the relationship types are divided into cooperative loading relationships and prohibited loading relationships. This constructs an ontology concept layer that can be used for logistics goods classification and subsequent sorting and transportation. In subsequent use, with the accumulation of experience and knowledge, these physical and chemical property classifications can be further refined to improve system applicability.
[0049] It is understood that the above classification of the physical and chemical properties of logistics goods can be set as needed. For example, in another embodiment, they can be classified as fragile goods, toxic chemicals, refrigerated goods, frozen goods, liquids, granular powders, easily oxidized goods, easily corrosive goods, sensitive goods, volatile goods, flammable and explosive goods, etc., or any combination of these categories. This invention does not impose any restrictions.
[0050] For example, considering that logistics management data includes not only the physical and chemical attributes of goods as described above, but also logistics participants, geographical locations, service providers, and document information, in order to establish a comprehensive logistics knowledge graph and logistics park sorting system, and to facilitate subsequent logistics sorting and transportation, the logistics ontology design should also include various entity types and relationship types related to logistics management. For instance, in addition to the physical and chemical attribute classifications mentioned above, entity types can also include: shipper, consignee, contact number, goods name, quantity, weight, place of origin, destination, courier company, waybill number, and other basic entity types. Relationship types can include: [Shipper] sends [goods], [Goods A] coordinates with [Goods B], [Goods], [Goods A] prohibits [Goods B] from being loaded, etc.
[0051] In step S120-2 above, this embodiment mainly adopts a joint extraction method for extracting logistics entities and relationships. The traditional pipeline approach, which performs Entity Recognition (NER) first and then Relationship Classification (RE), has the following drawbacks: error accumulation (incorrect entity recognition directly leads to relationship extraction failure); exposure of bias (using real labels during training and predicted labels during inference, resulting in inconsistent distribution); and lack of interaction (failing to utilize the implicit associations between entities and relationships (e.g., the subject of "sent" must be a person). The joint extraction technique used in this embodiment unifies Entity Recognition (NER) and Relationship Extraction (RE) into a "fragment extraction" task, directly outputting structured triples (head entity, relationship, tail entity) through a pre-trained information extraction model.
[0052] In some possible embodiments, a pre-trained information extraction model is used to jointly extract logistics entities and relationships from standardized logistics management data, resulting in multiple logistics triples, including: Determine the types of logistics entities and relationships to be extracted, and construct prompt words; among them, the relationship types include the transportation and loading relationships between two logistics entity types, and the transportation and loading relationships include prohibited loading and cooperative loading; The logistics information text and prompts from standardized logistics management data are input into a pre-trained information extraction model to jointly extract logistics entities and relationships, resulting in multiple logistics triples corresponding to the standardized logistics management data.
[0053] Specifically, the pre-trained information extraction model can adopt the SiameseUIE model, which uses a prompt mechanism for unified modeling. Its implementation mainly includes three steps: defining extraction fields, constructing prompt templates, and inputting text from logistics management data for reasoning. First, based on the classification system of the ontology concept layer, the entity types and relation types to be extracted are clarified. Next, prompt words are constructed, such as designing task prompts based on JSON format encoding, to instruct the SiameseUIE model to extract logistics entities and relations according to requirements. For example, for basic entity types and relation types, the following structured output requirements based on JSON format encoding can be used when constructing prompt words: { "schema": { "Sent out": ["Shipper", "Goods"], "To": ["Goods", "Destination"], "Received": ["Consignee", "Goods"] } } Finally, the logistics information text from the standardized logistics management data and the constructed prompt words are input into the SiameseUIE model to extract logistics entities and relations, resulting in multiple logistics triples with a head entity-relationship-tail entity structure.
[0054] In step S120-3 above, under the guidance of the ontology concept layer in step S120-1, the logistics triplet in step S120-2 is imported to achieve instantiation. Combining the ontology concept layer and the instance layer, a two-layer logistics knowledge graph can be constructed.
[0055] In this embodiment, the Neo4j graph database is used to store the logistics knowledge graph. In the Neo4j graph database, entities in the knowledge graph correspond to nodes in the graph database, and relationships are represented by edges. Unlike a knowledge graph that only contains triples, nodes and relationship edges in the Neo4j graph database can both contain properties. After the knowledge graph is placed in the Neo4j graph database, it essentially becomes a property graph. Entity and relationship properties are stored in key-value pairs. In reality, entities and relationships in the knowledge graph may not have properties, meaning their properties are empty. In the knowledge graph, entities are typically represented as nodes, and relationships between entities are represented as edges. For example, "Beijing" is an entity node, and "located in" is a relationship edge between two entities.
[0056] Unlike traditional relational databases that simulate graph structures through table joins, Neo4j graph database supports native graphs at both the storage and query levels. At the storage level, it employs a native graph storage mechanism, storing graph data directly on disk in graph structure form. It utilizes features such as index-free adjacency, fixed-size records, and data separation to improve storage and query speed. At the query level, it provides a specially optimized query language, Cypher, focused on describing connection patterns in graphs, simplifying the query efficiency of complex graphs. This native graph storage mechanism makes Neo4j well-suited for handling large-scale, high-density graph data, such as in social network analysis, recommendation systems, and supply chain analysis. Furthermore, Neo4j provides a Python interface, py2neo, which allows for easy and rapid import of small-scale knowledge graphs into the graph database. This embodiment prioritizes writing simple Python code to store a logistics knowledge graph into the graph database. As the logistics knowledge graph grows, the dedicated import tool neo4j-import provided by Neo4j can be used.
[0057] In step S120-4 above, different names referring to the same entity in the knowledge graph are aligned to further integrate knowledge from different logistics management systems or different data source systems. These different logistics management systems include those with different organizational structures and geographical locations; for example, in a cross-border e-commerce system, the same entity may have different names due to language differences. In this case, knowledge fusion is necessary through entity alignment.
[0058] Suppose there are two knowledge graphs and , where the symbol Let i represent the set of entities, relations, and triples of the knowledge graph numbered i. Let j represent the sets of entities, relations, and triples in the knowledge graph. The task of entity alignment is to align the knowledge graph. and The basic idea behind unaligned entities is to continuously train and learn the embedding vectors of entities and relations in the knowledge graph, so that the embedding vectors of the same entity and relation in two knowledge graphs are closer together, while the embedding vectors of different entities are farther apart.
[0059] To more fully learn the embedding vectors of entities and relations, it is necessary to model the subgraph structure surrounding the entity. Since a knowledge graph is a directed graph, for a given entity... Its subgraph structure includes two types of links: outgoing links with the entity as the head and incoming links with the entity as the tail. Furthermore, to ensure that each node in the subgraph structure retains its own information during information transmission, each node introduces a self-loop relationship pointing to itself. For entities... Modeling a knowledge graph involves aggregating its surrounding inbound and outbound links. Since a knowledge graph is a multi-relationship graph, relationship labels are crucial for modeling its structure. Therefore, unlike ordinary undirected graphs, it's essential to address how to incorporate relationship information during the link aggregation process.
[0060] For example, we can draw on the Translational Embedding (TransE) model from the field of knowledge graph modeling, viewing relations as a translation process from head entity to tail entity. The head entity h, tail entity t, and relation r can be formally represented by formulas. or This can be represented as follows. Ultimately, the graph structure modeling process based on graph neural networks can be formally represented as follows:
[0061] in, It is the normalization coefficient, equal to the entity. The degree, This indicates the number of layers in a graph neural network. It is a trainable matrix. Representative Entity In the Layer embedding vector representation, This represents an activation function (such as ReLU, Sigmoid, etc.). Representative Entity In the Layer embedding vector; Represents the current entity The set of relation types involved in the subgraph structure; Representative in relationship Below, in physical form It is the set of all tail entities of the head entity; Representative in relationship Below, in physical form The set of all head entities of the tail entity; Represents the tail entity In the Layer embedding vector; Represents the head entity In the Layer embedding; Representative relationship In the Layer embedding vector; Representing the The learnable weight matrix of the layer.
[0062] By using the formulas in the above structural modeling process, we can learn about solids. Fusion Embedded Representation that Incorporates Subgraph Structure Ultimately, the training objective of the entity alignment task can be achieved through iterative iteration of the loss function, resulting in an entity alignment model. The main function of the loss function is to make the embedding vectors of aligned entities as close as possible, rather than making the embedding vectors of the same entity as far apart as possible. With a well-trained entity alignment model, entities from different logistics management systems or data source systems can be automatically aligned, greatly reducing the cost of relying on manual alignment and improving the fusion efficiency of knowledge graphs.
[0063] In step S120-5 above, the basic knowledge graphs of different logistics management systems are fused based on the results of entity alignment. After knowledge extraction and fusion, a logistics knowledge graph with rich information will be obtained.
[0064] In the above implementation, the purpose of step S130 is to transform the symbolic representation of the logistics knowledge graph into numerical knowledge, so as to provide support for the accurate classification of logistics entities in the future.
[0065] Symbolic representation here means that the entities and relationships in the logistics knowledge graph are represented by natural language text, which is not conducive to the subsequent analysis and reasoning of the logistics knowledge graph. Therefore, this invention uses the knowledge graph embedding method to fully explore the semantic features of the entities and relationships in the knowledge graph, form hidden feature vectors in the form of embedded vector representation, and store them in the vector database Milvus.
[0066] Through in-depth exploration by academic researchers, various knowledge graph embedding methods have emerged, such as the TransE method based on translation operations, the TuckER method based on tensor decomposition, and the ConvE method based on convolutional neural networks. Research shows that analyzing and modeling the input of the central node can better capture the global semantic information of the knowledge graph. Traditional graph neural network methods, such as convolutional neural networks (GCNs), learn the semantic information of nodes in the graph network by aggregating information from surrounding neighboring nodes. This approach essentially only models the local structural information around the entity, neglecting the global structural information of the graph network, and therefore cannot fully model the graph structure data.
[0067] In some possible embodiments, semantic features of each logistics entity and relationship in the logistics knowledge graph are extracted using a knowledge graph embedding method to obtain hidden feature vectors for each logistics entity and relationship, and these hidden feature vectors are stored in a vector database, including: S130-1. Structural semantics are extracted from the subgraph structure around each logistics entity in the logistics knowledge graph using a graph neural network. Based on the translation hypothesis, the relationship is integrated into the calculation process of structural semantics extraction to obtain the structural semantics embedding vector. S130-2. Determine the degree centrality feature vector of each logistics entity using a webpage ranking algorithm; S130-3. Random Gaussian noise is introduced on the basis of degree centrality feature vectors to simulate variable centrality features and generate trainable variable centrality feature vectors. S130-4. The basic embedding vector, degree centrality feature vector, and variable centrality feature vector are respectively fused to obtain the hidden feature vectors corresponding to each logistics entity.
[0068] Specifically, in order to fully learn the hidden features of entities and relationships containing the knowledge graph structure, this embodiment proposes a knowledge graph embedding method based on variable centrality. That is, the concept of centrality modeling is introduced into the semantic modeling stage of the knowledge graph, in order to further improve the level of existing knowledge graph embedding methods.
[0069] Currently, various statistical methods exist for measuring centrality in graph neural networks. These include: Degree Centrality, which measures the number of direct connections a node has; a higher degree indicates more neighbors and stronger information acquisition and propagation capabilities; Closeness Centrality, which reflects the reciprocal of the average shortest path distance from a node to all other nodes in the network; a higher value indicates faster access to other nodes; Betweenness Centrality, which measures how many pairs of nodes a node acts as a "bridge" on the shortest path; a higher value indicates stronger control over information flow; Eigenvector Centrality, which considers not only the node's own number of connections but also the importance of its neighbors, with nodes connected to high-influence nodes receiving higher scores; and an improved Eigenvector Centrality proposed by the PageRank algorithm, which introduces a random jump mechanism to prevent weight concentration on a few nodes. Other metrics include Katz Centrality and Delta Centrality. However, these methods are fixed statistical measures and cannot capture the changes in centrality observed from different observer perspectives.
[0070] To facilitate understanding, let's consider a real-world scenario to illustrate the concept of variable centrality. Take a computer science department at a university as an example. The faculty members can form a small social network. If we analyze the central nodes of this social network from the perspective of administrative leaders, then the department secretary and dean are the central nodes. However, if we analyze it from the perspective of academic leaders, then a professor with profound academic achievements becomes the central node. This demonstrates that the central nodes of the network change depending on the perspective from which the same graph network is observed. Therefore, simply using fixed centrality measures from statistics is insufficient to meet the requirements for modeling variable centrality.
[0071] To address this challenge, this invention introduces a learnable centrality metric parameter to model variable centrality. Specifically, it integrates centrality embedding vectors and centrality statistics into traditional graph neural networks, training a knowledge graph embedding model to achieve the variable centrality knowledge graph embedding method of this invention. This overcomes the limitation of traditional graph neural networks in modeling the global semantics of graph structures.
[0072] The following describes the specific implementation of the knowledge graph embedding method based on variable centrality proposed in this invention, in conjunction with the above steps S130-1 to S130-4.
[0073] In step S130-1 above, similar to the entity alignment method described above, graph structure modeling is performed using a graph neural network. The graph neural network extracts structural semantics from the subgraph structures surrounding each logistics entity in the logistics knowledge graph, and incorporates relationships into the structural semantic extraction calculation process based on translation hypotheses.
[0074] Specifically, for a given logistics entity in a logistics knowledge graph First, the surrounding subgraph structure (including inbound and outbound links) must be modeled to extract its structural semantics. Then, based on the translation hypothesis, the relationships are integrated into the graph structure modeling process. The graph structure modeling process is described in formula (1), and will not be repeated here. Through this process, a structural semantic embedding vector incorporating structural features can be obtained. .
[0075] In step S130-2 above, the degree centrality feature vector of the entity can be obtained using the PageRank algorithm. .
[0076] Specifically, the core idea of the PageRank algorithm is that "nodes pointed to by important nodes are more important," and its basic formula is:
[0077] in, Representing logistics entities The node PageRank value; Representative node PageRank value; Represents the damping factor, indicating the probability. Spreading along the border, with Random redirection; Represents the total number of nodes. Represents a pointer to a node The set of all nodes (in-neighborhood). Representative node out-degree (from) (Number of starting edges). The iterative calculation process is as follows: initialize the PageRank value of all nodes to 1 / N, repeatedly apply the above formula until convergence, and finally obtain a probability distribution vector (the sum of the PageRank values of all nodes is 1), that is, the logistics entity. Degree centrality eigenvectors It can be viewed as a logistics entity. The distribution vector of “importance / authority”.
[0078] In step S130-3 above, in this centrality eigenvector Random Gaussian noise is added to simulate variable centrality features, and a trainable variable centrality feature vector is generated. .
[0079] In step S130-4 above, the logistics entity The hidden feature vector (embedding vector) is , It contains both local and global structural information of the entity.
[0080] For example, a knowledge graph embedding model can be trained based on a graph neural network to implement the knowledge graph embedding methods corresponding to steps S130-1 to S130-4 above.
[0081] Compared to the name information of entities and relationships, their hidden feature vectors contain richer and more complete information, such as graph structure information and association information with other entities, which are the cornerstone of subsequent entity analysis and classification. After comprehensively analyzing the advantages and disadvantages of existing graph neural network technologies, this invention further improves the performance of graph structure modeling by adding graph centrality measures, variable graph centrality hidden feature learning, and key subgraph mining algorithms.
[0082] After obtaining the hidden feature vectors of all entities and relationships, they are stored in the Milvus vector database for direct retrieval and use in subsequent classification tasks. The Milvus vector database is specifically designed for efficient similarity searching of large-scale, high-dimensional vector data, offering significant advantages in performance, scalability, ease of use, and ecosystem support. It supports over 10 index types, such as HNSW, IVF_FLAT, Annoy, and GPU-accelerated CAGRA, allowing users to flexibly choose based on their accuracy, speed, and resource consumption requirements. HNSW is suitable for high-precision real-time retrieval, while GPU indexing can increase index building speed by up to 50 times. Vector databases inherently possess efficient retrieval capabilities, and the database itself has efficient data search or matching algorithms, easily finding similar goods for subsequent modules to aid in analysis and judgment. The Milvus vector database natively supports Cosine similarity, which greatly improves the efficiency of logistics entity classification in this invention.
[0083] In the above implementation, for step S140, when a new shipment arrives, the hidden feature vector with the highest similarity is determined by querying the vector database, thereby extracting the corresponding logistics entity type and obtaining the physical and chemical attribute classification result of the shipment.
[0084] Specifically, based on the basic information of the goods to be sorted, the embedding vector of the goods to be sorted is determined, the vector database is queried, the similarity between the embedding vector of the goods to be sorted and each hidden feature vector in the vector database is calculated, and the hidden feature vector with the highest similarity is determined. The physical and chemical attribute classification of the logistics entity corresponding to the hidden feature vector with the highest similarity is the physical and chemical attribute classification result of the goods to be sorted.
[0085] In the above implementation, for step S150, the goods to be sorted are sorted according to the classification results of physical and chemical properties.
[0086] Specifically, the physicochemical properties of the goods to be sorted are classified as fragile, toxic chemicals, refrigerated goods, frozen goods, liquids, granular or powdery goods, easily oxidized or corrosive goods, or sensitive goods. Based on this classification, goods belonging to the same physicochemical property category can be grouped together, quickly identifying and clustering similar goods, which helps in adopting a unified sorting and packaging strategy.
[0087] For example, since the logistics knowledge graph constructed in step S120 can also include basic information such as destination, transportation conditions and transportation mode, the classification results of the physical and chemical properties of the goods to be sorted can be combined with these basic information to carry out logistics sorting, which facilitates subsequent secondary sorting or loading and transportation.
[0088] In some possible embodiments, the goods to be sorted are sorted according to their physicochemical properties, including: Logistics sorting is carried out based on the classification results of the physical and chemical properties of the goods to be sorted, as well as the transportation and loading relationship between the goods to be sorted and other logistics goods.
[0089] Specifically, the logistics knowledge graph constructed in step S120 also models the transportation and loading relationships between logistics goods. Therefore, after determining the physicochemical property classification result of the current goods to be sorted, the transportation and loading relationships between the goods and other logistics goods can be combined to coordinate logistics sorting. For example, if the physicochemical property classification result of the current goods to be sorted is fragile, and according to the transportation and loading relationships between logistics goods, fragile goods can be coordinated with soft cushioning materials. Therefore, during logistics sorting, the current goods to be sorted can be sorted together with other soft cushioning materials that have already been classified and have the same destination, which facilitates subsequent loading and transportation.
[0090] Furthermore, once a logistics entity is correctly categorized, the Milvus vector database can quickly find similar logistics entities and their transportation and loading relationships with other logistics entities, avoiding duplicate classifications and queries and reducing unnecessary computing power and energy consumption.
[0091] In some possible embodiments, the step S140 described above is followed by: S160. If querying the vector database fails, the pre-trained classifier is used to classify the physical and chemical properties of the goods to be sorted, and a preliminary classification result of the physical and chemical properties is obtained. S170. The preliminary physicochemical property classification results are confirmed and corrected using a large language model to obtain the final physicochemical property classification results.
[0092] Specifically, in cases where querying the vector database fails or the classification result of the physicochemical properties of the goods to be sorted cannot be determined, the goods can be automatically classified using a pre-trained classifier. Building upon this, to further improve classification accuracy and persuasiveness, a large-model classification confirmation mechanism is introduced.
[0093] The classifier can be implemented using a simple fully connected neural network model followed by a softmax operation. Based on the logistics ontology design, logistics goods encompass a wide variety of categories, which can be viewed as a multi-classification task. The loss function for a multi-classification task... It can be formally expressed as the following formula:
[0094] in, This indicates the total number of samples. Indicates the total number of categories. It is the true class label of the entity in the i-th sample. is the class label of the entity in the i-th sample predicted by the classifier, i=1,2,…,N, c=1,2,…,C.
[0095] Through continuous iterative training on the dataset, a classifier with accurate classification can be trained. After inputting the basic information of the unclassified goods to be sorted into the classifier model, the category of its logistics entity can be initially determined.
[0096] Next, a set of refined prompt words needs to be constructed and input into a large language model to further confirm and correct the classification results of the pre-trained classifier. The prompt words contain graph structure contextual information of relevant entities and information on similar items, which can help the large language model make more accurate judgments.
[0097] For example, a prompt word could be: "Suppose you are an expert in the logistics field, and you are very familiar with the physical and chemical properties of common logistics goods, such as wine glasses being fragile and Coca-Cola being a liquid. Now, you are given an entity A, and the classifier classifies it as XXX. Do you think this is correct? Please provide detailed reasons." Inputting this prompt word into the large model will confirm the final result of the entity classification.
[0098] Pre-trained classifiers and large language models can serve as important supplementary means in the process of improving logistics knowledge graphs, thereby gradually improving and enhancing the performance of logistics knowledge graphs and intelligent logistics sorting systems.
[0099] At this point, the intelligent classification task for logistics entities has been completed, and the final classification results based on physical and chemical properties can be used for subsequent logistics sorting.
[0100] In some possible embodiments, step S170 is followed by: S180. Update the logistics knowledge graph and vector database based on the classification results of the physicochemical properties of the goods to be sorted or the final classification results of the physicochemical properties. S190. If the preliminary physicochemical property classification results fail to predict, optimize the pre-trained classifier.
[0101] Specifically, this embodiment introduces an adaptive optimization mechanism based on the physicochemical property classification results of step S150 or step S170. For entities with abnormal classifications, it is analyzed whether they are newly added logistics entities or existing entities in the logistics knowledge graph, and then different optimization strategies are adopted according to the results. If the abnormally classified entity is a newly added entity, its hidden features need to be fully trained.
[0102] For example, firstly, verify whether the classification in step S150 is correct. If the classification is incorrect, analyze the relevant entities in depth to determine whether they are new or existing entities. If they are new entities, further train the knowledge graph embedding model to fully extract their hidden feature vectors to improve the classification quality of subsequent goods to be sorted. This is because if frequently misclassified entities are new entities not in the constructed logistics knowledge graph, it means that the hidden features of that entity have not been fully explored or have not been trained well. It is necessary to continue iteratively training its embedding vectors until the classification model can correctly classify it. After further training the knowledge graph embedding model, it is also necessary to update the logistics knowledge graph and vector database. For example, add newly appearing entities to the logistics knowledge graph and delete outdated entities. The update of the vector database involves two aspects: firstly, updating the hidden feature vectors of newly added logistics entities and relationships; secondly, due to optimization adjustments, the hidden feature vectors of some entities and relationships have been further optimized and need to be updated synchronously in the vector database.
[0103] If the misclassified entity is an entity that already exists in the knowledge graph, it is first assumed that its hidden features have been fully mined by the knowledge graph embedding submodule, that is, the knowledge graph embedding model is considered to have been fully trained. Then, the misclassified entity and its label are input into the classifier to update the classifier parameters and optimize the classifier.
[0104] Similarly, if the large language model in step S170 determines that the classifier in step S160 has made a misclassification, then the classifier needs to be optimized. Furthermore, the prompt words of the large language model can be progressively optimized based on the final classification result of step S170, so that the large language model can more accurately assist in analysis, reasoning, and judgment.
[0105] Through the above adaptive optimization mechanisms, the intelligent logistics sorting system will continuously improve its sorting performance.
[0106] Please see Figure 4 , Figure 4 A schematic diagram of a logistics intelligent sorting device 400 based on dual database driving is provided for an embodiment of the present invention. The device includes: The data preprocessing module 410 is used to acquire logistics management data and perform preprocessing and standardization to obtain standardized logistics management data. The knowledge graph construction module 420 is used to jointly extract logistics entities and relationships from standardized logistics management data and construct a logistics knowledge graph through a graph database. The vector database construction module 430 is used to extract the semantic features of each logistics entity and relationship in the logistics knowledge graph through the knowledge graph embedding method, obtain the hidden feature vectors of each logistics entity and relationship, and store each hidden feature vector into the vector database. The logistics sorting module 440 is used to query the vector database through similarity calculation to determine the classification results of the physical and chemical attributes of the goods to be sorted; and to perform logistics sorting of the goods to be sorted according to the classification results of the physical and chemical attributes.
[0107] The intelligent logistics sorting device based on dual databases and the intelligent logistics sorting method based on dual databases described above can be referred to and correspond to each other.
[0108] Figure 5 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 5 As shown, the electronic device may include a processor 510, a communications interface 520, a memory 530, and a communication bus 540. The processor 510, communications interface 520, and memory 530 communicate with each other via the communication bus 540. The processor 510 can call logical instructions stored in the memory 530 to execute a dual-database-driven intelligent logistics sorting method provided in the above embodiments.
[0109] Furthermore, the logical instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0110] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute a logistics intelligent sorting method based on dual database drive provided in the above-described method embodiments.
[0111] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform a dual-database-driven intelligent logistics sorting method provided by the methods described above.
[0112] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0113] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0114] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A logistics intelligent sorting method based on dual databases, characterized in that, include: Acquire logistics management data and perform preprocessing and standardization to obtain standardized logistics management data; The standardized logistics management data is used to jointly extract logistics entities and relationships, and a logistics knowledge graph is constructed using a graph database. The semantic features of each logistics entity and relationship in the logistics knowledge graph are extracted by the knowledge graph embedding method, the hidden feature vectors of each logistics entity and relationship are obtained, and the hidden feature vectors are stored in the vector database. By querying the vector database using similarity calculations, the classification results of the physicochemical properties of the goods to be sorted are determined. The goods to be sorted are sorted according to the classification results of the physicochemical properties.
2. The intelligent logistics sorting method based on dual database driving according to claim 1, characterized in that, The process of acquiring logistics management data and performing preprocessing and standardization to obtain standardized logistics management data includes: The logistics management data is acquired and then cleaned, format-converted, and noise-reduced to obtain preprocessed logistics management data. The logistics management data originates from multiple components of the logistics management system, including at least the order system, inventory system, and transportation system. The logistics entities in the preprocessed logistics management data are standardized to obtain standardized logistics management data.
3. The intelligent logistics sorting method based on dual database driving according to claim 1, characterized in that, The process of jointly extracting logistics entities and relationships from the standardized logistics management data and constructing a logistics knowledge graph using a graph database includes: The logistics entity design is based on the classification of the physical and chemical properties of logistics entities and the transportation and loading relationships between logistics goods; wherein, the classification of the physical and chemical properties of logistics entities includes at least one of fragile goods, toxic chemicals, refrigerated goods, frozen goods, liquid goods, granular and powder goods, easily oxidized goods, easily corrosive goods, and sensitive goods; the transportation and loading relationships include prohibited loading and cooperative loading. The standardized logistics management data is subjected to joint extraction of logistics entities and relationships using a pre-trained information extraction model, resulting in multiple logistics triples. The multiple logistics triples are instantiated and stored in a graph database to construct a two-layer basic knowledge graph. Entity alignment is performed on the basic knowledge graphs of different logistics management systems through graph structure modeling; wherein, the different logistics management systems include logistics management systems with different organizational structures, different regions, or different data sources; The entity alignment results are used to fuse the basic knowledge graphs of the different logistics management systems to construct a logistics knowledge graph.
4. The intelligent logistics sorting method based on dual database driving according to claim 3, characterized in that, The standardized logistics management data is subjected to joint extraction of logistics entities and relationships using a pre-trained information extraction model, resulting in multiple logistics triples, including: Based on the logistics ontology design, determine the types of logistics entities and relationships to be extracted, and construct prompt words; The logistics information text and the prompt words in the standardized logistics management data are input into a pre-trained information extraction model to jointly extract logistics entities and relationships, thereby obtaining multiple logistics triples corresponding to the standardized logistics management data.
5. The intelligent logistics sorting method based on dual database driving according to claim 1, characterized in that, The step involves extracting semantic features of each logistics entity and relationship in the logistics knowledge graph using a knowledge graph embedding method, obtaining hidden feature vectors for each logistics entity and relationship, and storing each hidden feature vector in a vector database, including: The structural semantics of the subgraph structure surrounding each logistics entity in the logistics knowledge graph are extracted using a graph neural network, and the relationships are incorporated into the calculation process of the structural semantics extraction based on the translation hypothesis to obtain the structural semantics embedding vector. The degree centrality feature vector of each logistics entity is determined using a webpage ranking algorithm; Random Gaussian noise is introduced into the degree centrality feature vector to simulate variable centrality features, and trainable variable centrality feature vectors are generated. The basic embedding vector, the degree centrality feature vector, and the variable centrality feature vector are respectively fused to obtain the hidden feature vectors corresponding to each of the logistics entities.
6. The intelligent logistics sorting method based on dual database driving according to claim 1, characterized in that, The method further includes: If querying the vector database fails, the pre-trained classifier is used to classify the physical and chemical properties of the goods to be sorted, and a preliminary classification result of the physical and chemical properties is obtained. The preliminary physicochemical property classification results are confirmed and corrected using a large language model to obtain the final physicochemical property classification results. The goods to be sorted are sorted according to the final physicochemical property classification results.
7. The intelligent logistics sorting method based on dual database driving according to claim 6, characterized in that, The method further includes: Update the logistics knowledge graph and the vector database based on the classification results of the physicochemical properties of the goods to be sorted or the final classification results of the physicochemical properties. If the preliminary physicochemical property classification results fail to predict, the pre-trained classifier is optimized.
8. The intelligent logistics sorting method based on dual database driving according to claim 3, characterized in that, The process of sorting the goods to be sorted according to the classification results of the physicochemical properties includes: Logistics sorting is carried out based on the classification results of the physical and chemical properties of the goods to be sorted, as well as the transportation and loading relationship between the goods to be sorted and other logistics goods.
9. A logistics intelligent sorting device based on dual databases, characterized in that, include: The data preprocessing module is used to acquire logistics management data and perform preprocessing and standardization to obtain standardized logistics management data. The knowledge graph construction module is used to jointly extract logistics entities and relationships from the standardized logistics management data, and construct a logistics knowledge graph through a graph database. The vector database construction module is used to extract the semantic features of each logistics entity and relationship in the logistics knowledge graph through the knowledge graph embedding method, obtain the hidden feature vectors of each logistics entity and relationship, and store each of the hidden feature vectors into the vector database. The logistics sorting module is used to query the vector database through similarity calculation to determine the classification results of the physical and chemical attributes of the goods to be sorted; and to perform logistics sorting on the goods to be sorted according to the classification results of the physical and chemical attributes.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the intelligent logistics sorting method based on dual database driving as described in any one of claims 1 to 8.