Logistics business efficient response processing system based on big data and big model
Through the efficient response processing system of logistics services of big data and large models, the hash mapping function and Euclidean distance matrix are used for data classification and dimensionality reduction, and combined with the distributed computing framework, the problem of inefficiency of traditional logistics services data processing systems is solved, and faster response speed and higher processing frequency are achieved.
Patent Information
- Application Number
- CN202510930286.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-08-19
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional logistics business data processing systems are inefficient in high concurrency and low latency applications, have weak analysis capabilities, and slow response speed.
The efficient response processing system for logistics services based on big data and large models is adopted. Through the data acquisition module, classification processing module, data processing module and storage module, the hash mapping function and Euclidean distance matrix are used to classify and reduce data dimensionality, combined with the MapReduce or Spark distributed computing framework, subprograms are dynamically allocated to accelerate data processing.
It improves the processing efficiency of logistics business data, enhances the system's analysis ability and response speed, optimizes file transfer efficiency, and solves the problem of low processing frequency in traditional systems.
Smart Images

Figure CN120512428A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing systems, and in particular to a logistics business efficient response processing system based on big data and big models. Background Art
[0002] With the rapid development of the logistics and transportation industry, the business data of various logistics business parties have also shown exponential growth. These data have become resources for the development of logistics companies. By conducting statistical analysis on them, corresponding decisions can be made to promote the development of the company.
[0003] In the era of big data, the scale of logistics business data is growing. Enterprises use processing systems to centrally process logistics business data. In some application systems that require high concurrency and low latency, this processing method can affect application system performance, resulting in low logistics business data processing efficiency. Summary of the Invention
[0004] The purpose of the present invention is to provide a logistics business efficient response processing system based on big data and big models to solve the problems raised in the above background technology.
[0005] To achieve the above object, the present invention provides the following technical solutions: An efficient logistics business response processing system based on big data and big models, including: A data acquisition module, wherein the data acquisition module is used to acquire logistics business data, which includes transportation information data, warehousing information data, and order information data; a classification processing module, the classification processing module being connected to the data acquisition module, and receiving and classifying the logistics business data, wherein transportation information data is assigned to the first category of data, warehousing information data is assigned to the second category of data, and order information data is assigned to the third category of data; The data processing module is connected to the classification processing module, and a preset program is set in the data processing module. The preset program includes three subroutines, which are used to independently calculate and process the first type of data, the second type of data and the third type of data.
[0006] Furthermore, it also includes a storage module, which is connected to the data processing module and the server and is used to store the data calculated and processed by the data processing module.
[0007] Furthermore, it also includes an interaction module, which is connected to the data processing module, the application end and the server respectively. The user submits the data processing task through the interaction module and obtains the data processing result of the data processing module.
[0008] Furthermore, the classification processing module receives and classifies the logistics business data, wherein the transportation information data is classified as the first category of data, the warehousing information data is classified as the second category of data, and the order information data is classified as the third category of data. Specifically, the classification processing module receives and classifies the logistics business data, wherein the transportation information data is classified as the first category of data, the warehousing information data is classified as the second category of data, and the order information data is classified as the third category of data. Classify the data types of transportation information data, warehousing information data, and order information data; Design hash mapping functions; Use hash mapping function to represent the relationship between classified data types; According to their mutual relationships, the classified data types are divided into first-category data, second-category data, and third-category data.
[0009] Furthermore, the three subroutines independently calculate and process the first type of data, the second type of data, and the third type of data as follows: Set the Euclidean distance and distance matrix between subroutines; When the value of the position corresponding to the data in the subroutine is less than the threshold, the distance matrix is merged to achieve dimensionality reduction; Set the two data clusters closest to each other in the program, merge the two data clusters, and the merged cluster is a new cluster; Set up distributed data processing subroutines based on the new class cluster.
[0010] Furthermore, the storage module is based on a distributed storage system of HDFS or object storage service, which is used to achieve data access and availability.
[0011] Furthermore, the data processing module adopts MapReduce or Spark distributed computing framework.
[0012] Furthermore, a management control center is preset in the data processing module, and the management control center is used to dynamically allocate subroutines according to new clusters.
[0013] Compared with the prior art, the present invention has the following beneficial effects: The present invention classifies the data types of transportation information data, warehousing information data and order information data; then designs a hash mapping function and uses the hash mapping function to represent the mutual relationship of the classified data types; and then allocates the classified data types into first-category data, second-category data and third-category data according to the mutual relationship.
[0014] This invention addresses the problem of low processing frequency in traditional systems, resulting from weak analytical capabilities and slow response times. This improves the processing efficiency of logistics business data.
[0015] This invention treats warehouse information data and order information data as smaller files to be transmitted, compared to transportation information data. Transportation information data is then treated as an oversized file. These smaller files can be consolidated and transmitted using a single channel (implemented at the distribution and queuing layers). This significantly improves file transmission efficiency. To reduce the transmission time associated with oversized files, this embodiment employs a fragmented transmission strategy. During transmission, oversized files are segmented into several smaller files according to the fragmentation strategy. These files are then collectively transmitted to the classification processing module using the distribution and queuing layers. This improves the efficiency of the data acquisition module in transmitting logistics data. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 This is a schematic diagram of the efficient response processing system for logistics business based on big data and big models of the present invention.
[0017] Figure 2 This is a schematic diagram of linear classification of the present invention.
[0018] Figure 3 This is a schematic diagram of the overall framework of file transmission of the data acquisition module of the present invention. DETAILED DESCRIPTION
[0019] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention. Example 1
[0020] See also Figures 1 to 2 , the present invention provides a technical solution: An efficient logistics business response processing system based on big data and big models, including: A data acquisition module, which is used to acquire logistics business data, including transportation information data, warehousing information data, and order information data; a classification processing module, the classification processing module being connected to the data acquisition module, and receiving and classifying the logistics business data, wherein transportation information data is assigned to the first category of data, warehousing information data is assigned to the second category of data, and order information data is assigned to the third category of data; The data processing module is connected to the classification processing module, and a preset program is set in the data processing module. The preset program includes three subroutines, which are used to independently calculate and process the first type of data, the second type of data and the third type of data.
[0021] In the present invention, transportation information data refers to relevant information data about the transportation of goods during the logistics process, including data such as the starting point, destination, means of transportation, transportation time, and transportation distance.
[0022] In the present invention, storage information data refers to relevant information data about goods storage during the logistics process, including data such as warehouse location, storage time, and storage costs.
[0023] In the present invention, order information data refers to information data related to goods orders during the logistics process, including order number, goods quantity, shipper, consignee and other data.
[0024] The classification processing module receives and classifies the logistics business data, wherein the transportation information data is classified as the first category data, the warehousing information data is classified as the second category data, and the order information data is classified as the third category data. Specifically, the classification processing module includes: Classify the data types of transportation information data, warehousing information data, and order information data; Design hash mapping functions; Use hash mapping function to represent the relationship between classified data types; According to the mutual relationship, the classified data types are allocated into first-category data, second-category data and third-category data.
[0025] Specifically, the above data types are classified as the first, second and third types of data streams, respectively, A ( x ), B ( x ), C ( x ), the three use 48-order programmable functions, then: (1) (2) (3) Where, a , b These are the data of the first and second types of data streams, respectively. The upper and lower subscripts represent the data order. x Represents the data in the third type of data stream. A ( x ), B ( x )and C ( x ) can be expressed as: A ( x ) B ( x )= F ( x )+C ( x )(4) in, F ( x )express A ( x ), B ( x )and C ( x ) is the hash mapping relationship between them, which is also the final classification scheme.
[0026] The three subroutines independently calculate and process the first, second, and third types of data as follows: Set the Euclidean distance and distance matrix between subroutines; When the value of the position corresponding to the data in the subroutine is less than the threshold, the distance matrix is merged to achieve dimensionality reduction; Set the two data clusters closest to each other in the program, merge the two data clusters, and the merged cluster is a new cluster; Set up distributed data processing subroutines based on the new class cluster.
[0027] The present invention utilizes a preset program in a data processing module, which includes at least three subroutines, thereby converting the calculation and processing of data into multiple subroutines, thereby speeding up the data processing frequency within the same processing time.
[0028] Specifically, the Euclidean distance between subroutines is preset, as expressed in the following formula: (5) Where: d uv Represents any two adjacent subroutines u and v The Euclidean distance between R us , R vs Indicates the program u , v No. s variables. Set the distance matrix at this time to K , determine the minimum distance element d min , when the data in the subroutine corresponds to the position a 1, b 1, and its value is less than the threshold Y When , the distance matrix is merged to achieve dimensionality reduction. Set the two data clusters closest to each other in the program as C a1 (s) , C b1(s) , the new cluster after merging is C a1b1 (s) ={ C a1 (s) , C b1 (s)}, so the new classification is C 1 (s) , C 2 (s) , … , C m (s) The processing system sets up distributed data processing subroutines based on the new clusters.
[0029] After the distributed processing mode is set up, the data is linearly classified so that the data in each processing unit belongs to the same type, or has the same purpose, or has similar value. Randomly select two sample data and divide them into two categories. k dimensional vector X Represents the corresponding data features, y Represents the corresponding classification mark. The calculation expression of the linear classification hyperplane and the classification function is: (6) in: w T Represents the transposed matrix established for the data type; B represents a fixed constant; f ( X ) represents the classification function. f ( X )>0, the corresponding classification mark y =1; when f ( X )<0, y =-1; when f ( X )=0, the support vector of the data is above the hyperplane. According to formula (6), the linear classification is as follows Figure 2 shown.
[0030] Figure 2 In the figure, the triangle and circle represent two randomly selected sample vectors. Set constraints and establish a high-speed processing function based on cloud computing technology. u , achieving instantaneous processing of large amounts of data.
[0031] (7) in:c is the total amount of data; ω is the error correction coefficient; f ( y ) is a processing constraint; k i To process the frequency i The limit value of the data segment; q To control the harmonic coefficient; t is the instantaneous reaction time; n is the processing path. According to the above formula (7), a large number of data processing programs are set to control the processing process of the subprograms. This realizes the design of a distributed processing system for large amounts of logistics business data.
[0032] The data processing module adopts a distributed computing framework such as MapReduce or Spark. A management control center is also preset in the data processing module, and the management control center is used to dynamically allocate subroutines according to new clusters.
[0033] In order to improve the dynamic allocation subroutine of the management control center and thus improve the performance of the distributed computing framework for data processing rate, the present invention optimizes the distributed computing framework. The specific method is as follows: Preset Dataset J ,Will J All data records in k If adjacent categories are the same, the adjacent arrays need to be merged into one interval. Then the best data segmentation point must be between different arrays. At this time, KU ( D ) can be defined as follows ( D is the set of optimal split points): (8) Pick D Any two split points in a 11 and a 22 , get the segmentation points through the continuous histogram a 11 and a 22 The corresponding Gini coefficient is expressed as follows: (9) (10) At this time, the dataset J The Gini index value is KU ( a 11 )and KU ( a21 ) between a 11 and a 22 There are one or more optimal split points outside the interval. For continuous attributes in the dataset, the maximum attribute value within the merged interval is taken as the candidate split point. The Gini coefficient value of each split point is calculated, and the split points are sorted one by one.
[0034] The present invention, J 1 and J 2 is a collection J Split into two parts and swap J 1 and J 2, the Gini index function KU ( D ) remains unchanged. Based on this, in order to reduce the complexity of the system algorithm and the amount of calculation for the split point screening, the distributed computing framework is optimized and improved: first, the data attribute type partitioning table is initialized, and the data set and J 2. J 1 Randomly pick i Data in a collection J 1, the total number of times delivered is C i n , at the same time, traverse J All values of 2. Determine this i If the value combination of the two sets is the same, the delivery is invalid; if they are different, the calculation starts KU ( D ) value. When the delivery is executed C i n Afterwards, judge i The value of n Value size, if i The value of n If the values are the same, the next loop is not executed. The Gini value obtained at this time is the minimum. After the above optimization, the parallel performance and execution rate of the distributed computing framework are improved, the loop calculation process is prevented from falling into the local optimal solution, and the excessive occupation of local memory is avoided, which reduces resource consumption. It effectively improves the dynamic allocation subroutine of the management and control center, thereby improving the performance of the distributed computing framework in terms of data processing rate.
[0035] The system also includes a storage module connected to the data processing module for storing data processed by the data processing module. The storage module is based on a distributed storage system such as HDFS or an object storage service to achieve data access and availability.
[0036] It also includes an interaction module, which is connected to the data processing module and the application end respectively. The user submits data processing tasks through the interaction module and obtains the data processing results of the data processing module.
[0037] The hardware device of the interactive module can be a host computer (PC), and the application end can be a smart phone, etc.
[0038] The present invention classifies the data types of transportation information data, warehousing information data and order information data; then designs a hash mapping function and uses the hash mapping function to represent the mutual relationship of the classified data types; and then allocates the classified data types into first-category data, second-category data and third-category data according to the mutual relationship.
[0039] The present invention, starting with data processing frequency, breaks down large amounts of data into several subroutines, accelerating the system's analytical capabilities and response speed. The system of the present invention solves the problem of low processing frequency caused by weak analytical capabilities and slow response speed in traditional systems. Example 2
[0040] See also Figure 3 The present invention provides a technical solution that is basically the same as Example 1, with the following slight differences: During system operation, the system needs to receive logistics business data in real time, which puts pressure on batch data transmission. When the aforementioned data processing module has a good data processing rate, it is necessary for the logistics business data to enter the data processing module in a timely manner. In this process, the data acquisition module must have high transmission efficiency, which is an essential function.
[0041] Furthermore, the file sizes (measured in bytes, hereinafter referred to as "file sizes") of transportation information data, warehousing information data, and order information data are different: Transport information data: Transport information data typically includes detailed information such as transport method, transport time, and transport company. This data typically exists in the form of text and a small number of numbers, occupying relatively little storage space. For example, the transport method (such as road, rail, or air) can be represented by a few characters.
[0042] Warehouse information data: Warehouse information data includes inventory data, inbound and outbound records, etc. This data usually includes a large number of entries and detailed records, and therefore takes up a large amount of storage space.
[0043] Order Information Data: Order information data includes customer orders, purchase orders, and return orders, and includes information such as product type, quantity, and delivery time. This data is primarily structured and includes customer information, product information, and transaction details. The size of order data depends on the number of orders and the level of detail in each order. For example, a customer order may include customer information, product information, and transaction details. This information is stored in both text and numeric form, taking up a lot of space.
[0044] Therefore, the file size of the transportation information data can be set to be smaller than the file size of the warehousing information data, and the file size of the warehousing information data can be set to be smaller than the file size of the order information data.
[0045] Therefore, the warehouse information data and order information data can be regarded as small files to be transmitted relative to the transportation information data, and the transportation information data can be regarded as an overly large file.
[0046] Based on this, the data acquisition module structure is improved to realize that the data acquisition module is used to increase the transmission rate of logistics business data to the classification processing module; The data acquisition module includes: The interface layer is used to access transportation information data, warehousing information data and order information data, and provides interface services for connecting to the server; The distribution and queuing layer, according to the transportation information data, storage information data and order information data accessed by the interface layer, classifies and packages the transportation information data, storage information data and order information data according to different sizes using a fragmented transmission strategy to obtain packaged data, and then transmits the packaged data to the classification processing module, specifically including: integrating the batch files to be transmitted into a set, using SS If it is expressed as follows: (11) ss i Indicates the size of the file to be transferred; N and ss N Respectively represent the maximum value of file transfer and the largest file. MM Transfer files of different sizes simultaneously and measure their respective transfer rates (in MB / s). VV Expressed as: (12) Since the size of the transferred file and the transfer efficiency are related to each other, vv N Indicates transmission MM The size is ss N The transfer rate of the files and the collection UU Expressed as: (13) When the transfer file size is ss i When the transmission efficiency is When the measured file transfer efficiency is at its maximum, it can be concluded that the file block size is exactly matched to the server buffer size, which can be expressed as: (14) Formula (14) cannot tell when the file transfer efficiency reaches its maximum value, so formula (15) is needed to express it as: (15) That is, the server buffer size at the data receiving end (data processing module end) Center buf Expressed as Center buf = ss i .
[0047] Set all files that need to be transferred to a collection FF express, JJ Indicates the maximum number of files transferred ff i Indicates the size of any unpackaged file in the collection. Set it to be larger than the collection FF The threshold value of any file size in KK ,but: (16) gather FF Expressed as: (17) The size of the packed file block is set PP Expressed as: (18) Among them, GG, pp GG,FF and pp i,FF They represent the maximum number of times all files are packed, the size of the packed block, and the size of the packed file. (19) The time to pack all files is used as a collection TT Expressed as: (20) in, Indicates that when the file block size is pp i,FFThe transmission time required. This indicates that the transmission time is the shortest, that is, the server buffer size is ss i , the file block size is pp i,FF When both and ss i ≤ pp i,FF The transmission rate reaches the fastest.
[0048] In order to maximize the file transfer rate (analyze the relationship between the server buffer size and the file size, and analyze the specific impact of the buffer size on the file transfer rate using the fragmented transmission strategy.) FF It can be expressed by formula (17) to get the exact value of the buffer, using the set BB Expressed as: (twenty one) HH Indicates the maximum number of transfer files that can be written into the buffer; bb HH and bb e Represents the maximum and minimum values of the buffer respectively. Assume that there are several buffers of different sizes transferring files of the same size at the same time. The time taken is QQ Indicates that the file is bb e The time it takes to transfer in the buffer; Indicates the time required to transfer the file after packaging.
[0049] (twenty two) when When the minimum value of the buffer is bb e In order to increase the transmission rate, the size of the packed file block is set to pp i,FF ,and pp i,FF The value of is equal to the size of the buffer. pp i,FF ≥ bb e and ss i ≤ pp i,FF These two conditions and ensure and Both equations hold true at the same time.
[0050] In this embodiment, to improve the efficiency of logistics data transmission within the data acquisition module, small files to be transmitted can be consolidated and transmitted using a single channel (implemented at the allocation and queuing layers). This significantly improves file transmission efficiency. To reduce the time required to transmit large files, this embodiment employs a fragmented transmission strategy. During transmission, oversized files are split into several smaller files according to the fragmentation strategy and collectively transmitted to the classification processing module using the allocation and queuing layers.
[0051] The present invention, the undescribed part is the prior art.
[0052] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. An efficient response processing system for logistics business based on big data and big models, characterized by: include: A data acquisition module, wherein the data acquisition module is used to acquire logistics business data, which includes transportation information data, warehousing information data, and order information data; a classification processing module, the classification processing module being connected to the data acquisition module, and receiving and classifying the logistics business data, wherein transportation information data is assigned to the first category of data, warehousing information data is assigned to the second category of data, and order information data is assigned to the third category of data; A data processing module, wherein the data processing module is connected to the classification processing module and has a preset program in the data processing module. The preset program includes three subroutines for independently calculating and processing the first category of data, the second category of data, and the third category of data; The data acquisition module is used to increase the transmission rate of logistics business data to the classification processing module, and the data acquisition module includes: Interface layer, which is used to access transportation information data, warehousing information data and order information data; The distribution and queuing layer classifies and packages the transportation information data, warehousing information data and order information data according to the transportation information data, warehousing information data and order information data accessed by the interface layer according to different sizes using a fragmented transmission strategy to obtain packaged data, and then transmits the packaged data to the classification processing module.
2. The logistics business efficient response processing system based on big data and big models according to claim 1 is characterized in that: It also includes a storage module, which is connected to the data processing module and the server and is used to store the data calculated and processed by the data processing module.
3. The logistics business efficient response processing system based on big data and big models according to claim 2 is characterized in that: It also includes an interaction module, which is connected to the data processing module, the application end and the server respectively. The user submits the data processing task through the interaction module and obtains the data processing result of the data processing module.
4. The logistics business efficient response processing system based on big data and big models according to claim 1 is characterized in that: The classification processing module receives and classifies the logistics business data, wherein the transportation information data is classified as the first category data, the warehousing information data is classified as the second category data, and the order information data is classified as the third category data. Specifically, the classification processing module includes: Classify the data types of transportation information data, warehousing information data, and order information data; Design hash mapping functions; Use hash mapping function to represent the relationship between classified data types; According to the mutual relationship, the classified data types are allocated into first-category data, second-category data and third-category data.
5. The logistics business efficient response processing system based on big data and big models according to claim 1 is characterized in that: The three subroutines independently calculate and process the first, second, and third types of data as follows: Set the Euclidean distance and distance matrix between subroutines; When the value of the position corresponding to the data in the subroutine is less than the threshold, the distance matrix is merged to achieve dimensionality reduction; Set the two data clusters closest to each other in the program, merge the two data clusters, and the merged cluster is a new cluster; Set up distributed data processing subroutines based on the new class cluster.
6. The logistics business efficient response processing system based on big data and big models according to claim 2 is characterized in that: The storage module is based on a distributed storage system of HDFS or object storage service, and is used to achieve data access and availability.
7. The logistics business efficient response processing system based on big data and big models according to claim 1 is characterized in that: The data processing module adopts the distributed computing framework of MapReduce or Spark.
8. The logistics business efficient response processing system based on big data and big models according to claim 5 is characterized in that: The data processing module is also preset with a management control center, which is used to dynamically allocate subroutines according to new clusters.
Citation Information
Patent Citations
Logistics information management system based on cloud computing
CN111639889A
Platform architecture based on computer big data
CN118228001A
KR20250093894A