Intelligent extraction system and method for cargo multimedia information
By collecting, scanning, and semantically bridging multimedia information about goods, the problem of integrating multimedia data in existing technologies has been solved. This has enabled automatic structured conversion of cargo information and accurate matching of transportation capacity, thereby improving the intelligence and efficiency of the logistics matching platform.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- 上海新颐科技软件股份有限公司
- Filing Date
- 2026-01-15
- Publication Date
- 2026-05-08
AI Technical Summary
Existing logistics information processing technologies cannot effectively integrate multimedia data, resulting in repetitive work in entering cargo information, response delays, and misjudgments. They also fail to achieve a deep match between cargo characteristics and driver carrying capacity, thus hindering the intelligent and precise development of logistics matching platforms.
By collecting cargo images, videos, and text annotations, and combining them with equipment positioning signals to generate location coordinate bindings, we perform content layered scanning and cross-semantic bridging to build an integrated semantic chain, extract structured attribute sets, generate driver query sequences, and perform adaptive sorting.
It enables the automatic integration and structured conversion of multimedia information, improves the accuracy of freight matching and the efficiency of transportation resource allocation, avoids repetitive manual processing and misjudgment, and achieves deep matching between cargo characteristics and driver carrying capacity.
Smart Images

Figure CN121526262B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of smart logistics technology, and more specifically, to a system and method for intelligent extraction of multimedia information from goods. Background Technology
[0002] With the deep integration of mobile internet and smart logistics, freight matching platforms have become core information hubs connecting shippers and carriers. In modern logistics services, the accurate collection and efficient transmission of cargo information plays a crucial role in achieving precise matching of supply and demand and optimizing the allocation of transportation resources. To improve the convenience and completeness of cargo information entry, mainstream logistics platforms generally support shippers to upload cargo-related information in multimedia formats such as taking photos, recording videos, and editing text notes. This multimedia data can intuitively record information such as the appearance, stacking status, packaging, and on-site environment of the cargo, providing original basis for subsequent capacity scheduling and transportation plan formulation.
[0003] Existing logistics information processing technologies have significant limitations in automatically converting multimedia content uploaded by shippers into structured freight data suitable for intelligent matching. This problem is particularly prominent in real-time freight matching scenarios that require rapid response. Specifically, when shippers simultaneously upload images, videos, and text descriptions containing cargo information, these data from different modalities differ significantly in their expression, information granularity, and semantic level. The system struggles to automatically identify and associate content fragments describing the same cargo attribute across different modalities, and it is even more unable to effectively integrate and structure key information such as cargo type, specifications, quantity characteristics, loading and unloading requirements, and transportation time constraints scattered across different media. Existing systems typically only store and display various multimedia files independently, relying on manual review and subsequent information aggregation and secondary entry. This processing mode not only causes a large amount of repetitive work and response delays, but also leads to the omission or misjudgment of key cargo characteristics due to subjective differences in human understanding. More importantly, due to the lack of technical capabilities to automatically extract structured freight elements from multimedia content, the platform cannot construct accurate order feature vectors based on the actual attributes of goods and transportation needs. This means that capacity matching can only rely on simple geographical proximity relationships, failing to achieve a deep adaptation between cargo characteristics and driver carrying capacity. Ultimately, this restricts the progress of logistics matching platforms towards intelligence and precision.
[0004] In view of this, the present invention proposes an intelligent system and method for extracting multimedia information from goods to solve the above problems. Summary of the Invention
[0005] To overcome the aforementioned deficiencies of the prior art and to achieve the above objectives, the present invention provides the following technical solution: a method for intelligent extraction of multimedia information from goods, comprising:
[0006] Step S1: Collect cargo images, video clips, and text annotations uploaded by the cargo owner, and combine them with the device positioning signals to generate location coordinate bindings, thus obtaining a multimedia location set;
[0007] Step S2: Perform content layer scanning on the images and video clips in the corresponding media location set, separate visual elements and dynamic sequences to form a hierarchical structure, and obtain layered content groups;
[0008] Step S3: Perform cross-semantic bridging on the hierarchical content groups and text annotations to construct the association path between elements and obtain the integrated semantic chain;
[0009] Step S4: Apply an attribute extraction loop to the integrated semantic chain to extract cargo specification details and transportation constraint fragments to build an attribute set, thus obtaining a structured attribute set;
[0010] Step S5: Generate a driver query sequence based on the structured attribute set and location coordinate binding, incorporate route-following factors for adaptation and sorting, and obtain a list of preferred drivers.
[0011] A cargo multimedia information intelligent extraction system, comprising:
[0012] Location module: Collects cargo images, video clips, and text labels uploaded by the cargo owner, and combines them with the device's positioning signals to generate location coordinate bindings, resulting in a multimedia location set;
[0013] Content parsing module: Performs content layer scanning on images and video clips within the corresponding media location set, separates visual elements and dynamic sequences to form a hierarchical structure, and obtains layered content groups;
[0014] Semantic bridging module: performs cross-semantic bridging on hierarchical content groups and text annotations, constructs the association path between elements, and obtains an integrated semantic chain;
[0015] Attribute extraction module: Apply an attribute extraction loop to the integrated semantic chain to extract cargo specification details and transportation constraint fragments to build an attribute set, resulting in a structured attribute set;
[0016] The matching recommendation module generates a driver query sequence based on the binding of structured attribute sets and location coordinates, incorporates route-following factors for matching and sorting, and obtains a list of preferred drivers.
[0017] The technical effects and advantages of the intelligent extraction system and method for multimedia information of goods according to the present invention are as follows:
[0018] This invention performs content-layered scanning on images and video clips within a multimedia location set, separating visual elements and dynamic sequences to form a hierarchical structure and obtain layered content groups. This hierarchically organizes the appearance features, stacking status, and dynamic changes of goods across different media carriers, reducing the interference of differences in information granularity between different modal data on subsequent semantic analysis. Furthermore, it constructs an integrated semantic chain by cross-semantically bridging the layered content groups and text annotations to establish inter-element association paths. Through cross-modal semantic association, it effectively connects content fragments describing the same goods attribute scattered across images, videos, and text, solving the problem of information isolation caused by differences in expression and semantic level between different modal data. Finally, it applies attributes to the integrated semantic chain. The system extracts loops, extracts cargo specification details and transportation constraint fragments to form an attribute set, resulting in a structured attribute set. This enables the automatic conversion of multimedia content into structured freight data that can be used for intelligent matching, avoiding repetitive work, response delays, and feature omissions or misjudgments caused by subjective differences due to manual review and secondary input. Based on the structured attribute set and location coordinates, a driver query sequence is generated and route proximity factors are incorporated for adaptation and sorting to obtain a list of preferred drivers. This allows capacity matching to be deeply adapted by comprehensively considering the actual attributes of the cargo, transportation needs, and driver trajectories, breaking through the limitations of matching based solely on simple geographical proximity relationships. This improves the accuracy of supply and demand matching and the efficiency of transportation resource allocation on the freight matching platform. Attached Figure Description
[0019] Figure 1 This is a schematic diagram of a method for intelligent extraction of multimedia information from goods according to the present invention;
[0020] Figure 2 This is a schematic diagram of a cargo multimedia information intelligent extraction system according to the present invention. Detailed Implementation
[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] Example 1
[0023] Please see Figure 1 As shown in the figure, this embodiment of a method for intelligent extraction of multimedia information of goods includes:
[0024] Step S1: Collect cargo images, video clips, and text labels uploaded by the cargo owner, and combine them with the device positioning signal to generate location coordinate binding, thus obtaining a multimedia location set.
[0025] In practical applications of freight matching platforms, cargo owners typically use smartphones or tablets to take photos and videos of their goods and input text descriptions. This multimedia data records information such as the appearance, stacking status, packaging, and on-site environment of the goods from different dimensions. However, because different modal data may differ in collection time and location, without establishing a unified spatiotemporal correlation, it will be impossible to accurately determine whether each media element describes the same batch of goods during subsequent information integration. Therefore, this step establishes location coordinate binding for multimedia data by fusing device positioning signals, laying the foundation for subsequent cross-modal information association.
[0026] Preferably, in some possible implementations of the embodiments of the present invention, the method for obtaining the multimedia location set includes:
[0027] Step S11: Receive cargo images and video clips from the cargo owner's device, simultaneously capture text annotation input, attach upload timing tags to organize media elements, and obtain the original multimedia package.
[0028] When a cargo owner uploads cargo-related information through the logistics platform application, the system's backend server receives cargo image files and video clips transmitted from the cargo owner's device in real time, while simultaneously capturing text annotations edited by the cargo owner in the input boxes. To ensure consistency in the timing of subsequent processing, the system attaches an upload time tag to each received media element. This tag is recorded using a Unix timestamp format, accurate to the millisecond level. For example, if a cargo owner uploads 3 cargo images, 15-second video clips, and a text description "200 steel pipes, each 6 meters long, requiring flatbed truck transport" at 10:15:32 AM, the system will attach time tags 1718234132001 to 1718234132005 to these 5 media elements, numbered sequentially according to the order of receipt, forming the original multimedia package.
[0029] By attaching upload time-series tags, the system can accurately record the acquisition order of each media element, providing a time-dimensional reference for subsequent judgment on whether different media elements describe the same batch of goods, effectively solving the problem that existing systems only store multimedia files independently and lack time-series correlation.
[0030] Step S12: The original multimedia package is fused with the device positioning signal, the binding position coordinates are calibrated by coordinate accuracy, and auxiliary anchor points are constructed with the help of device sensor data to obtain the coordinate enhancement package.
[0031] The equipment's positioning signal originates from the GPS module of the cargo owner's mobile device. In this embodiment, a combination of GPS and BeiDou dual-mode positioning is used to obtain the location coordinates. Since indoor environments or signal obstruction may cause a decrease in positioning accuracy, coordinate accuracy calibration is required.
[0032] Specifically, step S12 includes:
[0033] Step S121: Isolate cargo images, video clips, and text labels from the original multimedia package, integrate the device positioning signal, and record the integration time point to obtain the signal integration group.
[0034] The system extracts three types of media elements from the original multimedia package: cargo images, video clips, and text annotations. For each type of element, it integrates the device location signal corresponding to its upload time. The recording accuracy of the integrated time point is consistent with the upload time sequence tag, ensuring accurate time correspondence between location information and media elements. By precisely mapping the location signal to the media element's time, the system can establish a reliable geographic location identifier for each media element, avoiding location information mismatches caused by time deviations.
[0035] Step S122: Perform coordinate accuracy calibration on the signal integration group, compare the device sensor data and adjust the binding strength, correct the coordinate values through cross-verification of multiple signal sources, and obtain the anchor point reinforcement group.
[0036] The core of coordinate accuracy calibration lies in cross-validation using data from multiple built-in sensors. In this embodiment, the device's sensor data includes accelerometer data, gyroscope data, and barometer data. The system compares the changes in sensor data within 5 seconds before and after the positioning signal acquisition. If the displacement change detected by the accelerometer exceeds a preset threshold of 10 meters, it is determined that the positioning signal may have a drift error, and the binding strength is reduced to 70% of the baseline value. If the sensor data shows that the device is stationary and the coordinate difference between multiple positioning signals is less than 5 meters, the binding strength is increased to 120% of the baseline value. The preset threshold of 10 meters is based on the fact that in cargo loading and unloading scenarios, the normal activity range of the cargo owner in a short period of time usually does not exceed 10 meters, and positional changes exceeding this range may indicate abnormal positioning signals.
[0037] The method for cross-validation of multiple signal sources is as follows: GPS signals, BeiDou signals, and base station positioning signals are collected simultaneously, and the weighted average coordinates of the three are calculated as the corrected coordinate values. The weighting rules are as follows: the positioning source with the highest signal strength has a weight of 0.5, the second highest has a weight of 0.3, and the lowest has a weight of 0.2.
[0038] Through a multi-signal source cross-verification mechanism, the system can effectively eliminate occasional errors from a single positioning source, significantly improve the accuracy and reliability of location coordinates, and provide accurate spatial reference for subsequent geographic location-based capacity matching.
[0039] Step S123: Extend the element binding within the anchor point reinforcement group to adjacent media elements, construct a coordinate chain connection, and obtain a chain binding group.
[0040] For media elements with high binding strength within the anchor point reinforcement group, the system extends their position coordinates to temporally adjacent media elements. The extension rule is: if the upload interval between adjacent media elements is less than 30 seconds and their own binding strength is less than 80% of the baseline value, then the position coordinates of the higher-strength element are used for supplementary binding. The 30-second time interval threshold is based on the scenario of cargo information collection, where the time interval between multiple photos or videos taken by the cargo owner of the same batch of goods is usually no more than 30 seconds. Media elements exceeding this time interval may describe goods in different locations.
[0041] The chain connection mechanism allows media elements with lower positioning accuracy to supplement their positioning with the position information of adjacent high-precision elements, effectively solving the problem of positioning loss caused by signal blockage of some media elements and ensuring that all media elements have usable position coordinates.
[0042] Step S124: Apply positional mutation scan to the chained binding group, add mutation compensation markers, and convert it into a coordinate augmentation package.
[0043] Positional variation scanning is used to detect changes in the positional coordinates of adjacent media elements within a chained group. If the positional coordinate distance between adjacent elements exceeds 50 meters, the system determines that a positional variation exists and adds a variation compensation marker to the subsequent element. This marker records the variation direction and distance, which is used in subsequent steps to identify different batches of goods. The variation distance threshold of 50 meters is based on the fact that the typical distance between different goods stacking areas in a large logistics park or factory is approximately 50 to 200 meters. Positional changes of less than 50 meters are usually considered normal movement within the same stacking area.
[0044] The location variation scanning mechanism enables the system to automatically identify cargo information taken by the cargo owner at different locations, avoiding the incorrect merging of information from different batches of cargo and improving the accuracy of subsequent information integration.
[0045] Step S13: Divide the video segments in the corresponding coordinate enhancement package into time axis markers, bind the associated position coordinates and embed the mutation tracking points to obtain the temporal position group.
[0046] Compared to still images, video clips contain continuous information in the time dimension, and different moments within a video clip may correspond to different shooting positions. This step divides the video clip into timeline markers at fixed time intervals; in this embodiment, 2 seconds is used as the division interval. For example, a 15-second video is divided into 8 timeline markers, corresponding to video frames at seconds 0, 2, 4, 6, 8, 10, 12, and 14.
[0047] For each timeline marker, the system associates its corresponding position coordinates with changes in the device's positioning signal during video recording. If the position coordinates of adjacent timeline markers change significantly (i.e., the distance exceeds a preset threshold of 10 meters), a variation tracking point is embedded at that time point. The variation tracking point records the start and end coordinates of the position change and the rate of change, which is used for subsequent analysis of the spatial distribution of goods in the video.
[0048] Through timeline marking and mutation tracking point mechanisms, the system can finely break down continuous information in video clips, accurately linking content at different times within the video with corresponding geographical locations, and fully mining the spatiotemporal information value of video data.
[0049] Step S14: Merge the temporal location group with the cargo image and text label, expand the binding range to cover the positional variations between elements, and obtain a multimedia location set.
[0050] The system fuses the temporal location group obtained in step S13 with the cargo images and text annotations in the coordinate enhancement package. During the fusion process, for media elements with location variations, the system expands their binding range so that the binding area covers the entire interval from the start position to the end position. Each element in the multimedia location set carries location coordinate binding information, temporal tags, and variation markers, providing a unified spatiotemporal reference framework for subsequent content layering scanning and semantic bridging.
[0051] The construction of the multimedia location set enables unified spatiotemporal management of images, videos, and text annotations uploaded by cargo owners, solving the problem of lack of correlation between different modal data in the existing system and laying a solid foundation for subsequent cross-modal information integration.
[0052] Step S2: Perform content layer scanning on the images and video clips in the corresponding media location set, separate visual elements and dynamic sequences to form a hierarchical structure, and obtain layered content groups.
[0053] The images and videos uploaded by cargo owners contain rich visual information, but this information exists in an unstructured form and needs to be transformed into a hierarchical structure that can be analyzed through content layered scanning. The core idea of layered scanning is to decompose complex visual content into multiple levels of visual elements, which facilitates semantic bridging with subsequent text annotations. Without layered processing, the system will have difficulty identifying specific cargo objects and their attribute characteristics in images or videos, thus failing to achieve the technical goal of automatically extracting structured freight elements from multimedia content.
[0054] Preferably, in some possible implementations of the embodiments of the present invention, the method for obtaining hierarchical content groups includes:
[0055] Step S21: Select cargo images from the multimedia location set, perform content layer scanning to separate static visual elements, split the outline texture and background components, and obtain image layers.
[0056] This embodiment employs deep learning-based image segmentation technology to perform content-layered scanning of cargo images. The scanning process separates the image content into three layers: a foreground cargo layer, a mid-ground auxiliary layer, and a background environment layer. The foreground cargo layer contains the outline, texture, and color features of the cargo; the mid-ground auxiliary layer contains surrounding packaging materials, pallets, forklifts, and other auxiliary items; and the background environment layer contains environmental components such as warehouse walls, floors, and ceilings.
[0057] Contour segmentation uses an edge detection algorithm to extract the outline of the goods, while texture segmentation uses a local binary mode algorithm to extract the texture feature vectors of the goods' surface. For example, for an image of stacked steel pipes, the system separates the cylindrical outline of the steel pipes as the contour component, the rust marks and reflective features on the surface of the steel pipes as the texture component, and the concrete floor and metal shelves of the warehouse as the background component.
[0058] Through layered scanning technology, the system can accurately extract cargo-related visual elements from complex image content, filter out background information unrelated to the cargo, and significantly improve the accuracy and efficiency of subsequent attribute recognition.
[0059] Step S22: Apply dynamic sequence segmentation to the video segments within the corresponding multimedia location set, capture motion trajectories and layer dynamic elements, and connect the sequences through inter-frame transition markers to obtain video layering.
[0060] The video clip contains time-series information. This step uses dynamic sequence segmentation technology to capture moving elements in the video. The system analyzes the video frame by frame, detecting differences between frames and identifying the trajectories of moving objects. Dynamic sequence segmentation divides the video content into four layers: static background layer, slowly changing layer, fast motion layer, and human activity layer.
[0061] Inter-frame transition markers are used to record moments of scene changes or camera movement in a video. When the pixel difference rate between consecutive frames exceeds a preset threshold of 40%, the system determines that a scene change has occurred and inserts a transition marker. The 40% pixel difference rate threshold is based on the fact that the inter-frame difference caused by normal camera shake or changes in lighting is usually below 40%. Changes exceeding this percentage indicate that the camera is focused on a different subject or that a scene change has occurred.
[0062] The dynamic sequence segmentation mechanism enables the system to make full use of the temporal dimension information of the video, identify the movement state of the goods and the loading and unloading process, extract dynamic features that cannot be presented by static images, and enrich the descriptive dimensions of the goods information.
[0063] Step S23: Align elements in the image layer and video layer, build cross-layer connections and set alignment anchor points to obtain the aligned layer.
[0064] Since the same batch of goods may appear in both images and videos, it is necessary to align similar visual elements in the image and video layers. Element alignment uses a feature matching algorithm to calculate the feature similarity between the outline components in the image and the corresponding components in the video frame. When the similarity exceeds a preset threshold of 0.75, the system determines that the two elements describe the same goods object, establishes a cross-layer connection, and sets alignment anchor points.
[0065] The similarity threshold of 0.75 is based on the differences in shooting angle and lighting conditions. The visual feature similarity of the same goods in different media carriers is usually between 0.7 and 0.9. Using 0.75 as a threshold can ensure matching accuracy while avoiding missing valid matches.
[0066] The element alignment mechanism enables the automatic association of the same goods objects in images and videos, allowing the integration and analysis of goods information scattered across different media carriers. This solves the problem that existing systems cannot identify and associate fragments of content describing the same goods attribute in different modal data.
[0067] Step S24: Bind the alignment layer to the position coordinates and extend the connection to the temporal position group element to obtain the extended layer.
[0068] The system aligns each visual element in the hierarchical alignment layer with the location coordinates of its source media file. For dynamic elements in video clips, the system associates them with corresponding location coordinates based on their timeline markers, achieving a precise correspondence between visual elements and geographical locations.
[0069] By binding visual elements with location coordinates, the system can determine the geographical location of each identified cargo object, providing a spatial reference for subsequent location-based capacity matching and achieving deep integration of cargo characteristics and geographic information.
[0070] Step S25: Perform element clustering loop for the corresponding extended layer, merge visual elements and attach clustering labels to obtain layered content groups.
[0071] The element clustering loop is used to merge multiple visual elements describing the same goods object into a single cluster unit. The clustering algorithm employs a density-based spatial clustering method, grouping visual elements whose feature vector distance is less than a preset threshold and whose positional coordinate distance is less than 50 meters into the same cluster. After clustering is complete, the system attaches a cluster label to each cluster unit, which includes the cluster number, the number of elements contained, and the coordinates of the cluster center.
[0072] The element clustering mechanism integrates similar visual elements from different image and video frames to form a comprehensive description of the goods object, avoiding redundant processing caused by information fragmentation and improving the efficiency of subsequent semantic bridging and attribute extraction. The hierarchical content group, as the output of this step, provides a structured set of visual elements for subsequent semantic bridging.
[0073] Step S3: Perform cross-semantic bridging on the hierarchical content groups and text annotations to construct the association path between elements and obtain the integrated semantic chain.
[0074] Visual elements and text labels in the hierarchical content group describe cargo information from the visual and linguistic dimensions, respectively. However, these two modalities differ significantly in their expression and semantic level. There is no direct correlation between the text label "200 steel pipes" and the outline of the steel pipes in the image, making it difficult for the system to automatically determine whether they describe the same cargo. Therefore, this step establishes a semantic association path between visual elements and text labels through cross-semantic bridging technology, achieving effective integration of cross-modal information and solving the core problem in existing technologies where systems struggle to automatically identify and associate content fragments describing the same cargo attribute across different modalities.
[0075] Preferably, in some possible implementations of the embodiments of the present invention, the method for obtaining the integrated semantic chain includes:
[0076] Step S31: Extract visual elements from the hierarchical content group, perform semantic bridging with text annotations to generate preliminary association paths, and attach bridging nodes to obtain a set of bridging paths.
[0077] The system extracts visual elements of the foreground cargo layer from the hierarchical content group and uses a visual semantic encoder to convert them into semantic vector representations. Simultaneously, natural language processing techniques are used for word segmentation and entity recognition of the text annotations to extract key entities such as cargo name, specifications, and quantity.
[0078] Semantic bridging employs a cross-modal similarity calculation method, calculating the cosine similarity between the semantic vectors of visual elements and text entities. When the similarity exceeds a preset threshold of 0.6, the system determines that a semantic association exists, generates a preliminary association path, and adds a bridging node at the midpoint of the path. The bridging node records the visual element number, text entity content, and similarity value.
[0079] The similarity threshold of 0.6 is based on the fact that cross-modal semantic matching has an inherent semantic gap compared to single-modal matching, and the similarity is usually lower than that of single-modal matching. Using 0.6 as a threshold can ensure that most valid associations are captured while controlling the false match rate.
[0080] The semantic bridging mechanism breaks down the modal barriers between visual and textual information, enabling the system to automatically identify the correspondence between goods objects in images and goods names in text descriptions, laying the foundation for subsequent attribute integration.
[0081] Step S32: Apply cross-validation loop to the bridging path set, check the semantic consistency between elements and set the bridging strength value, expand the validation range, and obtain the validation path set.
[0082] The cross-validation loop is used to verify the reliability of the initial association path and quantify the bridging strength. Specifically, step S32 includes:
[0083] Step S321: Select a path from the bridging path set, introduce text annotations as reference anchors, perform semantic consistency checks and record the checked nodes to obtain the checked path group.
[0084] The system sequentially selects each associated path from the bridging path set, using the text annotation entities connected at both ends of the path as reference anchors. The semantic consistency check method verifies whether the visual attributes of a visual element match the semantic attributes of the text entity. For example, if a text annotation contains the description "red" while the dominant color of a visual element is blue, a semantic inconsistency is determined, and a check node is recorded on that path with the inconsistency type marked.
[0085] The semantic consistency check mechanism can identify and mark contradictions between visual features and textual descriptions, avoid misjudgments caused by incorrect associations, and improve the accuracy of cross-modal information integration.
[0086] Step S322: Set bridging strength values for the inspection path group, dynamically adjust the strength based on the distance between elements and semantic overlap, apply the strength propagation mechanism, and obtain the strength-enhanced path set.
[0087] The bridging strength value is calculated by comprehensively considering two factors: inter-element distance and semantic overlap. Inter-element distance refers to the geographical distance between the location coordinates of the visual element and the location where the text annotation was uploaded; semantic overlap refers to the degree of overlap between the semantic features contained in the visual element and the attributes described by the text entity.
[0088] The formula for calculating the bridging strength value is: In the formula, This represents the bridging strength value. The geographical distance between elements is expressed in meters. As a distance normalization parameter, this embodiment uses an empirical value of 200 meters, which means that the correlation strength between elements with a distance of more than 200 meters is reduced to zero; This represents the semantic overlap, with a value ranging from 0 to 1. and The weighting coefficients are 0.4 and 0.6, respectively, based on empirical values. This means that semantic overlap has a greater weight in the bridging strength calculation because the same batch of goods may be photographed separately, with geographical distances but close semantic connections.
[0089] The intensity propagation mechanism is used to transfer the credibility of a high-intensity bridging path to its adjacent low-intensity paths. The propagation rule is: if two paths share the same visual element or text entity, the intensity value of the low-intensity path is determined according to the formula... Enhancement, among which For path enhancement value, This represents the intensity value of adjacent high-intensity paths.
[0090] By quantifying the bridging strength value and using a strength propagation mechanism, the system can accurately assess the reliability of each associated path, prioritize the retention of high-reliability associations, effectively filter noisy associations, and improve the overall quality of semantic bridging.
[0091] Step S323: Perform branch scans on the paths within the strength enhancement path set, apply mutation compensation, and transform them into a validation path set.
[0092] Branch scanning is used to identify multi-branch path structures that originate from the same visual element and connect to multiple text entities. For multi-branch paths, the system analyzes the bridging strength distribution of each branch. If there are branches with significantly different strengths, a variation compensation tag is added to the low-strength branches, indicating that the association should be treated with caution in subsequent processing. The branch scanning mechanism enables the system to identify complex one-to-many or many-to-one association structures, accurately handling situations where an image contains multiple goods or multiple text descriptions correspond to the same goods, enhancing the flexibility and adaptability of semantic bridging.
[0093] Step S324: Perform a path compression loop on the corresponding verification path set, merge adjacent nodes and attach compression labels.
[0094] The path compression loop simplifies the structure of the validation path set. For adjacent paths with similar bridging strength values and semantically similar connected objects, the system merges them into a single composite path, reducing the computational complexity of subsequent processing. The merged path is labeled with a compression tag, recording the number of original paths and their strength value ranges. This path compression mechanism simplifies the semantic network structure while preserving key association information, reducing the computational overhead of subsequent processing and improving the overall processing efficiency of the system.
[0095] Step S33: Construct an integrated semantic chain based on the verification path set, expand the path to cover multi-level element associations and embed branch points within the chain.
[0096] The system sorts all paths in the verification path set in descending order of bridging strength value, and gradually expands to build an integrated semantic chain starting from the path with the highest strength. The chain expansion rule is: each time, the system selects the unadded path with the highest strength value that shares elements with the current chain endpoint node and is not yet included in the chain for connection. The integrated semantic chain covers multiple levels of visual elements and multiple entities in text annotations within the hierarchical content group, forming a complete cross-modal semantic association network. For nodes with multiple association directions in the chain, the system embeds branch points within the chain, marking the branch direction and the strength distribution of each branch. The integrated semantic chain integrates cargo information scattered in images, videos, and text into a unified semantic network, achieving deep fusion of multimodal data and providing a complete information source for subsequent structured attribute extraction. This fundamentally solves the problem that existing technologies cannot effectively integrate key cargo information scattered across different media carriers.
[0097] Step S4: Apply an attribute extraction loop to the integrated semantic chain to extract cargo specification details and transportation constraint fragments to build an attribute set, thus obtaining a structured attribute set.
[0098] The integrated semantic chain establishes semantic relationships between visual elements and text annotations, but the information in the chain still exists in its raw form and has not been transformed into structured freight data that can be used for capacity matching. This step extracts cargo specification details and transportation constraint fragments from the integrated semantic chain through attribute extraction loops, transforming unstructured multimedia information into a structured attribute set containing fields such as cargo type, size specifications, quantity characteristics, loading and unloading requirements, and transportation timeliness constraints. This addresses the problem that existing technologies lack the ability to automatically extract structured freight elements from multimedia content.
[0099] Preferably, in some possible implementations of the embodiments of the present invention, the method for obtaining the structured attribute set includes:
[0100] Step S41: Start the attribute extraction loop from the integrated semantic chain, scan the cargo specification details to break down the size, material and quantity components, and obtain specification fragment groups.
[0101] The attribute extraction process iterates through each node in the integrated semantic chain, applying different attribute extraction strategies to visual element nodes and text entity nodes. For visual element nodes, the system estimates the dimensions of the goods using image measurement technology and determines the material category using a material recognition model. For text entity nodes, the system extracts dimension values, material names, and quantity information using regular expressions and entity recognition technology.
[0102] The extraction method for the size component is as follows: detect the bounding box of the goods from visual elements and estimate the actual size by combining it with the size of a reference object. If the text label contains a clear size description, such as "6 meters long each," the size value in the text label is used first. The extraction method for the material component is as follows: determine the material category through the texture features of visual elements, including five major categories: metal, wood, plastic, textiles, and paper products. The extraction method for the quantity component is as follows: extract quantity words and combinations of quantity words from the text label.
[0103] By processing attribute information from both visual and textual modalities separately and performing cross-validation, the system can obtain a more accurate and complete description of cargo specifications, effectively compensating for potential omissions or ambiguities in information from a single modality.
[0104] Step S42: Extract transportation constraint segments from the specification segment group, and fuse time-sensitive loading and unloading and route markings to construct constraint sub-chains, thus obtaining the constraint extension group.
[0105] Transportation constraints are categorized into three types: timeliness constraints, loading / unloading constraints, and route constraints. Timeliness constraints extract delivery time requirements from textual annotations, such as "delivery before 3 PM tomorrow" or "within 48 hours." Loading / unloading constraints extract loading / unloading method requirements from visual elements and textual annotations, such as visual elements indicating the need for mechanical loading / unloading when a forklift or lifting equipment is identified, and the textual description "flatbed truck required" indicating vehicle type constraints. Route constraints extract origin, destination, and waypoint information from location coordinate bindings and textual annotations. Constraint sub-chains link multiple constraint fragments related to the same goods to form a complete transportation constraint description. For example, the constraint sub-chain for goods A might include: origin coordinates → timeliness constraint (delivery within 24 hours) → loading / unloading constraint (forklift loading / unloading required) → destination coordinates.
[0106] The automatic extraction of transportation constraints enables the system to fully understand the transportation needs of goods, providing complete constraints for subsequent accurate capacity matching and avoiding matching failures or transportation disputes caused by missing constraint information.
[0107] Step S43: Merge the specification fragment group and the constraint extension group, establish links between attributes and set the fusion node to obtain the fused attribute group.
[0108] The system integrates the size, material, and quantity components in the specification fragment group with the timeliness, loading / unloading, and route constraints in the constraint extension group. The integration process establishes logical links between attributes; for example, the link between the size component and the loading / unloading constraint indicates that large-sized goods require special loading / unloading equipment. The integration node records the linked attribute component number and the link type. This attribute integration mechanism establishes a logical association between cargo specifications and transportation constraints, enabling the system to understand the interrelationships between different attributes and providing richer contextual information for subsequent intelligent matching decisions.
[0109] Step S44: Perform attribute priority sorting on the corresponding fusion attribute group, adjust the fragment weights and attach sorting anchors to obtain the sorted attribute group.
[0110] Different attributes have varying degrees of importance for capacity matching. This step prioritizes the attribute components in the fusion attribute group according to their importance. The ranking rule is: location coordinates (weight 0.30) > time constraints (weight 0.25) > cargo type (weight 0.20) > size specifications (weight 0.15) > loading and unloading requirements (weight 0.10). The weight allocation is based on the following: capacity matching first requires meeting geographical accessibility, then considering the feasibility of time constraints, and finally considering the compatibility between cargo and vehicles.
[0111] The attribute priority sorting mechanism ensures that the subsequent matching process considers various constraints in the correct priority order, avoids secondary conditions from having an excessive impact on the matching results, and improves the rationality and success rate of capacity matching.
[0112] Step S45: Bind the sorting attribute group to the end of the integrated semantic chain and extend it to the position coordinate binding to obtain the extended attribute group.
[0113] The system binds sorted attribute groups to integrated semantic chains, making each attribute component traceable to its source visual element or text entity. Simultaneously, extended attribute groups inherit location coordinate binding information, ensuring that structured attributes remain associated with geographic locations.
[0114] The attribute traceability mechanism enables the system to trace the original source of each attribute value when needed, making it easier for cargo owners or drivers to verify the accuracy of attribute information and enhancing the credibility and transparency of the information.
[0115] Step S46: Apply an attribute validation loop to the extended attribute group, check link consistency and attach validation tags to obtain a structured attribute set.
[0116] The attribute validation loop checks the logical consistency between attribute components within the extended attribute group. Validation rules include: whether the volume calculations for size and quantity exceed the loading capacity of common vehicle models; whether the origin-destination distance and time constraints are within reasonable limits; and whether loading / unloading requirements match the cargo type. Attribute components that pass validation are labeled with a validation pass tag, while attribute components with potential conflicts are labeled with a validation warning tag.
[0117] The attribute verification mechanism can automatically identify logical contradictions or irrationalities in cargo information, detect potential problems in advance and issue warnings, reducing matching failures and subsequent disputes caused by information errors. The structured attribute set, as the output of this step, contains complete cargo specification details and fragments of transportation constraints. Each attribute field is verified and carries priority weights and verification tags, and can be directly used for subsequent driver queries and capacity matching.
[0118] Step S5: Generate a driver query sequence based on the structured attribute set and location coordinate binding, incorporate route-following factors for adaptation and sorting, and obtain a list of preferred drivers.
[0119] The structured attribute set provides a complete feature description of the goods. This step generates a driver query sequence based on these features and intelligently sorts them using a route proximity factor to achieve a deep fit between goods characteristics and driver carrying capacity. The route proximity factor considers the spatial relationship between the driver's current location, planned route, and the origin and destination of the goods, prioritizing drivers with the shortest detour distance. This improves the utilization efficiency of transportation resources and solves the problem in existing technologies where capacity matching relies solely on simple geographical proximity and cannot achieve a deep fit between goods characteristics and driver carrying capacity.
[0120] Preferably, in some possible implementations of the embodiments of the present invention, the preferred method for obtaining the driver list includes:
[0121] Step S51: Extract cargo specification details and transportation constraint fragments from the structured attribute set, bind them with location coordinates to generate an initial driver query sequence, attach attribute labels, and obtain a sequence draft.
[0122] The system extracts key attributes such as cargo type, size specifications, quantity characteristics, loading and unloading requirements, and transportation time constraints from a structured attribute set, and generates an initial driver query sequence by combining these with the location coordinates of the cargo's origin and destination. The query sequence uses a structured query statement format, including filtering conditions and sorting criteria.
[0123] For example, for a freight demand of "200 steel pipes, each 6 meters long, requiring flatbed truck transport, to be delivered from Beijing to Shanghai within 24 hours," the system generates a query sequence that includes: vehicle type (flatbed truck), load capacity (estimated total weight based on steel pipe specifications), current location (within 200 kilometers of the origin), and timeliness (transportation can be completed within 24 hours). Each filter condition in the query sequence is accompanied by an attribute label indicating its source attribute and priority weight.
[0124] The mechanism of automatically generating query sequences based on structured attributes eliminates the tedious process of manually reviewing each item, summarizing information, and re-entering it, significantly improving the efficiency of freight information processing and reducing errors and delays caused by manual operation.
[0125] Step S52: Incorporate the path alignment factor into the sequence draft, simulate the driver's trajectory and embed the path alignment matching points, expand the trajectory branches, and obtain the factor-enhanced sequence.
[0126] The route proximity factor is the core innovation of this step, used to identify the spatial overlap between the cargo transportation route and the driver's existing transportation plan. For drivers who are performing other transportation tasks or planning to return, if the origin and destination of the cargo happen to be near their driving route, the detour cost for that driver to accept the order is lower, and they should be given a higher matching priority.
[0127] Specifically, step S52 includes:
[0128] Step S521: Select a query sequence from the sequence draft, introduce the path alignment factor as the trajectory simulation input, construct the simulation sub-path, and obtain the simulation trajectory group.
[0129] The system retrieves the current location and planned route information of candidate drivers from the platform database. For each candidate driver, the system constructs a simulated sub-path from the current location, through the origin and destination of the goods, to the driver's destination. The simulated sub-path is generated using a road network shortest path algorithm, taking into account the differences in traffic capacity among highways, national roads, and provincial roads.
[0130] By simulating sub-path construction, the system can predict the actual driving route of each driver after accepting an order, providing accurate route information for subsequent assessment of route convenience.
[0131] Step S522: Embed matching points along the route into the simulated trajectory group, bind the dynamic expansion factor range based on the position coordinates, and apply the matching point diffusion to obtain the expanded trajectory group.
[0132] A route matching point is the closest point between the simulated sub-path and the origin and destination of the goods. The system calculates the shortest distance between the simulated sub-path of each candidate driver and the origin and destination of the goods, and marks the path point corresponding to the shortest distance as a route matching point.
[0133] The matching point diffusion mechanism is used to handle situations where multiple feasible transfer points exist near the origin and destination of goods. The system searches for other feasible loading and unloading locations within a 30-kilometer radius of the nearest matching point, adding these locations to the extended trajectory group. The 30-kilometer diffusion radius is based on the fact that the average distance between loading and unloading sites within a city is approximately 20 to 40 kilometers, and this 30-kilometer radius covers most alternative loading and unloading points.
[0134] The matching point diffusion mechanism increases matching flexibility, enabling the system to consider multiple feasible loading and unloading locations, providing cargo owners and drivers with more options and improving the matching success rate.
[0135] Step S523: Branch and integrate the trajectories within the extended trajectory group, set integration nodes and record branch mutations to obtain the integrated trajectory group.
[0136] For goods with multiple feasible loading and unloading points, the system integrates the simulated sub-paths corresponding to each loading and unloading point to construct an integrated trajectory group containing multiple optional paths. The integrated node records the length of each branch path, the estimated travel time, and information on toll stations along the way.
[0137] The branch integration mechanism enables the system to comprehensively evaluate multiple feasible options and recommend the optimal loading and unloading locations and routes for cargo owners and drivers, thereby achieving overall optimization of the transportation plan.
[0138] Step S524: Execute the factor enhancement loop for the corresponding integrated trajectory group, adjust the weight of the follow-the-path factor and attach enhancement labels to obtain the enhanced trajectory group.
[0139] The factor reinforcement loop is used to quantify the route affinity of each candidate driver and adjust its matching weight. The formula for calculating the route affinity factor weight is: In the formula, This is the weight of the route factor, with a value ranging from 0 to 1. The larger the value, the higher the degree of route compatibility. This refers to the total distance the driver travels after picking up the goods, which is the path length from the current location through the origin and destination of the goods to the driver's destination. This represents the original planned driving distance when the driver does not accept cargo, i.e., the path length from the current location directly to the driver's destination. Less than or equal to When this is the case, it indicates that the goods are being accepted along the entire route or require only a very minor detour. The value is close to or equal to 1; when Significantly greater than At that time, it indicated that receiving the goods required a significant detour. The value approaches 0. The system adds a reinforcement label to drivers with high route affinity factors and gives them priority in subsequent rankings.
[0140] The quantitative calculation of the route factor weight enables the system to accurately assess the detour cost of each driver accepting an order, and prioritize recommending drivers with the shortest detours, effectively reducing transportation costs and improving the utilization efficiency of transportation resources.
[0141] Step S525: Transform the enhanced trajectory group into a factor enhancement sequence and bind it to a structured attribute set element.
[0142] The system encapsulates the simulated trajectory information and route-following factor weights of each candidate driver in the enhanced trajectory group into a factor enhancement sequence, and binds it to the corresponding attribute elements in the structured attribute set. The binding relationship records the degree of matching between cargo attributes and driver carrying capacity.
[0143] The binding mechanism between factor-enhanced sequences and attribute sets enables a deep correlation between cargo characteristics and driver capabilities, providing comprehensive information support for integrated matching assessment.
[0144] Step S53: Perform adaptive sorting based on the factor-enhanced sequence, calculate path overlap and adjust ranking, attach sorting tags, and obtain a sorted sequence group.
[0145] The matching and ranking process comprehensively considers three dimensions: cargo attribute matching degree, route proximity factor weight, and driver historical service rating. Path overlap is calculated as the percentage of common road segment length between the driver's planned route and the cargo transportation route.
[0146] The overall matching score is calculated by weighting the attribute matching degree, the route-following factor weight, and the historical service score, with weights of 0.4, 0.35, and 0.25, respectively. The weighting is based on the following: attribute matching degree ensures that the vehicle can carry goods and is a basic condition; the route-following factor weight affects transportation costs and drivers' willingness to accept orders; and the historical service score reflects the driver's service quality and reliability.
[0147] The system sorts candidate drivers in descending order of their comprehensive matching scores, generating a sorted sequence group. Each driver entry in the sorted sequence group is marked with a sorting tag, recording its ranking position and detailed scores for each dimension. The top 10 drivers in the sorted sequence group form a preferred driver list, which is then sent to cargo owners for selection.
[0148] The multi-dimensional comprehensive ranking mechanism achieves a deep match between cargo characteristics and driver carrying capacity. Compared with the traditional matching method that only relies on geographical proximity, it can more accurately identify the most suitable driver to take on specific cargo, significantly improving the accuracy of matching and the satisfaction of both parties.
[0149] It should be noted that the number of drivers included in the preferred driver list can be dynamically adjusted based on the urgency of the freight demand and the number of candidate drivers. For urgent orders, the system can expand the push to the top 20 drivers to increase the probability of a successful transaction; for regular orders, pushing the top 10 drivers can strike a balance between matching quality and response efficiency.
[0150] This embodiment collects multimedia data uploaded by cargo owners and establishes location coordinate bindings. It performs layered content scanning of images and videos, and cross-semantically bridges this data with text annotations. This enables the automatic extraction of structured freight data from unstructured multimedia content, solving the problem that existing systems can only independently store and simply display various multimedia files, relying on manual information aggregation. The structured attribute set contains cargo specification details and transportation constraint fragments, which can be directly used for intelligent capacity matching, avoiding the omission or misjudgment of key cargo features due to subjective differences in human understanding. The adaptation and sorting mechanism incorporating route proximity factors achieves deep matching of cargo characteristics and driver carrying capacity. Compared to traditional matching methods that rely solely on geographical proximity, this significantly improves the matching accuracy and capacity resource utilization efficiency of the logistics matching platform, powerfully promoting the evolution of the logistics matching platform towards intelligence and precision.
[0151] Example 2
[0152] Please see Figure 2 As shown, parts not described in detail in this embodiment are described in Embodiment 1. A cargo multimedia information intelligent extraction system is provided, including:
[0153] Location module: Collects cargo images, video clips, and text labels uploaded by the cargo owner, and combines them with the device's positioning signals to generate location coordinate bindings, resulting in a multimedia location set;
[0154] Content parsing module: Performs content layer scanning on images and video clips within the corresponding media location set, separates visual elements and dynamic sequences to form a hierarchical structure, and obtains layered content groups;
[0155] Semantic bridging module: performs cross-semantic bridging on hierarchical content groups and text annotations, constructs the association path between elements, and obtains an integrated semantic chain;
[0156] Attribute extraction module: Apply an attribute extraction loop to the integrated semantic chain to extract cargo specification details and transportation constraint fragments to build an attribute set, resulting in a structured attribute set;
[0157] The matching recommendation module generates a driver query sequence based on the binding of structured attribute sets and location coordinates, incorporates route-following factors for matching and sorting, and obtains a list of preferred drivers.
[0158] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for intelligent extraction of multimedia information from goods, characterized in that, include: Step S1: Collect cargo images, video clips, and text annotations uploaded by the cargo owner, and combine them with the device positioning signals to generate location coordinate bindings, thus obtaining a multimedia location set; Step S2: Perform content layer scanning on the images and video clips in the corresponding media location set, separate visual elements and dynamic sequences to form a hierarchical structure, and obtain layered content groups; Step S3: Perform cross-semantic bridging between hierarchical content groups and text annotations to construct inter-element association paths and obtain an integrated semantic chain. This includes: extracting visual elements from hierarchical content groups, semantically bridging them with text annotations to generate preliminary association paths, and attaching bridging nodes to obtain a set of bridging paths. The semantic bridging uses a cross-modal similarity calculation method to calculate the cosine similarity between the semantic vector of the visual element and the semantic vector of the text entity. When the similarity exceeds a preset threshold, the system determines that there is a semantic association, generates a preliminary association path, and attaches a bridging node at the midpoint of the path. The bridging node records the visual element number, text entity content, and similarity value. Apply cross-validation loops to the bridging path set, check the semantic consistency between elements and set the bridging strength value, expand the validation scope, and obtain the validation path set, including: selecting paths from the bridging path set, introducing text annotations as reference anchors, performing semantic consistency checks and recording the checked nodes, and obtaining the checked path group. A bridging strength value is set for each inspection path group, and the strength is dynamically adjusted based on the distance between elements and semantic overlap. An intensity propagation mechanism is applied to obtain a set of enhanced paths. This mechanism transfers the credibility of high-intensity bridging paths to their adjacent low-intensity paths. The propagation rule is: if two paths share the same visual element or text entity, the intensity value of the low-intensity path is calculated according to the formula... Enhancement, among which For path enhancement value, This represents the bridging strength value. The intensity value is the value of the adjacent high-intensity path; Branch scanning is performed on the paths within the strength enhancement path set, mutation compensation is added, and the path set is transformed into a validation path set. Perform path compression loop on the verification path set, merging adjacent nodes and appending compression labels; An integrated semantic chain is built based on the verification path set, and the path is extended to cover multi-level element associations and embed branch points within the chain. Step S4: Apply an attribute extraction loop to the integrated semantic chain to extract cargo specification details and transportation constraint fragments to build an attribute set, thus obtaining a structured attribute set; Step S5: Generate a driver query sequence based on the structured attribute set and location coordinate binding, incorporate route-following factors for adaptation and sorting, and obtain a list of preferred drivers.
2. The intelligent extraction method for multimedia information of goods according to claim 1, characterized in that, Step S1 includes: Step S11: Receive cargo images and video clips from the cargo owner's equipment, simultaneously capture text annotation input, attach upload timing tags to organize media elements, and obtain the original multimedia package; Step S12: Integrate the device positioning signal into the original multimedia package, bind the position coordinates through coordinate accuracy calibration, and construct auxiliary anchor points with the help of device sensor data to obtain a coordinate enhancement package, including: isolating cargo images, video clips and text labels from the original multimedia package, integrating the device positioning signal and recording the integration time point to obtain a signal integration group; The coordinate accuracy of the signal integration group is calibrated, the data of the device sensors are compared and the binding strength is adjusted, and the coordinate values are corrected by cross-verification of multiple signal sources to obtain the anchor point reinforcement group; Extend the binding of elements within the anchor point reinforcement group to adjacent media elements to construct a coordinate chain connection, resulting in a chain binding group. The method for constructing the coordinate chain connection includes: for media elements with high binding strength within the anchor point reinforcement group, extend their position coordinates to temporally adjacent media elements; the extension binding rule is: if the upload time interval between adjacent media elements is less than 30 seconds and their own binding strength is less than 80% of the baseline value, then the position coordinates of the high-strength element are used for supplementary binding. Apply positional mutation scans to chained binding groups, add mutation compensation markers, and convert them into coordinate augmentation packages; Step S13: Divide the video segments within the corresponding coordinate enhancement package into timeline markers, associate and bind the associated position coordinates, and embed mutation tracking points to obtain a temporal position group; wherein, the method of embedding mutation tracking points includes: for each timeline marker point, the system associates the corresponding position coordinates according to the changes in the device positioning signal during video recording; if the position coordinates between adjacent timeline marker points change significantly, i.e., the distance exceeds a preset value, then a mutation tracking point is embedded at that time point; the mutation tracking point records the start coordinates, end coordinates, and change rate of the position change; Step S14: Merge the temporal location group with the cargo image and text label, expand the binding range to cover the positional variations between elements, and obtain a multimedia location set.
3. The intelligent extraction method for multimedia information of goods according to claim 2, characterized in that, Step S2 includes: Step S21: Select cargo images from the multimedia location set, perform content layer scanning to separate static visual elements, split the outline texture and background components, and obtain image layers; Step S22: Apply dynamic sequence segmentation to the video segments within the corresponding multimedia location set, capture motion trajectories and layer dynamic elements, and connect the sequences through inter-frame transition markers to obtain video layering. Step S23: Align elements in the image layer and video layer, construct cross-layer connections and set alignment anchor points to obtain the aligned layer; Step S24: Bind the alignment layer to the position coordinates and extend the connection to the temporal position group element to obtain the extended layer; Step S25: Perform element clustering loop for the corresponding extended layer, merge visual elements and attach clustering labels to obtain layered content groups.
4. The intelligent extraction method for multimedia information of goods according to claim 3, characterized in that, Step S4 includes: Step S41: Start the attribute extraction loop from the integrated semantic chain, scan the cargo specification details to break down the size, material and quantity components, and obtain specification fragment groups; Step S42: Extract transportation constraint segments from the specification segment group, and fuse time-sensitive loading and unloading and route markings to construct constraint sub-chains to obtain the constraint extension group; Step S43: Merge the specification fragment group and the constraint extension group, establish links between attributes and set the fusion node to obtain the fused attribute group; Step S44: Perform attribute priority sorting on the corresponding fusion attribute group, adjust the fragment weights and attach sorting anchors to obtain the sorted attribute group; Step S45: Bind the sorting attribute group to the end of the integrated semantic chain and extend it to the position coordinate binding to obtain the extended attribute group; Step S46: Apply an attribute validation loop to the extended attribute group, check link consistency and attach validation tags to obtain a structured attribute set.
5. The intelligent extraction method for multimedia information of goods according to claim 4, characterized in that, Step S5 includes: Step S51: Extract cargo specification details and transportation constraint fragments from the structured attribute set, bind them with location coordinates to generate an initial driver query sequence, attach attribute labels, and obtain a sequence draft; Step S52: Incorporate the path alignment factor into the draft sequence, simulate the driver's trajectory and embed the path alignment matching points, expand the trajectory branches, and obtain the factor-enhanced sequence; Step S53: Perform adaptive sorting based on the factor-enhanced sequence, calculate path overlap and adjust ranking, attach sorting tags, and obtain a sorted sequence group.
6. The intelligent extraction method for multimedia information of goods according to claim 5, characterized in that, Step S52 includes: Step S521: Select a query sequence from the sequence draft, introduce the path alignment factor as the trajectory simulation input, construct the simulation sub-path, and obtain the simulation trajectory group; Step S522: Embed matching points along the route into the simulated trajectory group, bind the dynamic expansion factor range based on the position coordinates, and apply the matching point diffusion to obtain the expanded trajectory group; Step S523: Branch and integrate the trajectories within the extended trajectory group, set integration nodes and record branch mutations to obtain the integrated trajectory group; Step S524: Execute the factor enhancement loop for the corresponding integrated trajectory group, adjust the weight of the along-path factor and attach enhancement labels to obtain the enhanced trajectory group; Step S525: Transform the enhanced trajectory group into a factor enhancement sequence and bind it to a structured attribute set element.
7. A cargo multimedia information intelligent extraction system, used to implement the cargo multimedia information intelligent extraction method according to any one of claims 1 to 6, characterized in that, include: Location module: Collects cargo images, video clips, and text labels uploaded by the cargo owner, and combines them with the device's positioning signals to generate location coordinate bindings, resulting in a multimedia location set; Content parsing module: Performs content layer scanning on images and video clips within the corresponding media location set, separates visual elements and dynamic sequences to form a hierarchical structure, and obtains layered content groups; Semantic bridging module: This module performs cross-semantic bridging between hierarchical content groups and text annotations to construct inter-element association paths and obtain an integrated semantic chain. This includes: extracting visual elements from hierarchical content groups, semantically bridging them with text annotations to generate preliminary association paths, and attaching bridging nodes to obtain a set of bridging paths. The semantic bridging uses a cross-modal similarity calculation method to calculate the cosine similarity between the semantic vectors of visual elements and the semantic vectors of text entities. When the similarity exceeds a preset threshold, the system determines that there is a semantic association, generates a preliminary association path, and attaches a bridging node at the midpoint of the path. The bridging node records the visual element number, text entity content, and similarity value. Apply cross-validation loops to the bridging path set, check the semantic consistency between elements and set the bridging strength value, expand the validation scope, and obtain the validation path set, including: selecting paths from the bridging path set, introducing text annotations as reference anchors, performing semantic consistency checks and recording the checked nodes, and obtaining the checked path group. A bridging strength value is set for each inspection path group, and the strength is dynamically adjusted based on the distance between elements and semantic overlap. An intensity propagation mechanism is applied to obtain a set of enhanced paths. This mechanism transfers the credibility of high-intensity bridging paths to their adjacent low-intensity paths. The propagation rule is as follows: if two paths share the same visual element or text entity, the intensity value of the low-intensity path is calculated according to the formula... Enhancement, among which For path enhancement value, This represents the bridging strength value. The intensity value is the value of the adjacent high-intensity path; Branch scanning is performed on the paths within the strength enhancement path set, mutation compensation is added, and the path set is transformed into a validation path set. Perform path compression loop on the verification path set, merging adjacent nodes and appending compression labels; An integrated semantic chain is built based on the verification path set, and the path is extended to cover multi-level element associations and embed branch points within the chain. Attribute extraction module: Apply an attribute extraction loop to the integrated semantic chain to extract cargo specification details and transportation constraint fragments to build an attribute set, resulting in a structured attribute set; The matching recommendation module generates a driver query sequence based on the binding of structured attribute sets and location coordinates, incorporates route-following factors for matching and sorting, and obtains a list of preferred drivers.
Citation Information
Patent Citations
System and method for managing whole course of entire vehicle logistics transportation
CN107679814A
Method and system for displaying video map dynamic tag based on mobile terminal
CN109063039A