Multimodal Cargo Flow Intelligent Visual Analysis System and Method
By constructing a 3D freight scene and adding labels to multimodal data, and combining image algorithms to identify key frames, the problem of poor multimodal data fusion was solved, enabling multi-dimensional analysis and efficient data processing, reducing system load, and improving analysis efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- XIAN AERONAUTICAL UNIV
- Filing Date
- 2026-02-02
- Publication Date
- 2026-04-21
AI Technical Summary
Existing freight flow monitoring and analysis technologies suffer from poor multimodal data fusion, resulting in severe data silos, making multi-dimensional collaborative analysis impossible, and increasing the data processing volume, thus increasing system load and operating costs.
By constructing a 3D freight scene, static and dynamic entity models are generated, and spatial, entity, time, and event labels are added to the multimodal data. Image algorithms are used to identify keyframes, enabling the binding and rendering of multimodal data with 3D models, reducing computational load, and improving the accuracy and efficiency of data fusion.
It achieves accurate correlation and 3D visualization of multimodal data, reduces system computation, improves synchronous response efficiency and analysis efficiency, and reduces data processing load.
Smart Images

Figure CN121616177B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of freight flow technology, and in particular to a multimodal intelligent visual analysis system and method for freight flow. Background Technology
[0002] With the rapid development of the logistics industry, the scale of freight flow is constantly expanding, and freight scenarios are becoming increasingly complex (such as port-road intermodal transport, multimodal transport, etc.), placing higher demands on the real-time monitoring, accurate analysis, and efficient management of freight flow. Current freight flow monitoring and analysis technologies mainly rely on single-modal data, such as video surveillance or completed waybills. However, neither of these methods can meet the needs of practical applications. Although some freight flow analysis methods incorporate multimodal data, shortcomings still exist.
[0003] For example, there is a problem of poor multimodal data fusion. In existing technologies, video stream data, sensor data, and text data are mostly stored and processed independently, lacking a unified correlation mechanism. This leads to severe data silos, making multi-dimensional collaborative analysis impossible. For instance, GPS location data of freight vehicles cannot be accurately matched with video surveillance data of corresponding road segments, making it difficult to trace the entire process of abnormal vehicle movement. Fusion processing, on the other hand, increases the data volume, leading to a heavier system processing load, thus increasing operating costs and reducing analysis efficiency.
[0004] Therefore, there is an urgent need for a multimodal freight flow intelligent visual analysis technology that can achieve efficient fusion of multimodal data and accurate 3D visualization to address the shortcomings of existing technologies. Summary of the Invention
[0005] The purpose of this invention is to provide a multimodal intelligent visual analysis system and method for freight flow to solve the problems of the traditional freight flow data being presented in a single way and having a large amount of data processing. It has the advantages of comprehensive and accurate visualization, convenient interaction, and high efficiency.
[0006] Firstly, to achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0007] A multimodal intelligent visual analysis method for freight flow includes the following steps:
[0008] A three-dimensional freight scenario is constructed based on the road network data of freight flow, and a corresponding three-dimensional model is generated based on the original freight data. The generation of the corresponding three-dimensional model based on the original freight data includes generating a three-dimensional model based on the size parameters of the physical devices involved in logistics, and the three-dimensional model is divided into a static physical model and a dynamic physical model.
[0009] Acquire multimodal data from the freight flow, including video stream data, sensor data, and text data; add spatial labels, entity labels, time labels, and event labels to the multimodal data; and assign values to each label.
[0010] All 3D models are bound to various tags and then rendered. The binding and rendering methods are as follows:
[0011] The multimodal data is bound to all three-dimensional models using the spatial labels and the entity labels;
[0012] Acquire keyframes from the video stream data, set the keyframes as anchor points, set the rendering time according to the time tags corresponding to the anchor points, and set the rendering content of the dynamic entity model according to the event tags.
[0013] The rendered content is displayed through an interactive device.
[0014] Preferably, the method for adding spatial labels, entity labels, time labels, and event labels to the multimodal data is as follows:
[0015] The addition of spatial labels includes: establishing the reference coordinates of each 3D model in the 3D freight scene, acquiring the location data of the entity device associated with the multimodal data in real time, converting the location data into spatial coordinates based on the corresponding reference coordinates, and using the converted spatial coordinates as spatial labels.
[0016] The addition of entity labels includes: assigning an entity ID to each entity device, adding the entity ID to the multimodal data associated with the entity device, adding a classification symbol, with each classification symbol corresponding to a multimodal data category, and using the entity ID with the added classification symbol as the entity label;
[0017] Adding time stamps includes: establishing a time baseline, obtaining the upload time of multimodal data in real time, converting the upload time according to the time baseline to obtain a unified timestamp, and adding a timestamp to each multimodal data as a time stamp;
[0018] Adding event tags includes: setting judgment rules for business events and abnormal events, pre-setting business thresholds and abnormal thresholds for multimodal data, judging multimodal data exceeding the business threshold as business completion status, judging multimodal data exceeding the abnormal threshold as business abnormal status, and adding corresponding statuses as event tags to multimodal data.
[0019] Preferably, the method for obtaining keyframes of the video stream data includes: pre-setting feature frames, which are used as a reference group for input image algorithms; identifying each frame image in the video stream data through image algorithms; and outputting the frame image as a keyframe when the frame image is detected to match the reference group.
[0020] Preferably, the method further includes setting motion constraint parameters for the dynamic entity model, including thresholds for movement trajectory and loading / unloading position, and generating a warning signal for the 3D model that exceeds the motion constraint parameters and displaying it in the 3D freight scene.
[0021] Preferably, the method for constructing a three-dimensional freight scene based on road network data of freight flow includes:
[0022] Acquire vector data, which includes the latitude and longitude coordinates of freight routes and import / export geographical boundaries;
[0023] The vector data is converted to three-dimensional spatial coordinates and stitched together to form a global geographic base map.
[0024] Preferably, the method for acquiring multimodal data in freight flow is as follows:
[0025] The raw bitstream of the container loading and unloading video stream captured by the camera is used as the video stream data;
[0026] The container's GPS data and temperature and humidity sensor data are used as sensor data.
[0027] The shipping information in the backend is retrieved as text data using the shipping tracking number.
[0028] Preferably, the method further includes generating a floating window in the three-dimensional freight scene when receiving an instruction from the user to select a three-dimensional model, and displaying video footage and text from the multimodal data corresponding to the three-dimensional model in the floating window.
[0029] Secondly, to achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0030] A multimodal intelligent visual analysis system for freight flow includes:
[0031] A 3D scene construction unit is used to construct a 3D freight scene based on the road network data of freight flow and generate a corresponding 3D model based on the original freight data. The generation of the corresponding 3D model based on the original freight data includes generating a 3D model based on the size parameters of the physical devices involved in logistics and dividing the 3D model into a static physical model and a dynamic physical model.
[0032] A multimodal data processing unit is used to acquire multimodal data in freight flow, including video stream data, sensor data and text data, add spatial labels, entity labels, time labels and event labels to the multimodal data, and assign values to each label;
[0033] The binding and rendering unit is used to bind and render all 3D models with various tags. The binding and rendering method is as follows:
[0034] The multimodal data is bound to all three-dimensional models using the spatial labels and the entity labels;
[0035] Acquire keyframes from the video stream data, set the keyframes as anchor points, set the rendering time according to the time tags corresponding to the anchor points, and set the rendering content of the dynamic entity model according to the event tags.
[0036] The display unit is used to display the rendered content through an interactive device.
[0037] Preferably, the method for adding spatial labels, entity labels, time labels, and event labels to the multimodal data in the multimodal data processing unit is as follows:
[0038] The addition of spatial labels includes: establishing the reference coordinates of each 3D model in the 3D freight scene, acquiring the location data of the entity device associated with the multimodal data in real time, converting the location data into spatial coordinates based on the corresponding reference coordinates, and using the converted spatial coordinates as spatial labels.
[0039] The addition of entity labels includes: assigning an entity ID to each entity device, adding the entity ID to the multimodal data associated with the entity device, adding a classification symbol, with each classification symbol corresponding to a multimodal data category, and using the entity ID with the added classification symbol as the entity label;
[0040] Adding time stamps includes: establishing a time baseline, obtaining the upload time of multimodal data in real time, converting the upload time according to the time baseline to obtain a unified timestamp, and adding a timestamp to each multimodal data as a time stamp;
[0041] Adding event tags includes: setting judgment rules for business events and abnormal events, pre-setting business thresholds and abnormal thresholds for multimodal data, judging multimodal data exceeding the business threshold as business completion status, judging multimodal data exceeding the abnormal threshold as business abnormal status, and adding corresponding statuses as event tags to multimodal data.
[0042] Thirdly, to achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0043] A storage medium having a computer program stored thereon, which, when executed by a processor, implements the multimodal freight flow intelligent visual analysis method as described in the first aspect.
[0044] Compared with the prior art, the present invention has the following beneficial effects:
[0045] 1. This invention visualizes the status of freight flow by generating a 3D freight scene. Furthermore, by adding spatial, entity, temporal, and event tags to multimodal data, it achieves precise correlation between video streams, sensor data, text data, and the 3D model, ensuring the accuracy of data fusion. Moreover, the 3D model is generated only through size parameters, and its content is displayed through event tags, eliminating the need for in-depth processing of the multimodal data. Although the amount of multimodal data is large, this reduces the system's computational load.
[0046] 2. This invention automatically identifies key frames in video using image algorithms, and uses key frames as anchor points to achieve rapid association with other multimodal data, which greatly shortens the location time of business events or abnormal events, thereby improving synchronous response efficiency. Attached Figure Description
[0047] Figure 1 This is a flowchart of the multimodal freight flow intelligent visual analysis method of the present invention.
[0048] Figure 2 This is a schematic diagram of the multimodal freight flow intelligent visual analysis system of the present invention. Detailed Implementation
[0049] To make the technical means, creative features, objectives and effects of this invention easier to understand, the invention will be further described below in conjunction with specific embodiments.
[0050] Example 1
[0051] like Figure 1 As shown, the intelligent visual analysis method for multimodal freight flow includes the following steps:
[0052] Step S1: Construct a three-dimensional freight scene based on the road network data of the freight flow, and generate a corresponding three-dimensional model based on the original freight data. The generation of the corresponding three-dimensional model based on the original freight data includes generating a three-dimensional model based on the size parameters of the physical devices involved in the logistics. The three-dimensional model is divided into a static physical model and a dynamic physical model. The three-dimensional freight scene is a 1:1 virtual scene generated according to the entire freight flow. In order to ensure the accuracy of the scale, road network data must first be obtained. The road network data is vector data. Taking the GIS vector data of a port and surrounding highways as an example, the road network data needs to include the freight route: from port A to highway hub B, with a latitude and longitude range of 120.56°-120.58° east longitude and 30.12°-30.14° north latitude. The geographical boundary of the port's entrance and exit is 120.565° east longitude and 30.125° north latitude as the boundary starting point, and the latitude and longitude coordinate sequence in a clockwise direction.
[0053] The method for constructing a three-dimensional freight scene based on road network data of freight flow includes:
[0054] Acquire vector data, which includes the latitude and longitude coordinates of freight routes and import / export geographical boundaries;
[0055] The vector data is converted to three-dimensional spatial coordinates and then stitched together to form a global geographic base map. As we already know, road network data is GIS vector data. Using the formula for converting latitude and longitude to Mercator 3D coordinates, the acquired road network vector data is converted into three-dimensional spatial coordinates. The Mercator coordinate conversion formula is as follows:
[0056] , ;
[0057] Where R is the Earth's radius, λ is longitude, and φ is latitude. Taking the port's import / export boundary point (120.565°E, 30.125°N) as an example, the conversion process is as follows: Latitude and longitude to radians: λ = 120.565° × π / 180 ≈ 2.104 rad; φ = 30.125° × π / 180 ≈ 0.526 rad; Calculate x and y coordinates: x = 6378137 × 2.104 ≈ 13420000 m; y = 6378137 × ln(tan(π / 4 + 0.526 / 2)) ≈ 6378137 × ln(tan(0.963)) ≈ 6378137 × 0.482 ≈ 3074000 m. The converted three-dimensional spatial coordinates are then stitched together according to the hierarchy of the entire road network - operating area - key nodes to form a global geographic base map.
[0058] Once the 3D freight scene is determined, generating the 3D model becomes simple and quick. Measuring equipment can be used to measure the actual dimensions of the physical structures, in a length × width × height format. For example, a gantry crane is 50 × 10 × 30 m, a warehouse wall is 100 × 20 × 8 m, a container is 12 × 2.5 × 2.5 m, and a freight truck is 12 × 2.5 × 3.5 m. These dimensions are then converted to 3D spatial coordinates using the aforementioned conversion method to generate the corresponding 3D models. The gantry crane and warehouse wall are defined as static solid models, while the container and freight truck are defined as dynamic solid models. To make the 3D model more realistic, texture mapping can be added to colorize the 3D freight scene. The final 3D freight scene can reproduce the actual freight flow at a 1:1 scale.
[0059] Step S2: Acquire multimodal data from the freight flow. The multimodal data includes video stream data, sensor data, and text data. Add spatial labels, entity labels, time labels, and event labels to the multimodal data and assign values to each label. Acquiring multimodal data is similar to acquiring raw freight data, both using a data acquisition terminal. The method for acquiring multimodal data from the freight flow is as follows:
[0060] The raw bitstream of the container loading and unloading video stream captured by the camera is used as the video stream data;
[0061] The container's GPS data and temperature and humidity sensor data are used as sensor data.
[0062] The shipping information in the backend is retrieved as text data using the shipping tracking number.
[0063] Using spatial tags, entity tags, time tags, and event tags can ensure the consistency of data from video streams, sensors, and text, facilitating accurate association with subsequent 3D models. This eliminates the need to calculate multimodal data for each entity device; by binding the 3D model to each tag, synchronization between the virtual 3D model and the real entity device can be achieved, reducing computational load.
[0064] The method for adding spatial labels, entity labels, time labels, and event labels to the multimodal data is as follows:
[0065] The addition of spatial labels includes: establishing the reference coordinates of each 3D model in the 3D freight scene, acquiring the location data of the entity device associated with the multimodal data in real time, converting the location data into spatial coordinates based on the corresponding reference coordinates, and using the converted spatial coordinates as spatial labels; taking container number J-01 as an example, obtaining the initial position of container J-01 in reality, calculating the reference coordinates based on it, and calculating the reference coordinates (13421000, 3074500, 5.2) using the Mercator transformation formula in step S1. If container J-01 is constantly moving, then obtain the GPS position of container J-01 at 10:00:00 (longitude 120.566°, latitude 30.127°), and convert it to spatial coordinates (13421050, 3074580, 5.2). These coordinates are the spatial tags, labeled as the base coordinates (13421000, 3074500, 5.2) and the current coordinates (13421050, 3074580, 5.2). Bind the video stream data, sensor data, and text data to the above spatial tags.
[0066] Adding entity tags involves: assigning an entity ID to each entity device, adding the entity ID to the multimodal data associated with the entity device, and adding a classification symbol. Each classification symbol corresponds to a multimodal data category, and the entity ID with the added classification symbol serves as the entity tag. Again, taking container J-01 as an example, classification symbols are added to the multimodal data: video stream data (V), sensor data (D), and text data (E). Example entity tags representing the container: video stream data tag "J-01-V", sensor data tag "J-01-D", and text data tag "J-01-E".
[0067] Adding time stamps involves: establishing a time baseline, acquiring the upload time of multimodal data in real time, converting the upload time according to the time baseline to obtain a unified timestamp, and adding a timestamp to each multimodal data point as a time stamp; converting the local upload time of the sensor / camera (e.g., camera local time 10:00:00.123) to a UTC timestamp using the formula: Timestamp = UTC Base Time (seconds) + Milliseconds / 1000. For example, the corresponding UTC time for local time is calculated to be 1746060000.123, which is then assigned as a time stamp to the video stream data, sensor data, and text data. A unified time reduces the generation of errors.
[0068] Adding event tags includes: setting judgment rules for business events and abnormal events, pre-setting business thresholds and abnormal thresholds for multimodal data, judging multimodal data exceeding the business threshold as business completion status, judging multimodal data exceeding the abnormal threshold as business abnormal status, and adding corresponding statuses as event tags to multimodal data. Business events and abnormal events correspond to two states. Taking a container as an example, its temperature and humidity abnormality thresholds can be set as follows: temperature [5-28℃], humidity [20%-60%RH]. If the temperature is >28℃ or <5℃, and the humidity is >60%RH or <20%RH, then the container's temperature and humidity can be determined to be in an abnormal state, and the container can be marked as a business abnormal state. The business threshold for a business event is the change in the container's height coordinate. For example, a container lifting height ≥2m is considered as loading / unloading in progress, and a height ≤0.5m is considered as loading / unloading completed. For example, at 10:00:00, the temperature of J-01 is 23.5℃ and the humidity is 45%RH, which does not exceed the abnormal threshold. The lifting height is 1.2m, which does not exceed the business threshold. The event label is assigned as "Transporting - Normal". Abnormal example: At 10:10:00, the temperature and humidity sensor collects a temperature of 29.2℃, exceeding the abnormal threshold. The event label is assigned as "Transporting - Temperature and Humidity Abnormal". Event tags allow for real-time viewing of the operational status of physical devices in a 3D scene, eliminating the need to collect and process all data for identification, thus greatly improving identification efficiency. Moreover, this clear display can intuitively show the status of freight flow.
[0069] Step S3: Bind and render all 3D models to their respective tags. The binding and rendering methods are as follows:
[0070] Step S31: Bind the multimodal data to all 3D models using the spatial tags and entity tags. The binding method involves constructing a tag-model association index. This index can be built using a Redis cluster, which is existing technology and its principles will not be elaborated here. The final data format names in the index are, for example, J-01-V, J-01-D, and J-01-E. These names indicate that the data is associated with the J-01 container, representing video stream data, sensor data, and text data.
[0071] Step S32: Obtain keyframes from the video stream data, set the keyframes as anchor points, set the rendering time according to the time tag corresponding to the anchor point, and set the rendering content of the dynamic entity model according to the event tag. The advantages of choosing keyframes from the video stream as anchor points are as follows: 1. Visual intuitiveness: The video stream directly records the physical movement and scene environment of the logistics entity. Its keyframes can intuitively correspond to the core business events of the freight flow. It is the only modal data that can visually restore the entire process of the event, making it easy for users to intuitively perceive the relationship between the data and the scene; 2. Core business events of the freight flow can be directly identified through video visual features, while sensor data (such as vibration peaks) and text information (such as abnormal notes) are only indirect feedback of the event. Therefore, using keyframes as anchor points can realize the event-level association of multimodal data; 3. Temporal continuity: The video stream is collected at a fixed frame rate, and the time dimension is continuous and uninterrupted. However, the large sampling interval of sensor data and the time dispersion of text information both lead to inaccurate time. The rendered content is the synchronization between the 3D model and the real physical device. For example, the rendering time determines whether the position of the physical device has changed in the current time. If the coordinates change, the 3D model also changes its coordinate position synchronously to dynamically display the changes of the physical device. For example, a container can be used as a physical label to identify the 3D model as a container. The time label can be determined based on the anchor point. The spatial label, physical label, time label and event label are unified. Therefore, when the time label is determined, the spatial label and event label can be known. During rendering, the position change of the container 3D model can be rendered to mark the transportation or loading and unloading status.
[0072] In summary, by determining the rendering time through anchor points and then using event tags to determine the rendering content, the dynamics of the 3D scene can be updated in a timely manner.
[0073] Step S4: Display the rendered content through an interactive device.
[0074] Step S32 mentions the importance of keyframes in video stream data. The acquisition of keyframes affects rendering time and content. Therefore, to ensure the effectiveness of keyframes, the method for acquiring keyframes in the video stream data includes: pre-setting feature frames, which are used as a reference group for input image algorithms. The image algorithm identifies each frame in the video stream data, and when a frame matches the reference group, it is output as a keyframe. Setting the feature frame reference group: Taking the feature frame at the moment of container lifting as an example, it contains feature points of the hook, container corners, and gantry crane boom in the video image. The coordinates of these feature points are extracted using OpenCV's SIFT algorithm. SIFT is a classic image feature detection and description algorithm proposed by David Lowe. It can extract local keypoints with scale and rotation invariance and generate highly discriminative descriptors. It is widely used in image stitching, target recognition, 3D reconstruction, robot navigation, and other fields, and is an existing algorithm, so it will not be elaborated further here.
[0075] The SIFT algorithm is used to calculate the similarity between each frame of the video stream and the feature frames of the reference group. The similarity formula is as follows:
[0076] ;in: For the descriptor of the i-th feature point of the reference frame, Let n be the descriptor of the i-th feature point in the current frame. Assuming there are 50 feature points, then n=50. A similarity threshold S ≥ 0.8 is set to determine a match. Taking the video frame at 10:00:10.500 as an example, S = 0.85 ≥ 0.8 is calculated, so it is determined to be a keyframe and set as an anchor point.
[0077] The method also includes setting motion constraint parameters for the dynamic entity model. These parameters include thresholds for the movement trajectory and loading / unloading positions. For 3D models exceeding the motion constraint parameters, a warning signal is generated and displayed within the 3D freight scene. The motion constraint parameters can provide warnings for trajectory deviations. For example, the preset trajectory for a freight truck from the port to highway hub B is a latitude-longitude sequence (120.565°, 30.125°) → (120.570°, 30.130°) → (120.575°, 30.135°), with a trajectory deviation threshold ≤ 50m. If the deviation threshold is exceeded, a warning is issued.
[0078] The method also includes generating a floating window in the 3D freight scene when a user selects a 3D model. The floating window displays video footage and text from the multimodal data corresponding to the 3D model. The instruction to select a 3D model can be a manual click on the 3D model, which displays the floating window, showing the video stream, sensor data, and text data of the physical device corresponding to the current 3D model. The floating window ensures the cleanliness of the entire 3D freight scene and allows for a complete display of the corresponding multimodal data content when the user selects to view it.
[0079] Example 2
[0080] like Figure 2 As shown, the multimodal freight flow intelligent visual analysis system includes:
[0081] A 3D scene construction unit is used to construct a 3D freight scene based on the road network data of freight flow and generate a corresponding 3D model based on the original freight data. The generation of the corresponding 3D model based on the original freight data includes generating a 3D model based on the size parameters of the physical devices involved in logistics and dividing the 3D model into a static physical model and a dynamic physical model.
[0082] A multimodal data processing unit is used to acquire multimodal data in freight flow, including video stream data, sensor data and text data, add spatial labels, entity labels, time labels and event labels to the multimodal data, and assign values to each label;
[0083] The binding and rendering unit is used to bind and render all 3D models with various tags. The binding and rendering method is as follows:
[0084] The multimodal data is bound to all three-dimensional models using the spatial labels and the entity labels;
[0085] Acquire keyframes from the video stream data, set the keyframes as anchor points, set the rendering time according to the time tags corresponding to the anchor points, and set the rendering content of the dynamic entity model according to the event tags.
[0086] The display unit is used to display the rendered content through an interactive device.
[0087] The method for adding spatial labels, entity labels, time labels, and event labels to the multimodal data in the multimodal data processing unit is as follows:
[0088] The addition of spatial labels includes: establishing the reference coordinates of each 3D model in the 3D freight scene, acquiring the location data of the entity device associated with the multimodal data in real time, converting the location data into spatial coordinates based on the corresponding reference coordinates, and using the converted spatial coordinates as spatial labels.
[0089] The addition of entity labels includes: assigning an entity ID to each entity device, adding the entity ID to the multimodal data associated with the entity device, adding a classification symbol, with each classification symbol corresponding to a multimodal data category, and using the entity ID with the added classification symbol as the entity label;
[0090] Adding time stamps includes: establishing a time baseline, obtaining the upload time of multimodal data in real time, converting the upload time according to the time baseline to obtain a unified timestamp, and adding a timestamp to each multimodal data as a time stamp;
[0091] Adding event tags includes: setting judgment rules for business events and abnormal events, pre-setting business thresholds and abnormal thresholds for multimodal data, judging multimodal data exceeding the business threshold as business completion status, judging multimodal data exceeding the abnormal threshold as business abnormal status, and adding corresponding statuses as event tags to multimodal data.
[0092] Example 2 is essentially the same as Example 1, so the principles of each module will not be described in detail here.
[0093] The present invention also discloses a storage medium storing a computer program thereon, which, when executed by a processor, implements the multimodal freight flow intelligent visual analysis method as described in Example 1.
[0094] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the invention can be implemented in other specific forms without departing from its spirit or essential characteristics. Therefore, the embodiments should be considered in all respects as exemplary and non-limiting, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be included within the present invention. No reference numerals in the claims should be construed as limiting the scope of the claims.
[0095] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A multimodal intelligent visual analysis method for freight flow, characterized in that, Includes the following steps: A three-dimensional freight scene is constructed based on road network data of freight flow, and a corresponding three-dimensional model is generated based on the original freight data. Generating the corresponding three-dimensional model based on the original freight data includes generating a three-dimensional model based on the size parameters of the physical devices involved in the logistics, and dividing the three-dimensional model into a static physical model and a dynamic physical model. The method for constructing a three-dimensional freight scene based on road network data of freight flow includes: Acquire vector data, which includes the latitude and longitude coordinates of freight routes and import / export geographical boundaries; The vector data is converted to three-dimensional spatial coordinates and stitched together to form a global geographic base map; Multimodal data from freight flow is acquired, including video stream data, sensor data, and text data. Spatial labels, entity labels, time labels, and event labels are added to the multimodal data, and values are assigned to each label. The method for adding spatial labels, entity labels, time labels, and event labels to the multimodal data is as follows: The addition of spatial labels includes: establishing the reference coordinates of each 3D model in the 3D freight scene, acquiring the location data of the entity device associated with multimodal data in real time, converting the location data into spatial coordinates based on the corresponding reference coordinates, and using the converted spatial coordinates as spatial labels. The addition of entity labels includes: assigning an entity ID to each entity device, adding the entity ID to the multimodal data associated with the entity device, adding a classification symbol, with each classification symbol corresponding to a multimodal data category, and using the entity ID with the added classification symbol as the entity label; Adding time stamps includes: establishing a time baseline, obtaining the upload time of multimodal data in real time, converting the upload time according to the time baseline to obtain a unified timestamp, and adding a timestamp to each multimodal data as a time stamp; Adding event tags includes: setting judgment rules for business events and abnormal events, pre-setting business thresholds and abnormal thresholds for multimodal data, judging multimodal data exceeding the business threshold as business completion status, judging multimodal data exceeding the abnormal threshold as business abnormal status, and adding corresponding statuses as event tags to multimodal data; All 3D models are bound to and rendered with various tags. The binding and rendering methods are as follows: The multimodal data is bound to all three-dimensional models using the spatial labels and the entity labels; Acquire keyframes from the video stream data, set the keyframes as anchor points, set the rendering time according to the time tags corresponding to the anchor points, and set the rendering content of the dynamic entity model according to the event tags. The rendered content is displayed through an interactive device.
2. The multimodal freight flow intelligent visual analysis method according to claim 1, characterized in that, The method for obtaining keyframes of the video stream data includes: pre-setting feature frames, which are used as a reference group for input image algorithms; identifying each frame image in the video stream data through image algorithms; and outputting the frame image as a keyframe when the frame image is detected to match the reference group.
3. The intelligent visual analysis method for multimodal freight flow according to claim 1, characterized in that, The method also includes setting motion constraint parameters for the dynamic entity model, including thresholds for movement trajectory and loading / unloading position, and generating early warning signals for the 3D model that exceeds the motion constraint parameters and displaying them in the 3D freight scene.
4. The intelligent visual analysis method for multimodal freight flow according to claim 1, characterized in that, The method for acquiring multimodal data in freight flow is as follows: The raw bitstream of the container loading and unloading video stream captured by the camera is used as the video stream data; The container's GPS data and temperature and humidity sensor data are used as sensor data. The shipping information in the backend is retrieved as text data using the shipping tracking number.
5. The intelligent visual analysis method for multimodal freight flow according to claim 1, characterized in that, The method also includes generating a floating window in the 3D freight scene when a user selects a 3D model, and displaying video footage and text from the multimodal data corresponding to the 3D model in the floating window.
6. A multimodal freight flow intelligent visual analysis system, characterized in that, include: A 3D scene construction unit is used to construct a 3D freight scene based on road network data of freight flow, and to generate a corresponding 3D model based on the original freight data. Generating the corresponding 3D model based on the original freight data includes generating a 3D model based on the size parameters of the physical devices involved in the logistics, and dividing the 3D model into a static physical model and a dynamic physical model. The method for constructing a 3D freight scene based on road network data of freight flow includes: Acquire vector data, which includes the latitude and longitude coordinates of freight routes and import / export geographical boundaries; The vector data is converted to three-dimensional spatial coordinates and stitched together to form a global geographic base map; A multimodal data processing unit is used to acquire multimodal data from freight flow, including video stream data, sensor data, and text data. The unit adds spatial labels, entity labels, time labels, and event labels to the multimodal data and assigns values to each label. The method for adding spatial labels, entity labels, time labels, and event labels to the multimodal data in the multimodal data processing unit is as follows: The addition of spatial labels includes: establishing the reference coordinates of each 3D model in the 3D freight scene, acquiring the location data of the entity device associated with multimodal data in real time, converting the location data into spatial coordinates based on the corresponding reference coordinates, and using the converted spatial coordinates as spatial labels. The addition of entity labels includes: assigning an entity ID to each entity device, adding the entity ID to the multimodal data associated with the entity device, adding a classification symbol, with each classification symbol corresponding to a multimodal data category, and using the entity ID with the added classification symbol as the entity label; Adding time stamps includes: establishing a time baseline, obtaining the upload time of multimodal data in real time, converting the upload time according to the time baseline to obtain a unified timestamp, and adding a timestamp to each multimodal data as a time stamp; Adding event tags includes: setting judgment rules for business events and abnormal events, pre-setting business thresholds and abnormal thresholds for multimodal data, judging multimodal data exceeding the business threshold as business completion status, judging multimodal data exceeding the abnormal threshold as business abnormal status, and adding corresponding statuses as event tags to multimodal data; The binding and rendering unit is used to bind and render all 3D models with various tags. The binding and rendering method is as follows: The multimodal data is bound to all three-dimensional models using the spatial labels and the entity labels; Acquire keyframes from the video stream data, set the keyframes as anchor points, set the rendering time according to the time tags corresponding to the anchor points, and set the rendering content of the dynamic entity model according to the event tags. The display unit is used to display the rendered content through an interactive device.
7. A storage medium, characterized in that, It stores a computer program that, when executed by a processor, implements the multimodal freight flow intelligent visual analysis method as described in any one of claims 1-5.
Citation Information
Patent Citations
Multi-modal event data processing and collaborative circulation method and device, equipment and medium
CN121032421A
Coal mine operation and maintenance monitoring method and device, electronic equipment and storage medium
CN121173827A