Methods, devices, equipment, media, and products for multimodal data semantic representation and compression based on a hierarchical goal-attribute-relationship model.

By using a hierarchical target-attribute-relationship model to perform hierarchical semantic encoding on multimodal video data, semantic information streams of different granularities are generated, solving the problems of low data transmission efficiency and semantic distortion in wide-area video networks, and realizing efficient video data transmission and reconstruction.

CN120512558BActive Publication Date: 2025-10-28TSINGHUA UNIVERSITY +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511007575.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2025-10-28
Estimated Expiration
2045-07-22

AI Technical Summary

Technical Problem

In wide-area video network scenarios, the real-time processing requirements of multimodal video data conflict with network bandwidth, resulting in low transmission efficiency and semantic distortion, making it difficult to meet the demand for efficient transmission.

Method used

A multimodal data semantic representation and compression method based on a hierarchical target-attribute-relationship model is adopted. The multimodal video data is subjected to hierarchical semantic encoding at the pixel level, visual level, object level and event level through an edge server to generate semantic information streams of data layer, visual layer, object layer and event layer, and the corresponding semantic information streams are transmitted according to the semantic granularity requirements of the cloud server.

Benefits of technology

It significantly reduces bandwidth requirements, improves transmission efficiency, reduces semantic distortion, meets the diverse application needs of different scenarios and tasks, and achieves efficient multimodal video encoding and transmission.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120512558B_ABST
    Figure CN120512558B_ABST
Patent Text Reader

Abstract

This invention provides a method, apparatus, device, medium, and product for multimodal data semantic representation and compression based on a hierarchical target-attribute-relationship model, relating to the field of data processing. It includes: acquiring multimodal video data from an edge server; performing pixel-level semantic encoding, visual-level semantic encoding, object-level semantic encoding, and event-level semantic encoding on the multimodal video data to obtain data-layer semantic information streams, visual-layer semantic information streams, object-layer semantic information streams, and event-layer semantic information streams; and then, according to the semantic granularity requirements of a cloud server, transmitting one or more of these streams to the cloud server to improve transmission efficiency and reduce semantic distortion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing, and in particular to a method, apparatus, device, medium and product for multimodal data semantic representation and compression based on a hierarchical target-attribute-relationship model. Background Technology

[0002] The integrated space-ground wide-area video network refers to the comprehensive utilization of facilities such as towers, drones, and satellites, leveraging their advantages of high location and wide coverage to provide high-definition video surveillance and information services within an observation range of tens to thousands of kilometers. Currently, the integrated space-ground wide-area video network has a large monitoring range, a large volume of multimodal video data, and the demand for real-time processing of multimodal video data is also growing rapidly.

[0003] Multimodal video data in wide-area video network scenarios exhibits high-dimensional heterogeneous data characteristics and complex dynamic scene distributions. The conflict between the massive real-time transmission requirements and network bandwidth has become a key technical bottleneck restricting wide-area video network monitoring and recognition. Therefore, how to resolve the contradiction between existing network transmission capabilities and the real-time transmission requirements of multimodal video data, improve transmission efficiency, and reduce semantic distortion is the technical problem that this invention urgently needs to solve. Summary of the Invention

[0004] In view of the above-mentioned technical problems, the present invention provides a method, apparatus, device, medium and product for multimodal data semantic representation and compression based on a hierarchical target-attribute-relationship model, which aims to overcome the above problems or at least partially solve the above problems.

[0005] The first aspect of this invention provides a method for multimodal data semantic representation and compression based on a hierarchical target-attribute-relationship model, the method comprising:

[0006] The edge server acquires multimodal video data of the target scene simultaneously captured by cameras from remote sensing satellites, drones, and ground towers. The multimodal video data includes: satellite video data, drone video data, and ground tower video data.

[0007] The edge server aims to preserve the original semantic information of the multimodal video data by performing pixel-level semantic encoding on the multimodal video data to obtain a data layer semantic information stream.

[0008] The edge server aims to characterize the visual information of the foreground and background in the multimodal video data by performing visual-level semantic encoding on the multimodal video data to obtain a visual layer semantic information stream.

[0009] The edge server performs object-level semantic encoding on the multimodal video data to obtain an object-layer semantic information stream, which is used to characterize the target-attribute-relationship in the multimodal video data related to the target task.

[0010] The edge server performs event-level semantic encoding on the multimodal video data to obtain an event-layer semantic information stream, with the goal of representing the events related to the target task contained in the multimodal video data.

[0011] The edge server transmits one or more of the data layer semantic information stream, the visual layer semantic information stream, the object layer semantic information stream, and the event layer semantic information stream to the cloud server according to the semantic granularity requirements of the cloud server.

[0012] A second aspect of this invention provides a multimodal data semantic representation and compression system based on a hierarchical target-attribute-relationship model, the system comprising at least: a terminal, an edge server, and a cloud server; the terminal comprising at least: a remote sensing satellite, a drone, and a ground tower;

[0013] The terminal is used to simultaneously capture images of the target scene using its camera, obtain multimodal video data, and send it to the edge server.

[0014] The edge server is used for:

[0015] The multimodal video data is acquired, including satellite video data, UAV video data, and ground tower video data.

[0016] With the goal of preserving the original semantic information of the multimodal video data, pixel-level semantic encoding is performed on the multimodal video data to obtain a data layer semantic information stream;

[0017] With the goal of characterizing the visual information of the foreground and background in the multimodal video data, visual-level semantic encoding is performed on the multimodal video data to obtain a visual layer semantic information stream;

[0018] The multimodal video data is subjected to object-level semantic encoding to obtain an object-layer semantic information stream, which is used to characterize the target-attribute-relationship in the multimodal video data related to the target task.

[0019] With the goal of characterizing the events contained in the multimodal video data that are related to the target task, event-level semantic encoding is performed on the multimodal video data to obtain an event-layer semantic information stream;

[0020] Based on the semantic granularity requirements of the cloud server, one or more of the data layer semantic information stream, the visual layer semantic information stream, the object layer semantic information stream, and the event layer semantic information stream are transmitted to the cloud server.

[0021] A third aspect of the present invention provides an electronic device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the multimodal data semantic representation and compression method based on a hierarchical target-attribute-relationship model as described in the first aspect of the present invention.

[0022] A fourth aspect of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the multimodal data semantic representation and compression method based on a hierarchical target-attribute-relationship model as described in the first aspect of the present invention.

[0023] The fifth aspect of the present invention provides a computer program product, including a computer program that, when executed by a processor, implements the multimodal data semantic representation and compression method based on a hierarchical target-attribute-relationship model as described in the first aspect of the present invention.

[0024] In the multimodal data semantic representation and compression method based on a hierarchical target-attribute-relationship model proposed in this invention, the edge server performs hierarchical semantic encoding on multimodal video data. Following a progression from fine to coarse semantic granularity, pixel-level, visual-level, object-level, and event-level semantic encoding are performed on the multimodal video data, resulting in data-layer semantic information streams, visual-layer semantic information streams, object-layer semantic information streams, and event-layer semantic information streams corresponding to the multimodal video data. Based on the semantic granularity requirements of the cloud server, one or more of these streams are transmitted to the cloud server. Thus, this invention transmits semantic information streams, significantly reducing bandwidth requirements. Through multi-level semantic representation, hierarchical semantic information stream transmission can be performed according to actual semantic granularity requirements for diverse application needs in different scenarios and tasks, achieving efficient multimodal video encoding and transmission. This effectively reduces redundant information processing overhead, resolves the contradiction between current network transmission capacity and the real-time transmission requirements of multimodal video data, improves transmission efficiency, and reduces semantic distortion. Attached Figure Description

[0025] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0026] Figure 1 This is a flowchart illustrating the steps of a multimodal data semantic representation and compression method based on a hierarchical target-attribute-relationship model, as shown in an embodiment of the present invention.

[0027] Figure 2 This is a schematic diagram illustrating a hierarchical semantic representation according to an embodiment of the present invention;

[0028] Figure 3 This is a flowchart illustrating the workflow of a discriminator in a generative adversarial network according to an embodiment of the present invention;

[0029] Figure 4 This is a structural block diagram of a multimodal data semantic representation and compression system based on a hierarchical target-attribute-relationship model provided by an embodiment of the present invention;

[0030] Figure 5 This is a schematic diagram of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0031] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0032] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating the steps of a multimodal data semantic representation and compression method based on a hierarchical target-attribute-relationship model, as shown in an embodiment of the present invention. Figure 1 As shown, the multimodal data semantic representation and compression method based on a hierarchical target-attribute-relationship model provided in this embodiment includes at least the following steps:

[0033] Step S11: The edge server acquires multimodal video data of the target scene simultaneously captured by cameras from remote sensing satellites, drones, and ground towers.

[0034] In this embodiment, for the integrated space-ground wide-area video network, terminals such as remote sensing satellites, drones, and ground towers can simultaneously capture images of the same target scene through cameras, obtaining multimodal video data for the same target scene. The multimodal video data includes: satellite video data, drone video data, and ground tower video data.

[0035] Among them, remote sensing satellites offer wide coverage and macroscopic perspective; drones provide flexibility and maneuverability; and ground-based towers offer high resolution and strong real-time performance. Remote sensing video data, characterized by its large observation range and strong macroscopic perspective, is suitable for applications such as large-scale geomorphological analysis and land use classification. However, remote sensing platforms are limited by spatial resolution, resulting in insufficient accuracy in perceiving small-scale targets; simultaneously, their temporal resolution is typically only a few hours or even days, making it difficult to achieve continuous, all-time coverage. Drones are flexible and maneuverable, suitable for rapid response and localized high-resolution observation. However, drones are limited by battery capacity and quantity, with flight distances typically only tens of kilometers and endurance of about half an hour, resulting in insufficient spatiotemporal reach. Ground-based towers offer high spatial resolution and strong real-time performance, capable of capturing information on small-sized features, subtle geomorphological characteristics, and rapidly changing targets. However, their observation range is relatively small and their perspective is limited, making them unsuitable for the needs of wide-area, complex scenarios.

[0036] It is understood that the multimodal video data in this embodiment consists of video data with different fields of view and resolutions. Specifically, the first field of view corresponding to the satellite video data is larger than the second field of view corresponding to the UAV video data, the first field of view corresponding to the satellite video data is larger than the third field of view corresponding to the ground tower video data, and the first resolution corresponding to the satellite video data is smaller than the second resolution corresponding to the UAV video data, and the first resolution corresponding to the satellite video data is smaller than the third resolution corresponding to the ground tower video data.

[0037] Step S12: The edge server performs pixel-level semantic encoding on the multimodal video data with the goal of preserving the original semantic information of the multimodal video data, and obtains a data layer semantic information stream.

[0038] In this embodiment, the edge server can perform hierarchical semantic representation of multimodal video data according to the semantic granularity from fine to coarse, namely the data layer, visual layer, object layer, and event layer, so as to realize multi-level semantic encoding of multimodal video data.

[0039] Specifically, regarding the hierarchical semantic representation of the data layer, the data layer represents the original multimodal video data. The edge server, aiming to preserve the original semantic information of the multimodal video data, performs pixel-level semantic encoding on the multimodal video data to obtain the semantic information stream of the data layer. In this way, high-fidelity pixel-level representation and semantic encoding of multimodal video data can maintain pixel consistency of the multimodal video frames to the greatest extent possible, effectively preserving the rich original semantic information in the multimodal video data.

[0040] Step S13: The edge server performs visual-level semantic encoding on the multimodal video data with the goal of characterizing the visual information of the foreground and background in the multimodal video data, and obtains a visual layer semantic information stream.

[0041] In this embodiment, for the hierarchical semantic representation of the visual layer, the edge server aims to represent the visual information of the foreground and background in the multimodal video data by performing visual-level semantic encoding on the multimodal video data to obtain a visual layer semantic information stream. This visual layer semantic information stream includes at least the texture, color, edge, and shape features of the foreground (i.e., the target) and background in the multimodal video data; it represents preliminary semantic features. The semantic information contained in this visual layer semantic information stream is the most comprehensive among the multi-level semantic information streams (data layer semantic information stream, visual layer semantic information stream, object layer semantic information stream, and event layer semantic information stream), but it also has a large data volume.

[0042] Step S14: The edge server performs object-level semantic encoding on the multimodal video data to obtain an object-level semantic information stream.

[0043] In this embodiment, for the hierarchical semantic representation of the object layer, the edge server performs object-level semantic encoding on the multimodal video data to obtain an object-level semantic information stream. Specifically, the edge server can perform target recognition on the multimodal video data based on the target task to obtain a set of target objects, and extract the attribute features of each target object in the target object set and the relationship features between the target objects. Then, based on the target objects in the target object set, the attribute features of the target objects, and the relationship features between the target objects, the multimodal video data is subjected to object-level semantic encoding to obtain an object-level semantic information stream. This object-level semantic information stream is used to represent the target-attribute-relationship in the multimodal video data related to the target task. Thus, this embodiment, based on the identification of target objects in the multimodal video data based on the target task, achieves object-level semantic representation through target, attribute, and relationship triples, enabling efficient semantic description of objects related to the target task while further reducing the amount of data.

[0044] Step S15: The edge server performs event-level semantic encoding on the multimodal video data to obtain an event-layer semantic information stream, with the goal of representing the events related to the target task contained in the multimodal video data.

[0045] In this embodiment, for the hierarchical semantic representation of the event layer, the edge server performs event-level semantic encoding on the multimodal video data, aiming to represent the events related to the target task contained in the multimodal video data, to obtain an event-layer semantic information stream. This event-layer semantic information stream is a comprehensive description of the video content or the events related to the target task contained in the multimodal video data, and can express the semantic information related to the target task with a very low data volume.

[0046] In this embodiment, according to the semantic granularity from fine to coarse, the data layer semantic information flow, visual layer semantic information flow, object layer semantic information flow, and event layer semantic information flow have progressively decreasing amounts of representational data and progressively increasing levels of semantic abstraction. For example... Figure 2 As shown, Figure 2 This is a schematic diagram illustrating a hierarchical semantic representation according to an embodiment of the present invention. Figure 2 In this embodiment, the data layer semantic information flow includes the raw semantic information of satellite video data (i.e., remote sensing data), drone video data (i.e., drone data), and ground tower video data (i.e., tower monitoring data); the visual layer semantic information flow includes the texture features, color features, edge features, and shape features of the foreground and background in the multimodal video data; the object layer semantic information flow represents the target-attribute-relationship related to the target task in the multimodal video data; and the event layer semantic information flow represents the overall description of the events (such as forest fire events) contained in the multimodal video data. Thus, the hierarchical semantic representation in this embodiment presents a feature pyramid structure, with the data layer, visual layer, object layer, and event layer having progressively decreasing amounts of representational data and progressively increasing levels of semantic abstraction.

[0047] It should be noted that this embodiment does not impose any restrictions on the execution order of the above steps S12, S13, S14 and S15. Steps S12, S13, S14 and S15 can be executed simultaneously or in any order.

[0048] Step S16: The edge server transmits one or more of the data layer semantic information stream, the visual layer semantic information stream, the object layer semantic information stream, and the event layer semantic information stream to the cloud server according to the semantic granularity requirements of the cloud server.

[0049] In this embodiment, the cloud server can send corresponding semantic granularity requirements to the edge server according to the target task. The edge server, based on the semantic granularity requirements of the cloud server, transmits one or more of the data layer semantic information stream, visual layer semantic information stream, object layer semantic information stream, and event layer semantic information stream to the cloud server.

[0050] For example, when the semantic granularity requirement falls within the first semantic scope, the edge server sends the event layer semantic information stream to the cloud server; when the semantic granularity requirement falls within the second semantic scope, the edge server sends at least the object layer semantic information stream to the cloud server; when the semantic granularity requirement falls within the third semantic scope, the edge server sends at least the visual layer semantic information stream to the cloud server; and when the semantic granularity requirement falls within the fourth semantic scope, the edge server sends at least the data layer semantic information stream to the cloud server. The semantic granularity requirements corresponding to the first, second, third, and fourth semantic scopes are arranged from coarse to fine in that order.

[0051] In this embodiment, through multi-level semantic representation, hierarchical semantic information stream transmission can be carried out according to the actual semantic granularity requirements for diverse application needs of different scenarios and tasks, thereby achieving efficient multimodal video encoding transmission. Moreover, since the transmitted information stream is semantic information, this embodiment also significantly reduces bandwidth requirements, effectively reduces redundant information processing overhead, resolves the contradiction between current network transmission capabilities and the real-time transmission requirements of multimodal video data, improves transmission efficiency, and reduces semantic distortion.

[0052] In conjunction with the above embodiments, in one implementation, the present invention also provides a method for multimodal data semantic representation and compression based on a hierarchical target-attribute-relationship model. In this method, in addition to the steps described above, step S21 may also be included:

[0053] Step S21: Upon receiving the object layer semantic information stream, the cloud server processes the object layer semantic information stream using a pre-trained generator to obtain reconstructed multimodal video data.

[0054] In this embodiment, the object-layer semantic information flow includes target-attribute-relation triples. A pre-trained generator is pre-deployed on the cloud server. This pre-trained generator takes the target-attribute-relation triples in the object-layer semantic information flow as input and, through multimodal feature fusion and spatiotemporal modeling, restores the abstract object-layer semantic information flow into pixel-level multimodal video data that conforms to physical laws. The pre-trained generator is obtained by training a generative adversarial network using multimodal video data samples and corresponding object-layer semantic information flow samples as training data.

[0055] The cloud server can receive semantic information streams sent by the edge server. Upon receiving the object-layer semantic information stream, the cloud server can process the object-layer semantic information stream through a pre-trained generator to obtain the reconstructed multimodal video data output by the pre-trained generator, thereby realizing the reconstruction of multimodal video data.

[0056] In one alternative example, the generator can reconstruct multimodal video data through structured semantics-driven and spatiotemporal adversarial optimization: The generator first parses the target-attribute-relationship triples and reconstructs the visual content layer by layer through a multimodal feature fusion architecture: In the object feature reconstruction layer, basic geometric structures are generated based on the object category (i.e., the target category) and the object position; in the attribute refinement layer, attribute details (such as dynamic and / or static attributes) are injected; in the relation constraint layer, the interaction logic between objects is dynamically adjusted based on the spatiotemporal relation graph, thereby obtaining reconstructed keyframes based on multi-layer fusion; then, temporal modeling is performed, using the target-attribute-relationship triples corresponding to the keyframes as the starting point to predict intermediate frames, ensuring the continuity of actions between frames, and finally realizing the reconstruction of multimodal video data.

[0057] In conjunction with any of the above embodiments, the present invention also provides a method for multimodal data semantic representation and compression based on a hierarchical target-attribute-relationship model. In addition to the steps described above, this method may further include the following steps S31 to S35:

[0058] Step S31: The cloud server inputs the object layer semantic information stream sample into the generator in the generative adversarial network to obtain the reconstructed multimodal video data sample.

[0059] In this embodiment, the multimodal video data sample is the multimodal video data used to train the generator, and the object-layer semantic information stream sample is the object-layer semantic information stream used to train the generator. This object-layer semantic information stream sample is the object-layer semantic information stream corresponding to the multimodal video data sample. The cloud server can input the object-layer semantic information stream sample into the generator in the generative adversarial network to obtain the reconstructed multimodal video data sample output by the generator in the generative adversarial network.

[0060] Step S32: The cloud server inputs the reconstructed multimodal video data sample and the multimodal video data sample into the discriminator in the generative adversarial network.

[0061] In this embodiment, the input to the discriminator in the generative adversarial network is the reconstructed multimodal video data sample output by the generator in the generative adversarial network and the real multimodal video data sample. Based on this, the cloud server can input the reconstructed multimodal video data sample and the multimodal video data sample into the discriminator in the generative adversarial network.

[0062] Step S33: The cloud server compares the objects related to the target task in the reconstructed multimodal video data samples with the objects related to the target task in the multimodal video data samples at a local scale through the first branch of the discriminator, and obtains a first evaluation result.

[0063] In this embodiment, the discriminator includes a first branch and a second branch. The first branch is a visual realism branch. The cloud server focuses on the detail realism of objects related to the target task at a local scale through the first branch of the discriminator: by comparing the objects related to the target task in the reconstructed multimodal video data sample with the objects related to the target task in the multimodal video data sample at a local scale through the first branch, a first evaluation result is obtained.

[0064] In one optional embodiment, a first high-resolution feature map and a second high-resolution feature map of objects related to the target task in the reconstructed multimodal video data samples can be obtained using a 3D convolutional network. Then, the first and second high-resolution feature maps are compared region by region to obtain an adversarial loss, which serves as the first evaluation result. This adversarial loss characterizes the microscopic consistency between the reconstructed object and the real object.

[0065] Step S34: The cloud server, through the second branch of the discriminator, uses the multimodal video data samples as a benchmark to analyze the rationality of the relationship between objects related to the target task in the reconstructed multimodal video data samples on a global scale, and obtains a second evaluation result.

[0066] In this embodiment, the second branch is a semantic branch. The cloud server analyzes the semantic rationality of the overall scene on a global scale through the second branch of the discriminator: through the second branch, based on the multimodal video data samples, it analyzes the rationality of the relationship between objects related to the target task in the reconstructed multimodal video data samples on a global scale, and obtains the second evaluation result.

[0067] In an optional embodiment, graph similarity comparison can be performed by reverse-engineering the OAR (target-attribute-relationship) graph to obtain a second evaluation result: verify whether the dynamic interaction between objects related to the target task in the reconstructed multimodal video data sample follows the relation constraints defined by the target-attribute-relationship triplet corresponding to the multimodal video data sample, and verify the attribute correlation to obtain the second evaluation result.

[0068] In one embodiment, such as Figure 3 As shown, Figure 3 This is a flowchart illustrating the workflow of a discriminator in a generative adversarial network according to an embodiment of the present invention. Figure 3 In this process, the input frames (i.e., the reconstructed multimodal video data samples and the multimodal video data samples mentioned above) are input to the discriminator. The discriminator performs dual evaluation through two branches: through the visual realism branch (i.e., the first branch), the adversarial loss between the reconstructed multimodal video data samples and the multimodal video data samples is output based on the 3D convolutional network as the first evaluation result; through the semantic branch (i.e., the second branch), graph similarity is compared by back-inferring the OAR graph to obtain the second evaluation result; and then, based on the first evaluation result and the second evaluation result, the generator in the generative adversarial network is adversarially trained.

[0069] Step S35: The cloud server, through the discriminator, drives the generator to reconstruct the multimodal video data sample with the goal of semantically approximating it, based on the first evaluation result and the second evaluation result, to obtain the pre-trained generator.

[0070] In this embodiment, the cloud server can use a discriminator to calculate a loss function based on the first evaluation result output by the first branch and the second evaluation result output by the second branch. By minimizing the loss function, the generator is driven to perform video reconstruction with the goal of semantically approximating multimodal video data samples. The network parameters of the generator in the generative adversarial network are updated until the generator outputs video content that is both visually realistic and conforms to structured logic, that is, until the generated frames output by the generator closely approximate the distribution of real video at the semantic level (such as consistency of object attributes, relational logic, etc.), thus obtaining a trained generator, i.e., a pre-trained generator. In an optional embodiment, the loss function is... Loss function.

[0071] In this embodiment, the discriminator performs dual evaluation using a hierarchical convolutional neural network. At the local scale, it focuses on the realism of object details, generating microscopic consistency with real objects through region-by-region comparison of high-resolution feature maps. At the global scale, it analyzes the semantic rationality of the overall scene, verifying whether dynamic interactions between objects follow the relational constraints defined by the target-attribute-relationship triplet and checking attribute correlations. This drives the generator to repair local details and enhance global physical rationality, ultimately outputting a reconstructed video that strictly adheres to OAR semantic logic and possesses visual realism. Thus, through the collaborative mechanism of explicit semantic supervision (i.e., global semantic supervision) and implicit distribution matching (i.e., local detail allocation matching), the discriminator ensures that the reconstructed video maintains high task adaptability and strong noise resistance even under low-bandwidth transmission. While reducing bandwidth, it significantly improves the fidelity of key object details and the temporal consistency of complex interactions.

[0072] In conjunction with any of the above embodiments, in one implementation, the present invention also provides a method for multimodal data semantic representation and compression based on a hierarchical target-attribute-relationship model. In this method, in addition to the steps described above, the following steps S41 to S45 may also be included:

[0073] Step S41: The cloud server uses the first and last two frames of the multimodal video data sample as two key frames to obtain the manual annotation results of the key frames.

[0074] In this embodiment, after acquiring the multimodal sample data, the cloud server can use the first and last frames of the multimodal video data sample as two keyframes. Manual annotation is then performed on these keyframes to obtain the manually annotated results. The manually annotated results include the bounding boxes of targets related to the target task in the first and last frames; each bounding box is a boundary box that surrounds the target.

[0075] Step S42: The cloud server determines the direction of the vector of the center coordinate of the bounding box of the target related to the target task in the two key frames as the direction of movement of the target in the video frame.

[0076] In this embodiment, after the manual annotation results of the first and last keyframes, a spatiotemporal continuity annotation algorithm can be used to automatically generate bounding boxes for targets related to the target task in the intermediate frames. Specifically, the cloud server determines the direction of movement of the target related to the target task in the video frame (e.g., speed, direction) based on the motion characteristics of the target with the video frame as a reference. For example, the direction of the vector formed by the center coordinates (center coordinates A1 and A2) of the bounding box of target A in the first and last keyframes is determined as the direction of movement of target A in the video frame; similarly, the direction of the vector formed by the center coordinates (center coordinates B1 and B2) of the bounding box of target B in the first and last keyframes is determined as the direction of movement of target B in the video frame.

[0077] Step S43: The cloud server determines the target's movement speed in the video frame by the ratio of the length of the vector of the center coordinates of the bounding box of the target related to the target task in the two key frames to the time interval between the two key frames.

[0078] In this embodiment, the cloud server can determine the target's movement speed in the video frame (such as the overall video frame) by the ratio of the length of the vector formed by the center coordinates of the bounding box of the target related to the target task in the first and last keyframes to the interval time between the first and last keyframes. For example, the ratio l1 / t of the length l1 of the vector formed by the center coordinates (center coordinates A1 and A2) of the bounding box of target A in the first and last keyframes to the interval time t between the first and last keyframes can be determined as the movement speed of target A in the video frame; the ratio l2 / t of the length l2 of the vector formed by the center coordinates (center coordinates B1 and B2) of the bounding box of target B in the first and last keyframes to the interval time t between the first and last keyframes can be determined as the movement speed of target B in the video frame.

[0079] Step S44: The cloud server annotates the bounding boxes of the targets related to the target task in the video frame based on the movement speed and direction of the targets related to the target task in the video frame.

[0080] In this embodiment, the cloud server can annotate bounding boxes of targets related to the target task in multiple intermediate frames between the first and last keyframes based on the target task’s movement speed and direction in the video frame.

[0081] In an optional example, the cloud server can determine the vertex vectors one by one based on the corresponding vertex coordinates of the bounding boxes of the targets related to the target task in the first and last keyframes. Based on the vertex vectors, the server can automatically label the bounding boxes of the targets related to the target task in multiple intermediate frames between the first and last keyframes, taking into account the movement speed and direction of the targets related to the target task in the video frame.

[0082] For example, given the vertex coordinates of the bounding box of target A in the first and last keyframes (vertex coordinates A3, A4, A5, and A6 in the first keyframe, and vertex coordinates A7, A8, A9, and A10 in the last keyframe), vertex coordinates A3 and A7 correspond to vertex vector X1; vertex coordinates A4 and A8 correspond to vertex vector X2; vertex coordinates A5 and A9 correspond to vertex vector X3; and vertex coordinates A6 and A10 correspond to vertex vector X4. Therefore, based on vertex vectors X1, X2, X3, and X4, and considering the movement speed and direction of target A in the video frame, the bounding boxes of target A can be labeled for multiple intermediate frames between the first and last keyframes. Based on this, the bounding box of each target can be labeled for multiple intermediate frames between the first and last keyframes.

[0083] Step S45: The cloud server determines the target-attribute-relationship related to the target task in the multimodal video data sample based on the bounding boxes corresponding to the two key frames and the multiple intermediate frames, and uses it as the object layer semantic information flow sample.

[0084] In this embodiment, the cloud server can determine the target-attribute-relationship related to the target task in the multimodal video data sample based on the bounding boxes of the targets related to the target task corresponding to the first and last keyframes and multiple intermediate frames, and use this as an object-layer semantic information flow sample. Specifically, it can perform target recognition on the multimodal video data sample based on the bounding boxes of the targets related to the target task corresponding to the first and last keyframes and multiple intermediate frames, obtain a target object set sample, and extract the sample attribute features of each target object sample and the association relationship feature samples between each target object sample. Then, based on the target object samples in the target object set sample, the sample attribute features of the target object samples, and the association relationship feature samples between each target object sample, it performs object-level semantic encoding on the multimodal video data sample to obtain an object-layer semantic information flow sample. This object-layer semantic information flow sample is used to represent the target-attribute-relationship related to the target task in the multimodal video data sample.

[0085] In this embodiment, a spatiotemporal continuity annotation algorithm is used to automatically annotate the bounding boxes of targets related to the target task in the intermediate frames, avoiding the waste of computational resources and potential false detection problems caused by using the target detection model frame by frame. For example, in a multimodal video including a leopard, if the positions of the leopard in the start and end frames are manually annotated, this embodiment can reasonably infer the position of the leopard in the intermediate frames and annotate it based on the leopard's running speed and direction.

[0086] In conjunction with any of the above embodiments, in one implementation, the present invention also provides a method for multimodal data semantic representation and compression based on a hierarchical target-attribute-relationship model. In addition to the steps described above, this method may further include steps S51 to S53:

[0087] Step S51: The cloud server inputs the reconstructed multimodal video data into the event detection model to obtain the event detection result.

[0088] In this embodiment, after obtaining the reconstructed multimodal video data, the cloud server can input the reconstructed multimodal video data into a pre-trained event detection model to obtain event detection results. The pre-trained event detection model is trained based on multimodal video data samples and their corresponding historical event data, and is used for event detection on the multimodal video data.

[0089] Step S52: If the presence of a target event is determined based on the event detection results, the cloud server sends a scheduling instruction to the drone to schedule the drone to focus on the target event for multi-view video acquisition.

[0090] In this embodiment, the target event is the event to be monitored. For example, the target event includes, but is not limited to, abnormal events (such as forest fires, oil and gas leaks, etc.), animal migration, animal reproduction, etc., without limitation. When the cloud server determines that a target event exists based on the event detection results corresponding to the reconstructed multimodal video data, it sends a scheduling instruction to the drone. This scheduling instruction is used to direct the drone to focus on the target event and perform multi-view video acquisition.

[0091] Based on received scheduling instructions, the drone identifies the target object corresponding to the target event, performs multi-view video capture of the target object, obtains multi-view video data of the target object, and sends it to a cloud server. In an optional embodiment, the drone can send the multi-view video data to the cloud server through an edge server.

[0092] Step S53: The cloud server continuously monitors the target event based on the multi-view video data collected by the drone.

[0093] In this embodiment, the cloud server can receive multi-view video data collected by the drone to continuously monitor the target event, further clarify the information of the target event, and form a real-time application closed loop.

[0094] In conjunction with any of the above embodiments, in one implementation, the present invention also provides a method for multimodal data semantic representation and compression based on a hierarchical target-attribute-relationship model. In this method, step S16 may specifically include steps S61 to S65:

[0095] Step S61: When the edge server receives the target scene event-level monitoring request sent by the cloud server, it transmits the event layer semantic information stream to the cloud server.

[0096] In this embodiment, the cloud server can perform event-level monitoring of the target scene, such as monitoring for abnormal events. In this case, it can send an event-level monitoring request for the target scene to the edge server. Upon receiving the event-level monitoring request from the cloud server, the edge server determines that the cloud server's semantic granularity requirement is only for event-layer semantic information streams (e.g., the semantic granularity requirement belongs to the first semantic range). Therefore, upon receiving the event-level monitoring request from the cloud server, the edge server can transmit the event-layer semantic information stream to the cloud server.

[0097] Step S62: When the cloud server determines that an abnormal event has occurred based on the event layer semantic information stream, it sends a target scene object-level monitoring request to the edge server.

[0098] In this embodiment, the cloud server receives the event layer semantic information stream sent by the edge server and performs event detection on the event layer semantic information stream. When the event detection is performed based on the event layer semantic information stream and it is determined that an abnormal event has occurred, the cloud server can then perform object set monitoring on the target scene and send the target scene object-level monitoring request to the edge server.

[0099] Step S63: Upon receiving the target scene object-level monitoring request, the edge server transmits the object-layer semantic information stream to the cloud server.

[0100] In this embodiment, when the edge server receives the target scene object-level monitoring request sent by the cloud server, it determines that the semantic granularity requirement of the cloud server is only a requirement for object-level semantic information flow (such as the semantic granularity requirement belonging to the second semantic range). Therefore, the edge server can transmit the object-level semantic information flow to the cloud server when it receives the target scene object-level monitoring request sent by the cloud server.

[0101] Step S64: When the cloud server needs to analyze the cause of the abnormal event, it sends the target scene data-level and visual-level monitoring requirements to the edge server.

[0102] In this embodiment, when the cloud server needs to analyze the cause of an abnormal event that occurs in the target scene, the cloud server can perform data-level and visual-level monitoring of the target scene. At this time, the cloud server can send the data-level and visual-level monitoring request of the target scene to the edge server.

[0103] Step S65: Upon receiving the target scene data-level and visual-level monitoring requirements, the edge server transmits the visual layer semantic information stream and the data layer semantic information stream to the cloud server.

[0104] In this embodiment, when the edge server receives the target scene data-level and visual-level monitoring requirements sent by the cloud server, it determines that the semantic granularity requirement of the cloud server is a requirement for both data layer semantic information streams and visual layer semantic information streams (e.g., the semantic granularity requirement belongs to the union of the second semantic range and the third semantic range). Therefore, when the edge server receives the target scene data-level and visual-level monitoring requirements sent by the cloud server, it can transmit the visual layer semantic information stream and the data layer semantic information stream to the cloud server, so that the cloud server can analyze the cause of the abnormal event based on the visual layer semantic information stream and the data layer semantic information stream.

[0105] In conjunction with any of the above embodiments, in one implementation, the target task in this embodiment includes, but is not limited to, target identification task, event warning task, target event monitoring task, etc. This embodiment does not limit the specific type of target task.

[0106] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of the present invention are not limited to the described order of actions, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.

[0107] Based on the same inventive concept, one embodiment of the present invention provides a multimodal data semantic representation and compression system based on a hierarchical target-attribute-relationship model. (Reference) Figure 4 , Figure 4 This is a structural block diagram of a multimodal data semantic representation and compression system based on a hierarchical target-attribute-relationship model, provided by an embodiment of the present invention. For example... Figure 4 As shown, the system includes at least: terminals, edge servers, and cloud servers; the terminals include at least: remote sensing satellites, drones, and ground towers;

[0108] The terminal is used to simultaneously capture images of the target scene using its camera, obtain multimodal video data, and send it to the edge server.

[0109] The edge server is used for:

[0110] The multimodal video data is acquired, including satellite video data, UAV video data, and ground tower video data.

[0111] With the goal of preserving the original semantic information of the multimodal video data, pixel-level semantic encoding is performed on the multimodal video data to obtain a data layer semantic information stream;

[0112] With the goal of characterizing the visual information of the foreground and background in the multimodal video data, visual-level semantic encoding is performed on the multimodal video data to obtain a visual layer semantic information stream;

[0113] The multimodal video data is subjected to object-level semantic encoding to obtain an object-layer semantic information stream, which is used to characterize the target-attribute-relationship in the multimodal video data related to the target task.

[0114] With the goal of characterizing the events contained in the multimodal video data that are related to the target task, event-level semantic encoding is performed on the multimodal video data to obtain an event-layer semantic information stream;

[0115] Based on the semantic granularity requirements of the cloud server, one or more of the data layer semantic information stream, the visual layer semantic information stream, the object layer semantic information stream, and the event layer semantic information stream are transmitted to the cloud server.

[0116] Optionally, the cloud server is configured to process the object layer semantic information stream using a pre-trained generator upon receiving the object layer semantic information stream to obtain reconstructed multimodal video data;

[0117] The pre-trained generator is obtained by training a generative adversarial network using multimodal video data samples and corresponding object-layer semantic information stream samples as training data.

[0118] Optionally, the cloud server is further used for:

[0119] The semantic information stream sample of the object layer is input into the generator in the generative adversarial network to obtain the reconstructed multimodal video data sample;

[0120] The reconstructed multimodal video data samples and the multimodal video data samples are input into the discriminator in the generative adversarial network;

[0121] By using the first branch of the discriminator, the consistency between the objects related to the target task in the reconstructed multimodal video data samples and the objects related to the target task in the multimodal video data samples is compared at a local scale to obtain a first evaluation result;

[0122] Through the second branch of the discriminator, using the multimodal video data samples as a benchmark, the rationality of the relationship between objects related to the target task in the reconstructed multimodal video data samples is analyzed on a global scale to obtain a second evaluation result;

[0123] The discriminator drives the generator to reconstruct the multimodal video data sample with the goal of semantically approximating it, based on the first evaluation result and the second evaluation result, thus obtaining the pre-trained generator.

[0124] Optionally, the cloud server is further used for:

[0125] The first and last two frames of the multimodal video data sample are used as two key frames to obtain the manual annotation results of the key frames. The manual annotation results include the bounding boxes of the targets related to the target task in the first and last frames.

[0126] The direction of the vector of the center coordinate of the bounding box of the target related to the target task in the two key frames is determined as the direction of movement of the target in the video frame;

[0127] The ratio of the length of the vector of the center coordinates of the bounding box of the target related to the target task in the two key frames to the time interval between the two key frames is determined as the moving speed of the target in the video frame.

[0128] Using the speed and direction of movement of the target related to the target task in the video frame, the bounding boxes of the target related to the target task are labeled for multiple intermediate frames between the two key frames;

[0129] Based on the bounding boxes corresponding to the two keyframes and the plurality of intermediate frames, the target-attribute-relationships related to the target task in the multimodal video data samples are determined and used as the object layer semantic information flow samples.

[0130] Optionally, the cloud server is also used for:

[0131] The reconstructed multimodal video data is input into the event detection model to obtain the event detection results;

[0132] If a target event is determined to exist based on the event detection results, a scheduling command is sent to the drone to schedule the drone to focus on the target event for multi-view video acquisition.

[0133] The target event is continuously monitored based on the multi-view video data collected by the drone.

[0134] Optionally, the edge server is configured to transmit the event-layer semantic information stream to the cloud server upon receiving a target scene event-level monitoring request sent by the cloud server.

[0135] The cloud server is used to send a target scene object-level monitoring request to the edge server when an abnormal event is determined to have occurred based on the event layer semantic information stream.

[0136] The edge server is used to transmit the object-level semantic information stream to the cloud server when it receives the target scene object-level monitoring request.

[0137] The cloud server is used to send target scene data-level and visual-level monitoring requests to the edge server when it is necessary to analyze the cause of abnormal events.

[0138] The edge server is used to transmit the visual layer semantic information stream and the data layer semantic information stream to the cloud server when it receives the target scene data-level and visual-level monitoring requirements.

[0139] Based on the same inventive concept, another embodiment of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in the multimodal data semantic representation and compression method based on a hierarchical target-attribute-relationship model as described in any of the above embodiments of the present invention.

[0140] Based on the same inventive concept, another embodiment of the present invention provides an electronic device, such as... Figure 5 As shown, Figure 5 This is a schematic diagram of an electronic device according to an embodiment of the present invention. The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When executed by the processor, the program implements the steps of the multimodal data semantic representation and compression method based on a hierarchical target-attribute-relationship model described in any of the above embodiments of the present invention.

[0141] Based on the same inventive concept, another embodiment of the present invention provides a computer program product, including a computer program that, when executed by a processor, implements the steps in the multimodal data semantic representation and compression method based on a hierarchical target-attribute-relationship model described in any of the above embodiments of the present invention.

[0142] As the system implementation is basically similar to the method implementation, it is described in a relatively simple way. For relevant details, please refer to the description of the method implementation.

[0143] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0144] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, apparatus, or computer program products. Therefore, embodiments of the present invention can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of the present invention can take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0145] Embodiments of the present invention are described with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, as well as combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0146] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0147] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0148] Although preferred embodiments of the present invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.

[0149] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.

[0150] The foregoing has provided a detailed description of a multimodal data semantic representation and compression method, apparatus, device, medium, and product based on a hierarchical target-attribute-relationship model provided by the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A multimodal data semantic representation and compression method based on a hierarchical goal-attribute-relationship model, characterized in that, The method includes: The edge server acquires multimodal video data of the target scene simultaneously captured by cameras from remote sensing satellites, drones, and ground towers. The multimodal video data includes: satellite video data, drone video data, and ground tower video data. The edge server aims to preserve the original semantic information of the multimodal video data by performing pixel-level semantic encoding on the multimodal video data to obtain a data layer semantic information stream. The edge server aims to characterize the visual information of the foreground and background in the multimodal video data by performing visual-level semantic encoding on the multimodal video data to obtain a visual layer semantic information stream. The edge server performs object-level semantic encoding on the multimodal video data to obtain an object-layer semantic information stream, which is used to characterize the target-attribute-relationship in the multimodal video data related to the target task. The edge server performs event-level semantic encoding on the multimodal video data to obtain an event-layer semantic information stream, with the goal of representing the events related to the target task contained in the multimodal video data. The edge server transmits one or more of the data layer semantic information stream, the visual layer semantic information stream, the object layer semantic information stream, and the event layer semantic information stream to the cloud server according to the semantic granularity requirements of the cloud server.

2. The multimodal data semantic representation and compression method based on a hierarchical target-attribute-relationship model according to claim 1, characterized in that, The method further includes: Upon receiving the object-layer semantic information stream, the cloud server processes the object-layer semantic information stream using a pre-trained generator to obtain reconstructed multimodal video data. The pre-trained generator is obtained by training a generative adversarial network using multimodal video data samples and corresponding object-layer semantic information stream samples as training data.

3. The multimodal data semantic representation and compression method based on a hierarchical target-attribute-relationship model according to claim 2, characterized in that, The method further includes: The cloud server inputs the object layer semantic information stream sample into the generator in the generative adversarial network to obtain reconstructed multimodal video data samples; The cloud server inputs the reconstructed multimodal video data sample and the multimodal video data sample into the discriminator in the generative adversarial network; The cloud server compares the objects related to the target task in the reconstructed multimodal video data samples with the objects related to the target task in the multimodal video data samples at a local scale through the first branch of the discriminator, and obtains a first evaluation result; The cloud server, through the second branch of the discriminator, uses the multimodal video data samples as a benchmark to analyze the rationality of the relationship between objects related to the target task in the reconstructed multimodal video data samples on a global scale, and obtains a second evaluation result; The cloud server, through the discriminator, drives the generator to reconstruct the multimodal video data sample with the goal of semantically approximating it, based on the first evaluation result and the second evaluation result, to obtain the pre-trained generator.

4. The multimodal data semantic representation and compression method based on a hierarchical target-attribute-relationship model according to claim 2, characterized in that, The method further includes: The cloud server uses the first and last two frames of the multimodal video data sample as two key frames to obtain the manual annotation results of the key frames. The manual annotation results include the bounding boxes of the targets related to the target task in the first and last frames. The cloud server determines the direction of the vector of the center coordinate of the bounding box of the target related to the target task in the two key frames as the direction of movement of the target in the video frame. The cloud server determines the target's movement speed in the video frame by the ratio of the length of the vector of the center coordinates of the bounding box of the target related to the target task in the two key frames to the time interval between the two key frames. The cloud server uses the speed and direction of movement of the target related to the target in the video frame to annotate the bounding boxes of the target related to the target in multiple intermediate frames between the two keyframes. The cloud server determines the target-attribute-relationship related to the target task in the multimodal video data sample based on the bounding boxes corresponding to the two key frames and the multiple intermediate frames, and uses it as the object layer semantic information flow sample.

5. The multimodal data semantic representation and compression method based on a hierarchical target-attribute-relationship model according to claim 2, characterized in that, The method further includes: The cloud server inputs the reconstructed multimodal video data into the event detection model to obtain the event detection results; If the presence of a target event is determined based on the event detection results, the cloud server sends a scheduling command to the drone to schedule the drone to focus on the target event for multi-view video acquisition. The cloud server continuously monitors the target event based on the multi-view video data collected by the drone.

6. The multimodal data semantic representation and compression method based on a hierarchical target-attribute-relationship model according to any one of claims 1 to 5, characterized in that, The edge server transmits one or more of the data layer semantic information stream, the visual layer semantic information stream, the object layer semantic information stream, and the event layer semantic information stream to the cloud server according to the semantic granularity requirements of the cloud server, including: Upon receiving a target scene event-level monitoring request from the cloud server, the edge server transmits the event-layer semantic information stream to the cloud server. When the cloud server determines that an abnormal event has occurred based on the event layer semantic information stream, it sends a target scene object-level monitoring request to the edge server. Upon receiving the target scene object-level monitoring request, the edge server transmits the object-layer semantic information stream to the cloud server. When the cloud server needs to analyze the cause of abnormal events, it sends target scene data-level and visual-level monitoring requests to the edge server. Upon receiving the data-level and visual-level monitoring requirements for the target scene, the edge server transmits the visual layer semantic information stream and the data layer semantic information stream to the cloud server.

7. A multimodal data semantic representation and compression system based on a hierarchical goal-attribute-relationship model, characterized in that, The system includes at least: terminals, edge servers, and cloud servers; the terminals include at least: remote sensing satellites, drones, and ground towers. The terminal is used to simultaneously capture images of the target scene using its camera, obtain multimodal video data, and send it to the edge server. The edge server is used for: The multimodal video data is acquired, including satellite video data, UAV video data, and ground tower video data. With the goal of preserving the original semantic information of the multimodal video data, pixel-level semantic encoding is performed on the multimodal video data to obtain a data layer semantic information stream; With the goal of characterizing the visual information of the foreground and background in the multimodal video data, visual-level semantic encoding is performed on the multimodal video data to obtain a visual layer semantic information stream; The multimodal video data is subjected to object-level semantic encoding to obtain an object-layer semantic information stream, which is used to characterize the target-attribute-relationship in the multimodal video data related to the target task. With the goal of characterizing the events contained in the multimodal video data that are related to the target task, event-level semantic encoding is performed on the multimodal video data to obtain an event-layer semantic information stream; Based on the semantic granularity requirements of the cloud server, one or more of the data layer semantic information stream, the visual layer semantic information stream, the object layer semantic information stream, and the event layer semantic information stream are transmitted to the cloud server.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the computer program is executed by the processor, it implements the multimodal data semantic representation and compression method based on the hierarchical target-attribute-relationship model as described in any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program thereon, characterized in that, When the computer program is executed by the processor, it implements the multimodal data semantic representation and compression method based on the hierarchical target-attribute-relationship model as described in any one of claims 1 to 6.

10. A computer program product, characterized in that, The method includes a computer program that, when executed by a processor, implements the multimodal data semantic representation and compression method based on a hierarchical target-attribute-relationship model as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Low-bandwidth crowd scene security monitoring method and system based on semantic coding and decoding

    CN116708725A

  • Task processing system and method based on OAR semantic knowledge base

    CN117194698A