An adaptive decoding method and device based on multi-modal heterogeneous data and a medium

CN120561862BActive Publication Date: 2026-08-28SHANDONG INSPUR ULTRA HD INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510710451.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-29
Publication Date
2026-08-28
Estimated Expiration
2045-05-29

AI Technical Summary

Technical Problem

[0004]本申请实施例提供了一种基于多模态异构数据的自适应解码方法、设备及介质,解决了现有技术中智能终端无法对多模态异构数据进行智能解码的技术问题

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120561862B_ABST
    Figure CN120561862B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on multi-modal heterogeneous data adaptive decoding method, equipment and medium, it is related to embedded artificial intelligence technical field, method includes: to multi-modal sensor is identified type, to determine sensor access driving parameter, and based on sensor access driving parameter, obtain multi-modal sensor data;Multi-modal sensor data is fused to multi-modal, to determine cross-modal fusion data;According to cross-modal fusion data, by data dynamic decoding, obtain sensor decoding data;Obtain sensor environment perception data, and based on sensor environment perception data and sensor decoding data, by load closed loop optimization, determine load optimization update strategy.The application solves the technical problem that intelligent terminal cannot be decoded to multi-modal heterogeneous data in prior art by the above method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of embedded artificial intelligence technology, and in particular to an adaptive decoding method, device and medium based on multimodal heterogeneous data. Background Technology

[0002] The integration of multimodal sensors in smart terminals involves the cross-integration of multiple fields, including sensor technology, signal processing, artificial intelligence, and embedded systems. With the popularization of IoT and AI technologies, the demand for multimodal sensor integration in smart terminals has increased dramatically.

[0003] Due to significant differences in data formats and interface protocols among different sensors, traditional decoding systems require customized development of adaptation modules, resulting in low hardware resource utilization. Furthermore, existing technologies for data analysis from multimodal sensors in smart terminals still have some shortcomings. On the one hand, existing data fusion methods rely on deep networks with fixed structures, failing to optimize for the spatiotemporal characteristics of sensor data, leading to low terminal operating efficiency. On the other hand, existing technologies for smart terminals lack dynamic adjustment mechanisms, making it difficult to optimize decoding parameters based on network bandwidth or lighting conditions. Summary of the Invention

[0004] This application provides an adaptive decoding method, device, and medium based on multimodal heterogeneous data, which solves the technical problem that smart terminals in the prior art cannot intelligently decode multimodal heterogeneous data.

[0005] In a first aspect, embodiments of this application provide an adaptive decoding method based on multimodal heterogeneous data. The method includes: identifying the type of a multimodal sensor to determine sensor access driving parameters, and obtaining multimodal sensor data based on these parameters; performing multimodal fusion on the multimodal sensor data to determine cross-modal fused data; obtaining sensor decoded data through dynamic data decoding based on the cross-modal fused data; acquiring sensor environmental perception data, and determining a load optimization update strategy through load closed-loop optimization based on the sensor environmental perception data and the sensor decoded data.

[0006] In one implementation of this application, type identification of the multimodal sensor is performed to determine the sensor access driving parameters. Specifically, this includes: performing data access to the multimodal sensor through interface adaptation analysis to determine the sensor type; and determining the sensor access driving parameters based on the sensor type through driving parameter configuration.

[0007] In one implementation of this application, multimodal fusion is performed on multimodal sensor data to determine cross-modal fused data. Specifically, this includes: performing cross-modal alignment on the multimodal sensor data to obtain aligned multimodal sensor data; performing feature dimensionality reduction on the aligned multimodal sensor data, and then performing sparse voxelization on the dimensionality-reduced data to obtain effective multimodal sensor data; determining multimodal sensor feature data based on the effective multimodal sensor data through bi-branch Transformer feature analysis; wherein the bi-branch Transformer feature analysis includes: intra-modal Transformer feature extraction and cross-modal attention weight fusion; and performing three-level data fusion on the multimodal sensor feature data to determine cross-modal fused data; wherein the three-level data fusion includes: feature concatenation, channel weighting, and graph network aggregation.

[0008] In one implementation of this application, sensor decoded data is obtained by dynamically decoding the cross-modal fusion data. Specifically, this includes: determining the data format of the cross-modal fusion data to identify the cross-modal fusion data format; determining the decoding algorithm parameters by dynamically configuring FPGA logic based on the cross-modal fusion data format; and decoding the cross-modal fusion data according to the decoding algorithm parameters to obtain the sensor decoded data.

[0009] In one implementation of this application, based on the cross-modal fusion data format, the decoding algorithm parameters are determined by dynamic configuration of FPGA logic. Specifically, this includes: performing FPGA decoding analysis on the cross-modal fusion data format to determine the FPGA decoding logic parameters; synchronizing the FPGA decoding logic parameters to the ASIC fixed function unit to obtain a multi-format compatible interface; and determining the decoding algorithm parameters based on the multi-format compatible interface through decoding algorithm configuration.

[0010] In one implementation of this application, a load optimization update strategy is determined based on sensor environmental perception data and sensor decoded data through load closed-loop optimization. Specifically, this includes: determining the optimization execution condition based on the sensor environmental perception data to determine the optimization execution type; wherein the execution condition determination includes: light intensity determination and bandwidth determination; obtaining the load optimization strategy through dynamic parameter adjustment based on the optimization execution type; determining the dynamic power management state through DVFS control according to the load optimization strategy; and performing closed-loop feedback updates on the dynamic power management state to determine the load optimization update strategy.

[0011] In one implementation of this application, a closed-loop feedback update is performed on the dynamic power management state to determine a load optimization update strategy. Specifically, this includes: feeding back the dynamic power management state to a preset fuzzy controller to obtain first feedback data; determining second feedback data based on the first feedback data through frame rate adjustment analysis; and determining a load optimization update strategy based on the first feedback data and the second feedback data through a closed-loop strategy update.

[0012] In one implementation of this application, after determining the load optimization update strategy based on sensor environmental perception data and sensor decoding data through load closed-loop optimization, the method further includes: obtaining periodic load optimization data through periodic load monitoring based on the load optimization update strategy; evaluating the optimization status of the periodic load optimization data to determine the load optimization trend; and obtaining load optimization parameters by adjusting the load closed-loop optimization parameters according to the load optimization trend.

[0013] Secondly, embodiments of this application also provide an adaptive decoding device based on multimodal heterogeneous data, characterized in that the device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to: perform type identification on multimodal sensors to determine sensor access driving parameters, and obtain multimodal sensor data based on the sensor access driving parameters; perform multimodal fusion on the multimodal sensor data to determine cross-modal fused data; obtain sensor decoded data through dynamic data decoding based on the cross-modal fused data; acquire sensor environmental perception data, and determine a load optimization update strategy through load closed-loop optimization based on the sensor environmental perception data and the sensor decoded data.

[0014] Thirdly, embodiments of this application also provide a non-volatile computer storage medium for adaptive decoding based on multimodal heterogeneous data, storing computer-executable instructions. The computer-executable instructions are characterized by: performing type identification on multimodal sensors to determine sensor access driving parameters, and obtaining multimodal sensor data based on the sensor access driving parameters; performing multimodal fusion on the multimodal sensor data to determine cross-modal fused data; obtaining sensor decoded data through dynamic data decoding based on the cross-modal fused data; acquiring sensor environmental perception data, and determining a load optimization update strategy through load closed-loop optimization based on the sensor environmental perception data and the sensor decoded data.

[0015] This application provides an adaptive decoding method, device, and medium based on multimodal heterogeneous data. By calling driving parameters based on automatic sensor type identification, multimodal data fusion, and adaptive closed-loop optimization, it solves the technical problem that smart terminals cannot intelligently decode multimodal heterogeneous data in the prior art. It realizes low-latency decoding and fusion of multimodal sensor data, improves the environmental adaptability and energy efficiency of smart terminals, and optimizes the compatibility of multimodal data processing. Attached Figure Description

[0016] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 A flowchart of an adaptive decoding method based on multimodal heterogeneous data provided in this application embodiment; Figure 2 A system architecture diagram for adaptive decoding based on multimodal heterogeneous data is provided in this application embodiment; Figure 3 A schematic diagram of a dynamic decoding engine logic unit configuration provided in an embodiment of this application; Figure 4 This is a schematic diagram of feature extraction and data fusion provided in an embodiment of this application; Figure 5 A flowchart of a fuzzy control algorithm for a self-calibration module provided in an embodiment of this application; Figure 6 This is a schematic diagram of the internal structure of an adaptive decoding device based on multimodal heterogeneous data, provided in an embodiment of this application. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0018] This application provides an adaptive decoding method, device, and medium based on multimodal heterogeneous data. By calling driving parameters based on automatic sensor type identification, multimodal data fusion, and adaptive closed-loop optimization, it solves the technical problem that smart terminals cannot intelligently decode multimodal heterogeneous data in the prior art. It realizes low-latency decoding and fusion of multimodal sensor data, improves the environmental adaptability and energy efficiency of smart terminals, and optimizes the compatibility of multimodal data processing.

[0019] The technical solutions proposed in the embodiments of this application will be described in detail below with reference to the accompanying drawings.

[0020] Figure 1 This application provides an embodiment of an adaptive decoding flowchart based on multimodal heterogeneous data. Figure 1 As shown in the figure, the adaptive decoding method based on multimodal heterogeneous data provided in this application embodiment specifically includes the following steps: Step 101: Perform type identification on the multimodal sensor to determine the sensor access driving parameters, and obtain multimodal sensor data based on the sensor access driving parameters.

[0021] For example, due to the significant differences in data formats and interface protocols (I²C / SPI) of different sensors, the compatibility of sensors in the prior art is subject to many limitations. This application performs type identification of multimodal sensors to determine the sensor access driving parameters, and obtains multimodal sensor data based on the sensor access driving parameters, thereby achieving efficient compatibility of multimodal sensors.

[0022] Type identification of multimodal sensors is performed to determine the sensor access driving parameters. Specifically, this includes: performing data access to multimodal sensors through interface adaptation analysis to determine the sensor type; and determining the sensor access driving parameters based on the sensor type through driving parameter configuration.

[0023] In one embodiment, the data formats and interface protocols (I²C / SPI) of different sensors differ significantly. Multimodal sensors are connected to the system through an interface adapter layer, which automatically identifies the sensor type and loads the corresponding sensor access drive parameters.

[0024] By accessing the sensor's driving parameters and calling the interface that meets the interface protocol and data format of the sensor's access data, multimodal sensor data can be obtained.

[0025] Furthermore, by automatically identifying sensor types through the interface layer, it can support plug-and-play multimodal sensors such as cameras, infrared / TOF / radar, and realize device ID recognition and driver configuration through automatic enumeration and pre-stored parameter library of the interface adapter layer.

[0026] Step 102: Perform multimodal fusion on the multimodal sensor data to determine the cross-modal fusion data.

[0027] For example, existing data fusion methods rely on deep networks with fixed structures and are not optimized for the spatiotemporal characteristics of sensor data; fusing camera and radar data requires multiple feature alignment and dimensionality reduction operations, resulting in low computational efficiency. This application utilizes multimodal fusion, combining sparse voxelization of multimodal sensor data, bi-branch Transformer feature analysis, and three-level data fusion, to specifically enhance key information in cross-modal fused data, eliminate useless information, thereby reducing computational overhead and improving the computational efficiency of smart terminals.

[0028] Specifically, multimodal fusion is performed on multimodal sensor data to determine cross-modal fused data. This includes: cross-modal alignment of multimodal sensor data to obtain aligned multimodal sensor data; feature dimensionality reduction of the aligned multimodal sensor data, followed by sparse voxelization of the dimensionality-reduced data to obtain effective multimodal sensor data; determination of multimodal sensor feature data based on the effective multimodal sensor data through bi-branch Transformer feature analysis; wherein the bi-branch Transformer feature analysis includes: intra-modal Transformer feature extraction and cross-modal attention weight fusion; and three-level data fusion of the multimodal sensor feature data to determine cross-modal fused data; wherein the three-level data fusion includes: feature concatenation, channel weighting, and graph network aggregation.

[0029] Figure 2 This application provides a system architecture diagram for adaptive decoding based on multimodal heterogeneous data; wherein, multimodal fusion is achieved through a multimodal data fusion module. Figure 4 This is a schematic diagram of feature extraction and data fusion provided in an embodiment of this application.

[0030] In one embodiment, the multimodal sensor data is first aligned across modalities to obtain multimodal sensor aligned data; the multimodal sensor aligned data is then subjected to feature dimensionality reduction processing, and the data after feature dimensionality reduction is then subjected to sparse voxelization processing.

[0031] By using sparse voxelization, point cloud and image data are uniformly encoded into a sparse voxel mesh, processing only the regions containing valid data, thus reducing feature dimension and memory usage.

[0032] Then, multimodal sensor feature data are determined through bi-branch Transformer feature analysis. Bi-branch Transformer feature analysis includes: intra-modal Transformer feature extraction and cross-modal attention weight fusion. The intra-modal branch extracts spatiotemporal features from a single sensor (such as camera images or radar point clouds); the cross-modal branch uses an attention mechanism to model the correlation between multi-source data (such as images and point clouds), reducing the computational cost of feature alignment.

[0033] Finally, a three-level fusion strategy (feature splicing / channel weighting / graph network aggregation) is adopted to output the fusion results and determine the cross-modal fusion data.

[0034] Step 103: Based on the cross-modal fusion data, obtain the sensor decoded data through dynamic data decoding.

[0035] For example, the decoding engine dynamically configures FPGA logic units according to data types (such as H.265 video streams or radar point clouds) to perform parallel decoding. Through a dynamically reconfigurable hardware architecture that coordinates FPGA and ASIC, it achieves low-latency decoding and multi-format parallel processing of data, thereby improving the efficiency of data decoding.

[0036] Specifically, based on the cross-modal fusion data, sensor decoded data is obtained through dynamic data decoding. This includes: determining the data format of the cross-modal fusion data to identify the cross-modal fusion data format; determining the decoding algorithm parameters through dynamic configuration of FPGA logic based on the cross-modal fusion data format; and decoding the cross-modal fusion data according to the decoding algorithm parameters to obtain sensor decoded data.

[0037] Furthermore, based on the cross-modal fusion data format, the decoding algorithm parameters are determined through dynamic configuration of FPGA logic. Specifically, this includes: performing FPGA decoding analysis on the cross-modal fusion data format to determine the FPGA decoding logic parameters; synchronizing the FPGA decoding logic parameters to the ASIC fixed functional unit to obtain a multi-format compatible interface; and determining the decoding algorithm parameters through decoding algorithm configuration based on the multi-format compatible interface.

[0038] Figure 3 This is a schematic diagram of a dynamic decoding engine logic unit configuration provided in an embodiment of this application.

[0039] In one embodiment, the FPGA programmable unit adopts a modular hardware design, supporting dynamic switching between video formats such as H.264 / H.265 / AV1. Its core modules include motion estimation, intra-frame prediction, and entropy coding / decoding units, which adjust algorithm parameters (such as quantization step size and search window size) in real time through configuration files.

[0040] The ASIC fixed-function unit integrates general-purpose computing circuits such as DCT / IDCT transformation and quantization / dequantization, improves energy efficiency through dedicated hardware acceleration, and works in collaboration with FPGA through a standardized interface.

[0041] The interface adapter layer has a built-in I²C / SPI bus controller to realize automatic enumeration of multimodal sensors and complete plug-and-play configuration.

[0042] The dynamic encoding / decoding logic loading layer in the FPGA programmable unit synchronizes parameters with the fixed algorithm acceleration engine of the ASIC fixed function unit. That is, the FPGA performs decoding analysis on the cross-modal fusion data format to determine the FPGA decoding logic parameters and determine whether the data can be accelerated by the fixed algorithm.

[0043] The core encoding module interacts with multi-format compatible interfaces, with the data exchange sourced from external configuration or sensor data via an interface adaptation layer. The dynamically reconfigurable decoding engine supports parallel processing of multiple formats such as H.264 / H.265 / AV1, and by combining sparse voxelization technology with a lightweight feature fusion strategy, the decoding latency is controlled to within 8ms.

[0044] Step 104: Obtain sensor environmental perception data, and based on the sensor environmental perception data and sensor decoding data, determine the load optimization update strategy through load closed-loop optimization.

[0045] For example, Specifically, based on sensor environmental perception data and sensor decoded data, a load optimization update strategy is determined through load closed-loop optimization. This includes: determining the optimization execution condition based on the sensor environmental perception data to identify the optimization execution type; the execution condition determination includes: light intensity determination and bandwidth determination; obtaining the load optimization strategy through dynamic parameter adjustment based on the optimization execution type; determining the dynamic power management state through DVFS control according to the load optimization strategy; and performing closed-loop feedback updates on the dynamic power management state to determine the load optimization update strategy.

[0046] Furthermore, a closed-loop feedback update is performed on the dynamic power management status to determine the load optimization update strategy. Specifically, this includes: feeding back the dynamic power management status to a preset fuzzy controller to obtain first feedback data; determining second feedback data based on the first feedback data through frame rate adjustment analysis; and determining the load optimization update strategy through closed-loop strategy update based on the first and second feedback data.

[0047] Furthermore, after determining the load optimization update strategy based on sensor environmental perception data and sensor decoding data through load closed-loop optimization, the method also includes: obtaining periodic load optimization data through periodic load monitoring based on the load optimization update strategy; evaluating the optimization status of the periodic load optimization data to determine the load optimization trend; and obtaining load optimization parameters by adjusting the load closed-loop optimization parameters according to the load optimization trend.

[0048] Figure 5 A flowchart of a fuzzy control algorithm for a self-calibration module provided in an embodiment of this application.

[0049] In one embodiment, firstly, when the light sensor detects that the ambient brightness is <50 lux or the network bandwidth is <10 Mbps, the resolution reduction mode is activated; Then, based on environmental perception, dynamic adjustments are made: the down-resolution module reduces the video resolution from 1080p to 720p to reduce data transmission; the fuzz controller dynamically adjusts the frame rate based on real-time latency (threshold 8ms); and the DVFS controller reduces the processor frequency and voltage according to load requirements.

[0050] Finally, the adjusted system status parameters are fed back to the controller, forming a continuous optimization loop. Dynamic power management feeds back the power consumption status to the fuzzy controller, which dynamically controls the resolution reduction parameters of the down-resolution mode execution module based on the execution results and power consumption feedback. Based on the fuzzy control algorithm and DVFS technology, the resolution, frame rate, and processor frequency are dynamically adjusted according to illumination and network bandwidth, resulting in an overall power consumption reduction of over 40% compared to traditional solutions.

[0051] This application addresses industry pain points such as poor compatibility, computational redundancy, and insufficient environmental adaptability in multimodal data processing. It can be widely applied to edge computing scenarios such as intelligent security (real-time monitoring of multiple sensors), autonomous driving (fusion perception of LiDAR and cameras), and industrial inspection (high-precision visual quality inspection), providing a feasible technical solution for efficient intelligent perception in complex environments.

[0052] The above are embodiments of the method proposed in this application. Based on the same inventive concept, embodiments of this application also provide an adaptive decoding device based on multimodal heterogeneous data, the structure of which is as follows: Figure 6 As shown.

[0053] Figure 6 This is a schematic diagram of the internal structure of an adaptive decoding device based on multimodal heterogeneous data, provided as an embodiment of this application. Figure 6 As shown, the device includes: At least one processor 601; And a memory 602 that is communicatively connected to at least one processor; The memory 602 stores instructions executable by at least one processor, which are executed by at least one processor 601 to enable at least one processor 601 to: Multimodal sensor type identification is performed to determine sensor access driving parameters, and multimodal sensor data is obtained based on sensor access driving parameters; multimodal fusion is performed on multimodal sensor data to determine cross-modal fusion data; sensor decoded data is obtained through dynamic data decoding based on cross-modal fusion data; sensor environmental perception data is acquired, and load optimization update strategy is determined through load closed-loop optimization based on sensor environmental perception data and sensor decoded data.

[0054] Some embodiments of this application provide corresponding to Figure 1 A non-volatile computer storage medium for adaptive decoding of multimodal heterogeneous data, storing computer-executable instructions, wherein the computer-executable instructions are configured as follows: Multimodal sensor type identification is performed to determine sensor access driving parameters, and multimodal sensor data is obtained based on sensor access driving parameters; multimodal fusion is performed on multimodal sensor data to determine cross-modal fusion data; sensor decoded data is obtained through dynamic data decoding based on cross-modal fusion data; sensor environmental perception data is acquired, and load optimization update strategy is determined through load closed-loop optimization based on sensor environmental perception data and sensor decoded data.

[0055] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments for IoT devices and media are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0056] The systems, media, and methods provided in this application are one-to-one correspondences. Therefore, the systems and media also have similar beneficial technical effects as their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the systems and media will not be repeated here.

[0057] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0058] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0059] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0060] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0061] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0062] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0063] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0064] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0065] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. An adaptive decoding method based on multimodal heterogeneous data, characterized in that, The method includes: The multimodal sensor is type-identified to determine the sensor access driving parameters, and multimodal sensor data is obtained based on the sensor access driving parameters. Multimodal fusion is performed on the multimodal sensor data to determine cross-modal fused data; Based on the cross-modal fusion data, sensor decoded data is obtained through dynamic data decoding; Acquire sensor environmental perception data, and based on the sensor environmental perception data and the sensor decoded data, determine the load optimization update strategy through load closed-loop optimization; Type identification of multimodal sensors is performed to determine the sensor access driving parameters, specifically including: An I²C / SPI bus controller is built into the interface adaptation layer. Through interface adaptation analysis, the multimodal sensor is connected to the data to determine the sensor type. Based on the sensor type, the driver parameters are configured and the sensor access driver parameters are determined through the automatic enumeration and pre-stored parameter library of the interface adapter layer.

2. The adaptive decoding method based on multimodal heterogeneous data according to claim 1, characterized in that, Multimodal fusion of the multimodal sensor data is performed to determine cross-modal fusion data, specifically including: The multimodal sensor data is aligned across modes to obtain multimodal sensor aligned data. The multimodal sensor aligned data is subjected to feature dimensionality reduction processing, and the data after feature dimensionality reduction processing is subjected to sparse voxelization processing to obtain effective multimodal sensor data. Based on the effective data from the multimodal sensor, the feature data of the multimodal sensor is determined through bi-branch Transformer feature analysis; wherein, the bi-branch Transformer feature analysis includes: intra-modal Transformer feature extraction and cross-modal attention weight fusion; The feature data of the multimodal sensors are fused in three levels to determine cross-modal fused data; wherein the three-level data fusion includes: feature stitching, channel weighting, and graph network aggregation.

3. The adaptive decoding method based on multimodal heterogeneous data according to claim 1, characterized in that, Based on the cross-modal fusion data, sensor decoded data is obtained through dynamic data decoding, specifically including: The cross-modal fusion data is subjected to data format determination to determine the cross-modal fusion data format; Based on the aforementioned cross-modal fusion data format, the decoding algorithm parameters are determined through dynamic configuration of FPGA logic; Based on the decoding algorithm parameters, the cross-modal fusion data is decoded to obtain the sensor decoded data.

4. The adaptive decoding method based on multimodal heterogeneous data according to claim 1, characterized in that, Based on the aforementioned cross-modal fusion data format, the decoding algorithm parameters are determined through dynamic configuration of FPGA logic, specifically including: FPGA decoding analysis is performed on the cross-modal fusion data format to determine the FPGA decoding logic parameters; The FPGA decoding logic parameters are synchronized to the ASIC fixed function unit to obtain a multi-format compatible interface; Based on the multi-format compatible interface, the decoding algorithm parameters are determined through decoding algorithm configuration.

5. The adaptive decoding method based on multimodal heterogeneous data according to claim 1, characterized in that, Based on the sensor environmental perception data and the sensor decoded data, a load optimization update strategy is determined through load closed-loop optimization, specifically including: The sensor's environmental perception data is used to determine the optimization execution conditions to identify the optimization execution type; wherein, the execution condition determination includes: light intensity determination and bandwidth determination; Based on the aforementioned optimized execution type, a load optimization strategy is obtained through dynamic parameter adjustment; Based on the load optimization strategy, the dynamic power management state is determined through DVFS control; The dynamic power management status is updated via closed-loop feedback to determine the load optimization update strategy.

6. The adaptive decoding method based on multimodal heterogeneous data according to claim 5, characterized in that, The dynamic power management state is updated via closed-loop feedback to determine the load optimization update strategy, specifically including: The dynamic power consumption management status is fed back to a preset fuzzy controller to obtain the first feedback data; Based on the first feedback data, the second feedback data is determined through frame rate adjustment analysis; Based on the first feedback data and the second feedback data, the load optimization update strategy is determined through closed-loop strategy updates.

7. The adaptive decoding method based on multimodal heterogeneous data according to claim 1, characterized in that, After determining the load optimization update strategy based on the sensor environment perception data and the sensor decoded data through load closed-loop optimization, the method further includes: Based on the aforementioned load optimization and update strategy, periodic load optimization data is obtained through periodic load monitoring. The periodic load optimization data is used to evaluate the optimization status in order to determine the load optimization trend; Based on the load optimization trend, load optimization parameters are obtained by adjusting the load closed-loop optimization parameters.

8. An adaptive decoding device based on multimodal heterogeneous data, characterized in that, The device includes: At least one processor; And, a memory communicatively connected to the at least one processor; The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enable the at least one processor to: The multimodal sensor is type-identified to determine the sensor access driving parameters, and multimodal sensor data is obtained based on the sensor access driving parameters. Multimodal fusion is performed on the multimodal sensor data to determine cross-modal fused data; Based on the cross-modal fusion data, sensor decoded data is obtained through dynamic data decoding; Acquire sensor environmental perception data, and based on the sensor environmental perception data and the sensor decoded data, determine the load optimization update strategy through load closed-loop optimization; Type identification of multimodal sensors is performed to determine the sensor access driving parameters, specifically including: An I²C / SPI bus controller is built into the interface adaptation layer. Through interface adaptation analysis, the multimodal sensor is connected to the data to determine the sensor type. Based on the sensor type, the driver parameters are configured and the sensor access driver parameters are determined through the automatic enumeration and pre-stored parameter library of the interface adapter layer.

9. A non-volatile computer storage medium for adaptive decoding of multimodal heterogeneous data, storing computer-executable instructions, characterized in that, The computer-executable instructions are set as follows: The multimodal sensor is type-identified to determine the sensor access driving parameters, and multimodal sensor data is obtained based on the sensor access driving parameters. Multimodal fusion is performed on the multimodal sensor data to determine cross-modal fused data; Based on the cross-modal fusion data, sensor decoded data is obtained through dynamic data decoding; Acquire sensor environmental perception data, and based on the sensor environmental perception data and the sensor decoded data, determine the load optimization update strategy through load closed-loop optimization; Type identification of multimodal sensors is performed to determine the sensor access driving parameters, specifically including: An I²C / SPI bus controller is built into the interface adaptation layer. Through interface adaptation analysis, the multimodal sensor is connected to the data to determine the sensor type. Based on the sensor type, the driver parameters are configured and the sensor access driver parameters are determined through the automatic enumeration and pre-stored parameter library of the interface adapter layer.

Citation Information

Patent Citations

  • Multi-modal fusion method for heterogeneous data of intelligent networked vehicle multi-source sensor

    CN118445748A

  • Multi-modal fusion

    US20220405578A1