Method and apparatus for co-optimization of software and hardware for multi-modal perception and communication systems
By integrating multimodal data and dynamically adjusting resources, the problem of low resource utilization efficiency in multimodal sensing systems was solved, achieving hardware and software co-optimization and improving sensing accuracy and communication performance.
Patent Information
- Application Number
- CN202510161468.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2026-01-16
- Estimated Expiration
- 2045-02-13
AI Technical Summary
Existing multimodal sensing systems lack coordinated optimization of hardware and software, resulting in low resource utilization efficiency when sensing, processing and transmitting information, and making it difficult to achieve dynamic adjustment and power consumption optimization.
Multiple modal data are collected by various sensor modules, and the data are integrated into a single representation using a pre-defined inter-modal data fusion algorithm. Through pattern recognition and data prediction processing, a resource adjustment strategy is formulated to achieve dynamic adjustment of hardware and software resources.
It improves the system's perception accuracy and robustness in complex environments, optimizes the coordinated use of hardware and software resources, solves the reliability and latency control problems of communication links, and achieves a high-efficiency, low-power, and low-latency operating state.
Smart Images

Figure CN119966820B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, and particularly relates to a hardware and software collaborative optimization method and device of a multi-modal perception and communication system. BACKGROUND
[0002] In recent years, multi-modal perception and communication systems have been widely applied in the fields of artificial intelligence, Internet of Things, autonomous driving, etc. Such systems can achieve more comprehensive data acquisition and information transmission by fusing multiple perception modes (such as vision, hearing, touch, etc.) and communication methods. However, there are still many challenges in the design and application of existing multi-modal systems.
[0003] For example, existing multi-modal perception systems are often designed separately at the hardware and software levels, lacking sufficient collaborative optimization. This leads to efficiency bottlenecks in the system when perceiving, processing, and transmitting information, and the hardware resources cannot be fully utilized to improve overall performance. In addition, the limitations of existing multi-modal systems are also obvious in specific application scenarios. For example, in the field of autonomous driving, multi-modal perception systems need to process data from multiple sensors such as cameras, radars, and lidar in real time. However, due to the lack of close collaboration between hardware and software in existing systems, information processing delays often occur, affecting the timeliness and reliability of decision-making. Similarly, in smart home applications, the response speed of multi-modal systems and the processing efficiency of multi-sensor information are limited, making it difficult to achieve precise control in complex scenarios.
[0004] In addition, current hardware and software collaborative optimization faces several specific technical challenges. First, real-time issues, multi-modal systems need to process and fuse data from multiple sensors in real time, which puts high demands on the computing speed and response time of the system. The lack of collaboration between hardware and software often leads to processing delays. Second, power consumption issues, multi-modal perception and communication systems often need to run in resource-constrained environments, so how to reduce energy consumption while ensuring performance is an important challenge. Existing systems often have high energy consumption due to the lack of dynamic management of hardware resources, which cannot meet the needs of mobile devices or other low-power scenarios. Third, the complexity of resource allocation is also a major problem. Multi-modal systems involve multiple perception and communication tasks, and different tasks have different demands for hardware resources. How to reasonably allocate resources among multiple tasks and dynamically adjust according to the running state is a problem that needs to be solved urgently.
[0005] Current system design methods are usually limited to simply allocating resources between perception tasks and communication tasks without considering the different modal data processing requirements and bandwidth limitations at the hardware level, and cannot dynamically adjust hardware resources at the software level, resulting in the overall efficiency and power consumption of the system cannot be optimized. These limitations severely restrict the performance of multi-modal perception and communication systems in high real-time and high performance applications. SUMMARY
[0006] The main purpose of the present application is to provide a multi-modal perception and communication system software and hardware co-optimization method and device to solve the technical problems of the lack of hardware and software co-optimization in existing multi-modal perception systems, resulting in low resource utilization efficiency when perceiving, processing and transmitting information, and difficulty in achieving dynamic adjustment and power optimization.
[0007] To achieve the above purpose, the present application provides a multi-modal perception and communication system software and hardware co-optimization method, comprising the following steps: acquiring data through a plurality of sensor modules to obtain a plurality of modal data, wherein each sensor module corresponds to a different perception mode, and the perception mode at least includes any of the following: vision, hearing and touch; using a preset inter-modal data fusion algorithm to integrate the plurality of modal data into the same representation to obtain modal set data; performing pattern recognition and data prediction processing on the modal set data to obtain resource adjustment strategies for various modal data; dynamically adjusting the software and hardware resources according to the resource adjustment strategies for various modal data, and processing the data collected by the plurality of sensor modules based on the dynamically adjusted software and hardware resources.
[0008] The present application also provides a multi-modal perception and communication system software and hardware co-optimization device, comprising: an acquisition unit for acquiring data through a plurality of sensor modules to obtain a plurality of modal data, wherein each sensor module corresponds to a different perception mode, and the perception mode at least includes any of the following: vision, hearing and touch; an integration unit for using a preset inter-modal data fusion algorithm to integrate the plurality of modal data into the same representation to obtain modal set data; a prediction unit for performing pattern recognition and data prediction processing on the modal set data to obtain resource adjustment strategies for various modal data; an adjustment unit for dynamically adjusting the software and hardware resources according to the resource adjustment strategies for various modal data, and processing the data collected by the plurality of sensor modules based on the dynamically adjusted software and hardware resources.
[0009] The present application also provides a computer device comprising a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of the method of any one of the above.
[0010] The application further provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the steps of the method.
[0011] The application provides a multi-modal perception and communication system software and hardware collaborative optimization method and device, which integrates multi-modal data through a preset inter-modal data fusion algorithm, and performs mode recognition and data prediction processing on the integrated modal set data, so that the resource demand and dynamic change trend of different modal data can be accurately obtained, and an effective resource adjustment strategy can be formulated. The strategy is used for real-time dynamic adjustment of software and hardware resources, fully matches the processing demand of different modal data and the hardware bandwidth limit, and ensures efficient use of resources and minimization of power consumption. Based on the dynamically adjusted resources, the system can optimize the overall performance in the information acquisition, processing and transmission links, and improve the efficiency of the multi-modal perception and communication system. BRIEF DESCRIPTION OF DRAWINGS
[0012] Figure 1 is a multi-modal perception and communication system software and hardware collaborative optimization method step schematic diagram in an embodiment of the application;
[0013] Figure 2 is a multi-modal perception and communication system software and hardware collaborative optimization device structure block diagram in an embodiment of the application;
[0014] Figure 3 is a structure schematic block diagram of a computer device in an embodiment of the application.
[0015] The implementation of the object, functional features and advantages of the application will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION
[0016] In order to make the object, technical scheme and advantages of the application more clear, the application will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the application, and are not used to limit the application.
[0017] With reference to Figure 1 The embodiment of the application provides a multi-modal perception and communication system software and hardware collaborative optimization method, which comprises the following steps:
[0018] S1, data acquisition is performed through a plurality of sensor modules to obtain a plurality of modal data, wherein each sensor module corresponds to a different perception mode, and the perception mode at least comprises any one of the following: vision, hearing and touch.
[0019] S2, a plurality of modal data is integrated into the same representation by using a preset inter-modal data fusion algorithm to obtain modal set data.
[0020] S3, performing mode recognition and data prediction processing on the modal set data to obtain resource adjustment strategies for various modal data.
[0021] S4, dynamically adjusting the software and hardware resources according to the resource adjustment strategies of the various modal data, and processing the data collected by the multiple sensor modules based on the dynamically adjusted software and hardware resources.
[0022] That is, in the context of a multi-modal perception and communication system, the current technical challenges focus on perception accuracy, system robustness, software and hardware collaboration efficiency, and data transmission reliability, etc. The multi-modal perception and communication system software and hardware collaboration optimization method provided by the present application is proposed to solve these technical problems. Specifically, the technical solution collects data of different modalities through multiple sensor modules, integrates multiple perception modes into a unified representation through inter-modal data fusion algorithms, then performs pattern recognition and prediction processing on these data, and finally dynamically adjusts the software and hardware resources based on the analysis results to ensure that the system can maintain efficient operation in complex dynamic environments.
[0023] First, to address the problem of insufficient perception ability of the multi-modal system in complex environments mentioned in the background art, the present application collects data through multiple sensor modules, covering visual, auditory, tactile and other perception modes. This design directly improves the system's perception ability in various environments. Different perception modes can collect information from different angles, such as visual sensors that can capture objects in the environment, auditory sensors that can capture sound characteristics, and tactile sensors that can sense the contact feedback of objects. Under the joint action of these perception modes, the system can obtain more comprehensive perception information, thereby improving the accuracy of perception in complex environments. For example, in the automatic driving scenario, visual sensors can recognize vehicles or pedestrians in front, auditory sensors can detect the surrounding horn sounds, and tactile sensors can sense the road bumps. Through the fusion of multi-modal data, these information can form a more complete perception picture, thereby helping the automatic driving system to make more accurate decisions in complex environments.
[0024] Secondly, the existing multi-modal system has obvious deficiencies in software and hardware collaboration, often with separate hardware and software design, making it difficult to effectively optimize resources. To this end, the key of the present application is to integrate the data of multiple modalities into a unified modal set through inter-modal data fusion algorithms. This integration not only simplifies the processing flow of multi-modal data, but also provides a basis for subsequent resource adjustment strategies. By performing pattern recognition and data prediction on the modal set data, the present application can generate resource adjustment strategies according to the characteristics of each modal data, thereby dynamically optimizing the software and hardware resources. This dynamic adjustment can break the rigid mode of resource allocation between perception tasks and communication tasks in traditional systems, and truly realize the collaborative optimization of software and hardware. For example, the data processing demand of visual sensors is usually high, especially when high-resolution image processing is required, which will occupy a large amount of computing resources, while the data processing of auditory sensors is relatively light. When the system detects that the visual data flow is large, it can allocate more computing resources to the visual data processing unit according to the current resource adjustment strategy, while reducing the resource occupation of auditory processing, thereby improving the efficiency of the overall system.
[0025] In addition, it is pointed out in the background art that current systems often have reliability and latency control problems in data transmission due to the explosive growth of data volume. The technical solution provided by the present application realizes more efficient data processing and transmission through dynamic adjustment of software and hardware resources. Through resource adjustment strategies, the present application can prioritize communication bandwidth to critical data streams in the case of limited hardware resources, ensuring low latency and high reliability output of the system. For example, in autonomous driving, the system may need to process a large amount of data from visual, auditory and tactile sensors simultaneously. If the communication link is congested, the priority of visual data may be raised because it is crucial for environmental perception and driving decision-making. While the transmission priority of some secondary data such as environmental sound can be temporarily reduced to ensure real-time transmission of core perception data. This data processing and transmission mechanism with priority effectively solves the congestion problem in traditional communication links and ensures the overall response speed of the system.
[0026] In addition, in the process of dynamic adjustment of software and hardware, the present application does not simply allocate fixed resources between different modalities, but through data prediction and pattern recognition, it predicts future perception needs and optimizes resource allocation in advance. For example, when the system predicts that the environment ahead is complex, it can increase resource allocation to the visual sensor in advance to ensure real-time processing of high-resolution images; while in a simple environment, it can reduce the resource occupation of the visual sensor and improve the processing capacity of other sensors, thereby achieving the best balance between power consumption and performance.
[0027] Overall, the present application solves a plurality of key problems in the existing multi-modal perception and communication system through a hardware and software collaborative optimization method. Through the integration and dynamic adjustment of multi-modal data, the perception accuracy and robustness of the system in a complex environment are improved, the collaborative use of software and hardware resources is optimized, the reliability and delay control of the communication link are solved, and finally the system can still maintain an efficient, low-power and low-latency running state under the condition of limited hardware resources.
[0028] It should be noted that the commonly used parameters of multi-modal data can be as shown in the following table.
[0029]
[0030] In one example, a plurality of modal data is integrated into the same representation by using a preset inter-modal data fusion algorithm to obtain modal set data, including: determining the relative positions of each modal data in space by a calibration algorithm, and mapping the each modal data to a unified coordinate system based on the relative positions to obtain a plurality of first modal data; performing dynamic time warping processing on the plurality of first modal data to interpolate different first modal data to consistent time steps to obtain a plurality of second modal data that are completely aligned in time; performing feature extraction on the plurality of second modal data, and performing encoding processing on the features extracted from different second modal data to obtain feature data with a unified data format; calculating the Pearson correlation coefficients of the feature data corresponding to each second modal data, and evaluating the correlation of each feature data based on the Pearson correlation coefficients to obtain a nonlinear correlation relationship of each feature data; using a preset inter-modal data fusion algorithm, integrating the feature data into a shared representation space based on the nonlinear correlation relationship of each feature data to obtain preliminary modal set data; and continuing to perform back propagation processing on the preliminary modal set data by using the inter-modal data fusion algorithm to eliminate inconsistent features and repeated features in the preliminary modal set data to obtain final modal set data without information conflict and repetition.
[0031] That is, through this data fusion and processing method, the problems existing in the multi-modal perception system mentioned in the background art can be effectively solved, especially the resource allocation problem caused by inconsistent modal data. Specifically, the unified dimension data representation allows the system to compare the relative importance of each modal data on the same scale.
[0032] Traditional multi-modal perception systems are usually designed separately for perception and communication tasks, with low matching between perception data and hardware resources, resulting in the inefficiency of the overall system. However, through the above method, the system not only realizes the fusion of each modal data at the perception level, but also makes the allocation of software and hardware resources more flexible and accurate through unified feature representation. When the system can dynamically adjust the hardware resources according to the importance of different modal data and task requirements, not only can the existing hardware capabilities be fully utilized, but also the waste of resources can be greatly reduced, and the overall efficiency and power consumption of the system can be improved.
[0033] It should be noted that the current other software and hardware resource allocation techniques do not take these problems into account.
[0034] Specifically, first, the starting point of the whole process is to determine the relative positions of each modal data in space through a calibration algorithm. Data from different modalities may come from different sensors at different locations, so they must be mapped to the same spatial coordinate system. For example, in the autonomous driving scenario, the vision sensor may be installed on the roof to capture the scene in front of the car, the hearing sensor is located on both sides of the car body to collect the surrounding sound, and the touch sensor may be installed near the tire to sense the road conditions. Because of the different physical locations of these sensors, the information they collect is inconsistent in space. Through the calibration algorithm, we can determine their relative positions in space and map each modal data to the same unified coordinate system, thereby generating a plurality of first modal data. Taking autonomous driving as an example, the image data collected by the vision sensor, the audio data collected by the hearing sensor, and the vibration feedback data collected by the touch sensor will be mapped to the same coordinate system to ensure that these modal data can be effectively integrated and aligned in the subsequent data processing process.
[0035] Next, through dynamic time warping processing, the system can interpolate different first modal data to consistent time steps. Because the data sampling frequencies of different sensors may be different, for example, the frame rate of the vision sensor is 30 frames / second, while the sampling rate of the hearing sensor may be 44.1 kHz, there is a significant difference in the time steps of the two. If these modal data are not aligned in time, their data content will not be synchronized, resulting in information bias when fusion. Through dynamic time warping processing, the system can automatically calculate the time gap between modal data and interpolate low-frequency modal data to make all data consistent at the same time step. For example, in an autonomous driving system, a vision sensor may capture 30 frames of images per second, while a hearing sensor may capture thousands of sound samples per second. Through dynamic time warping processing, the vision data will be interpolated to align with the sound data in time, thereby generating second modal data that are completely consistent in time.
[0036] After the time alignment is completed, the next step is to perform feature extraction on the second modality data. Different modalities of data each contain rich information, but this information is usually in different formats. Therefore, the system needs to perform feature extraction on the data of each modality and further encode these features. The process of feature extraction is to extract the key information of the modality data according to their properties, for example, visual data may extract edge and color features of the image, auditory data extract frequency and waveform features, and tactile data extract vibration amplitude and contact pressure features. Through feature extraction and encoding, the data of different modalities are converted into feature data with a unified data format, so that they can be better integrated and analyzed in subsequent steps.
[0037] After the generation of feature data, the system further calculates the Pearson correlation coefficient between each modality data to evaluate the correlation of these feature data. The Pearson correlation coefficient can measure the degree of linear correlation between different modalities of features. Through this calculation, the system can find the association between different modalities of data, such as whether there is some connection between visual features and auditory features, or whether tactile features have certain correlation with environmental visual features. These correlations provide a reference for subsequent data fusion, so that the data of different modalities can be better integrated together.
[0038] Next, the system evaluates the nonlinear correlation association of each feature data based on the Pearson correlation coefficient of each feature data, and uses a preset inter-modal data fusion algorithm to integrate these feature data into a shared representation space. This shared representation space can be regarded as the "language" common to different modalities of data, which converts features from different modalities into a format that can be uniformly processed. Through this integration, the system can better understand the mutual relationship between different modalities of data and ensure that the final data fusion result is rich and consistent in information.
[0039] Finally, the system continues to perform backpropagation processing on the preliminary modal set data through the inter-modal data fusion algorithm. The process of backpropagation is mainly to eliminate inconsistent features and redundant features in the preliminary modal set data. Since different modalities of data sources may have redundant information or inconsistencies, for example, visual sensors and auditory sensors may collect similar information in some cases, or due to sensor noise, some feature data may be inaccurate, backpropagation can eliminate these redundant and incorrect information through iterative optimization, ensuring that the final modal set data is conflict-free and without duplication.
[0040] In one example, the resource adjustment strategy of the sensor module includes: differentiating the modal set data into different modal categories to obtain data subsets of different modal categories; performing time series analysis on the data subsets of different modal categories, and performing data prediction processing on different sensor modules based on the results of the time series analysis; determining the weight values of the modal data collected by the different sensor modules in the next time period based on the data prediction results of the different sensor modules; and determining the adjustment strategy of the different sensor modules based on the weight values of the modal data collected by the different sensor modules in the next time period.
[0041] The resource adjustment strategy of this sensor module can effectively solve the problems of resource allocation and performance optimization of the multi-modal perception system mentioned in the background art by subdividing the modal set data into different modal categories and performing time series analysis and prediction processing on these data. Specifically, this strategy first classifies the modal set data, meaning that after the unified representation of different perception modes (such as vision, hearing, and touch), each mode still retains its unique characteristics. By dividing these data into data subsets of different modal categories, the system can process and analyze each modal data separately, thereby having higher flexibility and accuracy in time series analysis and future data prediction. One significant advantage of this approach is that it allows the system to more accurately judge the future importance of each modal data, thereby guiding the dynamic adjustment of software and hardware resources.
[0042] In actual operation, by performing time series analysis on the data subsets of different modal categories, the system can predict in advance that certain modal data may occupy a larger proportion or become more critical in the future time period. For example, in the context of autonomous driving, visual data has a high weight under normal lighting conditions, but if the system discovers through time series analysis that the future road conditions enter a tunnel or night, the light conditions may deteriorate, and the accuracy of visual data will decrease, then the system can adjust the resource allocation by prediction to allocate more resources to auditory or radar sensors. This ability to pre-adjust resource allocation ensures that the system can be prepared in advance for environmental changes, rather than passively responding to unexpected situations.
[0043] Based on the prediction of the temporal analysis, the system can determine the weight values of the data collected by different sensor modules in the next time period. This process helps the system determine which modal data will be more important in the upcoming time period, thereby optimizing resource allocation. For example, in a multi-modal perception system, if auditory data predicts that environmental noise will become more complex and critical, the system can allocate more processing power and communication bandwidth to auditory sensors in advance to ensure that auditory data can be accurately acquired and processed during that time period. For modal data that is less important during the current time period, such as tactile data, the system can accordingly reduce resource allocation, thereby effectively saving hardware and software resources. This dynamic adjustment mechanism not only optimizes the efficiency of sensors, but also improves the overall performance of the system and reduces resource waste.
[0044] In addition, the adjustment strategy is optimized based on the prediction of future data, which can significantly improve the robustness and response speed of the system in handling complex perception scenarios. For example, in autonomous driving, if it is predicted that visual data may be affected by weather conditions (such as rain and fog) during a certain period, the system can allocate more processing power to radar sensors in advance to ensure that accurate perception of the road and obstacles can be maintained in bad weather. This strategy ensures that the system will not have a perception blind area due to the failure of a single modal data, thereby improving the stability and safety of the system in complex environments.
[0045] The key advantage of this method in solving the resource allocation problem in the background art is that it can dynamically adjust resources based on the future importance of modal data. Traditional multi-modal perception systems often cannot real-time re-allocate hardware and software resources according to environmental changes, resulting in some modal data occupying too many resources in actual scenarios, while other important modal data does not get enough computing power and communication bandwidth support. By analyzing and predicting the future importance of each modal data in a unified dimension, the system can more accurately allocate resources to different sensor modules. For example, when the system predicts that visual data will become critical in the future, it can increase resource allocation for visual processing, including processor time and storage bandwidth, to ensure the accuracy of visual perception. For modalities with lower future importance, the system can accordingly reduce resource investment, which can effectively avoid resource waste and ensure that high-priority tasks are fully supported.
[0046] This strategy can also have a positive impact on power consumption control. Since multi-modal perception systems are often deployed on power-sensitive devices (such as drones, smartphones, smart home devices, etc.), how to minimize energy consumption while ensuring accurate perception is a very important issue. By predicting the future importance of each modality data and dynamically adjusting resources, the system can avoid unnecessary computational and communication resource overhead, reducing power consumption. For example, in an indoor environment, the weight of tactile data may be lower than that of visual data, and the system can reduce the resource allocation of tactile sensors and reduce their energy consumption through power management mechanisms to improve the overall energy efficiency of the system.
[0047] In one example, the data subsets of different modality categories are subjected to time series analysis, and based on the results of the time series analysis, data prediction processing is performed on different sensor modules, including: dividing the data subsets of different modality categories into different time windows, wherein the data in each time window represents the modality information in a preset time period; extracting features from the data in each time window of each modality category to obtain first feature information of the data of different modality categories, and identifying the time series characteristics of the first feature information, wherein the time series characteristics at least include: periodicity and trend of the feature information; based on the time series characteristics of the first feature information, generating a data prediction model, and performing prediction analysis on the data subsets of different modality categories through the prediction model to determine the predicted data information of different sensor modules in a future time period.
[0048] Further, based on the data prediction results of the different sensor modules, the weight values of the modality data collected by the different sensor modules in the next time period are predicted, including: determining the weight values of different sensor modules based on the predicted data information of different sensor modules in a future time period; normalizing the weight values of the different sensor modules to integrate the first weight values of the different sensor modules into the interval range of 0-1 to obtain the weight values of the modality data collected by the sensor modules in the next time period.
[0049] In this example, first, time series analysis is performed on the data subsets of different modality categories, which is achieved by dividing these data into different time windows. Each time window contains modality information in a preset time period. The purpose of this division is to capture the changing trend and characteristic performance of the data in a specific time period. Then, feature extraction is performed on the data in each time window of each modality category to generate first feature information. This process aims to extract representative features from the data for further analysis of the time series characteristics of these features. Time series characteristics include periodicity and trend, which describe the variation law and pattern of data in the time dimension.
[0050] Next, based on the extracted feature information and its time series characteristics, a data prediction model is generated. This model analyzes the periodicity and trend of the features to predict the performance of different modal category data in the future time period. The purpose of the prediction analysis is to determine the data that each sensor module may generate in the future time period. This step ensures the prediction and planning of future modal data, so as to adjust the working mode and resource allocation of the sensor module in advance.
[0051] After obtaining the data prediction results of different sensor modules, the weight values of the modal data collected by the sensor modules in the next time period are further predicted. This weight value is determined according to the predicted data information of the sensor module in the future, reflecting the importance of a certain sensor module in the future time period. In order to make these weight values more comparable and consistent, they need to be normalized. Normalization integrates the weight values of each sensor module into the range of 0-1, ensuring that the weight values are expressed on the same scale. This enables the final determined weight values to be used to reasonably allocate sensor resources, thereby optimizing the collection and processing process of different modal data.
[0052] In one example, the hardware resources are dynamically adjusted according to the resource adjustment strategy of the various modal data, including: dividing the hardware resources to obtain first type hardware resources and second type hardware resources, wherein the first type hardware resources are quantifiable resources, and the second type hardware resources are non-quantifiable resources; determining the current allocation of the first type hardware resources, and based on the weight values of the modal data collected by the sensor modules in the next time period, re-allocating the first type hardware resources based on the current allocation of the first type hardware resources; determining a plurality of preset configurations of the second type hardware resources and a weight value range corresponding to each of the preset configurations, and based on the weight values of the modal data collected by the sensor modules in the next time period, reconfiguring the second type hardware resources.
[0053] In this example, the system dynamically adjusts the hardware resources according to the resource adjustment strategy of the modal data. This adjustment process can be divided into the processing of two types of hardware resources: the first type of hardware resources is quantifiable resources, and the second type of hardware resources is non-quantifiable resources.
[0054] The first type of soft and hardware resources, as quantifiable resources, refers to those hardware resources that can be accurately measured and allocated. For example, CPU processing power, memory capacity, bandwidth, and storage space, etc. This type of resource can explicitly determine its current allocation through numerical means. According to the weight values of the modal data collected by each sensor module of the system in the future time period, the system can dynamically adjust the allocation of these resources. Specifically, based on the current resource allocation state, the system will consider the importance of each sensor in the future and preferentially allocate more processing power, memory, bandwidth, etc. to those sensor modules with higher weights to improve overall performance. For example, if the vision sensor needs to process a large amount of video data in the future, the system will allocate more CPU and memory resources to it.
[0055] The second type of soft and hardware resources is non-quantifiable resources. This type of resource refers to those hardware or software configurations that cannot be simply quantified and accurately allocated, usually including configuration options of hardware structure, scheduling strategies, I / O interface priority, and selection of different algorithms, etc. Because the adjustment of non-quantifiable resources often depends on the system's preset multiple configuration options, the system first determines different preset configurations of these resources and the corresponding weight value range of each configuration. Then, according to the weight values of the sensor modules, the system will select the most suitable configuration combination in the preset configuration to reconfigure. For example, the priority of the I / O interface is a type of non-quantifiable resource. If the system discovers through data prediction that the data transmission demand of the haptic sensor will increase in the future, it may adjust the I / O interface priority to prioritize the data transmission speed and stability of the haptic sensor, rather than simply allocating more bandwidth to solve the problem.
[0056] In summary, the adjustment of the first type of soft and hardware resources is based on quantified allocation, such as allocating more CPU, memory, etc., while the second type of resource is based on configuration selection, optimizing the running mode of different hardware modules to improve system efficiency. Such dynamic adjustment strategy ensures that the system can effectively optimize resource allocation when facing different modal data, thereby improving the overall system's perception ability and communication performance.
[0057] Specifically, the current allocation of the first type of soft and hardware resources is determined, and based on the weight values of the modal data collected by the sensor modules in the next time period, the first type of soft and hardware resources is re-allocated based on the current allocation of the first type of soft and hardware resources, including: determining the current allocation of each soft and hardware resource in the first type of soft and hardware resources; based on the weight values of the modal data collected by the sensor modules in the next time period, the current allocation of each soft and hardware resource in the first type of soft and hardware resources is re-allocated.
[0058] That is, the system dynamically adjusts the allocation of the first type of software and hardware resources by analyzing the weight values of the modal data collected by the sensor modules in future time periods. Specifically, the first type of software and hardware resources includes quantifiable resources such as CPU, memory, bandwidth, and storage space. To ensure that the system can allocate sufficient resources to each sensor module in different time periods, it is necessary to first determine the current resource allocation, and then perform reallocation processing according to the future weight values of the sensor modules.
[0059] First, the system determines the current allocation of each software and hardware resource. For example, in one scenario, the visual sensor, auditory sensor, and tactile sensor are allocated 20%, 30%, and 50% of the CPU resources, respectively. This allocation is based on the current workload and historical data. However, over time, future perception needs may change. To address these changes, the system will adjust resource allocation by analyzing the weight values of the sensor modules in the next time period.
[0060] In a specific case, suppose through time series analysis and data prediction, the system learns that the visual sensor needs to handle a large amount of video stream in the future time period, and is expected to occupy more processing power. At the same time, the demand of the auditory sensor may decrease, while the tactile sensor remains relatively stable. In this case, the system will adjust the current resource allocation based on the future weight values of these sensors. Assuming that the weight value of the visual sensor increases from 0.2 to 0.5, the weight value of the auditory sensor decreases from 0.3 to 0.2, and the tactile sensor remains at 0.3, the system will reallocate CPU resources, allocating more processing power to the visual sensor while reducing the CPU occupancy of the auditory sensor.
[0061] This reallocation process is not limited to CPU resources, but also applies to memory and bandwidth. For example, if future predictions show that the visual sensor will handle higher resolution video data, the system can allocate more memory and bandwidth to it in advance, ensuring that video data processing and transmission will not be delayed or lag. For the auditory sensor, since its future data demand decreases, the system can accordingly reduce its memory and bandwidth allocation, allocating these resources to sensor modules with higher weights.
[0062] This dynamic adjustment mechanism ensures that the system can reasonably allocate quantifiable hardware resources according to the importance of sensor modules in different time periods, thereby improving the overall efficiency and perception ability of the system. Through this weight-based resource reallocation, the system not only meets the real-time needs of each sensor module, but also avoids resource waste, ensuring optimal performance in the case of limited hardware resources.
[0063] Specifically, a plurality of preset configurations of the second type of soft and hardware resources and a corresponding weight value range of each of the preset configurations are determined, and the second type of soft and hardware resources are reconfigured based on the weight values of the modal data collected by the respective sensor modules in the next time period, including: splitting the second type of soft and hardware resources into a first type of soft and hardware sub-resources and a second type of soft and hardware sub-resources, wherein the first type of soft and hardware sub-resources are used to match different preset configurations for different sensor modules, and the second type of soft and hardware sub-resources are used to match the same preset configuration for different sensor modules; determining a plurality of preset configurations of each soft and hardware resource in the first type of soft and hardware sub-resources, and matching a preset configuration for each soft and hardware resource in the first type of soft and hardware sub-resources for different sensor modules based on the weight values of the modal data collected by the respective sensor modules in the next time period; determining a plurality of preset configurations of each soft and hardware resource in the first type of soft and hardware sub-resources, weighting and summing the weight values of the modal data collected by the respective sensor modules in the next time period, and determining a preset configuration for each soft and hardware resource in the second type of soft and hardware sub-resources based on the weighted sum result.
[0064] It should be noted that the second type of soft and hardware resources refers to those non-quantifiable resources, such as system running strategy, sensor working mode, network connection quality or transmission protocol priority, etc. The adjustment of these resources cannot be as simple as numerical allocation as the first type of resources (such as CPU, memory, etc.), but needs to be flexibly configured based on specific strategies. In order to realize reasonable allocation of these non-quantifiable resources, the system will split these resources into two types: the first type of soft and hardware sub-resources and the second type of soft and hardware sub-resources.
[0065] First, the first type of soft and hardware sub-resources are used to match different preset configurations for different sensor modules. The so-called preset configuration refers to the running mode of a certain sensor module in a specific scenario. For example, a visual sensor may need to enable a specific image enhancement algorithm in a low light environment, and an auditory sensor may need to activate a noise filtering function in a noisy environment. Through time series analysis and future data prediction, the system can match the most suitable preset configuration for each sensor module based on the weight values of the modal data collected by the respective sensor modules in the next time period. For example, if it is predicted that the visual sensor will process a large amount of night video data in the future, the system will match a low light mode preset configuration for the sensor module to ensure that it can collect the clearest image in this specific environment.
[0066] In this process, each software and hardware sub-resource will have multiple preset configurations that provide different operating strategies for different sensor modules. For example, if the visual sensor is collecting daytime data in the future environment, it can choose a normal mode without wasting resources to enable a low-light mode. Therefore, based on the weight values of the modal data, the system will select the appropriate running configuration for each sensor module.
[0067] Next is the second type of software and hardware sub-resource, which is usually a resource that needs to match the same preset configuration for multiple sensor modules. For example, network bandwidth priority allocation or overall system energy management strategy, which usually provides uniform services for all sensor modules. In order to determine the allocation of these resources, the system will first perform weighted summation processing on the weight values of each sensor module. The purpose of weighted summation is to integrate the importance of future data of different sensor modules, so as to select the optimal configuration for the unified resource. For example, if the data weight of the visual sensor and the tactile sensor is high, and the data weight of the auditory sensor is low in the future time period, the system may prioritize higher network bandwidth for the visual and tactile sensors, or adjust the energy management strategy to ensure that these two sensor modules can work continuously.
[0068] In this way, the configuration of the second type of software and hardware resource is adjusted according to the weight of each sensor module. The first type of software and hardware sub-resource provides independent configuration for each sensor module to ensure that they can perform best in their respective working environments; while the second type of software and hardware sub-resource selects the optimal configuration for the shared resources of the entire system to ensure that multiple sensor modules work cooperatively under a unified strategy. Through this flexible resource configuration mechanism, the system can reasonably schedule the software and hardware resources in the future working scenario, and improve the overall efficiency of the sensor network.
[0069] As shown in Figure 2 The embodiment of the present application provides a multi-modal perception and communication system software and hardware collaborative optimization device, which comprises:
[0070] The acquisition unit 1 acquires data through multiple sensor modules to obtain multiple modal data, wherein each sensor module corresponds to a different perception mode, and the perception mode at least includes any of the following: vision, hearing and touch;
[0071] The integration unit 2 is used for integrating multiple modal data into the same representation by using a preset inter-modal data fusion algorithm to obtain modal set data;
[0072] The prediction unit 3 is used for mode recognition and data prediction processing on the modal set data to obtain resource adjustment strategies for various modal data;
[0073] An adjusting unit 4 is configured to dynamically adjust the software and hardware resources according to a resource adjustment strategy of the various modal data, and process the data collected by the various sensor modules based on the dynamically adjusted software and hardware resources.
[0074] In the embodiment, the implementation of each unit in the above device embodiment can refer to the description in the above method embodiment, and will not be repeated here.
[0075] With reference to Figure 3 In the embodiment, a computer device can be a server, and the internal structure of the computer device can be as shown in Figure 3 The computer device includes a processor, a memory, a display screen, an input device, a network interface and a database connected through a system bus. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium. The database of the computer device is configured to store corresponding data in the embodiment. The network interface of the computer device is configured to communicate with an external terminal through a network connection. The computer program is executed by the processor to implement the above method.
[0076] Those skilled in the art can understand Figure 3 that the structure shown in the embodiment is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied.
[0077] The embodiment of the present application also provides a computer readable storage medium having a computer program stored thereon, and the computer program is executed by the processor to implement the above method. It can be understood that the computer readable storage medium in the embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.
[0078] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiment methods. Any reference to memory, storage, database or other medium provided by the present application and used in the embodiments can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM, etc.
[0079] It should be noted that in this document, the terms "comprising", "including", or any other variant thereof are intended to cover a non-exclusive inclusion, such that a process, device, article or method that comprises a list of elements does not only include those elements, but can also include other elements not expressly listed or inherent to such process, device, article or method. Without more limitations, the element defined by the statement "comprising a" does not exclude the presence of additional identical elements in the process, device, article or method that includes the element.
[0080] The above description is only the preferred embodiment of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation, or direct or indirect application in other related technical fields, based on the content of the present application specification and drawings, are also included in the patent protection scope of the present application.
Claims
1. A method for hardware and software co-optimization of a multimodal sensing and communication system, characterized in that, The method comprises the following steps: Data acquisition is performed through various sensor modules to obtain various modal data, wherein each sensor module corresponds to different perception modes, and the perception modes at least include any of the following: vision, hearing and touch; A preset inter-modal data fusion algorithm is used to integrate the various modal data into the same representation to obtain modal set data; The modal set data is divided into different modal categories to obtain data subsets of different modal categories; time series analysis is performed on the data subsets of different modal categories, and data prediction processing is performed on different sensor modules based on the results of the time series analysis; the weight values of the modal data collected by the different sensor modules in the next time period are determined based on the data prediction results of the different sensor modules; and the adjustment strategy of the different sensor modules is determined based on the weight values of the modal data collected by the different sensor modules in the next time period; According to the resource adjustment strategy of various modal data, the software and hardware resources are dynamically adjusted, and the data collected by the various sensor modules are processed based on the dynamically adjusted software and hardware resources.
2. The method of claim 1, wherein, A preset inter-modal data fusion algorithm is used to integrate the various modal data into the same representation to obtain modal set data, comprising: The relative positions of various modal data in space are determined through a calibration algorithm, and the various modal data are mapped to a unified coordinate system based on the relative positions to obtain a plurality of first modal data; The plurality of first modal data are subjected to dynamic time warping processing, and different first modal data are interpolated to consistent time steps to obtain a plurality of second modal data that are completely aligned in time; Feature extraction is performed on the plurality of second modal data, and the features extracted from different second modal data are subjected to encoding processing to obtain feature data with a unified data format; Pearson correlation coefficients of the feature data corresponding to each second modal data are calculated, and the correlation of each feature data is evaluated based on the Pearson correlation coefficients to obtain a nonlinear correlation association relationship of each feature data; A preset inter-modal data fusion algorithm is used to integrate the feature data into a shared representation space based on the nonlinear correlation association relationship of each feature data to obtain preliminary modal set data; The preliminary modal set data is subjected to back propagation processing through the inter-modal data fusion algorithm to eliminate inconsistent features and repeated features in the preliminary modal set data to obtain final modal set data without information conflict and repetition.
3. The method of claim 1, wherein, Time series analysis is performed on the data subsets of different modal categories, and data prediction processing is performed on different sensor modules based on the results of the time series analysis, comprising: The data subsets of different modal categories are divided into different time windows, wherein the data in each time window represents modal information in a preset time period; extracting feature information of data in a time window of each modality category, to obtain first feature information of data of different modality categories, and identifying time sequence characteristics of the first feature information, wherein the time sequence characteristics at least include periodicity and trend of the feature information; generating a digital prediction model based on the time sequence characteristics of the first feature information, and performing prediction analysis on a subset of data of different modality categories through the prediction model to determine predicted data information of different sensor modules in a future time period.
4. The method of claim 1, wherein, Based on the data prediction results of the different sensor modules, determine the weight values of the modality data collected by the different sensor modules in the next time period, including: determining the weight values of different sensor modules based on the predicted data information of different sensor modules in the future time period; normalizing the weight values of the different sensor modules to integrate the first weight values of the different sensor modules into the interval range of 0-1, to obtain the weight values of the modality data collected by the sensor modules in the next time period.
5. The method of claim 1, wherein, According to the resource adjustment strategy of various modal data, dynamically adjust the software and hardware resources, including: performing division processing on the software and hardware resources to obtain first and second types of software and hardware resources, wherein the first type of software and hardware resources is quantifiable resources, and the second type of software and hardware resources is non-quantifiable resources; determine the current allocation of the first type of software and hardware resources, and based on the weight values of the modality data collected by the sensor modules in the next time period, re-allocate the first type of software and hardware resources based on the current allocation of the first type of software and hardware resources; determine a plurality of preset configurations of the second type of software and hardware resources, and a weight value range corresponding to each of the preset configurations, and based on the weight values of the modality data collected by the sensor modules in the next time period, reconfigure the second type of software and hardware resources.
6. The multi-modal perception and communication system software and hardware co-optimization method of claim 5, wherein determining the current allocation of the first type of software and hardware resources, and based on the weight values of the modality data collected by the sensor modules in the next time period, re-allocate the first type of software and hardware resources based on the current allocation of the first type of software and hardware resources, including: determining the current allocation of each software and hardware resource in the first type of software and hardware resources; based on the weight values of the modality data collected by the sensor modules in the next time period, re-allocate the current allocation of each software and hardware resource in the first type of software and hardware resources; and / or determining a plurality of preset configurations of the second type of software and hardware resources, and a weight value range corresponding to each of the preset configurations, and based on the weight values of the modality data collected by the sensor modules in the next time period, reconfigure the second type of software and hardware resources, including: The second type of software and hardware resources is split into first type of software and hardware sub-resources and second type of software and hardware sub-resources, wherein the first type of software and hardware sub-resources are used to match different preset configurations for different sensor modules, and the second type of software and hardware sub-resources are used to match the same preset configuration for different sensor modules; A plurality of preset configurations of each software and hardware resource in the first type of software and hardware sub-resources are determined, and for each software and hardware resource in the first type of software and hardware sub-resources, a preset configuration is matched for different sensor modules based on the weight values of the modal data collected by the respective sensor modules in the next time period; A plurality of preset configurations of each software and hardware resource in the first type of software and hardware sub-resources are determined, and the weight values of the modal data collected by the respective sensor modules in the next time period are processed by weighted summation, and based on the result of the weighted summation, a preset configuration is determined for each software and hardware resource in the second type of software and hardware sub-resources.
7. A device for soft hardware co-optimization of a multi-modal sensing and communication system, characterized in that, Comprise: The acquisition unit acquires data through a plurality of sensor modules to obtain a plurality of modal data, wherein each sensor module corresponds to a different perception mode, and the perception mode at least includes any of the following: vision, hearing and touch; The integration unit is configured to integrate the plurality of modal data into the same representation by using a preset inter-modal data fusion algorithm to obtain modal set data; The prediction unit is configured to divide the modal set data into different modal categories to obtain data subsets of different modal categories; perform time series analysis on the data subsets of different modal categories, and perform data prediction processing on different sensor modules based on the result of the time series analysis; determine the weight values of the modal data collected by the different sensor modules in the next time period based on the data prediction results of the different sensor modules; and determine the adjustment strategy of the different sensor modules based on the weight values of the modal data collected by the different sensor modules in the next time period; The adjustment unit is configured to dynamically adjust the software and hardware resources according to the resource adjustment strategies of the various modal data, and process the data collected by the plurality of sensor modules based on the dynamically adjusted software and hardware resources.
8. A computer device comprising a memory and a processor, the memory having stored therein a computer program, characterized in that, The processor executes the computer program to implement the steps of the method of any one of claims 1 to 6.
9. A computer readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 6.
Citation Information
Patent Citations
Data fusion method and system based on multi-modal sensor
CN118097352A
Edge adaptive control system based on multiple modes
CN118778454A
Metacosmic police service processing system based on multiple modes
CN119272943A