Optimization method and system for embedded image recognition algorithm

By obtaining multimodal data and generating dynamic risk levels, dynamically configuring the parameters of embedded image recognition algorithms, the problem of the contradiction between computing power and real-time in fire warning by embedded image recognition algorithms is solved, and efficient fire recognition and real-time response are achieved.

CN120339847AActive Publication Date: 2025-07-18HANGZHOU ZIPENG TECH CO LTD

Patent Information

Application Number
CN202510803522.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-07-18
Estimated Expiration
2045-06-17

AI Technical Summary

Technical Problem

Embedded image recognition algorithms face the contradiction between computing power and real-time in fire warning, and it is difficult to take into account the conflict between high-resolution image processing requirements and limited computing resources, as well as the mismatch between the computing delay of complex models and the real-time requirements of actual scenarios.

Method used

By obtaining multimodal monitoring data, including scene feature data, multimodal environment data and hardware status data, deep mining of risk assessment feature data, generating dynamic risk levels, and dynamically configure initial image processing parameters in combination with computational performance data and operational performance data, adjusting parameters in real time to meet real-time requirements.

Benefits of technology

While meeting the real-time requirements, ensuring fire recognition accuracy, alleviating the contradiction between high resolution and limited computing power, and ensuring the reliable operation of embedded image recognition algorithms in fire early warning scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339847A_ABST
    Figure CN120339847A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to an optimization method and system for an embedded image recognition algorithm. According to the method, the multi-modal environment data is deeply mined to obtain the risk assessment feature data, the scene feature data is combined to generate the dynamic risk level, and then the hardware state data is combined to obtain the initial image processing parameter, so that the initial parameter is dynamically configured according to the hardware actual capability and the risk demand; according to the method, fire identification precision can be guaranteed while real-time requirements are met, the contradiction between high resolution and finite computing power is relieved, and then a real-time matching value is obtained according to computing performance data, risk assessment feature data and initial image processing parameters, so that whether hardware resources under current parameter configuration can meet algorithm processing requirements or not is assessed; when the calculation delay does not meet the real-time requirement, the initial image processing parameters are dynamically adjusted to ensure the reliable operation of the embedded image recognition algorithm in the fire early warning scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly relates to an optimization method and system for an embedded image recognition algorithm. Background Art

[0002] An embedded image recognition algorithm refers to an image recognition technology deployed on embedded devices (such as microcontrollers, ARM processors, FPGA / ASIC chips and other terminal devices with limited computing resources). It directly processes, analyzes and extracts feature data from the collected image data on the device locally to achieve tasks such as classification, detection, and segmentation of image content. Currently, embedded image recognition algorithms have been widely applied in the field of fire warning. Relying mainly on the edge computing ability of the embedded system, through image processing algorithms deployed in front-end devices (such as cameras, sensor nodes), real-time analysis of image data in the monitored scene is carried out to achieve early detection and warning of fires.

[0003] In the scenario where an embedded image recognition algorithm is applied to fire warning, the contradiction between computing power and real-time performance is a key bottleneck problem. This contradiction is mainly reflected in the conflict between the requirement for high-resolution image processing and the limited computing resources of embedded devices, and the mismatch between the computing delay of complex models and the real-time requirements of the actual scenario. These two aspects of conflicts make it difficult for the embedded image recognition algorithm to balance accuracy and timeliness in fire warning. Summary of the Invention

[0004] The main object of the present invention is to provide an optimization method and system for an embedded image recognition algorithm, aiming to solve the technical problems in the prior art.

[0005] The present invention proposes an optimization method for an embedded image recognition algorithm, including: Obtaining multi-modal monitoring data monitored by sensors, where the multi-modal monitoring data includes scene feature data, multi-modal environmental data, and hardware status data; Obtaining risk assessment feature data according to the multi-modal environmental data; Obtaining a dynamic risk level according to the risk assessment feature data and the scene feature data; Obtaining computing performance data and running performance data according to the hardware status data, and obtaining initial image processing parameters according to the computing performance data, running performance data, and dynamic risk level; Obtaining a real-time matching value according to the computing performance data, risk assessment feature data, and initial image processing parameters; Judging whether the real-time matching value is greater than a preset threshold; If the real-time matching value is greater than the preset threshold, it is determined that the computing delay meets the real-time requirement; If the real-time matching value is not greater than the preset threshold, it is determined that the calculation delay does not meet the real-time requirement, and the initial image processing parameters are adjusted according to the real-time matching value and the dynamic risk level until the real-time matching value is greater than the preset threshold.

[0006] Preferably, the step of obtaining risk assessment feature data according to the multi-modal environment data includes: Obtain spectral image data, thermal radiation image data and gaseous molecular data according to the multi-modal environment data, and obtain a spectral alignment image and a thermal radiation alignment image according to the spectral image data and the thermal radiation image data; Obtain the global chromaticity variance according to the spectral alignment image, and obtain a spectral feature vector according to the global chromaticity variance by principal component analysis; Obtain timestamp feature points according to the spectral alignment image and the thermal radiation alignment image, and obtain time-aligned gas data according to the timestamp feature points and the gaseous molecular data; Obtain a CO concentration sequence and a concentration gradient according to the time-aligned gas data, and obtain a color feature vector according to the CO concentration sequence, the concentration gradient and the spectral feature vector; Obtain high-temperature centroid coordinates and a temperature gradient matrix according to the thermal radiation alignment image, and obtain a spatio-temporal adjacency matrix according to the high-temperature centroid coordinates; Obtain a thermodynamic gradient tensor according to the spatio-temporal adjacency matrix and the temperature gradient matrix, and obtain a dynamic texture entropy value matrix according to the spatio-temporal adjacency matrix.

[0007] Preferably, the step of obtaining a dynamic risk level according to the risk assessment feature data and the scene feature data includes: Obtain spatial positioning data, optical feature parameters and structural topology parameters according to the scene feature data, and obtain a regional function type according to the spatial positioning data and the structural topology parameters; Obtain a material risk level according to the regional function type and the optical feature parameters, and obtain a spatial attribute category according to the regional function type, the material risk level and the structural topology parameters; Perform a correlation analysis on the color feature vector, the dynamic texture entropy value matrix and the thermodynamic gradient tensor with the spatial attribute category respectively to obtain corresponding modal correlation coefficients, and obtain corresponding associated weight indices according to the modal correlation coefficients; Obtain a normalized feature data set according to the color feature vector, the dynamic texture entropy value matrix and the thermodynamic gradient tensor, and obtain a dynamic risk level according to the normalized feature data set and the associated weight index.

[0008] Preferably, the step of obtaining initial image processing parameters according to the computing performance data, operating performance data, and dynamic risk level includes: Obtain the core frequency, reference frequency, architecture coefficient, utilization margin, remaining memory capacity, and memory fragmentation rate according to the computing performance data, and obtain a comprehensive resource index according to the core frequency, reference frequency, architecture coefficient, and utilization margin; Obtain the memory health according to the remaining memory capacity and the memory fragmentation rate; Obtain the real-time operating temperature, safety threshold temperature, and maximum allowable temperature according to the operating performance data, and obtain the computing health according to the real-time operating temperature, safety threshold temperature, and maximum allowable temperature; Obtain the resolution parameter according to the comprehensive resource index and the computing health; Obtain a risk-resource mapping table, and obtain a computing requirement benchmark and a memory requirement benchmark according to the risk-resource mapping table and the dynamic risk level; Obtain the model architecture parameter according to the memory requirement benchmark and the memory health, and obtain the computing precision parameter according to the comprehensive resource index and the computing requirement benchmark.

[0009] Preferably, the step of obtaining a real-time matching value according to the computing performance data, risk assessment feature data, and initial image processing parameters includes: Obtain the effective value of hardware floating-point computing power, chip junction temperature, cycle operation count, memory bandwidth, and processing cycle according to the computing performance data, and obtain the real-time CPU computing power according to the effective value of hardware floating-point computing power, core frequency, cycle operation count, chip junction temperature, and safety threshold temperature; Obtain the processing requirement data according to the risk assessment feature data, and obtain the processing time threshold and algorithm computing amount according to the processing requirement data; Obtain the theoretical processing time according to the algorithm computing amount and the real-time CPU computing power, and obtain the algorithm matching degree according to the theoretical processing time and the processing time threshold; Obtain the available memory space according to the remaining memory capacity, memory fragmentation rate, memory bandwidth, and processing cycle; Obtain the memory requirement space and required bandwidth according to the resolution parameter, model architecture parameter, and computing precision parameter, and obtain the memory matching degree according to the memory requirement space, available memory space, memory bandwidth, and required bandwidth; Obtain the real-time matching value according to the algorithm matching degree and the memory matching degree.

[0010] Preferably, the step of adjusting the initial image processing parameters according to the real-time matching value and the dynamic risk level includes: Obtain an adjustment intensity level based on the real-time matching value and the dynamic risk level, and obtain an adjustment constraint condition according to the dynamic risk level; Obtain a parameter adjustment data set according to the adjustment intensity level and the adjustment constraint condition, where the parameter adjustment data set includes a plurality of resolution adjustment data, model architecture adjustment data, and computing precision adjustment data; Arrange and combine the plurality of resolution adjustment data, model architecture adjustment data, and computing precision adjustment data to obtain a plurality of parameter adjustment combination sets; Screen the plurality of parameter adjustment combination sets according to the real-time matching value and the adjustment constraint condition to obtain an optimal parameter adjustment combination, and return the optimal parameter adjustment combination as the initial image processing parameter to the step of obtaining the real-time matching value according to the computing performance data, risk assessment feature data, and initial image processing parameter.

[0011] This application also provides an optimization system for an embedded image recognition algorithm, including: A first acquisition module for acquiring multi-modal monitoring data monitored by a sensor, where the multi-modal monitoring data includes scene feature data, multi-modal environment data, and hardware status data; A second acquisition module for obtaining risk assessment feature data according to the multi-modal environment data; A third acquisition module for obtaining a dynamic risk level according to the risk assessment feature data and the scene feature data; A fourth acquisition module for obtaining computing performance data and operating performance data according to the hardware status data, and obtaining an initial image processing parameter according to the computing performance data, operating performance data, and dynamic risk level; A fifth acquisition module for obtaining a real-time matching value according to the computing performance data, risk assessment feature data, and initial image processing parameter; A judgment module for judging whether the real-time matching value is greater than a preset threshold; If the real-time matching value is greater than the preset threshold, it is determined that the calculation delay meets the real-time requirement; If the real-time matching value is not greater than the preset threshold, it is determined that the calculation delay does not meet the real-time requirement, and the initial image processing parameter is adjusted according to the real-time matching value and the dynamic risk level until the real-time matching value is greater than the preset threshold.

[0012] Preferably, the third acquisition module includes: A first acquisition unit for obtaining spatial positioning data, optical feature parameters, and structural topology parameters according to the scene feature data, and obtaining a regional function type according to the spatial positioning data and the structural topology parameters; A second acquisition unit, configured to acquire a material risk level according to the area function type and the optical feature parameter, and acquire a spatial attribute category according to the area function type, the material risk level, and the structural topology parameter; A third acquisition unit, configured to perform a correlation analysis on the color feature vector, the dynamic texture entropy value matrix, and the thermodynamic gradient tensor with the spatial attribute category respectively, obtain corresponding modal correlation coefficients, and obtain corresponding associated weight indexes according to the modal correlation coefficients; A fourth acquisition unit, configured to acquire a normalized feature data set according to the color feature vector, the dynamic texture entropy value matrix, and the thermodynamic gradient tensor, and acquire a dynamic risk level according to the normalized feature data set and the associated weight index.

[0013] The present invention further provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the above optimization method for an embedded image recognition algorithm are implemented.

[0014] The present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above optimization method for an embedded image recognition algorithm are implemented.

[0015] The beneficial effects of the present invention are as follows: the present invention collects multimodal data from different types of sensors, including scene feature data, multimodal environmental data and hardware status data, to provide a comprehensive data basis for subsequent analysis, avoiding the limitations of a single data source, and then obtains risk assessment feature data by deeply mining the multimodal environmental data, which can more accurately identify signs of fire, detect weak flames and temperature changes in the early stages of a fire, buy time for subsequent processing, and reduce real-time risks caused by delays in complex model calculations. The scene feature data is then further combined to generate a dynamic risk level, and then the computing performance data, operating performance data and dynamic risk level are combined to obtain initial image processing parameters. This method of dynamically configuring initial parameters based on actual hardware capabilities and risk requirements , which can ensure the accuracy of fire identification while meeting the real-time requirements, alleviate the contradiction between high resolution and limited computing power, and then obtain the real-time matching value based on the computing performance data, risk assessment feature data and initial image processing parameters, so as to evaluate whether the hardware resources can meet the algorithm processing requirements under the current parameter configuration. This quantitative evaluation method enables the system to dynamically adjust according to the actual situation, avoiding the waste of computing power or insufficient real-time performance caused by blindly configuring parameters. When the calculation delay does not meet the real-time requirements, the initial image processing parameters are dynamically adjusted according to the real-time matching value and dynamic risk level until the real-time requirements are met. This closed-loop adjustment mechanism continuously optimizes the parameter configuration, which can effectively solve the contradiction between computing power and real-time performance, and ensure the reliable operation of the embedded image recognition algorithm in the fire warning scenario. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 The figure is a schematic diagram of a method flow according to an embodiment of the present invention.

[0017] Figure 2 FIG. 1 is a schematic diagram of a system structure according to an embodiment of the present invention.

[0018] Figure 3 A schematic diagram of the internal structure of a computer device according to an embodiment of the present application.

[0019] The realization of the purpose, functional features and advantages of the present invention will be further explained in conjunction with embodiments and with reference to the accompanying drawings. DETAILED DESCRIPTION

[0020] It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention.

[0021] like Figure 1 As shown, the present application provides an optimization method for an embedded image recognition algorithm, comprising: S1. Acquire multimodal monitoring data monitored by a sensor, wherein the multimodal monitoring data includes scene feature data, multimodal environment data, and hardware status data; S2. Obtain risk assessment feature data according to the multi-modal environmental data; S3. Obtain a dynamic risk level according to the risk assessment feature data and the scenario feature data; S4. Obtain computing performance data and running performance data according to the hardware status data, and obtain initial image processing parameters according to the computing performance data, the running performance data, and the dynamic risk level; S5. Obtain a real-time matching value according to the computing performance data, the risk assessment feature data, and the initial image processing parameters; S6. Determine whether the real-time matching value is greater than a preset threshold; If the real-time matching value is greater than the preset threshold, it is determined that the calculation delay meets the real-time requirement; If the real-time matching value is not greater than the preset threshold, it is determined that the calculation delay does not meet the real-time requirement, and the initial image processing parameters are adjusted according to the real-time matching value and the dynamic risk level until the real-time matching value is greater than the preset threshold.

[0022] As described in the above steps S1-S6, the present invention collects multimodal data from different types of sensors, wherein multimodal data refers to a data set reflecting information of different dimensions of the monitoring target, including scene feature data, multimodal environmental data and hardware status data, wherein scene feature data refers to image or video data that directly characterizes the visual characteristics of the monitoring scene, multimodal environmental data refers to environmental physicochemical feature data collected by non-visual sensors, and hardware status data refers to state parameters of the embedded device hardware during operation. These data provide a comprehensive data basis for subsequent analysis and avoid the limitations of a single data source. Then, feature extraction and analysis are performed on the multimodal environmental data to obtain risk assessment feature data, wherein the risk assessment feature data refers to a set of key features extracted from the multimodal environmental data for assessing fire risks, including spectral feature vectors, color feature vectors and thermodynamic gradient tensors. Through deep mining Data features can more accurately identify signs of fire. For example, by using the color features of spectral image data (global chromaticity variance, spectral feature vector) and the temperature distribution of thermal radiation image data (temperature gradient matrix, thermodynamic gradient tensor), weak flames and temperature changes can be detected in the early stages of a fire. Compared with the traditional method that relies only on a single image feature, it can detect fires earlier, buy time for subsequent processing, and reduce the real-time risk caused by delays in complex model calculations. Then, it is further combined with scene feature data to generate a dynamic risk level. The dynamic risk level refers to the fire risk level (such as low, medium, and high risk) calculated in real time based on risk assessment feature data and scene feature data. The dynamic risk level provides priority guidance for subsequent parameter adjustments. For example, in high-risk scenarios, the real-time requirements are higher, and the system can prioritize the allocation of computing resources and use higher resolution and more complex models to ensure the accuracy of fire identification.In a low-risk scenario, resource consumption can be appropriately reduced. A lightweight model and a lower resolution are adopted to balance computational latency and resource utilization, avoiding excessive occupation of limited computing power by high-resolution image processing. Then, computational performance data and operating performance data are extracted from the hardware status data. Among them, the computational performance data refers to static parameters characterizing the hardware computing power of the embedded device, such as the number of cores, main frequency, peak computing power, etc. The operating performance data refers to the real-time status parameters during the operation of the embedded device. Initial image processing parameters are obtained in combination with the dynamic risk level. Among them, the initial image processing parameters refer to a set of image processing algorithm parameters pre-configured based on hardware performance and risk level, including resolution parameters, model architecture parameters, and computational precision parameters. This method of dynamically configuring initial parameters according to the actual hardware capabilities and risk requirements can ensure the fire recognition accuracy while meeting the real-time requirements, alleviating the contradiction between high resolution and limited computing power. Then, a real-time matching value is obtained based on the computational performance data, risk assessment feature data, and initial image processing parameters. Among them, the real-time matching value is a comprehensive index measuring the matching degree between the current image processing parameters and the hardware performance. The real-time matching value is compared with a preset threshold to determine whether the current computational latency meets the real-time requirements. If the real-time matching value is greater than the threshold, it indicates that the computational latency is within an acceptable range, and the system continues to run with the current configuration. If the real-time matching value is not greater than the threshold, it indicates that the hardware resources cannot meet the algorithm processing requirements under the current parameter configuration, and there may be a problem of excessive computational latency, and parameter adjustment is required. This quantitative evaluation method enables the system to dynamically adjust according to the actual situation, avoiding waste of computing power or insufficient real-time performance caused by blind parameter configuration. When the computational latency does not meet the real-time requirements, the initial image processing parameters are dynamically adjusted according to the real-time matching value and the dynamic risk level (such as reducing the resolution, switching to a lightweight model, adjusting the computational precision), and the real-time matching value is recalculated until the real-time requirements are met. This closed-loop adjustment mechanism continuously optimizes the parameter configuration, effectively solving the contradiction between computing power and real-time performance, and ensuring the reliable operation of the embedded image recognition algorithm in the fire warning scenario.

[0023] In one embodiment, step S2 of obtaining risk assessment feature data according to the multi-modal environment data includes: S21. Obtain spectral image data, thermal radiation image data, and gaseous molecular data according to the multi-modal environment data, and obtain a spectral alignment image and a thermal radiation alignment image according to the spectral image data and the thermal radiation image data; S22. Obtain the global chromaticity variance according to the spectral alignment image, and obtain a spectral feature vector according to the global chromaticity variance through principal component analysis; S23. Obtain timestamp feature points according to the spectral alignment image and the thermal radiation alignment image, and obtain time-aligned gas data according to the timestamp feature points and the gaseous molecular data; S24. Obtain the CO concentration sequence and concentration gradient according to the time-aligned gas data, and obtain the color feature vector according to the CO concentration sequence, concentration gradient, and spectral feature vector; S25. Obtain the high-temperature centroid coordinates and temperature gradient matrix according to the heat radiation-aligned image, and obtain the spatio-temporal adjacency matrix according to the high-temperature centroid coordinates; S26. Obtain the thermodynamic gradient tensor according to the spatio-temporal adjacency matrix and the temperature gradient matrix, and obtain the dynamic texture entropy value matrix according to the spatio-temporal adjacency matrix.

[0024] As described in the above steps S21 - S26, in the present invention, fire risk characteristics in the environment are extracted from multiple dimensions by obtaining spectral image data, thermal radiation image data, and gaseous molecule data. Among them, spectral image data refers to the scene image data collected by a spectral sensor (such as a multispectral / hyperspectral camera), which records the radiation intensity distribution at different wavelengths (such as visible light, near-infrared). Thermal radiation image data refers to the scene temperature distribution image collected by a thermal imager, which reflects the thermal radiation intensity of the object surface (usually the temperature level is displayed in grayscale or pseudo-color). Gaseous molecule data refers to the concentration data of gaseous molecules in the environment collected by a gas sensor (such as an infrared gas detector, an electrochemical sensor), such as the real-time concentration values of fire-related gases such as CO, CO2, and smoke particles. Then, feature points in the spectral image data and the thermal radiation image data are extracted, and mismatched points are removed by RANSAC, and the visible light and infrared images are aligned to obtain a spectrally aligned image and a thermally radiated aligned image. Then, the spectrally aligned image is converted from RGB to the HSV color space, the chromaticity (H) and saturation (S) channels are separated, and the global variances of the H channel and the S channel are calculated respectively to obtain the global chromaticity variance. Among them, the global chromaticity variance refers to calculating the variance of the chromaticity values (such as saturation and hue in the HSV space) of all pixels in the spectral image, which is used to measure the dispersion degree of the overall color distribution of the image. Then, principal component analysis is performed on the global chromaticity variance to further obtain a spectral feature vector. Among them, the spectral feature vector refers to the feature vector obtained after dimensionality reduction of spectral features such as the global chromaticity variance. This method of deeply mining the color features of spectral images in a fire scene can capture these features more accurately, thereby improving the accuracy of using spectral image information in fire risk assessment and solving the deficiencies in the extraction of spectral color features in the prior art. Immediately afterwards, the pyramid LK optical flow method is used to track FAST corner points in the spectrally aligned image and the thermally radiated aligned image to generate a trajectory with a timestamp, and points with a movement speed > 5 pixels / frame and a trajectory length > 5 frames are retained as timestamp feature points. Among them, the timestamp feature point refers to the spatial feature point with a timestamp extracted by a feature point detection algorithm (such as SIFT, ORB) in multiple frames of spectral images and thermally radiated images. Then, the linear interpolation method is used to synchronize the gaseous molecule data to the timestamp of each image frame to generate time-aligned gas data. In a fire scene, the generation and change of gases are temporally related to the spectral and thermal radiation changes of the flame. The prior art may not strictly align the timestamps of multimodal data, resulting in a timing deviation during feature fusion. This solution deeply correlates different modal data in the time and space dimensions, taking into account the consistency and relevance of different physical phenomena (optical, thermal, gas changes) during a fire, and can more comprehensively understand the physical and chemical changes during the fire development process, improving the accuracy and comprehensiveness of risk assessment. Then, the CO concentration sequence and concentration gradient are obtained according to the time-aligned gas data, whereThe CO concentration sequence refers to the sequence of CO concentration values arranged in chronological order. The concentration gradient refers to the rate of change of the CO concentration in space. Since the CO concentration gradient can reflect the speed and trend of fire development, it is combined with the spectral feature vector to obtain the color feature vector. Among them, the color feature vector is a multi-dimensional feature vector formed by fusing the CO concentration sequence, the concentration gradient, and the spectral feature vector, which can more comprehensively describe the characteristics of the fire scene. Then, threshold segmentation is performed on the thermal radiation aligned image, and the area where the temperature > the ambient temperature + 30°C is extracted as the high-temperature area. The connected component analysis is used to label each area to obtain the high-temperature centroid coordinates. Among them, the high-temperature centroid coordinates refer to the centroid (geometric center) coordinates of the high-temperature area in the thermal radiation image, which are used to locate the spatial position of the heat source. Then, a 3×3 Gaussian derivative kernel is used to calculate the gradient matrix of the high-temperature area to obtain the temperature gradient matrix. Among them, the temperature gradient matrix is a matrix composed of the temperature gradient values of each pixel point in the thermal radiation image, and each element represents the rate of change of the temperature of that point in the x and y directions (such as the gradient amplitude and direction). Then, each high-temperature centroid coordinate is regarded as a node, and the Euclidean distance between the nodes is calculated. If the distance < 50 pixels (empirical threshold), an undirected edge is established, and the weight is the absolute value of the temperature difference. The Hungarian algorithm is used to match the high-temperature centroids of adjacent frames. If the match is successful, a time edge is established, and the weight is the centroid displacement speed, so as to obtain the spatio-temporal adjacency matrix. Among them, the spatio-temporal adjacency matrix is a matrix constructed based on the spatio-temporal correlation of the high-temperature centroid coordinates, which represents the spatial adjacency relationship of the high-temperature areas in adjacent time frames. Through this method, the propagation, diffusion, and dynamic changes of heat distribution in the fire can be understood more accurately. Then, the temperature gradient matrix and the spatio-temporal adjacency matrix are subjected to a tensor product operation to generate a thermodynamic gradient tensor. Among them, the thermodynamic gradient tensor is a multi-dimensional physical quantity that describes the spatial temperature gradient and the time adjacency relationship, so as to more accurately characterize the thermodynamic behavior of the fire heat source. Finally, for each high-temperature area, the local binary pattern texture feature is calculated to generate a dynamic texture entropy value matrix. Among them, the dynamic texture entropy value matrix is a matrix formed by calculating the entropy value of the texture features (such as the flicker of the flame and the turbulence of the smoke) of the thermal radiation image sequence, which can be used as a key feature to distinguish "dynamic flame" from "static high temperature".

[0025] In one embodiment, step S3 of obtaining the dynamic risk level according to the risk assessment feature data and the scene feature data includes: S31. Obtain spatial positioning data, optical feature parameters, and structural topology parameters according to the scene feature data, and obtain the regional function type according to the spatial positioning data and the structural topology parameters; S32. Obtain the material risk level according to the regional function type and the optical feature parameters, and obtain the spatial attribute category according to the regional function type, the material risk level, and the structural topology parameters; S33. Perform correlation analysis on the color feature vector, the dynamic texture entropy value matrix, and the thermodynamic gradient tensor with the spatial attribute category respectively to obtain the corresponding modal correlation coefficients, and obtain the corresponding associated weight indexes according to the modal correlation coefficients; S34. Obtain a normalized feature data set according to the color feature vector, the dynamic texture entropy value matrix, and the thermodynamic gradient tensor, and obtain a dynamic risk level according to the normalized feature data set and the associated weight index.

[0026] As described in the above steps S31 - S34, the present invention provides a three - dimensional environmental awareness ability for the fire monitoring system of scene perception by obtaining spatial positioning data, optical feature parameters, and structural topology parameters. Among them, the spatial positioning data refers to the data describing the position, size, and geometric relationship of objects or regions in a three - dimensional space. The optical feature parameters refer to the visual features extracted from images or spectra, reflecting the physical properties of the object surface. The structural topology parameters refer to the data describing the spatial connection relationship and network structure between objects in the scene. The traditional fire monitoring system uses a "one - size - fits - all" algorithm, using a fixed feature set for all scenes (such as a unified monitoring temperature threshold, smoke concentration), ignoring the differences in fire characteristics of different scenes. For example, the key features for forest fire monitoring are flame spectrum, heat radiation diffusion speed, and wind direction. The key features for mall fire monitoring are smoke particle concentration, change in personnel density, and electrical equipment abnormalities. Therefore, the existing system cannot automatically switch the monitoring strategy, resulting in a high false negative rate for forest fires (relying on smoke recognition) and a high false positive rate for malls (air - conditioner hot air triggering the temperature threshold). Therefore, according to the spatial positioning data and structural topology parameters, this solution uses a support vector machine to obtain the regional function type. Among them, the regional function type refers to the scene function category divided according to spatial positioning and topological structure. Then, according to the regional function type and optical feature parameters, the material risk level is obtained. Among them, the material risk level is a quantitative level for evaluating the flammability and post - combustion harmfulness of object materials. Finally, considering the regional function type, material risk level, and structural topology parameters comprehensively, the spatial attribute category is obtained. Among them, the spatial attribute category is a composite semantic label integrating regional function type, material risk level, and topological relationship. For example: First, through sensor collection and image analysis, the spatial positioning data (such as three - dimensional coordinates, regional area, contour range) and structural topology parameters (such as adjacency relationship between objects, connectivity index) of the target area in the scene are obtained. Based on the above data, a support vector machine (SVM) classifier is used to determine the regional function type, and it is determined to belong to categories such as kitchen, electrical equipment room, storage area, or ordinary office area. On the basis of clarifying the regional function type, combined with optical feature parameters such as color mean, texture roughness, and edge density extracted from the spectral image, the fire risk level of the object materials in the region is classified through a decision tree algorithm, identifying flammable material areas, medium - risk material areas, or low - risk material areas. Finally, considering the spatial adjacency relationship (such as the distance from a high - risk heat source) in the regional function type, material risk level, and structural topology parameters, the spatial risk is corrected, thereby determining the final spatial attribute category of the region. Then, the color feature vector, dynamic texture entropy matrix, and thermodynamic gradient tensor are respectively analyzed for correlation with the spatial attribute category to obtain the corresponding modal correlation coefficients. Among them, the modal correlation coefficient is a statistical index measuring the degree of association between features and the spatial attribute category. The higher the coefficient, the more important the feature is for the risk assessment of the current scene.And obtain the corresponding associated weight index according to the modal correlation coefficient. Here, the associated weight index refers to the weight factor dynamically assigned to each feature according to the modal correlation coefficient. This step, through the dynamic weight allocation mechanism of scene perception, solves the fundamental defect of "fixed weights being unable to adapt to scene changes" in traditional multimodal fusion algorithms, thereby significantly improving the accuracy and adaptability of fire warning. For example, if the correlation coefficient between the color feature vector and the spatial attribute category is , the correlation coefficient between the dynamic texture entropy value matrix and the spatial attribute category is , and the correlation coefficient between the thermodynamic gradient tensor and the spatial attribute category is , then the modal correlation coefficients corresponding to the color feature vector, the dynamic texture entropy value matrix, and the thermodynamic gradient tensor are , , . The color feature vector, the dynamic texture entropy value matrix, and the thermodynamic gradient tensor have different data structures and dimensions. Direct splicing by traditional methods will lead to chaotic feature expressions. Therefore, this solution normalizes the multimodal data to obtain a normalized feature data set. For example: perform Min-Max normalization on the color feature vector to retain the relative relationship of spectral features, use Z-score standardization for the texture matrix to highlight abnormal texture changes, perform norm normalization on the thermodynamic gradient tensor to maintain the physical meaning of spatial heat distribution, and finally obtain the dynamic risk level through weighted summation combined with the weight index.

[0027] In one embodiment, the step S4 of obtaining the initial image processing parameters according to the calculated performance data, the running performance data, and the dynamic risk level includes: S41. Obtain the core frequency, the reference frequency, the architecture coefficient, the utilization margin, the remaining memory capacity, and the memory fragmentation rate according to the calculated performance data, and obtain the resource comprehensive index according to the core frequency, the reference frequency, the architecture coefficient, and the utilization margin; S42. Obtain the memory health according to the remaining memory capacity and the memory fragmentation rate; S43. Obtain the real-time running temperature, the safety threshold temperature, and the maximum allowable temperature according to the running performance data, and obtain the computing health according to the real-time running temperature, the safety threshold temperature, and the maximum allowable temperature; S44. Obtain the resolution parameter according to the resource comprehensive index and the computing health; S45. Obtain the risk-resource mapping table, and obtain the computing requirement benchmark and the memory requirement benchmark according to the risk-resource mapping table and the dynamic risk level; S46. Obtain model architecture parameters according to the memory requirement benchmark and the memory health, and obtain computing precision parameters according to the resource comprehensive index and the computing requirement benchmark.

[0028] As described in the above steps S41 - S46, in the present invention, by obtaining the core frequency, architecture coefficient, and utilization margin, the parameters of three independent dimensions are fused into a unified index, breaking through the traditional single evaluation method based only on frequency or utilization. Among them, the core frequency refers to the clock frequency of the processor core (such as CPU, GPU, or NPU) in the embedded device, the architecture coefficient refers to the quantitative index representing the advancement of the processor architecture, and the utilization margin refers to the difference between the current utilization of the processor core and the maximum utilization (usually 100%). Then, the resource comprehensive index is calculated through the formula "resource comprehensive index = 0.4 * (core frequency / reference frequency) + 0.3 * architecture coefficient + 0.7 * utilization margin". Among them, the resource comprehensive index is a quantitative value used to evaluate the comprehensive computing power. Since the traditional method only focuses on the remaining memory amount, this solution also considers the memory fragmentation rate, avoiding the problem of "having memory but unable to allocate" caused by memory fragmentation, and is particularly suitable for the scenario where the embedded system memory is tight. And the memory health is calculated through the formula "memory health = 0.6 * remaining memory capacity + 0.4 * (1 - memory fragmentation rate)". Among them, the remaining memory capacity refers to the size of the physical memory space not occupied in the embedded device, the memory fragmentation rate refers to the proportion of discontinuous free blocks in the total free memory in the memory, and the memory health is a comprehensive index to measure the availability of memory resources. Then, the temperature is incorporated into the resource evaluation system, and the computing load is actively reduced when approaching the safety threshold, avoiding the traditional passive mode of "overheating frequency reduction → sudden performance drop". First, the real-time operating temperature is mapped to the [0, 1] interval through the formula "temperature stress value = (real-time operating temperature - safety threshold temperature) / (maximum allowable temperature - safety threshold temperature)" to obtain the temperature stress value, and the degree to which the current temperature approaches the dangerous area is reflected by the obtained temperature stress value. Then, the exponential decay model is adopted: computing health = is calculated to obtain the computing health. Among them, the computing health is an evaluation index of the hardware operating state based on temperature. This method of converting the physical constraint of the hardware temperature into a computable resource index through the health function realizes the closed-loop optimization of "hardware state → algorithm parameters". Then, through the formula "resolution parameter = reference resolution "Calculate the resolution parameter, where the resolution parameter refers to the input image resolution processed by the image recognition algorithm. Immediately afterwards, according to the dynamic risk level and the risk-resource mapping table, look up the corresponding calculation requirement benchmark and memory requirement benchmark. The risk-resource mapping table refers to the mapping relationship table between predefined dynamic risk levels (such as low, medium, and high risks) and hardware resource requirements. The calculation requirement benchmark refers to the minimum computing power requirement determined according to the risk level, and the memory requirement benchmark refers to the minimum memory requirement determined according to the risk level, so that the algorithm configuration is more in line with the actual risk scenario. Then, obtain the model architecture parameter according to the memory requirement benchmark and the memory health. The model architecture parameter refers to the parameter characterizing the structural complexity of the image recognition model, such as the number of convolutional layers, the number of channels, the attention module configuration, etc. For example: if the memory health ≥ 0.5: when the memory requirement benchmark ≤ 512MB, select a lightweight or medium-weight model (such as EfficientNetB0) to balance accuracy and performance; when the memory requirement benchmark = 512MB - 1024MB, select a medium-weight model (such as ResNet50y) to ensure the detection accuracy of common risks; when the memory requirement benchmark ≥ 1024MB, select a complex model (such as YOLOv5s). Through the above method, the selection of the model architecture parameter no longer depends on a fixed configuration, but is dynamically adjusted according to the real-time memory state and risk requirements, significantly improving the environmental adaptability of the embedded image recognition system. Finally, according to the formula "Calculation accuracy parameter = "Calculate the calculation accuracy parameter, where the calculation accuracy parameter refers to the data accuracy used in the algorithm calculation process (such as FP32, FP16, INT8). Through this method, when resources are insufficient, the accuracy can be gradually reduced instead of directly switching to the low-precision mode, thereby reducing the detection performance loss caused by a sudden drop in accuracy.

[0029] In one embodiment, the step S5 of obtaining the real-time matching value according to the calculation performance data, the risk assessment feature data, and the initial image processing parameters includes: S51. Obtain the effective value of the hardware floating-point computing power, the chip junction temperature, the number of cycle operations, the memory bandwidth, and the processing cycle according to the calculation performance data, and obtain the CPU real-time computing power according to the effective value of the hardware floating-point computing power, the core frequency, the number of cycle operations, the chip junction temperature, and the safety threshold temperature; S52. Obtain the processing requirement data according to the risk assessment feature data, and obtain the processing time threshold and the algorithm calculation amount according to the processing requirement data; S53. Obtain the theoretical processing time according to the algorithm calculation amount and the CPU real-time computing power, and obtain the algorithm matching degree according to the theoretical processing time and the processing time threshold; S54. Obtain the available memory space according to the remaining memory capacity, the memory fragmentation rate, the memory bandwidth, and the processing cycle; S55. Obtain the memory requirement space and required bandwidth according to the resolution parameter, model architecture parameter, and computing precision parameter, and obtain the memory matching degree according to the memory requirement space, available memory space, memory bandwidth, and required bandwidth; S56. Obtain the real-time matching value according to the algorithm matching degree and the memory matching degree.

[0030] As described in the above steps S51 - S56, the theoretical computing power of the CPU of the present invention is obtained by multiplying the effective value of the hardware floating - point computing power, the core frequency, and the number of cycle operations. Among them, the effective value of the hardware floating - point computing power refers to the floating - point operation ability that the hardware (such as CPU / GPU) can stably output during actual operation, and the number of cycle operations refers to the number of floating - point operations that the CPU can execute within each clock cycle. And according to the difference between the chip junction temperature and the safety threshold temperature, the theoretical computing power of the CPU is dynamically adjusted (such as reducing the frequency proportionally when the junction temperature exceeds the safety threshold) to obtain the real - time computing power of the CPU. Among them, the chip junction temperature refers to the actual temperature of the semiconductor junction inside the chip, and the real - time computing power of the CPU refers to the actual available computing power of the CPU at the current moment. By obtaining the real - time computing power of the CPU in this way, it is possible to break through the traditional static evaluation mode that relies on the nominal computing power, and quantify the real hardware computing power in real time through the frequency utilization rate and the temperature - reduction frequency coefficient, solving the problem of computing power fluctuation caused by heat - dissipation limitations in embedded devices. Then, the algorithm complexity is deduced from the risk - assessment feature data (such as color feature vectors, thermodynamic gradient tensors, and dynamic texture entropy matrices), so as to obtain the processing - requirement data. Among them, the processing - requirement data refers to the requirements of the image - processing task deduced from the risk - assessment feature data, including indicators such as the amount of calculation, timeliness, and memory occupancy. Then, according to the processing - requirement data, the processing - time threshold and the algorithm calculation amount are obtained. Among them, the processing - time threshold refers to the maximum time allowed for the image - processing task by the system, and the algorithm calculation amount refers to the number of floating - point operations required to execute a specific image - processing algorithm. And according to the ratio of the algorithm calculation amount to the real - time computing power of the CPU, the theoretical processing time is obtained. Immediately afterwards, according to the formula "algorithm matching degree = 1 - max((theoretical processing time - processing - time threshold), 0) / processing - time threshold", the algorithm matching degree is calculated. Among them, the algorithm matching degree is a quantitative index that measures the matching degree between the algorithm calculation requirements and the real - time computing power of the CPU. The result range of the algorithm matching degree is between 0 and 1, and the higher the value, the more it meets the real - time requirements. By this processing method of associating multi - modal risk features with processing requirements (such as forcibly reducing the time threshold to 50 ms in high - risk scenarios), the computing - power allocation is more in line with the actual threat level, thus avoiding the problems of "wasting high computing power in low - risk scenarios and missing detections with low computing power in high - risk scenarios". Then, the available memory space is calculated according to the formula "available memory space = remaining memory capacity * (1 - memory fragmentation rate)+memory bandwidth * processing cycle". Among them, the available memory space refers to the actual available memory resources considering memory fragmentation and bandwidth utilization rate, the memory bandwidth refers to the data - transfer rate between the memory and the CPU, and the processing cycle refers to the time cycle required to complete a complete image - processing task. Then, further, the memory - requirement space is obtained according to the resolution parameter, the model - architecture parameter, and the computing - precision parameter. Among them, the memory - requirement space refers to the memory resources required to execute a specific image - processing algorithm. And according to the formula "memory matching degree=(available memory space / memory - requirement space)*min(1, memory bandwidth / required bandwidth)", the memory matching degree is calculated. Among them,The required bandwidth refers to the data transfer rate requirement between memory and CPU during the operation of the algorithm. The memory matching degree is a quantitative indicator for measuring the matching degree between the memory requirement of the algorithm and the hardware memory resources. By incorporating the fragmentation rate (spatial continuity) and bandwidth (transmission efficiency) into the evaluation system, the problem of "processing delay caused by sufficient capacity but insufficient bandwidth" in the embedded system can be solved, thus breaking through the traditional "capacity-first" memory management mode. Finally, the algorithm matching degree and the memory matching degree are fused according to the preset weights to obtain a real-time matching value. Through the real-time matching value, the adaptation degree between the hardware resources of the embedded image recognition system and the algorithm processing requirements can be dynamically evaluated, thereby providing a quantitative basis for system optimization, task scheduling, or parameter adjustment.

[0031] In one embodiment, step S6 of adjusting the initial image processing parameters according to the real-time matching value and the dynamic risk level includes: S61. Obtain an adjustment intensity level according to the real-time matching value and the dynamic risk level, and obtain an adjustment constraint condition according to the dynamic risk level; S62. Obtain a parameter adjustment data set according to the adjustment intensity level and the adjustment constraint condition, where the parameter adjustment data set includes a plurality of resolution adjustment data, model architecture adjustment data, and calculation accuracy adjustment data; S63. Arrange and combine the plurality of resolution adjustment data, model architecture adjustment data, and calculation accuracy adjustment data to obtain a plurality of parameter adjustment combination sets; S64. Screen the plurality of parameter adjustment combination sets according to the real-time matching value and the adjustment constraint condition to obtain an optimal parameter adjustment combination, and return the optimal parameter adjustment combination as the initial image processing parameter to the step of obtaining the real-time matching value according to the calculation performance data, risk assessment feature data, and initial image processing parameter.

[0032] As described in the above steps S61 - S64, the present invention breaks through the limitation of a single factor determining the adjustment strategy by constructing a two - dimensional mapping table of real - time matching values and dynamic risk levels, and realizes a more refined division of adjustment intensities. For example, the real - time matching values are divided into three intervals: low (0 - 0.3), medium (0.3 - 0.7), and high (0.7 - 1), and the dynamic risk levels are divided into levels 1 - 5. By cross - indexing the mapping table, the corresponding adjustment intensity levels are determined. Among them, the adjustment intensity level refers to the grading of the parameter adjustment strength comprehensively determined according to the real - time matching value and the dynamic risk level. For example, when the real - time matching value is 0.2 and the dynamic risk level is level 4, the highest adjustment intensity level is obtained by looking up the table. At the same time, the adjustment constraint conditions are obtained according to the dynamic risk level. Among them, the adjustment constraint conditions refer to the hard limit conditions determined by the dynamic risk level, which are used to ensure that the parameter adjustment is within a safe and feasible range. For example, in a low - risk scenario, a lower resolution (such as not less than 320×240), a lower model accuracy (such as INT8 quantization), and a longer delay (such as not exceeding 500 ms) are allowed; in a high - risk scenario, a higher resolution (such as not less than 1080P), a higher model accuracy (such as FP16), and a shorter delay (such as not exceeding 100 ms) are required. In this way, it is possible to ensure the necessary image - processing quality and system response speed in high - risk scenarios, and avoid problems such as missed detections caused by excessive adjustment. Then, according to the adjustment intensity level and the adjustment constraint conditions, a parameter adjustment data set is obtained. Among them, the parameter adjustment data set refers to the basic data set containing various adjustable parameters, including multiple resolution adjustment data, model architecture adjustment data, and calculation accuracy adjustment data. This acquisition method avoids blindly enumerating all possible parameters and reduces invalid calculations. Then, a permutation and combination algorithm is used to combine the resolution adjustment data, model architecture adjustment data, and calculation accuracy adjustment data. For example, if there are 3 resolution values, 2 model architectures, and 2 calculation accuracy values, then 3×2×2 = 12 parameter adjustment combination sets can be generated. Among them, the parameter adjustment combination set refers to multiple groups of candidate adjustment schemes generated by permuting and combining parameters such as resolution, model architecture, and calculation accuracy in the parameter adjustment data set. Each combination set contains a complete set of resolution, model architecture, and calculation accuracy parameters. Finally, for each parameter adjustment combination set, in combination with the real - time matching value and the adjustment constraint conditions, and using a multi - objective optimization algorithm (such as the non - dominated sorting genetic algorithm NSGA - II) for evaluation, the combinations that meet the constraints are sorted, and the combination with the highest real - time matching value is selected as the optimal parameter adjustment combination, and the optimal parameter adjustment combination is returned to the step of obtaining the real - time matching value as the initial image - processing parameters for recalculating the real - time matching value to verify the adjustment effect. Through this closed - loop adjustment method, it is possible to ensure the continuous optimization of the system, which is beneficial to adapting to the dynamic changes of the environment and system state.

[0033] The present application also provides an optimization system for an embedded image recognition algorithm, including: A first acquisition module, configured to acquire multimodal monitoring data monitored by a sensor, where the multimodal monitoring data includes scene feature data, multimodal environment data, and hardware status data; A second acquisition module, configured to acquire risk assessment feature data according to the multimodal environment data; A third acquisition module, configured to acquire a dynamic risk level according to the risk assessment feature data and the scene feature data; A fourth acquisition module, configured to acquire computing performance data and operating performance data according to the hardware status data, and acquire initial image processing parameters according to the computing performance data, operating performance data, and dynamic risk level; A fifth acquisition module, configured to acquire a real-time matching value according to the computing performance data, risk assessment feature data, and initial image processing parameters; A judgment module, configured to judge whether the real-time matching value is greater than a preset threshold; If the real-time matching value is greater than the preset threshold, it is determined that the calculation delay meets the real-time requirement; If the real-time matching value is not greater than the preset threshold, it is determined that the calculation delay does not meet the real-time requirement, and the initial image processing parameters are adjusted according to the real-time matching value and the dynamic risk level until the real-time matching value is greater than the preset threshold.

[0034] In one embodiment, the third acquisition module includes: A first acquisition unit, configured to acquire spatial positioning data, optical feature parameters, and structural topology parameters according to the scene feature data, and acquire a regional function type according to the spatial positioning data and the structural topology parameters; A second acquisition unit, configured to acquire a material risk level according to the regional function type and the optical feature parameters, and acquire a spatial attribute category according to the regional function type, material risk level, and structural topology parameters; A third acquisition unit, configured to perform a correlation analysis on the color feature vector, dynamic texture entropy value matrix, and thermodynamic gradient tensor with the spatial attribute category respectively to obtain corresponding modal correlation coefficients, and acquire corresponding associated weight indexes according to the modal correlation coefficients; A fourth acquisition unit, configured to acquire a normalized feature data set according to the color feature vector, dynamic texture entropy value matrix, and thermodynamic gradient tensor, and acquire a dynamic risk level according to the normalized feature data set and the associated weight index.

[0035] The present invention also provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the above-mentioned optimization method for the embedded image recognition algorithm are implemented.

[0036] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned optimization method for the embedded image recognition algorithm are implemented.

[0037] Those of ordinary skill in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above embodiments of the methods. Among them, any reference to a memory, storage, database, or other medium provided in the present application and used in the embodiments can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0038] It should be noted that in this article, the term "including", "comprising", or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, device, article, or method including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such a process, device, article, or method. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of another identical element in the process, device, article, or method including that element.

[0039] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, shall be similarly included in the patent protection scope of the present invention.

Claims

1. An optimization method for an embedded image recognition algorithm, characterized in that Including: Obtain multi-modal monitoring data monitored by sensors, where the multi-modal monitoring data includes scene feature data, multi-modal environment data, and hardware status data; Obtain risk assessment feature data according to the multi-modal environment data; Obtain a dynamic risk level according to the risk assessment feature data and the scene feature data; Obtain computing performance data and operating performance data according to the hardware status data, and obtain initial image processing parameters according to the computing performance data, operating performance data, and dynamic risk level; Obtain a real-time matching value according to the computing performance data, risk assessment feature data, and initial image processing parameters; Judge whether the real-time matching value is greater than a preset threshold; If the real-time matching value is greater than the preset threshold, it is determined that the calculation delay meets the real-time requirement; If the real-time matching value is not greater than the preset threshold, it is determined that the calculation delay does not meet the real-time requirement, and the initial image processing parameters are adjusted according to the real-time matching value and the dynamic risk level until the real-time matching value is greater than the preset threshold.

2. The optimization method for an embedded image recognition algorithm according to claim 1, characterized in that The step of obtaining risk assessment feature data according to the multi-modal environment data includes: Obtain spectral image data, thermal radiation image data, and gaseous molecular data according to the multi-modal environment data, and obtain a spectral alignment image and a thermal radiation alignment image according to the spectral image data and the thermal radiation image data; Obtain the global chromaticity variance according to the spectral alignment image, and obtain a spectral feature vector according to the global chromaticity variance through principal component analysis; Obtain timestamp feature points according to the spectral alignment image and the thermal radiation alignment image, and obtain time-aligned gas data according to the timestamp feature points and the gaseous molecular data; Obtain a CO concentration sequence and a concentration gradient according to the time-aligned gas data, and obtain a color feature vector according to the CO concentration sequence, concentration gradient, and spectral feature vector; Obtain high-temperature centroid coordinates and a temperature gradient matrix according to the thermal radiation alignment image, and obtain a spatio-temporal adjacency matrix according to the high-temperature centroid coordinates; Obtain a thermodynamic gradient tensor according to the spatio-temporal adjacency matrix and the temperature gradient matrix, and obtain a dynamic texture entropy value matrix according to the spatio-temporal adjacency matrix.

3. The optimization method for an embedded image recognition algorithm according to claim 1, characterized in that The step of obtaining a dynamic risk level according to the risk assessment feature data and the scene feature data includes: Obtain spatial positioning data, optical feature parameters, and structural topology parameters according to the scene feature data, and obtain a regional function type according to the spatial positioning data and the structural topology parameters; Obtain a material risk level according to the regional function type and the optical feature parameters, and obtain a spatial attribute category according to the regional function type, material risk level, and structural topology parameters; Perform a correlation analysis on the color feature vector, dynamic texture entropy value matrix, and thermodynamic gradient tensor with the spatial attribute category respectively to obtain corresponding modal correlation coefficients, and obtain corresponding associated weight indices according to the modal correlation coefficients; Obtain a normalized feature dataset based on the color feature vector, the dynamic texture entropy value matrix, and the thermodynamic gradient tensor, and obtain a dynamic risk level based on the normalized feature dataset and the associated weight index.

4. The optimization method for an embedded image recognition algorithm according to claim 1, characterized in that The step of obtaining initial image processing parameters according to the computing performance data, the running performance data, and the dynamic risk level includes: Obtain the core frequency, the reference frequency, the architecture coefficient, the utilization margin, the remaining memory capacity, and the memory fragmentation rate according to the computing performance data, and obtain a comprehensive resource index according to the core frequency, the reference frequency, the architecture coefficient, and the utilization margin; Obtain the memory health according to the remaining memory capacity and the memory fragmentation rate; Obtain the real-time operating temperature, the safety threshold temperature, and the maximum allowable temperature according to the running performance data, and obtain the computing health according to the real-time operating temperature, the safety threshold temperature, and the maximum allowable temperature; Obtain the resolution parameter according to the comprehensive resource index and the computing health; Obtain a risk-resource mapping table, and obtain a computing requirement benchmark and a memory requirement benchmark according to the risk-resource mapping table and the dynamic risk level; Obtain the model architecture parameter according to the memory requirement benchmark and the memory health, and obtain the computing precision parameter according to the comprehensive resource index and the computing requirement benchmark.

5. The optimization method for an embedded image recognition algorithm according to claim 1, characterized in that The step of obtaining the real-time matching value according to the computing performance data, the risk assessment feature data, and the initial image processing parameters includes: Obtain the effective value of the hardware floating-point computing power, the chip junction temperature, the number of cycle operations, the memory bandwidth, and the processing cycle according to the computing performance data, and obtain the real-time CPU computing power according to the effective value of the hardware floating-point computing power, the core frequency, the number of cycle operations, the chip junction temperature, and the safety threshold temperature; Obtain the processing requirement data according to the risk assessment feature data, and obtain the processing time threshold and the algorithm computing amount according to the processing requirement data; Obtain the theoretical processing time according to the algorithm computing amount and the real-time CPU computing power, and obtain the algorithm matching degree according to the theoretical processing time and the processing time threshold; Obtain the available memory space according to the remaining memory capacity, the memory fragmentation rate, the memory bandwidth, and the processing cycle; Obtain the memory requirement space and the required bandwidth according to the resolution parameter, the model architecture parameter, and the computing precision parameter, and obtain the memory matching degree according to the memory requirement space, the available memory space, the memory bandwidth, and the required bandwidth; Obtain the real-time matching value according to the algorithm matching degree and the memory matching degree.

6. The optimization method for an embedded image recognition algorithm according to claim 1, wherein The step of adjusting the initial image processing parameters according to the real-time matching value and the dynamic risk level includes: Obtain the adjustment intensity level according to the real-time matching value and the dynamic risk level, and obtain the adjustment constraint condition according to the dynamic risk level; Obtain a parameter adjustment dataset according to the adjustment intensity level and the adjustment constraint condition, where the parameter adjustment dataset includes multiple resolution adjustment data, model architecture adjustment data, and computing precision adjustment data; Arrange and combine the multiple resolution adjustment data, model architecture adjustment data, and computing precision adjustment data to obtain multiple parameter adjustment combination sets; Screen the multiple parameter adjustment combination sets according to the real-time matching value and the adjustment constraint conditions to obtain the optimal parameter adjustment combination, and return the optimal parameter adjustment combination as the initial image processing parameter to the step of obtaining the real-time matching value according to the computing performance data, risk assessment feature data, and initial image processing parameter.

7. An optimization system for an embedded image recognition algorithm, characterized in that, Including: A first acquisition module for acquiring multi-modal monitoring data monitored by a sensor, where the multi-modal monitoring data includes scene feature data, multi-modal environment data, and hardware status data; A second acquisition module for acquiring risk assessment feature data according to the multi-modal environment data; A third acquisition module for acquiring a dynamic risk level according to the risk assessment feature data and the scene feature data; A fourth acquisition module for acquiring computing performance data and operating performance data according to the hardware status data, and acquiring an initial image processing parameter according to the computing performance data, operating performance data, and dynamic risk level; A fifth acquisition module for acquiring a real-time matching value according to the computing performance data, risk assessment feature data, and initial image processing parameter; A judgment module for judging whether the real-time matching value is greater than a preset threshold; If the real-time matching value is greater than the preset threshold, it is determined that the calculation delay meets the real-time requirement; If the real-time matching value is not greater than the preset threshold, it is determined that the calculation delay does not meet the real-time requirement, and the initial image processing parameter is adjusted according to the real-time matching value and the dynamic risk level until the real-time matching value is greater than the preset threshold.

8. The optimization system for an embedded image recognition algorithm according to claim 7, wherein, The third acquisition module includes: A first acquisition unit for acquiring spatial positioning data, optical feature parameters, and structural topology parameters according to the scene feature data, and acquiring a regional function type according to the spatial positioning data and the structural topology parameters; A second acquisition unit for acquiring a material risk level according to the regional function type and the optical feature parameters, and acquiring a spatial attribute category according to the regional function type, material risk level, and structural topology parameters; A third acquisition unit for performing a correlation analysis on the color feature vector, dynamic texture entropy value matrix, and thermodynamic gradient tensor with the spatial attribute category respectively to obtain corresponding modal correlation coefficients, and acquiring corresponding associated weight indices according to the modal correlation coefficients; A fourth acquisition unit for acquiring a normalized feature data set according to the color feature vector, dynamic texture entropy value matrix, and thermodynamic gradient tensor, and acquiring a dynamic risk level according to the normalized feature data set and the associated weight index.

9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Flue fire hazard early warning method based on image intelligent identification

    CN111223262A

  • Multi-modal expert collaborative fire risk monitoring method for large-scale forestry resources

    CN119151301A

  • Active flame identification and positioning method, system and equipment

    CN119723342A

  • Kitchen safety monitoring method and system based on multi-modal AI large model

    CN119892545A

  • Forest fire monitoring method based on edge intelligence

    CN119992733A

Cited By

  • Early fire point identification method and system based on multispectrum

    CN120766429A

  • A method and system for early fire detection based on multispectral spectrum

    CN120766429B

  • OLED defect detection equipment based on machine vision

    CN120997177A

  • Lightweight real-time image recognition system and method for intelligent terminal

    CN122019202A

  • A lightweight real-time image recognition system and method for intelligent terminals

    CN122019202B