Image calculation method, framework and image calculation model training method

By dynamically identifying dynamic areas through optical flow intensity calculation and sparse computing modules, combined with lightweight FlowNetS and mask optimization models, the problems of high energy consumption and redundant computing in high-resolution image processing are solved, and energy efficiency balance and reasonable allocation of computing resources in intelligent image acquisition equipment are achieved.

CN120689371APending Publication Date: 2025-09-23EAPIL
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510823485.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

The high energy consumption and redundant computing problems caused by high-resolution image processing have limited the development and application scenarios of intelligent image acquisition equipment, especially in mobile devices that rely on battery power, where battery life has become a bottleneck. At the same time, the existing fixed ROI solution has a high missed detection rate in complex scenarios, reducing the reliability of monitoring.

Method used

By calculating the optical flow intensity of the target image, dynamically identifying dynamic areas, combining the sparse computing module and the mask optimization model, rationally allocating computing resources, and adopting a lightweight FlowNetS variant and mask optimization model, energy consumption is reduced and computing efficiency is improved.

Benefits of technology

It achieves flexible adjustment of computing resources in complex scenarios, reduces energy consumption, improves computing efficiency and accuracy, adapts to the needs of different application scenarios, and reduces the amount of calculation and missed detection rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689371A_ABST
    Figure CN120689371A_ABST
Patent Text Reader

Abstract

The invention provides an image calculation method, a framework and an image calculation model training method, which are applied to a hardware system in image acquisition equipment, and the method comprises the following steps: calculating the corresponding optical flow intensity of a target object in a target image; determining a dynamic region in the target image according to a relationship between the optical flow intensity and a motion intensity threshold; and determining a calculation result of the target image by calculating the dynamic region. According to the embodiment of the invention, the optical flow intensity of the target image is calculated, the dynamic region in the target image is dynamically identified based on the optical flow intensity, and the target image is calculated by calculating the dynamic region, so that the flexible adjustment of the dynamic region, the reasonable allocation of calculation resources, the reduction of the calculation amount and the reduction of the energy consumption can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing, and more specifically, to an image computing method, architecture, and image computing model training method. Background Art

[0002] With the rapid development of technologies such as the Internet of Things (IoT), smart security, and mobile terminals, image acquisition devices, as core sensing devices, are widely used in smart homes, unmanned inspections, wearable devices, and autonomous driving. High-resolution image processing has become a core competitive advantage for intelligent image acquisition devices, with cameras offering 4K and even 8K resolutions becoming increasingly common. However, the high energy consumption, redundant computation, and lack of real-time performance associated with high-resolution video processing have severely hindered the further development of intelligent image acquisition devices. Summary of the Invention

[0003] In view of this, the purpose of the embodiments of the present application is to provide an image computing method, architecture and image computing model training method, which can reduce the computing power of the image acquisition device and reduce the energy consumption of the image acquisition device.

[0004] In a first aspect, an embodiment of the present application provides an image calculation method, which is applied to a hardware system in an image acquisition device. The method includes: calculating the optical flow intensity corresponding to the target object in the target image; determining the dynamic area in the target image based on the relationship between the optical flow intensity and the motion intensity threshold; and determining the calculation result of the target image by calculating the dynamic area.

[0005] In the above implementation process, by calculating the optical flow intensity of the target image, and dynamically identifying the dynamic area in the target image based on the optical flow intensity, and then realizing the calculation of the target image by calculating the dynamic area, flexible adjustment of the dynamic area can be achieved, computing resources can be reasonably allocated, the amount of calculation can be reduced, and energy consumption can be reduced.

[0006] In one embodiment, calculating the optical flow intensity corresponding to the target object in the target image includes: expanding each pixel point in the target image through a quadratic polynomial; solving the polynomial coefficients of the quadratic polynomial by weighted least squares method; and determining the optical flow vector by minimizing the difference between the local polynomial coefficients in two adjacent frames of images; wherein the optical flow vector is configured to reflect the optical flow intensity.

[0007] In the above implementation, the Farneback algorithm is used to calculate optical flow intensity. The entire calculation process is based on traditional mathematical principles, making it simple and easy to implement. Furthermore, the optical flow of each frame relies on the comparison between the current and previous frames, which is both real-time and dynamic. This improves the accuracy of optical flow intensity and, in turn, the real-time determination of dynamic regions.

[0008] In one embodiment, calculating the optical flow intensity corresponding to the target object in the target image includes: downsampling the target image to obtain an image pyramid of different levels; wherein the image pyramid is arranged layer by layer from high to low resolution, with the bottom layer of the image pyramid having the highest resolution and the top layer having the lowest resolution; starting from the top layer of the image pyramid, passing the optical flow result of the previous layer as an initial value to the next layer, and calculating the optical flow vector of each layer; when calculating the bottom layer of the image pyramid, adjusting the optical flow vector based on the details of the high-resolution image to output a dense optical flow field.

[0009] In the above implementation, by leveraging the powerful feature extraction capabilities of convolutional neural networks, we can more accurately process optical flow calculations in complex motion scenes, expanding application scenarios. In addition, layer-by-layer calculations can improve the robustness and accuracy of the algorithm.

[0010] In one embodiment, the calculation of the optical flow of each level includes: taking the optical flow vector of the previous level as the initial value; minimizing the grayscale difference between adjacent frame images through an iterative algorithm; solving the polynomial coefficients through weighted least squares method; and updating the optical flow vector according to the polynomial coefficients.

[0011] In the above implementation process, when calculating the optical flow of each layer, the optical flow vector of the previous layer is used as the initial value to update the optical flow vector of the current layer, which can make the calculated optical flow vector more accurate and improve the accuracy of the optical flow vector.

[0012] In one embodiment, determining the dynamic area in the target image based on the relationship between the optical flow intensity and the motion intensity threshold includes: adjusting the motion intensity threshold according to sparsity; generating a binary mask according to the adjusted motion intensity threshold, the optical flow intensity and a sparse mask generation formula; and determining the dynamic area in the target image based on the binary mask.

[0013] In the above implementation process, by adjusting the motion intensity threshold according to the sparsity, a dynamic mask can be further generated based on the motion intensity threshold, thereby realizing flexible adjustment of the dynamic area. The size of the dynamic area can be flexibly controlled according to actual needs, thereby increasing the application scenarios of the image calculation method.

[0014] In one embodiment, adjusting the motion intensity threshold according to sparsity includes: obtaining an environmental complexity parameter of the target environment corresponding to the target image; mapping the environmental complexity parameter into a dynamic adjustment coefficient according to a preset rule; and adjusting the motion intensity threshold according to the dynamic adjustment coefficient and a set calculation formula.

[0015] In the above implementation process, by adjusting the motion intensity threshold according to the environmental complexity parameter, the dynamic area segmentation can be made more suitable for complex scenes, thereby increasing the applicability of the image calculation method for complex scenes. In one embodiment, the method further includes: obtaining the target detection accuracy, actual power consumption, and actual delay of the hardware system; adjusting the power consumption weight coefficient and the delay weight coefficient based on the target detection accuracy, the actual power consumption, the actual delay, and the objective function; optimizing the background sparsity based on the adjusted power consumption weight coefficient; and / or optimizing the target area coverage based on the adjusted delay weight coefficient.

[0016] In the above implementation process, since the objective function takes into account the three key factors of average target detection accuracy, actual power consumption and actual delay, by adjusting the power consumption weight coefficient and the delay weight coefficient, the optimal energy efficiency ratio can be found according to the requirements of different application scenarios, thereby achieving the energy efficiency balance of the system.

[0017] In a second aspect, an embodiment of the present application also provides an image computing model training method, including: the computing model is configured to execute the image computing method of the above-mentioned first aspect, or any possible implementation of the first aspect; the computing model includes an optical flow network and a mask optimization model; the method includes: lightweight processing of FlowNetS to obtain a FlowNetS variant; training the FlowNetS variant through training data to obtain an optical flow network; wherein, the optical flow network is configured to output a dense optical flow field; adjusting the model parameters of the mask optimization model through the Adam optimizer body; determining the target area coverage based on the intersection-and-union ratio between the predicted mask and the true mask; continuing to adjust the model parameters of the mask optimization model through the Adam optimizer body until the target area coverage reaches the target coverage; wherein, the input of the mask optimization model is the optical flow calculation result and the target image, and the output of the mask optimization model is the optimized mask.

[0018] In the above implementation, when training the optical flow network, the corresponding FlowNetS is lightweighted, which can reduce the computational complexity of the optical flow network and improve operational efficiency. When training the mask optimization model, the model parameters are iteratively updated to ensure that the target area coverage determined by the mask optimization model reaches the target coverage. This effectively balances target detection accuracy and computational complexity, improving model performance.

[0019] In one embodiment, the training of the FlowNetS variant using training data to obtain an optical flow network includes: adjusting the weighted importance of motion boundaries during the training of the FlowNetS variant using a loss function; wherein the loss function combines endpoint errors and motion boundary weights; using an edge detection algorithm to calculate the edges of objects in a target image; and adjusting the learning rate of the optical flow network using a cosine annealing learning rate; wherein the learning rate is configured to adjust the prediction accuracy of the optical flow network.

[0020] In the above implementation process, pre-training on a large-scale dataset enables the optical flow network to accurately learn the characteristics and patterns of optical flow. Furthermore, by using an edge detection algorithm to calculate the edges of objects in the target image, motion boundaries can be effectively highlighted, improving the adaptability and accuracy of the optical flow network for complex motion scenes. Furthermore, by adjusting the learning rate of the optical flow network using a cosine annealing learning rate, the model can be slowly adjusted in areas close to the optimal solution, avoiding missing the optimal solution, improving the model's prediction accuracy, and preventing overfitting.

[0021] In a third aspect, an embodiment of the present application further provides an image computing architecture, which is configured to execute the image computing method of the first aspect above, or any possible implementation of the first aspect; the image computing architecture includes: a storage unit, an optical flow perception module, a sparse computing module and an interface circuit; the optical flow perception module is connected to the sparse computing module through the interface circuit; wherein, the optical flow perception module is configured to calculate the corresponding optical flow intensity of the target object in the target image; the sparse computing module is configured to determine the dynamic area in the target image; the storage unit is configured to store the weights and biases of the optical flow network and / or mask optimization model.

[0022] In the above implementation process, by setting an optical flow perception module and a sparse computing module in the image computing architecture, performing optical flow calculation based on the optical flow perception module, and controlling the computing area based on the sparse computing module, it is possible to achieve a reasonable allocation of image computing resources, improve image computing efficiency, and reduce image computing energy consumption.

[0023] In one embodiment, the interface circuit includes: a buffer, a level converter, and a timing controller; the buffer is configured to store data output by the optical flow perception module; the level converter is configured to convert the level signal output by the optical flow perception module into the level signal required by the sparse computing module; and the timing controller is configured to coordinate the timing of data transmission.

[0024] In the above implementation process, by setting up the interface circuit, data transmission between the optical flow perception module and the sparse computing module can be realized, thereby realizing the reasonable allocation of image computing resources through optical flow, improving image computing efficiency, and reducing image computing energy consumption.

[0025] In a fourth aspect, an embodiment of the present application further provides an electronic device comprising: a processor and a memory, wherein the memory stores machine-readable instructions executable by the processor, and when the electronic device is running, the machine-readable instructions are executed by the processor to perform the steps of the method in the above-mentioned first aspect, or any possible implementation of the first aspect, the second aspect, or any possible implementation of the second aspect.

[0026] In a fifth aspect, an embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the computer program executes the steps of the method in the first aspect, or any possible implementation of the first aspect, the second aspect, or any possible implementation of the second aspect.

[0027] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the following embodiments are given in conjunction with the accompanying drawings for detailed description as follows. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. It should be understood that the following drawings only show certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without creative work.

[0029] Figure 1 A schematic diagram of an image computing architecture provided in an embodiment of the present application; Figure 2 A flowchart of the image calculation method provided in an embodiment of the present application; Figure 3 A flowchart of the image computing model training method provided in an embodiment of the present application; Figure 4 A schematic diagram of the functional modules of the image computing device provided in an embodiment of the present application; Figure 5 Schematic diagram of the functional modules for image computing model training provided in an embodiment of the present application. DETAILED DESCRIPTION

[0030] The technical solutions in the embodiments of the present application will be described below in conjunction with the accompanying drawings in the embodiments of the present application.

[0031] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings. At the same time, in the description of this application, the terms "first", "second", etc. are only used to distinguish the description and should not be understood as indicating or implying relative importance.

[0032] With the rapid development of the Internet of Things, smart security, autonomous driving, and portable devices, image acquisition devices, as core sensor components, are seeing their applications continue to expand. However, traditional image acquisition device technologies generally face the technical bottleneck of high energy consumption.

[0033] On the one hand, traditional image acquisition devices consume excessive power when performing AI calculations on full-frame images and processing 4K resolution (3840×2160 pixels) video. This poses a battery life bottleneck for battery-powered mobile devices. With the widespread adoption of 5G technology, the data transmission volume of intelligent image acquisition devices has increased significantly. This high energy consumption requires frequent charging or battery replacement, significantly limiting their application scenarios and ease of use.

[0034] On the other hand, in scenarios where static backgrounds account for over 70% of the image (such as indoor surveillance), approximately 90% of computing resources are wasted on unchanging background areas. While existing fixed-area ROI (Region of Interest) solutions can reduce computational complexity, they suffer from a high miss detection rate of 15%. When moving objects exceed the preset area, the system cannot detect them in a timely manner, reducing monitoring reliability. For example, in intelligent security systems, improper fixed ROI settings can prevent intrusions into critical areas from being detected promptly, posing a safety hazard.

[0035] In view of this, the present application proposes an image calculation method, which calculates the optical flow intensity of the target image, dynamically identifies the dynamic area in the target image based on the optical flow intensity, and then calculates the target image by calculating the dynamic area. This can achieve flexible adjustment of the dynamic area, reasonable allocation of computing resources, reduce the amount of calculation, and reduce energy consumption.

[0036] Before starting the embodiments of the present application, several parameters in the embodiments of the present application are explained: Farneback algorithm: a classic algorithm for optical flow estimation.

[0037] FlowNetS: An optical flow estimation model based on convolutional neural networks. The input of FlowNetS is two consecutive target images, and the output is an optical flow field.

[0038] AdamW optimizer: An adaptive learning rate optimization algorithm. It aims to solve the problem of improper coupling between weight decay and adaptive learning rate mechanisms in the traditional Adam optimizer.

[0039] To facilitate understanding of this embodiment, the application scenario of an image calculation method disclosed in the embodiment of this application is first introduced in detail. The image calculation method is applied to an image acquisition device, which is used to acquire a target image.

[0040] Optionally, the image acquisition device may be a camera, a skynet, a driving recorder, a pan / tilt camera, etc. The image acquisition device may be selected according to actual conditions.

[0041] Among them, the image acquisition device can be used in security monitoring, autonomous driving, smart home and other fields, and the application scenario of the image acquisition device can be selected according to actual conditions.

[0042] The image acquisition device here is provided with a corresponding image computing architecture, which is used to calculate the optical flow of the target image and determine the image computing area according to the optical flow.

[0043] Optionally, the image computing architecture is implemented in a heterogeneous computing chip, which can combine the advantages of multiple computing units to efficiently process different types of tasks. For example, the chip can integrate one or more of a CPU, NPU, DSP, etc. The units integrated in the chip can be selected based on actual conditions.

[0044] Among them, the CPU is responsible for system management and complex logic control, the NPU is specifically used for neural network calculations, and the DSP is good at digital signal processing such as optical flow calculation and mask generation.

[0045] To facilitate understanding of this embodiment, an image computing architecture disclosed in the embodiment of this application is described in detail below.

[0046] like Figure 1 FIG. 1 is a schematic diagram of an image computing architecture provided by an embodiment of the present application. The image computing architecture 100 includes: a storage unit 111 , an optical flow sensing module 112 , a sparse computing module 113 and an interface circuit 114 .

[0047] The optical flow perception module 112 and the sparse calculation module 113 are connected via an interface circuit 114 .

[0048] The optical flow perception module 112 here is configured to calculate the optical flow intensity corresponding to the target object in the target image, and the optical flow perception module 112 outputs an optical flow vector.

[0049] Optionally, an optical flow calculation unit may be embedded in the optical flow perception module 112 , and the optical flow calculation unit supports optical flow calculation modes such as the Farneback algorithm and lightweight deep learning.

[0050] The Farneback algorithm has an independent computational process. It first preprocesses the two input images, converting the color image into grayscale and smoothing it with a Gaussian filter to reduce computational complexity and the effects of noise. It then uses a polynomial to approximate the image grayscale variation within the local neighborhood of each pixel. The polynomial coefficients are solved using weighted least squares. The polynomial coefficients for the corresponding local neighborhoods of two adjacent image frames are then compared to calculate the optical flow vector. This entire process is based on traditional mathematical algorithms and does not rely on deep learning models.

[0051] Lightweight deep learning utilizes the powerful feature extraction capabilities of convolutional neural networks to more accurately handle optical flow calculations in complex motion scenes. This module can output dense optical flow fields, support a certain resolution range, have low computing latency and low power consumption.

[0052] The lightweight deep learning and Farneback algorithm here use different optical flow calculation methods. The two complement each other and can be applied to different application scenarios.

[0053] It should be understood that when designing the image computing architecture 100, the optical flow perception module 112 can be encapsulated into an independent function or class to ensure that its input and output interfaces are clear and fully dependent on the input image data, so as not to have excessive coupling with subsequent mask generation and detection modules, thereby improving the independence of the optical flow perception block.

[0054] The sparse computing module 113 is configured to determine the dynamic area in the target image. The sparse computing module 113 is configured to control the computing area based on the mask, thereby achieving reasonable allocation of computing resources.

[0055] In one embodiment, a mask generation module is further provided, which can be provided in the image computing architecture 100 or in a corresponding software structure. The location of the mask generation module can be selected according to actual conditions.

[0056] The mask generation module here uses the optical flow output by the optical flow perception module 112 as input, combines it with the target image, and generates a mask through a specific algorithm (for example, combining the optical flow result with the target image and using an optimization algorithm to find the optimal mask).

[0057] Optionally, the mask generation module can be designed as an independent functional unit, with the input being the optical flow and the target image, and the output being the mask.

[0058] In another embodiment, a target detection module may also be provided. The target detection module receives the target image and the mask output by the mask generation module as input and performs target detection. The target detection algorithm in the target detection module may be YOLO, Faster R-CNN, etc., and the target detection algorithm may be selected based on actual conditions.

[0059] Optionally, the target detection module can also be encapsulated as an independent function or class.

[0060] It should be understood that input and output data queues can be set for independent modules in each independent stage. For example, the output queue of the optical flow perception module 112 serves as the input queue of the mask generation module, and the output queue of the mask generation module serves as the input queue of the object detection module. This allows data transfer and buffering between different stages.

[0061] Each independent stage can be assigned an independent thread or process for processing, which can fully utilize hardware resources and achieve parallel computing.

[0062] The storage unit 111 is configured to store weights and biases of the optical flow network and / or the mask optimization model.

[0063] The interface circuit 114 here is used to realize the physical connection between the optical flow perception module 112 and the sparse calculation module 113.

[0064] In the sparse computing module 113, the dense optical flow field data output by the optical flow perception module 112 needs to be interfaced with the sparse computing module 113. To implement this interface, data format matching must first be performed. The output of the optical flow perception module 112 is usually dense optical flow field data, generally presented in the form of a two-dimensional matrix or tensor, where each element in the matrix represents the optical flow vector of the corresponding pixel point (including horizontal and vertical motion information). However, dense optical flow field data may have specific format requirements for input data, such as specific data types (e.g., fixed-point or floating-point numbers), data layout (e.g., row-first or column-first), etc. Therefore, data format conversion is required at the interface. For example, if the optical flow perception module 112 outputs an optical flow field in floating-point format, and the sparse computing module 113 requires a fixed-point format, a quantization operation is required to convert the floating-point number into a fixed-point number of appropriate bit width.

[0065] To ensure that the output data from the optical flow sensing module 112 is accurately transmitted to the sparse computation module 113, a suitable communication protocol must be defined. This may involve factors such as the data transmission rate, transmission method (e.g., parallel or serial transmission), and handshake signals. For example, when using a serial communication protocol, parameters such as the start bit, stop bit, data bits, and parity bit must be defined to ensure error-free data transmission. Furthermore, handshake signals (e.g., the sender's "data ready" signal and the receiver's "receive ready" signal) are used to coordinate data transmission.

[0066] In the above implementation process, by setting the optical flow perception module 112 and the sparse computing module 113 in the image computing architecture 100, performing optical flow calculation based on the optical flow perception module 112, and controlling the computing area based on the sparse computing module 113, it is possible to achieve reasonable allocation of image computing resources, improve image computing efficiency, and reduce image computing energy consumption.

[0067] In a possible implementation, the interface circuit 114 includes a buffer, a level converter, and a timing controller.

[0068] The buffer is configured to store data output by the optical flow sensing module 112. The buffer can be used to temporarily store or permanently store the data output by the optical flow sensing module 112 to match the processing speed of the mask generation module.

[0069] The level converter here is configured to convert the level signal output by the optical flow perception module 112 into the level signal required by the sparse calculation module 113 .

[0070] The timing controller is configured to coordinate the timing of data transmission to ensure that data can be transmitted and processed at the correct time.

[0071] In the above implementation process, by setting the interface circuit 114, data transmission between the optical flow perception module 112 and the sparse computing module 113 can be realized, thereby realizing the reasonable allocation of image computing resources through optical flow, improving image computing efficiency, and reducing image computing energy consumption.

[0072] The image computing architecture in this embodiment can be used to execute each step in each method provided in the embodiments of this application. The implementation process of the image computing method is described in detail below through several embodiments.

[0073] See also Figure 2 , is a flow chart of the image calculation method provided in the embodiment of the present application. Figure 2 The specific process shown is explained in detail.

[0074] Step 201: Calculate the optical flow intensity corresponding to the target object in the target image.

[0075] Among them, optical flow refers to the motion information of an object in an image between two consecutive frames of images, and a two-dimensional vector is usually used to represent the movement speed and direction of each pixel.

[0076] The target image here refers to one or more images acquired by the image acquisition device.

[0077] Optionally, the optical flow intensity may be calculated by a Farneback algorithm or by lightweight deep learning. The calculation method of the optical flow intensity may be selected according to actual conditions.

[0078] Among them, the Farneback algorithm uses a polynomial to approximate the image grayscale changes in the local neighborhood of each pixel point, solves the polynomial coefficients by weighted least squares method, and then compares the polynomial coefficients of the local neighborhood corresponding to two adjacent frames of images to calculate the optical flow vector.

[0079] Lightweight deep learning uses the powerful feature extraction capabilities of convolutional neural networks to handle optical flow calculations in complex motion scenes.

[0080] In one embodiment, before step 201 , the method further includes: pre-processing the target image.

[0081] The preprocessing steps here may include: converting the color image into a grayscale image, smoothing the image using a Gaussian filter, and removing noise from the image.

[0082] It should be understood that since optical flow calculation mainly focuses on the grayscale changes of the image, color information has little impact on optical flow calculation. The amount of calculation can be reduced by converting the color image into a grayscale image.

[0083] In addition, since noise may cause errors in optical flow calculation, smoothing can improve the accuracy of the calculation.

[0084] Step 202: Determine the dynamic area in the target image according to the relationship between the optical flow intensity and the motion intensity threshold.

[0085] The motion intensity threshold here is a preset threshold (eg, 0.5-5 pixels / frame), and the motion intensity threshold is used to distinguish between still and moving areas.

[0086] In one embodiment, the dynamic area in the target image is determined based on the relationship between the optical flow intensity and the motion intensity threshold, which can be determined by the optical flow-guided sparse mask generation formula: ; in, Pixel The optical flow vector, is the generated binary mask, is the exercise intensity threshold, represents the dynamic area, and 0 represents the static area.

[0087] The dynamic area here refers to the area where active calculation is required, that is, the pixels in this area are more moving and need to be processed later. The static area refers to the area where calculation can be skipped, that is, the pixels in this area are not moving significantly or are relatively static.

[0088] Step 203: Determine the calculation result of the target image by calculating the dynamic area.

[0089] It should be understood that since the dynamic area refers to the area in the target image where the pixel motion is more obvious, the target image can be calculated by calculating the dynamic area.

[0090] The calculation of the target image here refers to calculating the actions, abnormal situations, etc. occurring in the corresponding area of ​​the target image based on the target image.

[0091] It is understandable that after acquiring the target image, the image acquisition device needs to process the target image to determine the actions, abnormal conditions, etc. occurring in the target image, and then monitor the corresponding area monitored by the image acquisition device.

[0092] In the above implementation process, by calculating the optical flow intensity of the target image, and dynamically identifying the dynamic area in the target image based on the optical flow intensity, and then realizing the calculation of the target image by calculating the dynamic area, flexible adjustment of the dynamic area can be achieved, computing resources can be reasonably allocated, the amount of calculation can be reduced, and energy consumption can be reduced.

[0093] In one possible implementation, step 201 includes: expanding each pixel in the target image through a quadratic polynomial; solving the polynomial coefficients of the quadratic polynomial by weighted least squares method; and determining the optical flow vector by minimizing the difference between the polynomial coefficients in two adjacent frames of image.

[0094] Among them, for a two-dimensional image , can be expanded into a quadratic polynomial in the local neighborhood: .in, , , , , , are the polynomial coefficients.

[0095] The polynomial coefficients here can be solved as follows: For a pixel point The error function can be defined as: ; in, is a weight function, and Gaussian weight is usually used to emphasize the importance of the center pixel of the neighborhood.

[0096] By the error function about , , , , , By taking the partial derivatives and setting them to zero, we can get a system of linear equations, and solving this system of equations can give us the polynomial coefficients.

[0097] After obtaining the polynomial coefficients of each pixel in two adjacent frames, the optical flow is calculated by comparing the polynomial coefficients of the corresponding local neighborhoods in the two adjacent frames. Assume that in the first frame, the pixel The polynomial coefficients are , the corresponding pixel in the second frame image The polynomial coefficients are ,in is the optical flow vector.

[0098] By minimizing the difference between the polynomial coefficients in two adjacent frames, the optical flow vector can be obtained. Specifically, the optical flow can be calculated by solving the following minimization problem to obtain the optical flow vector : ; in, is the pixel coordinate in the target image, A local neighborhood centered on a pixel. The size and shape of the neighborhood are determined by the algorithm. Calculations are performed within this local neighborhood to obtain motion information for the pixels within it. In the first frame image, the pixel In the local neighborhood centered on , the coefficients of image grayscale changes are expressed by quadratic polynomial expansion; The corresponding pixel in the second frame image is In the local neighborhood of , the coefficients of the image grayscale change are expressed by quadratic polynomial expansion. Here is based on the optical flow vector ( , ), from the first frame image The position moves to the corresponding position after the second frame image.

[0099] Furthermore, the optical flow vector can be calculated as follows: [ ] represents the polynomial approximation of the grayscale values ​​of pixels in the local neighborhood of the first frame image. [ ] represents the polynomial approximation of the grayscale value of the pixel in the local neighborhood of the corresponding position in the second frame image. By taking the difference between the polynomial approximations of the corresponding local neighborhoods in the two frames of images and taking the square of the difference in the local neighborhood The internal summation is: ; The sum value reflects the difference between the two frames in the local neighborhood. By minimizing the sum value, the optimal optical flow vector ( , ), so that the difference between the corresponding local neighborhoods in the two frames of images is minimized, thereby determining the motion information of the pixels and realizing optical flow calculation.

[0100] The optical flow vector is configured to reflect the optical flow intensity.

[0101] In the above implementation, the Farneback algorithm is used to calculate optical flow intensity. The entire calculation process is based on traditional mathematical principles, making it simple and easy to implement. Furthermore, the optical flow of each frame relies on the comparison between the current and previous frames, which is both real-time and dynamic. This improves the accuracy of optical flow intensity and, in turn, the real-time determination of dynamic regions.

[0102] In one possible implementation, step 201 includes: downsampling the target image to obtain image pyramids of different levels; starting from the top level of the image pyramid, passing the optical flow results of the previous level as initial values ​​to the next level, and calculating the optical flow vectors of each level; when calculating the bottom level of the image pyramid, adjusting the optical flow vectors based on the details of the high-resolution image to output a dense optical flow field.

[0103] The downsampling process here is as follows: First, convolution is performed on the input target image using a hardware-integrated Gaussian filter to suppress high-frequency noise and prevent aliasing after downsampling. The filter parameters can be dynamically adjusted via hardware configuration registers to adapt to different resolutions (e.g., a 5×5 kernel is used for 4K images), thereby reducing noise. The smoothed image is then subsampled with a fixed 2×2 step size, with pixel decimation achieved via a data selector or shift register, directly outputting a lower-resolution image. The number of pyramid levels is dynamically determined by the input resolution. For example, a 4K image generates a three-level pyramid (4K → 2K → 1080P), while a 1024×768 image generates a two-level pyramid (1024×768 → 512×384).

[0104] The image pyramid is arranged layer by layer from high to low resolution, with the bottom layer at the highest resolution and the top layer at the lowest resolution. The resolutions of adjacent layers decrease by a factor of 2. The number of layers is determined by the input image resolution.

[0105] The low-resolution image at the top of the image pyramid contains the overall image structure and global motion trends, requiring minimal computation and making it suitable for quickly estimating initial optical flow values. The low-resolution image at the bottom of the pyramid retains detailed textures and local motion features, which are used to refine optical flow estimation. Through multi-scale computation, the algorithm balances robustness to global motion with accuracy in local details.

[0106] The calculation begins at the top level of the pyramid and calculates optical flow at each level. The calculation process begins at the top level, where a coarse-grained optical flow estimate is obtained (for example, the optical flow vector is initially set to zero) due to its low resolution and minimal computational effort. The optical flow result from the top level is passed to the next level as the initial search starting point for the optical flow calculation of the corresponding pixel area in the next level.

[0107] It should be understood that because the bottom layer of the image pyramid is the original resolution image, it retains rich texture details and local motion features. When performing optical flow vector correction, the initial optical flow value passed down from the top layer is used as a basis, and an iterative algorithm such as the Gauss-Newton method is used to minimize the grayscale differences between adjacent images. This grayscale difference calculation fully utilizes the precise pixel information in the high-resolution image.

[0108] The process of iteratively optimizing the optical flow vector can be as follows: During the underlying calculation process, the polynomial coefficients are continuously solved through iterative algorithms, such as weighted least squares, to update the optical flow estimate. Each iteration adjusts the optical flow vector based on the details of the high-resolution image. As the number of iterations increases, the optical flow vector gradually converges to a more accurate value, ultimately outputting a high-precision dense optical flow field. In each iteration, the pixel difference between corresponding positions in adjacent frames is calculated based on the current optical flow vector. The optical flow vector is then adjusted based on this difference, gradually reducing the difference. This continuously optimizes the optical flow estimate, and the output optical flow field more accurately reflects the actual movement of objects in the image.

[0109] In the above implementation, by leveraging the powerful feature extraction capabilities of convolutional neural networks, we can more accurately process optical flow calculations in complex motion scenes, expanding application scenarios. In addition, layer-by-layer calculations can improve the robustness and accuracy of the algorithm.

[0110] In one possible implementation, the optical flow of each level is calculated, including: taking the optical flow vector of the previous level as the initial value; minimizing the grayscale difference between adjacent frame images through an iterative algorithm; solving the polynomial coefficients through a weighted least squares method; and updating the optical flow vector according to the polynomial coefficients.

[0111] The iterative algorithm here can be a Gauss-Newton method, LM algorithm, quasi-Newton method, conjugate gradient method or the like, and the iterative algorithm can be selected according to actual conditions.

[0112] Among them, minimizing the grayscale difference between adjacent frame images through iterative algorithm can be expressed as: ; in, and are two adjacent frames of images, is a local neighborhood.

[0113] In the above implementation process, when calculating the optical flow of each layer, the optical flow vector of the previous layer is used as the initial value to update the optical flow vector of the current layer, which can make the calculated optical flow vector more accurate and improve the accuracy of the optical flow vector.

[0114] In one possible implementation, step 202 includes: adjusting a motion intensity threshold according to sparsity; generating a binary mask according to the adjusted motion intensity threshold, optical flow intensity, and a sparse mask generation formula; and determining a dynamic area in the target image according to the binary mask.

[0115] Sparsity is a quantitative metric used to represent the proportion of areas skipped (or inactive) during the computation process, a quantitative description of the allocation of computing resources. A higher sparsity means more unnecessary calculations can be skipped, significantly reducing power consumption. For example, in the sparse computation module, the sparse mask generated based on optical flow information can determine which areas require computation and which areas can be skipped. As sparsity increases, fewer pixels are involved in the computation, reducing the hardware's computational workload and power consumption.

[0116] Optionally, the sparsity can be adjusted in the range of 10%-90%. By adjusting the threshold, the size of the activation area can be flexibly controlled to meet the needs of different scenarios.

[0117] It should be understood that the process of dynamically adjusting the motion intensity threshold according to the motion intensity of the current scene is closely related to the adjustment of sparsity. When the scene motion intensity is low, the optical flow changes less, and the value of the motion intensity threshold can be appropriately increased. In this way, the generated sparse mask will be sparser, that is, more pixels are judged as not requiring calculation, thereby increasing the overall sparsity. On the contrary, when the scene motion intensity is high, the value of the motion intensity threshold is lowered to ensure that important motion information is not missed, while also maintaining sparsity to a certain extent and reducing unnecessary calculations.

[0118] The above binary mask can be expressed as: ; in, Pixel The optical flow vector, is the generated binary mask, is the exercise intensity threshold, represents the dynamic area, and 0 represents the static area.

[0119] In one embodiment, the objective function can be optimized by an online feedback mechanism. and Weights can also indirectly affect sparsity. and Respectively represent the importance of target area coverage and background sparsity in the objective function. When the system detects that the power consumption is high, it can be appropriately increased. The weight of , emphasizes the optimization of background sparsity, thus generating a sparser mask and reducing power consumption. On the contrary, when the accuracy of target detection needs to be guaranteed, the weight of The weight of , appropriately reduce the sparsity to improve the coverage of the target area.

[0120] In the above implementation process, by adjusting the motion intensity threshold according to the sparsity, a dynamic mask can be further generated based on the motion intensity threshold, thereby realizing flexible adjustment of the dynamic area. The size of the dynamic area can be flexibly controlled according to actual needs, thereby increasing the application scenarios of the image calculation method.

[0121] In one possible implementation, the motion intensity threshold is adjusted according to sparsity, including: obtaining an environmental complexity parameter of a target environment corresponding to a target image; mapping the environmental complexity parameter into a dynamic adjustment coefficient according to a preset rule; and adjusting the motion intensity threshold according to the dynamic adjustment coefficient and a set calculation formula.

[0122] The environment complexity parameter may include the density of dynamic objects in the target scene, the range of illumination variation, etc. The environment complexity parameter may be selected according to actual conditions.

[0123] The preset rules here can be to match the dynamic adjustment coefficient corresponding to the environmental complexity parameter through a preset environment-threshold mapping table, or to determine the dynamic adjustment coefficient corresponding to the environmental complexity parameter through a preset formula, or to output the dynamic adjustment coefficient corresponding to the environmental complexity parameter by setting the model, etc. The preset rules can be selected according to actual conditions.

[0124] The above setting calculation formula can be: ; in, is the adjusted exercise intensity threshold, is the initial exercise intensity threshold, is the adaptive weight factor, is the dynamic adjustment coefficient.

[0125] In the above implementation process, by adjusting the motion intensity threshold according to the environment complexity parameter, the dynamic area division can be made more suitable for complex scenes, thereby increasing the applicability of the image calculation method to complex scenes.

[0126] In one possible implementation, the method also includes: obtaining the target detection accuracy, actual power consumption and actual delay of the hardware system; adjusting the power consumption weight coefficient and the delay weight coefficient based on the target detection accuracy, actual power consumption, actual delay and the objective function; optimizing the background sparsity based on the adjusted power consumption weight coefficient; and / or optimizing the target area coverage based on the adjusted delay weight coefficient.

[0127] Among them, the objective function is the energy efficiency optimization objective function. This objective function can be expressed as: ; in, is the actual power consumption, is the actual delay, is the average precision of target detection, is the power consumption weight coefficient, is the delay weight coefficient. The power consumption weight coefficient here is used to indicate the importance of target area coverage in the objective function. The delay weight coefficient is used to indicate the importance of background sparsity in the objective function.

[0128] Among them, when When reduced due to sparse activation, the average precision of target detection Maintain stability (e.g., no less than 80%) and end-to-end processing delay (Unit: milliseconds) When the real-time requirements are met, the denominator The value of is reduced, so that the entire energy efficiency optimization objective function Increasing the value of can optimize energy efficiency. That is, the energy efficiency ratio of the system can be improved by reducing computing power.

[0129] In one embodiment, .

[0130] in, is the power consumption of the optical flow perception module, is the power consumption of the sparse computing module.

[0131] It should be understood that in the image computing architecture of the embodiment of the present application, the storage unit is used to store the power consumption monitoring model. The power consumption monitoring model is used to collect the computing power consumption of the optical flow perception module and the processing power consumption of the sparse computing module in real time. Based on the target detection accuracy, actual power consumption and actual delay, the system energy efficiency ratio can be expressed as: ; when When the energy efficiency falls below the preset threshold, the power consumption weight coefficient and delay weight coefficient are adaptively adjusted.

[0132] The adjustment strategy here can be: , then increase the power consumption weight coefficient to reduce the power consumption of optical flow calculation; like If the real-time requirement is exceeded, the delay weight coefficient is increased to optimize the processing delay.

[0133] Understandably, the objective function comprehensively considers three key factors: average target detection accuracy, actual power consumption, and actual latency, to evaluate the system's energy efficiency. We identify the optimal energy efficiency ratio based on the requirements of different application scenarios, balancing the relationship between average target detection accuracy and average target detection precision to achieve optimal energy efficiency in different scenarios.

[0134] The following examples further demonstrate that the image calculation method in the embodiment of the present application can reduce power consumption and improve detection accuracy: Indoor static scene (90% background): The traditional solution consumes 11.2W of power for target image calculation, and the average target detection accuracy is 82.1%. The solution in the embodiment of the present application reduces power consumption to 3.5W (a 68.7% reduction), with an average target detection accuracy of 80.3% (only a 1.8% loss in accuracy), and a computational sparsity of 72%. In indoor monitoring scenarios, human activity is relatively rare, and the background mostly remains static. The embodiment of the present application identifies static background areas through optical flow detection and skips calculations for these areas, significantly reducing power consumption while maintaining high detection accuracy.

[0135] Street vehicle scene (vehicle speed ≤ 40 km / h): The traditional solution consumes 12.8W of power for target image calculation, with an average target detection accuracy of 85.7%. The embodiment of the present application consumes 6.2W of power for target image calculation (a 51.6% reduction), with an average target detection accuracy of 84.1% and a computational sparsity of 55%. In street vehicle scenes, vehicles move at high speeds, placing high demands on real-time performance. The embodiment of the present application can quickly and accurately detect the vehicle's motion area, rationally allocate computing resources, and ensure accurate vehicle target detection while reducing power consumption. Dense crowd scene (30% motion area): The traditional solution consumes 13.5W of power for target image calculation, with an average target detection accuracy of 79.5%. The embodiment of this application consumes 8.9W of power for target image calculation (a 34.1% reduction), with an average target detection accuracy of 77.8% and a calculation sparsity of 40%. In dense crowd scenes, people move in complex ways, with background and foreground intertwined. This embodiment of the application uses optical flow-guided sparse calculations to effectively identify the motion areas of the crowd, reduce unnecessary calculations, and lower power consumption.

[0136] In the above implementation process, since the objective function takes into account the three key factors of average target detection accuracy, actual power consumption and actual delay, by adjusting the power consumption weight coefficient and the delay weight coefficient, the optimal energy efficiency ratio can be found according to the requirements of different application scenarios, thereby achieving the energy efficiency balance of the system.

[0137] See also Figure 3 , is a flow chart of the image computing model training method provided in the embodiment of the present application. Figure 3 The specific process shown is explained in detail.

[0138] Step 301: Perform lightweight processing on FlowNetS to obtain a FlowNetS variant.

[0139] FlowNetS is an optical flow estimation model based on convolutional neural networks. The input of FlowNetS is two consecutive target images, and the output is an optical flow field.

[0140] The lightweight processing here can include simplifying the FlowNetS structure, compressing parameters, optimizing network design, and other processing methods.

[0141] Simplifying the FlowNetS structure involves simplifying some of the complex and redundant layers that FlowNetS may have originally contained during the lightweight improvement. For example, the number of convolutional layers is reduced, and some auxiliary layers that contribute less to performance improvement are removed. At the same time, the size of the convolution kernel is adjusted, using smaller kernels. For example, larger kernels are replaced with 3×3, 1×1, and other sizes. This not only reduces the number of parameters but also the amount of computation, as smaller kernels require fewer multiplication and addition operations during convolution, thereby improving operational efficiency.

[0142] Parameter compression involves using pruning algorithms to remove redundant connections and parameters from a model. By setting a threshold, weight parameters that have little impact on the model's output can be set to zero. The model can then be retrained, allowing it to maintain good performance despite the reduction in parameters. Quantization techniques can also be used to convert model parameters from higher-precision representations (such as 32-bit floating point) to lower-precision representations (such as 16-bit half-precision or 8-bit integers). This reduces the space required for model storage and computational resources without significantly compromising accuracy, thereby achieving lightweighting and improved efficiency.

[0143] Optimizing network design involves introducing more efficient network modules. For example, the depthwise separable convolution module in MobileNet (a lightweight convolutional neural network designed specifically for mobile and embedded devices, whose core goal is to address the deployment challenges of traditional deep learning models on resource-constrained devices. Through technological innovation, it significantly reduces computational complexity and parameter count while maintaining high accuracy) decomposes traditional convolution into depthwise and pointwise convolutions, reducing computational complexity while maintaining feature extraction capabilities. Alternatively, the channel shuffling operation in ShuffleNet (a lightweight convolutional neural network designed specifically for mobile devices, whose core goal is to optimize computational efficiency and parameter count to adapt to low-computing resource environments while maintaining accuracy) enhances information exchange between different channels and improves model performance without significantly increasing computational complexity. Furthermore, using grouped convolution, which divides the input channels into multiple groups and performs a separate convolution operation on each group, can reduce the complexity of the convolution operation and the number of parameters.

[0144] Step 302: Train the FlowNetS variant using the training data to obtain an optical flow network.

[0145] Among them, the optical flow network is configured to output a dense optical flow field.

[0146] Optionally, during the pre-training phase of the optical flow network, the number of training rounds can be set to 50 rounds.

[0147] In preliminary experiments, we tested 30, 50, and 80 training rounds. The results showed that after 30 rounds of training, the model had not yet fully converged, resulting in low optical flow prediction accuracy. While 80 rounds of training further improved accuracy, overfitting occurred, resulting in a decrease in the model's ability to generalize to new data. However, after 50 rounds of training, the model was able to better learn the characteristics and patterns of optical flow while maintaining good generalization ability, achieving a relatively balanced accuracy and loss on the validation set.

[0148] In one embodiment, the AdamW optimizer can be selected when training a FlowNetS variant. Combining the Adam algorithm's fast convergence characteristics with a weight decay mechanism, the AdamW optimizer effectively prevents model overfitting, enabling the model to achieve good performance even with less training data and fewer training rounds. Furthermore, by properly adjusting training hyperparameters (e.g., the initial learning rate) and finding the optimal initial learning rate through experimentation, the model can converge to better results more quickly during training, reducing training time and improving overall operational efficiency.

[0149] Understandably, when training an optical flow network, a quantization-aware training approach can be employed to adapt the trained model parameters to the hardware's quantization requirements. When deploying the optical flow network on hardware, the quantized model parameters are mapped to the corresponding storage units of the hardware's mask calculation module. For example, the model's weights and biases are stored in a suitable format in the hardware's registers or memory so that they can be quickly read and used during hardware runtime. Secondly, hardware resources are allocated appropriately based on the computational complexity of the optical flow network and the resource availability of the hardware's mask calculation module. For example, the number of multipliers, adders, and registers required to implement the optical flow network's calculations is determined, as well as how to optimize the use of these resources to improve hardware computational efficiency. Furthermore, considering the parallelism of the hardware, the computational tasks of the optical flow network are rationally divided and processed in parallel to fully utilize the hardware's parallel computing capabilities. Through this interface and connection approach, the training results of the optical flow network can be effectively applied to the hardware architecture, enabling optical flow-guided sparse event-driven computation.

[0150] Step 303: Adjust the model parameters of the mask optimization model through the Adam optimizer.

[0151] The input of the mask optimization model is the optical flow calculation result and the target image, and the output of the mask optimization model is the optimized mask.

[0152] The Adam optimizer here combines the advantages of momentum gradient and adaptive learning rate, which can quickly and stably adjust model parameters during training.

[0153] In one embodiment, the hyperparameters of the Adam optimizer can be set as follows: the learning rate is ,This smaller learning rate enables the mask optimization model to adjust parameters more finely when optimizing the mask, avoiding model instability or missing the optimal solution due to excessive learning rate; , used to calculate the first-order moment estimate of the gradient, controlling the influence of past gradient information on the current update; 99, used to calculate the second-order moment estimate of the gradient, plays an important role in the adaptive adjustment of the learning rate; , is a very small constant used to prevent the denominator from being zero and ensure the stability of the algorithm. These hyperparameter settings can effectively balance the target detection accuracy and computational cost reduction in the sparse mask joint optimization stage, enabling the model to achieve better performance.

[0154] Step 304 : determining the target area coverage rate based on the intersection-over-union ratio between the predicted mask and the true mask.

[0155] In one embodiment, the specific calculation formula for the target area coverage can be expressed as: ; Among them, when When the sparsity is greater than or equal to 0.5, the predicted mask is considered to cover the target area. During the sparse mask joint optimization process, the model parameters are continuously adjusted to ensure that the target area coverage of the predicted mask is ≥ 95%. This ensures that the generated mask covers the target area as accurately as possible and avoids missed detections. At the same time, combined with the constraint of background sparsity ≥ 70%, the computational effort can be minimized while maintaining target detection accuracy, achieving model optimization.

[0156] Step 305 : Continue adjusting the model parameters of the mask optimization model through the Adam optimizer until the target area coverage reaches the target coverage.

[0157] It should be understood that by iteratively adjusting the model parameters of the mask optimization model, the target area coverage of the predicted mask can be made to reach a higher probability, thereby covering the target area as accurately as possible and achieving model optimization.

[0158] The optical flow network is used to perform the calculation of light intensity in the image calculation method, and the mask optimization model is used to determine the mask in the image calculation method, thereby determining the dynamic area in the target area.

[0159] As you can understand, after training the optical flow network and mask optimization model, they can be applied to key components of image acquisition devices (such as smart cameras and dashcams). This process must balance multiple requirements, including computational efficiency, power consumption, real-time performance, and hardware compatibility. In terms of technical implementation, model quantization is a key step. Its purpose is to convert floating-point models to a low-precision format, thereby reducing storage and computational overhead.

[0160] Optionally, the method of applying the optical flow network and mask optimization model to an image acquisition device may include mixed-precision quantization. For example, by converting weights from 32-bit floating point to 16-bit half-precision, the model size can be halved, the inference speed can be increased by 30%-50%, and the accuracy loss can be controlled within 1%. The model can also be further quantized to INT8 or INT4, which can increase the inference speed by 2-4 times. In actual operation, you can choose to use frameworks such as TensorRT and ONNXRuntime, which can automatically complete quantization and graph optimization.

[0161] In the above implementation, when training the optical flow network, the corresponding FlowNetS is lightweighted, which can reduce the computational complexity of the optical flow network and improve operational efficiency. When training the mask optimization model, the model parameters are iteratively updated to ensure that the target area coverage determined by the mask optimization model reaches the target coverage. This effectively balances target detection accuracy and computational complexity, improving model performance.

[0162] In one possible implementation, step 302 includes: adjusting the weighted importance of motion boundaries during the training of the FlowNetS variant through a loss function; calculating the edges of objects in the target image using an edge detection algorithm; and adjusting the learning rate of the optical flow network using a cosine annealing learning rate; wherein the learning rate is configured to adjust the prediction accuracy of the optical flow network.

[0163] The loss function here combines the endpoint error and the motion boundary weight. For example, the loss function can be expressed as: ; in, is the predicted optical flow vector, is the true optical flow vector, is the weight parameter, is the loss function.

[0164] It should be understood that the method used to adjust the importance of motion boundary weighting during the pre-training stage of the optical flow network can be: in the early stages of training, the importance of motion boundary weighting can be reduced to allow the optical flow network to first learn the overall optical flow characteristics; and in the later stages of training, the importance of motion boundary weighting can be gradually increased to make the optical flow network pay more attention to the details of the motion boundary area. Of course, an adaptive adjustment strategy can also be designed to automatically adjust the importance of motion boundary weighting based on certain indicators during the training process (such as the loss function value, verification set accuracy, etc.). For example, when the loss function decreases slowly in the motion boundary area, the importance of motion boundary weighting is automatically increased; otherwise, its importance is reduced).

[0165] The above edge detection algorithm has good edge detection performance and can effectively extract the edges of objects in the image.

[0166] In one embodiment, the low threshold of the edge detection algorithm can be set to 30, and the high threshold can be set to 90. The low threshold is used to initially detect weak edges, and the high threshold is used to determine strong edges. The combination of the two can ensure that the detected edges are both complete and accurate. Set to 0.6. In the loss function, Controls the weighted importance of motion boundaries. If it is too small, the model will not pay enough attention to the motion boundary, which may lead to inaccurate optical flow calculation at the edge of the object; If it is too large, the model will focus too much on the motion boundary and may ignore the accuracy of the overall optical flow. A value of 0.6 can effectively highlight motion boundaries while ensuring the overall accuracy of the optical flow, thereby improving the adaptability and accuracy of the optical flow network to complex motion scenes.

[0167] Understandably, when using a cosine annealing learning rate to adjust the learning rate of the optical flow network, the initial learning rate can be set to 1e-4. In the early stages of training, the model rapidly explores the parameter space with a large learning rate, accelerating convergence. As the number of training rounds increases, the learning rate gradually decreases according to the cosine function. In the middle stages of training, a moderate reduction in the learning rate allows the model to more finely adjust parameters and optimize optical flow prediction. Towards the end of training, the learning rate becomes very small, allowing the model to slowly adjust parameters near the optimal solution, avoiding missing the optimal solution, further improving the model's prediction accuracy, and preventing overfitting.

[0168] In the above implementation process, pre-training on a large-scale dataset enables the optical flow network to accurately learn the characteristics and patterns of optical flow. Furthermore, by using an edge detection algorithm to calculate the edges of objects in the target image, motion boundaries can be effectively highlighted, improving the adaptability and accuracy of the optical flow network for complex motion scenes. Furthermore, by adjusting the learning rate of the optical flow network using a cosine annealing learning rate, the model can be slowly adjusted in areas close to the optimal solution, avoiding missing the optimal solution, improving the model's prediction accuracy, and preventing overfitting.

[0169] Based on the same application concept, an image calculation device corresponding to the image calculation method is also provided in the embodiment of the present application. Since the principle of solving the problem by the device in the embodiment of the present application is similar to that of the aforementioned image calculation method embodiment, the implementation of the device in this embodiment can refer to the description in the embodiment of the aforementioned method, and the repeated parts will not be repeated.

[0170] See also Figure 4, is a functional module diagram of the image computing device provided in an embodiment of the present application. Each module in the image computing device in this embodiment is used to perform each step in the above method embodiment. The image computing device includes a computing module 401, a first determining module 402, and a second determining module 403; wherein, The calculation module 401 is used to calculate the optical flow intensity corresponding to the target object in the target image.

[0171] The first determining module 402 is configured to determine a dynamic area in the target image according to a relationship between the optical flow intensity and a motion intensity threshold.

[0172] The second determining module 403 is configured to determine a calculation result of the target image by calculating the dynamic area.

[0173] In one possible implementation, the calculation module 401 is further used to: expand each pixel point in the target image through a quadratic polynomial; solve the polynomial coefficients of the quadratic polynomial by weighted least squares method; and determine the optical flow vector by minimizing the difference between the local polynomial coefficients in two adjacent frames of images; wherein the optical flow vector is configured to reflect the optical flow intensity.

[0174] In one possible implementation, the computing module 401 is further configured to: downsample the target image to obtain an image pyramid of different levels; wherein the image pyramid is arranged layer by layer from high to low resolution, with the bottom layer of the image pyramid having the highest resolution and the top layer having the lowest resolution; starting from the top layer of the image pyramid, passing the optical flow result of the previous layer as an initial value to the next layer to calculate the optical flow vector of each layer; and when calculating the bottom layer of the image pyramid, adjusting the optical flow vector based on the details of the high-resolution image to output a dense optical flow field.

[0175] In one possible implementation, the calculation module 401 is specifically used to: use the optical flow vector of the previous level as the initial value; minimize the grayscale difference between adjacent frame images through an iterative algorithm; solve the polynomial coefficients through weighted least squares method; and update the optical flow vector according to the polynomial coefficients.

[0176] In one possible implementation, the first determination module 402 is further used to: adjust the motion intensity threshold according to sparsity; generate a binary mask according to the adjusted motion intensity threshold, the optical flow intensity and a sparse mask generation formula; and determine the dynamic area in the target image based on the binary mask.

[0177] In one possible embodiment, the image calculation method also includes a first adjustment module for obtaining the target detection accuracy, actual power consumption and actual delay of the hardware system; adjusting the power consumption weight coefficient and the delay weight coefficient according to the target detection accuracy, the actual power consumption, the actual delay and the objective function; optimizing the background sparsity according to the adjusted power consumption weight coefficient; and / or optimizing the target area coverage according to the adjusted delay weight coefficient.

[0178] Based on the same application concept, the embodiment of the present application also provides an image computing model training device corresponding to the image computing model training method. Since the principle of solving the problem by the device in the embodiment of the present application is similar to the aforementioned image computing model training method embodiment, the implementation of the device in this embodiment can refer to the description in the embodiment of the above method, and the repeated parts will not be repeated.

[0179] See also Figure 5 , is a functional module diagram of the image computing model training device provided in an embodiment of the present application. The various modules in the image computing model training device in this embodiment are used to perform the various steps in the above method embodiment. The image computing model training device includes a lightweight module 501, a training module 502, a second adjustment module 503, a third determination module 504, and an iteration module 505; wherein, The lightweight module 501 is used to perform lightweight processing on FlowNetS to obtain a FlowNetS variant.

[0180] The training module 502 is used to train the FlowNetS variant using training data to obtain an optical flow network; wherein the optical flow network is configured to output a dense optical flow field.

[0181] The second adjustment module 503 is used to adjust the model parameters of the mask optimization model through the Adam optimizer.

[0182] The third determining module 504 is configured to determine the target area coverage rate according to an intersection-over-union ratio between the predicted mask and the true mask.

[0183] The iterative module 505 is used to continue adjusting the model parameters of the mask optimization model through the Adam optimizer until the target area coverage reaches the target coverage; wherein the input of the mask optimization model is the optical flow calculation result and the target image, and the output of the mask optimization model is the optimized mask.

[0184] In one possible embodiment, the training module 502 is further used to: adjust the weighted importance of motion boundaries during the training process of the FlowNetS variant through a loss function; wherein the loss function combines endpoint error and motion boundary weighting; use an edge detection algorithm to calculate the edges of objects in the target image; use a cosine annealing learning rate to adjust the learning rate of the optical flow network; wherein the learning rate is configured to adjust the prediction accuracy of the optical flow network.

[0185] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions and operations of the devices, methods and computer program products according to the multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and the module, program segment or a part of the code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented using a dedicated hardware-based system that performs the specified function or action, or can be implemented using a combination of dedicated hardware and computer instructions.

[0186] In addition, the functional modules in each embodiment of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0187] If the functions are implemented in the form of software modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, or the portion that contributes to the prior art, or the portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage media include various media that can store program code, such as USB flash drives, mobile hard drives, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical disks. It should be noted that, in this document, relational terms such as first and second, etc., are used solely to distinguish one entity or operation from another, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising..." does not preclude the presence of additional identical elements in the process, method, article, or apparatus that includes the element. The foregoing description is merely a preferred embodiment of the present application and is not intended to limit the present application. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included within the scope of protection of the present application. It should be noted that similar reference numerals and letters represent similar items in the following figures. Therefore, once an item is defined in one figure, it does not need to be further defined or explained in subsequent figures.

[0188] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. An image calculation method, characterized in that: A hardware system applied to an image acquisition device, the method comprising: Calculate the corresponding optical flow intensity of the target object in the target image; determining a dynamic area in the target image according to a relationship between the optical flow intensity and a motion intensity threshold; The calculation result of the target image is determined by calculating the dynamic area.

2. The method according to claim 1, characterized in that Calculating the optical flow intensity corresponding to the target object in the target image includes: Expand each pixel in the target image using a quadratic polynomial; solving the polynomial coefficients of the quadratic polynomial by weighted least squares method; An optical flow vector is determined by minimizing a difference between local polynomial coefficients in two adjacent frames of image; wherein the optical flow vector is configured to reflect the optical flow intensity.

3. The method according to claim 1, characterized in that Calculating the optical flow intensity corresponding to the target object in the target image includes: Downsampling the target image to obtain image pyramids of different levels; wherein the image pyramids are arranged layer by layer from high to low resolution, with the bottom layer of the image pyramid having the highest resolution and the top layer of the image pyramid having the lowest resolution; Starting from the top layer of the image pyramid, the optical flow result of the previous layer is passed to the next layer as an initial value, and the optical flow vector of each layer is calculated; When calculating the bottom layer of the image pyramid, the optical flow vector is adjusted according to the details of the high-resolution image, and a dense optical flow field is output.

4. The method according to claim 3, characterized in that The calculation of optical flow at each level includes: Use the optical flow vector of the previous level as the initial value; Minimize the grayscale difference between adjacent frame images through an iterative algorithm; Solve the polynomial coefficients by weighted least squares method; The optical flow vector is updated according to the polynomial coefficients.

5. The method according to claim 1, wherein The determining of the dynamic area in the target image according to the relationship between the optical flow intensity and the motion intensity threshold comprises: adjusting the motion intensity threshold according to sparsity; generating a binary mask according to the adjusted motion intensity threshold, the optical flow intensity, and a sparse mask generation formula; A dynamic area in the target image is determined according to the binary mask.

6. The method according to claim 5, characterized in that The adjusting the motion intensity threshold according to the sparsity includes: Obtaining an environmental complexity parameter of a target environment corresponding to the target image; Mapping the environmental complexity parameter into a dynamic adjustment coefficient according to a preset rule; The exercise intensity threshold is adjusted according to the dynamic adjustment coefficient and a set calculation formula.

7. The method according to claim 1, characterized in that The method further comprises: Obtaining target detection accuracy, actual power consumption, and actual latency of the hardware system; Adjusting a power consumption weight coefficient and a delay weight coefficient according to the target detection accuracy, the actual power consumption, the actual delay, and an objective function; Optimizing background sparsity according to the adjusted power consumption weight coefficient; and / or, Optimize target area coverage based on the adjusted delay weight coefficient.

8. A method for training an image computing model, characterized in that: in, The computing model is configured to execute the image computing method according to any one of claims 1 to 7; the computing model includes an optical flow network and a mask optimization model; The method comprises: Lightweight FlowNetS to obtain a FlowNetS variant; Training the FlowNetS variant using training data to obtain an optical flow network; wherein the optical flow network is configured to output a dense optical flow field; Adjust the model parameters of the mask optimization model through the Adam optimizer; Determine the target area coverage based on the intersection-over-union ratio between the predicted mask and the true mask; Continuing to adjust the model parameters of the mask optimization model through the Adam optimizer until the target area coverage reaches the target coverage; The input of the mask optimization model is the optical flow calculation result and the target image, and the output of the mask optimization model is the optimized mask.

9. The method according to claim 8, characterized in that The FlowNetS variant is trained with training data to obtain an optical flow network, including: Adjusting the weighted importance of motion boundaries during training of the FlowNetS variant using a loss function; wherein the loss function combines endpoint errors and motion boundary weights; Use edge detection algorithm to calculate the edges of objects in the target image; A cosine annealing learning rate is used to adjust the learning rate of the optical flow network; wherein the learning rate is configured to adjust the prediction accuracy of the optical flow network.

10. An image computing architecture, characterized in that: The image computing architecture is configured to execute the image computing method according to any one of claims 1 to 7; The image computing architecture includes: a storage unit, an optical flow perception module, a sparse computing module and an interface circuit; The optical flow perception module is connected to the sparse calculation module via the interface circuit; The optical flow perception module is configured to calculate the optical flow intensity corresponding to the target object in the target image; the sparse calculation module is configured to determine the dynamic area in the target image; and the storage unit is configured to store the weights and biases of the optical flow network and / or mask optimization model.

11. The architecture according to claim 10, wherein: The interface circuit includes: a buffer, a level converter, and a timing controller; The buffer is configured to store data output by the optical flow perception module; The level converter is configured to convert the level signal output by the optical flow perception module into the level signal required by the sparse calculation module; The timing controller is configured to coordinate the timing of data transmission.

Citation Information

Patent Citations

  • Video processing method and device, electronic equipment and readable storage medium

    CN115063365A

  • Accurate tracking method for video moving object

    CN115294500A

  • Real-time optical flow processing system based on FPGA and RAFT algorithm

    CN116503496A

  • Dynamic environment SLAM method based on vision

    CN117635660A

  • Dense optical flow calculation system and method based on FPGA

    US20220383521A1