Computer vision target detection method and system
Through dynamic selection of sub-network and adaptive adjustment of computational volume, the problem of wasted computing resources and insufficient detection accuracy in the prior art is solved, and efficient and accurate recognition of computer vision object detection is achieved.
Patent Information
- Application Number
- CN202510616557.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-08-08
AI Technical Summary
The existing object detection model is difficult to adaptively adjust the calculation amount according to the complexity of the image, resulting in wasted computing resources in simple images and insufficient detection accuracy in complex images.
A dynamic network based on complexity score is adopted to dynamically select lightweight, medium complexity or high complexity subnets for detection. Combining dynamic weights and adaptive thresholds, the model calculation volume is adaptively adjusted, multi-level semantic features are extracted using the pyramid network structure of the convolutional neural network, and the duplicate detection box is removed through a non-maximum suppression algorithm.
The balance between computing resources and detection effects in different scenarios is achieved, detection efficiency and accuracy are improved, the adaptability of the model is enhanced, simple images can be quickly processed and targets in complex scenarios are accurately identified.
Smart Images

Figure CN120451510A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of visual detection technology, and in particular to a computer vision target detection method and system. Background Art
[0002] Visual inspection is the use of machines to replace human eyes for measurement and judgment. Visual inspection refers to the use of machine vision products (i.e., image capture devices, divided into CMOS and CCD) to convert the captured target into image signals, which are transmitted to a dedicated image processing system. Based on pixel distribution, brightness, color and other information, these signals are converted into digital signals. The image system performs various operations on these signals to extract the characteristics of the target, and then controls the operation of the equipment on site based on the judgment results.
[0003] Existing target detection models find it difficult to adaptively adjust the amount of computation according to the complexity of the image. When faced with simple images, they may still use complex network structures for detection, resulting in a waste of computing resources. When processing complex images, the network structure may not be strong enough to accurately extract features, affecting detection accuracy. Summary of the Invention
[0004] The purpose of the present invention is to overcome the problems raised by the background technology and provide a computer vision target detection method and system.
[0005] To achieve the above object, the present invention provides the following technical solutions:
[0006] A computer vision target detection method comprises the following steps:
[0007] Acquire image and video data;
[0008] Preprocess image and video data: use image enhancement algorithms to improve image clarity and use filtering algorithms to reduce noise in video data;
[0009] Use deep learning model feature extraction algorithms to extract features from pre-processed image and video data;
[0010] The extracted features are fused and transformed through the fully connected layer, and then the regressor predicts the complexity score of the image based on the fused features;
[0011] Through a pre-designed dynamic network, the corresponding sub-network is dynamically selected for target detection based on the image complexity score, thereby adaptively adjusting the computational load of the model. During the detection process, each sub-network outputs the detection result by predicting the category and location information of the target object;
[0012] The non-maximum suppression algorithm is used to remove duplicate detection frames through the set threshold and retain the detection results that meet the settings.
[0013] As a further solution of the present invention: after acquiring image and video data, an adaptive image enhancement algorithm is used to automatically adjust the enhancement parameters according to the brightness and contrast statistical characteristics of the image, and a Gaussian function is used to perform weighted averaging of pixels in the video data to remove Gaussian noise in the video data.
[0014] As a further solution of the present invention: Using a deep learning model feature extraction algorithm, the method for extracting features from preprocessed image and video data is as follows: adopting a deep learning model based on a convolutional neural network, the deep learning model uses a pyramid network structure, and inputting the image and video data preprocessed by the data acquisition module into the deep learning model, extracting multi-level semantic features on feature maps of different scales, the bottom-level feature map retains the detailed information of the image, and the high-level feature map contains more abstract semantic information, and the features of different levels are combined through cross-layer connection and fusion operations.
[0015] As a further solution of the present invention, the extracted features are fused and transformed by a fully connected layer, and then a regressor predicts the complexity score of the image based on the fused features. Each neuron in the fully connected layer of the convolutional neural network is connected to all elements of the input features, and the input features are weighted and summed using the formula: ; in, Representative The output of a neuron, Represents the connection input element and the The weights of neurons, Represents the input feature elements, Representative The bias of a neuron, Represents the number of neurons, fuses features of different dimensions, performs nonlinear transformation on the weighted summation result by applying activation function, and then inputs the transformed features into the fully connected layer for fusion.
[0016] As a further solution of the present invention: the complexity scoring method is trained using a data set containing images of different complexities and their corresponding complexity labels, and by continuously adjusting the parameters of the regressor, it predicts the complexity score of the image. The complexity score is a scalar value between 0 and 1, 0≤complexity score<0.3 corresponds to low complexity, 0.3≤complexity score<0.7 corresponds to medium complexity, and 0.7≤complexity score≤1 corresponds to high complexity.
[0017] As a further solution of the present invention: the dynamic network includes a lightweight subnetwork, a medium-complexity subnetwork and a high-complexity subnetwork. The dynamic network is designed using the three subnetworks. The lightweight subnetwork adopts BiFPN-Lite, the medium-complexity subnetwork adopts the standard BiFPN, and the high-complexity subnetwork adopts the enhanced version of BiFPN. The dynamic network dynamically selects the corresponding subnetwork for target detection based on the image complexity score output by the complexity evaluation module, and adaptively adjusts the computational complexity of the model. During the detection process, each subnetwork outputs the detection result by predicting the category and location information of the target object.
[0018] As a further solution of the present invention, the dynamic network introduces dynamic weights and adaptive thresholds. The dynamic weights smoothly adjust the participation of different sub-networks according to changes in the image complexity score. The adaptive threshold dynamically adjusts the boundary of sub-network switching according to actual conditions. The adjustment equation uses a logical function to adjust the dynamic weights. At the same time, an adaptive threshold adjustment mechanism based on historical detection results and current complexity scores is combined. The adjustment equation is: ; ; ; ; in, 、 and Represent the selection weights of lightweight sub-network, medium complexity sub-network and high complexity sub-network respectively, and The parameter representing the rate of change of the control logic function, and represents the adaptive threshold, and The initial values of are set to 0.3 and 0.7, Represents the complexity score.
[0019] As a further solution of the present invention: a non-maximum suppression algorithm is used to remove duplicate detection frames through a set threshold, and a method for retaining detection results that meet the settings is as follows: a non-maximum suppression algorithm is used, and a threshold is set so that the intersection-and-union ratio of the detection frame and the retained detection frame is greater than the set threshold. The detection frame highly overlaps with the retained detection frame, which is a repeated detection of the same object. The frame is removed and this process is repeated until all detection frames are processed. The retained detection frame is the accurate and non-duplicate detection result after processing.
[0020] A second aspect of the present invention provides a computer vision target detection system, comprising a camera and a data processing terminal, wherein the computer vision target detection system further comprises the following modules:
[0021] Data acquisition module, which acquires image and video data through the camera,
[0022] The preprocessing module is used to preprocess the collected data, improve the clarity of the image using the image enhancement algorithm, and use the filtering algorithm to reduce the noise of the video data;
[0023] The feature extraction module uses the deep learning model feature extraction algorithm to extract features from the image preprocessed by the data acquisition module;
[0024] The complexity evaluation module fuses and transforms features through a fully connected layer, and then uses a regressor to predict the complexity score of the image based on the fused features;
[0025] The target detection module designs a dynamic network. Based on the image complexity score output by the complexity evaluation module, it dynamically selects the corresponding sub-network for target detection, thereby adaptively adjusting the computational complexity of the model. During the detection process, each sub-network predicts the category and location information of the target object and outputs the detection result.
[0026] The post-processing module uses a non-maximum suppression algorithm to remove duplicate detection frames through a set threshold and retain detection results that meet the settings.
[0027] By adopting the above technical solution, compared with the prior art, the beneficial effects of the present invention are:
[0028] 1. The target detection module of the present invention dynamically selects lightweight, medium complexity, or high complexity sub-networks for detection based on complexity scoring, with the help of dynamic weights and adaptive thresholds. It also adaptively adjusts the model computational load. When faced with simple images, lightweight sub-networks are selected for rapid processing, improving detection efficiency. For complex images, high complexity sub-networks are used to ensure detection accuracy. This balances computing resources and detection effects, enhancing the adaptability of the model in different scenarios.
[0029] 2. The feature extraction module of the present invention utilizes a pyramid network structure based on a convolutional neural network to extract multi-level semantic features on feature maps of different scales, and combines features through cross-layer connections and fusion operations, so that the model can not only capture image details but also understand the overall semantics, thereby improving the processing capabilities of complex scenes and multi-scale targets, and can more accurately locate and identify large and small targets in target detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 Schematic diagram of the process of the computer vision target detection system in an embodiment of the present invention. DETAILED DESCRIPTION
[0031] The specific embodiments of the present invention will be further described below in conjunction with the accompanying drawings. It should be noted that the description of these embodiments is used to help understand the present invention, but does not constitute a limitation of the present invention.
[0032] In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0033] The present invention provides a computer vision target detection method, comprising the following steps:
[0034] Acquire image and video data;
[0035] Preprocess image and video data: use image enhancement algorithms to improve image clarity and use filtering algorithms to reduce noise in video data;
[0036] Use deep learning model feature extraction algorithms to extract features from pre-processed image and video data;
[0037] The extracted features are fused and transformed through the fully connected layer, and then the regressor predicts the complexity score of the image based on the fused features;
[0038] Through a pre-designed dynamic network, the corresponding sub-network is dynamically selected for target detection based on the image complexity score, thereby adaptively adjusting the computational load of the model. During the detection process, each sub-network outputs the detection result by predicting the category and location information of the target object;
[0039] The non-maximum suppression algorithm is used to remove duplicate detection frames through the set threshold and retain the detection results that meet the settings.
[0040] In one embodiment of the present invention: data acquisition uses a camera as a front-end device for data acquisition, and the camera acquires image and video data in real time. After acquiring the image and video data, an adaptive image enhancement algorithm is used to automatically adjust the enhancement parameters based on the brightness and contrast statistical characteristics of the image, and a Gaussian function is used to perform weighted averaging of the pixels in the video data to remove Gaussian noise in the video data.
[0041] In one embodiment of the present invention: feature extraction adopts a deep learning model based on a convolutional neural network, and the deep learning model uses a pyramid network structure. The image and video data preprocessed by the data acquisition module are input into the deep learning model, and multi-level semantic features are extracted on feature maps of different scales. The bottom-level feature maps retain the detailed information of the image, and the high-level feature maps contain more abstract semantic information. Through cross-layer connection and fusion operations, features at different levels are combined.
[0042] In one embodiment of the present invention, the complexity evaluation includes a processing unit and a prediction unit. The image features extracted by the feature extraction module are input to the processing unit. Each neuron in the fully connected layer of the convolutional neural network is connected to all elements of the input features. The input features are weighted and summed using the formula: ; in, Representative The output of a neuron, Represents the connection input element and the The weights of neurons, Represents the input feature elements, Representative The bias of a neuron, Represents the number of neurons, fuses features of different dimensions, performs nonlinear transformation on the weighted summation result by applying activation function, and then inputs the transformed features into the fully connected layer for fusion.
[0043] In one embodiment of the present invention: the regressor in the prediction unit uses linear regression, and is trained using a data set containing images of different complexities and their corresponding complexity labels. By continuously adjusting the parameters of the regressor, it predicts the complexity score of the image. The complexity score is a scalar value between 0 and 1, 0≤complexity score <0.3 corresponds to low complexity, 0.3≤complexity score <0.7 corresponds to medium complexity, and 0.7≤complexity score ≤1 corresponds to high complexity.
[0044] In one embodiment of the present invention: target detection includes a lightweight subnetwork, a medium complexity subnetwork and a high complexity subnetwork. The three subnetworks are used to design a dynamic network. The lightweight subnetwork adopts BiFPN-Lite, the medium complexity subnetwork adopts standard BiFPN, and the high complexity subnetwork adopts an enhanced version of BiFPN. The dynamic network dynamically selects the corresponding subnetwork for target detection based on the image complexity score output by the complexity evaluation module, and adaptively adjusts the computational complexity of the model. During the detection process, each subnetwork outputs the detection result by predicting the category and location information of the target object.
[0045] In one embodiment of the present invention, a dynamic network introduces dynamic weights and adaptive thresholds. The dynamic weights smoothly adjust the participation of different sub-networks according to changes in the image complexity score. The adaptive threshold dynamically adjusts the boundary of sub-network switching according to actual conditions. The adjustment equation uses a logical function to adjust the dynamic weights. At the same time, an adaptive threshold adjustment mechanism based on historical detection results and current complexity scores is combined. The adjustment equation is: ; ; ; ; in, 、 and Represent the selection weights of lightweight sub-network, medium complexity sub-network and high complexity sub-network respectively, and The parameter representing the rate of change of the control logic function, and represents the adaptive threshold, and The initial values of are set to 0.3 and 0.7, Represents the complexity score.
[0046] In one embodiment of the present invention: the detection results are post-processed using a non-maximum suppression algorithm. By setting a threshold, the intersection-over-union ratio of the detection frame and the retained detection frame is greater than the set threshold. The detection frame that highly overlaps with the retained detection frame is a repeated detection of the same object, which is removed. This process is repeated until all detection frames are processed. The retained detection frame is the accurate and non-duplicate detection result after processing.
[0047] Embodiment 1: Features at different levels have different characteristics. Shallow features are close to the input data and contain rich details, such as local information such as edges and textures. They have high resolution but weak semantic information. Deep features are abstracted through multiple layers and have rich semantic information. They can express the overall category and high-level attributes of the target, but have low resolution and lose details. Cross-layer connection and fusion operations combine the two, allowing the model to capture details and understand overall semantics, improving the processing capabilities of complex scenes and multi-scale targets. In target detection, small targets rely on shallow detail features for positioning, and large targets rely on deep semantic features for recognition. After fusion, both large and small target detection can be taken into account.
[0048] Example 2: Complexity evaluation will output an image complexity score. When the dynamic network receives the score, it will select a suitable subnetwork based on the preset threshold range. If the image complexity score is low, it indicates that the image is relatively simple. The dynamic network will select a lightweight subnetwork (BiFPN-Lite) for target detection. BiFPN-Lite reduces the amount of computation and parameters by streamlining the structure, can quickly process simple images, and output the category and location information of the target object. If the image complexity score is at a medium level, the standard BiFPN subnetwork will be enabled. It achieves a good balance between computation and detection accuracy and is suitable for processing images with a certain degree of complexity. For images with a high complexity score, the enhanced BiFPN subnetwork will be called. By adding network layers and a complex connection structure, it has stronger feature extraction capabilities and can handle target detection tasks under complex backgrounds. It also outputs the category and location information of the target object as the detection result.
[0049] Example 3: When setting the detection frame threshold, if the threshold is set too high, more overlapping detection frames will be retained, resulting in a large number of duplicate frames in the detection results, affecting the accuracy and simplicity of the results; if the threshold is set too low, some detection frames that should be retained may be mistakenly deleted, resulting in some targets being missed. In actual applications, it is necessary to adjust the threshold according to the specific detection task and data set characteristics. For some scenes with dense target distribution, such as crowd detection, it may be necessary to appropriately lower the threshold to avoid mistakenly deleting too many detection frames; for scenes with sparse target distribution, such as parking lot vehicle detection, the threshold can be appropriately increased to ensure the simplicity of the detection results.
[0050] Example 4: Figure 1 As shown, this embodiment provides a computer vision target detection system, including a camera and a data processing terminal. The computer vision target detection system also includes the following modules:
[0051] The data acquisition module acquires image and video data through the camera, pre-processes the collected data, uses image enhancement algorithms to improve image clarity, and uses filtering algorithms to reduce noise in video data;
[0052] The feature extraction module uses the deep learning model feature extraction algorithm to extract features from the image preprocessed by the data acquisition module;
[0053] The complexity evaluation module fuses and transforms features through a fully connected layer, and then uses a regressor to predict the complexity score of the image based on the fused features;
[0054] The target detection module designs a dynamic network. Based on the image complexity score output by the complexity evaluation module, it dynamically selects the corresponding sub-network for target detection, thereby adaptively adjusting the computational complexity of the model. During the detection process, each sub-network predicts the category and location information of the target object and outputs the detection result.
[0055] The post-processing module uses a non-maximum suppression algorithm to remove duplicate detection frames through a set threshold and retain detection results that meet the settings.
[0056] The feature extraction module extracts features from the image and video data collected by the data acquisition module. The complexity evaluation module uses the features to score the complexity of the image. The target detection module uses the complexity score to dynamically select the corresponding sub-network for target detection.
[0057] Although the present invention is disclosed above with reference to preferred embodiments, this is not intended to limit the present invention. Any person skilled in the art may make possible changes and modifications without departing from the spirit and scope of the present invention. Therefore, any modifications, equivalent variations, and modifications made to the above embodiments in accordance with the technical essence of the present invention without departing from the content of the technical solution of the present invention shall fall within the scope of protection defined by the claims of the present invention.
Claims
1. A computer vision target detection method, characterized in that: The steps include: Acquire image and video data; Preprocess image and video data: use image enhancement algorithms to improve image clarity and use filtering algorithms to reduce noise in video data; Use deep learning model feature extraction algorithms to extract features from pre-processed image and video data; The extracted features are fused and transformed through the fully connected layer, and then the regressor predicts the complexity score of the image based on the fused features; Through a pre-designed dynamic network, the corresponding sub-network is dynamically selected for target detection based on the image complexity score, thereby adaptively adjusting the computational load of the model. During the detection process, each sub-network outputs the detection result by predicting the category and location information of the target object; The non-maximum suppression algorithm is used to remove duplicate detection frames through the set threshold and retain the detection results that meet the settings.
2. A computer vision target detection method according to claim 1, characterized in that: After acquiring image and video data, an adaptive image enhancement algorithm is used to automatically adjust the enhancement parameters according to the statistical characteristics of image brightness and contrast, and a Gaussian function is used to perform weighted averaging of pixels in the video data to remove Gaussian noise in the video data.
3. The computer vision target detection method according to claim 1, wherein: The method of extracting features from preprocessed image and video data using a deep learning model feature extraction algorithm is as follows: a deep learning model based on a convolutional neural network is adopted. The deep learning model uses a pyramid network structure. The image and video data preprocessed by the data acquisition module are input into the deep learning model. Multi-level semantic features are extracted on feature maps of different scales. The bottom-level feature maps retain the detailed information of the image, and the high-level feature maps contain more abstract semantic information. Features at different levels are combined through cross-layer connection and fusion operations.
4. The computer vision target detection method according to claim 1, wherein: The extracted features are fused and transformed through the fully connected layer, and then the regressor predicts the complexity score of the image based on the fused features. The method is as follows: each neuron in the fully connected layer of the convolutional neural network is connected to all elements of the input features, and the input features are weighted and summed using the formula: ; in, Representative The output of a neuron, Represents the connection input element and the The weights of neurons, Represents the input feature elements, Representative The bias of a neuron, Represents the number of neurons, fuses features of different dimensions, performs nonlinear transformation on the weighted summation result by applying activation function, and then inputs the transformed features into the fully connected layer for fusion.
5. The computer vision target detection method according to claim 1, wherein: The complexity scoring method is trained using a dataset containing images of different complexities and their corresponding complexity labels. By continuously adjusting the parameters of the regressor, it predicts the complexity score of the image. The complexity score is a scalar value between 0 and 1. 0 ≤ complexity score < 0.3 corresponds to low complexity, 0.3 ≤ complexity score < 0.7 corresponds to medium complexity, and 0.7 ≤ complexity score ≤ 1 corresponds to high complexity.
6. A computer vision target detection method according to claim 1, characterized in that: The dynamic network includes a lightweight subnetwork, a medium-complexity subnetwork and a high-complexity subnetwork. The dynamic network is designed using the three subnetworks. The lightweight subnetwork adopts BiFPN-Lite, the medium-complexity subnetwork adopts the standard BiFPN, and the high-complexity subnetwork adopts the enhanced version of BiFPN. The dynamic network dynamically selects the corresponding subnetwork for target detection based on the image complexity score output by the complexity evaluation module, and adaptively adjusts the computational complexity of the model. During the detection process, each subnetwork outputs the detection result by predicting the category and location information of the target object.
7. A computer vision target detection system according to claim 6, characterized in that: The dynamic network introduces dynamic weights and adaptive thresholds. The dynamic weights smoothly adjust the participation of different sub-networks according to the changes in the image complexity score. The adaptive threshold dynamically adjusts the boundary of sub-network switching according to the actual situation. The adjustment equation uses a logical function to adjust the dynamic weights. At the same time, it combines an adaptive threshold adjustment mechanism based on historical detection results and the current complexity score. The adjustment equation is: ; ; ; ; in, 、 and Represent the selection weights of lightweight sub-network, medium complexity sub-network and high complexity sub-network respectively, and The parameter representing the rate of change of the control logic function, and represents the adaptive threshold, and The initial values of are set to 0.3 and 0.7, Represents the complexity score.
8. The computer vision target detection method according to claim 1, wherein: The non-maximum suppression algorithm is used to remove duplicate detection frames through a set threshold, and the method for retaining detection results that meet the settings is as follows: the non-maximum suppression algorithm is used to set a threshold, and the intersection-and-union ratio of the detection frame and the retained detection frame is greater than the set threshold. The detection frame highly overlaps with the retained detection frame, which is a repeated detection of the same object. It is removed and this process is repeated until all detection frames are processed. The retained detection frame is the accurate and non-duplicate detection result after processing.
9. A computer vision target detection system, comprising a camera and a data processing terminal, characterized in that: The computer vision target detection system also includes the following modules: Data acquisition module, which acquires image and video data through the camera, The preprocessing module is used to preprocess the collected data, improve the clarity of the image using the image enhancement algorithm, and use the filtering algorithm to reduce the noise of the video data; The feature extraction module uses the deep learning model feature extraction algorithm to extract features from the image preprocessed by the data acquisition module; The complexity evaluation module fuses and transforms features through a fully connected layer, and then uses a regressor to predict the complexity score of the image based on the fused features; The target detection module designs a dynamic network. Based on the image complexity score output by the complexity evaluation module, it dynamically selects the corresponding sub-network for target detection, thereby adaptively adjusting the computational complexity of the model. During the detection process, each sub-network predicts the category and location information of the target object and outputs the detection result. The post-processing module uses a non-maximum suppression algorithm to remove duplicate detection frames through a set threshold and retain detection results that meet the settings.
Citation Information
Cited By
Machine vision-based landform contour intelligent acquisition system
CN120997430A
Data generation method and device based on artificial intelligence, computer equipment and medium
CN121262441A
Queen bee identification method and device
CN121482570A