Target detection system and method based on computer vision

Through the improved convolutional neural network architecture and the optimized YOLO algorithm, combined with multimodal data fusion and dynamic confidence threshold adjustment, the problem of insufficient real-time and small-objective detection accuracy in the prior art is solved, and a more efficient object detection effect is achieved.

CN120107535APending Publication Date: 2025-06-06CHINA THREE GORGES UNIV
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510221259.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-27
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing deep learning object detection methods have shortcomings in real-time and small object detection accuracy, especially in complex backgrounds and occlusions.

Method used

Using an improved convolutional neural network architecture and an optimized YOLO algorithm, combined with regional suggestion networks and multimodal data fusion, the confidence threshold is dynamically adjusted to improve the accuracy and real-timeness of object detection.

Benefits of technology

It significantly improves the accuracy and real-time nature of object detection, enhances the adaptability to small targets and complex scenarios, and is suitable for real-time monitoring and other scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107535A_ABST
    Figure CN120107535A_ABST
Patent Text Reader

Abstract

The invention discloses a target detection system and method based on computer vision, and relates to the technical field of vision technologies. Comprising a data acquisition module used for acquiring image or video data; the data preprocessing module is used for cutting, zooming, normalizing and enhancing the collected data; the feature extraction module is used for extracting feature information in the image; the target detection module is used for generating candidate target regions and carrying out classification and bounding box regression on the candidate regions in combination with a region suggestion network and a classifier; the multi-modal data fusion module is used for fusing the RGB image data and other modal data so as to improve the accuracy and robustness of target detection; the post-processing module is used for performing non-maximum suppression, confidence threshold screening and result visualization processing on the detection result; according to the invention, the accuracy and real-time performance of target detection are significantly improved; the robustness to small targets and complex scenes is improved; and the development cost and maintenance difficulty of the algorithm are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of visual technology, and in particular to a computer vision-based target detection system and method. Background Art

[0002] In the field of computer vision, target detection is one of the core tasks and is widely used in security monitoring, autonomous driving, industrial inspection, robot vision and other fields. Traditional target detection methods mainly rely on manual feature extraction and classic machine learning algorithms, such as HOG features combined with SVM classifiers. However, these methods often have problems such as low detection accuracy and poor generalization ability when facing complex scenes and diverse targets.

[0003] In recent years, the rapid development of deep learning technology has brought new opportunities to the field of computer vision. Object detection algorithms based on convolutional neural networks (CNNs), such as R-CNN, Fast R-CNN, Faster R-CNN, and YOLO, have significantly improved the performance of object detection by automatically learning image features. However, the existing deep learning object detection methods still have some shortcomings. For example, some algorithms perform poorly in real-time and are difficult to meet the needs of real-time monitoring and other scenarios; other algorithms have insufficient accuracy in small target detection, especially in complex backgrounds and occlusions. Therefore, a computer vision-based object detection system and method are proposed to solve the above problems. Summary of the invention

[0004] The present invention provides a computer vision-based target detection system and method, which improves the accuracy and real-time performance of target detection through an optimized deep learning model and an improved detection process, while enhancing the adaptability to small targets and complex scenes, so as to solve the problems in the background technology.

[0005] To achieve the above object, the present invention provides the following technical solution: a computer vision-based target detection system and method, comprising: A data acquisition module, used to acquire image or video data; Data preprocessing module, used to crop, scale, normalize and enhance the collected data; Feature extraction module, based on an improved convolutional neural network architecture, is used to extract feature information from images; The target detection module combines the region proposal network and the classifier to generate candidate target regions and classify and regress the candidate regions into bounding boxes to achieve accurate positioning and recognition of targets. Multimodal data fusion module, used to fuse RGB image data with other modal data (such as infrared images, depth images or lidar data) to improve the accuracy and robustness of target detection; The post-processing module is used to perform non-maximum suppression, confidence threshold screening and result visualization on the detection results; Output module: used to output detection results, including target category and location information.

[0006] Furthermore, the data acquisition module is used to collect images or video streams to be detected, and the data acquisition module includes a data source interface module, a data cache module and a data transmission module; The data source interface module is responsible for connecting to different data sources and reading data; The data cache module is responsible for caching the preprocessed data to improve data reading efficiency; The data transmission module is responsible for transmitting the cached data to the data preprocessing module.

[0007] Furthermore, the improved convolutional neural network architecture in the feature extraction module includes: Feature highlighting module: used to highlight the features of the target area, for the detection of small targets and review scenes; Multi-scale feature fusion module: used to fuse feature maps of different scales to better capture the global and local information of the target.

[0008] Furthermore, the salient feature module calculates the weight distribution of the feature map by the following formula: ; Where: W att represents the attention weight, σ represents the activation function, W q and W k is a learnable weight matrix, X is the input feature map; The multi-scale feature fusion module realizes the fusion of feature maps of different scales through the following formula: ; Among them: F msf Represents the fused feature map, F s1 and F s2 They represent feature maps of different scales respectively, and α is the weight coefficient, which is used to balance the contribution of feature maps of different scales.

[0009] Furthermore, the target detection module includes: Region Proposal Network (RPN): used to generate candidate target regions; Classifier: Target detection is performed on the feature map based on the improved YOLO algorithm to classify the candidate target area and determine which target category it belongs to; Bounding box regressor: used to perform bounding box regression on the candidate target area and accurately locate the position of the target.

[0010] Furthermore, the improved YOLO algorithm introduces a category balancing mechanism, which is implemented by the following formula: ; Among them, L class Represents the loss of class probability, L ce represents the cross entropy loss, w i is the class weight, which is used to balance the loss contribution of different classes, and C is the total number of classes.

[0011] Furthermore, the multimodal data fusion module adopts a weighted fusion strategy to dynamically adjust the weights of different modal data according to the characteristics of the target. The multimodal data fusion module includes: Data alignment module: responsible for the spatial and temporal alignment of data of different modalities; Feature extraction submodule: responsible for extracting features from the data of each modality; Feature fusion module: responsible for secondary fusion of features of different modalities.

[0012] Furthermore, the post-processing module includes the following modules: Non-maximum suppression module: used to remove redundant detection boxes and retain the detection results with the highest confidence; Bounding box regression module: used to optimize the position and size of the target bounding box; The post-processing module increases the accuracy and robustness of the detection results by dynamically adjusting the confidence threshold. The dynamic confidence threshold adjustment mechanism is as follows: When the complexity of the detection scene increases, false detections are reduced by increasing the confidence threshold; When the complexity of the detection scene decreases, the recall rate of detection is improved by lowering the confidence threshold.

[0013] Furthermore, the mathematical expression of the dynamic confidence threshold adjustment mechanism is as follows: ; in, T thres represents the adjusted confidence threshold, T base represents the basic threshold, β is the adjustment factor, N obj Indicates the number of detected targets, N total Indicates the total number of targets.

[0014] The computer vision-based target detection method comprises the following steps: Step S1: collecting the image or video stream to be detected through the data acquisition module; Step S2: using a preprocessing module to perform preprocessing operations such as normalization, cropping, and scaling on the collected images; Step S3: extracting feature information from the image through a feature extraction module, wherein the feature extraction module adopts an improved convolutional neural network architecture; Step S4: Detect the target on the feature map using a target detection module, and output the category and location information of the target. The target detection module is based on an improved YOLO algorithm, optimizes the loss function of target detection, and introduces a category balance mechanism. Step S5: perform multimodal data fusion, combine RGB image data with other modal data, and perform feature fusion through weighted fusion strategy; Step S6: Optimize the detection results through the post-processing module, perform non-maximum suppression, confidence threshold screening and result visualization on the detection results.

[0015] Step S7: Output the detection results, including target category and location information.

[0016] Compared with the prior art, the present invention provides a computer vision-based target detection system and method, which has the following beneficial effects: The computer vision-based target detection system and method can detect targets more accurately through an improved CNN architecture and an optimized YOLO algorithm, and perform excellently in small targets and complex scenes. The improved YOLO algorithm significantly improves the detection speed while maintaining high precision, and is suitable for scenes such as real-time monitoring. The target detection system of the present invention can adapt to detection scenes of different complexities by dynamically adjusting the confidence threshold, and has wide applicability. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0018] Figure 1 A system control schematic diagram of a computer vision-based target detection system and method of the present invention; Figure 2 A schematic diagram of a convolutional neural network architecture of a computer vision-based target detection system and method of the present invention; Figure 3A schematic diagram of a target detection module of a computer vision-based target detection system and method of the present invention; Figure 4 A schematic diagram of a multimodal data fusion module system of a computer vision-based target detection system and method of the present invention; Figure 5 A schematic diagram of a post-processing module of a computer vision-based target detection system and method of the present invention; Figure 6 The present invention is a schematic diagram of a dynamic confidence threshold adjustment mechanism of a computer vision-based target detection system and method. DETAILED DESCRIPTION

[0019] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings.

[0020] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.

[0021] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor is it a separate or selective embodiment that is mutually exclusive with other embodiments.

[0022] See also Figure 1-6The present invention discloses a computer vision-based target detection system and method, including: a data acquisition module for acquiring image or video data; a data preprocessing module for cropping, scaling, normalizing and enhancing the collected data; a feature extraction module, based on an improved convolutional neural network architecture, for extracting feature information from an image; a target detection module, in combination with a region proposal network and a classifier, for generating candidate target regions and classifying and regressing the candidate regions into bounding boxes, so as to achieve accurate positioning and recognition of targets; a multimodal data fusion module, for fusing RGB image data with other modal data (such as infrared images, depth images or lidar data) to improve the accuracy and robustness of target detection; a post-processing module, for performing non-maximum suppression, confidence threshold screening and result visualization processing on the detection results; an output module: for outputting the detection results, including target categories and location information; the present application improves the accuracy of target detection: in complex backgrounds, occlusions and multi-target scenes, the recognition accuracy is significantly improved; enhances real-time performance: compared with traditional algorithms, the target detection speed is faster and more suitable for application scenarios such as real-time monitoring; improves the ability to detect small targets: It can identify small targets more accurately and solve problems commonly encountered by traditional algorithms.

[0023] An improved convolutional neural network architecture is used to strengthen the extraction of target features and enhance the adaptability to small targets and complex scenes.

[0024] Combined with multimodal data fusion, the complementary information from different data sources is utilized to improve the accuracy and robustness of target detection.

[0025] Based on the dynamic confidence threshold adjustment mechanism, the overall accuracy of target detection results is improved.

[0026] Data acquisition module: responsible for collecting image or video data, just like the eyes collect visual information.

[0027] Data preprocessing module: organize the data to improve the efficiency of subsequent algorithms, just like cleaning up pictures.

[0028] Feature extraction module: responsible for analyzing data and extracting key information, just like our brain recognizes objects.

[0029] Target detection module: Based on the extracted information, it identifies the target and determines the category and location, just like the brain analyzes and marks the purpose of the object.

[0030] Multimodal data fusion module: Combines other data sources, such as depth information, audio information, etc., to improve recognition accuracy, just like multi-party confirmation, which is more accurate.

[0031] Post-processing module: optimizes the recognition results to improve accuracy and reliability, just like checking the recognition results to ensure that there are no errors.

[0032] The technical effects of the present application are as follows: the computer vision-based target detection system and method, through the improved CNN architecture and the optimized YOLO algorithm, can detect targets more accurately, especially perform well in small targets and complex scenes; the improved YOLO algorithm significantly improves the detection speed while maintaining high precision, and is suitable for scenes such as real-time monitoring; the target detection system of the present invention can adapt to detection scenes of different complexities by dynamically adjusting the confidence threshold, has a wide range of applicability, significantly improves the accuracy and real-time performance of target detection; improves the robustness to small targets and complex scenes; and reduces the development cost and maintenance difficulty of the algorithm.

[0033] Specifically, the data acquisition module is used to collect images or video streams to be detected, and the data acquisition module includes a data source interface module, a data cache module and a data transmission module; The data source interface module is responsible for connecting to different data sources and reading data; For example: Camera interface: supports connecting USB cameras, webcams, etc. to obtain video stream data in real time; Video file interface: supports reading locally stored video files, such as MP4, AVI and other formats; Image file interface: supports reading locally stored image files, such as JPEG, PNG and other formats; Network streaming media interface: supports obtaining real-time video stream data from network streaming media servers, such as RTSP, RTMP and other protocols.

[0034] The data cache module is responsible for caching the preprocessed data to improve data reading efficiency; For example: Memory cache: caches data in memory to increase data reading speed; Disk cache: caches data on disk to cope with large data volumes; The data transmission module is responsible for transmitting the cached data to the data preprocessing module; For example: shared memory transmission: efficient data transmission between modules is achieved through shared memory; message queue transmission: asynchronous data transmission between modules is achieved through message queue.

[0035] The data preprocessing module is responsible for preprocessing the collected raw data; For example: Image / video decoding: decode compressed image / video data into raw pixel data; Image / video format conversion: convert image / video data into a system-specified format, such as RGB, YUV, etc.; Image / video resizing: adjust image / video data to a system-specified resolution; Image / video enhancement: perform enhancement processing on image / video data, such as denoising, sharpening, contrast adjustment, etc., to improve the accuracy of target detection.

[0036] Specifically, the improved convolutional neural network architecture in the feature extraction module includes: Feature highlighting module: used to highlight the features of the target area, for the detection of small targets and review scenes; Multi-scale feature fusion module: used to fuse feature maps of different scales to better capture the global and local information of the target.

[0037] The salient feature module calculates the weight distribution of the feature map by the following formula: ; Where: W att represents the attention weight, σ represents the activation function, W q and W k is a learnable weight matrix and X is the input feature map.

[0038] The multi-scale feature fusion module realizes the fusion of feature maps of different scales through the following formula: ; Among them: F msf Represents the fused feature map, F s1 and F s2 They represent feature maps of different scales respectively, and α is the weight coefficient, which is used to balance the contribution of feature maps of different scales.

[0039] Specifically, the target detection module includes: Region Proposal Network (RPN): used to generate candidate target regions; Classifier: Target detection is performed on the feature map based on the improved YOLO algorithm to classify the candidate target area and determine which target category it belongs to; Bounding box regressor: used to perform bounding box regression on the candidate target area and accurately locate the position of the target.

[0040] Specifically, the improved YOLO algorithm introduces a category balancing mechanism, which is implemented by the following formula: Among them, L class Represents the loss of class probability, Lce represents the cross entropy loss, w i is the class weight, which is used to balance the loss contribution of different classes, and C is the total number of classes.

[0041] Specifically, the multimodal data fusion module adopts a weighted fusion strategy to dynamically adjust the weights of different modal data according to the characteristics of the target. The multimodal data fusion module includes: Data alignment module: responsible for the spatial and temporal alignment of data of different modalities; For example: time synchronization: ensure that data of different modes are collected at the same time point; spatial alignment: map data of different modes to the same coordinate system.

[0042] Feature extraction submodule: responsible for extracting features from the data of each modality; For example: RGB image feature extraction: use convolutional neural network (CNN) to extract color and texture features; infrared image feature extraction: use CNN to extract thermal radiation features; depth image feature extraction: use CNN to extract depth features; lidar data feature extraction: use PointNet or PointNet++ and other networks to extract 3D point cloud features.

[0043] Feature fusion module: responsible for secondary fusion of features of different modalities; For example: Early Fusion: fusing data of different modalities before feature extraction, such as splicing RGB images and depth images together, and then inputting them into CNN to extract features; Middle Fusion: fusing features of different modalities during feature extraction, such as using the Cross-Attention mechanism to interact RGB image features and depth image features; Late Fusion: fusing features of different modalities after feature extraction, such as splicing RGB image features, depth image features and LiDAR features together, and then inputting them into the classifier for classification.

[0044] Specifically, the post-processing module includes the following modules: Non-maximum suppression module: used to remove redundant detection boxes and retain the detection results with the highest confidence; Bounding box regression module: used to optimize the position and size of the target bounding box; The post-processing module increases the accuracy and robustness of the detection results by dynamically adjusting the confidence threshold. The dynamic confidence threshold adjustment mechanism is as follows: When the complexity of the detection scene increases, false detections are reduced by increasing the confidence threshold; When the complexity of the detection scene decreases, the recall rate of detection is improved by lowering the confidence threshold.

[0045] The mathematical expression of the dynamic confidence threshold adjustment mechanism is as follows: in, T thres represents the adjusted confidence threshold, T base represents the basic threshold, β is the adjustment factor, N obj Indicates the number of detected targets. N total Indicates the total number of targets.

[0046] Specifically, the output module outputs the detection results in JSON format, including target category, confidence, and bounding box coordinate information.

[0047] The accuracy of this application is improved: by improving the convolutional neural network architecture, highlighting the characteristics of the target area, integrating features of different scales, and adjusting the mechanism based on dynamic confidence thresholds, the target recognition precision and accuracy are improved.

[0048] Improved real-time performance: Optimize algorithm performance, reduce computational complexity, shorten recognition time, and improve real-time performance of target detection.

[0049] Improved robustness: Optimized strategies for small target detection and multimodal data fusion to enhance the ability to resist interference in complex backgrounds and occlusion situations.

[0050] A computer vision-based target detection method, characterized in that it comprises the following steps: Step S1: collecting the image or video stream to be detected through the data acquisition module; Step S2: using a preprocessing module to perform preprocessing operations such as normalization, cropping, and scaling on the collected images; Step S3: extracting feature information from the image through a feature extraction module, wherein the feature extraction module adopts an improved convolutional neural network architecture; Step S4: Detect the target on the feature map using a target detection module, and output the category and location information of the target. The target detection module is based on an improved YOLO algorithm, optimizes the loss function of target detection, and introduces a category balance mechanism. Step S5: perform multimodal data fusion, combine RGB image data with other modal data, and perform feature fusion through weighted fusion strategy; Step S6: Optimize the detection results through the post-processing module, perform non-maximum suppression, confidence threshold screening and result visualization on the detection results.

[0051] Step S7: Output the detection results, including target category and location information.

[0052] An improved convolutional neural network architecture highlights the characteristics of the target area and integrates features of different scales; a multimodal data fusion strategy combines data of different modalities to improve recognition accuracy and robustness; a dynamic confidence threshold adjustment mechanism dynamically adjusts the threshold according to the complexity of the scene to improve the detection effect; the optimization of the data preprocessing module and the post-processing module improves data processing efficiency and result accuracy. This application relates to a target detection system and method based on computer vision, which aims to improve the accuracy and real-time performance of target detection, while enhancing the adaptability to small targets and complex scenes; the system is widely used in security monitoring, autonomous driving, industrial inspection, robot vision and other fields.

[0053] In summary, the target detection system and method based on computer vision can detect targets more accurately through the improved CNN architecture and the optimized YOLO algorithm, especially for small targets and complex scenes; the improved YOLO algorithm significantly improves the detection speed while maintaining high precision, and is suitable for scenes such as real-time monitoring; the target detection system of the present invention can adapt to detection scenes of different complexities by dynamically adjusting the confidence threshold, and has wide applicability.

[0054] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in conjunction with such instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in conjunction with such instruction execution systems, devices or apparatuses.

[0055] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk case (magnetic devices), a random access memory (RAM), a read-only memory (ROM), an erasable and programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be a paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering or, if necessary, processing in another suitable manner, and then stored in a computer memory.

[0056] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit with a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit with a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.

[0057] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.

Claims

1. A computer vision-based target detection system, characterized in that: include: A data acquisition module, used to acquire image or video data; Data preprocessing module, used to crop, scale, normalize and enhance the collected data; Feature extraction module, based on an improved convolutional neural network architecture, is used to extract feature information from images; The target detection module combines the region proposal network and the classifier to generate candidate target regions and classify and regress the candidate regions into bounding boxes to achieve accurate positioning and recognition of targets. Multimodal data fusion module, used to fuse RGB image data with other modal data to improve the accuracy and robustness of target detection; The post-processing module is used to perform non-maximum suppression, confidence threshold screening and result visualization on the detection results; The output module is used to output the detection results, including target category and location information.

2. The computer vision-based target detection system according to claim 1, characterized in that: The data acquisition module is used to collect images or video streams to be detected, and the data acquisition module includes a data source interface module, a data cache module and a data transmission module; The data source interface module is responsible for connecting to different data sources and reading data; The data cache module is responsible for caching the preprocessed data to improve data reading efficiency; The data transmission module is responsible for transmitting the cached data to the data preprocessing module.

3. The computer vision-based target detection system according to claim 1, characterized in that: The improved convolutional neural network architecture in the feature extraction module includes: Feature highlighting module: used to highlight the features of the target area, for the detection of small targets and review scenes; Multi-scale feature fusion module: used to fuse feature maps of different scales to better capture the global and local information of the target.

4. The computer vision-based target detection system according to claim 3, characterized in that: The salient feature module calculates the weight distribution of the feature map by the following formula: ; Where: W att represents the attention weight, σ represents the activation function, W q and W k is a learnable weight matrix, X is the input feature map; The multi-scale feature fusion module realizes the fusion of feature maps of different scales through the following formula: ; Among them: F msf Represents the fused feature map, F s1 and F s2 They represent feature maps of different scales respectively, and α is the weight coefficient, which is used to balance the contribution of feature maps of different scales.

5. The computer vision-based target detection system according to claim 1, characterized in that: The target detection module comprises: Region Proposal Network: used to generate candidate target regions; Classifier: Target detection is performed on the feature map based on the improved YOLO algorithm to classify the candidate target area and determine which target category it belongs to; Bounding box regressor: used to perform bounding box regression on the candidate target area and accurately locate the position of the target.

6. The computer vision-based target detection system according to claim 5, characterized in that: The improved YOLO algorithm introduces a category balancing mechanism, which is implemented by the following formula: Among them, L class Represents the loss of class probability, L ce represents the cross entropy loss, w i is the class weight, which is used to balance the loss contribution of different classes, and C is the total number of classes.

7. The computer vision-based target detection system according to claim 1, characterized in that: The multimodal data fusion module adopts a weighted fusion strategy to dynamically adjust the weights of different modal data according to the characteristics of the target. The multimodal data fusion module includes: Data alignment module: responsible for the spatial and temporal alignment of data of different modalities; Feature extraction submodule: responsible for extracting features from the data of each modality; Feature fusion module: responsible for secondary fusion of features of different modalities.

8. The computer vision-based target detection system according to claim 1, characterized in that: The post-processing module includes the following modules: Non-maximum suppression module: used to remove redundant detection boxes and retain the detection results with the highest confidence; Bounding box regression module: used to optimize the position and size of the target bounding box; The post-processing module increases the accuracy and robustness of the detection results by dynamically adjusting the confidence threshold. The dynamic confidence threshold adjustment mechanism is as follows: When the complexity of the detection scene increases, false detections are reduced by increasing the confidence threshold; When the complexity of the detection scene decreases, the recall rate of detection is improved by lowering the confidence threshold.

9. The computer vision-based target detection system according to claim 8, characterized in that: The mathematical expression of the dynamic confidence threshold adjustment mechanism is as follows: in, T thres represents the adjusted confidence threshold, T base represents the basic threshold, β is the adjustment factor, N obj Indicates the number of detected targets. N total Indicates the total number of targets.

10. A computer vision-based target detection method according to any one of claims 1 to 9, characterized in that: The following steps are involved: Step S1: collecting the image or video stream to be detected through the data acquisition module; Step S2: using a preprocessing module to perform preprocessing operations such as normalization, cropping, and scaling on the collected images; Step S3: extracting feature information from the image through a feature extraction module, wherein the feature extraction module adopts an improved convolutional neural network architecture; Step S4: Detect the target on the feature map using a target detection module, and output the category and location information of the target. The target detection module is based on an improved YOLO algorithm, optimizes the loss function of target detection, and introduces a category balance mechanism. Step S5: perform multimodal data fusion, combine RGB image data with other modal data, and perform feature fusion through weighted fusion strategy; Step S6: Optimize the detection results through the post-processing module, perform non-maximum suppression, confidence threshold screening and result visualization on the detection results. Step S7: Output the detection results, including target category and location information.

Citation Information

Cited By

  • Industrial flexible object detection method based on machine vision

    CN120876399A

  • An industrial flexible object detection method based on machine vision

    CN120876399B

  • Battery car detection method and system based on edge calculation and elevator

    CN121392730A