Detection method and device based on attention scene, equipment and storage medium
By building a hierarchical network architecture, first detecting the attention scene in the visible light image, and then performing object detection on the attention scene, solving the problem that the existing technology cannot effectively detect undisclosed hot spots, and achieving efficient detection efficiency and accuracy.
Patent Information
- Application Number
- CN202411882215.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-19
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-12-19
AI Technical Summary
The prior art cannot effectively detect undisclosed hot spots in visible light images without prior knowledge, resulting in waste of resources and low detection efficiency.
A method based on attention scene detection is proposed. By constructing a hierarchical network architecture, the attention scene is first detected, and then object detection is performed on the attention scene. The method includes obtaining the public visible light image dataset, constructing a scene blind inspection model, performing preliminary blind inspection, determining all scenes contained in the image, and scheduling a single-class object detection model for object detection according to the scene type and priority.
By first detecting the scene, the invalid calculation of the entire image is reduced, the detection efficiency is improved, and object detection is carried out for specific scenes, which improves the detection accuracy. This method can flexibly schedule object detection models according to different needs and adapt to a variety of application scenarios.
Smart Images

Figure CN120047660A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of digital image processing technology, and particularly to a method, device, equipment and storage medium for detecting a scene of interest. Background Art
[0002] Classic visible light image object detection usually refers to closed-set fixed-class detection, which is a detection model centered on a single class of objects and maintains state-of-the-art detection performance. However, in visible light data with coexisting open-world multi-class scenes, the single-class object detection model can only detect one class of object regions, and cannot monitor the states of other class scenes at the same time. Other regions of interest are ignored or deleted in this task. This causes a waste of various resources in a large number of visible light images and greatly reduces the actual role that visible light data should play.
[0003] In summary, how to detect many unpublicized and prior-knowledge-free hot regions in visible light images is an urgent problem to be solved at present. Summary of the Invention
[0004] This application aims to solve at least one of the technical problems in the related art to some extent.
[0005] To this end, the first object of this application is to propose a method for detecting a scene of interest to solve problems such as the inability of existing technical means to detect many unpublicized and prior-knowledge-free hot regions in visible light images.
[0006] The second object of this application is to propose a device.
[0007] The third object of this application is to propose an electronic device.
[0008] The fourth object of this application is to propose a computer-readable storage medium.
[0009] The fifth object of this application is to propose a computer program product.
[0010] To achieve the above object, the first aspect embodiment of this application proposes a method for detecting a scene of interest, including:
[0011] Obtain a publicly available visible light image dataset;
[0012] Based on a multi-scene supervised learning detection neural network model, construct a scene blind detection model, and use the publicly available visible light image dataset to train the scene blind detection model to obtain a trained scene blind detection model;
[0013] Use the trained scene blind detection model to perform a preliminary blind detection on the visible light image to be detected, and obtain all scenes included in the visible light image to be detected;
[0014] Call the corresponding single-class object detection model for all scenes included in the visible light image to be detected, and obtain the detection result.
[0015] Preferably, the obtaining of the public visible light image dataset includes:
[0016] Obtain high-resolution visible light images with different resolutions and public visible light images with large characteristic differences, including various complex scenes and visible light data samples with serious image texture degradation caused by different sources, different resolutions, and different imaging conditions.
[0017] Preferably, the constructing of the scene blind detection model based on the multi-scene supervised learning detection neural network model, and training the scene blind detection model with the public visible light image dataset to obtain the trained scene blind detection model includes:
[0018] Construct a scene blind detection model using a deep convolutional neural network, and the scene blind detection model is used to extract feature information from the input image;
[0019] Use the open visible light image dataset, label all scenes of interest by means of supervised learning, and train the scene blind detection model to obtain the trained scene blind detection model.
[0020] Preferably, the calling of the corresponding single-class object detection model for all scenes included in the visible light image to be detected and obtaining the detection result includes:
[0021] Determine the scenes included in the image according to the detection result of the scene blind detection model;
[0022] For the detected scenes, schedule the corresponding single-class object detection model for object detection based on the single-class detection model scheduling mechanism, and the scheduling can be carried out according to the priority of the scenes.
[0023] Preferably, the single-class detection model scheduling mechanism includes:
[0024] Scheduling based on scene types, calling different single-class detection models for different scene types;
[0025] Scheduling based on priorities, preset the scene priorities, and schedule the object detection model to detect the scene with the highest priority;
[0026] Scheduling based on the importance of the targets, sort the importance of the scenes, and give priority to detecting the scene with the highest importance.
[0027] Preferably, the single-class object detection model adopts the YOLO network, and an attention mechanism and a multi-scale fusion mechanism are added for model construction.
[0028] To achieve the above object, an embodiment of the second aspect of the present application provides a device for detecting a scene of interest, including:
[0029] A data acquisition module, which acquires a public visible light image dataset;
[0030] A training module, which constructs a scene blind detection model based on a multi-scene supervised learning detection neural network model, and uses the public visible light image dataset to train the scene blind detection model to obtain a trained scene blind detection model;
[0031] A blind detection module, which uses the trained scene blind detection model to perform a preliminary blind detection on a visible light image to be detected, and obtains all scenes included in the visible light image to be detected;
[0032] A detection model calling module, which calls a corresponding single-class object detection model based on all scenes included in the visible light image to be detected to obtain a detection result.
[0033] To achieve the above object, an embodiment of the third aspect of the present application provides an electronic device, including: a processor and a memory communicatively connected to the processor;
[0034] The memory stores computer execution instructions;
[0035] The processor executes the computer execution instructions stored in the memory to implement the method described in any one of the above.
[0036] To achieve the above object, an embodiment of the fourth aspect of the present application provides a computer-readable storage medium, including computer execution instructions stored in the computer-readable storage medium, and the computer execution instructions are used to implement the method described in any one of the above when executed by a processor.
[0037] To achieve the above object, an embodiment of the fifth aspect of the present application provides a computer program product, including computer instructions, and the computer instructions are used to cause a computer to execute the method described in the first aspect or any corresponding embodiment thereof.
[0038] A method for detecting a scene of interest provided by the present application constructs a hierarchical network architecture. After detecting the scene of interest, object detection is performed on the scene of interest. First, the first layer designs a scene blind detection model to simultaneously detect scenes for all categories of interest in visible light data. That is, through supervised learning, a preliminary blind detection is performed on the regions of interest in the open visible light image dataset to determine all scenes of interest included in the dataset. Secondly, the second layer schedules single-class object detection models simultaneously or according to priorities based on the blind detection results. By detecting the scene first, the invalid calculations of the entire image are reduced, improving the detection efficiency. Object detection is performed for specific scenes, improving the accuracy. The object detection models can be flexibly scheduled according to different requirements to adapt to a variety of application scenarios.
[0039] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present application. Brief Description of the Drawings
[0040] The above and / or additional aspects and advantages of the present application will become apparent and be readily understood from the following description of the embodiments in conjunction with the accompanying drawings, where:
[0041] Figure 1 It is a flowchart of the first specific embodiment of a method for detecting a scene of interest provided by the present invention;
[0042] Figure 2 It is a schematic diagram of an arbitrary visible light image scene detection strategy;
[0043] Figure 3 It is a schematic diagram summarizing intelligent object detection algorithms;
[0044] Figure 4 It is a schematic diagram of a scene detection strategy model;
[0045] Figure 5 It is a structural block diagram of a device for detecting a scene of interest provided by an embodiment of the present invention. Detailed Embodiments
[0046] The core of the present invention is to provide a method, device, equipment, and storage medium for detecting a scene of interest. By constructing a hierarchical network architecture, the scene of interest is detected first, and then object detection is performed on the scene of interest, reducing the invalid calculations of the entire image and improving the detection efficiency.
[0047] To enable those skilled in the art to better understand the solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0048] Please refer to Figure 1 , Figure 1 which is a flowchart of the first specific embodiment of a method for detecting a focus scenario provided by the present invention; the specific operation steps are as follows:
[0049] Step S101: Obtain a publicly available visible light image dataset.
[0050] Obtain high-resolution visible light images with different resolutions and publicly available visible light images with large characteristic differences, including various complex scenarios and visible light data samples with serious image texture degradation caused by different sources, different resolutions, and different imaging conditions.
[0051] Step S102: Build a scene blind detection model based on a multi-scenario supervised learning detection neural network model, and use the publicly available visible light image dataset to train the scene blind detection model to obtain a trained scene blind detection model.
[0052] Build a scene blind detection model using a deep convolutional neural network. The scene blind detection model is used to extract feature information from the input image;
[0053] Use the publicly available visible light image dataset, label all scenes of interest using the method of supervised learning, and train the scene blind detection model to obtain a trained scene blind detection model.
[0054] Step S103: Use the trained scene blind detection model to perform a preliminary blind detection on the visible light image to be detected, and obtain all the scenes included in the visible light image to be detected.
[0055] Step S104: Call the corresponding single-class object detection model based on all the scenes included in the visible light image to be detected to obtain the detection result.
[0056] According to the detection result of the scene blind detection model, determine the scenes included in the image;
[0057] For the detected scenes, schedule the single-class object detection model corresponding to the scene based on the single-class detection model scheduling mechanism for object detection, and the scheduling can be performed according to the priority of the scene.
[0058] Among them, the single-class detection model scheduling mechanism includes:
[0059] Scene-type-based scheduling, where different single-class detection models are called for different scene types;
[0060] Priority-based scheduling, where scene priorities are preset, and the scheduling object detection model detects the scene with the highest priority;
[0061] Importance-of-object-based scheduling, where the scene importance is sorted, and the scene with the highest importance is preferentially detected.
[0062] The single-class object detection model uses the YOLO network and incorporates an attention mechanism and a multi-scale fusion mechanism for model construction.
[0063] This embodiment provides a method for detecting a scene of interest, constructing a hierarchical network architecture. After detecting the scene of interest, object detection is then performed on the scene of interest. First, in the first layer, a scene blind detection model is designed to simultaneously detect scenes for all categories of interest in visible light data. That is, through supervised learning, a preliminary blind detection is performed on the regions of interest in the open visible light image dataset to determine all scenes of interest contained in the dataset; secondly, in the second layer, according to the blind detection results, single-class object detection models are scheduled simultaneously or according to priorities. By detecting the scene first, the invalid calculations of the entire image are reduced, improving the detection efficiency. Object detection is performed for specific scenes, improving the accuracy. The object detection model can be flexibly scheduled according to different requirements to adapt to a variety of application scenarios.
[0064] Based on the above embodiment, this embodiment describes the method for detecting a scene of interest as Figure 2 shown, specifically as follows:
[0065] Before introducing the embodiments of the present application, the following content is first introduced:
[0066] As Figure 3 shown, deep learning object detection algorithms are mainly divided into two categories: two-stage detection and single-stage detection. Two-stage algorithms first generate candidate boxes as samples and apply image classification algorithms to the candidate regions; while single-stage detection algorithms directly perform regression on the predicted object objects.
[0067] Two-stage object detection algorithm
[0068] The two-stage object detection algorithm has a higher detection accuracy, mainly because the RPN is used to generate accurate candidate boxes, reducing the false detection rate. In large objects and complex scenes, it can provide better detection performance. Representative algorithms include Faster R-CNN, Mask R-CNN, Cascade R-CNN, etc. It requires two stages of calculation, so the speed is slower. It is not suitable for real-time object detection processing.
[0069] Single-stage object detection algorithm
[0070] Single-stage object detection algorithms perform well in different application scenarios and usually have a relatively fast inference speed. The choice of which algorithm depends on the specific requirements of the application, including factors such as accuracy, speed, and resource consumption. In addition, these algorithms usually provide pre-trained models in deep learning frameworks (such as TensorFlow and PyTorch), which can be used for rapid development of custom object detection applications. Online real-time object detection can be achieved, but the detection accuracy is not as high as that of two-stage object detection algorithms.
[0071] Anchor-free detection algorithm
[0072] In both classic single-stage and two-stage object detection, the bounding boxes of objects, i.e., Anchors, need to be preset, which brings additional operations to the detection process. While Anchor-free object detection algorithms do not require presetting Anchors and directly predict the positions and categories of objects in the image. Representatives of such algorithms include YOLOv1, CornerNet, CenterNet, etc. The advantage of Anchor-free algorithms is that they avoid the process of setting and tuning Anchors and simplify the algorithm structure.
[0073] Single-stage object detection can also be called the YOLO network, which defines object detection as a regression problem. Detection research based on it has rapidly gained popularity, and it is also widely used in scene detection such as focusing on scenes. Its greatest advantage is its fast detection speed. Although the detection accuracy rate is lower than that of two-stage object detection models, through other technical means such as attention mechanisms and multi-scale fusion, it has also reached the level of two-stage detection models. Therefore, single-stage object detection models are more valuable in applications.
[0074] As Figure 4 shown, this embodiment includes the following two-layer architecture:
[0075] The first layer: Scene blind detection model
[0076] Model structure: Use a deep convolutional neural network (CNN) to design the scene blind detection model, which can effectively extract image features.
[0077] Training data: Utilize an open visible light image dataset for supervised learning and label all scene categories of interest.
[0078] Blind detection process: Through preliminary scene detection of all categories of interest in visible light data, determine all scene categories of interest included in the dataset.
[0079] Output: The output of this layer is a list of detected scene categories and corresponding confidence scores.
[0080] The second layer: Object detection model
[0081] Policy scheduling: According to the blind detection results of the first layer, preferentially schedule the targeted object detection model (i.e., single-class object detection model).
[0082] Detection execution: On the determined target scenario, perform object detection to identify and locate the objects of interest.
[0083] Priority scheduling: According to the importance or priority of the scenario, schedule different object detection models to achieve the optimal allocation of resources.
[0084] In one embodiment, the policy scheduling includes:
[0085] The scene blind detection model analyzes the input image to identify various scene information in the image. The recognition results are not only the type of the scene, but also the importance or priority of the scene in the application. Different scenes may have different impacts on the detection task. Therefore, scheduling the subsequent detection tasks according to the priority of the scene is a key optimization of the present invention.
[0086] Scene type: The scene can be of various types such as city, nature, indoor, etc. Each type of image will contain different backgrounds and objects.
[0087] Priority setting: The priorities of different scenes may be set based on various factors. For example, in a security monitoring system, outdoor scenes may have a higher priority than indoor scenes because emergencies are more likely to occur in outdoor areas; in medical image processing, specific areas (such as the location of tumors) may require a higher detection priority.
[0088] Based on the importance of the scene, the scene categories output by the first-layer scene detection system will be weighted or sorted by priority according to specific rules. These priorities can be static (such as preset rules) or dynamic (such as adjusted according to real-time data feedback or historical experience).
[0089] Model scheduling mechanism
[0090] The scene detection and scheduling process is a dynamic decision-making process, and its core is to select the appropriate single-class object detection model according to the scene priority. Specifically, the scheduling mechanism can be implemented in the following ways:
[0091] Scheduling based on scene type: When a specific scene (such as an urban traffic scene) is detected, schedule the corresponding traffic monitoring object detection model, which is specifically used to identify targets such as traffic signals, vehicles, and pedestrians. Different scene types may require different single-class detection models.
[0092] Priority-based Scheduling: When a scenario is determined to have a high priority (e.g., a scenario where an emergency occurs), schedule a higher-performance or faster-responsive object detection model for rapid detection. For example, for a sudden scenario in security monitoring, a single-class detection model with low latency can be selected to ensure the early identification of possible abnormal events.
[0093] Scheduling Based on Target Importance: For each scenario, the model can schedule models according to the detection importance of different targets. For example, in natural scenarios, it may be necessary to prioritize the detection of rare animals or dangerous species, while in other scenarios, ordinary targets (such as ordinary pedestrians) can choose a detection model with a lower priority.
[0094] Scheduling Algorithms and Decision Rules
[0095] To make the scheduling mechanism efficient and accurate, a set of scheduling algorithms can be designed to dynamically determine which object detection models to call based on the type and priority of the scenario. The following are several possible scheduling strategies:
[0096] Static Rule Scheduling: Based on the category and priority of the scenario, adopt static rules to select the corresponding single-class detection model. For example, when the scenario type is "urban street" and the priority is high, schedule the "pedestrian detection" model for detection.
[0097] Dynamic Feedback Scheduling: In real-time applications, the scheduling mechanism can also be dynamically adjusted according to historical data and the current state. For example, if a certain scenario has shown frequent changes in historical data (such as traffic conditions changing due to weather changes), a real-time feedback scheduling detection model can be selected.
[0098] Hybrid Scheduling Strategy: Combine static rules and dynamic adjustment to form a hybrid scheduling system. Initially select the detection model according to the scenario type and priority, and adjust the priority of the model or select other alternative models according to real-time feedback (such as the uncertainty or confidence of the detection results) during real-time operation.
[0099] Priority and Performance Optimization
[0100] The scheduling process not only needs to focus on the priority of the scenario but also consider the allocation of computing resources. To ensure that the system can detect high-priority scenarios in a timely manner while not causing the detection of low-priority scenarios to lag, the following performance optimization strategies can be adopted:
[0101] Dynamic Resource Allocation: When detecting multiple scenarios simultaneously, computing resources can be dynamically allocated according to the priority of the scenarios. For example, allocate more GPU resources to high-priority scenarios and reduce the detection resources for low-priority scenarios.
[0102] Multi-threading and Parallel Processing: To improve efficiency, multi-threading or parallel computing technologies can be adopted. In scenarios with different priorities, multi-threading is used for detection, enabling high-priority tasks to be completed in the shortest possible time while ensuring that other tasks can also proceed in parallel.
[0103] For example, in an intelligent transportation monitoring system, the system first uses a scene blind detection model to identify the "urban road" scene in the image, and the system sets it as high priority according to preset rules. Next, the system schedules the "vehicle detection" model to perform object detection on the road scene to quickly identify whether there are illegal parking, traffic congestion, etc. At the same time, if the scene is a "parking lot", another single-class object detection model dedicated to "parking space detection" is scheduled.
[0104] Subsequent Processing and Optimization
[0105] After object detection is completed, the detection results can also be optimized through post-processing. For example, the non-maximum suppression (NMS) method can be used to remove redundant detection boxes, or the detection priority can be further adjusted according to the confidence level of the detection results.
[0106] RSBD consists of an image encoder, a text encoder, and an object decoder. The text encoder processes any descriptions related to the task, including object categories, any form of name, captions about the object, and reference expressions. The prior information acts as a prompt or a tag, that is, the location information of the input scene, the object scene name, etc. Then they are integrated into the detector, and the object scene is extracted from the image according to the text and prior information input.
[0107] For the Image backbone model, we use the ResNET network to achieve multi-scale feature extraction of image objects; the TextEncoder model encodes the text using the seq2seq idea and establishes a correspondence with the objects in the Image backbone; finally, it is fed into YOLOX with a dynamic class head together with the prior object location information to construct an object decoder to detect the object scene.
[0108] This embodiment provides a method for detecting a focus scenario, constructs a hierarchical network architecture, first detects the focus scenario, and then performs object detection on the focus scenario. First, the first layer designs a scene blind detection model to simultaneously detect scenes for all categories of interest in visible light data. That is, through supervised learning, a preliminary blind detection is performed on the regions of interest in the open visible light image dataset to determine all the scenes of interest included in the dataset; secondly, the second layer, according to the blind detection results, simultaneously or according to priorities, specifically schedules single-class object detection models. By detecting the scene first, the ineffective calculations of the entire image are reduced, the detection efficiency is improved, and object detection is performed for a specific scene, improving the accuracy. The object detection model can be flexibly scheduled according to different requirements to adapt to various application scenarios.
[0109] Please refer to Figure 5 , Figure 5 which is a structural block diagram of a device for detecting a focus scenario provided by an embodiment of the present invention; the specific device may include:
[0110] A data acquisition module 100, which acquires a public visible light image dataset;
[0111] A training module 200, which constructs a scene blind detection model based on a neural network model for multi-scene supervised learning, and uses the public visible light image dataset to train the scene blind detection model to obtain a trained scene blind detection model;
[0112] A blind detection module 300, which uses the trained scene blind detection model to perform a preliminary blind detection on the visible light image to be detected, and obtains all the scenes included in the visible light image to be detected;
[0113] A detection model calling module 400, which calls a corresponding single-class object detection model based on all the scenes included in the visible light image to be detected to obtain a detection result.
[0114] The device for detecting a focus scenario in this embodiment is used to implement the foregoing method for detecting a focus scenario. Therefore, the specific implementation manners in the device for detecting a focus scenario can be seen in the embodiment part of the foregoing method for detecting a focus scenario. For example, the data acquisition module 100, the training module 200, the blind detection module 300, and the detection model calling module 400 are respectively used to implement steps S101, S102, S103, and S104 in the foregoing method for detecting a focus scenario. Therefore, the specific implementation manners can refer to the descriptions of the corresponding parts of the embodiments, and will not be repeated here.
[0115] To implement the above embodiments, the present application further provides an electronic device, including: a processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the method provided in the foregoing embodiments.
[0116] To implement the above embodiments, the present application further provides a computer-readable storage medium storing computer-executable instructions, which are used to implement the method provided in the foregoing embodiments when executed by a processor.
[0117] To implement the above embodiments, the present application further provides a computer program product including a computer program, which implements the method provided in the foregoing embodiments when executed by a processor.
[0118] The collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved in the present application all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0119] It should be noted that personal information from users should be collected for legal and reasonable purposes and not shared or sold outside of these legal uses. In addition, such collection / sharing should be carried out after obtaining the user's informed consent, including but not limited to notifying the user to read the user agreement / user notice and sign an agreement / authorization including authorizing relevant user information before the user uses the function. In addition, any necessary steps should be taken to protect and safeguard access to such personal information data and ensure that others with access to personal information data comply with their privacy policies and procedures.
[0120] The present application anticipates providing embodiments that allow users to selectively block the use or access of personal information data. That is, the present disclosure anticipates providing hardware and / or software to prevent or block access to such personal information data. Once personal information data is no longer needed, the risk can be minimized by restricting data collection and deleting the data. In addition, when applicable, personal identifiers are removed from such personal information to protect the privacy of the user.
[0121] In the descriptions of the foregoing embodiments, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0122] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present application, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0123] Any process or method description shown in the flowchart or described in other ways herein may be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a customized logic function or process, and the scope of the preferred embodiments of the present application includes additional implementations, where the functions may be executed in a substantially simultaneous manner or in a reverse order according to the involved functions, rather than in the order shown or discussed, which should be understood by those skilled in the art to which the embodiments of the present application pertain.
[0124] The logic and / or steps represented in the flowchart or otherwise described herein can, for example, be considered as a definitional sequence list of executable instructions for implementing logical functions, which can be embodied specifically in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or used in conjunction with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection part with one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing when necessary, and then stored in a computer memory.
[0125] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGA), field programmable gate arrays (FPGA), etc.
[0126] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0127] In addition, each functional unit in various embodiments of the present application may be integrated into one processing module, may exist separately as individual physical units, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0128] The above-mentioned storage medium may be a read-only memory, a magnetic disk, an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A method for detecting a scene of interest, characterized in that: include: Obtain public visible light image datasets; Building a scene blind detection model based on a multi-scene supervised learning detection neural network model, and training the scene blind detection model using the public visible light image dataset to obtain a trained scene blind detection model; Using the trained scene blind detection model to perform a preliminary blind detection on the visible light image to be detected, and obtaining all scenes included in the visible light image to be detected; Based on all the scenes included in the visible light image to be detected, a corresponding single-class object detection model is called to obtain a detection result.
2. The method for detecting scenes of interest according to claim 1, characterized in that: The obtaining of a public visible light image dataset comprises: Acquire high-resolution visible light images of different resolutions and public visible light images with large differences in characteristics, including visible light data samples of various complex scenes and those with severe image texture degradation caused by different sources, different resolutions, and different imaging conditions.
3. The method for detecting scenes of interest according to claim 1, characterized in that: The scene blind detection model is constructed based on the multi-scene supervised learning detection neural network model, and the scene blind detection model is trained using the public visible light image data set to obtain the trained scene blind detection model, which includes: A scene blind detection model is constructed using a deep convolutional neural network, wherein the scene blind detection model is used to extract feature information from an input image; By using an open visible light image dataset and adopting a supervised learning method, all scenes of interest are labeled, and the scene blind detection model is trained to obtain a trained scene blind detection model.
4. The method for detecting scenes of interest according to claim 1, characterized in that: The method of calling a corresponding single-class object detection model based on all scenes included in the visible light image to be detected and obtaining a detection result includes: Determining the scene contained in the image according to the detection result of the scene blind detection model; For the detected scene, a single-class object detection model corresponding to the scene is scheduled to perform object detection based on a single-class detection model scheduling mechanism. The scheduling can be performed according to the priority of the scene.
5. The method for detecting scenes of interest according to claim 4, characterized in that: The single-class detection model scheduling mechanism includes: Scheduling based on scene types, calling different single-class detection models for different scene types; Priority-based scheduling, preset scene priorities, and schedule the object detection model to detect the scene with the highest priority; Based on the scheduling of target importance, the scene importance is sorted and the scene with the highest importance is detected first.
6. The method for detecting scenes of interest according to claim 1, characterized in that: The single-class object detection model adopts the YOLO network, and adds an attention mechanism and a multi-scale fusion mechanism to build the model.
7. A detection device based on a focus scene, characterized in that: include: A data acquisition module, which acquires public visible light image datasets; A training module, which constructs a scene blind detection model based on a multi-scene supervised learning detection neural network model, and trains the scene blind detection model using the public visible light image data set to obtain a trained scene blind detection model; A blind inspection module, using the trained scene blind inspection model to perform a preliminary blind inspection on the visible light image to be detected, to obtain all scenes included in the visible light image to be detected; The detection model calling module calls the corresponding single-class object detection model based on all scenes included in the visible light image to be detected to obtain a detection result.
8. An electronic device, characterized in that: include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 6 when executed by a processor.
10. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Garbage detection method and device based on scene attention and related medium
CN115035474A
Dense scene text detection and recognition method
CN117218641A
Scene understanding-based method for detecting illegally aired articles along street
CN117830931A
Target detection method and device and electronic equipment
CN118781475A
Target detection method and device, equipment, storage medium and product
CN118968038A