Model training method, related device and storage medium
By selecting images with moderate confidence levels as candidate samples from the AI security model and obtaining real label information for training, the problem of insufficient data in the model is solved, the accuracy and generalization ability of the model are improved, and it is suitable for abnormal behavior recognition in complex scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING REALAI TECH CO LTD
- Filing Date
- 2022-12-13
- Publication Date
- 2026-04-28
AI Technical Summary
Existing AI security models suffer from low accuracy and insufficient generalization ability due to a lack of training data, making it difficult to effectively identify abnormal behavior, especially in complex scenarios.
By selecting images with confidence levels between the second and first confidence thresholds from the image acquisition device as candidate sample images, real sample label information is obtained and the model is trained. The allocation of computing resources is dynamically adjusted to optimize the model training process.
It improves the model's accuracy and generalization ability, enabling it to more effectively identify abnormal behaviors in complex scenarios.
Smart Images

Figure CN115984643B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a model training method, related equipment, and storage medium. Background Technology
[0002] Currently, the use of artificial intelligence (AI) for security monitoring is developing rapidly, driven by policies and technologies. In particular, in recent years, the demand for B-end enterprise security and C-end personal security in the civilian security field has been gradually expanding. More and more places are deploying AI security models in their monitoring equipment. For example, they can automatically monitor personnel violations or equipment violations in factories, ports, and industrial parks, and monitor special scenarios such as fires, prohibited entry, and regulatory violations.
[0003] In security scenarios, algorithms primarily focus on target detection and classification. However, real-world monitoring data is diverse, with training samples typically covering limited scenarios. Furthermore, much of this data is confidential, making it difficult for developers to obtain comprehensive data during the initial algorithm development phase. Additionally, much of this scenario data is often stored on internal networks, making it difficult to export. This lack of training data results in low accuracy for AI security models. Summary of the Invention
[0004] This application provides a model training method, related equipment, and storage medium, which can improve the accuracy of the model and enhance its generalization ability online.
[0005] In a first aspect, embodiments of this application provide a model training method, which includes:
[0006] Acquire multiple images to be identified that have been historically acquired by the image acquisition device in a preset location;
[0007] Each of the images to be identified is processed according to a preset security identification model to obtain the candidate confidence level corresponding to each image to be identified.
[0008] The image to be identified corresponding to the target confidence level is determined as a candidate sample image. The target confidence level is the confidence level among the candidate confidence levels that is greater than the second confidence level threshold and less than the first confidence level threshold. The first confidence level threshold is the confidence level threshold currently set by the security recognition model. The second confidence level threshold is less than the first confidence level threshold.
[0009] Obtain the real sample label information corresponding to each of the candidate sample images to obtain multiple target sample images, wherein each target sample image carries the corresponding real sample label information;
[0010] The security recognition model is trained based on multiple target sample images to obtain the trained security recognition model.
[0011] Secondly, embodiments of this application also provide a model training apparatus, which includes:
[0012] The transceiver module is used to acquire multiple images to be identified that have been historically acquired by the image acquisition device in a preset location;
[0013] The processing module is used to perform recognition processing on each of the images to be recognized according to a preset security recognition model, and obtain the candidate confidence scores corresponding to each of the images to be recognized; and to determine the image to be recognized corresponding to the target confidence score as a candidate sample image, wherein the target confidence score is the confidence score among the candidate confidence scores that is greater than a second confidence score threshold and less than a first confidence score threshold, wherein the first confidence score threshold is the confidence score threshold currently set by the security recognition model, and the second confidence score threshold is less than the first confidence score threshold;
[0014] The transceiver module is also used to obtain the real sample label information corresponding to each of the candidate sample images to obtain multiple target sample images, wherein the target sample images carry the corresponding real sample label information;
[0015] The processing module is further configured to train the security recognition model based on multiple target sample images to obtain the trained security recognition model.
[0016] In some embodiments, before the transceiver module performs the step of acquiring multiple images to be identified historically acquired by the image acquisition device in a preset location, the processing module is further configured to:
[0017] The target frame extraction frequency is determined based on the current scene of the image acquisition device and the preset frame extraction strategy.
[0018] According to the preset computing power allocation rules and the target frame extraction frequency, the preset computing power resources are divided into a first computing power resource and a second computing power resource. The first computing power resource is used to perform recognition processing on the images to be recognized, and the second computing power resource is used to train the security recognition model.
[0019] At this time, when the transceiver module performs the step of acquiring multiple images to be identified historically acquired by the image acquisition device in a preset location, it is specifically used for:
[0020] Based on the target frame rate, the image acquisition device acquires multiple images to be identified that have been historically acquired in the preset location.
[0021] In some embodiments, when the processing module performs the step of training the security recognition model based on multiple target sample images to obtain the trained security recognition model, it is specifically used for:
[0022] The multiple target sample images are divided into a training sample set and a validation sample set according to a preset sample allocation ratio.
[0023] The security identification model is trained with model parameters based on the training sample set to obtain an intermediate security identification model.
[0024] Based on the intermediate security identification model, the first confidence threshold is adjusted according to the verification sample set to obtain the trained security identification model.
[0025] In some embodiments, when the processing module performs the step of adjusting the first confidence threshold based on the intermediate security identification model and the verification sample set, it is specifically used for:
[0026] Based on the intermediate security identification model, the recall rate and false alarm rate corresponding to each of the candidate confidence thresholds are determined according to the verification sample set.
[0027] Based on the recall rate and the false positive rate, a target confidence threshold is selected from multiple candidate confidence thresholds that meet preset criteria, and the first confidence threshold is replaced with the target confidence threshold. The preset criteria are that the recall rate is higher than a preset recall rate threshold and the false positive rate is lower than a preset false positive rate threshold.
[0028] In some embodiments, before executing the step of determining the recall rate and false positive rate corresponding to each of the multiple candidate confidence thresholds based on the intermediate security identification model and the verification sample set, the processing module is further configured to:
[0029] The transceiver module receives the candidate confidence threshold input by the user.
[0030] In some embodiments, before the transceiver module performs the step of acquiring the real sample label information corresponding to each of the candidate sample images to obtain multiple target sample images, the processing module is further configured to:
[0031] The candidate sample images are clustered to obtain multiple cluster groups;
[0032] According to the preset filtering rules, the candidate sample images in each cluster group are filtered to obtain the filtered candidate sample images.
[0033] When the transceiver module performs the step of obtaining the real sample label information corresponding to each of the candidate sample images to obtain multiple target sample images, it is specifically used for:
[0034] Obtain the real sample label information corresponding to each of the filtered candidate sample images to obtain multiple target sample images.
[0035] In some embodiments, after performing the step of recognizing each of the images to be recognized according to a preset security recognition model to obtain the candidate confidence scores corresponding to each of the images to be recognized, the processing module is further configured to:
[0036] Images to be identified with a candidate confidence level greater than or equal to the first confidence level are determined as abnormal images;
[0037] Output the abnormal image.
[0038] In some embodiments, when the transceiver module performs the step of obtaining the real sample label information corresponding to each of the candidate sample images to obtain multiple target sample images, it is specifically used for:
[0039] The processing module displays each candidate sample image in the display interface;
[0040] The display interface receives real sample label information sent by the user for each of the candidate sample images.
[0041] The processing module adds corresponding real sample label information to each candidate sample image to obtain multiple target sample images.
[0042] Thirdly, embodiments of this application also provide a computer device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the above-described method.
[0043] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program, the computer program including program instructions that, when executed by a processor, can implement the above-described method.
[0044] Compared to existing technologies, the solution provided in this application has two advantages. First, it can select monitoring images (images to be identified) with a confidence level greater than a second confidence threshold and less than a first confidence threshold (the confidence threshold currently set by the model) from a real monitoring scene (preset location) of the image acquisition device as candidate sample images. Since the confidence level of the candidate sample images is close to the first confidence threshold, the obtained candidate sample images are very likely to include abnormal images that have not been filtered out by the first confidence threshold. Therefore, when used to train the model, real sample label information (such as abnormal labels or normal labels) can be added to these candidate sample images that may include abnormal images that have not been filtered out by the first confidence threshold to obtain target sample images. The target sample images can then be used to further train the model. It can be seen that this application can obtain training samples from real monitoring scenes and further train the model based on the obtained training samples, thereby improving the accuracy of the model. Second, the target sample images in this embodiment are obtained during the model's operation and can be used to train the model during the model's operation. It can be seen that this solution can improve the model's generalization ability online. Attached Figure Description
[0045] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0046] Figure 1 This is a schematic diagram illustrating an application scenario of the model training method provided in the embodiments of this application;
[0047] Figure 2 A schematic flowchart illustrating the model training method provided in this application embodiment;
[0048] Figure 3 A schematic block diagram of the model training apparatus provided in the embodiments of this application;
[0049] Figure 4 This is a schematic diagram of a server structure in one embodiment of this application;
[0050] Figure 5 This is a schematic diagram of the structure of a terminal in an embodiment of this application;
[0051] Figure 6 This is a schematic diagram of a server structure in one embodiment of this application. Detailed Implementation
[0052] The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments described herein can be implemented in a sequence other than that illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or modules is not necessarily limited to those explicitly listed, but may include other steps or modules not explicitly listed or inherent to these processes, methods, products, or devices. The division of modules in the embodiments of this application is merely a logical division; in actual applications, there may be other division methods. For example, multiple modules may be combined into or integrated into another system, or some features may be ignored or not performed. Additionally, the shown or discussed mutual coupling or direct coupling or communication connection may be through some interface, and the indirect coupling or communication connection between modules may be electrical or other similar forms, none of which are limited in the embodiments of this application. Furthermore, the modules or sub-modules described as separate components may or may not be physically separated, may or may not be physical modules, or may be distributed among multiple circuit modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the embodiments of this application.
[0053] This application provides a model training method, apparatus, and storage medium. The execution subject of the model training method can be the model training apparatus provided in this application, or a computer device integrating the model training apparatus, or a model training system including a model training apparatus and an image acquisition device. The model training apparatus can be implemented in hardware or software, and the computer device can be a terminal or a server.
[0054] When the computer device is a server, the server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.
[0055] When the computer device is a terminal, the terminal may include, but is not limited to, smart terminals with multimedia data processing functions (e.g., video data playback function, music data playback function), such as smartphones, tablets, laptops, desktop computers, smart TVs, smart speakers, personal digital assistants (PDAs), desktop computers, and smartwatches.
[0056] The solutions in this application can be implemented based on artificial intelligence technology, specifically involving computer vision technology in artificial intelligence technology and cloud computing, cloud storage and database in cloud technology, which will be described separately below.
[0057] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.
[0058] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0059] Computer vision (CV) is a science that studies how to enable machines to "see." More specifically, it refers to machine vision, which uses cameras and computers to replace human eyes in tasks such as target recognition, tracking, and measurement, and further performs image processing to create images more suitable for human observation or transmission to instruments. As a scientific discipline, computer vision studies related theories and technologies, attempting to build artificial intelligence systems capable of extracting information from images or multidimensional data. Computer vision technologies typically include image processing, face recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping (SLAM), and common biometric recognition technologies such as face recognition and fingerprint recognition.
[0060] With the research and advancement of artificial intelligence (AI) technology, AI is being studied and applied in various fields, such as smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, autonomous driving, drones, robots, smart healthcare, and smart customer service. It is believed that with the development of technology, AI will be applied in more fields and play an increasingly important role.
[0061] The solutions in this application can be implemented based on cloud technology, specifically involving cloud computing, cloud storage, and database technologies, which will be described below.
[0062] Cloud technology refers to a hosting technology that unifies hardware, software, and network resources within a wide area network (WAN) or local area network (LAN) to achieve data computation, storage, processing, and sharing. Cloud technology is a collective term for network technologies, information technologies, integration technologies, management platform technologies, and application technologies applied to cloud computing business models. It can form resource pools, providing flexible and convenient on-demand access. Cloud computing technology will become a crucial support. Backend services of technical network systems require substantial computing and storage resources, such as video websites, image websites, and many portal websites. With the rapid development and application of the internet industry, every item may have its own identification mark in the future, requiring transmission to a backend system for logical processing. Data at different levels will be processed separately, and various industry data will require robust system support, which can only be achieved through cloud computing. The embodiments of this application can use cloud technology to save identification results.
[0063] Cloud storage is a new concept that extends and develops from cloud computing. A distributed cloud storage system (hereinafter referred to as a storage system) refers to a storage system that uses cluster applications, grid technology, and distributed storage file systems to aggregate a large number of storage devices (also called storage nodes) of various types in a network through application software or application interfaces to work together and provide data storage and business access functions. In the embodiments of this application, network configuration and other information can be stored in this storage system for easy retrieval by the server.
[0064] Currently, the storage method of storage systems is as follows: Logical volumes are created. During the creation of a logical volume, physical storage space is allocated to each logical volume. This physical storage space may consist of a single storage device or the disks of several storage devices. Clients store data on a logical volume, which means storing the data on the file system. The file system divides the data into many parts, each part being an object. Each object contains not only the data but also additional information such as a data identifier (ID, ID entity). The file system writes each object to the physical storage space of that logical volume and records the storage location information of each object. Therefore, when a client requests access to data, the file system can allow the client to access the data based on the storage location information of each object.
[0065] The process by which a storage system allocates physical storage space to a logical volume is as follows: the physical storage space is pre-divided into strips according to the capacity estimate of the objects stored in the logical volume (this estimate often has a large margin relative to the actual capacity of the objects to be stored) and the grouping of Redundant Array of Independent Disks (RAID). A logical volume can be understood as a strip, thus allocating physical storage space to the logical volume.
[0066] A database, simply put, can be viewed as an electronic filing cabinet—a place to store electronic files, where users can perform operations such as adding, querying, updating, and deleting data. A "database" is a collection of data stored together in a certain way, capable of being shared by multiple users, with minimal redundancy, and independent of application programs.
[0067] A Database Management System (DBMS) is a computer software system designed to manage databases, generally possessing basic functions such as storage, retrieval, security, and backup. DBMSs can be classified according to the database model they support, such as relational or XML (Extensible Markup Language); or according to the type of computer they support, such as server clusters or mobile phones; or according to the query language used, such as SQL (Structured Query Language) or XQuery; or according to performance priorities, such as maximum scale or maximum operating speed; or other classification methods. Regardless of the classification method used, some DBMSs can cross categories, for example, simultaneously supporting multiple query languages. In this embodiment, the identification results can be stored in the DBMS for easy retrieval by the server.
[0068] It should be noted that the service terminal involved in the embodiments of this application can be a device that provides voice and / or data connectivity to the service terminal, a handheld device with wireless connectivity, or other processing devices connected to a wireless modem. Examples include mobile phones (or "cellular" phones) and computers with mobile terminals, such as portable, pocket-sized, handheld, computer-embedded, or vehicle-mounted mobile devices that exchange voice and / or data with a wireless access network. Examples include Personal Communication Service (PCS) phones, cordless phones, Session Initiation Protocol (SIP) phones, Wireless Local Loop (WLL) stations, and Personal Digital Assistants (PDAs).
[0069] Please see Figure 1 , Figure 1 This is a schematic diagram illustrating an application scenario of the model training method provided in an embodiment of this application. The model training method is applied to… Figure 1 The model training system includes a model training device and an image acquisition device. Each model training device can communicate with one or more image acquisition devices. The image acquisition devices are located in a preset location with high confidentiality, such as a scenario for monitoring personnel and / or equipment violations in a specific location (factory, port, or industrial park). Each image acquisition device is responsible for acquiring images from the same or different preset locations. The model training device and the image acquisition devices can be deployed centrally or separately. This application embodiment does not limit this, but only takes separate deployment as an example.
[0070] The technical solutions of the embodiments of this application will now be described in detail.
[0071] Please refer to Figure 2 The following describes a model training method provided by an embodiment of this application, such as... Figure 2 As shown, the method includes the following steps S110-S150.
[0072] S110, The model training device acquires multiple images to be identified that have been historically acquired by the image acquisition device in a preset location.
[0073] In this embodiment, the image acquisition device is installed in a preset location and is responsible for security monitoring of that location. This preset location can be a highly confidential location, such as a scenario monitoring personnel and / or equipment violations in a specific location (factory, port, or industrial park). In this embodiment, the model training device needs to extract the image to be identified from the image acquisition device and use the extracted image to perform subsequent model training and security identification.
[0074] In some embodiments, in order to make reasonable use of computing resources, it is necessary to dynamically adjust the frequency of the model training device extracting the image to be identified. In this case, before extracting the image to be identified, it is necessary to determine the current frame extraction frequency. In this case, the method further includes: determining the target frame extraction frequency according to the current scene of the image acquisition device and the preset frame extraction strategy. In this case, step S110 includes: acquiring multiple images to be identified that were historically acquired by the image acquisition device in the preset location according to the target frame extraction frequency.
[0075] That is, at this time, it is necessary to first determine the scene situation of the current scene in the preset location of the image acquisition device. For example, if the preset frame extraction strategy sets 12:00 am to 5:00 am as an idle scene (the scene does not change much at night) and other time periods as busy scenes, the idle scene corresponds to the first frame extraction frequency, and the busy scene corresponds to the second frame extraction frequency. The second frame extraction frequency is greater than the first frame extraction frequency. At this time, if the current scene is an idle scene, the first frame extraction frequency is determined as the target frame extraction frequency according to the frame extraction strategy. If the current scene is a busy scene, the second frame extraction frequency is determined as the target frame extraction frequency according to the frame extraction strategy.
[0076] Furthermore, if the duration for which no monitoring target (e.g., a person) exists in the current scene exceeds a preset duration (e.g., 10 minutes), then frame extraction of the current scene image is stopped according to the frame extraction strategy until a monitoring target appears in the current scene. Then, the frame extraction function is activated, and frame extraction is performed according to the frame extraction frequency corresponding to the current scene.
[0077] In this embodiment, since the computing power resources are fixed, in order to rationally allocate computing power resources between image recognition and model training, after determining the current target frame extraction frequency, the method further includes: dividing the preset computing power resources into a first computing power resource and a second computing power resource according to a preset computing power allocation rule and the target frame extraction frequency. The first computing power resource is used to perform recognition processing on the images to be recognized, and the second computing power resource is used to train the security recognition model. The lower the target frame extraction frequency, the smaller the amount of image recognition. In this case, more computing power can be allocated to model training; that is, the lower the target frame extraction frequency, the less the first computing power resource and the more the second computing power resource; the higher the target frame extraction frequency, the larger the first computing power resource and the less the second computing power resource. If the size of the second computing power resource is 0, training of the security recognition model is stopped.
[0078] It should be noted that, in order to avoid wasting computing resources, the preset computing resources are divided into first computing resources and second computing resources according to the preset computing power allocation rules only when a model training instruction is received. At this time, the preset computing resources are divided into first computing resources and second computing resources according to the preset computing power allocation rules and the target frame extraction frequency, including: when a model training instruction is received, the preset computing resources are divided into first computing resources and second computing resources according to the preset computing power allocation rules and the target frame extraction frequency.
[0079] S120. The model training device performs recognition processing on each of the images to be recognized according to the preset security recognition model, and obtains the candidate confidence scores corresponding to each of the images to be recognized.
[0080] In this embodiment, the security recognition model is set in the model training device. The model training device can use the built-in security recognition model to recognize the acquired image to be recognized, and train the built-in model training device according to the recognition result to improve the generalization ability of the model.
[0081] In this embodiment, after obtaining the image to be identified from the image acquisition device, each image to be identified is processed according to the preset security identification model, and the candidate confidence level corresponding to each candidate image is obtained. The candidate confidence level indicates the possibility that the corresponding candidate image has a violation. The higher the confidence level, the higher the possibility of a violation.
[0082] In some embodiments, after obtaining the candidate confidence scores of each image to be identified, images to be identified with candidate confidence scores greater than or equal to a first confidence score are further identified as anomalous images; the anomalous images are then output. The first confidence score is the confidence score corresponding to the anomalous image, and the anomalous image is an image that includes anomalous behavior, which includes at least one behavior (e.g., personnel misconduct, fire at a location, etc.).
[0083] For example, when the first confidence level is 0.5, the candidate images with a confidence level greater than or equal to 0.5 are identified as abnormal images, and the corresponding abnormal images are output. Furthermore, the abnormal images are displayed on a large screen, where the abnormal images include abnormal location markers (e.g., using a box to outline the abnormal location) and the corresponding confidence level, so that monitoring personnel can easily obtain abnormal information in the abnormal images.
[0084] In this embodiment, candidate sample images obtained online are subsequently used for training. This embodiment can train the model during its operation. It can be seen that this solution can improve the generalization ability of the model online.
[0085] S130, The model training device determines the image to be identified corresponding to the target confidence level as the candidate sample image.
[0086] Wherein, the target confidence level is the confidence level among the candidate confidence levels that is greater than the second confidence level threshold and less than the first confidence level threshold. The first confidence level threshold is the confidence level threshold currently set by the security recognition model. The second confidence level threshold is less than the first confidence level threshold. Specifically, both the first confidence level threshold and the second confidence level threshold are confidence level thresholds corresponding to abnormal images.
[0087] In this embodiment, specifically, images to be identified with a confidence level between the second confidence threshold and the first confidence threshold are selected as candidate sample images. For example, when the second confidence threshold is 0.4 and the first confidence threshold is 0.5, images to be identified with a confidence level between 0.4 and 0.5 are selected as candidate sample images.
[0088] It should be noted that in some embodiments, in order to save computing power, the model training device does not need to train the security recognition model in real time. The model training device will only train the security model after receiving user input or a model training instruction that is automatically triggered periodically. In this case, step S130 includes: when the model training instruction is received, the image to be identified corresponding to the target confidence level is determined as a candidate sample image. The image to be identified at this time is an image collected by the model training device within a preset historical period. The length of the historical period can be set according to actual needs, such as data within the previous week. The specific details are not limited here.
[0089] S140, The model training device obtains the real sample label information corresponding to each of the candidate sample images, and obtains multiple target sample images.
[0090] The target sample image carries the corresponding real sample label information, which indicates whether the target sample image is an image collected under abnormal conditions or abnormal behavior scenarios. Therefore, the target sample image carrying the real sample label information can be used to further train the model.
[0091] In some embodiments, step S140 specifically includes: displaying each of the candidate sample images in a display interface; receiving real sample label information sent by the user for each of the candidate sample images through the display interface; adding corresponding real sample label information to each of the candidate sample images to obtain a plurality of target sample images.
[0092] The candidate sample images displayed on the interface include multiple candidate recognition results and the confidence level of each candidate recognition result. If a candidate recognition result corresponds to the real sample label information, the user directly selects that candidate recognition result on the interface as the real recognition result for the candidate image sample image, obtaining a target sample image including the real sample label information. If no candidate recognition result corresponds to the real sample label information, the user inputs the real sample label information for the candidate image on the interface to obtain the target sample image. Obtaining the real sample label information corresponding to each candidate sample image includes: obtaining the real sample label information corresponding to each candidate sample image through a semi-automatic annotation method. Therefore, in this embodiment, the user can determine the real sample label information through the candidate recognition results displayed on the interface, achieving semi-automatic annotation of samples and thus improving the efficiency of sample annotation.
[0093] Specifically, when the model training device is deployed on a server (including a cloud server or a physical server), the server provides a display interface for candidate sample images to the user through a user terminal. The user determines the corresponding real sample label information of the candidate sample image through the display interface, enabling the user terminal to obtain the real sample label information of the candidate sample image. After obtaining the real sample label information, the user terminal sends the real sample label information to the server, enabling the server to obtain the real sample label information of the candidate sample image. In other embodiments, when the model training device is deployed on the user terminal, the candidate sample images are directly displayed in the display interface provided by the user terminal, allowing the user to input or determine the real sample label information of the candidate sample image in the display interface based on the displayed candidate sample image.
[0094] In some embodiments, the number of candidate sample images selected may be large, and there may be many similar images. In order to improve the training efficiency and reduce the number of labeled samples, this embodiment needs to select a small amount of common effective data from a large number of similar samples. At this time, the method further includes: performing clustering processing on the candidate sample images to obtain multiple cluster groups; and performing filtering processing on the candidate sample images in each cluster group according to preset filtering rules to obtain filtered candidate sample images.
[0095] The preset filtering rule is to select target candidate sample images (i.e. candidate sample images that do not need to be filtered) from each cluster group according to a specific ratio or number, and then filter out non-target candidate sample images.
[0096] At this time, the candidate sample image in step S140 is the filtered candidate sample image, that is, obtaining the real sample label information corresponding to each of the candidate sample images to obtain multiple target sample images includes: obtaining the real sample label information corresponding to each of the filtered candidate sample images to obtain multiple target sample images.
[0097] S150. The model training device trains the security recognition model based on multiple target sample images to obtain the trained security recognition model.
[0098] Specifically, step S150 includes steps A, B, and C, as follows:
[0099] Step A: Divide the multiple target sample images into the training sample set and the validation sample set according to the preset sample allocation ratio.
[0100] The ratio of training sample set to validation sample set can be 7:3 or other values. This ratio can be adjusted according to the user's actual needs, and is not limited here.
[0101] Step B: Train the model parameters of the security identification model based on the training sample set to obtain an intermediate security identification model.
[0102] Specifically, in this embodiment, the security identification model can be trained using the gradient descent algorithm based on the training sample set, and an intermediate security identification model with optimized parameters can be output.
[0103] Step C: Based on the intermediate security identification model, adjust the first confidence threshold according to the verification sample set to obtain the trained security identification model.
[0104] Specifically, in this embodiment, based on the intermediate security identification model, the recall rate and false positive rate corresponding to each of the multiple candidate confidence thresholds are determined according to the verification sample set; a target confidence threshold is selected from multiple candidate confidence thresholds that meet preset standards according to the recall rate and the false positive rate, and the first confidence threshold is replaced with the target confidence threshold, wherein the preset standard is that the recall rate is higher than a preset recall rate threshold and the false positive rate is lower than a preset false positive rate threshold.
[0105] In some embodiments, the candidate confidence threshold is a preset confidence threshold in the model training device. In this case, the model training device can arbitrarily select one candidate confidence threshold from multiple candidate confidence thresholds that meet the preset criteria as the target confidence threshold, or select the candidate confidence threshold with the highest recall rate or the lowest false positive rate from multiple candidate confidence thresholds that meet the preset criteria as the target confidence threshold.
[0106] In other embodiments, to improve the flexibility of confidence setting and to personalize the confidence setting according to user needs, the method further includes receiving the candidate confidence thresholds input by the user before determining the recall and false positive rates corresponding to each candidate confidence threshold among multiple candidate confidence thresholds based on the verification sample set. Specifically, an input interface for candidate confidence thresholds is provided to the user, who then inputs the candidate confidence thresholds through the input interface. The model training device then displays the recall and false positive rates corresponding to each candidate confidence threshold through the interface, and the user selects the target confidence threshold from multiple candidate confidence thresholds that meet preset standards according to their needs.
[0107] In some embodiments, when the model training device is deployed on a server (including a cloud server or a physical server), the server provides the user with an input interface for candidate confidence thresholds through a user terminal. After obtaining the candidate confidence thresholds through this input interface, the user terminal sends the candidate confidence thresholds to the server. In other embodiments, the model training device is deployed on a user terminal, whereby the user directly inputs the candidate confidence thresholds through the input interface provided by the user terminal, allowing the user terminal to obtain the user-input candidate confidence thresholds.
[0108] In summary, the solution provided in this application has two main advantages. First, it can filter out monitoring images (images to be identified) with a confidence level greater than a second confidence threshold and less than a first confidence threshold (the confidence threshold currently set by the model) from a real monitoring scene (preset location) of the image acquisition device as candidate sample images. Since the confidence level of the candidate sample images is close to the first confidence threshold, the obtained candidate sample images are very likely to include abnormal images that have not been filtered out by the first confidence threshold. Users can add real sample label information (such as abnormal labels or normal labels) to the candidate sample images to obtain target sample images, and then use the target sample images to further train the model. It can be seen that this application can obtain training samples from real monitoring scenes and further train the model based on the obtained training samples, thereby improving the accuracy of the model. Second, the target sample images in this embodiment are obtained during the model's operation and can be used to train the model during the model's operation. It can be seen that this solution can improve the generalization ability of the model online.
[0109] Figure 3 This is a schematic block diagram of a model training device provided in an embodiment of this application. Figure 3 As shown, corresponding to the above model training method, this application also provides a model training apparatus. This model training apparatus includes a unit for executing the above model training method, and the apparatus can be configured in a server or terminal.
[0110] For example, the device is configured on a user's (who may be a customer using the model training device or a person specifically training the model) computer, which is connected in communication with a surveillance camera (image acquisition device) to acquire the image to be identified captured by the surveillance camera, and to perform image recognition and online model updates based on the image to be identified.
[0111] Specifically, please refer to Figure 3 The model training device 300 includes a transceiver module 301 and a processing module 302, wherein:
[0112] The transceiver module 301 is used to acquire multiple images to be identified that have been historically acquired by the image acquisition device in a preset location;
[0113] The processing module 302 is used to perform recognition processing on each of the images to be recognized according to a preset security recognition model, and obtain the candidate confidence scores corresponding to each of the images to be recognized; and to determine the image to be recognized corresponding to the target confidence score as a candidate sample image, wherein the target confidence score is the confidence score among the candidate confidence scores that is greater than a second confidence score threshold and less than a first confidence score threshold, wherein the first confidence score threshold is the confidence score threshold currently set by the security recognition model, and the second confidence score threshold is less than the first confidence score threshold;
[0114] The transceiver module 301 is further configured to acquire the real sample label information corresponding to each of the candidate sample images to obtain multiple target sample images, wherein the target sample images carry the corresponding real sample label information.
[0115] The processing module 302 is further configured to train the security recognition model based on multiple target sample images to obtain the trained security recognition model.
[0116] In some embodiments, before the transceiver module 301 performs the step of acquiring multiple images to be identified historically acquired by the image acquisition device in a preset location, the processing module 302 is further configured to:
[0117] The target frame extraction frequency is determined based on the current scene of the image acquisition device and the preset frame extraction strategy.
[0118] According to the preset computing power allocation rules and the target frame extraction frequency, the preset computing power resources are divided into a first computing power resource and a second computing power resource. The first computing power resource is used to perform recognition processing on the images to be recognized, and the second computing power resource is used to train the security recognition model.
[0119] At this time, when the transceiver module 301 performs the step of acquiring multiple images to be identified historically acquired by the image acquisition device in a preset location, it is specifically used for:
[0120] Based on the target frame rate, the image acquisition device acquires multiple images to be identified that have been historically acquired in the preset location.
[0121] In some embodiments, when the processing module 302 performs the step of training the security recognition model based on multiple target sample images to obtain the trained security recognition model, it is specifically used for:
[0122] The multiple target sample images are divided into a training sample set and a validation sample set according to a preset sample allocation ratio.
[0123] The security identification model is trained with model parameters based on the training sample set to obtain an intermediate security identification model.
[0124] Based on the intermediate security identification model, the first confidence threshold is adjusted according to the verification sample set to obtain the trained security identification model.
[0125] In some embodiments, when the processing module 302 performs the step of adjusting the first confidence threshold based on the intermediate security identification model and the verification sample set, it is specifically used for:
[0126] Based on the intermediate security identification model, the recall rate and false alarm rate corresponding to each of the candidate confidence thresholds are determined according to the verification sample set.
[0127] Based on the recall rate and the false positive rate, a target confidence threshold is selected from multiple candidate confidence thresholds that meet preset criteria, and the first confidence threshold is replaced with the target confidence threshold. The preset criteria are that the recall rate is higher than a preset recall rate threshold and the false positive rate is lower than a preset false positive rate threshold.
[0128] In some embodiments, before executing the step of determining the recall rate and false positive rate corresponding to each of the multiple candidate confidence thresholds based on the intermediate security identification model and the verification sample set, the processing module 302 is further configured to:
[0129] The transceiver module 301 receives the candidate confidence threshold input by the user.
[0130] In some embodiments, before the transceiver module 301 performs the step of acquiring the real sample label information corresponding to each of the candidate sample images to obtain multiple target sample images, the processing module 302 is further configured to:
[0131] The candidate sample images are clustered to obtain multiple cluster groups;
[0132] According to the preset filtering rules, the candidate sample images in each cluster group are filtered to obtain the filtered candidate sample images.
[0133] When the transceiver module 301 performs the step of obtaining the real sample label information corresponding to each of the candidate sample images to obtain multiple target sample images, it is specifically used for:
[0134] Obtain the real sample label information corresponding to each of the filtered candidate sample images to obtain multiple target sample images.
[0135] In some embodiments, after performing the step of recognizing each of the images to be recognized according to a preset security recognition model to obtain the candidate confidence scores corresponding to each of the images to be recognized, the processing module 302 is further configured to:
[0136] Images to be identified with a candidate confidence level greater than or equal to the first confidence level are determined as abnormal images;
[0137] Output the abnormal image.
[0138] In some embodiments, when the transceiver module 301 performs the step of obtaining the real sample label information corresponding to each of the candidate sample images to obtain multiple target sample images, it is specifically used for:
[0139] The processing module 302 displays each candidate sample image in the display interface;
[0140] The display interface receives real sample label information sent by the user for each of the candidate sample images.
[0141] The processing module 302 adds corresponding real sample label information to each of the candidate sample images to obtain multiple target sample images.
[0142] In this embodiment, on the one hand, the transceiver module 301 can acquire the image to be identified (monitoring image) acquired by the image acquisition device, and the processing module 302 can filter out the images to be identified with a confidence level greater than the second confidence threshold and less than the first confidence threshold as candidate sample images. Since the confidence level of the candidate sample images is close to the first confidence threshold, the obtained candidate sample images are very likely to include abnormal images that have not been filtered out by the first confidence threshold. Therefore, when used to train the model, real sample label information can be added to these candidate sample images that may include abnormal images that have not been filtered out by the first confidence threshold to obtain the target sample image, and the model can be further trained using the target sample image. It can be seen that this application can obtain training samples from real monitoring scenarios and further train the model based on the obtained training samples, which can improve the accuracy of the model. On the other hand, the target sample image in this embodiment is obtained during the model's operation and can be used to train the model during the model's operation. It can be seen that this solution can improve the generalization ability of the model online.
[0143] The image information recognition system in this application embodiment has been described above from the perspective of modular functional entities. The image information recognition device in this application embodiment is described below from the perspective of hardware processing.
[0144] It should be noted that the various embodiments in this application (including) Figure 3 In the embodiments shown, the physical devices corresponding to all transceiver modules can be transceivers, and the physical devices corresponding to all processing modules can be processors. When one of the devices has such Figure 3 In the structure shown, the processor, transceiver, and memory implement the same or similar functions as the transceiver module and the processing module provided in the aforementioned device embodiments corresponding to this device. Figure 4 The memory stores the computer programs that the processor needs to call when executing the above image information recognition method.
[0145] Figure 3 The device shown can have, for example Figure 4 The structure shown, when Figure 3 The device shown has the following characteristics: Figure 4 When the structure shown is used, Figure 4 The processor in the device can perform the same or similar functions as the processing module provided in the aforementioned device embodiments. Figure 4 The transceiver in the device can perform the same or similar functions as the transceiver module provided in the aforementioned device embodiments corresponding to this device. Figure 4 The memory stores the computer programs that the processor needs to call when executing the above-described image information recognition method. (In this application embodiment) Figure 3 In the illustrated embodiment, the physical device corresponding to the transceiver module can be an input / output interface, and the physical device corresponding to the processing module can be a processor.
[0146] This application also provides a terminal, such as... Figure 5 As shown, for ease of explanation, only the parts related to the embodiments of this application are shown. For specific technical details not disclosed, please refer to the method section of the embodiments of this application. The terminal can be any terminal including mobile phones, tablets, personal digital assistants (PDAs), point-of-sale (POS) terminals, in-vehicle computers, etc. Taking a mobile phone as an example:
[0147] Figure 5 This is a block diagram illustrating a portion of the structure of a mobile phone related to the terminal provided in the embodiments of this application. (Reference) Figure 5The mobile phone includes: a radio frequency (RF) circuit 55, a memory 520, an input unit 530, a display unit 540, a sensor 550, an audio circuit 560, a wireless fidelity (Wi-Fi) module 570, a processor 580, and a power supply 590, among other components. Those skilled in the art will understand that... Figure 5 The mobile phone structure shown does not constitute a limitation on the mobile phone and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0148] The following is combined with Figure 5 A detailed introduction to each component of a mobile phone:
[0149] The RF circuit 55 can be used for receiving and transmitting signals during information transmission or calls. Specifically, it receives downlink information from the base station and processes it with the processor 580; additionally, it transmits uplink data to the base station. Typically, the RF circuit 55 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low-noise amplifier (LNA), a duplexer, etc. Furthermore, the RF circuit 55 can also communicate wirelessly with networks and other devices. The aforementioned wireless communications may use any communication standard or protocol, including but not limited to Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Long Term Evolution (LTE), email, and Short Messaging Service (SMS).
[0150] The memory 520 can be used to store software programs and modules. The processor 580 executes various mobile phone functions and data processing by running the software programs and modules stored in the memory 520. The memory 520 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, applications required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the mobile phone (such as audio data, phonebook, etc.). In addition, the memory 520 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0151] The input unit 530 can be used to receive input numerical or character information, and to generate key signal inputs related to user settings and function control of the mobile phone. Specifically, the input unit 530 may include a touch panel 531 and other input devices 532. The touch panel 531, also known as a touch screen, can collect touch operations performed by the user on or near it (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near the touch panel 531), and drive the corresponding connection devices according to a pre-set program. Optionally, the touch panel 531 may include two parts: a touch detection device and a touch controller. The touch detection device detects the user's touch position and the signal generated by the touch operation, and transmits the signal to the touch controller; the touch controller receives touch information from the touch detection device, converts it into touch point coordinates, and sends it to the processor 580, and can also receive and execute commands sent by the processor 580. In addition, the touch panel 531 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch panel 531, the input unit 530 may also include other input devices 532. Specifically, other input devices 532 may include, but are not limited to, one or more of the following: physical keyboard, function keys (such as volume control buttons, power buttons, etc.), trackball, mouse, joystick, etc.
[0152] Display unit 540 can be used to display information input by the user or information provided to the user, as well as various menus of the mobile phone. Display unit 540 may include display panel 541, optionally configured as a Liquid Crystal Display (LCD), Organic Light-Emitting Diode (OLED), or similar display panel 541. Further, touch panel 531 may cover display panel 541. When touch panel 531 detects a touch operation on or near it, it transmits the information to processor 580 to determine the type of touch event. Subsequently, processor 580 provides corresponding visual output on display panel 541 based on the type of touch event. Although in Figure 5 In this embodiment, the touch panel 531 and the display panel 541 are two separate components to realize the input and output functions of the mobile phone. However, in some embodiments, the touch panel 531 and the display panel 541 can be integrated to realize the input and output functions of the mobile phone.
[0153] The mobile phone may also include at least one sensor 550, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor. The ambient light sensor can adjust the brightness of the display panel 541 according to the ambient light level, and the proximity sensor can turn off the display panel 541 and / or backlight when the phone is moved to the ear. As a type of motion sensor, an accelerometer sensor can detect the magnitude of acceleration in various directions (generally three axes). When stationary, it can detect the magnitude and direction of gravity and can be used for applications that recognize the phone's posture (such as landscape / portrait switching, related games, magnetometer posture calibration), vibration recognition-related functions (such as pedometer, taps), etc. Other sensors that may be configured in the mobile phone, such as gyroscopes, barometers, hygrometers, thermometers, and infrared sensors, will not be described in detail here.
[0154] Audio circuit 560, speaker 561, and microphone 562 provide an audio interface between the user and the mobile phone. Audio circuit 560 converts received audio data into electrical signals and transmits them to speaker 561, where speaker 561 converts them into sound signals for output. On the other hand, microphone 562 converts collected sound signals into electrical signals, which are received by audio circuit 560, converted into audio data, and then processed by processor 580 before being transmitted via RF circuit 55 to, for example, another mobile phone, or the audio data can be output to memory 520 for further processing.
[0155] Wi-Fi is a short-range wireless transmission technology. Through the Wi-Fi module 570, mobile phones can help users send and receive emails, browse web pages, and access streaming media, providing users with wireless broadband internet access. Although Figure 5 The Wi-Fi module 570 is shown, but it is understood that it is not a necessary component of the mobile phone and can be omitted as needed without changing the nature of the application.
[0156] The processor 580 is the control center of the mobile phone, connecting various parts of the phone through various interfaces and lines. It executes software programs and / or modules stored in the memory 520, and calls data stored in the memory 520 to perform various functions and process data, thereby providing overall monitoring of the phone. Optionally, the processor 580 may include one or more processing units; preferably, the processor 580 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into the processor 580.
[0157] The mobile phone also includes a power supply 590 (such as a battery) that supplies power to various components. The power supply can be logically connected to the processor 580 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system.
[0158] Although not shown, mobile phones may also include a camera, Bluetooth module, etc., which will not be described in detail here.
[0159] In this embodiment of the application, the processor 580 included in the mobile phone also has the function of controlling and executing the above-mentioned... Figure 2 The flowchart shown illustrates the model training method.
[0160] Figure 6This is a schematic diagram of a server structure provided in an embodiment of this application. The server 620 can vary significantly due to different configurations or performance. It may include one or more central processing units (CPUs) 622 (e.g., one or more processors) and memory 632, and one or more storage media 630 (e.g., one or more mass storage devices) for storing application programs 642 or data 644. The memory 632 and storage media 630 can be temporary or persistent storage. The program stored in the storage media 630 may include one or more modules (not shown in the diagram), each module may include a series of instruction operations on the server. Furthermore, the CPU 622 may be configured to communicate with the storage media 630 and execute the series of instruction operations in the storage media 630 on the server 620.
[0161] Server 620 may also include one or more power supplies 626, one or more wired or wireless network interfaces 650, one or more input / output interfaces 658, and / or one or more operating systems 641, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, etc.
[0162] The steps performed by the server in the above embodiments can be based on this Figure 6 The structure of server 620 is shown. For example, in the above embodiment, it consists of... Figure 2 The steps of the server shown can be based on this Figure 6 The server architecture is shown. For example, the processor 622 performs the following operations by calling instructions in memory 632:
[0163] Acquire multiple images to be identified that have been historically acquired by the image acquisition device in a preset location;
[0164] Each of the images to be identified is processed according to a preset security identification model to obtain the candidate confidence level corresponding to each image to be identified.
[0165] The image to be identified corresponding to the target confidence level is determined as a candidate sample image. The target confidence level is the confidence level among the candidate confidence levels that is greater than the second confidence level threshold and less than the first confidence level threshold. The first confidence level threshold is the confidence level threshold currently set by the security recognition model. The second confidence level threshold is less than the first confidence level threshold.
[0166] Obtain the real sample label information corresponding to each of the candidate sample images to obtain multiple target sample images, wherein each target sample image carries the corresponding real sample label information;
[0167] The security recognition model is trained based on multiple target sample images to obtain the trained security recognition model.
[0168] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0169] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and modules described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0170] In the embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, apparatuses, or modules, and may be electrical, mechanical, or other forms.
[0171] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0172] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium.
[0173] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product.
[0174] The computer program product includes one or more computer instructions. When the computer program is loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that integrates one or more available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., a solid-state disk (SSD)).
[0175] The technical solutions provided in the embodiments of this application have been described in detail above. Specific examples have been used in the embodiments of this application to illustrate the principles and implementation methods of the embodiments of this application. The description of the above embodiments is only for the purpose of helping to understand the methods and core ideas of the embodiments of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the embodiments of this application. Therefore, the content of this specification should not be construed as a limitation on the embodiments of this application.
Claims
1. A model training method, characterized in that, include: Acquire multiple images to be identified that have been historically acquired by the image acquisition device in a preset location; Each of the images to be identified is processed according to a preset security identification model to obtain a candidate confidence level for each image to be identified. The candidate confidence level is a confidence level that indicates that the image to be identified is an abnormal image. The image to be identified corresponding to the target confidence level is determined as a candidate sample image. The target confidence level is the confidence level among the candidate confidence levels that is greater than the second confidence level threshold and less than the first confidence level threshold. The first confidence level threshold is the confidence level threshold currently set by the security recognition model. The second confidence level threshold is less than the first confidence level threshold. Both the first confidence level threshold and the second confidence level threshold are confidence level thresholds corresponding to abnormal images. Obtain the real sample label information corresponding to each of the candidate sample images to obtain multiple target sample images, wherein each target sample image carries the corresponding real sample label information; The security recognition model is trained based on multiple target sample images to obtain the trained security recognition model.
2. The method according to claim 1, characterized in that, Before the image acquisition device acquires multiple historical images to be identified in a preset location, the method further includes: The target frame extraction frequency is determined based on the current scene of the image acquisition device and the preset frame extraction strategy. According to the preset computing power allocation rules and the target frame extraction frequency, the preset computing power resources are divided into a first computing power resource and a second computing power resource. The first computing power resource is used to perform recognition processing on the images to be recognized, and the second computing power resource is used to train the security recognition model. The image acquisition device acquires multiple images to be identified historically in a preset location, including: Based on the target frame rate, the image acquisition device acquires multiple images to be identified that have been historically acquired in the preset location.
3. The method according to claim 1, characterized in that, The step of training the security recognition model based on multiple target sample images to obtain the trained security recognition model includes: The multiple target sample images are divided into a training sample set and a validation sample set according to a preset sample allocation ratio. The security identification model is trained with model parameters based on the training sample set to obtain an intermediate security identification model. Based on the intermediate security identification model, the first confidence threshold is adjusted according to the verification sample set to obtain the trained security identification model.
4. The method according to claim 3, characterized in that, The adjustment of the first confidence threshold based on the intermediate security identification model and the verification sample set includes: Based on the intermediate security identification model, the recall rate and false alarm rate corresponding to each of the candidate confidence thresholds are determined according to the verification sample set. Based on the recall rate and the false positive rate, a target confidence threshold is selected from multiple candidate confidence thresholds that meet preset criteria, and the first confidence threshold is replaced with the target confidence threshold. The preset criteria are that the recall rate is higher than a preset recall rate threshold and the false positive rate is lower than a preset false positive rate threshold.
5. The method according to claim 4, characterized in that, Before determining the recall rate and false positive rate corresponding to each candidate confidence threshold among multiple candidate confidence thresholds based on the intermediate security identification model and the verification sample set, the method further includes: The candidate confidence threshold is received from user input.
6. The method according to any one of claims 1 to 5, characterized in that, Before obtaining the real sample label information corresponding to each of the candidate sample images to obtain multiple target sample images, the method further includes: The candidate sample images are clustered to obtain multiple cluster groups; According to the preset filtering rules, the candidate sample images in each cluster group are filtered to obtain the filtered candidate sample images. The step of obtaining the real sample label information corresponding to each of the candidate sample images to obtain multiple target sample images includes: Obtain the real sample label information corresponding to each of the filtered candidate sample images to obtain multiple target sample images.
7. The method according to any one of claims 1 to 5, characterized in that, After performing recognition processing on each of the images to be recognized according to a preset security recognition model to obtain the candidate confidence scores corresponding to each of the images to be recognized, the method further includes: Images to be identified with a candidate confidence level greater than or equal to the first confidence level are determined as abnormal images; Output the abnormal image.
8. The method according to any one of claims 1 to 5, characterized in that, The step of obtaining the real sample label information corresponding to each of the candidate sample images to obtain multiple target sample images includes: The candidate sample images are displayed in the display interface; The display interface receives real sample label information sent by the user for each of the candidate sample images. Each candidate sample image is then labeled with a corresponding real sample label to obtain multiple target sample images.
9. A computer device, characterized in that, The computer device includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the method as described in any one of claims 1-8.
10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, which includes program instructions that, when executed by a processor, can implement the method as described in any one of claims 1-8.
Citation Information
Patent Citations
Image recognition method and device, computer equipment and storage medium
CN111523621A
Neural network training and image processing method and device
CN112381223A