A target detection method, apparatus and electronic device
By optimizing the labels of the security inspection model through weakly supervised learning and multi-dimensional feature information, the problem of low iterative training efficiency of existing security inspection models is solved, and efficient model iteration and accurate detection are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
- Filing Date
- 2023-05-11
- Publication Date
- 2026-05-26
AI Technical Summary
Existing security inspection models have weak generalization performance and require iterative training with on-site scenario data, resulting in low efficiency.
By determining the labels of target candidate regions based on weakly supervised learning, and combining multi-dimensional feature information and initial labels, the target detection model is optimized, reducing manual annotation and improving iteration efficiency.
It reduces the manpower and time costs of label annotation and improves the iteration efficiency and accuracy of the detection model.
Smart Images

Figure CN116630870B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data detection, and more particularly to a target detection method, apparatus, and electronic device. Background Technology
[0002] Security inspection machines are electronic devices that use X-ray scanning imaging technology to conduct security checks on luggage, parcels, and other items. They are typically installed in places requiring security checks, such as subways, airports, museums, and government offices.
[0003] When security screening machines inspect items such as luggage and parcels, they typically use pre-trained detection models. However, these models have relatively weak generalization capabilities and usually require iterative training using real-world scenario data. This involves the transfer, calibration, and full-scale training of the real-world scenario data, which is time-consuming, labor-intensive, and inefficient. Summary of the Invention
[0004] In view of this, embodiments of this application provide a target detection method, apparatus, and electronic device to improve the iterative efficiency of the target detection model.
[0005] According to a first aspect of the embodiments of this application, a target detection method is provided, comprising:
[0006] Based on the multi-dimensional feature information of the target candidate region in the current round of training images, the initial label of the target candidate region is determined; the target candidate region contains the target object, and the current round of training images are images containing the target object detected by the target detection model during the security check process. The current round of training images are bound to the weak supervision information generated by the target object during the security check process. The weak supervision information is used to indicate the operation that conforms to the security check operation specifications based on the target object.
[0007] Based on the weakly supervised information bound to the training images in the current round, the multi-dimensional feature information of the target candidate regions in the training images in the current round, and the initial labels of the target candidate regions, the target labels of the target candidate regions are determined. Among them, the target labels include at least: the position of the target candidate region, the label category, the sample attribute, and the sample weight. The sample attribute represents a positive sample or a negative sample, and the sample weight is determined based on the weakly supervised information bound to the training images in the current round and the sample attributes.
[0008] Target training samples are obtained based on the target labels of target candidate regions in the current training images. Target training samples refer to target candidate regions whose sample attributes are positive in the target label, and / or target candidate regions whose sample attributes are negative in the target label. The target detection model is optimized using the target labels of the target training samples and the built-in baseline data.
[0009] According to a second aspect of the embodiments of this application, a target detection apparatus is provided, comprising:
[0010] The initial label determination module is used to determine the initial label of the target candidate region based on the multi-dimensional feature information of the target candidate region in the current round of training images. The target candidate region contains the target object. The current round of training images are images containing the target object detected by the target detection model during the security check process. The current round of training images are bound to the weak supervision information generated by the target object during the security check process. The weak supervision information is used to indicate the operation that conforms to the security check operation specifications based on the target object.
[0011] The target label determination module is used to determine the target label of the target candidate region based on the weak supervision information bound to the training image in the current round, the multi-dimensional feature information of the target candidate region in the training image in the current round, and the initial label of the target candidate region. The target label includes at least: the position of the target candidate region, the label category, the sample attribute, and the sample weight. The sample attribute represents a positive or negative sample, and the sample weight is determined based on the weak supervision information bound to the training image in the current round and the sample attribute.
[0012] The model training module is used to obtain target training samples based on the target labels of target candidate regions in the current round of training images. Target training samples refer to target candidate regions whose sample attributes are positive in the target label, and / or target candidate regions whose sample attributes are negative in the target label. The target detection model is optimized using the target labels of the target training samples and the built-in baseline data.
[0013] According to a third aspect of the embodiments of this application, an electronic device is provided, the electronic device including: a processor and a memory; wherein, the memory is used to store machine-executable instructions; and the processor is used to read and execute the machine-executable instructions stored in the memory to implement the method described above.
[0014] The technical solutions provided in this application embodiment may include the following beneficial effects:
[0015] This embodiment uses weakly supervised learning to determine the labels of target candidate regions containing target objects. Weakly supervised learning does not require manual labeling of the type corresponding to each sample. It only needs to record the operation information of the user on the target candidate region. This process does not involve manual labeling of the target candidate region, thereby reducing the labor and time costs of label labeling and improving the iterative efficiency of the detection model.
[0016] Furthermore, in this embodiment, the target label of the target candidate region is determined by multi-dimensional fusion, namely weak supervision information, multi-dimensional feature information, and initial label fusion. Compared with the method of determining the label alone, the accuracy of the candidate region labeling is improved, thereby ensuring the accuracy of the training sample. Then, by using the training sample to optimize the target detection model, a detection model with higher detection accuracy can be obtained. Attached Figure Description
[0017] Figure 1 This is a structural diagram of a security inspection system shown in an embodiment of this application.
[0018] Figure 2 This is a flowchart illustrating a target detection method in an embodiment of this application.
[0019] Figure 3 This is a flowchart illustrating a target detection method in an embodiment of this application.
[0020] Figure 4 This is a schematic diagram of the structure for obtaining the target candidate region, as shown in an embodiment of this application.
[0021] Figure 5 This is a flowchart illustrating the process of obtaining multi-dimensional feature information of the target candidate region, as shown in an embodiment of this application.
[0022] Figure 6 This is a structural diagram illustrating the adjustment of the initial label of the target object in an embodiment of this application.
[0023] Figure 7 This is a flowchart illustrating the adjustment of the initial label of the target object in an embodiment of this application.
[0024] Figure 8 This is a schematic diagram of the structure of positive and negative sample decoupling training shown in the embodiments of this application.
[0025] Figure 9 This is a flowchart illustrating a method for determining the value of a target candidate region, as shown in an embodiment of this application.
[0026] Figure 10 This is a flowchart illustrating the model evaluation process in an embodiment of this application.
[0027] Figure 11 This is a schematic diagram of the sample inheritance structure shown in the embodiments of this application.
[0028] Figure 12 This is a flowchart illustrating sample inheritance as shown in an embodiment of this application.
[0029] Figure 13 This is an overall flowchart illustrating the iteration of the target detection model in an embodiment of this application.
[0030] Figure 14This is a block diagram of a target detection device shown in an embodiment of this application.
[0031] Figure 15 This is a hardware structure diagram of the electronic device containing the target detection device in an embodiment of this application. Detailed Implementation
[0032] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0033] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0034] It should be understood that although the terms first, second, third, etc., may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0035] The embodiments of this application will now be described in detail.
[0036] Before describing the method provided in the embodiments of this application, the security inspection system involved in the embodiments of this application will be described first:
[0037] like Figure 1 As shown, Figure 1 This is a structural diagram of a security inspection system shown in an embodiment of this application. Figure 1 As shown, the security inspection system mainly includes: a security inspection machine (also known as the terminal side) and a server (also known as the training side). As one embodiment, the security inspection machine and the server can be arranged in a 1:1 ratio or an n:1 ratio to form the security inspection system. Figure 1 Using an n:1 ratio as an example, describe a security inspection system consisting of n security inspection machines and a server.
[0038] Based on the security inspection system described above, the method provided in the embodiments of this application will now be described from the perspective of the security inspection machine:
[0039] See Figure 2 , Figure 2 This is a flowchart illustrating a method according to an embodiment of this application. This process is applied to the aforementioned security screening machine. Figure 2 As shown, the process may include the following steps:
[0040] Step 201: Input the acquired current frame image into the trained object detection model so that the object detection model can detect whether there is a target object in the current frame image.
[0041] Initially, in this embodiment, before the model is optimized on the server side, the object detection model is obtained based on deep learning training. During the optimization of the model on the server side, the object detection model can be the optimized model developed on the server side. The specific implementation of the server-side optimization model is described below and will not be repeated here.
[0042] In this embodiment, the target object refers to non-compliant or unwanted items. Once the target object is detected in the package, the security screening machine will send an alarm message to an external party, such as a security officer.
[0043] Step 202: Obtain the weak supervision information bound to the current frame image when the target detection model detects the target object.
[0044] In security inspection applications, security inspection tasks require external intervention, such as security personnel. For example, a security personnel would be stationed at a designated location, such as in front of the aforementioned security inspection machine or at a monitoring position, to verify the output information and execute the corresponding security inspection strategy. Once the target detection model detects a target object, the external personnel, such as the security personnel, would perform operations that comply with security inspection procedures based on the detected target object, such as pausing the security inspection machine, zooming in on the security inspection machine image, opening suspicious packages, clicking on problematic or prohibited items in the image, or clicking on incorrectly detected prohibited items in the image. The information from these operations is called weak supervision information. In this embodiment, this weak supervision information is bound to the current frame image (such as the package image).
[0045] Step 203: Send the current frame image and the weak supervision information bound to the current frame image to the server so that the server can self-learn and optimize the object detection model based on the received image and the weak supervision information bound to the image.
[0046] As an example, the security inspection machine and the server can be directly connected or relayed through a router. Based on this, in step 203, the current frame image and the weak supervision information bound to the current frame image can be transmitted to the server through direct connection or router.
[0047] This concludes the process. Figure 2 The process is shown below.
[0048] pass Figure 2 The process shown enables the security inspection machine to acquire images and weakly supervised information, so that the server can learn from the images and weakly supervised information and optimize the target detection model based on the high-value information verified by security inspectors during the model optimization process.
[0049] The following describes how the server self-learns and optimizes the object detection model based on images and weakly supervised information:
[0050] See Figure 3 , Figure 3 This is another method flowchart provided for an embodiment of this application. This process can be applied to the aforementioned server. Of course, if one of the security screening machines has relatively powerful performance, this process can also be applied to one of those machines.
[0051] like Figure 3 As shown, the process may include the following steps:
[0052] Step 301: Determine the initial label of the target candidate region based on the multi-dimensional feature information of the target candidate region in the training image of the current round.
[0053] In this embodiment, the training image for the current round is an image containing the aforementioned target object detected during the security check process, which can be an image containing the target object detected by the target detection model.
[0054] As an example, there are many ways to obtain target candidate regions from the current training images. For instance, one can use a large-scale teacher model trained offline that cannot be deployed on edge testing to obtain target candidate regions from the current training images. Specifically, the current training images are input into the aforementioned large-scale teacher model to obtain the target candidate regions from the current training images output by the large-scale teacher model. See [link to details]. Figure 4 As shown, the target candidate region here contains at least the target object. As an example, the aforementioned ultra-large-scale teacher model can be trained using a multimodal training method, characterized by performance far exceeding that of the deployment model. However, due to its excessive size, it cannot be deployed on platforms with limited computing power.
[0055] After obtaining the aforementioned target candidate regions, multi-dimensional feature information of the target candidate regions can be further obtained. Optionally, in this embodiment, the obtained target candidate regions can be input into multiple feature extraction models with different focuses (these models are pre-trained offline) to obtain multi-dimensional feature information of the target candidate regions. Here, the feature extraction models with different focuses emphasize different aspects; for example, some feature extraction models focus on image edge features, some focus on image color features, and some focus on image channel relationship features, etc. Figure 5 This example illustrates how to obtain multi-dimensional feature information of a target candidate region. For instance, the multi-dimensional feature information here is an n*m dimensional feature vector.
[0056] After obtaining the multi-dimensional feature information of the target candidate region in the current training image, as described in step 301, the initial label of the target object can be determined based on the multi-dimensional feature information of the target candidate region in the current training image. As an example, step 301 can be implemented in several ways to determine the initial label of the target object based on the multi-dimensional feature information of the target candidate region in the current training image. For instance, for the aforementioned target candidate region, one of the template multi-dimensional feature information that matches the multi-dimensional feature information of the target candidate region can be found from the template multi-dimensional feature information of the template category already obtained offline, and the template category of that template multi-dimensional feature information can be determined as the initial label of the target object.
[0057] As an example, if the multi-dimensional feature information is an n*m dimensional feature vector, the above-mentioned search for one of the template multi-dimensional feature information that matches the multi-dimensional feature information of the target candidate region from the template multi-dimensional feature information of the template categories already obtained offline may include: for each template category already obtained offline, calculating the distance between the n*m dimensional feature vector of the target candidate region and the n*m dimensional feature vector of the template category; if the distance is less than a preset threshold, then determining that the n*m dimensional feature vector of the template category matches the n*m dimensional feature vector of the target candidate region. Of course, if multiple template categories have n*m dimensional feature vectors that match the n*m dimensional feature vector of the target candidate region, then one of them can be randomly selected, or the template category with the smallest distance to the n*m dimensional feature vector of the target candidate region can be selected.
[0058] Step 302: Based on the weak supervision information bound to the training images in the current round, the multi-dimensional feature information of the target candidate region in the training images in the current round, and the initial label of the target candidate region, determine the target label of the target candidate region.
[0059] In this embodiment, when determining the target label of the target candidate area in step 302, the verification information of the target candidate area can also be combined. Figure 6This example illustrates a specific structure for adjusting the initial label of a target object. Figure 7 An example illustrates the specific process of adjusting the initial label of the target object, which will not be elaborated here.
[0060] In the above description, the verification information for the target candidate region is obtained when the target candidate region is verified. As an example, to save resources, this embodiment may selectively select some target candidate regions for verification, such as selecting high-value (value greater than a set threshold) target candidate regions for verification. Here, the value of the target candidate region may be determined based on the initial label and the confidence level of the target candidate region, which will be described with examples below and will not be elaborated here.
[0061] Optionally, this embodiment can simplify the review process into a judgment process of right or wrong. As an example, a target candidate region judged correctly during the review process, such as containing a target object that needs to be detected or a target object that is desired to be detected, can be defined as a positive sample. Conversely, a target candidate region judged incorrectly during the review process, such as containing a target that is not desired to be detected or a target object that does not need to be detected, is defined as a negative sample. Therefore, in this embodiment, the review information of the reviewed target candidate region can at least indicate whether the target candidate region is a positive or negative sample. Further, in this embodiment, the review information of the reviewed target candidate region can also include, as needed, an externally indicated label category for the target candidate region (which differs from the initial label of the target candidate region).
[0062] As an example, in step 302, the target label includes at least: the location of the target candidate region, the label category, the sample attribute, and the sample weight. The location of the target candidate region, for example, if the target candidate region is a rectangle, can be represented by the coordinates of the top-left and bottom-right points of the rectangle. The label category can be the initial label mentioned above, or based on... Figure 7 The adjusted label categories shown in the diagram are detailed below. Figure 7 The process can be mapped to integers from 1 to L. Sample attributes indicate the category of the sample, such as positive or negative. Sample weights are determined based on the weakly supervised information bound to the training images in the current round and the sample attributes; see [link to relevant documentation] for details. Figure 7 The process is shown below.
[0063] Step 303: Obtain target training samples based on the target labels of target candidate regions in the current round of training images, and optimize the target detection model using the target labels of the target training samples and the built-in baseline data; wherein, the target training samples refer to: target candidate regions whose sample attributes are positive in the target label, and / or, target candidate regions whose sample attributes are negative in the target label.
[0064] As described above, the target label of a target candidate region has the following elements: the position of the target candidate region, the label category, the sample attribute, and the sample weight. In this embodiment, target training samples are obtained based on the target labels of target candidate regions in the current round of training images. The target training samples refer to target candidate regions whose sample attributes are positive in the target label, and / or target candidate regions whose sample attributes are negative in the target label. Then, the target labels of the target training samples and the built-in baseline data (e.g., sampling from their respective training data according to the set n:m) are used to optimize the above target detection model.
[0065] As can be seen, in this embodiment, positive and negative samples are defined before training and optimizing the object detection model. These are concepts of positive and negative samples independent of the training process. Furthermore, when training and optimizing the object detection model, this embodiment can use training samples consisting only of positive samples or training samples consisting only of negative samples, achieving decoupled training of positive and negative samples; both positive and negative samples do not need to exist simultaneously. See details... Figure 8 The diagram shown is shown in the image.
[0066] It should be noted that when training with positive samples, the sample weights of the positive samples can be used when training and optimizing the object detection model (e.g., optimizing gradient update parameters). For example, the average loss will be multiplied by the sample weights of the positive samples. Similarly, when training with negative samples, the sample weights of the negative samples can be used when training and optimizing the object detection model (e.g., optimizing gradient update parameters). For example, the average loss will be multiplied by the sample weights of the negative samples.
[0067] This concludes the process. Figure 3 The process is shown below.
[0068] pass Figure 3 The process shown enables the server to optimize the target detection model based on the images sent by the security inspection machine and weakly supervised information. This process does not involve manual labeling of the target candidate areas, thereby reducing the manpower and time costs of labeling and improving the iteration efficiency of the detection model.
[0069] The following is about Figure 7 The process of adjusting the initial label of the target object is explained.
[0070] See Figure 7 The process includes the following steps:
[0071] Step 701: Based on the multi-dimensional feature information of the target candidate regions in the current round of training images, cluster each target candidate region to obtain at least one cluster.
[0072] In this embodiment, target candidate regions within the same cluster have matching label categories. The matching label categories indicate that the target objects contained in the target candidate regions within the cluster have the same or similar characteristics. For example, if the category of cluster 1 is apple, then the label types of the target candidate regions within the cluster are all apple, or some are apple and the rest are other fruits with the same or similar characteristics as apple.
[0073] As an example, when clustering each target candidate region, clustering can be performed based on the similarity between the multi-dimensional feature information corresponding to each target candidate region. For example, target candidate boxes with similarity greater than or equal to a set similarity can be aggregated into the same cluster.
[0074] Step 702: For each cluster, based on prior information and / or the verification information of the target candidate regions that have been verified in the cluster, determine the label category of each target candidate region in the cluster.
[0075] In this embodiment, the prior information is that the target candidate regions in the same cluster have matching label categories.
[0076] Optionally, in step 702, for each cluster, if the label category of any reviewed target candidate region within that cluster, as indicated by the review information, differs from the current label category of that target candidate region (e.g., the initial label mentioned above), then the current label category of that target candidate region is adjusted to match the label category indicated by the review information. For example, if the initial label of target candidate region 1 within a cluster is "apple," and the review information indicates that the label category of target candidate region 1 is "pear," then the label category of target candidate region 1 is adjusted from "apple" to "pear." Then, based on the aforementioned prior information and according to the voting method configured for the cluster, the label categories of unreviewed target candidate regions within that cluster are adjusted so that at least a specified proportion (e.g., 50%) of the target candidate regions in that cluster have their label categories adjusted to the same label category; for example, at least 50% of the target candidate regions in that cluster have the label category "pear."
[0077] It should be noted that in step 702, there is no fixed order between adjusting the label category of the target candidate region based on the verification information and adjusting the label category of the target candidate region in the cluster based on prior information. It is also possible to first adjust the label category of the unverified target candidate regions in the cluster based on the aforementioned prior information and according to the voting method configured for the cluster, so that at least a specified proportion (e.g., 50%) of the target candidate regions in the cluster have their label categories adjusted to the same label category, for example, at least 50% of the target candidate regions in the cluster have the label category "pear". Then, for each cluster, if the label category indicated by the verification information of any verified target candidate region in the cluster is different from the current label category of the target candidate region, such as the initial label mentioned above, then the current label category of the target candidate region is adjusted so that the adjusted label category matches the label category indicated by the verification information. For example, if the initial label of target candidate region 1 within a cluster is "apple", and the verification information indicates that the label category of target candidate region 1 is "pear", then the label category of target candidate region 1 will be adjusted from "apple" to "pear".
[0078] Optionally, in this embodiment, when making the above adjustments to any cluster, it is necessary to first check whether the number of target candidate regions and the feature variance within the cluster meet the corresponding conditions. For example, the number of target candidate regions within the cluster reaches a preset number, and the feature variance calculated based on the multi-dimensional feature information of the target candidate regions in the cluster is less than a preset variance threshold. Once the number of target candidate regions and the feature variance within the cluster meet the corresponding conditions, it indicates that the cluster has a high degree of aggregation, and the features of the target candidate regions in the cluster are relatively similar, and the above operations on the cluster can be performed.
[0079] Step 703: Determine the corresponding security level based on the weak supervision information of the target candidate region.
[0080] In this embodiment, the security level corresponding to the target candidate area represents the security level of the target object contained in the target candidate area. The higher the security level of the target candidate area, the higher the security level of the target object contained in the target candidate area. For example, the security level of a target candidate area containing books is higher than the security level of a target candidate area containing flammable materials.
[0081] As one embodiment, for each target candidate area, a security level assessment score can be calculated for the weak supervision information of that target candidate area under a preset security level assessment perspective. Based on the security level assessment score, a security level matching the score can be found in a preset mapping table. Different preset security level assessment perspectives can correspond to different security inspection scenarios, including but not limited to subway security inspections, airport security inspections, museum security inspections, and government agency security inspections. Under different preset security level assessment perspectives, the same weak supervision information corresponds to different security level assessment scores. For example, the security level assessment score for liquids in an airport security inspection scenario is higher than that for liquids in a government agency security inspection scenario.
[0082] Furthermore, the aforementioned rating score characterizes the risk level of non-compliant items within the target candidate area. A higher rating score indicates a higher risk level for the non-compliant item. Different weak surveillance information corresponds to different rating scores; for example, the rating score for information regarding the opening of a suspicious package is higher than the rating score for information regarding the magnified view of the security scanner.
[0083] Furthermore, after determining the level assessment score corresponding to the weak supervision information of the target candidate area, the level assessment scores corresponding to the weak supervision information of the target candidate area can be summed. For example, if the weak supervision information of the target candidate area includes the security inspection machine pause operation information and the security inspection machine screen magnification operation information, then the level assessment score of the security inspection machine pause operation information and the level assessment score of the security inspection machine screen magnification operation information are calculated to obtain the level assessment score corresponding to the target candidate area. Then, the security level corresponding to the level assessment score of the target candidate area is found from the preset mapping table to obtain the security level of the target candidate area.
[0084] Optionally, the aforementioned security levels may include high-risk level, suspicious level, and normal level, wherein the normal level indicates that the target objects contained in the target candidate area are compliant items; the suspicious level indicates that the target objects contained in the target candidate area may be non-compliant items; and the high-risk level indicates that the target objects contained in the target candidate area are non-compliant items.
[0085] Step 704: Based on the security level and sample attributes, determine the sample weights corresponding to the sample attributes of the target candidate region.
[0086] In this embodiment, the sample attribute indicates whether the target candidate region is a positive or negative sample. Optionally, in this embodiment, for unverified target candidate regions, the target candidate region can be directly defaulted to a negative sample.
[0087] The sample weights for different sample attributes and their corresponding security levels are also different. For example, when the target candidate region is a positive sample, the sample weights for the high-risk, suspicious, and normal security levels are 2.0, 1.0, and 0.5, respectively; when the target candidate region is a negative sample, the sample weights for the high-risk, suspicious, and normal security levels are 0.5, 1.0, and 2.0, respectively. It should be noted that the above sample weight values are only examples and are not specifically limited in this embodiment. In practical applications, they can be adjusted according to actual needs.
[0088] Step 705: Determine the target label of the target candidate region based on its location, label category, sample attributes, and sample weight.
[0089] In this embodiment, the target label of the target candidate region consists of four elements: the position of the target candidate region, the label category of the target candidate region, the sample attribute, and the sample weight. Therefore, after determining the position, label category, sample attribute, and sample weight of the target candidate region, the target label of the target candidate region is also determined.
[0090] This concludes the process. Figure 7 The process is shown below.
[0091] pass Figure 7 The process shown adjusts the initial labels of the target objects, improving the accuracy of labeling the target candidate regions. This process does not require manual labeling, thereby reducing the iteration cycle of the detection model and improving its iteration efficiency.
[0092] The following describes how to determine the value of the target candidate region:
[0093] See Figure 9 , Figure 9 A flowchart illustrating a method for determining the value of a target candidate region provided in an embodiment of this application. Figure 9 As shown, the process may include the following steps:
[0094] Step 901: Determine the confidence value corresponding to the target candidate region based on the obtained confidence level.
[0095] In this embodiment, the confidence level of the target candidate region can be determined when extracting the target candidate region.
[0096] As an example, in practical use, a mapping table between confidence level and confidence value can be constructed. This mapping table allows the determination of the confidence value of a target candidate region based on its confidence level. Optionally, a higher confidence level for a target candidate region corresponds to a lower confidence value, and vice versa.
[0097] Step 902: Determine the feature distance value corresponding to the target candidate region based on the feature distance between the multi-dimensional feature information corresponding to the target candidate region and the multi-dimensional feature information corresponding to the initial label.
[0098] In the above description, the distance between the multi-dimensional feature information of the target candidate region and the multi-dimensional feature information of the initial label can be Euclidean distance, Manhattan distance, Chebyshev distance, etc., and this embodiment is not specifically limited to it.
[0099] As an example, the aforementioned feature distance and feature distance value are negatively correlated; that is, the smaller the feature distance, the greater the feature distance value corresponding to the target candidate region. Similarly, a mapping relationship between feature distance and feature distance value can be constructed based on actual usage experience. Furthermore, after determining the feature distance between the multi-dimensional feature information corresponding to the target candidate region and the multi-dimensional feature information corresponding to the initial label, the feature distance value corresponding to the target candidate region can be determined according to the aforementioned mapping relationship.
[0100] Step 903: The confidence value and feature distance value are weighted and summed to obtain the value of the target candidate region.
[0101] As an example, the value of the target candidate region can be calculated by the following formula:
[0102] score all =a1*score dis +a2*score conf
[0103] In the above formula, score all The score represents the value of the target candidate region. dis The score represents the feature distance value of the target candidate regions mentioned above. conf a1 represents the confidence value corresponding to the target candidate region; a2 and a1 represent the weights of the feature distance value and the confidence value, respectively. These two weight values can be adjusted according to actual needs.
[0104] In this embodiment, the value corresponding to the target candidate region represents the priority of the target candidate region being recommended for review. The higher the value, the higher the priority of the target candidate region being recommended for review; the lower the value, the lower the priority of the target candidate region being recommended for review.
[0105] In this embodiment, the value of each target candidate region can be determined through the above method. After determining the value of each target candidate region, the server pushes the candidate regions to the review terminal according to the priority of the candidate regions being recommended for review based on their corresponding values, so that users can review the attribute information of the candidate regions based on their values. For example, in this embodiment, multiple target candidate regions can be sorted in descending order according to their values, with higher-value regions having a higher priority for being pushed to external review.
[0106] In specific implementation, when pushing the target candidate area to be reviewed, this embodiment first detects whether the value of the target candidate area is greater than or equal to the set value threshold. Only candidate areas that are greater than or equal to the set value threshold can be reviewed.
[0107] This concludes the process. Figure 9 The process is shown below.
[0108] pass Figure 9 The process shown determines the value of the target candidate region, and then determines whether to push it to an external review based on the value of the target candidate region, so as to further ensure the accuracy of the target label of the subsequent target candidate region determination, thereby improving the detection accuracy of the target detection model.
[0109] The following describes how the server evaluates the object detection model:
[0110] See Figure 10 , Figure 10 This is a model evaluation structure diagram provided in the embodiments of this application, such as... Figure 10 As shown, the server can optimize the target detection model before optimization (i.e., Figure 10 The initial detection model and the optimized target detection model were evaluated on the test dataset. The evaluation can use FPPI (False Positive per Image, target detection evaluation index) as the evaluation index to evaluate the two detection models, and the detection model with the best evaluation result is pushed to the security inspection machine so that the security inspection machine can load the best model.
[0111] It should be noted that by evaluating and testing the two detection models, it can be ensured that the detection model loaded by the security inspection machine is the optimal detection model, and that the performance of the detection model will not decrease after each incremental iteration of training.
[0112] In addition, it should be noted that the FPPI metric mainly includes two aspects of evaluation metrics: accuracy evaluation metrics and speed evaluation metrics. Accuracy evaluation metrics include, but are not limited to, mean accuracy, precision, confusion matrix, precision, recall, average correctness, crossover, division and union, nonmaximum suppression, etc. Speed evaluation metrics include, but are not limited to, FPS (Frames Per Second, the number of images processed per second or the time required to process each image).
[0113] The following describes how the server implements the inheritance of training samples:
[0114] See Figure 11 , Figure 11 This is a structural diagram of the training sample inheritance provided in the embodiments of this application. Figure 12 An exemplary flowchart illustrating the implementation of training sample inheritance is shown, such as... Figure 12 As shown, the method includes the following steps:
[0115] Step 1201: For each training image, obtain the first detection result of the object detection model before optimization and the second detection result of the object detection model after optimization.
[0116] In this embodiment, after optimizing the object detection model, the server inputs the same training image into both the unoptimized and optimized object detection models. Upon receiving the training image, each object detection model performs detection on the image to identify the target objects contained within it. For example, the two models can determine candidate regions containing objects in the training image and use the method provided in this application to identify the target labels corresponding to these candidate regions. These target labels represent the detection results of the object detection models for the target objects in the training image.
[0117] Step 1202: Determine the difference score of the training image based on the degree of difference between the first detection result and the second detection result.
[0118] In this embodiment, after obtaining the detection results of two object detection models on the same training image, the server can classify the target objects detected by each object detection model into multiple groups. The target objects in each group have the same category. For example, if the training image contains 10 target objects and each target object corresponds to a candidate region, the server divides the 10 candidate regions into two groups according to the different target labels of the 10 candidate regions. The target objects contained in the candidate regions of one group are flammable materials, and the target objects contained in the candidate regions of the other group are explosive materials.
[0119] After classifying the target objects in the training image according to the target labels corresponding to the candidate regions, the server matches the candidate regions belonging to the same group in the two detection results. If the candidate regions belonging to the same group in the two detection results are all the same, the difference score corresponding to the group is 0. If the candidate regions belonging to the same group in the two detection results are different, the number of different candidate regions is recorded, and the confidence of the unmatched candidate regions is recorded. For each pair of candidate regions that are different detected, the difference score corresponding to the group is incremented by 1.
[0120] The difference score for each group can be calculated using the above method. By calculating the difference scores for all groups, the difference score between the two test results can be obtained.
[0121] It should be noted that when candidate regions belonging to the same group differ in two detection results, an unmatched candidate region refers to a region that the original object detection model could detect but the optimized model could not, or vice versa. This unmatched candidate region may be a candidate region misjudged by either the original or optimized model, or it may be a candidate region that neither model could recognize. Therefore, such unmatched candidate regions have significant value. Recording the confidence level of these unmatched candidate regions allows them to be presented as high-value data for review during the next training phase.
[0122] Step 1203: Based on the difference scores of each training image and the preset maximum number of samples, delete invalid images from all stored training images; the difference scores of invalid images are less than the difference scores of other training images that have not been deleted.
[0123] In this embodiment, if the difference score between two detection results for the same training image is 0, it indicates that the detection results of the two target detection models for the same training image are consistent. In this case, if the storage unit in the security inspection system has sufficient capacity, that is, the number of training images stored in the storage unit is less than the preset maximum number of samples, the training image can be stored in the storage unit; if the storage unit in the security inspection system has insufficient capacity, that is, the number of training images stored in the storage unit is greater than or equal to the preset maximum number of samples, the training image does not need to be stored.
[0124] If the difference score between two detection results is large (e.g., greater than a certain difference threshold), it indicates that the detection results of the two detection models are inconsistent. In this case, the server stores the training image in the storage unit as a sample for the next training.
[0125] Before storing the training image in the storage unit, the server needs to check if the storage unit has sufficient capacity. If the storage unit has sufficient capacity to store the training image, the server directly stores the training image in the storage unit. Otherwise, the server deletes invalid images from all training images already stored in the storage unit and stores the training image in the storage unit after deleting the invalid images.
[0126] If two detection results differ, but the difference score is small (e.g., below a certain difference threshold), the server checks if the storage unit has sufficient capacity. If the storage unit has enough capacity to store the training image, the server directly stores the training image in the storage unit. Otherwise, the server compares the difference scores of the training images already stored in the storage unit with the difference score of the current training image to determine whether to continue storing the current training image. For example, if the difference score of the current training image is less than or equal to the difference scores of all training images already stored in the storage unit, the server will not store the current training image; otherwise, the server deletes the training image with the smallest difference score from the storage unit and stores the current training image.
[0127] This concludes the process. Figure 12 The process is shown below.
[0128] pass Figure 12 The process shown inherits the training images in the storage unit, retaining the more valuable images for subsequent model training. This not only improves the iteration efficiency of the detection model but also avoids the problem of storage failure caused by the gradual increase of training data in the storage unit.
[0129] As can be seen from the above, the target detection method provided in this embodiment adopts a multi-dimensional fusion pseudo-label technology that combines weakly supervised information and strongly supervised information. The operation information of security inspectors on security inspection machines and / or packages is used as weakly supervised information. Combined with low-workload verification information, the generation efficiency of target labels can be effectively improved, thereby improving the iteration efficiency of the detection model.
[0130] This concludes the description of the methods provided in the embodiments of this application. Additionally, Figure 13 The flowchart of iterative training of the object detection model on the server side is also provided. The specific process has been described above and will not be repeated here.
[0131] Corresponding to the embodiments of the aforementioned methods, this application also provides embodiments of a target detection device and the electronic devices and storage media used therein.
[0132] like Figure 14 As shown, Figure 14This is a block diagram of a target detection device according to an embodiment of this application. The target detection device includes: an initial label determination module, a target label determination module, and a model training module.
[0133] The initial label determination module is used to determine the initial label of the target candidate region based on the multi-dimensional feature information of the target candidate region in the current round of training images. The target candidate region contains the target object. The current round of training images are images containing the target object detected by the target detection model during the security check process. The current round of training images are bound to the weak supervision information generated by the target object during the security check process. The weak supervision information is used to indicate the operation that conforms to the security check operation specifications based on the target object.
[0134] The target label determination module is used to determine the target label of the target candidate region based on the weak supervision information bound to the training image in the current round, the multi-dimensional feature information of the target candidate region in the training image in the current round, and the initial label of the target candidate region. The target label includes at least: the position of the target candidate region, the label category, the sample attribute, and the sample weight. The sample attribute represents a positive or negative sample, and the sample weight is determined based on the weak supervision information bound to the training image in the current round and the sample attribute.
[0135] The model training module is used to obtain target training samples based on the target labels of target candidate regions in the current round of training images. Target training samples refer to target candidate regions whose sample attributes are positive in the target label, and / or target candidate regions whose sample attributes are negative in the target label. The target detection model is optimized using the target labels of the target training samples and the built-in baseline data.
[0136] Optionally, the initial label determination module is specifically used to find, for the target candidate region, one of the template multi-dimensional feature information that matches the multi-dimensional feature information of the target candidate region from the template multi-dimensional feature information of the template category already obtained offline; and to determine the template category of the found template multi-dimensional feature information as the initial label.
[0137] Optionally, determining the initial label of the target candidate region based on the multi-dimensional feature information of the target candidate region in the current round of training images includes:
[0138] For each target candidate region, find one of the template multi-dimensional feature information that matches the multi-dimensional feature information of the target candidate region from the template multi-dimensional feature information of the obtained template category.
[0139] The template category of one of the found templates is used to determine the initial label of the target candidate region.
[0140] Optionally, determining the target label of the target candidate region based on the weak supervision information bound to the training images in the current round, the multi-dimensional feature information of the target candidate region in the training images in the current round, and the initial label of the target candidate region includes:
[0141] Based on the multi-dimensional feature information of the target candidate regions in the current round of training images, each target candidate region is clustered to obtain at least one cluster.
[0142] For each cluster, based on prior information and / or the verification information of the target candidate regions that have been verified in the cluster, the label category of each target candidate region in the cluster is determined; the prior information is: target candidate regions in the same cluster have matching label categories, and the verification information of the target candidate regions at least indicates the label category of the target candidate regions;
[0143] For each target candidate region in each cluster, the corresponding security level is determined based on the weak supervision information of the target candidate region. Based on the security level and sample attributes, the sample weights corresponding to the sample attributes of the target candidate region are determined. Based on the position of the target candidate region, the label category of the target candidate region, the sample attributes, and the sample weights, the target label of the target candidate region is determined.
[0144] Optionally, determining the label category of each target candidate region in the cluster based on prior information and / or the verification information of the verified target candidate regions in the cluster includes:
[0145] For each cluster, if the label category of any reviewed target candidate region within that cluster, as indicated by the review information, differs from the current label category of that target candidate region, then the current label category of that target candidate region is adjusted to match the label category indicated by the review information; and / or,
[0146] Based on the prior information and in accordance with the voting method configured for the cluster, the label categories of the unreviewed target candidate regions in the cluster are adjusted so that at least a specified proportion of the target candidate regions in the cluster are adjusted to the same label category.
[0147] Optionally, determining the corresponding security level based on the weak supervision information of the target candidate region includes:
[0148] For each target candidate region, calculate the level evaluation score of the weak supervision information of that target candidate region under the preset level evaluation perspective;
[0149] Based on the assessment score, a security level matching the assessment score is searched in a preset mapping table.
[0150] Optionally, the verification information of the target candidate region is obtained when a verification operation is performed on the target candidate region;
[0151] The verification operation is also used to determine whether the target candidate region is a correct sample or an incorrect sample. When the target candidate region is determined to be a correct sample, the target candidate region is defined as a positive sample; when the target candidate region is determined to be an incorrect sample, the target candidate region is defined as a negative sample.
[0152] Optionally, the value of the target candidate region being reviewed is greater than or equal to a set value threshold;
[0153] The value of the target candidate region is determined through the following steps:
[0154] Based on the obtained confidence levels corresponding to the target candidate regions, the confidence value corresponding to the target candidate regions is determined, wherein the confidence value is used to characterize the priority of being reviewed;
[0155] The feature distance value corresponding to the target candidate region is determined based on the feature distance between the multi-dimensional feature information corresponding to the target candidate region and the multi-dimensional feature information corresponding to the initial label; the feature distance value is used to characterize the priority of being reviewed.
[0156] The value of the target candidate region is obtained by weighted summation of the confidence value and the feature distance value.
[0157] Optionally, the model training module further obtains, for each training image, the first detection result of the object detection model before optimization and the second detection result of the object detection model after optimization; and determines the difference score of the training image based on the degree of difference between the first detection result and the second detection result.
[0158] Based on the difference scores of each training image and the preset maximum number of samples, invalid images are deleted from all stored training images; the difference scores of the invalid images are less than the difference scores of the other training images that have not been deleted.
[0159] The specific implementation process of the functions and roles of each module in the above-mentioned device can be found in the implementation process of the corresponding steps in the above-mentioned target detection method, and will not be repeated here.
[0160] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this application according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0161] Correspondingly, embodiments of this application also provide Figure 15 The hardware structure diagram of the electronic device shown is as follows: Figure 15 As shown, the electronic device can be a device implementing the above-described method. Figure 15 As shown, the hardware structure includes a processor and a memory. The memory stores machine-executable instructions; the processor reads and executes the machine-executable instructions stored in the memory to implement the target detection method embodiment described above.
[0162] As one embodiment, the memory can be any electronic, magnetic, optical, or other physical storage device that can contain or store information such as executable instructions, data, etc. For example, the memory can be volatile memory, non-volatile memory, or similar storage media. Specifically, the memory can be RAM (Random Access Memory), flash memory, storage drives (such as hard disk drives), solid-state drives, any type of storage disk (such as optical discs, DVDs, etc.), or similar storage media, or combinations thereof.
[0163] This concludes the process. Figure 15 Description of the electronic device shown.
[0164] Based on the same inventive concept, this embodiment also provides a computer-readable storage medium. This computer-readable storage medium is used to store a computer program; when executed by a processor, the computer program implements the method embodiment shown above.
[0165] The foregoing has described specific embodiments of this application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired results. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0166] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention filed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not claimed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.
[0167] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
[0168] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A target detection method, characterized in that, The method includes: Based on the multi-dimensional feature information of the target candidate region in the current round of training images, the initial label of the target candidate region is determined; the target candidate region contains the target object, and the current round of training images are images containing the target object detected by the target detection model during the security check process. The current round of training images are bound to the weak supervision information generated by the target object during the security check process. The weak supervision information is used to indicate the operation that conforms to the security check operation specifications based on the target object. Based on the weakly supervised information bound to the training images in the current round, the multi-dimensional feature information of the target candidate regions in the training images in the current round, and the initial labels of the target candidate regions, the target labels of the target candidate regions are determined. Among them, the target labels include at least: the position of the target candidate region, the label category, the sample attribute, and the sample weight. The sample attribute represents a positive sample or a negative sample, and the sample weight is determined based on the weakly supervised information bound to the training images in the current round and the sample attributes. Target training samples are obtained based on the target labels of target candidate regions in the current training images. Target training samples refer to target candidate regions whose sample attributes are positive in the target label, and / or target candidate regions whose sample attributes are negative in the target label. The target detection model is optimized using the target labels of the target training samples and the built-in baseline data. The loss used to optimize the target detection model is calculated based on the sample weights in the target labels.
2. The method according to claim 1, characterized in that, The initial label determination of the target candidate region based on the multi-dimensional feature information of the target candidate region in the current training image includes: For each target candidate region, find one of the template multi-dimensional feature information that matches the multi-dimensional feature information of the target candidate region from the template multi-dimensional feature information of the obtained template category. The template category of one of the found templates is used to determine the initial label of the target candidate region.
3. The method according to claim 1, characterized in that, The determination of the target label for the target candidate region based on the weakly supervised information bound to the training images in the current round, the multi-dimensional feature information of the target candidate region in the training images in the current round, and the initial label of the target candidate region includes: Based on the multi-dimensional feature information of the target candidate regions in the current round of training images, each target candidate region is clustered to obtain at least one cluster. For each cluster, based on prior information and / or the verification information of the target candidate regions that have been verified in the cluster, the label category of each target candidate region in the cluster is determined; the prior information is: target candidate regions in the same cluster have matching label categories, and the verification information of the target candidate regions at least indicates the label category of the target candidate regions; For each target candidate region in each cluster, the corresponding security level is determined based on the weak supervision information of the target candidate region. Based on the security level and sample attributes, the sample weights corresponding to the sample attributes of the target candidate region are determined. Based on the position of the target candidate region, the label category of the target candidate region, the sample attributes, and the sample weights, the target label of the target candidate region is determined.
4. The method according to claim 3, characterized in that, The determination of the label category of each target candidate region in the cluster based on prior information and / or the verification information of the verified target candidate regions in the cluster includes: For each cluster, if the label category of any reviewed target candidate region within that cluster, as indicated by the review information, differs from the current label category of that target candidate region, then the current label category of that target candidate region is adjusted to match the label category indicated by the review information; and / or, Based on the prior information and in accordance with the voting method configured for the cluster, the label categories of the unreviewed target candidate regions in the cluster are adjusted so that at least a specified proportion of the target candidate regions in the cluster are adjusted to the same label category.
5. The method according to claim 3, characterized in that, The process of determining the corresponding security level for each target candidate region within each cluster, based on the weak supervision information of that target candidate region, includes: For each target candidate region, calculate the level evaluation score of the weak supervision information of that target candidate region under the preset level evaluation perspective; Based on the assessment score, a security level matching the assessment score is searched in a preset mapping table.
6. The method according to claim 3, characterized in that, The verification information of the target candidate region is obtained when the verification operation is performed on the target candidate region; The verification operation is also used to determine whether the target candidate region is a correct sample or an incorrect sample. When the target candidate region is determined to be a correct sample, the target candidate region is defined as a positive sample; when the target candidate region is determined to be an incorrect sample, the target candidate region is defined as a negative sample.
7. The method according to claim 3 or 6, characterized in that, The value of the target candidate region being reviewed is greater than or equal to a set value threshold. The value of the target candidate region is determined through the following steps: Based on the obtained confidence levels corresponding to the target candidate regions, the confidence value corresponding to the target candidate regions is determined, wherein the confidence value is used to characterize the priority of being reviewed; The feature distance value corresponding to the target candidate region is determined based on the feature distance between the multi-dimensional feature information corresponding to the target candidate region and the multi-dimensional feature information corresponding to the initial label; the feature distance value is used to characterize the priority of being reviewed. The value of the target candidate region is obtained by weighted summation of the confidence value and the feature distance value.
8. The method according to claim 1, characterized in that, The method further includes: For each training image, obtain the first detection result of the object detection model before optimization and the second detection result of the object detection model after optimization; determine the difference score of the training image based on the degree of difference between the first detection result and the second detection result. Based on the difference scores of each training image and the preset maximum number of samples, invalid images are deleted from all stored training images; the difference scores of the invalid images are less than the difference scores of the other training images that have not been deleted.
9. A target detection device, characterized in that, The device includes: The initial label determination module is used to determine the initial label of the target candidate region based on the multi-dimensional feature information of the target candidate region in the current round of training images. The target candidate region contains the target object. The current round of training images are images containing the target object detected by the target detection model during the security check process. The current round of training images are bound to the weak supervision information generated by the target object during the security check process. The weak supervision information is used to indicate the operation that conforms to the security check operation specifications based on the target object. The target label determination module is used to determine the target label of the target candidate region based on the weak supervision information bound to the training image in the current round, the multi-dimensional feature information of the target candidate region in the training image in the current round, and the initial label of the target candidate region. The target label includes at least: the position of the target candidate region, the label category, the sample attribute, and the sample weight. The sample attribute represents a positive or negative sample, and the sample weight is determined based on the weak supervision information bound to the training image in the current round and the sample attribute. The model training module is used to obtain target training samples based on the target labels of target candidate regions in the current round of training images. Target training samples refer to target candidate regions whose sample attributes are positive in the target label, and / or target candidate regions whose sample attributes are negative in the target label. The module optimizes the target detection model using the target labels of the target training samples and the built-in baseline data. The loss used to optimize the target detection model is calculated based on the sample weights in the target label.
10. An electronic device, characterized in that, Electronic devices include: processors and memory; The memory is used to store machine-executable instructions; The processor is configured to read and execute machine-executable instructions stored in the memory to implement the method as described in any one of claims 1 to 8.