Article detection method and device, equipment, storage medium and program product

By using an abnormal item detection model based on the fusion of multiple sub-models in the security inspection system, the problem of low recognition accuracy of abnormal item in express delivery in the prior art is solved, and higher detection accuracy is achieved.

CN120220039APending Publication Date: 2025-06-27SF TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311817422.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-26
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the prior art, the accuracy of identifying abnormal items in express parcels is low, mainly because the express parcel picture logic obtained by the security check machine is different from the acquisition logic of visible light, resulting in the reduction of the accuracy of manual recognition.

Method used

By acquiring the image to be detected and inputting it into an abnormal item detection model based on preset weights, the model is obtained by fusing the model parameters of multiple sub-models. These sub-models are trained based on different loss functions (such as pixel distance, item weights, and generalized cross-over ratios), improving the accuracy of detection.

Benefits of technology

The accuracy of identification of abnormal items in express delivery images is improved, and compared with the traditional manual image judgment method, it achieves higher detection accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220039A_ABST
    Figure CN120220039A_ABST
Patent Text Reader

Abstract

The invention relates to an article detection method and device, equipment, a storage medium and a program product. Multiple types of loss functions are utilized to train multiple sub-models, model parameters of the multiple sub-models are fused, an abnormal article detection model is constructed based on the fused model parameters, and a to-be-detected image is input into the abnormal article detection model. And detecting an abnormal article in the to-be-detected image through the abnormal article detection model and outputting an abnormal article detection result, thereby determining the abnormal article in the to-be-detected image based on the abnormal article detection result. Compared with a traditional method for identifying the abnormal article based on a manual image discrimination mode, the method provided by the invention has the advantages that the plurality of loss functions are utilized to train the plurality of sub-models, the sub-models are fused into the abnormal article detection model, and the fused model is utilized to perform abnormal article detection on the image acquired by the security inspection equipment; and the detection accuracy of the security inspection equipment on the abnormal object in the image is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of security inspection, and particularly to an article detection method, device, equipment, storage medium, and program product. Background Art

[0002] With the development of logistics technology, express delivery has been integrated into daily life. For each express item in logistics, it is necessary to be inspected to ensure the safety of express delivery. Currently, the common method for detecting objects in express items is to obtain images of express items through an X-ray security inspection machine, and then determine whether there are abnormal objects in the express items by manual image inspection. However, the method of manually identifying abnormal items in express items reduces the accuracy of manual identification of abnormal items because the logic of obtaining images of express items by the X-ray security inspection machine is different from the logic of obtaining visible light.

[0003] Therefore, the current method for identifying abnormal items in express items has the defect of low identification accuracy. Summary of the Invention

[0004] Based on this, it is necessary to provide an article detection method, device, computer equipment, computer-readable storage medium, and computer program product that can improve the identification accuracy for the above technical problems.

[0005] In a first aspect, the present application provides an article detection method, and the method includes:

[0006] Obtain an image to be detected;

[0007] Input the image to be detected into an abnormal item detection model, and obtain an abnormal item detection result output by the abnormal item detection model according to the image to be detected; the abnormal item detection model is obtained by fusing model parameters of multiple sub-models based on preset weights; the multiple sub-models are trained based on different loss functions;

[0008] Determine the abnormal items in the image to be detected according to the abnormal item detection result.

[0009] In one of the embodiments, the multiple sub-models are trained by the following method:

[0010] Obtain image samples, abnormal item samples corresponding to the image samples, and multiple sub-models to be trained;

[0011] For each sub-model to be trained, input the image sample into the sub-model to be trained, and the sub-model to be trained detects abnormal items in the image sample and outputs an abnormal item prediction result;

[0012] Input the abnormal item prediction result and the abnormal item sample into the loss function corresponding to the sub-model to be trained, and adjust the model parameters of the sub-model to be trained according to the function value of the loss function.

[0013] In one embodiment, the multiple sub-models include a first sub-model, which is trained based on a first loss function. The function value of the first loss function includes a pixel distance. The step of inputting the abnormal item prediction result and the abnormal item sample into the loss function corresponding to the sub-model and adjusting the model parameters of the sub-model according to the function value of the loss function includes:

[0014] Construct a coordinate system in the image sample, obtain a first region corresponding to the abnormal item in the image sample according to the abnormal item prediction result, and obtain a second region corresponding to the abnormal item in the image sample according to the abnormal item sample;

[0015] Obtain the predicted central coordinate result, the first axis distance prediction result, and the second axis distance prediction result of the first region in the coordinate system, and obtain the central coordinate sample, the first axis distance sample, and the second axis distance sample of the second region in the coordinate system;

[0016] Input the predicted central coordinate result, the central coordinate sample, the first axis distance prediction result, the second axis distance prediction result, the first axis distance sample, and the second axis distance sample into the first loss function to obtain the pixel distance between the first region and the second region;

[0017] Adjust the model parameters of the first sub-model according to the pixel distance.

[0018] In one embodiment, the multiple sub-models include a second sub-model, which is trained based on a second loss function. The step of inputting the abnormal item prediction result and the abnormal item sample into the loss function corresponding to the sub-model and adjusting the model parameters of the sub-model according to the function value of the loss function includes:

[0019] Obtain the area sample of the third region of the abnormal item sample in the image sample, and obtain at least two area prediction results corresponding to the fourth regions of at least two abnormal item prediction results in the image sample;

[0020] For each area prediction result, obtain the weight prediction result of the abnormal item prediction result corresponding to this area prediction result according to this area prediction result, the area sample corresponding to this area prediction result, and the average area sample of the at least two area samples corresponding to the at least two area prediction results;

[0021] Input the anomaly item detection result carrying the weight prediction result and the anomaly item sample into the second loss function, and adjust the model parameters of the second sub-model according to the function value output by the second loss function.

[0022] In one embodiment, the multiple sub-models include a third sub-model, which is trained based on a third loss function. The function value of the third loss function includes the generalized intersection over union (IoU). The step of inputting the anomaly item prediction result and the anomaly item sample into the loss function corresponding to the sub-model and adjusting the model parameters of the sub-model according to the function value of the loss function includes:

[0023] Obtain the area sample of the third region of the anomaly item sample in the image sample, and obtain the area prediction result corresponding to the fourth region of the anomaly item prediction result in the image sample;

[0024] Input the area sample and the area prediction result into the third loss function to obtain the generalized intersection over union (IoU) of the anomaly item sample and the anomaly item prediction result;

[0025] Adjust the model parameters of the sub-model according to the generalized intersection over union (IoU).

[0026] In one embodiment, the step of performing model parameter fusion on multiple sub-models based on the preset weights to obtain the anomaly item detection model includes:

[0027] For each sub-model, obtain the prediction accuracy of the sub-model and obtain multiple sub-model parameters corresponding to multiple layers of neural networks in the sub-model;

[0028] Determine multiple weights corresponding to the multiple sub-models according to the multiple prediction accuracies of the multiple sub-models, and perform weighted fusion on the multiple sub-model parameters belonging to the same layer of neural network in the multiple sub-models according to the multiple weights;

[0029] Obtain the fused model parameters according to the fused multiple sub-model parameters, and obtain the anomaly item detection model based on the fused model parameters.

[0030] In a second aspect, the present application provides an item detection device, and the device includes:

[0031] An acquisition module, configured to acquire an image to be detected;

[0032] A detection module, configured to input the image to be detected into an abnormal object detection model, and obtain an abnormal object detection result output by the object detection model according to the image to be detected; the abnormal object detection model is obtained by fusing model parameters of multiple sub-models based on preset weights; the multiple sub-models are trained based on different loss functions;

[0033] A determination module, configured to determine an abnormal object in the image to be detected according to the abnormal object detection result.

[0034] In a third aspect, the present application provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.

[0035] In a fourth aspect, the present application provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0036] In a fifth aspect, the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0037] In the above object detection method, device, computer device, storage medium and computer program product, by training multiple sub-models using multiple types of loss functions, fusing the model parameters of the multiple sub-models, constructing an abnormal object detection model based on the fused model parameters, inputting the image to be detected into the abnormal object detection model, detecting the abnormal object in the image to be detected through the abnormal object detection model and outputting an abnormal object detection result, so as to determine the abnormal object in the image to be detected based on the abnormal object detection result. Compared with the traditional method of identifying abnormal objects based on manual image judgment, in the embodiments of the present application, by training multiple sub-models using multiple loss functions and fusing the sub-models into an abnormal object detection model, and using the fused model to detect abnormal objects in the images collected by the security inspection equipment, the accuracy of detecting abnormal objects in the images in the security inspection scenario is improved. Description of the Drawings

[0038] Figure 1 It is an application environment diagram of the object detection method in an embodiment;

[0039] Figure 2 It is a schematic flowchart of the object detection method in an embodiment;

[0040] Figure 3 It is a schematic diagram of model recognition in an embodiment;

[0041] Figure 4 It is a structural block diagram of the object detection device in an embodiment;

[0042] Figure 5 It is the internal structure diagram of a computer device in an embodiment. Specific implementation manners

[0043] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0044] The article detection method provided by the embodiments of the present application can be applied to, for example Figure 1 the application environment shown. Among them, the terminal communicates with the security inspection device through the network. The security inspection device can collect the to-be-detected image of the express item, so that the terminal can obtain the to-be-detected image collected by the security inspection device, and detect the abnormal item in the to-be-detected image through the abnormal item detection model. Among them, the terminal can be but is not limited to various devices such as personal computers, laptop computers, and tablet computers.

[0045] In one embodiment, as Figure 2 shown, a method for detecting an article is provided. Taking the terminal in Figure 1 as an example, the method includes the following steps:

[0046] Step S202, obtain the to-be-detected image.

[0047] Among them, an image acquisition device can be set in the security inspection device. For example, a camera can be installed in the security inspection machine. The camera can be a device for image acquisition based on invisible light. The terminal can obtain the to-be-detected image collected by the security inspection device, and the to-be-detected image collected by the image acquisition device can be an invisible light image. For example, an infrared image or an X-ray image, etc. Then the terminal can detect the abnormal item in the above to-be-detected image.

[0048] Step S204, input the to-be-detected image into the abnormal item detection model, and obtain the abnormal item detection result output by the abnormal item detection model according to the to-be-detected image; the abnormal item detection model is obtained by fusing the model parameters of multiple sub-models based on preset weights; the multiple sub-models are trained based on different loss functions.

[0049] Among them, the terminal can obtain an abnormal item detection model through pre-training. Among them, the above abnormal item detection model can be obtained by fusing multiple sub-models. For example, the terminal can construct multiple types of loss functions, including at least two of the first loss function based on pixel distance, the second loss function based on item weight, and the third loss function based on intersection over union, and based on the above multiple loss functions, train multiple sub-models respectively, and fuse the multiple sub-model parameters of the multiple sub-models according to the weights corresponding to the multiple sub-models to obtain the fused model parameters. Thus, the terminal can obtain the above abnormal item detection model according to the fused model parameters.

[0050] When detecting abnormal items, the terminal can input the above image to be detected into the above abnormal item detection model, and the abnormal item detection model can detect the abnormal items in the image to be detected and output the corresponding abnormal item detection results. Among them, there can be one or more of the above abnormal item detection results. That is, the above image to be detected can contain one or more abnormal items.

[0051] Step S206, determine the abnormal items in the image to be detected according to the abnormal item detection results.

[0052] Among them, the above abnormal item detection results can be the detection results of abnormal items in the image to be detected corresponding to the express delivery. The terminal can determine the abnormal items in the image to be detected according to the above abnormal item detection results. For example, the abnormal item detection results can include one or more regions containing abnormal items obtained after bounding and marking the abnormal items in the image to be detected, so that the terminal can obtain the identification of the region where the abnormal items are located in the image to be detected according to the abnormal item detection results. And, in some embodiments, the above abnormal item detection results can also include the recognition results of the types of the detected abnormal items by the above abnormal item detection model, for example, marking the types of the abnormal items in the regions of the abnormal items in the image to visually display the types of the detected abnormal items.

[0053] In the above item detection method, by training multiple sub-models using multiple types of loss functions, fusing the model parameters of multiple sub-models, constructing an abnormal item detection model based on the fused model parameters, inputting the image to be detected into the abnormal item detection model, detecting the abnormal items in the image to be detected by the abnormal item detection model and outputting the abnormal item detection results, and thus determining the abnormal items in the image to be detected based on the abnormal item detection results. Compared with the traditional method of identifying abnormal items by manual image judgment, in the embodiments of the present application, by training multiple sub-models using multiple loss functions and fusing the sub-models into an abnormal item detection model, and using the fused model to detect abnormal items in the images collected by the security inspection equipment, the accuracy of the security inspection equipment in detecting abnormal objects in the images is improved.

[0054] In one embodiment, multiple sub-models are trained by the following method: obtaining image samples, abnormal object samples corresponding to the image samples, and multiple sub-models to be trained; for each sub-model to be trained, inputting the image samples into the sub-model to be trained, detecting abnormal objects in the image samples by the sub-model to be trained, and outputting abnormal object prediction results; inputting the abnormal object prediction results and the abnormal object samples into the loss function corresponding to the sub-model to be trained, and adjusting the model parameters of the sub-model to be trained according to the function value of the loss function.

[0055] In this embodiment, to solve the problem that there are few surface features of objects in the images to be detected collected by the image acquisition device in the security inspection machine and it is difficult to recognize small targets, the terminal can use an abnormal object detection model based on the fusion of multiple sub-models to detect abnormal objects in the images collected by the security inspection machine. Among them, a small target refers to an image of an object with a relatively small area in the above-mentioned image to be detected. For example, if the area of an object in the image to be detected is less than a preset area threshold, then the object can be called a small target. The terminal can pre-train multiple sub-models and obtain an abnormal object detection model by fusing the trained multiple sub-models. Among them, each of the above sub-models can be trained based on different types of loss functions.

[0056] For example, the terminal can obtain image samples, abnormal object samples corresponding to the image samples, and multiple sub-models to be trained. Among them, the abnormal object samples can be information such as the regions, positions, and types of abnormal objects pre-annotated in the image samples. For each sub-model to be trained, the terminal can input the above image samples into the sub-model to be trained, and the sub-model to be trained can detect the abnormal objects in the above image samples. The sub-model to be trained can perform operations such as bounding box selection and marking on the detected abnormal objects to determine information such as the position, region, and object type of the abnormal objects in the image. Thus, the sub-model to be trained outputs abnormal object detection results according to the information such as the regions, positions, and object types of the detected abnormal objects.

[0057] The terminal can also obtain the loss function corresponding to the sub-model to be trained. Input the above abnormal item detection results and the corresponding abnormal item samples into the loss function corresponding to the sub-model to be trained. The terminal can obtain the function value of the above loss function. Among them, the above loss function can be a loss function for verifying the detection effect of the model. In some embodiments, the above sub-model can also include a loss function for verifying the classification effect, such as a function for verifying the classification effect of the model on abnormal and non-abnormal items. The terminal can adjust the model parameters of the sub-model according to the function value output by the loss function for verifying the detection effect of the model until the preset training end condition is met. At this time, the terminal can obtain the trained sub-model. The preset training end condition can include but is not limited to that within the preset number of training times, the value of the above loss function is less than or equal to the preset threshold, or the above training reaches the preset number of training times.

[0058] The terminal can train multiple sub-models using different types of loss functions by adopting the above training steps. Thus, the terminal can obtain multiple trained sub-models. Each of the above sub-models can include a multi-layer neural network. Each layer of the neural network has corresponding model parameters, and the structures of the multi-layer neural networks of the above multiple sub-models can be the same. Then the terminal can obtain multiple model parameters of the above multiple sub-models. One sub-model can correspond to multiple model parameters. The terminal can also determine multiple weights corresponding to the above multiple trained sub-models. Among them, the above weights can be determined based on the detection accuracy of the sub-model. For example, the higher the accuracy of the sub-model in detecting abnormal items, the greater its weight. The terminal can fuse the multiple model parameters of the multiple sub-models according to the above multiple weights to form fused model parameters. Among them, the fused model parameters can include parameters corresponding to the multi-layer neural network, that is, the terminal can fuse the model parameters of each layer of the neural network in the multiple sub-models to obtain fused model parameters including the model parameters of the multi-layer neural network. Thus, the terminal can construct a trained abnormal item detection model according to the fused model parameters.

[0059] Through this embodiment, the terminal can utilize multiple sub-models trained with different loss functions, and then obtain an abnormal item detection model by fusing model parameters. Thus, the terminal can detect abnormal items in the images collected by the security inspection equipment based on the fused abnormal item detection model, improving the accuracy of abnormal item detection.

[0060] In one embodiment, the multiple sub-models include a first sub-model, which is trained based on a first loss function. The function value of the first loss function includes the pixel distance. Inputting the abnormal item prediction result and the abnormal item sample into the loss function corresponding to the sub-model, and adjusting the model parameters of the sub-model according to the function value of the loss function, including: constructing a coordinate system in the image sample, and obtaining a first region corresponding to the abnormal item in the image sample according to the abnormal item prediction result; obtaining a second region corresponding to the abnormal item in the image sample according to the abnormal item sample, obtaining the predicted result of the central coordinate, the predicted result of the first axis distance, and the predicted result of the second axis distance of the first region in the coordinate system, and obtaining the central coordinate sample, the first axis distance sample, and the second axis distance sample of the second region in the coordinate system; inputting the predicted result of the central coordinate, the central coordinate sample, the predicted result of the first axis distance, the predicted result of the second axis distance, the first axis distance sample, and the second axis distance sample into the first loss function to obtain the pixel distance between the first region and the second region; adjusting the model parameters of the sub-model according to the pixel distance to reduce the pixel distance.

[0061] In this embodiment, the above-mentioned sub-model includes a sub-model trained using the first loss function. The first loss function can be a loss function based on the pixel distance. Then the terminal can input the image sample into the above-mentioned sub-model, and the sub-model can identify the area where the abnormal item is located in the image sample based on the image sample using a preset algorithm, and output the abnormal item prediction result based on the area where the abnormal item is located.

[0062] The terminal can use the loss function corresponding to the sub-model to train the sub-model. For example, the terminal can construct a coordinate system in the above-mentioned image sample and detect a first region corresponding to the abnormal item in the coordinate system. Among them, the region corresponding to the abnormal item can be a region with a specific shape, such as an ellipse.

[0063] When the above-mentioned area is elliptical, the terminal can obtain the predicted result of the central coordinate, the predicted result of the first axis distance, and the predicted result of the second axis distance of the first area corresponding to the abnormal item in the coordinate system. Among them, the predicted result of the central coordinate represents the central coordinate of the ellipse represented by the above-mentioned area, the predicted result of the first axis distance can represent the length of the major axis of the above-mentioned ellipse, and the predicted result of the second axis distance can represent the length of the minor axis of the above-mentioned ellipse. At the same time, the terminal can also obtain the second area corresponding to the abnormal item in the image sample according to the abnormal item sample, and obtain the central coordinate sample, the first axis distance sample, and the second axis distance sample of the second area in the coordinate system. Among them, the central coordinate sample represents the central coordinate of the ellipse corresponding to the area of the abnormal item sample in the coordinate system, and the first axis distance sample and the second axis distance sample can respectively represent the major axis and the minor axis of the ellipse corresponding to the area of the abnormal item sample in the coordinate system. Among them, the specific lengths of the above-mentioned first axis distance prediction result, second axis distance prediction result, first axis distance sample, and second axis distance sample can all be determined based on the area size of the abnormal item.

[0064] The terminal can input the above-mentioned predicted result of the central coordinate, central coordinate sample, first axis distance prediction result, second axis distance prediction result, first axis distance sample, and second axis distance sample into the first loss function, so that the terminal can obtain the pixel distance between the first area corresponding to the abnormal item prediction result and the second area corresponding to the abnormal item sample based on the function value of the first loss function. That is, the pixel distance between the ellipse corresponding to the abnormal item prediction result and the ellipse corresponding to the abnormal item sample. The terminal can adjust the model parameters of the sub-model based on the above pixel distance to reduce the value of the above pixel distance.

[0065] Specifically, when training the sub-model, the terminal can construct the sub-model based on yolox and autoassign as the main structure, using the label smoothing loss function as the loss function for model classification and the NWD (Normalized Weighted Distance) loss function as the loss function for detection. Among them, the above-mentioned sub-model can classify abnormal and non-abnormal items in the image, and the loss function for classification can be used to verify the accuracy of the model for the above classification; the NWD loss function is a deep learning loss function, which is a loss function based on pixel distance.

[0066] The process of the terminal detecting abnormal objects in the image through the model can be as Figure 3 shown, Figure 3 which is a schematic diagram of model recognition in an embodiment. Figure 3 Taking the area corresponding to the abnormal item as a rectangle as an example. Among them, the loss function for model detection usually uses the intersection over union method, but for small targets, such as Figure 3As shown by 300 and 302 in [Figure / Illustration], the target recognized in area 300 is smaller, and the target recognized in area 302 is larger. In area 300, area B represents the correct area of the recognized object, and areas A and C are areas obtained by offsetting different pixel values based on the object to be recognized; in area 302, area B represents the correct area of the recognized object, and areas A and C are areas obtained by offsetting different pixel values based on the object to be recognized. Among them, the pixels offset by areas A and C in area 300 and area 302 are the same. It can be seen from this that under the same number of pixel offsets, the influence of the deviation of the intersection over union value of the small target is greater than that of the large target. Therefore, the terminal reduces the influence of the number of pixels on the result based on the above NWD loss function. As shown in area 304 for example, the terminal constructs a coordinate system in the above image sample, by respectively constructing an ellipse of the abnormal item prediction result and an ellipse of the abnormal item sample. The model training is realized by the way of approximating the ellipse of the abnormal item prediction result to the ellipse of the abnormal item sample. The above ellipse can be constructed based on the center coordinates and the major and minor axes. Then the loss function of this sub-model can be specifically expressed as: . Among them, represents the pixel distance between the above two ellipses, cx and cy are the center coordinates of the above abnormal item prediction result and the abnormal item sample respectively; and w is the length of the major axis of the ellipse, h is the length of the minor axis of the ellipse. The terminal can aim at reducing the function value of the above loss function, and then improve the value of the intersection over union to optimize the network.

[0067] Specifically, when training the above sub-model, the size of the above image to be detected can be adjusted to a preset size, such as 640*640, and after normalizing the image to be detected, it is input into the sub-model for training. The above sub-model can be trained multiple times, so that the terminal obtains multiple abnormal item prediction results output by multiple trainings. The terminal can obtain the MAP (Mean Average Precision) value corresponding to the sub-model after each output of the abnormal item prediction result. And the sub-model with the largest MAP value obtained from multiple trainings is taken as the final trained sub-model.

[0068] Through this embodiment, the terminal can use the coordinate system and the first loss function based on pixel distance to train the sub-model, so that the terminal obtains an abnormal item detection model based on the fusion of multiple sub-models, and detects the express image collected by the security inspection machine based on the abnormal item detection model, improving the recognition accuracy of abnormal items.

[0069] In one embodiment, the multiple sub-models include a second sub-model, which is trained based on a second loss function. The abnormal item prediction result and the abnormal item sample are input into the loss function corresponding to the sub-model, and the model parameters of the sub-model are adjusted according to the function value of the loss function, including: obtaining the area sample of the third region of the abnormal item sample in the image sample, and obtaining at least two area prediction results corresponding to the fourth regions of at least two abnormal item prediction results in the image sample; for each area prediction result, according to the area prediction result, the area sample corresponding to the area prediction result, and the average area sample of at least two area samples corresponding to at least two area prediction results, obtaining the weight prediction result of the abnormal item prediction result corresponding to the area prediction result; inputting the abnormal item detection result carrying the weight prediction result and the abnormal item sample into the second loss function, and adjusting the model parameters of the second sub-model according to the function value output by the second loss function, so that the area of the image sample region is inversely proportional to the weight prediction result.

[0070] In this embodiment, the above-mentioned sub-model may further include a second sub-model trained using the second loss function. The second loss function may be a loss function based on item weights. If there are at least two abnormal item prediction results corresponding to the above-mentioned image sample, after the above-mentioned sub-model identifies at least two abnormal item prediction results based on the image sample, the above-mentioned sub-model may obtain the fourth regions of at least two abnormal item prediction results in the above-mentioned image sample. The terminal may also obtain the third region of the above-mentioned abnormal item sample in the image sample. Thus, the terminal may obtain at least two area prediction results of the fourth regions of at least two abnormal item prediction results in the above-mentioned image sample, and the area sample corresponding to the third region of the above-mentioned abnormal item sample in the image sample.

[0071] The terminal may obtain the average area sample of at least two area samples corresponding to the above-mentioned at least two area prediction results. Then, for each area prediction result, the terminal may obtain the weight prediction result of the abnormal item prediction result corresponding to the area prediction result according to the area prediction result and its corresponding area sample, and the above-mentioned average area prediction result. That is, each of the above-mentioned area prediction results has a corresponding weight. The terminal may input the abnormal item detection result carrying the weight prediction result and the above-mentioned abnormal item sample into the second loss function to obtain the function value of the second loss function. Thus, the terminal may adjust the model parameters of the second sub-model according to the function value output by the second loss function, so that the area of the image sample region is inversely proportional to the weight prediction result. That is, for a smaller region area, the terminal may give it a relatively large weight to improve the learning ability of the model for small targets.

[0072] Specifically, the main structure and classification loss function in the second sub-model can be the same as those of the previous sub-model. The difference is that the second loss function detected by the second sub-model can use the PLB (Pixel Level Balance) loss function. Among them, by setting the PLB loss function, the terminal can make the weights of the loss functions for identifying large targets and small targets in the sub-model different, thereby improving the learning ability of the sub-model for small targets. Then the second loss function can be specifically expressed as:

[0073] 。

[0074] Among them, box_area represents the area of the region of the abnormal item sample corresponding to the prediction result of the abnormal item in the image sample, n represents the number of prediction results of the abnormal item, area_mean represents the average area sample corresponding to the regions of all abnormal item samples, area_predict represents the area corresponding to the region of the current abnormal item prediction result, and PLB_weight represents the weight prediction result. Thus, the terminal can add the weight prediction result to the loss function of the abnormal item to be predicted based on the above average area sample, so as to improve the learning ability of the model for targets with smaller areas. Among them, in some embodiments, the terminal can also train the model by combining the loss function corresponding to the original intersection over union and the above weight-based loss function. For example, based on the loss function corresponding to the original intersection over union, make the region of the above abnormal item prediction result coincide as much as possible with the region of the corresponding abnormal item sample, and based on the above weight prediction result, improve the learning ability of the model for small target prediction.

[0075] Specifically, when training the above sub-model, the size of the above image to be detected can be adjusted to a preset size, such as 640*640, and after normalizing the image to be detected, it is input into the sub-model for training. The above sub-model can be trained multiple times, so that the terminal obtains multiple prediction results of abnormal items output after multiple trainings. The terminal can obtain the MAP value corresponding to the sub-model after each output of the abnormal item prediction result. And select the sub-model with the largest MAP value from the sub-models obtained by multiple trainings as the final trained sub-model.

[0076] Through this embodiment, the terminal can use the second loss function based on item weights to train the sub-model, and fuse multiple sub-models to obtain an abnormal item detection model. Thus, when the terminal detects the express image collected by the security inspection machine based on the abnormal item detection model, under the influence of the model parameters of the second loss function based on item weights, the recognition weight of small targets in the express image can be improved, and thus the recognition accuracy of abnormal items is improved.

[0077] In one embodiment, the multiple sub-models include a third sub-model, which is trained based on a third loss function. The function value of the third loss function includes the generalized intersection over union (GIoU). Inputting the abnormal object prediction result and the abnormal object sample into the loss function corresponding to the sub-model, and adjusting the model parameters of the sub-model according to the function value of the loss function includes: obtaining the area sample of the third region of the abnormal object sample in the image sample, and obtaining the area prediction result corresponding to the fourth region of the abnormal object prediction result in the image sample; inputting the area sample and the area prediction result into the third loss function, and obtaining the generalized intersection over union of the abnormal object sample and the abnormal object prediction result according to the function value of the third loss function; adjusting the model parameters of the sub-model according to the generalized intersection over union to increase the value of the generalized intersection over union.

[0078] In this embodiment, the above sub-model further includes a model trained using a third loss function based on the intersection over union. When the terminal trains the model using the loss function, it can obtain the area sample of the third region of the abnormal object sample in the image sample corresponding to the above sub-model, and obtain the area prediction result corresponding to the fourth region of the abnormal object prediction result in the image sample. The terminal inputs the above area sample and area prediction result into the third loss function and obtains the function value output by the third loss function. The terminal can obtain the GIOU (generalized IoU) of the above abnormal object sample and the abnormal object prediction result according to this function value. Among them, GIOU is an index used to measure the overlapping degree between two object bounding boxes. The terminal can adjust the model parameters of the above sub-model according to this generalized intersection over union to increase the value of the above generalized intersection over union. Among them, the larger the value of the generalized intersection over union, the higher the coincidence degree of the above abnormal object sample and the abnormal object prediction result, that is, the better the prediction effect of the sub-model.

[0079] Among them, when the terminal trains the sub-model, it can adjust the size of the above image to be detected, adjust it to a preset size, such as 640*640, and after normalizing the image to be detected, input it into the sub-model for training. The above sub-model can be trained multiple times to obtain multiple abnormal object prediction results, so that the terminal can obtain the MAP value corresponding to the sub-model after each output of the abnormal object prediction result. Select the one with the largest MAP value from the sub-models obtained from multiple trainings as the final trained sub-model.

[0080] Through this embodiment, the terminal can use the GIOU loss function to train the sub-model, so that the sub-model can improve the ability to identify abnormal objects in images. Moreover, by using the abnormal object detection model obtained by fusing multiple sub-models to detect the express delivery images collected by the security inspection machine, the accuracy of abnormal object recognition is also improved.

[0081] In one embodiment, the step of fusing model parameters of multiple sub-models based on preset weights to obtain an abnormal item detection model includes: for each sub-model, obtaining the prediction accuracy of the sub-model and obtaining multiple sub-model parameters corresponding to the multi-layer neural network in the sub-model; determining multiple weights corresponding to the multiple sub-models according to the multiple prediction accuracies of the multiple sub-models, and performing weighted fusion on the multiple sub-model parameters belonging to the same layer of neural network in the multiple sub-models according to the multiple weights; obtaining fused model parameters according to the fused multiple sub-model parameters, and obtaining an abnormal item detection model based on the fused model parameters.

[0082] In this embodiment, the above abnormal item detection model can be obtained based on the fused model parameters. The fused model parameters can be obtained by fusing the sub-model parameters of the above multiple sub-models. Among them, each of the above sub-models can include a multi-layer neural network, and each layer of the neural network has corresponding sub-model parameters. For each sub-model, the terminal can obtain the prediction accuracy corresponding to the sub-model. Among them, the accuracy represents the accuracy of the abnormal item prediction result output by the sub-model, that is, the regional coincidence degree between the abnormal item prediction result output by the sub-model and the corresponding abnormal item sample. Among them, each of the above sub-models can be trained multiple times, and the terminal can obtain the prediction accuracy corresponding to each training. Then, in some embodiments, the prediction accuracy of the above sub-model can be the prediction accuracy obtained by averaging the prediction accuracies corresponding to multiple trainings. That is, it belongs to an average accuracy value.

[0083] The terminal can also determine multiple weights corresponding to the multiple sub-models according to the multiple prediction accuracies of the multiple sub-models. Among them, the greater the above prediction accuracy, the greater the weight corresponding to the sub-model can be. After the terminal determines the multiple weights corresponding to the multiple sub-models, it can perform weighted fusion on the multiple sub-model parameters belonging to the same layer of neural network in the multiple sub-models. That is, the terminal can take the layer as a unit and perform weighted fusion on the sub-model parameters belonging to the same layer in the multiple sub-models. Among them, the weight of each layer of sub-model parameters can be determined based on the weight corresponding to the sub-model where it is located. After the terminal fuses the sub-model parameters of the multiple sub-models layer by layer, it can obtain fused model parameters that fuse the sub-model parameters corresponding to the multi-layer neural network.

[0084] Specifically, take the example where there are three sub-models, and each sub-model is trained with a different loss function. When the terminal fuses the sub-model parameters, it can obtain the sub-model parameters of the sub-models belonging to the same layer in the above three sub-models. Thus, the terminal can obtain multiple groups of sub-model parameters, and each group of sub-model parameters includes multiple sub-model parameters corresponding to the same layer of the above three sub-models respectively. The terminal can also determine the weights based on the prediction accuracy. Taking the prediction accuracies of the above three sub-models as p1, p2, and p3 respectively, the weights of the three sub-models can be expressed as: p1 / (p1 + p2 + p3), p2 / (p1 + p2 + p3), p3 / (p1 + p2 + p3). Thus, the terminal can use the above weights to fuse each sub-model parameter in the above multiple groups of sub-model parameters respectively, and then obtain the fused model parameters. Thus, the terminal can obtain the above abnormal item prediction model based on the fused model parameters, and the scale of the abnormal item prediction model is the same as that of the above three sub-models.

[0085] Through this embodiment, the terminal can fuse the sub-model parameters of multiple sub-models in the way of weight fusion. Thus, the terminal can construct an abnormal item detection model based on the fused model parameters obtained by fusion, and detect the express image collected by the security inspection machine based on the abnormal item detection model obtained by fusing multiple sub-models, which improves the accuracy of abnormal item recognition.

[0086] It should be understood that although each step in the flowcharts involved in the above embodiments is shown in sequence according to the indication of the arrows, these steps do not necessarily need to be executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above embodiments may include multiple steps or multiple stages. These steps or stages do not necessarily need to be executed at the same moment, but can be executed at different moments. The execution order of these steps or stages does not necessarily need to be sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.

[0087] Based on the same inventive concept, the embodiment of the present application also provides an item detection device for implementing the above-mentioned item detection method. The implementation solution provided by this device to solve the problem is similar to the implementation solution described in the above method. Therefore, the specific limitations in one or more embodiments of the item detection device provided below can refer to the limitations on the item detection method in the above text, and will not be repeated here.

[0088] In one embodiment, as Figure 4 shown, an item detection device is provided, including: an acquisition module 500, a detection module 502, and a determination module 504, where:

[0089] An acquisition module 500, configured to acquire an image to be detected.

[0090] A detection module 502, configured to input the image to be detected into an abnormal object detection model, and obtain an abnormal object detection result output by the abnormal object detection model according to the image to be detected; the abnormal object detection model is obtained by fusing model parameters of multiple sub-models based on preset weights; the multiple sub-models are trained based on different loss functions.

[0091] A determination module 504, configured to determine an abnormal object in the image to be detected according to the abnormal object detection result.

[0092] In one embodiment, the above device further includes: a training module, configured to acquire an image sample, an abnormal object sample corresponding to the image sample, and multiple sub-models to be trained; for each sub-model to be trained, input the image sample into the sub-model to be trained, and the sub-model to be trained detects abnormal objects in the image sample and outputs an abnormal object prediction result; input the abnormal object prediction result and the abnormal object sample into a loss function corresponding to the sub-model to be trained, and adjust the model parameters of the sub-model to be trained according to the function value of the loss function.

[0093] In one embodiment, the above training module is configured to construct a coordinate system in the image sample, obtain a first region corresponding to an abnormal object in the image sample according to the abnormal object prediction result, and obtain a second region corresponding to the abnormal object in the image sample according to the abnormal object sample; obtain a central coordinate prediction result, a first axis distance prediction result, and a second axis distance prediction result of the first region in the coordinate system, and obtain a central coordinate sample, a first axis distance sample, and a second axis distance sample of the second region in the coordinate system; input the central coordinate prediction result, the central coordinate sample, the first axis distance prediction result, the second axis distance prediction result, the first axis distance sample, and the second axis distance sample into a first loss function to obtain a pixel distance between the first region and the second region; adjust the model parameters of the first sub-model according to the pixel distance.

[0094] In one embodiment, the above training module is configured to obtain an area sample of a third region of an abnormal object sample in the image sample, and obtain at least two area prediction results corresponding to a fourth region of at least two abnormal object prediction results in the image sample; for each area prediction result, obtain a weight prediction result of the abnormal object prediction result corresponding to the area prediction result according to the area prediction result, the area sample corresponding to the area prediction result, and an average area sample of at least two area samples corresponding to at least two area prediction results; input the abnormal object detection result carrying the weight prediction result and the abnormal object sample into a second loss function, and adjust the model parameters of the second sub-model according to the function value output by the second loss function.

[0095] In one embodiment, the above training module is configured to obtain an area sample of a third region of an abnormal item sample in an image sample, and obtain an area prediction result corresponding to a fourth region of a prediction result of an abnormal item in the image sample; input the area sample and the area prediction result into a third loss function to obtain the generalized intersection over union (IoU) of the abnormal item sample and the prediction result of the abnormal item; and adjust the model parameters of the sub-model according to the generalized IoU.

[0096] In one embodiment, the above training module is configured to, for each sub-model, obtain the prediction accuracy of the sub-model, and obtain a plurality of sub-model parameters corresponding to a multi-layer neural network in the sub-model; determine a plurality of weights corresponding to the plurality of sub-models according to the plurality of prediction accuracies of the plurality of sub-models, and perform weighted fusion on the plurality of sub-model parameters belonging to the same layer of neural network in the plurality of sub-models according to the plurality of weights; obtain a fused model parameter according to the fused plurality of sub-model parameters, and obtain an abnormal item detection model based on the fused model parameter.

[0097] Each module in the above item detection device can be implemented in whole or in part by software, hardware, or a combination thereof. The above modules can be embedded in or independent of a processor in a computer device in the form of hardware, or stored in a memory in the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above respective modules.

[0098] In one embodiment, a computer device is provided. The computer device can be a terminal, and its internal structure diagram can be as Figure 5 shown. The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication), or other technologies. The computer program, when executed by the processor, implements an item detection method. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer covering the display screen, or a button, a trackball, or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad, or mouse, etc.

[0099] Those skilled in the art can understand, Figure 5The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0100] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the above-mentioned article detection method is implemented.

[0101] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned article detection method is implemented.

[0102] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the above-mentioned article detection method is implemented.

[0103] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0104] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.

[0105] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.

[0106] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.

Claims

1. An article detection method, characterized in that, The method includes: Obtain the image to be detected; Input the image to be detected into the abnormal object detection model, and obtain the abnormal object detection result output by the abnormal object detection model according to the image to be detected; the abnormal object detection model is obtained by fusing model parameters of multiple sub-models based on preset weights; the multiple sub-models are trained based on different loss functions; Determine the abnormal object in the image to be detected according to the abnormal object detection result.

2. The method according to claim 1, wherein The step of fusing model parameters of multiple sub-models based on the preset weights to obtain the abnormal object detection model includes: For each sub-model, obtain the prediction accuracy of the sub-model, and obtain multiple sub-model parameters corresponding to multiple layers of neural networks in the sub-model; Determine multiple weights corresponding to the multiple sub-models according to the multiple prediction accuracies of the multiple sub-models, and perform weighted fusion on the multiple sub-model parameters belonging to the same layer of neural network in the multiple sub-models according to the multiple weights; Obtain the fused model parameters according to the fused multiple sub-model parameters, and obtain the abnormal object detection model based on the fused model parameters.

3. The method according to claim 1, characterized in that, The multiple sub-models are trained by the following method: Obtain image samples, abnormal object samples corresponding to the image samples, and multiple sub-models to be trained; For each sub-model to be trained, input the image samples into the sub-model to be trained, and the sub-model to be trained detects the abnormal objects in the image samples and outputs the abnormal object prediction results; Input the abnormal object prediction results and the abnormal object samples into the loss function corresponding to the sub-model to be trained, and adjust the model parameters of the sub-model to be trained according to the function value of the loss function.

4. The method according to claim 3, wherein The multiple sub-models include a first sub-model, the first sub-model is trained based on a first loss function, the function value of the first loss function includes pixel distance, and inputting the abnormal object prediction results and the abnormal object samples into the loss function corresponding to the sub-model, and adjusting the model parameters of the sub-model according to the function value of the loss function includes: Construct a coordinate system in the image sample, obtain a first region corresponding to the abnormal object in the image sample according to the abnormal object prediction result, and obtain a second region corresponding to the abnormal object in the image sample according to the abnormal object sample; Obtain the predicted center coordinate result, the first axis distance prediction result, and the second axis distance prediction result of the first region in the coordinate system, and obtain the center coordinate sample, the first axis distance sample, and the second axis distance sample of the second region in the coordinate system; Input the predicted center coordinate result, the center coordinate sample, the first axis distance prediction result, the second axis distance prediction result, the first axis distance sample, and the second axis distance sample into the first loss function to obtain the pixel distance between the first region and the second region; Adjust the model parameters of the first sub-model according to the pixel distance.

5. The method according to claim 3, wherein The multiple sub-models include a second sub-model, which is trained based on a second loss function. The step of inputting the abnormal item prediction result and the abnormal item sample into the loss function corresponding to the sub-model and adjusting the model parameters of the sub-model according to the function value of the loss function includes: Obtain the area sample of the third region of the abnormal item sample in the image sample, and obtain at least two area prediction results corresponding to the fourth regions of at least two abnormal item prediction results in the image sample; For each area prediction result, obtain the weight prediction result of the abnormal item prediction result corresponding to this area prediction result according to this area prediction result, the area sample corresponding to this area prediction result, and the average area sample of the at least two area samples corresponding to the at least two area prediction results; Input the abnormal item detection result carrying the weight prediction result and the abnormal item sample into the second loss function, and adjust the model parameters of the second sub-model according to the function value output by the second loss function.

6. The method according to claim 3, characterized in that, The multiple sub-models include a third sub-model, which is trained based on a third loss function. The function value of the third loss function includes the generalized intersection over union (GIoU). The step of inputting the abnormal item prediction result and the abnormal item sample into the loss function corresponding to the sub-model and adjusting the model parameters of the sub-model according to the function value of the loss function includes: Obtain the area sample of the third region of the abnormal item sample in the image sample, and obtain the area prediction result corresponding to the fourth region of the abnormal item prediction result in the image sample; Input the area sample and the area prediction result into the third loss function to obtain the generalized intersection over union (GIoU) of the abnormal item sample and the abnormal item prediction result; Adjust the model parameters of this sub-model according to the generalized intersection over union (GIoU).

7. An article detection device, characterized in that, The device includes: An acquisition module, configured to acquire an image to be detected; A detection module, configured to input the image to be detected into an abnormal item detection model, and acquire the abnormal item detection result output by the abnormal item detection model according to the image to be detected; the abnormal item detection model is obtained by fusing the model parameters of multiple sub-models based on a preset weight; the multiple sub-models are trained based on different loss functions; A determination module, configured to determine the abnormal item in the image to be detected according to the abnormal item detection result.

8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer program product comprising a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 6 are implemented.