Target identification method, target repair network training method and related equipment

By tracking and contour correction of target objects in the smart farming site, and using contour correction mask images for anomaly classification and identification, the problem of misidentification caused by shortening the tracking time is solved, and efficient and accurate anomaly target identification is achieved.

CN120852827APending Publication Date: 2025-10-28MASHANG CONSUMER FINANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410516366.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-04-26
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

In smart farming sites, existing technologies make it difficult to improve the recognition accuracy of abnormal target objects while shortening the tracking time of target objects, and existing methods are prone to misidentification and insufficient recognition timeliness.

Method used

By tracking target objects within a preset monitoring area, suspected abnormal objects are initially identified, and their outlines are corrected. Anomaly classification and identification are then performed using the outline-corrected mask image, filtering out false alarms and improving the accuracy of identification.

Benefits of technology

While shortening the tracking time, it improves the accuracy of identifying abnormal target objects, balancing the timeliness and accuracy of identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120852827A_ABST
    Figure CN120852827A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a target recognition method, a target repair network training method and related equipment, and the method comprises the steps: carrying out the tracking of each target object in a preset monitoring region, preliminarily determining a suspected abnormal object, segmenting a contour correction mask image from a target image for each suspected abnormal object, and carrying out the recognition of a target repair network. And carrying out anomaly classification and recognition based on the contour correction mask graph to obtain a final anomaly classification and recognition result. Wherein in the preliminary screening process of the suspected abnormal objects, if the tracking duration of the target object is shortened, the number of the suspected abnormal objects which are wrongly identified is increased, so that accurate abnormal identification is further carried out on the basis of the contour correction mask graph of the suspected abnormal objects; according to the method, the number of misinformation caused by shortening the tracking duration is filtered out, so that the recognition accuracy of the abnormal target object can be improved under the condition of shortening the tracking duration, and the effect of considering the recognition timeliness and accuracy of the abnormal target object at the same time is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a target recognition method, a target repair network training method, and related equipment. Background Art

[0002] Currently, with the rapid development of image processing and artificial intelligence technologies, the application of intelligent farming is becoming increasingly widespread. This necessitates monitoring the health status of target animals in smart farming environments to issue alerts for any abnormalities. Specifically, by collecting images of the smart farming site, the system identifies target animals exhibiting abnormal health conditions, enabling farmers to accurately grasp the real-time status of these animals. Summary of the Invention

[0003] The purpose of this application is to provide a target recognition method, a target repair network training method, and related equipment, which can improve the recognition accuracy of abnormal target objects while shortening the tracking time, thereby achieving the effect of simultaneously taking into account the timeliness and accuracy of abnormal target object recognition.

[0004] To achieve the above technical solution, the embodiments of this application are implemented as follows:

[0005] In a first aspect, embodiments of this application provide a target recognition method, the method comprising:

[0006] Identify the target image containing a suspected abnormal object within the image to be processed;

[0007] The target image is segmented to obtain a sub-image containing the suspected abnormal object;

[0008] The suspected abnormal object in the sub-image is subjected to contour correction processing to obtain a contour correction mask image;

[0009] Based on the contour-corrected mask image and the sub-image, the suspected abnormal object is classified and identified to obtain the abnormal classification and identification result.

[0010] Secondly, this application provides a target repair network training method, the method comprising:

[0011] Obtain the training sample set of the target repair network to be trained; the training sample set includes target contour mask images corresponding to multiple key point labeled sample images respectively;

[0012] The target inpainting network is used to perform key point correction processing on the target contour mask image to obtain an intermediate convolutional feature map and a contour correction mask image of the key point labeled sample image;

[0013] Based on the intermediate convolutional feature map and the preset standard feature map, a first loss function is determined;

[0014] Based on the contour-corrected mask and the preset labeled mask, a second loss function is determined;

[0015] Based on the first loss function and the second loss function, the target repair network is iteratively updated to obtain the trained target repair network; the target repair network is used to perform contour correction processing on suspected abnormal objects in sub-images to obtain contour correction mask maps for anomaly classification and recognition, the sub-images include the image regions where the suspected abnormal objects are located in the target image, and the target image is the image containing the suspected abnormal objects in the image to be processed.

[0016] Thirdly, an embodiment of this application provides a target recognition device, the device comprising:

[0017] The abnormal target screening module is used to identify target images in the image to be processed that contain suspected abnormal objects;

[0018] The sub-image determination module is used to segment the target image to obtain a sub-image containing the suspected abnormal object;

[0019] An object contour correction module is used to perform contour correction processing on the suspected abnormal objects in the sub-image to obtain a contour correction mask image.

[0020] An abnormal target recognition module is used to perform anomaly classification and recognition on the suspected abnormal object based on the contour correction mask image and the sub-image, and obtain the anomaly classification and recognition result.

[0021] Fourthly, this application provides a target repair network training device, the device comprising:

[0022] The sample set acquisition module is used to acquire the training sample set of the target repair network to be trained; the training sample set includes target contour mask images corresponding to multiple key point labeled sample images respectively;

[0023] The key point correction module is used to perform key point correction processing on the target contour mask map using the target repair network to obtain an intermediate convolutional feature map and a contour correction mask map of the key point labeled sample image;

[0024] The first loss determination module is used to determine a first loss function based on the intermediate convolutional feature map and the preset standard feature map;

[0025] The second loss determination module is used to determine a second loss function based on the contour correction mask and the preset annotation mask;

[0026] The repair network training module is used to iteratively update the parameters of the target repair network based on the first loss function and the second loss function to obtain the trained target repair network; the target repair network is used to perform contour correction processing on suspected abnormal objects in sub-images to obtain contour correction mask maps for anomaly classification and recognition, the sub-images include the image regions where the suspected abnormal objects are located in the target image, and the target image is the image containing the suspected abnormal objects in the image to be processed.

[0027] Fifthly, an electronic device is provided in the embodiments of this application, the device comprising:

[0028] A processor; and a memory arranged to store computer-executable instructions configured to be executed by the processor, the executable instructions including steps for performing the methods described above.

[0029] Sixthly, embodiments of this application provide a computer-readable storage medium, wherein the storage medium is used to store computer-executable instructions that cause a computer to perform the steps in the method described above.

[0030] Seventhly, an embodiment of this application provides a computer program product, wherein the computer program product includes a computer program, which, when executed by a processor, implements some or all of the steps in the method described above.

[0031] As can be seen, in this embodiment, by first tracking each target object within a preset monitoring area, suspected abnormal objects are initially identified. Then, for each suspected abnormal object, a contour correction mask is segmented from the target image, and anomaly classification and recognition are performed based on the contour correction mask to obtain the final anomaly classification and recognition result. Considering that shortening the tracking time of the target object during the initial screening of suspected abnormal objects would inevitably lead to an increase in the number of false positives, further precise anomaly recognition is performed based on the contour correction mask of the suspected abnormal object, filtering out the number of false positives caused by shortening the tracking time. This improves the accuracy of identifying abnormal target objects while shortening the tracking time, thus achieving a balance between the timeliness and accuracy of abnormal target object identification. Attached Figure Description

[0032] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in one or more of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0033] Figure 1 This is a schematic diagram of a first flowchart of the target recognition method provided in the embodiments of this application;

[0034] Figure 2 This is a schematic diagram illustrating the first implementation principle of the target recognition method provided in the embodiments of this application;

[0035] Figure 3 This is a second flowchart illustrating the target recognition method provided in the embodiments of this application;

[0036] Figure 4 This is a schematic diagram illustrating the second implementation principle of the target recognition method provided in the embodiments of this application;

[0037] Figure 5 A schematic diagram illustrating the implementation principle of the contour correction mask determination process in the target recognition method provided in this application embodiment;

[0038] Figure 6a A flowchart illustrating the target repair network training method provided in this application embodiment;

[0039] Figure 6b A schematic diagram illustrating the implementation principle of the training process of the target repair network provided in this application embodiment;

[0040] Figure 7 This is a schematic diagram of the module composition of the target recognition device provided in the embodiments of this application;

[0041] Figure 8 A schematic diagram of the module composition of the target repair network training device provided in the embodiments of this application;

[0042] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0043] To enable those skilled in the art to better understand the technical solutions in one or more of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of one or more of this application, and not all embodiments. Based on the embodiments of one or more of this application, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of this application.

[0044] It should be noted that, unless otherwise specified, one or more embodiments and features described in this application can be combined with each other. The embodiments of this application will now be described in detail with reference to the accompanying drawings.

[0045] This application provides one or more embodiments of a target recognition method, a target repair network training method, and related equipment. Considering that if multiple frames of on-site captured images are used to track various target objects within a preset monitoring area, and then the tracking results of each target object are directly analyzed to determine whether the target object is an abnormal target object (e.g., the target object's current state is dead or sick), then the longer the tracking time of the target object, the higher the accuracy of identifying abnormal target objects. Therefore, if the set tracking time of the target object is relatively short, there is an inevitable possibility of misjudging normal target objects as abnormal target objects, resulting in low accuracy in identifying abnormal target objects. Conversely, if the set tracking time of the target object is relatively long, although the accuracy of identifying abnormal target objects can be ensured, the timeliness of identifying abnormal target objects cannot be guaranteed. Furthermore, considering that if a method is adopted where suspected abnormal objects are first screened out, and then anomaly confirmation is performed directly based on the texture image of the abnormal target object, since the texture image contains overlapping areas of multiple target objects (i.e., it not only contains a suspected abnormal object but also image areas of other target objects), it may interfere with the accuracy of identifying the suspected abnormal object, thus also failing to ensure the accuracy of identifying abnormal target objects.

[0046] To address the aforementioned issues, this technical solution first tracks each target object within a preset monitoring area to initially identify suspected anomalies. Then, for each suspected anomaly, a contour-corrected mask image is segmented from the target image. Anomaly classification and identification are then performed based on these contour-corrected masks to obtain the final anomaly classification and identification result. Considering that shortening the tracking time during the initial screening of suspected anomalies would inevitably increase the number of false positives, this solution further utilizes the contour-corrected masks of the suspected anomalies for precise anomaly identification. This filters out false positives caused by shortened tracking time. Furthermore, the contour mask used for precise anomaly identification is a corrected version of the target object's contour boundaries to ensure comprehensive filtering of false positives and accurate anomaly identification. Thus, while shortening the tracking time, the accuracy of anomaly target object identification can be improved, achieving a balance between timeliness and accuracy in anomaly target object identification.

[0047] in, Figure 1 This is a flowchart illustrating a target recognition method provided in one or more embodiments of this application. Figure 1 The methods described can be executed by an electronic device, which can be a terminal device or a designated server; such as Figure 1 As shown, the method includes at least the following steps:

[0048] S102, Determine that there is a target image with a suspected abnormal object in the image to be processed.

[0049] Specifically, target tracking is performed based on the image to be processed to obtain target tracking information, and target images containing suspected abnormal objects are determined based on the target tracking information. In some example embodiments, images in the image to be processed where the state of suspected abnormal objects changes can be determined as target images. In other example embodiments, images in the image to be processed where suspected abnormal objects exist and which are the most recently captured images can be determined as target images.

[0050] The images to be processed may include a first image and a second image, wherein the second image is captured later than the first image. Both the first and second images are images of the monitored site captured at different times according to a preset shooting time interval using image acquisition devices located at monitoring points within a preset monitoring area. The preset monitoring area may be a preset breeding site or other designated monitoring areas. For example, the preset breeding site may be a poultry breeding site. Images of the preset breeding site are captured at preset shooting time intervals T to obtain images containing multiple target objects within the preset breeding site. The shooting time interval T between the second image and the first image, i.e., the shooting time interval between the second shooting time of the second image and the first shooting time of the first image, is equal to T. The preset shooting time interval T is greater than a preset time threshold and less than a maximum tracking duration threshold, enabling rapid screening of suspected abnormal objects within the preset breeding site even with a relatively short tracking duration.

[0051] In practice, the shooting time interval can be constrained. Considering that if the shooting time interval is greater than the maximum tracking time threshold, it means that the shooting time interval is too long, which will affect the timeliness of identifying suspected abnormal objects, the shooting time interval can be set to be less than the maximum tracking time threshold.

[0052] In some example embodiments, target tracking is performed based on two images captured sequentially to determine the target tracking information of any target object within a preset monitoring area. Then, based on the target tracking information, suspected abnormal objects within the preset monitoring area are identified, and the image containing the suspected abnormal object is used as the target image (i.e., the second image). Specifically, during the process of capturing images of the preset monitoring area at preset shooting time intervals T, for each non-first captured current frame image, target tracking is performed on the target object within the preset monitoring area based on the current frame image (i.e., the second image) and the previous frame image (i.e., the first image) to obtain target tracking information. The target tracking information includes the coordinate information of the target object in the first image (representing the position information of the target object in the preset monitoring area at the first shooting time) and the coordinate information of the target object in the second image (representing the position information of the target object in the preset monitoring area at the second shooting time). Since the target tracking information can be combined to determine the target tracking duration and target tracking distance, if the target tracking distance is relatively small under certain constraints, it indicates that the target object may be in a pathological or resting state. Therefore, such target objects can be identified as suspected abnormal objects.

[0053] S104, the target image is segmented to obtain a sub-image containing a suspected abnormal object.

[0054] After initially identifying suspected abnormal objects, for each suspected abnormal object, a sub-image of the suspected abnormal object is extracted from the target image based on the coordinate information of the suspected abnormal object in the target image. The sub-image can be a rectangular frame sub-image containing the suspected abnormal object. In some example embodiments, a second image containing suspected abnormal objects is determined as the target image, and a sub-image of the suspected abnormal object in the second image is extracted based on the coordinate information of the suspected abnormal object in the second image.

[0055] S106, Perform contour correction processing on the suspected abnormal objects in the above sub-images to obtain a contour correction mask image.

[0056] After identifying the sub-image of each suspected anomalous object, target segmentation is performed based on the sub-image of the suspected anomalous object to obtain an initial contour mask; then contour correction is performed based on the initial contour mask of the suspected anomalous object to obtain a contour correction mask.

[0057] In some example embodiments, for the process of contour correction based on the initial contour mask, a target contour template corresponding to the initial contour mask can be determined from a preset contour template set. The initial contour mask is then corrected based on the target contour template to obtain a corrected contour mask. However, considering that the accuracy of contour correction directly affects the accuracy of the next step of anomaly classification and recognition, in order to improve the accuracy of contour correction and thus improve the accuracy of the final anomaly classification and recognition result, in other example embodiments, a neural network model can be introduced. A target repair network for contour correction is pre-trained, and the initial contour mask is input into the target repair network for contour correction to obtain a corrected contour mask of the suspected anomaly object.

[0058] S108, Based on the above contour correction mask and sub-image, perform anomaly classification and identification on the suspected abnormal object to obtain the anomaly classification and identification result.

[0059] After obtaining the contour-corrected mask image of the suspected anomalous object, the texture image of the target object (i.e., the contour-corrected texture image) can be reconstructed based on the contour-corrected mask image and its sub-images. Then, anomaly classification and recognition are performed based on the contour-corrected texture image to obtain the anomaly classification and recognition result. Since the contour-corrected mask image is a corrected contour mask image, it can more accurately represent the pose information of the target object. Furthermore, since the contour-corrected mask image is a corrected contour mask image, it can resolve the situation of overlapping target objects, so that the contour-corrected texture image only contains the image region of the suspected anomalous object. This allows more useful feature information belonging to the suspected anomalous object to be extracted during anomaly classification and recognition, while reducing the interference feature information of other target objects. Therefore, using the contour-corrected mask image as the basis for anomaly classification and recognition can further improve the accuracy of the final anomaly classification and recognition result.

[0060] For the process of anomaly classification and recognition based on contour-corrected texture images, in some example embodiments, feature extraction is directly performed on the contour-corrected texture image to obtain image feature information of suspected anomalies. Anomaly classification and recognition are then performed based on the extracted image feature information to obtain anomaly classification and recognition results. In other example embodiments, a neural network model can be introduced, and a discriminant network for anomaly classification is pre-trained. The discriminant network is then used to perform anomaly classification and recognition based on the contour-corrected texture image to obtain anomaly classification and recognition results. Specifically, the above-mentioned anomaly classification and recognition of suspected anomalies based on the contour-corrected mask image and sub-images to obtain anomaly classification and recognition results includes: performing color restoration processing on the contour-corrected mask image and sub-images of the suspected anomalies to obtain the contour-corrected texture image of the suspected anomalies; inputting the contour-corrected texture image into the discriminant network for anomaly classification and recognition to obtain the anomaly classification and recognition results of the suspected anomalies; in specific implementations, the discriminant network is used to perform feature extraction processing on the contour-corrected texture image to obtain image feature information of the suspected anomalies, and anomaly classification and recognition are then performed based on this image feature information to obtain the anomaly classification and recognition results.

[0061] In this embodiment, by first tracking each target object within a preset monitoring area, suspected abnormal objects are initially identified. Then, for each suspected abnormal object, a contour-corrected mask image is segmented from the target image, and anomaly classification and identification are performed based on the contour-corrected mask image to obtain the final anomaly classification and identification result. Considering that shortening the tracking time of target objects during the initial screening of suspected abnormal objects would inevitably lead to an increase in the number of false positives, further precise anomaly identification is performed based on the contour-corrected mask image of the suspected abnormal objects. This filters out the number of false positives caused by shortening the tracking time. Thus, while shortening the tracking time, the accuracy of identifying abnormal target objects can be improved, thereby achieving the effect of simultaneously considering the timeliness and accuracy of abnormal target object identification.

[0062] In one specific embodiment, a preset monitoring area is taken as a preset breeding site, and the preset breeding site contains N target objects as an example; for example... Figure 2 As shown, a specific implementation process of a target recognition method is presented, which mainly includes:

[0063] S21, acquire the first image at the first shooting time of the preset breeding site and the second image at the second shooting time; the second image is the latest image captured during the shooting process of the preset monitoring area according to the preset shooting time interval T, that is, the second image is the current frame image and the first image is the previous frame image of the second image.

[0064] S22, target tracking is performed based on the first and second images mentioned above to identify suspected abnormal objects in the preset breeding site, and the second image is identified as the target image; for example, among the N target objects, there are M suspected abnormal objects, where M and N are both integers greater than 1, and N is greater than M.

[0065] S23, for each suspected abnormal object, extract the sub-image of the suspected abnormal object in the second image.

[0066] S24, Perform target segmentation processing based on the sub-image of the suspected abnormal object to obtain the initial contour mask image.

[0067] S25. Perform contour correction processing on the initial contour mask image of the suspected abnormal object to obtain the contour correction mask image.

[0068] S26. Based on the contour correction mask and sub-image of the suspected abnormal object, color restoration processing is performed to obtain the contour correction texture image.

[0069] S27. Based on the above contour-corrected texture image, perform anomaly classification and recognition to obtain the anomaly classification and recognition result of the suspected abnormal object.

[0070] The process for identifying suspected abnormal objects mainly includes target tracking within a preset monitoring area and screening suspected targets based on the tracking results. For the target tracking process, to achieve target tracking based on a smaller number of images, for example, only two images, the movement orientation of the target object to be identified is predicted based on the first coordinate information. Then, coordinate matching is performed based on the predicted second coordinate information and the actual coordinate information of each target object in the second image. The target objects in the second image are associated with the target objects in the first image to obtain target tracking information for the target objects within the preset monitoring area. This eliminates the need to wait for three or more images to achieve target tracking. Specifically, the images to be processed can include the first image and the second image, with the second image being captured later than the first image. Based on this, in step S102, the images to be processed are determined to contain target images with suspected abnormal objects, specifically including:

[0071] Step A1: Input the first image into the target detection network to obtain the first coordinate information of the target object to be identified.

[0072] The target object to be identified is any aquaculture object in the preset aquaculture site. Target detection is performed based on the first image to obtain the actual coordinate information (i.e., the first coordinate information) of any target object in the preset aquaculture site in the first image. The first coordinate information is used to predict the movement direction of the target object, that is, as the basis for predicting the predicted coordinate information of the target object in the next frame image (i.e., the second image).

[0073] Step A2: Based on the first coordinate information mentioned above, predict the movement orientation of the target object to be identified, and determine the second coordinate information of the target object in the second image.

[0074] The aforementioned second coordinate information represents the predicted coordinate information of the target object to be identified in the second image. After determining the first coordinate information, the Kalman filter algorithm can be used to predict the movement orientation based on the first coordinate information to determine the predicted coordinate information (i.e., the second coordinate information) of the target object to be identified in the second image. The second coordinate information is used to match the predicted coordinates with the actual coordinates to associate the target object in the second image with the target object in the first image, and to preset the target tracking information of the target object in the monitoring area.

[0075] Step A3: Based on the second coordinate information and the second image mentioned above, target tracking is performed on the target object to be identified to obtain target tracking information.

[0076] After determining the second coordinate information, the Hungarian matching algorithm can be used to perform coordinate matching based on the second coordinate information and the actual coordinate information of each target object in the second image. Target objects whose coordinate overlap between the second coordinate information and the actual coordinate information is greater than a preset overlap threshold are associated with the corresponding target objects in the first image (the target objects corresponding to the first coordinate information used to determine the second coordinate information). Based on the first coordinate information of the associated target objects and the actual coordinate information in the second image, the target tracking information of the target object is determined.

[0077] Step A4: Based on the above target tracking information, identify suspected abnormal objects and determine the second image where the suspected abnormal object is located as the target image.

[0078] The target tracking information mentioned above may include target tracking duration and target tracking distance. If the target tracking duration is greater than the minimum tracking duration threshold but less than the maximum tracking duration threshold, and the target tracking distance is less than the preset distance threshold, it indicates that the target object may be in a sick or resting state. Therefore, such target objects can be identified as suspected abnormal objects.

[0079] Specifically, in the process of target tracking based on the second coordinate information, step A3 above, based on the second coordinate information and the second image, performs target tracking on the target object to be identified to obtain target tracking information, including:

[0080] Step A31: Based on the second coordinate information and the actual coordinate information of the target object in the second image, perform coordinate matching processing to determine the coordinate matching result of the target object to be identified.

[0081] In a specific implementation, the coordinate overlap degree is calculated between the second coordinate information and the actual coordinate information of the target object appearing in the second image. Therefore, the above coordinate matching result can include the coordinate overlap degree between the second coordinate information and the actual coordinate information.

[0082] Step A33: Based on the coordinate matching results above, determine the target tracking information of the target object to be identified.

[0083] In some example embodiments, the above-mentioned actual coordinate information may be obtained by inputting the second image into the target detection network for target detection; if the coordinate overlap between the second coordinate information and a certain actual coordinate information is greater than a preset overlap threshold, it indicates that the target object corresponding to the actual coordinate information and the target object corresponding to the second coordinate information (i.e. the target object corresponding to the first coordinate information used to determine the second coordinate information) belong to the same target object, and then the target tracking information of the target object can be determined based on the first coordinate information of the target object and the corresponding actual coordinate information.

[0084] In determining the second coordinate information, the initial predicted coordinates obtained based on the first coordinate information and the preset state transition matrix can be directly used as the second coordinate information. However, to improve the accuracy of the second coordinate information, a step of predicting the orientation based on the Kalman gain coefficient is added to improve the accuracy of target tracking. Specifically, step A2 above, based on the first coordinate information, predicts the movement orientation of the target object to be identified and determines the second coordinate information of the target object in the second image, specifically including:

[0085] Step A21: Based on the first coordinate information and the preset state transition matrix, predict the movement direction of the target object to be identified and determine the initial predicted coordinate information.

[0086] The aforementioned preset state transition matrix is ​​predetermined based on the orientation movement pattern of the target object; the first coordinate information can be represented as a first coordinate column matrix, in which each value represents (x, y, w, h). The second coordinate column matrix (i.e., the initial predicted coordinate information) can be obtained by multiplying the preset state transition matrix with the first coordinate column matrix.

[0087] Step A22: Determine the Kalman gain coefficients based on the target covariance matrix and the Kalman filtering algorithm; the target covariance matrix is ​​determined based on the state covariance matrix corresponding to the first image.

[0088] The aforementioned target covariance matrix is ​​determined using a preset covariance update formula based on the state covariance matrix corresponding to the previous frame image. If the state covariance matrix corresponding to the image obtained at the first shooting moment is known, then the preset covariance update formula can be used to determine the state covariance matrix corresponding to the image obtained at the second shooting moment, and so on, to determine the state covariance matrix (i.e., the target covariance matrix) corresponding to the second image. In specific implementation, the preset covariance update formula can be expressed as P1 = FP0F T +Q, where F represents the preset state transition matrix, P0 represents the state covariance matrix corresponding to the image obtained at the previous shooting time, and Q represents the process noise covariance matrix; next, the Kalman gain coefficient is determined based on the target covariance matrix using the preset Kalman gain coefficient determination formula; the Kalman gain coefficient is used to adjust the trade-off between the predicted and measured values ​​to minimize the mean square error between the estimated and actual values; in specific implementation, the preset Kalman gain coefficient determination formula can be expressed as K=P1H T (HP1H T +R) -1 Where P1 represents the target covariance matrix, H represents the observation matrix, and R represents the observation noise covariance matrix.

[0089] Step A23: Based on the Kalman gain coefficients mentioned above, the predicted orientation of the initial predicted coordinate information is corrected to obtain the second coordinate information of the target object to be identified in the second image.

[0090] After determining the Kalman gain coefficients, the corrected second coordinate information is calculated based on the initial predicted coordinate information and the Kalman gain coefficients using a preset prediction orientation correction formula. In specific implementation, the preset prediction orientation correction formula can be expressed as follows: in, Indicates the second coordinate information. x1 represents the initial predicted coordinate information, x1 represents the actual coordinate information of a target object in the second image, K represents the Kalman gain coefficient, and H represents the observation matrix.

[0091] Among them, after determining the target tracking information of the target object to be identified, the step of identifying whether the target object is a suspected abnormal object based on the target tracking information, as described in step A4 above, specifically includes:

[0092] Step A41: Based on the target tracking information above, determine the target tracking duration and target tracking distance of the target object to be identified.

[0093] In some example embodiments, if the shooting time interval is greater than or equal to the minimum tracking duration threshold, the target tracking duration is equal to the shooting time interval, and the target tracking distance is equal to the distance the target object moves from the first shooting time to the second shooting time. In other example embodiments, if the shooting time interval is less than the minimum tracking duration threshold, then the shooting time interval T < minimum tracking duration < target tracking duration < maximum tracking duration, i.e., the target tracking duration is greater than the shooting time interval. Therefore, the historical tracking information of the target object to be identified can be combined to determine whether the target object is a suspected abnormal object. That is, based on the target tracking information and the historical tracking information of the target object to be identified, the target tracking duration and target tracking distance of the target object to be identified are determined. Specifically, the actual coordinate information of the target object to be identified in the specified historical image is determined based on the historical tracking information, and the target tracking duration is equal to the shooting interval between the specified historical image and the second image. Based on the actual coordinate information of the specified historical image and the actual coordinate information of the target object to be identified in the second image, the target tracking distance (i.e., the actual distance the target object to be identified moves from the specified historical image to the second image during the shooting period) is determined.

[0094] Step A42: Identify the target objects to be identified as suspected abnormal objects if the target tracking time is greater than the first threshold and less than the second threshold, and the target tracking distance is less than the preset distance threshold.

[0095] In practical implementation, the target tracking duration can be constrained to further ensure the timeliness and accuracy of abnormal target object identification. The first threshold and the second threshold are determined based on preset requirements for the timeliness and accuracy of abnormal target object identification. The first threshold is related to the preset requirements for the accuracy of abnormal target object identification, and the second threshold is related to the preset requirements for the timeliness of abnormal target object identification. Therefore, setting the target tracking duration between the first threshold and the second threshold can further ensure the timeliness and accuracy of abnormal target object identification. The first threshold can be understood as the minimum tracking duration threshold, and the second threshold can be understood as the maximum tracking duration threshold.

[0096] To improve the accuracy of the contour-corrected mask image for suspected abnormal objects, a pre-trained target inpainting network can be used to perform contour correction processing on the initial contour mask image. Specifically, for example... Figure 3 As shown, in step S106 above, contour correction processing is performed on the suspected abnormal objects in the above sub-image to obtain a contour correction mask image, specifically including:

[0097] S1062, the sub-image of the suspected abnormal object is input into the target segmentation network for target contour extraction processing to obtain the initial mask image of the suspected abnormal object contour.

[0098] S1064, The above initial contour mask image is input into the target repair network for contour correction processing to obtain the contour correction mask image of the suspected abnormal object.

[0099] Both the target segmentation network and the target restoration network described above are pre-trained based on sample data. The target segmentation network can be implemented using an existing target segmentation network, while the target restoration network is obtained by modifying an existing target segmentation network. The target restoration network can be a neural network that adds a keypoint correction structure to the encoder-decoder structure. In some example embodiments, considering that the encoder may include multiple convolutional layers for encoding (i.e., first convolutional layers) and the decoder may include multiple convolutional layers for decoding (i.e., second convolutional layers), the target restoration network may include multiple cascaded first convolutional layers, multiple cascaded second convolutional layers, and at least one keypoint correction layer, wherein the keypoint... The correction layer is used to adjust the contour of the target object based on the position of preset key points in the feature map. The preset key points correspond to the key components of the target object. The contour correction mask map is obtained by feature encoding, feature decoding and at least one key point correction process based on the initial contour mask map. The number of key point correction layers and the connection relationship between the key point correction layers and the first and second convolutional layers can be set according to actual needs. For example, a key point correction layer can be added between the last first convolutional layer and the first second convolutional layer, and / or a key point correction layer can be added after the last second convolutional layer; or a key point correction layer can be added after each first convolutional layer, and / or a key point correction layer can be added after each second convolutional layer.

[0100] Furthermore, it should be noted that since the contour correction process is based on the initial contour mask, which is a black and white image and is not affected by image color and texture, it can solve the problem of low contour correction accuracy (such as incomplete or incorrect segmentation) caused by generalization factors such as color differences in images captured by different cameras and different target object types. For example, if the sample data used in the training phase of the target restoration network contains sample images of target objects of type A, then in the model application phase, i.e., during the contour correction process, for target objects of type B that are suspected of being abnormal, the contour correction accuracy of the target restoration network will not decrease due to the change in the target object type. Moreover, by adding the target restoration network, not only can the overlapping of target objects be resolved, ensuring that the contour-corrected texture image only contains the image region of the suspected abnormal object, thus enabling the extraction of more useful feature information belonging to the suspected abnormal object during anomaly classification and recognition, reducing the interference feature information of other target objects, but also further improving the contour correction accuracy.

[0101] In one specific embodiment, taking a pre-set aquaculture site containing N target objects as an example, and to further improve the accuracy of identifying abnormal target objects, the identification process utilizes multiple neural networks in cooperation. The model used for identifying abnormal target objects may include a target detection network, a target segmentation network, a target repair network, and a target anomaly discrimination network; such as Figure 4 As shown, a specific implementation process of a target recognition method is presented, which mainly includes:

[0102] S41, acquire the first image at the first shooting time and the second image at the second shooting time of the preset breeding site.

[0103] S42, the first image is input into the target detection network to obtain the first coordinate information of the target object to be identified.

[0104] S43, based on the first coordinate information mentioned above, perform motion orientation prediction to determine the second coordinate information of the target object to be identified in the second image.

[0105] S44. Based on the second coordinate information and the second image mentioned above, target tracking is performed on the target object to be identified to obtain target tracking information.

[0106] S45, based on the above target tracking information, identify suspected abnormal objects in the preset breeding site, and determine the second image as the target image; for example, among the N target objects, there are M suspected abnormal objects, where M and N are both integers greater than 1, and N is greater than M.

[0107] S46. For each suspected abnormal object, extract the sub-image of the suspected abnormal object in the second image, and input the sub-image of the suspected abnormal object into the target segmentation network for target segmentation processing to obtain the initial contour mask image.

[0108] S47. Input the initial mask image of the suspected abnormal object into the target repair network for contour correction processing to obtain the contour correction mask image.

[0109] S48. Color restoration processing is performed on the contour-corrected mask and sub-image of the suspected abnormal object to obtain the contour-corrected texture image.

[0110] S49, the above contour-corrected texture image is input into the discriminant network for anomaly classification and recognition, and the anomaly classification and recognition result of the suspected abnormal object is obtained.

[0111] Specifically, in the contour correction process, keypoint correction processing can be added to the feature processing based on the initial contour mask. For example, keypoint correction processing can be added in both the feature encoding and feature decoding stages to obtain a keypoint-corrected feature map (i.e., the fourth feature map). Then, the fourth feature map is converted into a contour correction mask. Specifically, in S1064 above, the initial contour mask is input into the target repair network for contour correction processing to obtain the contour correction mask of the suspected abnormal object, which specifically includes:

[0112] Step B1: Perform feature encoding on the initial mask image of the suspected abnormal object to obtain the first feature map; the feature encoding process is detailed in the existing downsampling feature processing process, and will not be repeated here.

[0113] Step B2 involves performing keypoint correction processing on the first feature map to obtain the second feature map. Specifically, the process of determining the second feature map includes: performing grouped convolution processing on the first feature map to obtain multiple discrete features; performing feature concatenation processing on the multiple discrete features to obtain concatenated features; performing convolution processing on the concatenated features to obtain an intermediate convolutional feature map; and generating the second feature map based on the first feature map and the intermediate convolutional feature map. The intermediate convolutional feature map can be a one-channel feature map obtained by convolution processing using a 3x3 convolution kernel.

[0114] Specifically, to reduce the computational load of the network and achieve further feature extraction, the first feature map is first subjected to grouped convolution operations. The first feature map is split into multiple groups according to channels, with each group corresponding to a convolution kernel. After the convolution operation, several discrete features for a preset number of channels are obtained. For example, if the size of the first feature map is (112, 112, 56), it is split into 56 groups, each group corresponding to a discrete feature with a size of (112, 112, 1). Then, the convolution results of the 56 groups (i.e., the discrete features) are concatenated. First, the concatenated feature map is obtained, with a size of (112, 112, 56). Next, the concatenated feature map is convolved with one convolution kernel to obtain an intermediate convolutional feature map, with a size of (112, 112, 1). Then, the first feature map is multiplied by the intermediate convolutional feature map to obtain a fused feature map, with a size of (112, 112, 56). Finally, the fused feature map is added to the first feature map to obtain a second feature map, with a size of (112, 112, 56).

[0115] Step B3: Perform feature decoding on the second feature map to obtain the third feature map; the feature decoding process is detailed in the existing upsampling feature processing process, and will not be repeated here.

[0116] Step B4 involves performing keypoint correction processing on the third feature map to obtain the fourth feature map. Specifically, the process of determining the fourth feature map includes: performing grouped convolution processing on the third feature map to obtain multiple discrete features; performing feature concatenation processing on the multiple discrete features to obtain concatenated features; performing convolution processing on the concatenated features to obtain an intermediate convolutional feature map; and generating the fourth feature map based on the aforementioned third feature map and intermediate convolutional feature map. The process of determining the fourth feature map can be referred to the process of determining the second feature map described above, and will not be repeated here.

[0117] Step B5: Based on the fourth feature map mentioned above, determine the contour correction mask map of the suspected abnormal object.

[0118] In this embodiment, the contour correction mask is determined based on the target feature map (i.e., the fourth feature map). The target feature map is obtained by performing feature encoding, key point correction processing, feature decoding, and key point correction processing on the initial contour mask. In other words, the process of determining the contour correction mask based on the initial contour mask involves at least two feature processing steps (feature encoding and feature decoding) and at least two key point correction processing steps. The feature processing and key point correction processing are performed alternately, thereby improving the accuracy of contour correction for suspected abnormal objects.

[0119] Specifically, the aforementioned target repair network may include an encoder and a decoder. Considering that the feature encoding process may include multiple convolutional processes, keypoint correction processing can be added after any convolutional process, or after each convolutional process. Similarly, considering that the feature decoding process may also include multiple convolutional processes, keypoint correction processing can be added after any convolutional process, or after each convolutional process. Therefore, the second feature map output by the encoder is obtained by performing n convolutional processes and m keypoint correction processes based on the initial contour mask map, where m and n are both positive integers and m is less than or equal to n; the decoder... The output fourth feature map is obtained by performing j convolutional processes and I keypoint correction processes based on the second feature map output by the encoder, where m and n are both positive integers and m is less than or equal to n; that is, the number of first keypoint correction layers in the encoder is less than or equal to the number of first convolutional layers, and the input of the first keypoint correction layer is the output of the first convolutional layer, and the output of the first keypoint correction layer is the input of another first convolutional layer or the first second convolutional layer. The number of second keypoint correction layers in the decoder is less than or equal to the number of second convolutional layers, and the input of the second keypoint correction layer is the output of the second convolutional layer, and the output of the second keypoint correction layer is another second convolutional layer or the fourth feature map.

[0120] In some example embodiments, if a keypoint correction layer is added between the last first convolutional layer and the first second convolutional layer, and a keypoint correction layer is added after the last second convolutional layer, then m equals 1 and I equals 1; in other example embodiments, if a keypoint correction layer is added after each first convolutional layer, and a keypoint correction layer is added after each second convolutional layer, then m equals n and I equals j.

[0121] In specific implementation, for adding keypoint correction processing after each convolutional process, i.e., m equals n and I equals j, that is, adding a keypoint correction layer after each first convolutional layer and after each second convolutional layer, specifically, the above-mentioned target restoration network may include an encoder and a decoder. The encoder includes multiple encoding units connected in sequence, each encoding unit including a first convolutional layer and a first keypoint correction layer connected in sequence. The decoder includes multiple decoding units connected in sequence, each decoding unit including a second convolutional layer and a second keypoint correction layer connected in sequence. The convolutional layer is used to perform convolution processing based on the feature map output by the previous keypoint correction layer, and the keypoint correction processing is used to perform keypoint correction processing based on the feature map output by the previous convolutional layer. Multiple first convolutional layers are used for downsampling, and multiple second convolutional layers are used for upsampling. Correspondingly, in S1064 above, the above-mentioned initial contour mask image is input into the target restoration network for contour correction processing to obtain the contour correction mask image of the suspected abnormal object, specifically including:

[0122] The initial mask image of the suspected abnormal object is input into the first convolutional layer for convolution processing to obtain the first feature map; the feature is encoded by using downsampling through multiple first convolutional layers.

[0123] The first feature map is input into the first keypoint correction layer for keypoint correction processing to obtain the second feature map; the first feature map output by the previous first convolutional layer is processed by the first keypoint correction layer to obtain the second feature map.

[0124] The second feature map output by the last first keypoint correction layer in the encoder is input into the first second convolutional layer for convolution processing to obtain the third feature map; feature decoding is achieved by performing convolution processing through multiple second convolutional layers and using an upsampling method.

[0125] The third feature map is input into the first second keypoint correction layer for keypoint correction processing to obtain the fourth feature map; the third feature map output by the previous second convolutional layer is processed by the second keypoint correction layer to obtain the fourth feature map.

[0126] Based on the fourth feature map output by the last second keypoint correction layer in the decoder, the contour correction mask map of the suspected abnormal object is determined.

[0127] In one specific embodiment, taking a target restoration network containing 3 encoding units and 3 decoding units, with each encoding unit containing a first keypoint correction layer and each decoding unit containing a second keypoint correction layer as an example, that is, the number of the first convolutional layer, the second convolutional layer, the first keypoint correction layer, and the second keypoint correction layer are all equal to 3. For example, the target restoration network includes encoding unit 1, encoding unit 2, encoding unit 3, decoding unit 1, decoding unit 2, and decoding unit 3. Encoding unit 1 includes a first convolutional layer 1 and a first keypoint correction layer 1; encoding unit 2 includes a first convolutional layer 2 and a first keypoint correction layer 2; encoding unit 3 includes a first convolutional layer 1 and a first keypoint correction layer 3; decoding unit 1 includes a second convolutional layer 1 and a second keypoint correction layer 1; decoding unit 2 includes a second convolutional layer 2 and a second keypoint correction layer 2; decoding unit 3 includes a second convolutional layer 3 and a second keypoint correction layer 3. Figure 5 As shown, the specific implementation process of determining the contour correction mask in the target recognition method is given, mainly including:

[0128] The initial mask image of the suspected abnormal object is input into the first convolutional layer 1 for convolution processing to obtain the first feature. Figure 1 .

[0129] The first feature Figure 1 The input is fed into the first keypoint correction layer 1 for keypoint correction processing to obtain the second feature. Figure 1 .

[0130] The second feature Figure 1 The input is fed into the first convolutional layer 2 for convolution processing to obtain the first feature. Figure 2 .

[0131] The first feature Figure 2 The input is fed into the first keypoint correction layer 2 for keypoint correction processing to obtain the second feature. Figure 2 .

[0132] The second feature Figure 2 The input is fed into the first convolutional layer 3 for convolution processing to obtain the first feature. Figure 3 .

[0133] The first feature Figure 3 The input is fed into the first keypoint correction layer 3 for keypoint correction processing to obtain the second feature. Figure 3 .

[0134] The second feature Figure 3 The input is fed into the second convolutional layer 1 for convolution processing to obtain the third feature. Figure 1 .

[0135] The third feature Figure 1 The input is fed into the second keypoint correction layer 1 for keypoint correction processing to obtain the fourth feature. Figure 1 .

[0136] The fourth feature Figure 1 The input is fed into the second convolutional layer 2 for convolution processing to obtain the third feature. Figure 2 .

[0137] The third feature Figure 2 The input is fed into the second keypoint correction layer 2 for keypoint correction processing to obtain the fourth feature. Figure 2 .

[0138] The fourth feature Figure 2 The input is fed into the second convolutional layer 3 for convolution processing to obtain the third feature. Figure 3 .

[0139] The third feature Figure 3 The input is fed into the second keypoint correction layer 3 for keypoint correction processing to obtain the fourth feature. Figure 3 .

[0140] Based on the aforementioned fourth feature Figure 3 The outline correction mask of the above suspected abnormal objects is determined.

[0141] In this regard, considering the correction process for the initial contour mask, to improve the accuracy of the contour-corrected mask used for anomaly classification, a target restoration network can be introduced for contour correction. Therefore, the parameters of the target restoration network need to be pre-trained. Regarding the training process of the aforementioned target restoration network, as follows... Figure 6a As shown, the main steps include the following:

[0142] S602, Obtain the training sample set of the target repair network to be trained; the training sample set includes the target contour mask map corresponding to multiple key point labeled sample images.

[0143] The aforementioned key point annotation sample image includes at least one sample target object. The target contour mask is the initial contour mask of any sample target object in the key point annotation sample image. The target contour mask can be obtained by extracting the target contour based on the key point annotation sample image using target segmentation processing.

[0144] S604 utilizes a target inpainting network to perform key point correction processing on the target contour mask image, obtaining the contour correction mask image of the intermediate convolutional feature map and the key point annotation sample image.

[0145] Considering that the target restoration network may include feature encoding and feature decoding, in some example embodiments, keypoint correction may be introduced only in either the feature encoding or feature decoding stage. Therefore, the intermediate convolutional feature map may include a first intermediate convolutional feature map or a second intermediate convolutional feature map. To improve the accuracy of the total loss function, thereby improving the accuracy of subsequent contour correction using the target restoration network, in other example embodiments, keypoint correction is introduced in both the feature encoding and feature decoding stages. For the loss term resulting from each keypoint correction process, the intermediate convolutional feature map may include a first intermediate convolutional feature map and a second intermediate convolutional feature map. Based on this, in S604, the target restoration network is used to perform keypoint correction processing on the target contour mask image to obtain the intermediate convolutional feature map and the contour correction mask image of the keypoint-annotated sample image, specifically including:

[0146] Feature encoding is performed on the target contour mask image of the suspected abnormal object to obtain the fifth feature map; key point correction processing is performed on the fifth feature map to obtain the first intermediate convolutional feature map and the sixth feature map; the sixth feature map is determined based on the first intermediate convolutional feature map; feature decoding is performed on the sixth feature map to obtain the seventh feature map; key point correction processing is performed on the seventh feature map to obtain the second intermediate convolutional feature map and the eighth feature map; the eighth feature map is determined based on the second intermediate convolutional feature map; based on the eighth feature map, the contour correction mask image of the key point annotation sample image is determined.

[0147] It should be noted that the process for determining the intermediate convolutional feature map can refer to the specific implementation process given in the above-described target recognition method embodiments. For example, the process for determining the intermediate convolutional feature map involved in determining the second or fourth feature map can be referred to. The process for determining the contour correction mask map can refer to the specific implementation process given in the above-described target recognition method embodiments. For example, the process for determining the contour correction mask map of a suspected abnormal object can be referred to, and will not be repeated here.

[0148] S606. Based on the above intermediate convolutional feature maps and preset standard feature maps, determine the first loss function.

[0149] Similarly, considering that the target restoration network may include feature encoding and feature decoding, in some example embodiments, keypoint correction may be introduced only in either the feature encoding or feature decoding stage. In this case, the first loss function (i.e., keypoint correction loss) is determined based on either the first sub-loss function or the second sub-loss function. To improve the accuracy of the total loss function, thereby improving the accuracy of subsequent contour correction using the target restoration network, in other example embodiments, keypoint correction is introduced in both the feature encoding and feature decoding stages. In this case, the first loss function is determined based on the first and second sub-loss functions. Specifically, the first sub-loss function is determined based on the first intermediate convolutional feature map and the first standard feature map. The first intermediate convolutional feature map is obtained by the target restoration network through keypoint correction processing based on the target contour mask map. The second sub-loss function is determined based on the second intermediate convolutional feature map and the second standard feature map. The second intermediate convolutional feature map is obtained by the target restoration network through keypoint correction processing based on the first intermediate convolutional feature map.

[0150] S608, Based on the contour correction mask and the preset annotation mask of the above key point labeled sample image, determine the second loss function.

[0151] In some example embodiments, the contour correction mask map described above may be determined based on the eighth feature map output by the last second keypoint correction layer in the target repair network, which is determined based on the second intermediate convolutional feature map obtained from the last second keypoint correction layer.

[0152] S610, based on the first and second loss functions mentioned above, the parameters of the target inpainting network are iteratively updated to obtain the trained target inpainting network; wherein, the target inpainting network is used to perform contour correction processing on suspected abnormal objects in the sub-image to obtain a contour correction mask map for anomaly classification and recognition, the sub-image includes the image region where the suspected abnormal object is located in the target image, and the target image is the image containing the suspected abnormal object in the image to be processed; the process of performing contour correction processing on the suspected abnormal object can be referred to the above embodiment, and will not be repeated here.

[0153] In this embodiment, the total loss function of the target restoration network is first determined, and then the parameters of the target restoration network are updated based on the total loss function. After multiple rounds of parameter updates, the target restoration network for contour correction (i.e., the trained target restoration network) can be obtained. The total loss function of the target restoration network is related to the keypoint correction loss and the contour mask map loss. The keypoint correction loss is calculated based on the intermediate convolutional feature map output by the intermediate network unit of the keypoint correction layer and the standard feature map of the sample target object. The contour mask map loss is calculated based on the contour correction mask map of the sample target object and the preset labeled mask map.

[0154] In some example embodiments, keypoint correction may be introduced only in either the feature encoding or feature decoding stage. In this case, the keypoint correction loss is determined based on either the first sub-loss function or the second sub-loss function. To improve the accuracy of the total loss function and thus the accuracy of subsequent contour correction using the target restoration network, in other example embodiments, keypoint correction is introduced in both the feature encoding and feature decoding stages. In this case, the keypoint correction loss is determined based on the first and second sub-loss functions. The first sub-loss function is calculated by performing loss calculations based on the first intermediate convolutional feature map output by the intermediate network unit of the first keypoint correction layer in the target restoration network and the first standard feature map of the sample target object. The second sub-loss function is calculated by performing loss calculations based on the second intermediate convolutional feature map output by the intermediate network unit of the second keypoint correction layer in the target restoration network and the second standard feature map of the sample target object. The first and second standard feature maps may be the same or different.

[0155] In practical implementation, the aforementioned first sub-loss function or second sub-loss function can be expressed as:

[0156]

[0157] Among them, the standard feature map represents M i (x,y) represents the intermediate convolutional feature map output by the intermediate network unit of the keypoint correction layer. The size of both the standard feature map and the intermediate convolutional feature map is (w,h,1), where i represents the index of the keypoint correction layer, and w and h represent the width and height of the feature map.

[0158] Additionally, it should be noted that, considering the dimensions (w, h, 1) of the first and second standard feature maps, and the dimensions (w, h, c) of the final output of the keypoint correction layer (i.e., the sixth or eighth feature map), where c represents the number of channels, the loss calculation is performed based on the output of the intermediate network units of the keypoint correction layer (i.e., the first or second intermediate convolutional feature map), rather than on the final output of the keypoint correction layer (i.e., the sixth or eighth feature map).

[0159] In some example embodiments, the target restoration network described above may include an encoder and a decoder. The encoder includes multiple encoding units, each encoding unit including a first convolutional layer and a first keypoint correction layer. The decoder includes multiple decoding units, each decoding unit including a second convolutional layer and a second keypoint correction layer. Correspondingly, the training process of the target restoration network mainly includes:

[0160] Step 1: Obtain the training sample set for the target repair network to be trained; the training sample set includes target contour mask images corresponding to multiple key point labeled sample images.

[0161] Step 2: Determine the first sub-loss function based on the first intermediate convolutional feature map and the first standard feature map output by the intermediate network units in each first keypoint correction layer; the aforementioned first intermediate convolutional feature map is determined based on the target contour mask map;

[0162] Step 3: Determine the second sub-loss function based on the second intermediate convolutional feature map and the second standard feature map output by the intermediate network unit in each second keypoint correction layer; the second intermediate convolutional feature map is determined based on the sixth feature map output by the last first keypoint correction layer in the encoder, and the sixth feature map is determined based on the first intermediate convolutional feature map output by the intermediate network unit in the last first keypoint correction layer.

[0163] Step 4: Based on the contour correction mask and the preset annotation mask of the above key point labeled sample image, determine the second loss function; the above contour correction mask is determined based on the eighth feature map output by the last second key point correction layer, and the eighth feature map is determined based on the second intermediate convolutional feature map output by the intermediate network unit in the last second key point correction layer.

[0164] Step 5: Based on the first sub-loss function, the second sub-loss function, and the second loss function mentioned above, the parameters of the target repair network are iteratively updated to obtain the trained target repair network. Specifically, the first sub-loss function, the second sub-loss function, and the second loss function are weighted and summed to obtain the total loss function. The parameters of the target repair network are iteratively updated based on the total loss function to obtain the trained target repair network.

[0165] The first intermediate convolutional feature map, the second intermediate convolutional feature map, and the contour-corrected mask map used to determine the loss function are all obtained during the process of determining the contour-corrected mask map of the key point annotation sample image based on the target contour mask map. Specifically, the above method also includes:

[0166] The target contour mask image above is subjected to feature encoding to obtain the fifth feature image; the feature encoding process is detailed in the existing downsampling feature processing process, and will not be repeated here.

[0167] The fifth feature map is subjected to keypoint correction processing to obtain the sixth feature map. Specifically, the process of determining the sixth feature map includes: performing grouped convolution processing on the fifth feature map to obtain multiple discrete features; performing feature concatenation processing on the multiple discrete features to obtain concatenated features; performing convolution processing on the concatenated features to obtain a first intermediate convolutional feature map; and generating the sixth feature map based on the fifth feature map and the first intermediate convolutional feature map. The first intermediate convolutional feature map is the intermediate convolutional feature map used to determine the first sub-loss function, i.e., the intermediate convolutional feature map output by the intermediate network unit in the first keypoint correction layer. The process of determining the sixth feature map can be referred to the process of determining the second feature map described above, and will not be repeated here.

[0168] The sixth feature map is then decoded to obtain the seventh feature map. The process of feature decoding is detailed in the existing upsampling feature processing documentation and will not be repeated here.

[0169] The seventh feature map is subjected to keypoint correction processing to obtain the eighth feature map. Specifically, the process of determining the eighth feature map includes: performing grouped convolution processing on the seventh feature map to obtain multiple discrete features; performing feature concatenation processing on the multiple discrete features to obtain concatenated features; performing convolution processing on the concatenated features to obtain a second intermediate convolutional feature map; and generating the eighth feature map based on the seventh feature map and the second intermediate convolutional feature map. The second intermediate convolutional feature map is the intermediate convolutional feature map used to determine the second sub-loss function, i.e., the intermediate convolutional feature map output by the intermediate network unit in the second keypoint correction layer. The process of determining the eighth feature map can be referred to the process of determining the fourth feature map, and will not be repeated here.

[0170] Based on the eighth feature map mentioned above, the contour correction mask map of the key point annotation sample image is determined.

[0171] In some example embodiments, if the target restoration network includes an encoder and a decoder, the encoder includes multiple encoding units connected in sequence, each encoding unit including a first convolutional layer and a first keypoint correction layer connected in sequence, and the decoder includes multiple decoding units connected in sequence, each decoding unit including a second convolutional layer and a second keypoint correction layer connected in sequence, then the input of the first keypoint correction layer is the feature map output by the first convolutional layer belonging to the same encoding unit, and the input of the first convolutional layer is the target contour mask or the feature map output by the first keypoint correction layer in the previous encoding unit; the input of the second keypoint correction layer is the feature map output by the second convolutional layer belonging to the same decoding unit, and the input of the second convolutional layer is the output of the encoder or the feature map output by the second keypoint correction layer in the previous decoding unit; correspondingly, the process of determining the contour correction mask of the target object in the above-mentioned keypoint annotation sample image specifically includes:

[0172] The target contour mask of the target object in the key point annotated sample image is input into the first convolutional layer for convolution processing to obtain the fifth feature map.

[0173] The fifth feature map is input into the first keypoint correction layer for keypoint correction processing to obtain the sixth feature map. The sixth feature map is determined by the first intermediate convolutional feature map output by the intermediate network layer of the first keypoint correction layer. The first intermediate convolutional feature map is obtained by performing group convolution processing, feature concatenation processing, and convolution processing on the fifth feature map. Specifically, the first intermediate convolutional feature map is obtained by the first keypoint correction layer performing keypoint correction processing on the fifth feature map output by the first convolutional layer. The fifth feature map is obtained by feature processing based on the target contour mask map.

[0174] The sixth feature map output by the last first keypoint correction layer in the encoder is input into the first second convolutional layer for convolution processing to obtain the seventh feature map.

[0175] The seventh feature map is input into the first second keypoint correction layer for keypoint correction processing, resulting in the eighth feature map. The eighth feature map is determined by the second intermediate convolutional feature map output by the intermediate network layer of the second keypoint correction layer. This second intermediate convolutional feature map is obtained by performing group convolution, feature concatenation, and convolution processing on the seventh feature map. Specifically, the second intermediate convolutional feature map is obtained by the second keypoint correction layer performing keypoint correction processing on the seventh feature map output by the second convolutional layer. The seventh feature map is determined by the sixth feature map output by the last first keypoint correction layer, and the sixth feature map is determined by the first intermediate convolutional feature map obtained from the last first keypoint correction layer.

[0176] Based on the eighth feature map output by the last second keypoint correction layer, a contour correction mask map of the target object in the keypoint-annotated sample image is determined. The contour correction mask map is determined based on the eighth feature map output by the last second keypoint correction layer, and the eighth feature map is determined based on the second intermediate convolutional feature map obtained from the last second keypoint correction layer.

[0177] In one specific embodiment, taking the example that the target inpainting network contains 3 encoding units and 3 decoding units, and that each encoding unit contains a first keypoint correction layer and each decoding unit contains a second keypoint correction layer, the training process of the target inpainting network is as follows: Figure 6b As shown, a specific implementation process of a target recognition method is presented, which mainly includes:

[0178] The initial mask image of the target object's contour is input into the first convolutional layer 1 for convolution processing to obtain the fifth feature. Figure 1 The initial mask image is the target contour mask image mentioned above.

[0179] The fifth feature Figure 1 The input is fed into the first keypoint correction layer 1 for keypoint correction processing to obtain the sixth feature. Figure 1 The first keypoint correction layer 1 is based on the fifth feature. Figure 1 The first intermediate convolutional feature is obtained by performing group convolution, feature concatenation, and convolution. Figure 1 (i.e., a one-channel feature map), and then based on the fifth feature Figure 1 and the first fusion feature Figure 1 (i.e., the fifth characteristic) Figure 1 With the first intermediate convolution feature Figure 1 The sum of the feature maps obtained by performing dot product determines the sixth feature. Figure 1 .

[0180] The sixth feature Figure 1 The input is fed into the first convolutional layer 2 for convolution processing to obtain the fifth feature. Figure 2 .

[0181] The fifth feature Figure 2 The input is fed into the first keypoint correction layer 2 for keypoint correction processing to obtain the sixth feature. Figure 2 The first keypoint correction layer 2 is based on the fifth feature. Figure 2 The first intermediate convolutional feature is obtained by performing group convolution, feature concatenation, and convolution. Figure 2 (i.e., a one-channel feature map), and then based on the fifth feature Figure 2 and the first fusion feature Figure 2 (i.e., the fifth characteristic) Figure 2 With the first intermediate convolution feature Figure 2 The sum of the feature maps obtained by performing dot product determines the sixth feature. Figure 2 .

[0182] The sixth feature Figure 2 The input is fed into the first convolutional layer 3 for convolution processing to obtain the fifth feature. Figure 3 .

[0183] The fifth feature Figure 3 The input is fed into the first keypoint correction layer 3 for keypoint correction processing to obtain the sixth feature. Figure 3 The first keypoint correction layer 3 is based on the fifth feature. Figure 3 The first intermediate convolutional feature is obtained by performing group convolution, feature concatenation, and convolution. Figure 3 (i.e., a one-channel feature map), and then based on the fifth feature Figure 3 and the first fusion feature Figure 3 (i.e., the fifth characteristic) Figure 3 With the first intermediate convolution feature Figure 3 The sum of the feature maps obtained by performing dot product determines the sixth feature. Figure 3 .

[0184] The sixth feature Figure 3 The input is fed into the second convolutional layer 1 for convolution processing to obtain the seventh feature. Figure 1 .

[0185] The seventh feature Figure 1 The input is fed into the second keypoint correction layer 1 for keypoint correction processing to obtain the eighth feature. Figure 1 The second keypoint correction layer 1 is based on the seventh feature. Figure 1 Group convolution, feature concatenation, and convolution are performed to obtain the second intermediate convolutional feature. Figure 1 (i.e., a one-channel feature map), and then based on the seventh feature Figure 1 Second fusion features Figure 1 (i.e., the seventh characteristic) Figure 1 With the second intermediate convolution feature Figure 1 The sum of the feature maps obtained by performing dot product determines the eighth feature. Figure 1 .

[0186] The eighth feature Figure 1 The input is fed into the second convolutional layer 2 for convolution processing to obtain the seventh feature. Figure 2 .

[0187] The seventh feature Figure 2 The input is fed into the second keypoint correction layer 2 for keypoint correction processing to obtain the eighth feature. Figure 2 The second keypoint correction layer 2 is based on the seventh feature. Figure 2 Group convolution, feature concatenation, and convolution are performed to obtain the second intermediate convolutional feature. Figure 2(i.e., a one-channel feature map), and then based on the seventh feature Figure 2 Second fusion features Figure 2 (i.e., the seventh characteristic) Figure 2 With the second intermediate convolution feature Figure 2 The sum of the feature maps obtained by performing dot product determines the eighth feature. Figure 2 .

[0188] The eighth feature Figure 2 The input is fed into the second convolutional layer 3 for convolution processing to obtain the seventh feature. Figure 3 .

[0189] The seventh feature Figure 3 The input is fed into the second keypoint correction layer 3 for keypoint correction processing to obtain the eighth feature. Figure 3 The second keypoint correction layer 3 is based on the seventh feature. Figure 3 Group convolution, feature concatenation, and convolution are performed to obtain the second intermediate convolutional feature. Figure 3 (i.e., a one-channel feature map), and then based on the seventh feature Figure 3 Second fusion features Figure 3 (i.e., the seventh characteristic) Figure 3 With the second intermediate convolution feature Figure 3 The sum of the feature maps obtained by performing dot product determines the eighth feature. Figure 3 .

[0190] Based on the aforementioned eighth feature Figure 3 Determine the contour correction mask of the above sample target object;

[0191] Based on the aforementioned first intermediate convolutional features Figure 1 First intermediate convolutional features Figure 2 First intermediate convolutional features Figure 3 Based on the standard feature maps of the target object and the sample, the first sub-loss function is determined. The first sub-loss function includes a first loss term, a second loss term, and a third loss term. The first loss term is based on the first intermediate convolutional features. Figure 1 The loss term is obtained by calculating the loss using the standard feature map of the target object in the sample. The second loss term is based on the first intermediate convolutional feature. Figure 2 The loss term is obtained by calculating the loss using the standard feature map of the target object in the sample. The third loss term is based on the first intermediate convolutional feature. Figure 3 The loss is calculated by combining the standard feature map of the target object in the sample.

[0192] Based on the aforementioned second intermediate convolution feature Figure 1 Second intermediate convolution features Figure 2 Second intermediate convolution features Figure 3Based on the standard feature maps of the target sample, the second sub-loss function is determined; the second sub-loss function includes a fourth loss term, a fifth loss term, and a sixth loss term, the fourth loss term being based on the second intermediate convolutional features. Figure 1 The loss term is calculated by performing loss calculations on the standard feature maps of the target object and the sample. The fifth loss term is based on the second intermediate convolutional features. Figure 2 The loss term is calculated by performing loss calculations on the standard feature maps of the target sample objects. The sixth loss term is based on the second intermediate convolutional features. Figure 3 The loss is calculated by combining the standard feature map of the target object in the sample.

[0193] Based on the contour correction mask and the preset annotation mask of the target object in the above sample, the second loss function is determined; based on the first sub-loss function, the second sub-loss function and the second loss function, the total loss function is determined.

[0194] Based on the above total loss function, the parameters of the target repair network are iteratively updated to obtain the trained target repair network.

[0195] The target recognition method in this embodiment first tracks each target object within a preset monitoring area to initially identify suspected abnormal objects. Then, for each suspected abnormal object, a contour-corrected mask image is segmented from the target image, and anomaly classification and recognition are performed based on the contour-corrected mask image to obtain the final anomaly classification and recognition result. Considering that shortening the tracking time of target objects during the initial screening of suspected abnormal objects would inevitably lead to an increase in the number of false positives, this method further performs precise anomaly recognition based on the contour-corrected mask image of the suspected abnormal objects, filtering out the number of false positives caused by shortening the tracking time. This improves the accuracy of abnormal target object recognition while shortening the tracking time, thus achieving a balance between the timeliness and accuracy of abnormal target object recognition.

[0196] Corresponding to the above Figures 1 to 5 Based on the same technical concept, this application also provides a target recognition device in its embodiments regarding the target recognition method described. Figure 7 This is a schematic diagram of the module composition of the target recognition device provided in the embodiments of this application. The device is used to perform... Figures 1 to 5 The described target recognition method, such as Figure 7 As shown, the device includes:

[0197] The abnormal target screening module 702 is used to determine the target images in the image to be processed that contain suspected abnormal objects;

[0198] The sub-image determination module 704 is used to perform segmentation processing on the target image to obtain a sub-image containing the suspected abnormal object;

[0199] The object contour correction module 706 is used to perform contour correction processing on the suspected abnormal object in the sub-image to obtain a contour correction mask image.

[0200] The abnormal target recognition module 708 is used to perform anomaly classification and recognition on the suspected abnormal object based on the contour correction mask image and the sub-image, and obtain the anomaly classification and recognition result.

[0201] The target recognition device in this embodiment first tracks each target object within a preset monitoring area to initially identify suspected abnormal objects. Then, for each suspected abnormal object, a contour-corrected mask image is segmented from the target image, and anomaly classification and recognition are performed based on the contour-corrected mask image to obtain the final anomaly classification and recognition result. Considering that shortening the tracking time of target objects during the initial screening of suspected abnormal objects would inevitably lead to an increase in the number of false positives, the device further performs precise anomaly recognition based on the contour-corrected mask image of the suspected abnormal objects, filtering out the number of false positives caused by shortening the tracking time. This improves the accuracy of abnormal target object recognition while shortening the tracking time, thus achieving a balance between the timeliness and accuracy of abnormal target object recognition.

[0202] It should be noted that the embodiments of the target recognition device in this application and the embodiments of the target recognition method in this application are based on the same inventive concept. Therefore, the specific implementation of this embodiment can be referred to the implementation of the corresponding target recognition method mentioned above, and the repeated parts will not be described again.

[0203] Corresponding to the target repair network training method described above, and based on the same technical concept, this application also provides a target repair network training device. Figure 8 This is a schematic diagram of the module composition of the target repair network training device provided in the embodiments of this application. This device is used to execute the target repair network training method described above, such as... Figure 8 As shown, the device includes:

[0204] The sample set acquisition module 802 is used to acquire the training sample set of the target repair network to be trained; the training sample set includes target contour mask images corresponding to multiple key point labeled sample images respectively;

[0205] The key point correction module 804 is used to perform key point correction processing on the target contour mask map using the target repair network to obtain the intermediate convolutional feature map and the contour correction mask map of the key point annotation sample image.

[0206] The first loss determination module 806 is used to determine a first loss function based on the intermediate convolutional feature map and the preset standard feature map;

[0207] The second loss determination module 808 is used to determine a second loss function based on the contour correction mask and the preset annotation mask;

[0208] The repair network training module 810 is used to iteratively update the parameters of the target repair network based on the first loss function and the second loss function to obtain the trained target repair network; the target repair network is used to perform contour correction processing on suspected abnormal objects in sub-images to obtain contour correction mask maps for anomaly classification and recognition, the sub-images include the image regions where the suspected abnormal objects are located in the target image, and the target image is the image containing the suspected abnormal objects in the image to be processed.

[0209] The target repair network training device in this embodiment of the application pre-trains the target repair network and then uses the pre-trained target repair network to perform contour correction processing on the initial contour mask image of the suspected abnormal object. This can improve the accuracy of the contour correction mask image of the suspected abnormal object, and further improve the accuracy of abnormal target identification.

[0210] It should be noted that the embodiments of the target repair network training device in this application and the embodiments of the target repair network training method in this application are based on the same inventive concept. Therefore, the specific implementation of this embodiment can refer to the implementation of the corresponding target repair network training method mentioned above, and the repeated parts will not be described again.

[0211] Corresponding to the above Figures 1 to 6b Based on the same technical concept, embodiments of this application also provide an electronic device for performing the above-described target recognition method, such as... Figure 9 As shown.

[0212] Electronic devices can vary considerably due to differences in configuration or performance. They may include one or more processors 901 and memories 902, with the memory 902 storing one or more application programs or data. The memory 902 can be temporary or persistent storage. The application programs stored in the memory 902 may include one or more modules (not shown in the figures), each module including a series of computer-executable instructions for the electronic device. Furthermore, the processor 901 may be configured to communicate with the memory 902 and execute the series of computer-executable instructions stored in the memory 902 on the electronic device. The electronic device may also include one or more power supplies 903, one or more wired or wireless network interfaces 904, one or more input / output interfaces 905, one or more keyboards 906, etc.

[0213] An electronic device includes a memory and one or more programs, wherein the one or more programs are stored in the memory, and the one or more programs may include one or more modules, and each module may include a series of computer-executable instructions for use in the electronic device, and is configured to be executed by one or more processors. The one or more programs include computer-executable instructions for performing the following:

[0214] Based on the first image and the second image, a suspected abnormal object is identified, wherein the second image was captured later than the first image;

[0215] Based on the sub-image of the suspected abnormal object in the second image, the contour correction processing is performed on the suspected abnormal object to obtain the contour correction mask image of the suspected abnormal object.

[0216] Based on the contour-corrected mask image and the sub-image, the suspected abnormal object is classified and identified to obtain the abnormal classification and identification result.

[0217] The program, configured to be executed by one or more processors, also includes computer-executable instructions for performing the following:

[0218] Obtain the training sample set of the target repair network to be trained; the training sample set includes target contour mask images corresponding to multiple key point labeled sample images respectively;

[0219] The target inpainting network is used to perform key point correction processing on the target contour mask image to obtain an intermediate convolutional feature map and a contour correction mask image of the key point labeled sample image;

[0220] Based on the intermediate convolutional feature map and the preset standard feature map, a first loss function is determined; based on the contour correction mask map and the preset annotation mask map, a second loss function is determined.

[0221] Based on the first loss function and the second loss function, the target repair network is iteratively updated to obtain the trained target repair network; the target repair network is used to perform contour correction processing on suspected abnormal objects in sub-images to obtain contour correction mask maps for anomaly classification and recognition, the sub-images include the image regions where the suspected abnormal objects are located in the target image, and the target image is the image containing the suspected abnormal objects in the image to be processed.

[0222] The electronic device in this embodiment first tracks each target object within a preset monitoring area to initially identify suspected abnormal objects. Then, for each suspected abnormal object, a contour-corrected mask image is segmented from the target image, and anomaly classification and recognition are performed based on the contour-corrected mask image to obtain the final anomaly classification and recognition result. Considering that shortening the tracking time of target objects during the initial screening of suspected abnormal objects would inevitably lead to an increase in the number of false positives, further precise anomaly recognition is performed based on the contour-corrected mask image of the suspected abnormal objects to filter out the number of false positives caused by shortening the tracking time. This improves the accuracy of abnormal target object recognition while shortening the tracking time, thus achieving a balance between the timeliness and accuracy of abnormal target object recognition. It should be noted that the embodiments of the electronic device and the embodiments of the target recognition method in this application are based on the same inventive concept. Therefore, the specific implementation of this embodiment can be referred to the implementation of the corresponding target recognition method described above, and repeated details will not be repeated.

[0223] Corresponding to the above Figures 1 to 6b Based on the same technical concept, this application also provides a computer-readable storage medium for storing computer-executable instructions. The storage medium can be a USB flash drive, optical disc, hard disk, etc. When the computer-executable instructions stored in the storage medium are executed by a processor, they can achieve the following process:

[0224] Based on the first image and the second image, a suspected abnormal object is identified, wherein the second image was captured later than the first image;

[0225] Based on the sub-image of the suspected abnormal object in the second image, the contour correction processing is performed on the suspected abnormal object to obtain the contour correction mask image of the suspected abnormal object.

[0226] Based on the contour-corrected mask image and the sub-image, the suspected abnormal object is classified and identified to obtain the abnormal classification and identification result.

[0227] When the computer-executable instructions stored on the aforementioned storage medium are executed by the processor, they can also perform the following processes:

[0228] Obtain the training sample set of the target repair network to be trained; the training sample set includes target contour mask images corresponding to multiple key point labeled sample images respectively;

[0229] The target inpainting network is used to perform key point correction processing on the target contour mask image to obtain an intermediate convolutional feature map and a contour correction mask image of the key point labeled sample image;

[0230] Based on the intermediate convolutional feature map and the preset standard feature map, a first loss function is determined;

[0231] Based on the contour-corrected mask and the preset labeled mask, a second loss function is determined;

[0232] Based on the first loss function and the second loss function, the target repair network is iteratively updated to obtain the trained target repair network; the target repair network is used to perform contour correction processing on suspected abnormal objects in sub-images to obtain contour correction mask maps for anomaly classification and recognition, the sub-images include the image regions where the suspected abnormal objects are located in the target image, and the target image is the image containing the suspected abnormal objects in the image to be processed.

[0233] When the computer-executable instructions stored in the storage medium in this embodiment are executed by the processor, they first track each target object within a preset monitoring area to initially identify suspected abnormal objects. Then, for each suspected abnormal object, a contour-corrected mask image is segmented from the target image, and anomaly classification and recognition are performed based on the contour-corrected mask image to obtain the final anomaly classification and recognition result. Considering that shortening the tracking time of target objects during the initial screening of suspected abnormal objects would inevitably lead to an increase in the number of false positives, further precise anomaly recognition is performed based on the contour-corrected mask image of the suspected abnormal objects to filter out the number of false positives caused by shortening the tracking time. This improves the accuracy of abnormal target object recognition while shortening the tracking time, thus achieving a balance between the timeliness and accuracy of abnormal target object recognition. It should be noted that the embodiments regarding the storage medium and the embodiments regarding the target recognition method in this application are based on the same inventive concept. Therefore, the specific implementation of this embodiment can be found in the implementation of the corresponding target recognition method described above, and repeated details will not be repeated.

[0234] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements some or all of the steps in the methods described above; wherein, when the computer program product is run on a computer, the computer performs the aforementioned related steps to implement the methods in the embodiments described above.

[0235] In addition, embodiments of this application also provide an apparatus, which may specifically be a chip, component, or module. The apparatus may include a connected processor and a memory; wherein the memory is used to store computer execution instructions, and when the apparatus is running, the processor may execute the computer execution instructions stored in the memory to cause the chip to execute the methods in the above-described method embodiments.

[0236] In this application, the electronic devices, servers, computer storage media, computer program products, or chips provided in the embodiments are all used to execute the corresponding methods provided above. Therefore, the beneficial effects they can achieve can be referred to the beneficial effects in the corresponding methods provided above, and will not be repeated here. Specific embodiments of this application have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired results. In some implementations, multitasking and parallel processing are also possible or may be advantageous. Those skilled in the art should understand that the embodiments of this application can be provided as methods, systems, or computer program products. Therefore, the embodiments of this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, this application can take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0237] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0238] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 One or more processes and / or boxes Figure 1 The functions specified in one or more boxes. These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable apparatus for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0239] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory. Memory may include non-persistent storage in computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media. Computer-readable media includes both permanent and non-persistent, removable and non-removable media that can store information by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transfer medium used to store information that can be accessed by the computing device. As defined herein, computer-readable media do not include transient media, such as modulated data signals and carrier waves. It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0240] The embodiments of this application can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. One or more embodiments of this application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can reside in local and remote computer storage media, including storage devices. The various embodiments in this application are described in a progressive manner, with reference to each other for similar or identical parts. Each embodiment focuses on describing the differences from other embodiments. In particular, for system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and relevant parts can be referred to in the description of the method embodiments. The above descriptions are merely embodiments of this document and are not intended to limit this document. Various modifications and variations can be made to this document by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this document should be included within the scope of the claims of this document.

Claims

1. A target recognition method, characterized in that, The method includes: Identify the target image containing a suspected abnormal object within the image to be processed; The target image is segmented to obtain a sub-image containing the suspected abnormal object; The suspected abnormal object in the sub-image is subjected to contour correction processing to obtain a contour correction mask image; Based on the contour-corrected mask image and the sub-image, the suspected abnormal object is classified and identified to obtain the abnormal classification and identification result.

2. The method according to claim 1, characterized in that, The images to be processed include a first image and a second image, wherein the second image was captured later than the first image. The step of determining that a target image containing a suspected abnormal object exists within the images to be processed includes: The first image is input into the target detection network to obtain the first coordinate information of the target object to be identified; Based on the first coordinate information, the movement orientation of the target object to be identified is predicted, and the second coordinate information of the target object to be identified in the second image is determined. Based on the second coordinate information and the second image, target tracking is performed on the target object to be identified to obtain target tracking information; Based on the target tracking information, the suspected abnormal object is identified, and the second image containing the suspected abnormal object is identified as the target image.

3. The method according to claim 2, characterized in that, The step of tracking the target object to be identified based on the second coordinate information and the second image to obtain target tracking information includes: Based on the second coordinate information and the actual coordinate information of the target object in the second image, coordinate matching processing is performed to determine the coordinate matching result of the target object to be identified; Based on the coordinate matching results, the target tracking information of the target object to be identified is determined.

4. The method according to claim 2, characterized in that, The step of predicting the movement orientation of the target object to be identified based on the first coordinate information, and determining the second coordinate information of the target object in the second image, includes: Based on the first coordinate information and the preset state transition matrix, the movement orientation of the target object to be identified is predicted, and the initial predicted coordinate information is determined. The Kalman gain coefficients are determined based on the target covariance matrix and the Kalman filtering algorithm; the target covariance matrix is ​​determined based on the state covariance matrix corresponding to the first image. Based on the Kalman gain coefficient, the initial predicted coordinate information is corrected for predicted orientation to obtain the second coordinate information of the target object to be identified in the second image.

5. The method according to claim 2, characterized in that, The step of determining the suspected abnormal object based on the target tracking information includes: Based on the target tracking information, the target tracking duration and target tracking distance are determined; The target objects whose target tracking duration is greater than a first threshold and less than a second threshold, and whose target tracking distance is less than a preset distance threshold, are identified as the suspected abnormal objects.

6. The method according to claim 1, characterized in that, The step of performing contour correction processing on the suspected abnormal objects in the sub-image to obtain a contour-corrected mask image includes: The sub-image is input into the target segmentation network for target contour extraction to obtain the initial mask image of the suspected abnormal object; The initial contour mask is input into the target repair network for contour correction processing to obtain the contour-corrected mask.

7. The method according to claim 6, characterized in that, The step of inputting the initial contour mask image into the target repair network for contour correction processing to obtain the contour-corrected mask image includes: The initial contour mask is feature-encoded to obtain a first feature map; The first feature map is subjected to key point correction processing to obtain the second feature map; The second feature map is decoded to obtain the third feature map; The third feature map is subjected to key point correction processing to obtain the fourth feature map; Based on the fourth feature map, the contour correction mask map is determined.

8. The method according to claim 7, characterized in that, The step of performing key point correction processing on the first feature map to obtain the second feature map includes: The first feature map is subjected to grouped convolution processing to obtain multiple discrete features; The discrete features are concatenated to obtain concatenated features. The spliced ​​features are convolved to obtain an intermediate convolutional feature map; A second feature map is generated based on the first feature map and the intermediate convolutional feature map.

9. The method according to claim 1, characterized in that, The step of classifying and identifying the suspected abnormal object based on the contour-corrected mask image and the sub-image to obtain the abnormal classification and identification result includes: Color restoration processing is performed based on the contour-corrected mask image and the sub-image to obtain the contour-corrected texture image of the suspected abnormal object. The contour-corrected texture image is input into a discriminative network to obtain the anomaly classification and recognition result of the suspected abnormal object.

10. A method for training a target repair network, characterized in that, The method includes: Obtain the training sample set of the target repair network to be trained; the training sample set includes target contour mask images corresponding to multiple key point labeled sample images respectively; The target inpainting network is used to perform key point correction processing on the target contour mask image to obtain an intermediate convolutional feature map and a contour correction mask image of the key point labeled sample image; Based on the intermediate convolutional feature map and the preset standard feature map, a first loss function is determined; Based on the contour-corrected mask and the preset labeled mask, a second loss function is determined; Based on the first loss function and the second loss function, the parameters of the target repair network are iteratively updated to obtain the trained target repair network.

11. The method according to claim 10, characterized in that, The step of determining the first loss function based on the intermediate convolutional feature map and the preset standard feature map includes: Based on the first intermediate convolutional feature map and the first standard feature map, a first sub-loss function is determined; the first intermediate convolutional feature map is obtained by the target repair network through key point correction processing based on the target contour mask map. The second sub-loss function is determined based on the second intermediate convolutional feature map and the second standard feature map; the second intermediate convolutional feature map is obtained by the target repair network through key point correction processing based on the first intermediate convolutional feature map. The first loss function is determined based on the first sub-loss function and the second sub-loss function.

12. The method according to claim 11, characterized in that, The target repair network includes an encoder and a decoder. The encoder includes multiple encoding units, each encoding unit including a first convolutional layer and a first keypoint correction layer. The decoder includes multiple decoding units, each decoding unit including a second convolutional layer and a second keypoint correction layer. The input to the first keypoint correction layer is the feature map output by the first convolutional layer belonging to the same coding unit, and the input to the first convolutional layer is the target contour mask or the feature map output by the first keypoint correction layer in the previous coding unit. The input to the second keypoint correction layer is the feature map output by the second convolutional layer belonging to the same decoding unit, and the input to the second convolutional layer is the output of the encoder or the feature map output by the second keypoint correction layer in the previous decoding unit.

13. The method according to claim 11, characterized in that, The intermediate convolutional feature map includes a first intermediate convolutional feature map and a second intermediate convolutional feature map; the step of using the target inpainting network to perform key point correction processing on the target contour mask map to obtain the intermediate convolutional feature map and the contour correction mask map of the key point labeled sample image includes: The target contour mask image is feature-encoded to obtain the fifth feature image; The fifth feature map is subjected to key point correction processing to obtain a first intermediate convolutional feature map and a sixth feature map; the sixth feature map is determined based on the first intermediate convolutional feature map. The sixth feature map is decoded to obtain the seventh feature map; The seventh feature map is subjected to key point correction processing to obtain a second intermediate convolutional feature map and an eighth feature map; the eighth feature map is determined based on the second intermediate convolutional feature map. Based on the eighth feature map, the contour correction mask map of the key point annotation sample image is determined.

14. A target recognition device, characterized in that, The device includes: The abnormal target screening module is used to identify target images in the image to be processed that contain suspected abnormal objects; The sub-image determination module is used to segment the target image to obtain a sub-image containing the suspected abnormal object; An object contour correction module is used to perform contour correction processing on the suspected abnormal objects in the sub-image to obtain a contour correction mask image. An abnormal target recognition module is used to perform anomaly classification and recognition on the suspected abnormal object based on the contour correction mask image and the sub-image, and obtain the anomaly classification and recognition result.

15. A target repair network training device, characterized in that, The device includes: The sample set acquisition module is used to acquire the training sample set of the target repair network to be trained; the training sample set includes target contour mask images corresponding to multiple key point labeled sample images respectively; The key point correction module is used to perform key point correction processing on the target contour mask map using the target repair network to obtain an intermediate convolutional feature map and a contour correction mask map of the key point labeled sample image; The first loss determination module is used to determine a first loss function based on the intermediate convolutional feature map and the preset standard feature map; The second loss determination module is used to determine a second loss function based on the contour correction mask and the preset annotation mask; The repair network training module is used to iteratively update the parameters of the target repair network based on the first loss function and the second loss function to obtain the trained target repair network; the target repair network is used to perform contour correction processing on suspected abnormal objects in sub-images to obtain contour correction mask maps for anomaly classification and recognition, the sub-images include the image regions where the suspected abnormal objects are located in the target image, and the target image is the image containing the suspected abnormal objects in the image to be processed.

16. An electronic device, characterized in that, The device includes: Processor; and A memory configured to store computer-executable instructions, which are configured to be executed by the processor, including steps for performing the target recognition method as claimed in any one of claims 1 to 9, or the target repair network training method as claimed in any one of claims 10 to 13.

17. A computer-readable storage medium, characterized in that, The storage medium is used to store computer-executable instructions that cause the computer to perform the steps of the target recognition method as described in any one of claims 1 to 9, or the target repair network training method as described in any one of claims 10 to 13.

18. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the target recognition method according to any one of claims 1 to 9, or the target repair network training method according to any one of claims 10 to 13.