Model training system, image recognition device, model training method and recognition method

By generating binarized image samples of multiple transforms and a behavior recognition model trained using a priori attention mechanism, the accuracy of the behavior recognition model when regional images are changed is solved, and the generalization ability and recognition accuracy of the model are improved.

CN120388250APending Publication Date: 2025-07-29LENS TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510394363.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

When the existing behavior recognition model processes changes in the number of regional images, it is impossible to accurately identify different behavioral scenarios. Due to the diversity of behavioral scenarios and low image quality, it is difficult to obtain sufficient sample image training algorithm models, resulting in poor model performance.

Method used

A model training system is adopted to generate multiple transformed binary image samples through image acquisition, processing and stitching modules, and combine residual networks and binary cross-entropy loss functions to train behavior recognition models, ignore the labels of preset values, and use the prior attention mechanism to adapt to the number and position changes of the area of attention.

Benefits of technology

The generalization ability of the model is improved, and it can accurately identify different behavioral scenarios while paying attention to changes in the number of regions and locations, reduce dependence on specific features, and avoid overfitting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388250A_ABST
    Figure CN120388250A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a model training system, an image recognition device, a model training method and a recognition method, and belongs to the field of image recognition. The model training system comprises an image processing module used for performing out-of-order processing on a first number of binarized images corresponding to each sample image, and randomly setting all pixel values of at least one binarized image as second pixel values to obtain a first number of processed binarized images; the image splicing module is used for splicing the sample image with the first number of processed binarized images for each sample image to obtain a to-be-input sample image; and the model training module is used for inputting all the to-be-input sample images into a preset model and training the preset model to obtain a behavior recognition model. And the model can still identify different behavior scenes if the number and the position of the concerned areas in the scene are changed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image recognition, and particularly to a model training system, an image recognition device, a model training method, and a recognition method. Background Art

[0002] With the rapid development of machine learning technology, behavior recognition models are widely used in various behavior recognition scenarios. Usually, it is necessary to bind the user's behavior to the regional image. For example, when recognizing the behavior scenario of a user adding detergent to a sink, the user's behavior is bound to the regional image of the sink part, so that the behavior recognition model can obtain the user's behavior recognition result based on the regional image.

[0003] However, in the process of training the algorithm model, the behavior scenarios corresponding to the sample images are usually relatively single, but the regional images to be recognized in the actual behavior scenarios may change at any time, such as increasing or decreasing. When the number of regional images to be processed changes, the accuracy of the results output by the algorithm model is poor, and thus different behavior scenarios cannot be accurately recognized. In addition, due to reasons such as the diversity of behavior scenarios, low image quality, and low data annotation efficiency, it is usually difficult to obtain a large number of sample images to train the algorithm model, resulting in low performance of the trained algorithm model, which further leads to the inability to accurately recognize different behavior scenarios. Summary of the Invention

[0004] The purpose of the embodiments of the present invention is to provide a model training system, an image recognition device, a model training method, and a recognition method to solve the problem of being unable to accurately recognize different behavior scenarios.

[0005] To achieve the above purpose, in a first aspect, the present application provides a model training system, which includes an image acquisition module, an image processing module, an image splicing module, and a model training module;

[0006] The image acquisition module is used to acquire a first number of binary images corresponding to each sample image, where the sample image includes a first number of regions of interest, each binary image has a first pixel value set to be recognizable for one region of interest, and the regions outside the regions of interest are set to a second pixel value that is not recognizable;

[0007] The image processing module is used to randomly arrange the first number of binary images corresponding to each sample image, and randomly set all pixel values of at least one binary image to the second pixel value to obtain a first number of processed binary images;

[0008] The image splicing module is used to splice the sample image and the first number of processed binary images for each sample image to obtain a sample image to be input;

[0009] A model training module, configured to input all to-be-input sample images into a preset model, train the preset model, and obtain a behavior recognition model.

[0010] In an embodiment of the present application, the sample image includes multiple labels, and the model training system further includes a label processing module;

[0011] The label processing module is configured to sort the multiple labels based on the order of the processed binary images of the first quantity, and set the label corresponding to the binary image with all pixel values being the second pixel value to a preset value.

[0012] In an embodiment of the present application, the model training system further includes a model construction module;

[0013] The model construction module is configured to construct a preset model based on a residual network structure;

[0014] The model training module is further configured to input all to-be-input sample images into the preset model, and train the preset model by using a binary cross-entropy loss function to obtain a behavior recognition model, wherein the preset model ignores the binary images with the label being the preset value.

[0015] In an embodiment of the present application, the image stitching module includes a region image extraction sub-module and a sample image obtaining sub-module;

[0016] The region image extraction sub-module is configured to extract a sample region image from the sample image, and the sample region image includes a first quantity of regions of interest;

[0017] The sample image obtaining sub-module is configured to stitch the sample region image with the first quantity of processed binary images to obtain a to-be-input sample image.

[0018] In an embodiment of the present application, the sample image obtaining sub-module is further configured to convert the sample region image into a sample region image of a preset shape;

[0019] Stitch the sample region image of the preset shape with the first quantity of processed binary images to obtain a to-be-input sample image.

[0020] In an embodiment of the present application, the first quantity is the maximum quantity of regions of interest determined by the behavior recognition model.

[0021] In an embodiment of the present application, the image processing module includes a scrambling processing sub-module and a pixel processing sub-module;

[0022] The disordered processing sub-module is used to disorder the arrangement order of the first number of binarized images corresponding to each sample image, and disorder the positions of the regions of interest in each binarized image, so as to obtain the first number of disordered binarized images;

[0023] The pixel processing sub-module is used to randomly set all pixel values of at least one disordered binarized image to the second pixel value, so as to obtain the first number of processed binarized images.

[0024] In the embodiments of the present application, the model training system further includes an image enhancement module;

[0025] The image enhancement module is used to perform image enhancement processing on each sample image to be input respectively, so as to obtain the enhanced sample image, wherein the image enhancement processing includes at least one of rotation, flipping, translation, scaling, affine transformation, brightness adjustment, contrast adjustment, adding Gaussian noise, adding Gaussian blur, and covering part of the pixel area;

[0026] The model training module is further used to input all the enhanced sample images into a preset model, train the preset model, and obtain a behavior recognition model.

[0027] In a second aspect, the present application provides an image recognition device, which includes a target area determination module, a first target image obtaining module, a second target image obtaining module, and a recognition result obtaining module;

[0028] The target area determination module is used to obtain the second number of target regions of interest in the target image, and the target binarized image corresponding to each target region of interest;

[0029] The first target image obtaining module is used to splice the target image and all the target binarized images to obtain the target image to be recognized when the first number is equal to the second number;

[0030] The second target image obtaining module is used to splice the target image, all the target binarized images, and the second pixel value images with the third number when the first number is greater than the second number, so as to obtain the target image to be recognized, wherein the third number is the difference between the first number and the second number;

[0031] The recognition result obtaining module is used to input the target image to be recognized into the behavior recognition model to obtain the behavior recognition result, wherein the behavior recognition model is obtained through the above-mentioned model training system.

[0032] In a third aspect, the present application provides a model training method, which is applied to the above-mentioned model training system, and the model training method includes:

[0033] Obtain a first number of binary images corresponding to each sample image, where the sample image includes a first number of regions of interest, each binary image has a corresponding region of interest set to a recognizable first pixel value, and regions outside the regions of interest are set to an unrecognizable second pixel value;

[0034] For the first number of binary images corresponding to each sample image, shuffle the first number of binary images, and randomly set all pixel values of at least one binary image to the second pixel value to obtain a first number of processed binary images;

[0035] For each sample image, splice the sample image with the first number of processed binary images to obtain a to-be-input sample image;

[0036] Input all to-be-input sample images into a preset model, train the preset model, and obtain a behavior recognition model.

[0037] In a fourth aspect, the present application provides an image recognition method applied to the above-mentioned image recognition device. The image recognition method includes:

[0038] Obtain a second number of target regions of interest in a target image, and a target binary image corresponding to each target region of interest;

[0039] When the first number is equal to the second number, splice the target image with all target binary images to obtain a to-be-recognized target image;

[0040] When the first number is greater than the second number, splice the target image, all target binary images, and a second pixel value image of a third number, where the third number is the difference between the first number and the second number, to obtain a to-be-recognized target image;

[0041] Input the to-be-recognized target image into the behavior recognition model to obtain a behavior recognition result, where the behavior recognition model is obtained through the above-mentioned model training system.

[0042] The present application provides a model training system, comprising: an image acquisition module for acquiring a first number of binary images corresponding to each sample image; an image processing module for shuffling the first number of binary images corresponding to each sample image, and randomly setting all pixel values of at least one binary image to a second pixel value, thereby obtaining a first number of processed binary images; an image splicing module for splicing the sample image with the first number of processed binary images for each sample image, thereby obtaining a sample image to be input; and a model training module for inputting all sample images to be input into a preset model, training the preset model, and obtaining a behavior recognition model. By providing a priori attention mechanism to process the image, the behavior recognition model is trained to adapt to any number of attention sub-regions within the first number. Even if the number and position of the attention sub-regions in the actual behavior scene change, the behavior recognition model can still output accurate recognition results, thereby accurately identifying different behavior scenes. In addition, by adjusting the order and pixels of the binary images, more sample images for the training model can be obtained, thereby improving the generalization ability and performance of the model, thereby more accurately identifying different behavior scenes.

[0043] Other features and advantages of the embodiments of the present invention will be described in detail in the subsequent detailed description. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] The accompanying drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. Together with the following detailed description, they are used to explain the embodiments of the present invention, but do not constitute a limitation of the embodiments of the present invention. In the accompanying drawings:

[0045] Figure 1 A schematic diagram of the structure of the model training system provided in an embodiment of the present application is shown;

[0046] Figure 2 An example diagram of a binarized image provided by an embodiment of the present application is shown;

[0047] Figure 3 An example diagram of a processed binary image provided by an embodiment of the present application is shown;

[0048] Figure 4 An example diagram of multiple tags provided in an embodiment of the present application is shown;

[0049] Figure 5 An example diagram of a sample area image provided by an embodiment of the present application is shown;

[0050] Figure 6 An example diagram showing a sample area image of a preset shape provided by an embodiment of the present application is shown;

[0051] Figure 7 Shows an example diagram of the sample image to be input provided by an embodiment of the present application;

[0052] Figure 8 Shows an example diagram of the residual network structure provided by an embodiment of the present application;

[0053] Figure 9 Shows a schematic structural diagram of an image recognition device provided by an embodiment of the present application;

[0054] Figure 10 Shows a flowchart of a model training method provided by an embodiment of the present application;

[0055] Figure 11 Shows a flowchart of an image recognition method provided by an embodiment of the present application. Detailed implementation manners

[0056] Next, the specific implementation manners of the embodiments of the present invention will be described in detail in conjunction with the accompanying drawings in the embodiments of the present invention. It should be understood that the specific implementation manners described herein are only used to illustrate and explain the embodiments of the present invention, and are not used to limit the embodiments of the present invention.

[0057] Generally, the components of the embodiments of the present invention described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0058] Hereinafter, the terms "including", "having" and their cognates that can be used in various embodiments of the present invention are only intended to represent specific features, numbers, steps, operations, elements, components or combinations of the foregoing items, and should not be construed as first excluding the existence of one or more other features, numbers, steps, operations, elements, components or combinations of the foregoing items or increasing the possibility of one or more features, numbers, steps, operations, elements, components or combinations of the foregoing items.

[0059] In addition, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0060] Unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by one of ordinary skill in the art to which various embodiments of the present invention belong. The terms (such as those defined in a commonly used dictionary) will be interpreted as having the same meaning as the contextual meaning in the relevant technical field and will not be interpreted as having an idealized meaning or an overly formal meaning unless clearly defined in various embodiments of the present invention.

[0061] A classification model is generally used for behavior recognition. However, the classification model is not applicable to scenarios where the number of sub-regions of interest increases or decreases. Specifically, assume that a classification model is used for behavior recognition and the number of target sub-regions of interest is 2. Then, when training the classification model, the sample images are labeled with 2 target sub-regions of interest by labels, and the sub-regions of interest other than the 2 target sub-regions of interest are not labeled. When the number of target sub-regions of interest becomes 3, since the sample images lack the labels of the newly added sub-regions of interest, the classification model is not applicable to the scenarios where the number of sub-regions of interest increases or decreases, and thus an accurate behavior recognition result cannot be obtained. Usually, it is necessary to re-label the sample images and train a new classification model to recognize the behavior scenarios where the number of target sub-regions of interest is 2. Since the generality of a single classification model is low and an accurate result cannot be output when the sub-regions of interest change, multiple classification models are generally used for behavior recognition. However, due to the number of classification models, when there is no trained classification model for a certain number of sub-regions of interest, an accurate behavior recognition result is still obtained. In addition, the classification model usually easily loses the interaction information of the environment, resulting in an inability to obtain an accurate behavior recognition result.

[0062] Embodiment 1

[0063] Please refer to Figure 1 , Figure 1 which shows a schematic structural diagram of the model training system provided by the embodiments of the present application. Figure 1 The model training system 100 in

[0064] The image acquisition module 110 is used to acquire a first number of binary images corresponding to each sample image, where the sample image includes a first number of regions of interest, each binary image has a first pixel value corresponding to a region of interest set to be recognizable, and the regions other than the regions of interest are set to a second pixel value that is not recognizable.

[0065] Obtain a first number of binarized images corresponding to each sample image, where the sample image includes a first number of regions of interest, each binarized image has a corresponding region of interest set to a recognizable first pixel value, and regions outside the region of interest are set to an unrecognizable second pixel value. It should be understood that the value of the first number is set according to actual needs and is not limited herein. The first pixel value and the second pixel value are also set according to actual needs and are not limited herein. For ease of understanding, in the embodiments of the present application, the first number is 3. The images of all sub-regions of interest are used as training samples, and the trained model is adapted to any number of sub-regions of interest within the first number, so that the trained model can recognize images with different numbers of sub-regions of interest.

[0066] A binarized image corresponding to a sub-region of interest is an image with a first color, and the region outside the sub-region of interest is a second color. The first color and the second color are both set according to actual needs and are not limited herein. The image type of the binarized image is also set according to actual needs and is not limited herein. For ease of understanding, in the embodiments of the present application, the binarized image is a Mask single-channel image, the first pixel value is 1, and the second pixel value is 0, that is, the binarized image is a grayscale image. The pixel values of the sub-region of interest are set to 1, that is, the first color is white, and the pixel values of the region outside the sub-region of interest are set to 0, that is, the second color is black.

[0067] The image processing module 120 is configured to perform a disordering process on the first number of binarized images corresponding to each sample image, and randomly set all pixel values of at least one binarized image to the second pixel value to obtain the first number of processed binarized images.

[0068] Please refer to Figure 2 , Figure 2 which shows an example diagram of the binarized image provided by the embodiments of the present application.

[0069] In this embodiment, a prior attention mechanism is provided to process the sample image. The attention mechanism can enable the model to focus on the important parts of the input image, thereby improving the overall recognition performance and efficiency of the model. As Figure 2 shown, taking any one of the binarized images as an example, set the pixel values of any one of the regions of interest in the first number to the first pixel value, and set the pixel values of the non-interest regions outside the region of interest and other regions of interest to the second pixel value to obtain a binarized image. The binarized image is an image with a white sub-region of interest and a black region outside the sub-region of interest. Specifically, for the first number of binarized images corresponding to each sample image, perform a disordering process on the first number of binarized images, and randomly set all pixel values of at least one binarized image to the second pixel value to obtain the first number of processed binarized images.

[0070] After shuffling the first quantity of binarized images, the position of the region of interest in the processed binarized images changes, corresponding to the change in the position of the region of interest in the actual behavior recognition scenario, so that the model can better adapt to the actual behavior recognition scenario. When all the pixel values of a binarized image are set to the second pixel value, that is, all the pixel values of the binarized image are zero, all regions of the binarized image are black images, which are not shown in the figure, corresponding to the change in the number of regions of interest in the actual behavior recognition scenario, so that the model can better adapt to the actual behavior recognition scenario.

[0071] The image stitching module 130 is configured to stitch each sample image with the first quantity of processed binarized images to obtain a sample image to be input.

[0072] For each sample image, the sample image is stitched with the first quantity of processed binarized images by channels to obtain a sample image to be input. The sample image to be input corresponds to the situation where the position and number of regions of interest in the actual behavior recognition scenario change at any time, so that the model can better adapt to the actual behavior recognition scenario.

[0073] In the embodiments of the present application, the image processing module 120 includes a shuffling sub-module and a pixel processing sub-module;

[0074] The shuffling sub-module is configured to shuffle the arrangement order of the first quantity of binarized images corresponding to each sample image, and shuffle the position of the region of interest in each binarized image to obtain the first quantity of shuffled binarized images;

[0075] The pixel processing sub-module is configured to randomly set all the pixel values of at least one shuffled binarized image to the second pixel value to obtain the first quantity of processed binarized images.

[0076] In this embodiment, the shuffling process includes shuffling the arrangement order and position. For the first quantity of binarized images corresponding to each sample image, the arrangement order of the first quantity of binarized images is shuffled, and the position of the region of interest in each binarized image is shuffled to obtain the first quantity of shuffled binarized images. Specifically, in this embodiment, the order of the first quantity of binarized images is the first binarized image, the second binarized image, and the third binarized image in sequence. Shuffling the arrangement order of the first quantity of binarized images can change the order of the binarized images to the second binarized image, the third binarized image, and the first binarized image.

[0077] Please refer toFigure 3 , Figure 3 shows an example diagram of the processed binary image provided by the embodiment of the present application.

[0078] Randomize the positions of the regions of interest in each binary image. As Figure 3 shown, after randomizing the positions, the positions of the regions of interest in the same binary image change. Randomly set all pixel values of at least one scrambled binary image to a second pixel value to obtain a first quantity of processed binary images. In an actual behavior recognition scenario, not only the number of regions of interest to be recognized changes, but also the arrangement order and positions of the regions of interest to be recognized change. By scrambling the binary images, the obtained input sample images can better correspond to the situations where the positions, quantities, and arrangement orders of the regions of interest change, thereby enabling the model to better adapt to the actual behavior recognition scenario.

[0079] It should be understood that each time the order of the binary images is adjusted differently, and / or the binary images with all pixels set to zero are different, so that multiple different sets of the first quantity of processed binary images can be obtained for the same sample image. For ease of understanding, in the embodiments of the present application, it is assumed that each scrambling process is the same. In the case of setting all pixel values of the first binary image to the second pixel value, the sample image is spliced with the first quantity of processed binary images to obtain a first input sample image. In the case of setting all pixel values of the second binary image to the second pixel value, the sample image is spliced with the first quantity of processed binary images to obtain a second input sample image. Each time a set of processed binary images is obtained, the binary images with all pixels set to zero are different, so that the same sample image can be spliced with multiple different sets of processed binary images. Usually, by obtaining multiple different sets of processed binary images, more input sample images for training the model can be obtained based on the same sample, thereby improving the generalization ability and performance of the model, reducing the over-dependence of the model on certain features, and avoiding model overfitting.

[0080] The model training module 140 is configured to input all the input sample images into a preset model, train the preset model, and obtain a behavior recognition model.

[0081] Input all the input sample images into the preset model to obtain the recognition result corresponding to each sample image output by the preset model. Since the input sample images are obtained by splicing binary images, the input dimension of the preset model is N*(3 + C)*Height*Width, where N is the batch size of the input input sample images, C is the first quantity, that is, the maximum number of sub-regions of interest, Height is the height of the input sample images, and Width is the width of the input sample images.

[0082] During the process of training the preset model, the preset model outputs the probability of each label, which is a real number in the range of [0, 1]. The model output dimension refers to the result of the output, that is, the number of features included in the output result. In this embodiment, the model output dimension is N*M, where N is the batch size of the input sample images to be input, and M is the number of labels of the sample images. For each sample image, based on the error between the recognition result and the label, the loss function is optimized, and the preset model is trained using the loss function to obtain a behavior recognition model. This application provides a prior attention mechanism to process the image, and trains a behavior recognition model to adapt to any number of attention sub-regions within the first number. Even if the number and position of the attention sub-regions change in the actual behavior scenario, the behavior recognition model can still output accurate recognition results, thereby accurately recognizing different behavior scenarios. In addition, by adjusting the order and pixels of the binary image, more sample images for training the model can be obtained, improving the generalization ability and performance of the model, and thus more accurately recognizing different behavior scenarios.

[0083] In this embodiment, the first number is the maximum number of attention regions determined by the behavior recognition model.

[0084] The first number is the maximum number of attention regions determined by the behavior recognition model. By presetting the first number, the number of attention regions for recognition can be adjusted within the maximum number to dynamically adapt to different behavior recognition scenarios. Specifically, when the first number is 3, in the actual behavior recognition scenario, the attention regions to be recognized can be 3, or 1 or 2. By splicing the sample image with the processed binary image of the first number, the input sample image to be obtained, and training the model with the input sample image to be obtained, only one behavior recognition model is required to meet the requirements of all behavior recognition scenarios when the number of attention regions is from 1 to 3.

[0085] In the embodiment of this application, the sample image includes multiple labels, and the model training system 100 further includes a label processing module;

[0086] The label processing module is used to sort the multiple labels based on the order of the processed binary images of the first number, and set the label corresponding to the binary image with all pixel values being the second pixel value to a preset value.

[0087] Please refer to Figure 4 , Figure 4 which shows an example diagram of multiple labels provided by the embodiment of this application.

[0088] As Figure 4As shown, each sample image includes multiple tags. The multiple tags can be global tags or tags for each attention sub-region, which will not be elaborated here. In this embodiment, the first quantity is C. For ease of understanding, it is assumed that the value of C is 3. The multiple tags are, in sequence, the global tag, the tag of the 1st attention sub-region, the tag of the (C - 1)th attention sub-region, and the tag of the Cth attention sub-region. Shuffle the binary images of the first quantity, and the order of the attention sub-regions is no longer the 1st, the (C - 1)th, and the Cth. Sort the multiple tags based on the order of the processed binary images of the first quantity.

[0089] Since the pixel values of at least one binary image are randomly set to zero, set the tag corresponding to the binary image with pixel values all set to zero to a preset value. During the model training process, the loss function ignores the tags with the preset value and the attention sub-regions corresponding to the tags with the preset value. During the process of training the preset model using the loss function, the label value of the sample image in the loss function is always not equal to the preset value, so that the model can ignore the tags of the attention sub-regions with pixel values all set to zero. It should be understood that the preset value is set according to actual needs and is not limited here. For ease of understanding, the preset value in the embodiments of this application is -1. As shown in the figure, the tags of the Cth attention sub-region are all -1, that is, all pixel values of the Cth attention sub-region are zero, and the black Cth attention sub-region is obtained. Establish a strong mutual relationship between the tags of the sample image and the attention sub-regions, so that the model can focus on the attention sub-regions of the prior attention mechanism, and enable the model to understand that when the attention sub-regions change, the tags will change dynamically, so that the trained model can adapt to the situation where the number and position of the attention sub-regions change.

[0090] In the embodiments of this application, the image stitching module 130 includes a regional image extraction sub-module and a sample image obtaining sub-module;

[0091] The regional image extraction sub-module is used to extract the sample regional image from the sample image, and the sample regional image includes the first quantity of attention regions;

[0092] The sample image obtaining sub-module is used to stitch the sample regional image with the processed binary images of the first quantity to obtain the sample image to be input.

[0093] Please refer to 's Figure 5 Figure 5 which shows an example diagram of the sample regional image provided by the embodiments of this application.

[0094] For ease of understanding, in the embodiments of this application, the sample regional image is an ROI image. ROI (Region of Interest) is an important image region for the model to perform inference and prediction. Extract the sample regions from each sample image respectively, and then obtain the sample regional image, such asFigure 5 As shown, in this example, the sample area is represented by a set of (x, y) pixel coordinate point sequences [(x1, y1), (x2, y2), (x3, y3) … (xn, yn)] ROI as a sample area in the shape of a polygon.

[0095] The sub - area of interest is any sub - area in the area of interest. Specifically, assuming that the area of interest includes a first number of washing troughs, when identifying the behavior scenario of the user adding detergent to the washing trough, each washing trough in the area of interest corresponds to a sub - area of interest. The sample area image is stitched with the first number of processed binary images to obtain the sample image to be input. By extracting the sample area image, unnecessary environmental information in the sample image is removed.

[0096] In the embodiment of the present application, the sample image obtaining sub - module is further configured to convert the sample area image into a sample area image of a preset shape;

[0097] The sample area image of the preset shape is stitched with the first number of processed binary images to obtain the sample image to be input.

[0098] Please refer to Figure 6 , Figure 6 which shows an example diagram of the sample area image of the preset shape provided by the embodiment of the present application.

[0099] Generally, the shape of the extracted sample area image is an irregular polygon. In this embodiment, a sample area of the preset shape as shown in Figure 6 is extracted from the sample area, and then the sample area image is converted into a sample area image of the preset shape. The preset shape is set according to actual needs and is not limited here. For the convenience of understanding the preset shape in the embodiment of the present application, it is a rectangle. In this embodiment, the pixel values of the sample area of the preset shape remain unchanged, and the pixel values of the areas outside the sample area of the preset shape are all set to the second pixel value, obtaining a sample area image of the preset shape with the same size as the sample image. The pixel values of the areas outside the sample area of the preset shape are all set to the unrecognizable second pixel value, removing a large amount of unnecessary environmental information in the sample image, and thus enabling the model to more accurately identify the area of interest.

[0100] Please refer to Figure 7 , Figure 7 which shows an example diagram of the sample image to be input provided by the embodiment of the present application.

[0101] The sample region image of the preset graphic is spliced with the first number of processed binary images to obtain the to-be-input sample image, further eliminating unnecessary environmental information in the sample image and reducing the interference of environmental information on the model. In this embodiment, the first number is C, and the value of C is 3. Then, the sample region image of the preset graphic is spliced with the first number of processed binary images to obtain a four-layer to-be-input sample image. As Figure 7 shown, the white image is the image that the model needs to focus on. The first-layer image is the sample region image of the preset graphic, and the second-layer to fourth-layer images are the first number of binary images. In this embodiment, the binary image is a single-channel image data with the same height and width as the sample region image. Since both the sample image and the sample region image are RGB (Red Green Blue) images, the dimension of the to-be-input sample image is (Height, Width, 3 + C), where Height is the height of the to-be-input sample image, Width is the width of the to-be-input sample image, RGB is 3-channel data, and the length of the to-be-input sample image is 3 + C, and C is the first number, that is, the number of sub-regions of interest.

[0102] In the embodiment of the present application, the region image extraction sub-module is further configured to retain the pixel values of the region of interest and set the pixel values of the image other than the region of interest in the sample image to the second pixel value, so as to extract the sample region image from the sample image.

[0103] Retain the pixel values of the region of interest and set the pixel values of the image other than the region of interest in the sample image to the second pixel value. Extracting the sample region image from the sample image includes 3 dimensions. The length of the first dimension is the height of the image, the length of the second dimension is the width of the image, and the length of the third dimension is the three channels of RGB, that is, the third dimension is 3. The extracted sample region image has the same length, width, and height as the sample image, and at the same time removes the environmental information in the sample image, avoiding the interference of too much complex environmental information on the model judgment.

[0104] In the embodiment of the present application, the model training system 100 further includes a model construction module;

[0105] The model construction module is configured to construct a preset model based on the residual network structure;

[0106] The model training module 140 is further configured to input all the to-be-input sample images into the preset model and train the preset model using the binary cross-entropy loss function to obtain a behavior recognition model, where the preset model ignores the binary images with the processing label being the preset value.

[0107] Please refer to Figure 8 , Figure 8 which shows an example diagram of the residual network structure provided by the embodiment of the present application.

[0108] Based on the Residual Network (ResNet) structure, a preset model is constructed. As Figure 8 shown, the ResNet structure passes the input signal to the subsequent layer structure of the network through residual connections. The network can learn the residuals instead of the global features, thereby enabling the network for constructing the preset model to be trained deeper and have stronger performance.

[0109] All the to-be-input sample images are input into the preset model, and the preset model is trained using the binary cross-entropy loss function to obtain a behavior recognition model. Since the loss function in this embodiment needs to ignore the preset value, the binary cross-entropy loss function in this embodiment is:

[0110]

[0111] where Loss is the binary cross-entropy loss function, i is the i-th sample image, n is the total number of sample images, y i is the label value of the i-th sample image, y i takes an integer value in [-1, 0, 1], -1 is the preset label value, p i is the predicted probability of the i-th sample image, p i takes a real number in [0, 1]. Since y i ≠ -1, the preset model ignores and processes the binary image with the label being the preset value.

[0112] In the embodiment of the present application, the model training module 140 includes a preset model training sub-module, a model evaluation sub-module, and an identification model determination sub-module;

[0113] The preset model training sub-module is configured to input all the to-be-input sample images into the preset model, train the preset model, and obtain the trained preset model;

[0114] The model evaluation sub-module is configured to evaluate the trained preset model based on a preset evaluation function to obtain the performance index of the trained preset model;

[0115] The identification model determination sub-module is configured to determine the trained preset model as the behavior recognition model when the performance index of the trained preset model is higher than the preset performance index.

[0116] The trained preset model is evaluated based on a preset evaluation function to obtain the performance index of the trained preset model. The type of the preset evaluation function is set according to actual requirements and is not limited herein. For ease of understanding, the preset evaluation function in the embodiment of the present application is:

[0117]

[0118] Among them, F β is the binary cross-entropy loss function, Precision is the accuracy of the model, Recall is the recall rate of the model, β is a fixed coefficient, and in this embodiment, the value of the fixed coefficient is 2.

[0119] In this embodiment, the positive samples are the samples that match the true target category, and the negative samples are the samples that do not match the true target category. In this embodiment, the positive samples can be sample images including the attention sub-region, and the negative samples can be sample images that do not include the attention sub-region. According to the number of positive and negative samples, the accuracy and recall rate of the model are obtained:

[0120]

[0121] Among them, Precision is the accuracy of the model, Recall is the recall rate of the model, TP is the number of samples where the model outputs a positive sample and the sample image is truly a positive sample, FN is the number of samples where the model outputs a positive sample and the sample image is truly a negative sample, and FN is the number of samples where the model outputs a negative sample and the sample image is truly a positive sample.

[0122] The performance indicators of the model are evaluated using an evaluation function. When the performance indicators of the trained preset model are higher than the preset performance indicators, the trained preset model is determined as the behavior recognition model. When the performance indicators of the trained preset model are higher than the preset performance indicators, the model is retrained until the trained preset model with performance indicators higher than the preset performance indicators. The evaluation function can measure the effectiveness, precision, robustness, etc. of the trained model, and then optimize the parameters of the model to avoid overfitting and underfitting of the model.

[0123] In the embodiment of the present application, the model training system 100 further includes an image enhancement module;

[0124] The image enhancement module is used to perform image enhancement processing on each sample image to be input respectively to obtain the enhanced sample image, where the image enhancement processing includes at least one of rotation, flipping, translation, scaling, affine transformation, brightness adjustment, contrast adjustment, adding Gaussian noise, adding Gaussian blur, and covering part of the pixel area;

[0125] The model training module 140 is further used to input all the enhanced sample images into the preset model to train the preset model to obtain the behavior recognition model.

[0126] After splicing the binary images to obtain the to-be-input sample images, image enhancement processing is performed on each to-be-input sample image respectively to obtain the enhanced sample images. It should be understood that the way of image enhancement processing can be set according to actual needs, and can be at least one of image rotation, flipping, translation, scaling, affine transformation, brightness adjustment, contrast adjustment, adding Gaussian noise, adding Gaussian blur, covering part of the pixel area, etc., which is not limited here.

[0127] All the enhanced sample images are input into a preset model, and the preset model is trained using a loss function to obtain a behavior recognition model. After the sample images are input into the model, preprocessing the sample images in advance can further improve the generalization ability and performance of the model, reduce the over-dependence of the model on specific features, and thus avoid model overfitting.

[0128] This application provides a model training system, including: an image acquisition module for acquiring a first number of binary images of each sample image; an image processing module for performing disorder processing on the first number of binary images corresponding to each sample image, and randomly setting all pixel values of at least one binary image to a second pixel value to obtain the first number of processed binary images; an image splicing module for splicing the sample image with the first number of processed binary images for each sample image to obtain the to-be-input sample image; a model training module for inputting all the to-be-input sample images into a preset model and training the preset model to obtain a behavior recognition model. By providing a prior attention mechanism to process the images, a behavior recognition model is trained to adapt to any number of attention sub-regions within the first number. Even if the number and position of the attention sub-regions change in the actual behavior scenario, the behavior recognition model can still output accurate recognition results, thereby accurately recognizing different behavior scenarios. In addition, by adjusting the order and pixels of the binary images, more sample images for training the model can be obtained, improving the generalization ability and performance of the model, and thus more accurately recognizing different behavior scenarios.

[0129] Embodiment 2

[0130] Please refer to Figure 9 , Figure 9 which shows a schematic structural diagram of an image recognition device provided by an embodiment of this application. Figure 9 The image recognition device 200 in

[0131] includes a target region determination module 210, a first target image obtaining module 220, a second target image obtaining module 230, and a recognition result obtaining module 240.

[0132] When it is necessary to recognize a target image, determine the target attention sub-region in the acquired target image and the number of target attention sub-regions. Determine whether the number of actually concerned target attention sub-regions is less than the first number, that is, determine whether the number of target attention sub-regions is less than the maximum number of attention sub-regions.

[0133] The first target image obtaining module 220 is configured to splice the target image with all target binary images to obtain a target image to be recognized when the first number is equal to the second number.

[0134] When training a behavior recognition model based on a sample image to be input, the sample image is an RGB image. Since the RGB image corresponds to 3 channels, the height of the spliced sample image to be input is 3 + C, where C is the first number. When obtaining a behavior recognition result through the recognition model, the height of the target image to be recognized needs to correspond to 3 + C for the sample image to be input. When the first number is equal to the second number, the target image can be directly spliced with all target binary images to obtain a target image to be recognized with a height of 3 + C.

[0135] The second target image obtaining module 230 is configured to splice the target image, all target binary images and a second pixel value image of a third number to obtain a target image to be recognized when the first number is greater than the second number, where the third number is the difference between the first number and the second number.

[0136] In this embodiment, a prior attention mechanism is provided to process the target image. Specifically, when the number of target attention sub-regions is less than the first number, if the target image is directly spliced with all target binary images, the height of the obtained image will be less than 3 + C. In this embodiment, the first number is C and the second number of target attention sub-regions is Ct, and a second pixel value image of a third number with the same size as the target image is obtained. The target image, all target binary images and the second pixel value image of the third number are spliced to obtain a target image to be recognized. Specifically, when the first number is 3 and the second number is 1, the height of the target image to be recognized is 6, and the third number is 2. One target binary image is obtained, and two second pixel value images with the same size as the target image are obtained. The target image, one target binary image and two second pixel value images are spliced to obtain a target image to be recognized with a height of 6.

[0137] The recognition result obtaining module 240 is configured to input the target image to be recognized into the behavior recognition model to obtain a behavior recognition result, where the behavior recognition model is obtained through the above model training system 100.

[0138] After obtaining the target image to be recognized, directly input the target image to be recognized into the behavior recognition model to obtain the behavior recognition result. The trained behavior recognition model is adapted to any number of attention sub-regions within the first quantity. Even if the quantity and position of the attention sub-regions in the actual behavior scenario change, the behavior recognition model can still output accurate recognition results, thereby accurately recognizing different behavior scenarios.

[0139] Please refer to Figure 10 , Figure 10 which shows a flowchart of the model training method provided by an embodiment of the present application. An embodiment of the present application also provides a model training method, which is applied to the above-mentioned model training system 100. Figure 10 The model training method in

[0140] S310, obtain the first quantity of binary images corresponding to each sample image, where the sample image includes the first quantity of attention regions, each binary image corresponds to a first pixel value that can be recognized for one attention region, and the regions outside the attention region are set to a second pixel value that cannot be recognized.

[0141] S320, for the first quantity of binary images corresponding to each sample image, shuffle the first quantity of binary images, and randomly set all pixel values of at least one binary image to the second pixel value to obtain the first quantity of processed binary images.

[0142] S330, for each sample image, splice the sample image with the first quantity of processed binary images to obtain the sample image to be input.

[0143] S340, input all the sample images to be input into a preset model, train the preset model, and obtain a behavior recognition model.

[0144] Please refer to Figure 11 , Figure 11 which shows a flowchart of the image recognition method provided by an embodiment of the present application. An embodiment of the present application also provides an image recognition method, which is applied to the above-mentioned image recognition device 200. The image recognition method includes:

[0145] S410, determine the target attention sub-region in the target image and the quantity of the target attention sub-region.

[0146] S420, when the first quantity is equal to the second quantity, splice the target image with all the target binary images to obtain the target image to be recognized.

[0147] S430. When the first quantity is greater than the second quantity, splice the target image, all target binary images, and the second pixel value images of the third quantity to obtain the target image to be recognized, where the third quantity is the difference between the first quantity and the second quantity.

[0148] S440. Input the target image to be recognized into the behavior recognition model to obtain the behavior recognition result, where the behavior recognition model is obtained by the above-mentioned model training system 100.

[0149] The embodiments of the present application also provide a computing device, including:

[0150] A memory configured to store instructions;

[0151] A processor configured to call instructions from the memory and capable of implementing the above-mentioned model training system or the above-mentioned image recognition device when executing the instructions.

[0152] The processor contains a kernel, and the kernel retrieves the corresponding program unit from the memory. One or more kernels can be set, and by adjusting the kernel parameters, the problem of being unable to accurately identify different behavior scenarios can be solved.

[0153] The memory may include non-permanent memory in a computer-readable medium, forms such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM), and the memory includes at least one storage chip.

[0154] The embodiments of the present application also provide a machine-readable storage medium, on which instructions are stored, and the instructions are used to make the machine execute the above-mentioned model training system or the above-mentioned image recognition device.

[0155] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0156] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, as well as the combination of flows and / or blocks in the flowchart and / or block diagram. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device produce a means for implementing the functions specified in one or more flows Figure 1 one or more flows and / or blocks Figure 1 or a means for implementing the functions specified in one or more blocks.

[0157] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction means that implements the functions specified in one or more flows Figure 1 one or more flows and / or blocks Figure 1 or a means for implementing the functions specified in one or more blocks.

[0158] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows Figure 1 one or more flows and / or blocks Figure 1 or a means for implementing the functions specified in one or more blocks.

[0159] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and a memory.

[0160] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.

[0161] Machine-readable storage media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0162] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.

[0163] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included within the scope of the claims of the present application.

Claims

1. A model training system, characterized in that, The model training system includes an image acquisition module, an image processing module, an image stitching module, and a model training module; The image acquisition module is configured to acquire a first number of binary images corresponding to each sample image, where the sample image includes a first number of regions of interest, each of the binary images has a first pixel value set to be recognizable for one of the regions of interest, and regions other than the regions of interest are set to a second pixel value that is not recognizable; The image processing module is configured to, for the first number of binary images corresponding to each sample image, scramble the order of the first number of binary images, and randomly set all pixel values of at least one of the binary images to the second pixel value, to obtain a first number of processed binary images; The image stitching module is configured to, for each sample image, stitch the sample image and the first number of processed binary images together to obtain a to-be-input sample image; The model training module is configured to input all the to-be-input sample images into a preset model, train the preset model, and obtain a behavior recognition model.

2. The model training system according to claim 1, wherein The sample image includes a plurality of labels, and the model training system further includes a label processing module; The label processing module is configured to sort the plurality of labels based on the order of the first number of processed binary images, and set the label corresponding to a binary image with all pixel values being the second pixel value to a preset value.

3. The model training system according to claim 2, characterized in that, The model training system further includes a model construction module; The model construction module is configured to construct a preset model based on a residual network structure; The model training module is further configured to input all the to-be-input sample images into the preset model, and train the preset model using a binary cross-entropy loss function to obtain a behavior recognition model, where the preset model ignores the binary images with labels being the preset value.

4. The model training system according to claim 1, wherein The image stitching module includes a region image extraction sub-module and a sample image obtaining sub-module; The region image extraction sub-module is configured to extract a sample region image from the sample image, where the sample region image includes the first number of regions of interest; The sample image obtaining sub-module is configured to stitch the sample region image and the first number of processed binary images together to obtain a to-be-input sample image. Preferably, the sample image obtaining sub-module is further configured to convert the sample region image into a sample region image of a preset shape; Stitch the sample region image of the preset shape and the first number of processed binary images together to obtain a to-be-input sample image.

5. The model training system according to claim 1, wherein The first number is the maximum number of regions of interest determined by the behavior recognition model.

6. The model training system according to claim 1, wherein The image processing module includes a scrambling processing sub-module and a pixel processing sub-module; The scrambling processing sub-module is configured to, for the first number of binary images corresponding to each sample image, scramble the arrangement order of the first number of binary images, and scramble the positions of the regions of interest in each of the binary images, to obtain a first number of scrambled binary images; A pixel processing sub-module, configured to randomly set all pixel values of at least one of the shuffled binary images to a second pixel value, obtaining a first number of processed binary images.

7. The model training system according to claim 1, wherein The model training system further includes an image enhancement module; The image enhancement module is configured to perform image enhancement processing on each of the to-be-input sample images respectively, obtaining enhanced sample images, wherein the image enhancement processing includes at least one of rotation, flipping, translation, scaling, affine transformation, brightness adjustment, contrast adjustment, Gaussian noise addition, Gaussian blur addition, and partial pixel region covering; The model training module is further configured to input all the enhanced sample images into a preset model, training the preset model to obtain a behavior recognition model.

8. An image recognition device, characterized in that, The image recognition device includes a target area determination module, a first target image obtaining module, a second target image obtaining module, and a recognition result obtaining module; The target area determination module is configured to obtain a second number of target attention sub-areas in a target image, and a target binary image corresponding to each target attention sub-area; The first target image obtaining module is configured to, when the first number is equal to the second number, splice the target image and all the target binary images to obtain a target image to be recognized; The second target image obtaining module is configured to, when the first number is greater than the second number, splice the target image, all the target binary images, and a second pixel value image with a third number, where the third number is the difference between the first number and the second number, to obtain a target image to be recognized; The recognition result obtaining module is configured to input the target image to be recognized into the behavior recognition model to obtain a behavior recognition result, where the behavior recognition model is obtained by the model training system according to any one of claims 1 to 7.

9. A model training method, characterized in that, Applied to the model training system according to any one of claims 1 to 7, the model training method includes: Obtaining a first number of binary images corresponding to each sample image, where the sample image includes a first number of attention areas, each binary image has one of the attention areas set to a recognizable first pixel value, and areas other than the attention areas are set to an unrecognizable second pixel value; For the first number of binary images corresponding to each sample image, performing a shuffling process on the first number of binary images, and randomly setting all pixel values of at least one of the binary images to a second pixel value, obtaining a first number of processed binary images; For each of the sample images, splicing the sample image and the first number of processed binary images to obtain a to-be-input sample image; Inputting all the to-be-input sample images into a preset model, training the preset model to obtain a behavior recognition model.

10. An image recognition method, characterized in that, Applied to the image recognition device according to claim 8, the image recognition method includes: Obtain the second quantity of target attention sub-regions in the target image, and the target binary images corresponding to each target attention sub-region; When the first quantity is equal to the second quantity, splice the target image with all the target binary images to obtain the target image to be recognized; When the first quantity is greater than the second quantity, splice the target image, all the target binary images and the second pixel value images of the third quantity to obtain the target image to be recognized, where the third quantity is the difference between the first quantity and the second quantity; Input the target image to be recognized into the behavior recognition model to obtain the behavior recognition result, where the behavior recognition model is obtained by the model training system according to any one of claims 1 to 7.