Method, apparatus, and product for processing a screen image with a long-tail distribution characteristic
By constructing the initial image set, expanding the sample set, extracting intermediate features, remixing and enhancing features, building target loss functions, and iteratively training the flower screen recognition model, the problem of difficult to subclassify and detect flower screen images in the existing technology is solved, and high-accurate flower screen type recognition is achieved.
Patent Information
- Application Number
- CN202410994312.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-23
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-07-23
AI Technical Summary
The prior art is difficult to realize subclassification detection of flower screen images, which affects the visual analysis model's identification of target objects in video surveillance images.
A method of flower screen image processing with long tail distribution characteristics is proposed. By constructing the initial image set, expanding the sample set, extracting intermediate features, remixing and enhancing features, constructing target loss functions, and iteratively training the flower screen recognition model, the subclassification detection of flower screen image types is achieved.
The subclassification detection of the flower screen image type is realized, the generalization ability and classification accuracy of the model are improved, and the generalization error caused by uneven sample size is reduced.
Smart Images

Figure CN119068233B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of image processing, and in particular, to a method, device, and product for processing a mosaic image with a long-tailed distribution characteristic. Background Art
[0002] With the improvement of the development level of information technology, video surveillance systems based on computer vision analysis are widely used in engineering construction. The video surveillance system mostly adopts an installation method of multi-point layout in outdoor construction sites, and uses a vision analysis model to identify and analyze the quantity, behavior postures, etc. of target objects such as construction machinery and construction personnel in the picture. Due to the influence of bandwidth, network transmission, environmental interference sources, unstable video frame rate, or video decoding format, some data may be lost during the video transmission process, resulting in abnormal local or whole-image color blocks in the image, that is, a mosaic image. The mosaic image will affect the identification of target objects in the video surveillance image by the vision analysis model. Therefore, before performing vision analysis on the video surveillance image, it is necessary to first carry out discriminant analysis of the mosaic image.
[0003] The current mosaic image recognition scheme can only achieve binary classification detection of mosaic images, and most of the current mosaic model training schemes focus on reducing the dependence of the model on the training data set or improving the accuracy of binary classification detection of mosaic images, and cannot achieve fine classification detection of mosaic types. However, different types of mosaic images have different degrees of influence on whether the subsequent vision analysis model can identify the quantity, behavior postures, etc. of target objects in the image, and also directly affect the training effect of the subsequent vision analysis model for target objects. Therefore, it is very important to achieve subdivision detection of specific types of mosaic images. Summary of the Invention
[0004] In view of this, the present application aims to propose a method, device, and product for processing a mosaic image with a long-tailed distribution characteristic to achieve fine classification detection of the type of mosaic image.
[0005] To achieve the above object, the technical solution of the present application is as follows:
[0006] The first aspect of the embodiment of the present application provides a method for processing a mosaic image with a long-tailed distribution characteristic, and the method includes:
[0007] Construct an initial image set, where the initial image set includes the following types of mosaic images: large-area solid-color mosaic images, small-area solid-color mosaic images, large-area striped mosaic images, small-area striped mosaic images, color-distorted mosaic images, and noisy mosaic images;
[0008] Expand the quantity of color-distorted mosaic images and noisy mosaic images in the initial image set to generate a first sample set;
[0009] Randomly extract multiple different types of screen - flower images from the first sample set, and extract the intermediate features of each screen - flower image;
[0010] Re - mix and enhance the intermediate features of different types of screen - flower images to generate mixed features;
[0011] Based on the intermediate features and the mixed features, construct an objective loss function;
[0012] Based on the objective loss function and the first sample set, perform iterative training on the screen - flower recognition model to obtain a trained screen - flower recognition model;
[0013] Input the image to be recognized into the screen - flower recognition model, and use the screen - flower recognition model to determine whether the image to be recognized is a screen - flower image and the type of the screen - flower image.
[0014] Optionally, construct an initial image set, including:
[0015] Extract frames from the sample video at a first time interval to obtain an initial image set; the sample video is a video with screen - flower scenes;
[0016] Remove non - screen - flower images from the initial image set;
[0017] Add labels corresponding to the screen - flower types to all images.
[0018] Optionally, expand the number of color - distorted screen - flower images and noisy screen - flower images in the initial image set to generate a first sample set, including:
[0019] Collect images with color - distorted screen - flowers, crop the screen - flower pixel regions in the images to obtain multiple small - sized regional images; randomly mix the multiple small - sized regional images with non - screen - flower images to generate new color - distorted screen - flower images;
[0020] Add noise to non - screen - flower images to generate new noisy screen - flower images;
[0021] Add the new color - distorted screen - flower images and the new noisy screen - flower images to the initial image set to generate a first sample set.
[0022] Optionally, before adding labels corresponding to the screen - flower types to all images, further include:
[0023] In the case where the coverage area of screen - flower pixels in a screen - flower image of the solid - color screen - flower type is not less than a first threshold, determine that the screen - flower image is a large - area solid - color screen - flower;
[0024] In a solid - color screen - like image, when the coverage area of the screen - like pixels is less than the first threshold, it is determined that the screen - like image is a small - area solid - color screen - like image;
[0025] In a stripe screen - like image, when the coverage area of the screen - like pixels is not less than the first threshold, it is determined that the screen - like image is a large - area stripe screen - like image;
[0026] In a stripe screen - like image, when the coverage area of the screen - like pixels is less than the first threshold, it is determined that the screen - like image is a small - area stripe screen - like image.
[0027] Optionally, the screen - like image recognition model includes: a first feature extraction network, a second feature extraction network, and a classification module; the first feature extraction network is used to extract intermediate features of each screen - like image;
[0028] Based on the intermediate features and the mixed features, a target loss function is constructed, including:
[0029] Through the second feature extraction network and the classification module, each intermediate feature is processed to obtain a first probability distribution of the screen - like image corresponding to each intermediate feature belonging to each type of screen - like image;
[0030] The intermediate features of images of different screen - like types are re - mixed and enhanced to obtain mixed features; through the second feature extraction network and the classification module, each mixed feature is processed to obtain a second probability distribution of the screen - like image corresponding to each mixed feature belonging to each type of screen - like image;
[0031] Based on the first probability distribution and the second probability distribution, a target loss function is constructed.
[0032] Optionally, the method for processing screen - like images with long - tail distribution characteristics further includes:
[0033] For each mixed feature, obtain the label of the screen - like image corresponding to the intermediate feature that generates the mixed feature;
[0034] According to the label of the screen - like image corresponding to the intermediate feature, determine the label of the screen - like image corresponding to the mixed feature.
[0035] Optionally, the method for processing screen - like images with long - tail distribution characteristics further includes:
[0036] When the image to be recognized is a non - screen - like image, add the image to be recognized to the second sample set;
[0037] When the image to be recognized is a screen - like image and its type is a noise screen - like image, add the image to be recognized to the second sample set;
[0038] When the image to be recognized is a garbled screen image and its type is a small - area solid - color garbled screen, a small - area striped garbled screen, or a color - distorted garbled screen, further determine whether the target object in the image to be recognized is completely covered by garbled screen pixels; if the target object is not completely covered by garbled screen pixels, add the image to be recognized to the second sample set;
[0039] Based on the second sample set, train a visual analysis model; the visual analysis model is used to perform visual analysis tasks on the target object in the image to be recognized.
[0040] Optionally, the method for processing garbled screen images with long - tail distribution characteristics further includes:
[0041] Calculate the proportion of the number of garbled screen images to the number of images to be recognized as the first proportion;
[0042] Compare the first proportion with a second threshold, and generate a first warning prompt when the first proportion is not less than the second threshold.
[0043] Optionally, calculating the proportion of the number of garbled screen images to the number of images to be recognized as the first proportion includes:
[0044] Determine the number of images to be recognized within the current nearest target time period;
[0045] Determine the number of garbled screen images within the target time period;
[0046] Calculate the proportion of the number of garbled screen images to the number of images to be recognized within the target time period as the first proportion.
[0047] Optionally, the method for processing garbled screen images with long - tail distribution characteristics further includes:
[0048] Calculate the proportion of the number of garbled screen images of the target type to the number of images to be recognized as the second proportion; the target type includes: large - area solid - color garbled screen, large - area striped garbled screen, color - distorted garbled screen, and noisy garbled screen;
[0049] Compare the second proportion with a second threshold, and generate a second warning prompt when the second proportion is not less than the second threshold.
[0050] Optionally, calculating the proportion of the number of garbled screen images of the target type to the number of images to be recognized as the second proportion includes:
[0051] Determine the number of images to be recognized within the current nearest target time period;
[0052] Determine the number of screen freeze images of the target type within the target time period;
[0053] Calculate the proportion of the number of screen freeze images of the target type within the target time period to the number of images to be recognized, as the second proportion.
[0054] According to the second aspect of the embodiments of the present application, there is provided a screen freeze image processing device with long-tail distribution characteristics, which is used to implement the screen freeze image processing method with long-tail distribution characteristics provided in the first aspect of the embodiments of the present application. The device includes:
[0055] A first sample construction module, configured to construct an initial image set, where the initial image set includes the following types of screen freeze images: large-area solid color screen freeze images, small-area solid color screen freeze images, large-area striped screen freeze images, small-area striped screen freeze images, color distortion screen freeze images, and noise screen freeze images; expand the number of color distortion screen freeze images and noise screen freeze images in the initial image set to generate a first sample set;
[0056] A first training module, configured to randomly extract multiple different types of screen freeze images from the first sample set, extract intermediate features of each screen freeze image; remix and enhance the intermediate features of different types of screen freeze images to generate mixed features; based on the intermediate features and the mixed features, construct a target loss function; based on the target loss function and the first sample set, perform iterative training on the screen freeze recognition model to obtain a trained screen freeze recognition model;
[0057] A detection module, configured to input the image to be recognized into the screen freeze recognition model, and determine whether the image to be recognized is a screen freeze image and the type of the screen freeze image through the screen freeze recognition model.
[0058] According to the third aspect of the embodiments of the present application, there is provided a computer program product, including a computer program, which when executed by a processor, implements the steps of the method described in the first aspect of the embodiments of the present application.
[0059] According to the fourth aspect of the embodiments of the present application, there is provided a computer-readable storage medium, on which a computer program is stored, which when executed by a processor, implements the steps of the method described in the first aspect of the embodiments of the present application.
[0060] According to the fifth aspect of the embodiments of the present application, there is provided an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, which when executed by the processor, implements the steps of the method described in the first aspect of the embodiments of the present application.
[0061] Adopt the method for processing a flower screen image with a long-tail distribution characteristic provided by this application, and pre-train a flower screen recognition model to detect whether the image to be recognized is a flower screen image and determine the specific type of the flower screen image. During the process of training the flower screen recognition model, first construct an initial image set with various types of flower screen images, and then expand the number of samples of two types of flower screen images with fewer samples in the initial image set to generate a first sample set. Randomly extract multiple different types of flower screen images from the first sample set, extract intermediate features, and perform re-mixing enhancement based on the intermediate features of different types to obtain mixed features. Construct an objective loss function based on the intermediate features and the mixed features, and perform iterative training on the flower screen recognition model.
[0062] Using this method can perform a fine classification detection of the flower screen type of the image and determine the specific type to which the flower screen image belongs. Moreover, when constructing the model training samples by this method, the samples of the minority classes (color distortion flower screen, noise flower screen) of the flower screen type are expanded by synthesizing samples, so as to move the decision boundary towards the majority class, reduce the generalization error caused by the imbalance of the number of samples, and improve the classification accuracy of the model. Moreover, during the training process, the intermediate features extracted from different types of flower screen images are re-mixed and enhanced to further improve the generalization ability of the model and the classification accuracy of the model. Description of the Drawings
[0063] In order to more clearly illustrate the technical solutions of the embodiments of this application, the drawings required to be used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0064] Figure 1 It is a flowchart of the method for processing a flower screen image with a long-tail distribution characteristic proposed in an embodiment of this application;
[0065] Figure 2 It is a schematic diagram of some types of flower screen images in an embodiment of this application;
[0066] Figure 3 It is a schematic diagram of training a flower screen recognition model in an embodiment of this application;
[0067] Figure 4 It is a schematic diagram of the device for processing a flower screen image with a long-tail distribution characteristic proposed in an embodiment of this application;
[0068] Figure 5 It is a schematic diagram of an electronic device proposed in an embodiment of this application. Detailed Embodiments
[0069] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts belong to the scope of protection of the present application.
[0070] It should be understood that the "one embodiment" or "an embodiment" mentioned throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of the present application. Therefore, the "in one embodiment" or "in an embodiment" that appears throughout the specification does not necessarily refer to the same embodiment. In addition, these specific features, structures, or characteristics can be combined in one or more embodiments in any suitable manner.
[0071] In various embodiments of the present application, it should be understood that the magnitudes of the serial numbers of the following processes do not mean the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0072] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all the implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects detailed in the present application.
[0073] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.
[0074] Next, the present application will be described in detail with reference to the accompanying drawings and in conjunction with the embodiments.
[0075] Figure 1 is a flowchart of a method for processing a flower screen image with a long-tail distribution characteristic proposed in an embodiment of the present application. As Figure 1 shown, the method includes:
[0076] S1: Construct an initial image set, where the initial image set includes the following types of flower screen images: large-area solid-color flower screen images, small-area solid-color flower screen images, large-area striped flower screen images, small-area striped flower screen images, color-distorted flower screen images, and noisy flower screen images;
[0077] Expand the number of color-distorted flower screen images and noisy flower screen images in the initial image set to generate a first sample set;
[0078] Randomly extract multiple different types of screen - flower images from the first sample set, and extract the intermediate features of each screen - flower image;
[0079] Re - mix and enhance the intermediate features of different types of screen - flower images to generate mixed features;
[0080] Based on the intermediate features and the mixed features, construct a target loss function;
[0081] Based on the target loss function and the first sample set, iteratively train the screen - flower recognition model to obtain a trained screen - flower recognition model;
[0082] S2: Input the image to be recognized into the screen - flower recognition model, and use the screen - flower recognition model to determine whether the image to be recognized is a screen - flower image and the type of the screen - flower image.
[0083] In this embodiment, by pre - constructing and training a screen - flower recognition model, fine - classification detection of the screen - flower type of an image is realized. Given that the sample images of each type of screen - flower image are unevenly distributed and have a long - tail distribution characteristic. For example, color - distortion screen - flower images and noise screen - flower images rarely appear in actual application scenarios, resulting in a small number of training samples for these two types of screen - flower images. The long - tail distribution of training samples will cause the model to have a higher classification accuracy in the head categories (categories with a larger number of samples), while the classification accuracy in the tail categories (categories with a small number of samples) decreases, making the overall recognition accuracy of the model insufficient. In order to reduce the generalization error caused by the imbalance of the sample quantity, in this embodiment, when constructing the sample set, the screen - flower image samples of the types with a small number of samples are expanded to generate the first sample set.
[0084] In this embodiment, an initial image set is first created, which specifically includes the following six types of screen-failure images: large-area solid-color screen-failure images, small-area solid-color screen-failure images, large-area striped screen-failure images, small-area striped screen-failure images, color-distorted screen-failure images, and noisy screen-failure images. Further, the number of samples of two types of screen-failure images with relatively small sample numbers (i.e., color-distorted screen-failure images and noisy screen-failure images) is expanded to obtain a first sample set. Different types of screen-failure images are randomly selected from the first sample set to iteratively train the screen-failure recognition model. Specifically, intermediate features of each screen-failure image are extracted through the screen-failure recognition model, and then the intermediate features of different types of screen-failure images are further remixed and enhanced to generate mixed features. A target loss function is constructed based on the intermediate features and the mixed features, and then the target loss function and the screen-failure images randomly selected from the first sample set are used to iteratively train the screen-failure recognition model to calculate the optimal model parameters and obtain a high-precision screen-failure recognition model. The image to be recognized is input into the trained screen-failure recognition model to determine whether the image to be recognized is a screen-failure image and to determine the specific screen-failure type of the screen-failure image, so as to realize the fine classification detection of the screen-failure type of the image to be recognized.
[0085] By using this method, it is possible to realize the fine classification detection of the screen-failure type of the image. Moreover, the screen-failure recognition model trained in this method has strong generalization ability, which improves the accuracy of the screen-failure fine classification detection.
[0086] As an implementation manner of the present application, an initial image set is constructed, including:
[0087] Frames are extracted from the sample video at a first time interval to obtain an initial image set; the sample video is a video with a screen-failure picture;
[0088] Non-screen-failure images are removed from the initial image set;
[0089] Labels corresponding to the screen-failure type are added to all images.
[0090] In this embodiment, an initial image set is created based on the sample video with screen failure. Specifically, a video containing a screen-failure picture is collected, frames are extracted from the video at a first time interval to obtain an initial image set, and the images without screen failure (i.e., non-screen-failure images) in the initial image set are removed, and the images with screen failure are retained. Optionally, the sample video can be a video for computer vision analysis tasks. For example, videos for visual analysis tasks such as object detection, object analysis, and action classification can be used as the sample video.
[0091] In this embodiment, the possible types of screen freeze are divided into six types, and when constructing the initial image set, it is ensured that all types of screen freeze images are included in the initial image set. Specifically, the types of screen freeze include the following six types: large-area solid color screen freeze, small-area solid color screen freeze, large-area stripe screen freeze, small-area stripe screen freeze, color distortion screen freeze, and noise screen freeze. Figure 2 is a schematic diagram of some types of screen freeze images in an embodiment of the present application. As Figure 2 shown, where (a) is a large-area solid color screen freeze image, (b) is a small-area solid color screen freeze image, (c) is a large-area stripe screen freeze image, (d) is a small-area stripe screen freeze image, and (e) is a color distortion screen freeze image. In this embodiment, according to the above six types of screen freeze, labels of the screen freeze types to which all images in the initial image set belong are added.
[0092] As an implementation manner of the present application, before adding labels of the corresponding screen freeze types to all images, it further includes:
[0093] In the screen freeze image of the solid color screen freeze type, when the covered area of the screen freeze pixels is not less than the first threshold, it is determined that the screen freeze image is a large-area solid color screen freeze;
[0094] In the screen freeze image of the solid color screen freeze type, when the covered area of the screen freeze pixels is less than the first threshold, it is determined that the screen freeze image is a small-area solid color screen freeze;
[0095] In the screen freeze image of the stripe screen freeze type, when the covered area of the screen freeze pixels is not less than the first threshold, it is determined that the screen freeze image is a large-area stripe screen freeze;
[0096] In the screen freeze image of the stripe screen freeze type, when the covered area of the screen freeze pixels is less than the first threshold, it is determined that the screen freeze image is a small-area stripe screen freeze.
[0097] In this embodiment, when adding labels of the screen freeze types to which the images in the initial image set belong, it is necessary to determine the large-area screen freeze type and the small-area screen freeze type in the solid color screen freeze images and the stripe screen freeze images. Specifically, for the solid color screen freeze image or the stripe screen freeze image, the covered area of the screen freeze pixels in the image is compared with the first threshold. If the covered area of the screen freeze pixels reaches the first threshold, it is determined as a large-area screen freeze image. If the covered area of the screen freeze pixels is less than the first threshold, it is determined as a small-area screen freeze image. The first threshold can be set according to the needs in actual applications. In this embodiment, the first threshold can be set to 1 / 3 of the image area size. For example, for a stripe screen freeze image, if the covered area of the screen freeze pixels in it reaches 1 / 3 of the image area, it is determined that the image is a large-area stripe screen freeze image.
[0098] As an implementation manner of the present application, the number of color-distorted and noise-screened images in the initial image set is expanded to generate a first sample set, including:
[0099] Collect images with color-distorted and noise-screened, crop the screened pixel regions in the images to obtain multiple small-sized regional images; randomly mix the multiple small-sized regional images with non-screened images to generate new color-distorted and noise-screened images;
[0100] Add noise to non-screened images to generate new noise-screened images;
[0101] Add the new color-distorted and noise-screened images to the initial image set to generate a first sample set.
[0102] In this embodiment, in order to reduce the generalization error caused by the imbalance of the sample quantity, the number of image samples of two types of screened images with fewer samples in the initial image set is expanded to generate a first sample set. The number of image samples of each type of screened image in the first sample set is more balanced than that in the initial image set. Training the screened image recognition model based on the first image set can improve the generalization ability of the model.
[0103] Specifically, in the initial image set, the number of color-distorted and noise-screened images is small. In this embodiment, different methods are used to expand these two types of sample images.
[0104] For the case where the proportion of color-distorted and noise-screened images is small, collect image samples with color-distorted and noise-screened phenomena, crop the color-distorted and noise-screened pixel regions in the images to obtain multiple small-sized regional images, and then randomly fill and mix these small-sized regional images onto non-screened normal images to generate new color-distorted and noise-screened image samples, and add labels to identify them as color-distorted and noise-screened. It should be noted that the normal images used to randomly mix the small-sized screened pixel region images can be video frames without screening extracted from sample videos, or other images without screening.
[0105] For the case where the proportion of noise-screened images is small, add noise to non-screened normal images to generate new noise-screened image samples, and add labels to identify them as noise-screened. The added noise can be salt-and-pepper noise (impulse noise), noise obeying a Gaussian distribution.
[0106] In this embodiment, by expanding the image samples of the less common types of screened images, the long-tail distribution characteristic of the screened image types in the initial image set is alleviated, the imbalance of the training samples is reduced, and thus the generalization ability of the screened image recognition model is improved.
[0107] As an implementation manner of the present application, the screen freeze recognition model includes: a first feature extraction network, a second feature extraction network, and a classification module; the first feature extraction network is used to extract intermediate features of each screen freeze image;
[0108] Based on the intermediate features and the mixed features, a target loss function is constructed, including:
[0109] Through the second feature extraction network and the classification module, each intermediate feature is processed to obtain a first probability distribution of the screen freeze image corresponding to each intermediate feature belonging to each type of screen freeze image;
[0110] The intermediate features of images of different screen freeze types are re - mixed and enhanced to obtain mixed features; through the second feature extraction network and the classification module, each mixed feature is processed to obtain a second probability distribution of the screen freeze image corresponding to each mixed feature belonging to each type of screen freeze image;
[0111] Based on the first probability distribution and the second probability distribution, a target loss function is constructed.
[0112] In one embodiment, a screen freeze recognition model is constructed, including: a first feature extraction network, a second feature extraction network, and a classification module. Figure 3 It is a schematic diagram of training a screen freeze recognition model in an embodiment of the present application. As Figure 3 shown, each time the model is trained, a training image set is randomly sampled from the first sample set, and the training image set is input into the first feature extraction network to obtain intermediate features of each screen freeze image. The intermediate features of different screen freeze images are re - mixed and enhanced to obtain mixed features, and then the intermediate features and the mixed features are respectively passed through the second feature extraction network and the classification module to obtain a first probability distribution and a second probability distribution of the image corresponding to each of the intermediate features and the mixed features belonging to each type of screen freeze image. Based on the first probability distribution and the second probability distribution, a target loss function is constructed, and based on the target loss function, the parameters of the model are updated using the backpropagation algorithm. Through multiple iterative trainings, optimal model parameters are obtained.
[0113] Specifically, in this embodiment, the samples of 6 types of screen freeze are respectively marked with integers from 1 to 6 for the screen freeze types: large - area solid - color screen freeze is marked as 1; small - area solid - color screen freeze is marked as 2; large - area striped screen freeze is marked as 3; small - area striped screen freeze is marked as 4; color - distortion screen freeze is marked as 5; noise screen freeze is marked as 6. The images in the training image set are input into the first feature extraction network to obtain intermediate features feature, and the expression is as follows:
[0114] feature i = f(x i ) ;
[0115] feature j = f(x j );
[0116] where i and j are different integers from 1 to 6, and x i , x j represent the screen freeze image samples of the i-th and j-th types of screen freeze respectively; f() represents the feature extraction function of the first feature extraction network, and feature i , feature j represent the intermediate feature representations obtained after being calculated by the first feature extraction network. Optionally, the first feature extraction network can be constructed based on convolutional neural networks such as the MBConv (Mobile Inverted Bottleneck Convolution) network and the Resnet residual network.
[0117] Re - mix and enhance the intermediate features of different types of screen freeze images to obtain the mixed feature mixFeature as follows:
[0118] mixFeature = λ x feature i + (1 - λ x )feature j ;
[0119] where λ x represents the feature mixing factor, and its value range is [-1, 1].
[0120] Pass the re - mixed and enhanced mixed feature mixFeature and the intermediate feature feature of the original screen freeze image through the second feature extraction network and the classification module respectively, and obtain the screen freeze category feature vector mixZ after re - mixed enhancement and the screen freeze category feature vector Z of the original input screen freeze image based on the following expressions:
[0121] Z = G(F(feature));
[0122] mixZ = G(F(mixFeature));
[0123] where F() represents the feature extraction function of the second feature extraction network, and G() represents the classifier function of the classification module. Optionally, the second feature extraction network in this embodiment can be constructed based on convolutional neural networks such as the MBConv module and the Resnet residual network, and the classification module can be constructed based on a fully - connected layer.
[0124] Based on the screen freeze category feature vectors, calculate the probability distribution of the predicted screen freeze image belonging to each screen freeze type, which is used to analyze the possibility of the prediction result of the screen freeze type. The calculation expression is as follows:
[0125]
[0126] Among them, p represents the probability distribution of the screen freeze image belonging to a type calculated based on the intermediate features / mixed features (i.e., the first probability distribution and the second probability distribution). represents the predicted label of the screen freeze type of the image sample. Z i 、Z j respectively represent the values of the i-th and j-th items in the screen freeze category feature vector of the original input training image, and mixZ i 、mixZ j respectively represent the values of the i-th and j-th items in the screen freeze category feature vector after re-mixed enhancement; C represents the total number of screen freeze types, and n i 、n j represent the number of training image samples of the i-th and j-th screen freeze types, and n max represents the number of screen freeze image samples of the screen freeze type with the largest proportion of the image sample number.
[0127] Based on the first probability distribution and the second probability distribution, construct the target loss function L for training the screen freeze recognition model:
[0128]
[0129] In this embodiment, each time the screen freeze recognition model is trained, the screen freeze images in the first sample set are randomly selected to generate a training image set. Through the training image set and the target loss function, the screen freeze recognition model is iteratively trained. During the iterative training process, the target loss function is continuously optimized, and the optimal model parameters are calculated. Finally, a screen freeze recognition model for fine classification detection of screen freeze types is obtained.
[0130] After the screen freeze recognition model is trained, the image to be recognized is input into the model. The model calculates the screen freeze category feature vector and outputs the screen freeze type with the highest probability corresponding to the image to be recognized, so as to determine whether the image to be recognized is a screen freeze image and the specific screen freeze type it belongs to. Optionally, the image to be recognized can be a video frame image extracted from the real-time video stream at the second time interval. In this embodiment, the screen freeze recognition model can be used for screen freeze detection of images or screen freeze detection of video streams.
[0131] As an implementation manner of the present application, the method for processing screen freeze images with long-tailed distribution characteristics further includes:
[0132] For each mixed feature, obtain the label of the screen freeze image corresponding to the intermediate feature that generates the mixed feature.
[0133] Determine the label of the screen freeze image corresponding to the mixed feature according to the label of the screen freeze image corresponding to the intermediate feature.
[0134] In this embodiment, after mixing the intermediate features to obtain the mixed features, it is also necessary to determine the labels corresponding to the respective mixed features. The training image samples x of the i-th and j-th types of screen freeze i , x j are remixed and enhanced at the label level. First, determine the label mixing factor λ through the following expression y :
[0135]
[0136] where n i , n j respectively represent the numbers of screen freeze image samples of the i-th and j-th types of screen freeze, k is a set category ratio threshold, and τ is a set feature-level mixing factor threshold, and the value range is (0, 1). For example, k is taken as 3 and τ is taken as 0.5.
[0137] Based on the label mixing factor λ y , calculate the label mixY corresponding to the mixed feature according to the following expression:
[0138] mixY = λ y y i + (1 - λ y )y j ;
[0139] where y i , y j respectively represent the labels of the screen freeze types corresponding to the training image samples x of the i-th and j-th types of screen freeze i , x j (that is, the labels corresponding to the intermediate features feature i , feature j ), and mixY represents the label of the mixed feature mixFeature obtained after the remixing and enhancement calculation (that is, the label of the screen freeze type corresponding to the screen freeze image after mixing the i-th and j-th types of training images).
[0140] In this embodiment, by mixing the intermediate features of different types of screen freeze images to obtain the mixed features, and processing the mixed features through the second feature extraction network and the classification module, the screen freeze recognition model can be used to detect the mixed screen freeze images with multiple types of screen freeze existing simultaneously, further improving the generalization ability of the model.
[0141] As an implementation manner of the present application, the method for processing screen freeze images with long-tail distribution characteristics further includes:
[0142] When the image to be recognized is a non - blank screen image, add the image to be recognized to the second sample set;
[0143] When the image to be recognized is a blank screen image and its type is noise blank screen, add the image to be recognized to the second sample set;
[0144] When the image to be recognized is a blank screen image and its type is small - area solid - color blank screen, small - area stripe blank screen or color - distortion blank screen, further determine whether the target object in the image to be recognized is completely covered by blank - screen pixel points; if the target object is not completely covered by blank - screen pixel points, add the image to be recognized to the second sample set;
[0145] Based on the second sample set, train a visual analysis model; the visual analysis model is used to perform visual analysis tasks on the target object in the image to be recognized.
[0146] In one embodiment, use a blank - screen recognition model to detect the image to be recognized, and construct a second sample set according to the detection result. The second sample set can be used for training
[0147] Specifically, according to the detection result of the blank - screen recognition model for the image to be recognized, when the recognition result of the image to be detected is a normal image (i.e., non - blank screen image) or a noise blank - screen image, add the image to the second image set. For small - area solid - color blank - screen images, small - area stripe blank - screen images or color - distortion blank - screen images, it is necessary to further determine whether the target object (e.g.,) in the image is completely covered by blank - screen pixel points (i.e., solid - color pixel points, stripe pixel points, color - distortion pixel points). If it is not completely covered, add the image to the second sample set. When a certain amount of blank - screen image samples are included in the second sample set, the visual analysis model can be trained to identify and analyze the target object in the presence of a blank screen without affecting the learning of the target object's features, improving the robustness of the visual analysis model applied to the actual scenario. If the target object in the image is completely covered by blank - screen pixel points, remove the image. For large - area solid - color blank - screen images or large - area stripe blank - screen images, since the area of large - area solid - color blank - screen images or large - area stripe blank - screen images covering the image is large, it has a great impact on the visual analysis model to identify the target object and actions in the image. Therefore, large - area solid - color blank - screen images and large - area stripe blank - screen images are removed when constructing the second sample set.
[0148] As an implementation manner of the present application, the blank - screen image processing method with long - tail distribution characteristics further includes:
[0149] Calculate the ratio of the number of blank - screen images to the number of images to be recognized as the first ratio;
[0150] Compare the first proportion with a second threshold, and generate a first warning prompt when the first proportion is not less than the second threshold.
[0151] In one embodiment, a trained screen freeze recognition model is used to monitor and give early warnings about the quality of video images. In this embodiment, the screen freeze recognition model can be used to detect screen freeze for the video frame images extracted from the historical record video, or to detect screen freeze for the video frames extracted in real time from the real-time video stream. Specifically, video frame images are periodically extracted from the video, and it is detected whether they are screen freeze images. The proportion of the number of screen freeze images in all images is recorded as the first proportion. The first proportion is compared with the second threshold. When the second threshold is reached, the video monitoring system may determine that the quality of the collected video does not meet the actual supervision requirements due to factors such as the environment or signal transmission, and then generate a first warning prompt to give an alarm about the image quality. The second threshold can be set according to the actual application needs. In this embodiment, the second threshold can be set to 20%-30%.
[0152] By using the screen freeze recognition model to detect screen freeze for video frames, it is possible to monitor the quality of video images. For the screen freeze detection of the real-time video stream, it is possible to give an alarm in a timely manner when there are more screen freezes; for the screen freeze detection of the historical video, it is possible to conduct an overall analysis and evaluation of the quality of the video images, which is convenient for the operation and maintenance personnel to carry out corresponding maintenance and improve the quality of the video stream collected by the video monitoring system.
[0153] As an implementation manner of the present application, calculating the proportion of the number of screen freeze images in the number of images to be recognized as the first proportion includes:
[0154] Determine the number of images to be recognized in the currently nearest target time period;
[0155] Determine the number of screen freeze images in the target time period;
[0156] Calculate the proportion of the number of screen freeze images in the number of images to be recognized in the target time period as the first proportion.
[0157] In one embodiment, when monitoring the quality of video images of a real-time video stream, according to the preset time period length, screen freeze detection is performed on the video frames extracted in the currently nearest target time period, and it is judged whether the proportion (i.e., the first proportion) of the number of screen freeze images appearing in the target time period in the total number of all video frame images in the target time period reaches the second threshold. When the second threshold is reached, a first warning prompt is generated.
[0158] By judging the quality of video images in the recent target time period, it is possible to give an alarm in time when a large number of frozen frames suddenly occur in the real-time video stream in a short time, which is convenient for the operation and maintenance personnel to maintain in time.
[0159] As an implementation manner of the present application, the method for processing frozen-frame images with a long-tail distribution characteristic includes:
[0160] Calculating the proportion of the number of frozen-frame images of the target type in the number of images to be recognized as the second proportion; the target type includes: large-area solid-color frozen frames, large-area striped frozen frames, color distortion frozen frames, and noisy frozen frames;
[0161] Comparing the second proportion with a second threshold, and generating a second warning prompt when the second proportion is not less than the second threshold.
[0162] In one embodiment, considering that real-time video streams often generate frozen-frame images due to environmental or network signal factors, and in most cases, the generation of small-area frozen-frame images does not affect the identification of target objects and behaviors in the picture. Therefore, when counting the proportion of all video frame images of frozen-frame images, small-area frozen-frame images can be ignored.
[0163] To avoid frequent alarms caused by small-area frozen-frame images and affect the normal maintenance work of the video monitoring system, in this embodiment, only the proportion of frozen-frame images of the target type in all video frame images is counted, that is, the total number of large-area solid-color frozen frames, large-area striped frozen frames, color distortion frozen frames, and noisy frozen-frame images except small-area solid-color frozen-frame images and small-area striped frozen-frame images is counted as the proportion of the total number of video frames as the second proportion. When the second proportion reaches the second threshold, it is determined that the quality of the collected video does not meet the actual supervision requirements, and then a second warning message is generated, which is convenient for the operation and maintenance personnel to carry out corresponding maintenance.
[0164] As an implementation manner of the present application, calculating the proportion of the number of frozen-frame images of the target type in the number of images to be recognized as the second proportion includes:
[0165] Determining the number of images to be recognized in the current recent target time period;
[0166] Determining the number of frozen-frame images of the target type in the target time period;
[0167] Calculating the proportion of the number of frozen-frame images of the target type in the number of images to be recognized in the target time period as the second proportion.
[0168] In one embodiment, when monitoring the video picture quality of a real-time video stream, according to a preset time period length, perform a screen freeze detection on the video frames extracted within the target time period closest to the current time, and determine whether the proportion (i.e., the second proportion) of the number of screen freeze images of the target type in the target time period to the total number of all video frame images in the target time period reaches a second threshold. In the case where the second threshold is reached, generate a second warning prompt.
[0169] By judging the video picture quality within the nearest target time period, it is possible to give an alarm in a timely manner in the case of a large number of sudden screen freezes in a short time in the real-time video stream, which is convenient for the operation and maintenance personnel to maintain in a timely manner.
[0170] Based on the same inventive concept, an embodiment of the present application provides a screen freeze image processing device with a long-tail distribution characteristic. Refer to Figure 4 , Figure 4 FIG. is a schematic diagram of a screen freeze image processing device 100 with a long-tail distribution characteristic proposed in an embodiment of the present application. As Figure 4 shown, the device includes:
[0171] A first sample construction module 101, configured to construct an initial image set, where the initial image set includes the following types of screen freeze images: large-area solid-color screen freeze images, small-area solid-color screen freeze images, large-area stripe screen freeze images, small-area stripe screen freeze images, color distortion screen freeze images, and noise screen freeze images; expand the number of color distortion screen freeze images and noise screen freeze images in the initial image set to generate a first sample set;
[0172] A first training module 102, configured to randomly extract multiple different types of screen freeze images from the first sample set, extract intermediate features of each screen freeze image; remix and enhance the intermediate features of different types of screen freeze images to generate mixed features; construct a target loss function based on the intermediate features and the mixed features; and iteratively train the screen freeze recognition model based on the target loss function and the first sample set to obtain a trained screen freeze recognition model.
[0173] A detection module 103, configured to input the image to be recognized into the screen freeze recognition model, and determine whether the image to be recognized is a screen freeze image and the type of the screen freeze image through the screen freeze recognition model.
[0174] As an implementation manner of the present application, the first sample construction module 101, configured to construct an initial image set, specifically includes:
[0175] Extract frames from the sample video at a first time interval to obtain an initial image set; the sample video is a video with a screen freeze picture.
[0176] Remove non - screen - corrupted images from the initial image set;
[0177] Add labels corresponding to the screen - corrupted types to all images.
[0178] As an implementation manner of the present application, the first sample construction module 101 is configured to expand the number of color - distorted screen - corrupted images and noise screen - corrupted images in the initial image set to generate a first sample set, specifically including:
[0179] Collect images with color - distorted screen corruption, crop the screen - corrupted pixel regions in the images to obtain multiple small - size regional images; randomly mix the multiple small - size regional images with non - screen - corrupted images to generate new color - distorted screen - corrupted images;
[0180] Add noise to non - screen - corrupted images to generate new noise screen - corrupted images;
[0181] Add the new color - distorted screen - corrupted images and the new noise screen - corrupted images to the initial image set to generate a first sample set.
[0182] As an implementation manner of the present application, before adding labels corresponding to the screen - corrupted types to all images, the first sample construction module 101 further includes:
[0183] In the case where the coverage area of screen - corrupted pixels in the screen - corrupted image of the solid - color screen - corrupted type is not less than a first threshold, determine that the screen - corrupted image is a large - area solid - color screen corruption;
[0184] In the case where the coverage area of screen - corrupted pixels in the screen - corrupted image of the solid - color screen - corrupted type is less than the first threshold, determine that the screen - corrupted image is a small - area solid - color screen corruption;
[0185] In the case where the coverage area of screen - corrupted pixels in the screen - corrupted image of the stripe screen - corrupted type is not less than the first threshold, determine that the screen - corrupted image is a large - area stripe screen corruption;
[0186] In the case where the coverage area of screen - corrupted pixels in the screen - corrupted image of the stripe screen - corrupted type is less than the first threshold, determine that the screen - corrupted image is a small - area stripe screen corruption.
[0187] As an implementation manner of the present application, the screen - corruption recognition model includes: a first feature extraction network, a second feature extraction network, and a classification module; the first feature extraction network is used to extract intermediate features of each screen - corrupted image;
[0188] The first training module 102 is configured to construct a target loss function based on the intermediate features and the mixed features, specifically including:
[0189] Through the second feature extraction network and the classification module, each intermediate feature is processed to obtain the first probability distribution of the screen freeze images corresponding to each intermediate feature belonging to each type of screen freeze image;
[0190] The intermediate features of images of different screen freeze types are re-mixed and enhanced to obtain mixed features; through the second feature extraction network and the classification module, each mixed feature is processed to obtain the second probability distribution of the screen freeze images corresponding to each mixed feature belonging to each type of screen freeze image;
[0191] Based on the first probability distribution and the second probability distribution, a target loss function is constructed.
[0192] As an implementation manner of the present application, the first training module 102 is further configured to perform the following steps:
[0193] For each mixed feature, obtain the label of the screen freeze image corresponding to the intermediate feature that generates the mixed feature;
[0194] According to the label of the screen freeze image corresponding to the intermediate feature, determine the label of the screen freeze image corresponding to the mixed feature.
[0195] As an implementation manner of the present application, the screen freeze image processing device 100 with long-tail distribution characteristics further includes:
[0196] A second sample construction module, configured to add the to-be-recognized image to the second sample set when the to-be-recognized image is a non-screen freeze image; add the to-be-recognized image to the second sample set when the to-be-recognized image is a screen freeze image and the type is a noise screen freeze; when the to-be-recognized image is a screen freeze image and the type is a small-area solid color screen freeze, a small-area stripe screen freeze or a color distortion screen freeze, further determine whether the target object in the to-be-recognized image is completely covered by screen freeze pixels; if the target object is not completely covered by screen freeze pixels, add the to-be-recognized image to the second sample set;
[0197] A second training module, configured to train a visual analysis model based on the second sample set; the visual analysis model is used to perform a visual analysis task on the target object in the to-be-recognized image.
[0198] As an implementation manner of the present application, the screen freeze image processing device 100 with long-tail distribution characteristics further includes a warning module, configured to perform the following steps:
[0199] Calculate the ratio of the number of screen freeze images to the number of to-be-recognized images as the first ratio;
[0200] Compare the first ratio with the second threshold, and generate a first warning prompt when the first ratio is not less than the second threshold.
[0201] As an implementation manner of the present application, the warning module is configured to calculate the ratio of the number of the screen freeze images to the number of the images to be recognized as the first ratio, and specifically includes:
[0202] Determine the number of the images to be recognized within the current nearest target time period;
[0203] Determine the number of the screen freeze images within the target time period;
[0204] Calculate the ratio of the number of the screen freeze images to the number of the images to be recognized within the target time period as the first ratio.
[0205] As an implementation manner of the present application, the warning module is configured to perform the following steps:
[0206] Calculate the ratio of the number of the screen freeze images of the target type to the number of the images to be recognized as the second ratio; the target type includes: large-area solid color screen freeze, large-area stripe screen freeze, color distortion screen freeze, and noise screen freeze;
[0207] Compare the second ratio with the second threshold, and generate a second warning prompt when the second ratio is not less than the second threshold.
[0208] As an implementation manner of the present application, the warning module is configured to calculate the ratio of the number of the screen freeze images of the target type to the number of the images to be recognized as the second ratio, and specifically includes:
[0209] Determine the number of the images to be recognized within the current nearest target time period;
[0210] Determine the number of the screen freeze images of the target type within the target time period;
[0211] Calculate the ratio of the number of the screen freeze images of the target type to the number of the images to be recognized within the target time period as the second ratio.
[0212] Based on the same inventive concept, an embodiment of the present application provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, it implements the steps in the method for processing screen freeze images with long-tail distribution characteristics as described in any one of the above embodiments of the present application.
[0213] Based on the same inventive concept, an embodiment of the present application provides a readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the method for processing a flower screen image with a long-tail distribution characteristic as described in any of the above embodiments of the present application are implemented.
[0214] Based on the same inventive concept, an embodiment of the present application provides an electronic device. Refer to Figure 5 , Figure 5 is a schematic diagram of an electronic device 200 proposed in an embodiment of the present application. The electronic device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, the steps in the method for processing a flower screen image with a long-tail distribution characteristic as described in any of the above embodiments of the present application are implemented.
[0215] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment related to the method, and will not be elaborated here.
[0216] The above are only the preferred embodiments of the present application, and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
[0217] For the method embodiments, for simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and components involved are not necessarily essential to the present application.
[0218] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a device, or a computer program product. Therefore, the embodiments of the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0219] Embodiments of the present application are described with reference to the flowcharts and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing terminal device generate a device for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or a device for implementing the functions specified in multiple blocks.
[0220] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or a device for implementing the functions specified in multiple blocks.
[0221] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, such that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 or a device for implementing the functions specified in multiple blocks.
[0222] Although the preferred embodiments of the embodiments of the present application have been described, those skilled in the art can make additional changes and modifications to these embodiments once they learn the basic creative concept. Therefore, the present application is interpreted to include the preferred embodiments and all changes and modifications that fall within the scope of the embodiments of the present application.
[0223] Finally, it should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or terminal device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising the element.
[0224] The above has introduced in detail the method, apparatus and product for processing a flower screen image with a long-tail distribution characteristic provided by the present application. Specific examples are used in this text to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A method for processing a distorted screen image with a long-tail distribution characteristic, characterized in that: include: Constructing an initial image set, wherein the initial image set includes the following types of flower screen images: large-area pure color flower screen images, small-area pure color flower screen images, large-area striped flower screen images, small-area striped flower screen images, color-distorted flower screen images, and noise flower screen images; Expanding the number of color-distorted distorted screen images and noise-distorted screen images in the initial image set to generate a first sample set; Randomly extracting a plurality of different types of flower screen images from the first sample set, and extracting intermediate features of each flower screen image; The intermediate features of different types of flower screen images are remixed and enhanced to generate mixed features; Based on the intermediate features and the mixed features, construct a target loss function; Iteratively training the flower screen recognition model based on the target loss function and the first sample set to obtain a trained flower screen recognition model; Inputting the image to be identified into the flower screen recognition model, and judging whether the image to be identified is a flower screen image and the type of the flower screen image by the flower screen recognition model; The flower screen recognition model includes: a first feature extraction network, a second feature extraction network and a classification module; the first feature extraction network is used to extract the intermediate features of each flower screen image; Based on the intermediate features and the mixed features, a target loss function is constructed, including: Processing each intermediate feature through the second feature extraction network and the classification module to obtain a first probability distribution of the flower screen image corresponding to each intermediate feature belonging to each type of flower screen image; Processing each mixed feature through the second feature extraction network and the classification module to obtain a second probability distribution of the flower screen image corresponding to each mixed feature belonging to each type of flower screen image; Constructing a target loss function based on the first probability distribution and the second probability distribution; The expressions for calculating the first probability distribution and the second probability distribution are as follows: Among them, p represents the probability distribution of the type of the flower screen image calculated based on the intermediate features / mixed features. is the first probability distribution, is the second probability distribution, Indicates the predicted label of the image sample flower screen type, Z i , Z j Respectively represent the values of the i-th and j-th items in the feature vector of the original input training image, mixZ i 、mixZ j They represent the values of the i-th and j-th items in the feature vector of the flower screen category after remixing and enhancement, C represents the total number of flower screen types, and n i 、n j Indicates the number of training image samples of the i-th and j-th types of flower screen, n max Indicates the number of flower screen image samples of the flower screen type with the largest number of image samples; The calculation expression of the objective loss function is as follows: Among them, y i Indicates the label of the flower screen type of the original input flower screen image, and mixY indicates the label of the flower screen type of the flower screen image after remixing and enhancement.
2. The method for processing a distorted screen image with a long-tail distribution characteristic according to claim 1, characterized in that: Construct an initial set of images, including: Extract frames from a sample video at a first time interval to obtain an initial image set; the sample video is a video with a distorted screen; Remove non-flowered screen images from the initial image set; Add labels corresponding to the type of screen distortion to all images.
3. The method for processing a distorted screen image with a long-tail distribution characteristic according to claim 2, characterized in that: The number of color distorted distorted screen images and noise distorted screen images in the initial image set is expanded to generate a first sample set, including: An image with color distortion and flower screen is collected, and the flower screen pixel area in the image is cropped to obtain a plurality of small-sized regional images; the plurality of small-sized regional images are randomly mixed with the non-flower screen image to generate a new color distortion and flower screen image; Add noise to the non-flowering image to generate a new noisy flowering image; The new color distorted distorted screen image and the new noise distorted screen image are added to the initial image set to generate a first sample set.
4. The method for processing a distorted screen image with a long-tail distribution characteristic according to claim 2, characterized in that: Before adding labels of corresponding screen distortion types to all images, it also includes: In a flower screen image of a pure color flower screen type, when the flower screen pixel coverage area is not less than a first threshold, the flower screen image is determined to be a large-area pure color flower screen; In a flower screen image of a pure color flower screen type, when the flower screen pixel coverage area is less than a first threshold, the flower screen image is determined to be a small-area pure color flower screen; In a striped flower screen type flower screen image, when the flower screen pixel coverage area is not less than a first threshold, determining that the flower screen image is a large-area striped flower screen; In a striped screen type distorted image, when the distorted screen pixel coverage area is smaller than a first threshold, the distorted screen image is determined to be a small-area striped screen distorted image.
5. The method for processing a distorted screen image with a long-tail distribution characteristic according to claim 1, characterized in that: Also includes: For each mixed feature, obtain the label of the flower screen image corresponding to the intermediate feature that generates the mixed feature; According to the label of the flower screen image corresponding to the intermediate feature, the label of the flower screen image corresponding to the mixed feature is determined.
6. The method for processing a distorted screen image with a long-tail distribution characteristic according to claim 1, characterized in that: Also includes: In the case where the image to be identified is a non-flowered screen image, adding the image to be identified to the second sample set; In a case where the image to be identified is a distorted screen image and the type is a noise distorted screen image, adding the image to be identified to the second sample set; In the case where the image to be identified is a flower screen image, and the type is a small area pure color flower screen, a small area stripe flower screen or a color distortion flower screen, further determining whether the target object in the image to be identified is completely covered by the flower screen pixels; if the target object is not completely covered by the flower screen pixels, adding the image to be identified to the second sample set; Based on the second sample set, training a visual analysis model; The visual analysis model is used to perform a visual analysis task on a target object in an image to be identified.
7. The method for processing a distorted screen image with a long-tail distribution characteristic according to claim 1, characterized in that: Also includes: Calculate the ratio of the number of the distorted screen images to the number of the images to be identified as a first ratio; The first proportion is compared with a second threshold, and a first warning prompt is generated when the first proportion is not less than the second threshold.
8. The method for processing a distorted screen image with a long-tail distribution characteristic according to claim 7, characterized in that: Calculating the ratio of the number of the distorted screen images to the number of the images to be identified as a first ratio includes: Determine the number of images to be identified within the current most recent target time period; Determine the number of distorted screen images within the target time period; The ratio of the number of the distorted screen images to the number of the to-be-recognized images within the target time period is calculated as a first ratio.
9. The method for processing a distorted screen image with a long-tail distribution characteristic according to claim 1, characterized in that: Also includes: Calculate the ratio of the number of the target type of distorted screen images to the number of the to-be-recognized images as a second ratio; The target types include: large-area pure color flower screen, large-area striped flower screen, color distortion flower screen and noise flower screen; The second proportion is compared with a second threshold, and a second warning prompt is generated when the second proportion is not less than the second threshold.
10. The method for processing a distorted screen image with a long-tail distribution characteristic according to claim 9, characterized in that: Calculating the ratio of the number of the target type of distorted screen images to the number of the to-be-recognized images as a second ratio includes: Determine the number of images to be identified within the current most recent target time period; Determine the number of the target type of distorted screen images within the target time period; The ratio of the number of the target type of distorted screen images to the number of the to-be-recognized images within the target time period is calculated as a second ratio.
11. A device for processing a distorted screen image with a long-tail distribution characteristic, characterized in that: Used to perform the method according to any one of claims 1 to 10, comprising: The first sample construction module is configured to construct an initial image set, wherein the initial image set includes the following types of flower screen images: large-area pure color flower screen images, small-area pure color flower screen images, large-area striped flower screen images, small-area striped flower screen images, color-distorted flower screen images, and noise flower screen images; the number of color-distorted flower screen images and noise flower screen images in the initial image set is expanded to generate a first sample set; The first training module is configured to randomly extract a plurality of different types of flower screen images from the first sample set, extract intermediate features of each flower screen image; remix and enhance the intermediate features of different types of flower screen images to generate mixed features; construct a target loss function based on the intermediate features and the mixed features; iteratively train the flower screen recognition model based on the target loss function and the first sample set to obtain a trained flower screen recognition model; The detection module is configured to input the image to be identified into the flower screen recognition model, and determine whether the image to be identified is a flower screen image and the type of the flower screen image through the flower screen recognition model; The flower screen recognition model includes: a first feature extraction network, a second feature extraction network and a classification module; the first feature extraction network is used to extract the intermediate features of each flower screen image; The first training module is configured to construct a target loss function based on the intermediate features and the mixed features, specifically including: Processing each intermediate feature through the second feature extraction network and the classification module to obtain a first probability distribution of the flower screen image corresponding to each intermediate feature belonging to each type of flower screen image; Processing each mixed feature through the second feature extraction network and the classification module to obtain a second probability distribution of the flower screen image corresponding to each mixed feature belonging to each type of flower screen image; Constructing a target loss function based on the first probability distribution and the second probability distribution; The expressions for calculating the first probability distribution and the second probability distribution are as follows: Among them, p represents the probability distribution of the type of the flower screen image calculated based on the intermediate features / mixed features. is the first probability distribution, is the second probability distribution, Indicates the predicted label of the image sample flower screen type, Z i , Z j Respectively represent the values of the i-th and j-th items in the feature vector of the original input training image, mixZ i 、mixZ j They represent the values of the i-th and j-th items in the feature vector of the flower screen category after remixing and enhancement, C represents the total number of flower screen types, and n i 、n j Indicates the number of training image samples of the i-th and j-th types of flower screen, n max Indicates the number of flower screen image samples of the flower screen type with the largest number of image samples; The calculation expression of the objective loss function is as follows: Among them, y i Indicates the label of the flower screen type of the original input flower screen image, and mixY indicates the label of the flower screen type of the flower screen image after remixing and enhancement.
12. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.
13. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 10 are implemented.
14. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 10 are implemented.
Citation Information
Patent Citations
Long-tail image recognition method based on representation data enhancement and loss rebalance
CN116030302A
Method and system for adaptation of a trained object detection model to account for domain shift
US20230281974A1