Methods, apparatus, equipment, media, and program products for generating sample images

By generating enhanced sample images in the image processing model, and utilizing probability distribution conditions and image registration techniques, the method solves the problems of low accuracy and limitations of image processing models in small sample cases, and improves image diversity and model robustness.

CN115239590BActive Publication Date: 2025-10-31TENCENT TECHNOLOGY (SHENZHEN) CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210893870.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-27
Publication Date
2025-10-31
Estimated Expiration
2042-07-27

AI Technical Summary

Technical Problem

In existing technologies, image processing models struggle to learn rich image knowledge when the number of training samples is small, resulting in low accuracy in image analysis and significant model limitations.

Method used

By acquiring a specified sample image and candidate sample images, the specified sample image is divided into regions based on probability distribution conditions to determine the target sub-image region. Candidate sub-image regions are then matched from the candidate sample images and applied to the specified sample image to generate an enhanced sample image.

Benefits of technology

The generated enhanced sample images largely avoid destroying the integrity of the main image, improve the diversity of images and the robustness of the model, and can uncover more image distribution patterns in small sample learning, overcoming the limitations of small sample datasets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115239590B_ABST
    Figure CN115239590B_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, device, medium, and program product for generating sample images, relating to the field of machine learning. The method includes: dividing a specified sample image into multiple sub-image regions; determining a target sub-image region that meets probability requirements from the multiple sub-image regions based on probability distribution conditions determined by the distribution pattern of the image subject in the image; determining candidate sub-image regions that match the target sub-image region from the candidate sample images based on the registration relationship between the specified sample image and candidate sample images; and applying the candidate sub-image regions to the location of the target sub-image region in the specified sample image to obtain an enhanced sample image. Through this method, the integrity of the image subject can be largely avoided, resulting in a large batch of enhanced sample images similar to the specified sample image, thus improving the diversity of enhanced sample images. This application can be applied to various scenarios such as cloud technology, artificial intelligence, and intelligent transportation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of machine learning, and in particular to a method, apparatus, device, medium, and program product for generating sample images. Background Technology

[0002] With the development of network technology, the phenomenon of information overload has become increasingly obvious, and traditional information recommendation methods are finding it difficult to make personalized recommendations for users from massive amounts of information.

[0003] In related technologies, sample images obtained from the network or dataset are typically used to train a pre-trained model with certain image processing capabilities, thereby enabling the trained image processing model to perform image analysis on images similar to the sample images.

[0004] In the above process, although the trained image processing model can perform relatively effective image processing, when the number of training samples is small, the image processing model has difficulty learning rich image knowledge from the limited training samples, which makes the image processing model more limited and the accuracy of image analysis lower. Summary of the Invention

[0005] This application provides a method, apparatus, device, medium, and program product for generating sample images, which can largely avoid damaging the integrity of the image subject and obtain a large number of enhanced sample images similar to the specified sample image, thereby improving the diversity of enhanced sample images. The technical solution is as follows.

[0006] On the one hand, a method for generating sample images is provided, the method comprising:

[0007] Acquire a specified sample image and a candidate sample image, wherein the specified sample image is the image to be enhanced using the candidate sample image;

[0008] The specified sample image is divided into regions to obtain multiple sub-image regions in the specified sample image;

[0009] Based on probability distribution conditions, at least one target sub-image region that meets the probability requirements is determined from the plurality of sub-image regions as the sub-image region to be enhanced, wherein the probability distribution conditions are determined based on the distribution pattern of the image subject in the image;

[0010] Based on the registration relationship between the specified sample image and the candidate sample image, at least one candidate sub-image region that matches the at least one target sub-image region is determined from the candidate sample image;

[0011] The at least one candidate sub-image region is applied to the region location of the at least one target sub-image region in the specified sample image to obtain an enhanced sample image, wherein the enhanced sample image is a sample image generated after adjusting the specified sample image.

[0012] On the other hand, a sample image generation apparatus is provided, the apparatus comprising:

[0013] The acquisition module is used to acquire a specified sample image and a candidate sample image, wherein the specified sample image is the image to be sampled and enhanced using the candidate sample image;

[0014] The segmentation module is used to segment the specified sample image into regions to obtain multiple sub-image regions in the specified sample image.

[0015] The determination module is used to determine at least one target sub-image region that meets the probability requirements from the plurality of sub-image regions based on probability distribution conditions, as the sub-image region to be enhanced, wherein the probability distribution conditions are determined based on the distribution pattern of the image subject in the image;

[0016] A registration module is used to determine at least one candidate sub-image region that matches the at least one target sub-image region from the candidate sample image based on the registration relationship between the specified sample image and the candidate sample image;

[0017] An application module is used to apply the at least one candidate sub-image region to the region location of the at least one target sub-image region in the specified sample image to obtain an enhanced sample image, wherein the enhanced sample image is a sample image generated after adjusting the specified sample image.

[0018] On the other hand, a computer device is provided, the computer device including a processor and a memory, the memory storing at least one instruction, at least one program, code set or instruction set, the at least one instruction, the at least one program, the code set or instruction set being loaded and executed by the processor to implement the sample image generation method as described in any of the above embodiments of this application.

[0019] On the other hand, a computer-readable storage medium is provided, wherein at least one instruction, at least one program, code set, or instruction set is stored therein, wherein the at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by a processor to implement the sample image generation method as described in any of the embodiments of this application above.

[0020] On the other hand, a computer program product or computer program is provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the sample image generation method described in any of the above embodiments.

[0021] The beneficial effects of the technical solutions provided in this application include at least the following:

[0022] Based on probability distribution conditions, a target sub-image region is determined from multiple sub-image regions corresponding to a specified sample image, and candidate sub-image regions with registration relationships with the target sub-image region are identified from candidate sample images. These candidate sub-image regions are then applied to the corresponding regions of the target sub-image region to obtain the enhanced sample image. Since the probability distribution conditions are determined based on the distribution patterns of the image subject within the image, the target sub-image region to be enhanced can largely avoid damaging the integrity of the image subject. This not only better protects the image information of the image subject but also allows the candidate sub-image regions to be used to expand other image regions in the specified sample image besides the image subject, resulting in a large number of enhanced sample images similar to the specified sample image, thus making the enhanced sample images more diverse. In few-shot learning, the model can be trained using the specified sample image and similar enhanced sample images, thereby uncovering more image distribution patterns, overcoming the limitations of small sample datasets, and improving the robustness of the model. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 This is a schematic diagram of an implementation environment provided by an exemplary embodiment of this application;

[0025] Figure 2 This is a flowchart of a method for generating a sample image provided in an exemplary embodiment of this application;

[0026] Figure 3 This is a schematic diagram of a candidate sample image provided in an exemplary embodiment of this application;

[0027] Figure 4 This is a schematic diagram of a specified sample image provided in an exemplary embodiment of this application;

[0028] Figure 5 This is a flowchart of a method for generating a sample image provided in another exemplary embodiment of this application;

[0029] Figure 6 This is a schematic diagram of a binary Gaussian distribution probability function provided in an exemplary embodiment of this application;

[0030] Figure 7 This is a flowchart of a method for generating a sample image provided in another exemplary embodiment of this application;

[0031] Figure 8 This is a schematic diagram of an enhanced sample image provided in an exemplary embodiment of this application;

[0032] Figure 9 This is a schematic diagram of a long-tail distribution provided in an exemplary embodiment of this application;

[0033] Figure 10 This is a schematic diagram of image segmentation provided in an exemplary embodiment of this application;

[0034] Figure 11 This is an application diagram illustrating a sample image generation method provided in an exemplary embodiment of this application;

[0035] Figure 12 This is a structural block diagram of a sample image generation apparatus provided in an exemplary embodiment of this application;

[0036] Figure 13 This is a structural block diagram of a server provided in an exemplary embodiment of this application. Detailed Implementation

[0037] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0038] In related technologies, data on user preferences and needs is typically collected, and this collected training data is used to train a manually built personalized recommendation model. The trained model then recommends information to users. For example, based on a user's historical preference data, information matching their preferences is recommended. While the trained recommendation model can provide relatively effective recommendations, it is still inevitably subject to human bias. Furthermore, the model is highly correlated with the collected training data; when this model is used to analyze other relevant data, the predictive accuracy of the recommendations is significantly reduced.

[0039] This application provides a method for generating sample images that can largely avoid damaging the integrity of the main subject of the image, obtaining a large number of enhanced sample images similar to the specified sample image, and improving the diversity of the enhanced sample images. The method for generating sample images trained according to this application can be applied to at least one of image classification scenarios, object detection scenarios, and image segmentation scenarios.

[0040] It should be noted that all information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this application have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the image data such as designated sample images and candidate sample images involved in this application were obtained with full authorization.

[0041] It is worth noting that the above application scenarios are merely illustrative examples. The sample image generation method provided in this embodiment can also be applied to other scenarios, and this application embodiment does not limit it.

[0042] Secondly, the implementation environment involved in the embodiments of this application will be described, for illustrative purposes only. Please refer to [the relevant documentation]. Figure 1 The implementation environment involves a terminal 110 and a server 120, which are connected via a communication network 130.

[0043] In some embodiments, the terminal 110 is equipped with an application that has image acquisition capabilities. In some embodiments, the terminal 110 is used to send a sample image to the server 120. After receiving the sample image, the server 120 can recognize the sample image according to the image recognition model 121 to obtain an image recognition result. Optionally, the server 120 sends the image recognition result to the terminal 110 so that the terminal 110 can display the image recognition result after recognizing the sample image.

[0044] The image recognition model 121 is trained using the following method: The server 120, based on sample images stored locally or sent by the terminal 110 (at least one of a specified sample image and candidate sample images), performs sample enhancement on the specified sample image using the candidate sample image; after dividing the specified sample image into regions, multiple sub-image regions (multiple small squares) are obtained; based on the probability distribution of the image subject in the image, at least one target sub-image region (diagonal striped square) that meets the probability requirements is determined from the multiple sub-image regions as the sub-image region to be enhanced; based on the registration relationship between the specified sample image and the candidate sample image, at least one candidate sub-image region (vertical striped square) that matches the at least one target sub-image region is determined from the candidate sample image; the at least one candidate sub-image region is applied to the region location of the at least one target sub-image region in the specified sample image to obtain the enhanced sample image generated after adjusting the specified sample image. The above process, which trains the image recognition model 121 using enhanced sample images, is an example of a non-unique case of the training process for the image recognition model 121.

[0045] It is worth noting that the aforementioned terminals include, but are not limited to, mobile terminals such as mobile phones, tablets, portable laptops, smart voice interaction devices, smart home appliances, and in-vehicle terminals, as well as desktop computers; the aforementioned servers can be independent physical servers, server clusters or distributed systems composed of multiple physical servers, or cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.

[0046] Cloud technology refers to a hosting technology that unifies hardware, applications, networks, and other resources within a wide area network (WAN) or local area network (LAN) to achieve data computation, storage, processing, and sharing. Based on the cloud computing business model, cloud technology encompasses network technology, information technology, integration technology, management platform technology, and application technology. It can form resource pools, providing flexible and convenient on-demand access.

[0047] In some embodiments, the server described above can also be implemented as a node in a blockchain system.

[0048] Based on the above-described terminology and application scenarios, the method for generating sample images provided in this application will be explained, taking the application of this method in a server as an example. Figure 2 As shown, the method includes the following steps 210 to 250.

[0049] Step 210: Obtain the specified sample image and candidate sample images.

[0050] The specified sample image is the image to be augmented using the candidate sample image.

[0051] Indicatively, the specified sample image is used to indicate the image to be enhanced, while the candidate sample image is used to assist the specified sample image in image enhancement.

[0052] In an optional embodiment, a sample image set is obtained.

[0053] The sample image set stores multiple sample images. Optionally, the sample image set is used to indicate a collection of various images, such as: a collection of all images on the Internet as a sample image library; or, a collection of various landscape images, animal images, etc.

[0054] Optionally, candidate sample images and specified sample images can be obtained from the sample image set.

[0055] For example, after obtaining a set of sample images, one sample image is randomly selected from it as the designated sample image to be enhanced; or, a sample image is obtained from outside the set of sample images as the designated sample image to be enhanced, for example, a photograph is taken and used as the designated sample image to be enhanced.

[0056] In illustrative terms, after obtaining the sample image set, at least one sample image is arbitrarily selected from it as a candidate sample image for sample augmentation of the specified sample image.

[0057] For example: In addition to determining the specified sample image, at least one sample image with the same image size as the specified sample image is arbitrarily selected from the sample image set as a candidate sample image; or, at least one sample image is arbitrarily selected from the sample image set as a candidate sample image, so that after determining the specified sample image, the specified sample image is sample augmented by the selected at least one candidate sample image.

[0058] Optionally, sample enhancement is used to indicate the enhancement of the number of sample images. Illustratively, a specified sample image has relatively simple image information. The image information of the specified sample image is enhanced using candidate sample images. This allows the image representation of the specified sample image to be expanded using image information from some candidate sample images while retaining the original image information, thus achieving the process of sample enhancement for candidate sample images.

[0059] Step 220: Divide the specified sample image into regions to obtain multiple sub-image regions in the specified sample image.

[0060] In an optional embodiment, after obtaining the specified sample image, the specified sample image is divided into regions using an equal area division method to obtain multiple sub-image regions in the specified sample image.

[0061] The equal area division method is used to indicate that the areas of multiple sub-map regions obtained are the same.

[0062] Optionally, an equal-area grid division method can be used to divide the specified sample image into regions.

[0063] For example, the length of the grid is predetermined to be l and the width to be w. The specified sample image is divided into regions using l×w small squares as the dividing standard, thereby obtaining multiple sub-image regions. Each sub-image region has the same area, which is l×w.

[0064] Optionally, after obtaining the specified sample image, the specified sample image is divided into a certain number of blocks based on its length and width.

[0065] To illustrate, after obtaining a specified sample image, the specified sample image is divided into regions in the length and width directions. For example, the specified sample image is divided into N segments in the length direction and M segments in the width direction, thereby dividing the specified sample image into N*M small squares horizontally and vertically. Here, N is a positive integer, M is a positive integer, and N and M can be the same or different.

[0066] For example, taking a rectangular (e.g., square, rectangle, etc.) image as a specified sample image, when N and M are the same, the length of the specified sample image is divided into N segments at unit intervals (e.g., 1mm, 0.2cm, etc.), and the width of the specified sample image is divided into N segments at the same unit intervals, thus dividing the specified sample image into N segments. 2 Each small square is treated as a sub-image region, thus obtaining N corresponding to a specified sample image. 2 Individual sub-graph regions.

[0067] Alternatively, taking a rectangular image as an example, when N and M are different, the "length" of the specified sample image is divided into N segments at equal intervals; the "width" of the specified sample image is divided into M segments at equal intervals, thus dividing the specified sample image into N×M small squares. Each small square is used as a sub-image region, that is, N×M sub-image regions corresponding to the specified sample image are obtained.

[0068] Optionally, a sliding window method can be used to divide the specified sample image into regions.

[0069] The sliding window is used to indicate the movement of a window, and the area corresponding to the window is used to determine the sub-image region corresponding to the specified sample image.

[0070] To illustrate, when a window is created, its size is preset. During the window's sliding process, as the right boundary of the window slides a certain distance to the right, the left boundary also slides a certain distance to the right. Based on the image region of a specified sample image, the sliding window is moved within the image region to obtain multiple sub-image regions of equal area.

[0071] For example, when the specified sample image is an irregular image, it is determined whether image information corresponding to the specified sample image exists in each region traversed by the sliding window. If image information corresponding to the sample image exists in region A traversed by the sliding window, region A is treated as a sub-image region; if image information corresponding to the sample image does not exist in region B traversed by the sliding window, region B is not treated as a sub-image region, and so on. This is illustrated by using the above method to determine the regions traversed by the sliding window.

[0072] In an optional embodiment, after obtaining the specified sample image, the specified sample image is divided into regions using a non-equal area division method to obtain multiple sub-image regions in the specified sample image.

[0073] Optionally, after obtaining the specified sample image, the specified sample image is divided into regions using a random division method.

[0074] Indicatively, regions c and d are arbitrarily selected from a specified sample image as sub-image regions within the specified sample image. Optionally, regions c and d are two non-overlapping regions in the specified sample image; or, regions c and d are two regions in the specified sample image that partially overlap, etc.

[0075] Step 230: Based on the probability distribution conditions, determine at least one target sub-graph region that meets the probability requirements from multiple sub-graph regions, as the sub-graph region to be enhanced.

[0076] In a schematic manner, after obtaining multiple sub-image regions in a specified sample image, the multiple sub-image regions are analyzed by probability distribution conditions, and at least one target sub-image region that meets the probability requirements is determined.

[0077] Among them, the probability distribution condition is determined based on the distribution pattern of the main subject in the image.

[0078] Optionally, the image includes at least a subject, and may also include a background. The subject indicates the main image information presented in the image; the background indicates image information other than the subject.

[0079] Indicative, such as Figure 3 The image shown is of a lion. Since the image primarily depicts lion 310, lion 310 is used as the main subject, and the sky behind it serves as the background. Alternatively, as shown... Figure 4 As shown, this is an image of a kitten. Since the image is mainly used to represent kitten 410, kitten 410 is used as the subject of the image, and the grass behind kitten 410 is used as the background.

[0080] In an optional embodiment, multiple sub-graph regions are analyzed using pre-defined probability distribution conditions to determine at least one target sub-graph region.

[0081] Indicatively, the probability distribution conditions are predetermined distribution conditions. When applying the probability distribution conditions to analyze multiple subgraph regions, the probability of multiple subgraph regions being used as target subgraph regions is determined.

[0082] Optionally, the main image subject, as the primary representation of image information, is typically located in the central region of the image. Based on this distribution pattern, the image background is usually located in the surrounding region. Since the image subject carries a significant amount of important information, changes to the image subject often lead to changes in this important information. Therefore, while considering maintaining the image subject information without significant changes, the image background information is adjusted; that is, a higher probability is set for the image background as the target sub-image region to be enhanced.

[0083] Indicatively, based on the above distribution pattern, the probability distribution conditions are set as follows: sub-image regions located in the center of the image have a lower probability of being the target sub-image region; sub-image regions located at the edge of the image have a higher probability of being the target sub-image region, etc.

[0084] In an optional embodiment, the probability distribution condition is determined by multiple sample region images, which are pre-collected image data.

[0085] Optionally, a region recognition model is used to perform region recognition on multiple sample region images to determine the image subject region corresponding to each of the multiple sample region images.

[0086] The image subject region is used to indicate the image region where the main image is located in the sample region image.

[0087] Indicatively, a pre-trained image recognition model is used to perform image recognition on multiple sample regions to identify the main image region carrying more important information in each sample region image. For example, when considering... Figure 3 After the lion image shown is used as the sample region image, then... Figure 3 Image recognition was performed, and the image area where Lion 310 was located was used as... Figure 3 The corresponding main image region; or, when such as Figure 4 After the kitten image shown is used as the sample region image, then... Figure 4 Image recognition was performed, and the image region where the kitten 410 was located was used as... Figure 4 The corresponding main image area.

[0088] Optionally, the probability distribution conditions can be determined by comprehensively analyzing the regional locations of multiple image subject regions in the corresponding sample region images.

[0089] Schematic illustration: the multiple image subject regions include a first image subject region and a second image subject region. The first image subject region is used to indicate the image subject region corresponding to the first sample region image, and the second image subject region is used to indicate the image subject region corresponding to the second sample region image.

[0090] Determine the location of the first main body region of the first image in the first sample region image, and the location of the second main body region of the second image in the second sample region image; comprehensively analyze the locations of the first and second regions to determine the distribution pattern of the main body regions, and use the distribution pattern of the main body regions as a probability distribution condition.

[0091] To illustrate, after obtaining the positions of the first and second regions, the distribution pattern of the main subject regions corresponding to multiple image subject regions can be determined based on these positions. For example, if both the first and second region positions are located near the center of the corresponding sample region image, and the center position is taken as the distribution center of the subject regions, then the distribution pattern of the subject regions is: most of the subject regions are located near the center of the sample region image.

[0092] Alternatively, the first region is located at the center of the first sample region image, and the second region is located at the upper right corner of the second sample region image. Optionally, the distribution center of the main body region is determined by the center of the position points of the first and second regions. Then, the distribution pattern of the main body region is that most of the main body regions are located near the center of the position points in the upper right corner of the sample region image.

[0093] Optionally, based on the above-mentioned main region distribution pattern, the probability distribution conditions are determined, and the sub-image regions in the specified sample image are analyzed to determine the target sub-image regions that meet the probability distribution conditions.

[0094] To illustrate, consider the distribution pattern of the main subject area: it is mostly located near the center of the sample image. Optionally, when analyzing a specified sample image, first determine the main subject area corresponding to the main subject, and the background area excluding the main subject area. When selecting the target sub-image area, at least one sub-image area is selected from multiple sub-image areas corresponding to the background area as the target sub-image area, without selecting from the multiple sub-image areas corresponding to the main subject area. That is, since the main subject area is mostly located at the center of the sample image, in the specified sample image, the main subject area is mostly located at the center of the specified sample image. To avoid disrupting the main subject, the target sub-image area is determined from the background area.

[0095] Alternatively, different selection probabilities can be set when selecting target sub-image regions, performing the selection process for target sub-image regions on the image background region and the image subject region. For example, when selecting a target sub-image region, multiple sub-image regions corresponding to the image subject region have a 1 / 10 probability of being selected as the target sub-image region; multiple sub-image regions corresponding to the image background region have a 9 / 10 probability of being selected as the target sub-image region, and so on.

[0096] Optionally, when determining at least one target subgraph region from multiple subgraph regions, a selection process is performed based on probability requirements.

[0097] The probability requirement indicates the probability condition for selecting a region as the target subgraph region. Illustratively, after determining the probability of multiple subgraph regions being the target subgraph region through probability distribution conditions, the probabilities corresponding to each of the multiple subgraph regions are compared with the probability requirement, thereby determining the target subgraph region from the multiple subgraph regions.

[0098] This is illustrative, where the probability requirement is a pre-set probability threshold. When comparing the probabilities corresponding to multiple sub-regions with the probability requirement, that is, comparing the probabilities corresponding to each sub-region with the preset probability threshold. Optionally, sub-regions with probabilities greater than or equal to the preset probability threshold are identified and designated as target sub-regions.

[0099] Step 240: Based on the registration relationship between the specified sample image and the candidate sample image, determine at least one candidate sub-image region from the candidate sample image that matches at least one target sub-image region.

[0100] To illustrate, after obtaining the specified sample image and the candidate sample image, image registration is performed on the specified sample image and the candidate sample image.

[0101] The purpose of image registration is to compare or fuse the acquired images so that different images correspond one-to-one with points at the same location in space, thereby achieving the purpose of image information fusion.

[0102] Optionally, the candidate sample image is registered onto the specified sample image, using the specified sample image as the target registration image, thereby determining the registration relationship between the specified sample image and the candidate sample image.

[0103] In illustrative terms, the image center of a specified sample image is used as the registration center. The image center of a candidate sample image is then determined, and the image center of the candidate sample image is registered to the registration center, thereby realizing the image registration process.

[0104] Alternatively, a point can be arbitrarily selected from the specified sample image (e.g., an image vertex, any point in the lower right region of the image, etc.) as the registration center. Based on the relative position coordinates of this point in the specified sample image, the image coordinates corresponding to the relative position coordinates are determined from the candidate sample image. The corresponding points in the candidate sample image are then registered to the aforementioned registration center, thereby realizing the image registration process.

[0105] Optionally, during image registration, the image size of the candidate sample image is adjusted based on the image size of the specified sample image. Illustratively, the candidate sample image is scaled (reduced or enlarged) according to the length and width of the specified sample image to match the image size of the candidate sample image, thereby performing image registration between the scaled candidate sample image and the specified sample image.

[0106] In an optional embodiment, the registration relationship between the specified sample image and the candidate sample image is determined based on the image registration process described above.

[0107] Indicatively, based on the image registration process described above, the position coordinates of different points in the specified sample image are mapped to the position coordinates of different points in the candidate sample image. For example, a coordinate system is established with the registration center as the origin, the first position coordinates and the second position coordinates of different points in the specified sample image are determined, and based on the image registration process, the relative coordinate relationship between the first and second position coordinates is determined. This relative coordinate relationship is then used as the registration relationship between the specified sample image and the candidate sample image.

[0108] Optionally, after determining at least one target sub-image region, based on the position information of the target sub-image region in the specified sample image, the first position coordinates corresponding to at least one point in the target sub-image region are determined, and based on the registration relationship, the second position coordinates corresponding to the first position coordinates are determined, thereby determining the candidate sub-image region from the candidate sample image based on the second position coordinates, wherein the position coordinates of the target sub-image region and its corresponding candidate sub-image region correspond.

[0109] Step 250: Apply at least one candidate sub-image region to the region location of at least one target sub-image region in the specified sample image to obtain an enhanced sample image.

[0110] Schematic, after determining at least one candidate sub-image region from the candidate sample image, based on the registration relationship between the candidate sub-image region and the target sub-image region, when applying the candidate sub-image region to the specified sample image, the candidate sub-image region is applied to the corresponding target sub-image region in the specified sample image.

[0111] Indicatively, after selecting target sub-image regions T1 and T2 from the specified sample images, based on the above registration relationship, candidate sub-image regions C1 corresponding to target sub-image region T1 and candidate sub-image regions C2 corresponding to target sub-image region T2 in the candidate sample images are determined.

[0112] For example, if a coordinate system is established with the registration center as the origin, then the first position coordinates of the target sub-map region T1 are the same as the second position coordinates of the candidate sub-map region C1; the first position coordinates of the target sub-map region T2 are the same as the second position coordinates of the candidate sub-map region C2.

[0113] In an optional embodiment, a candidate sub-image region is applied to the corresponding target sub-image region in a specified sample image to obtain an enhanced sample image.

[0114] Indicatively, when obtaining an enhanced sample image, candidate sub-image region C1 is applied to target sub-image region T1 in a specified sample image to obtain enhanced sample image E1; candidate sub-image region C2 is applied to target sub-image region T2 in a specified sample image to obtain enhanced sample image E2, and so on.

[0115] In an optional embodiment, at least two of the multiple candidate sub-image regions are applied to the corresponding target sub-image region in the specified sample image to obtain an enhanced sample image.

[0116] Indicatively, after selecting target sub-image regions T1, T2, and T3 from the specified sample image, based on the above registration relationship, candidate sub-image regions C1 corresponding to target sub-image region T1, C2 corresponding to target sub-image region T2, and C3 corresponding to target sub-image region T3 in the candidate sample image are determined.

[0117] When obtaining the enhanced sample image, candidate sub-image region C1 is applied to the target sub-image region T1 in the specified sample image, and candidate sub-image region C2 is applied to the target sub-image region T2 in the specified sample image to obtain the enhanced sample image E3; or, candidate sub-image region C1 is applied to the target sub-image region T1 in the specified sample image, and candidate sub-image region C3 is applied to the target sub-image region T3 in the specified sample image to obtain the enhanced sample image E4; or, candidate sub-image region C1 is applied to the target sub-image region T1 in the specified sample image, candidate sub-image region C2 is applied to the target sub-image region T2 in the specified sample image, and candidate sub-image region C3 is applied to the target sub-image region T3 in the specified sample image to obtain the enhanced sample image E5, etc.

[0118] In an optional embodiment, when applying a candidate sub-image region to the corresponding target sub-image region in a specified sample image, a replacement method is used to replace the target sub-image region in the specified sample image with a candidate sub-image region corresponding to the target sub-image region, thereby obtaining an enhanced sample image; or, an overlay method is used to overlay the pixels of the target sub-image region in the specified sample image with the pixels of the candidate sub-image region corresponding to the target sub-image region, thereby obtaining an enhanced sample image.

[0119] The enhanced sample image is a sample image generated by adjusting a specified sample image. Illustratively, at least one target sub-image region in the specified sample image is adjusted, replacing at least one target sub-image region with a candidate sub-image region, thereby obtaining the adjusted enhanced sample image.

[0120] In illustrative terms, in the enhanced sample image, a small portion of the image information in the specified sample image is adjusted so that the resulting enhanced sample image retains most of the image information in the specified sample image while performing an image expansion process on the specified sample image. That is, based on a specified sample image, multiple enhanced sample images related to that specified sample image can be obtained.

[0121] For example, among multiple enhanced sample images, the main image region information corresponding to the main image region of the enhanced sample image is roughly the same, while the background image region information corresponding to the background image region of the enhanced sample image is different. Thus, based on a specified sample image, multiple enhanced sample images with different backgrounds can be obtained.

[0122] It is worth noting that the above are merely illustrative examples, and the embodiments of this application are not limited thereto.

[0123] In summary, based on probability distribution conditions, target sub-image regions are determined from multiple sub-image regions corresponding to a specified sample image, and candidate sub-image regions with registration relationships with the target sub-image regions are identified from candidate sample images. These candidate sub-image regions are then applied to the corresponding regions of the target sub-image regions to obtain enhanced sample images. Since the probability distribution conditions are determined based on the distribution patterns of the image subject within the image, the target sub-image region to be enhanced can largely avoid damaging the integrity of the image subject. This not only better protects the image information of the image subject but also allows the candidate sub-image regions to be used to expand other image regions in the specified sample image besides the image subject, resulting in a large number of enhanced sample images similar to the specified sample image, thus making the enhanced sample images more diverse. In few-shot learning, the model can be trained using specified sample images and similar enhanced sample images, thereby uncovering more image distribution patterns, overcoming the limitations of small sample datasets, and improving the robustness of the model.

[0124] In an optional embodiment, a two-dimensional normal distribution condition is used as the probability distribution condition, and at least one target sub-graph region that meets the probability requirements is determined from multiple sub-graph regions based on the two-dimensional normal distribution condition. (Illustrative example, such as...) Figure 5 As shown above, Figure 2 Step 230 in the illustrated embodiment can also be implemented as steps 510 to 530.

[0125] Step 510: Based on the two-dimensional normal distribution condition and the distance between the multiple sub-image regions and the center point of the specified sample image, determine the distribution probability corresponding to the multiple sub-image regions respectively.

[0126] In an optional embodiment, the image center point (sample image center point) of the specified sample image is determined based on the image shape of the specified sample image.

[0127] Optionally, when the image shapes corresponding to the specified sample images differ, the center points of the corresponding sample images may also differ.

[0128] For illustrative purposes, when a specified sample image is implemented as a symmetrical image, the image center point of the specified sample image is determined according to the graphics processing method corresponding to the symmetrical image. For example, when the specified sample image is implemented as a rectangle, the intersection of the diagonals of the specified sample image is taken as the image center point; or, when the specified sample image is implemented as an equilateral triangle, the intersection of the diagonals of the specified sample image is taken as the image center point.

[0129] Alternatively, when the specified sample image is implemented as an asymmetric image, the centroid of the asymmetric image can be used as the image center point of the specified sample image.

[0130] In an alternative embodiment, an image recognition model is used to determine the inferred center point of a given sample image.

[0131] Optionally, the image recognition model is a pre-trained model used for identifying subjects in an image. Illustratively, the image recognition model can analyze the image information of a specified sample image to roughly determine the range of the subject's region.

[0132] For example: The specified sample image is implemented as follows Figure 4 The kitten image shown is used to input a specified sample image into an image recognition model. The image recognition model then performs a subject recognition process on the kitten image to identify the subject kitten 410 from the kitten image.

[0133] Optionally, the image center point of a specified sample image can be determined based on the image subject region determined by the image recognition model.

[0134] Schematic illustration: Arbitrarily select a point from the main body region of the image as the center point of the specified sample image; or, use the centroid of the main body region of the image as the center point of the specified sample image, etc. For example: when using an image recognition model to analyze... Figure 4 After performing image recognition on the kitten image shown, a point 411 is randomly selected from the main subject kitten 410 in the image as the image center point of the specified sample image.

[0135] In an optional embodiment, the probability distribution of each of the multiple sub-image regions is determined based on the distance between the multiple sub-image regions and the center point of the specified sample image.

[0136] Optionally, when determining the distance between multiple sub-image regions and the image center point, the distance between the center point corresponding to the multiple sub-image regions and the image center point can be determined; or, the distance between the region vertices corresponding to the multiple sub-image regions and the image center point can be determined, etc.

[0137] Optionally, after determining the distances between multiple sub-image regions and the center point of a specified sample image, the distribution probabilities corresponding to the multiple sub-image regions are determined by combining the two-dimensional normal distribution condition.

[0138] Indicatively, the bivariate Gaussian probability function is used as a condition for the two-dimensional normal distribution; that is, the bivariate Gaussian probability function is used as a probability distribution condition. For example... Figure 6 The image shown is a schematic diagram of the probability function of a bivariate Gaussian distribution. Figure 6 It can be seen that the probability function of the bivariate Gaussian distribution is bowl-shaped, in which the middle region 610 has a lower probability of being selected, while the surrounding regions 620 have a higher probability of being selected.

[0139] Optionally, the binary Gaussian distribution probability function is applied to the process of determining the distribution probability corresponding to the sub-image region. The minimum value of the binary Gaussian distribution probability function is associated with the center point of the specified sample image. That is, the closer the sub-image region is to the center point of the specified sample image, the lower the probability that the sub-image region is used as the target sub-image region; the farther the sub-image region is from the center point of the specified sample image, the higher the probability that the sub-image region is used as the target sub-image region.

[0140] In an optional embodiment, the distribution probability corresponding to different sub-image regions is determined based on the two-dimensional normal distribution condition and the distance between multiple sub-image regions and the image center point.

[0141] Indicatively, a probability distribution image is constructed based on a two-dimensional normal distribution condition, where the probability distribution image corresponds to a specified sample image, for example, the probability distribution image and the specified sample image have the same image size.

[0142] Optionally, the explanation takes a two-dimensional normal distribution condition as a two-dimensional Gaussian probability function. Based on the division criteria of sub-image regions on a specified sample image, the probability distribution image corresponding to the two-dimensional normal distribution condition is divided to obtain multiple probability regions corresponding to the probability distribution image. Among them, there is a one-to-one correspondence between multiple probability regions and multiple sub-image regions.

[0143] Optionally, after aligning the center point of the probability distribution image with the center point of the specified sample image, the distance of the sub-image region from the center point of the specified sample image is determined, and the probability region corresponding to the sub-image region is determined in the probability distribution image. Thus, the probability distribution corresponding to the sub-image region is determined by the probability distribution corresponding to the probability region. Based on the above method, the probability distribution corresponding to multiple sub-image regions in the specified sample image is determined respectively.

[0144] Schematic, the multiple sub-image regions include a first sub-image region and a second sub-image region, wherein the first distribution probability of the first sub-image region is higher than the second distribution probability of the second sub-image region, and the first distance between the first sub-image region and the center point of the specified sample image is greater than the second distance between the second sub-image region and the center point of the specified sample image.

[0145] It is worth noting that the above are merely illustrative examples, and the embodiments of this application are not limited thereto.

[0146] Step 520: Obtain sub-graph regions whose distribution probability is higher than the probability threshold among multiple sub-graph regions.

[0147] In illustrative terms, a probability threshold is used to filter sub-image regions. Optionally, the probability threshold is a pre-set probability value condition, such as a pre-set probability threshold of 0.5; or, the probability threshold is the probability mean, for example, after determining the probability distribution corresponding to multiple sub-image regions, the probability mean of the multiple probability distributions is calculated, and the probability mean is used as the aforementioned probability threshold, etc.

[0148] To illustrate, after obtaining the distribution probabilities corresponding to multiple sub-regions, the distribution probabilities corresponding to the multiple sub-regions are compared with a probability threshold, and the sub-regions with distribution probabilities higher than the probability threshold are selected from the multiple sub-regions.

[0149] Step 530: Identify at least one target sub-graph region from the sub-graph regions whose distribution probability is higher than the probability threshold.

[0150] Indicatively, after obtaining multiple sub-graph regions with a probability distribution higher than a probability threshold, at least one target sub-graph region is determined from them. For example, at least one sub-graph region is selected as the target sub-graph region from multiple filtered sub-graph regions in a random manner, that is, the target sub-graph region is selected in an equally probable manner.

[0151] Alternatively, the probability distributions of the multiple selected sub-regions can be ranked, and the top g sub-regions with the highest probability distributions can be selected as the target sub-regions, where g is a positive integer. That is, the sub-region with the highest probability distribution can be selected as the target sub-region; or, multiple sub-regions with the highest probability distributions can be selected as the target sub-regions, etc.

[0152] It is worth noting that the above are merely illustrative examples, and the embodiments of this application are not limited thereto.

[0153] In summary, since the probability distribution conditions are determined based on the distribution patterns of the main subject in the image, they not only better protect the image information of the main subject, but also expand other image regions in the specified sample image besides the main subject using candidate sub-image regions, resulting in a large number of enhanced sample images. By enhancing the sample images, the diversity of the specified sample images is improved. When training the model, the obtained enhanced sample images can be used to discover more image distribution patterns, overcome the limitations of image datasets, and improve the training effect of the model.

[0154] In this embodiment, the case where the two-dimensional normal distribution condition is used as the probability distribution condition is described. Based on the two-dimensional normal distribution condition and the distance between multiple sub-image regions and the center point of the specified sample image, the distribution probability corresponding to each of the multiple sub-image regions is determined; sub-image regions with distribution probabilities higher than a probability threshold are obtained; and at least one target sub-image region is determined from the sub-image regions with distribution probabilities higher than the probability threshold. Since the two-dimensional normal distribution condition has a bowl-shaped probability distribution characteristic of "low in the middle and high around the edges," when using the two-dimensional normal distribution condition to determine the distribution probabilities corresponding to multiple sub-image regions, the closer the sub-image region is to the center point of the specified sample image, the lower the probability that it will be used as the target sub-image region, and the farther the sub-image region is from the center point of the specified sample image, the higher the probability that it will be used as the target sub-image region. Since the main body of the image is mostly located near the center point of the image, using the two-dimensional normal distribution condition as the probability distribution condition to determine the target sub-image region can effectively expand the specified sample image while avoiding damage to the main body of the image.

[0155] In an optional embodiment, a region replacement method is used to replace at least one target sub-image region in a specified sample image with at least one candidate sub-image region to obtain an enhanced sample image. (Illustrative example, such as...) Figure 7 As shown above, Figure 2 The illustrated embodiment can also be implemented as follows: steps 710 to 750.

[0156] Step 710: Obtain the specified sample image and candidate sample images.

[0157] The specified sample image is the image to be augmented using the candidate sample image.

[0158] Step 710 has already been explained in step 210 above, and will not be repeated here.

[0159] Step 720: Divide the specified sample image into regions to obtain multiple sub-image regions in the specified sample image.

[0160] Optionally, after obtaining the specified sample image, the specified sample image is divided into regions using an equal area division method to obtain multiple sub-image regions in the specified sample image; or, after obtaining the specified sample image, the specified sample image is divided into regions using a non-equal area division method to obtain multiple sub-image regions in the specified sample image.

[0161] In illustrative terms, multiple sub-image regions can be implemented as all or part of a specified sample image. For example, after dividing the specified sample image into regions, all the resulting image regions can be used as the aforementioned multiple sub-image regions. Then, by stitching together the multiple sub-image regions, a complete specified sample image can be obtained. Alternatively, after dividing the specified sample image into regions, a portion of the resulting image regions can be selected and used as the aforementioned multiple sub-image regions. Then, by stitching together the multiple sub-image regions, a partial specified sample image can be obtained, and so on.

[0162] Step 720 has been described in step 220 above, and will not be repeated here.

[0163] Step 730: Based on the probability distribution conditions, determine at least one target sub-graph region that meets the probability requirements from multiple sub-graph regions, as the sub-graph region to be enhanced.

[0164] Among them, the probability distribution condition is determined based on the distribution pattern of the main subject in the image.

[0165] In an optional embodiment, based on the two-dimensional normal distribution condition and the distance between the multiple sub-image regions and the center point of the specified sample image, the distribution probability corresponding to each of the multiple sub-image regions is determined; the sub-image regions with distribution probabilities higher than a probability threshold are obtained from the multiple sub-image regions; at least one target sub-image region is determined from the sub-image regions with distribution probabilities higher than the probability threshold.

[0166] Step 730 has already been explained in step 230 above, and will not be repeated here.

[0167] Step 740: Based on the registration relationship between the specified sample image and the candidate sample image, determine the pairing relationship between the n target sub-image regions and the n candidate sub-image regions.

[0168] In this context, the i-th target subgraph region is paired with the i-th candidate subgraph region, where 0 < i ≤ n, and i is an integer.

[0169] In an optional embodiment, at least one candidate sub-image region that matches at least one target sub-image region is determined from the candidate sample image.

[0170] To illustrate, after obtaining the specified sample image and the candidate sample image, image registration is performed on the specified sample image and the candidate sample image.

[0171] The purpose of image registration is to compare or fuse the acquired images so that different images correspond one-to-one with points at the same location in space, thereby achieving the purpose of image information fusion.

[0172] Optionally, a specified sample image is used as the target image for image registration, and a candidate sample image is registered onto the specified sample image; or, a candidate sample image is used as the target image for image registration, and a specified sample image is registered onto the candidate sample image.

[0173] Schematic illustration: Based on the image registration operation described above, the registration relationship between a specified sample image and a candidate sample image can be determined. Optionally, the registration relationship is used to indicate the positional coordinate correspondence between the specified sample image and the candidate sample image.

[0174] For example: After performing image registration on the specified sample image and the candidate sample image, a coordinate system is established with the registration center as the origin. After determining the position coordinates of point A in the specified sample image, the relative position relationship between the position coordinates of point A and the origin of the coordinate system is determined. Based on the registration relationship and the relative position relationship, the position coordinates of point B corresponding to the position coordinates of point A are determined from the candidate sample image.

[0175] Similarly, based on the relative relationship of position coordinates, after determining at least one target sub-image region in the specified sample image, candidate sub-image regions corresponding to the at least one target sub-image region are determined from the candidate sample images.

[0176] Schematic, at least one target sub-image region in the specified sample image includes target sub-image region T1 and target sub-image region T2. ​​Based on the relative position coordinates between the specified sample image and the candidate sample image, candidate sub-image region C1 corresponding to target sub-image region T1 and candidate sub-image region C2 corresponding to target sub-image region T2 are selected from the specified sample image.

[0177] Step 750: Replace the i-th target sub-image region with the i-th candidate sub-image region, and iterate through the replacement between the n target sub-image regions and the n candidate sub-image regions to obtain the enhanced sample image.

[0178] Indicatively, after obtaining n candidate sub-image regions corresponding to n target sub-image regions, the i-th candidate sub-image region corresponding to the i-th target sub-image region is determined. When obtaining the enhanced sample image through the specified sample image and the candidate sample image, the i-th target sub-image region in the specified sample image is replaced with the i-th candidate sub-image region.

[0179] Based on the above replacement method, n target subgraph regions are replaced, thereby replacing the n target subgraphs with their respective candidate subgraph regions, that is: replacing the n candidate subgraph regions with the corresponding target subgraph regions.

[0180] Since n is a positive integer and 0 < i ≤ n, i is an integer, the above replacement process includes at least two replacement forms as follows.

[0181] (1) Obtain a target sub-image region from the specified sample image. Based on the registration relationship between the specified sample image and the candidate sample image, determine the candidate sub-image region in the candidate sample image that corresponds to the target sub-image region. When the enhanced sample image corresponding to the specified sample image is obtained, replace the target sub-image region with the candidate sub-image region to obtain the enhanced sample image.

[0182] (2) Obtain at least two target sub-image regions from the specified sample image. Based on the registration relationship between the specified sample image and the candidate sample image, determine the candidate sub-image regions in the candidate sample image that correspond to the at least two target sub-image regions respectively. When obtaining the enhanced sample image corresponding to the specified sample image, replace the at least two target sub-image regions with the corresponding candidate sub-image regions respectively to obtain the enhanced sample image.

[0183] It is worth noting that the above are merely illustrative examples, and the embodiments of this application are not limited thereto.

[0184] Indicative, as Figure 3 The lion image shown (image 310) will be used as a candidate sample image, and will be as follows: Figure 4 The kitten image 410 shown is used as a specified sample image.

[0185] Optionally, the block division method is used for, for example Figure 4 After the specified sample image is divided into regions, multiple sub-image regions are obtained. The first target sub-image region 420, the second target sub-image region 430, and the third target sub-image region 440 in the specified sample image are used as examples for illustration.

[0186] Indicative, such as Figure 3 and Figure 4 As shown, based on the registration relationship between the candidate sample image and the specified sample image, in such a way... Figure 3 In the candidate sample image shown, a first candidate sub-image region 320 corresponding to the first target sub-image region 420 in the specified sample image is determined, in such a way as Figure 3 In the candidate sample image shown, a second candidate sub-image region 330 corresponding to the second target sub-image region 430 in the specified sample image is determined, in such a way as Figure 3The candidate sample image shown is used to determine the third candidate sub-image region 340 that corresponds to the third target sub-image region 440 in the specified sample image.

[0187] When obtaining enhanced sample images, in such cases... Figure 4 In the specified sample image shown, the first target sub-image region 420 is replaced with as shown. Figure 3 The first candidate sub-image region 320 shown is replaced with the second target sub-image region 430 as shown. Figure 3 The second candidate sub-map region 330 shown is used to replace the third target sub-map region 440 as shown. Figure 3 The third candidate subgraph region 340 shown is thus obtained as follows: Figure 8 The enhanced sample image shown, wherein the first region 810 in the enhanced sample image is as follows: Figure 3 The first candidate sub-image region 320 shown, and the second region 820 in the enhanced sample image are as follows: Figure 3 The second candidate sub-image region 330 shown, and the third region 830 in the enhanced sample image are as follows: Figure 3 The third candidate subgraph region 340 is shown.

[0188] In an optional embodiment, the specified sample image is labeled with a specified sample tag.

[0189] In a specified sample image, at least one candidate sub-image region is applied to the region location of at least one target sub-image region, and the specified sample label is used as the sample label corresponding to the enhanced sample image.

[0190] Indicative, such as Figure 4 The kitten image 410 shown is a specified sample image, and the specified sample label corresponding to this specified sample image is "kitten". In the following... Figure 3 After replacing the candidate sub-image region in the lion image 310 with the corresponding target sub-image region in the kitten image 410, the specified sample label "kitten" is used as... Figure 8 The sample labels corresponding to the enhanced sample images shown.

[0191] It is worth noting that the above are merely illustrative examples, and the embodiments of this application are not limited thereto.

[0192] In an optional embodiment, augmented sample images are used to train the image recognition model to be trained.

[0193] Optionally, the enhanced sample image is input into the image recognition model; the image prediction result output by the image recognition model is obtained; the loss value is determined based on the difference between the image prediction result and the specified sample label corresponding to the enhanced sample image; the image recognition model is trained with the loss value to obtain the target image recognition model.

[0194] Among them, the target image recognition model is used to perform image recognition on the image to be recognized.

[0195] In summary, since the probability distribution conditions are determined based on the distribution patterns of the main subject in the image, they not only better protect the image information of the main subject, but also expand other image regions in the specified sample image besides the main subject using candidate sub-image regions, resulting in a large number of enhanced sample images. By enhancing the sample images, the diversity of the specified sample images is improved. When training the model, the obtained enhanced sample images can be used to discover more image distribution patterns, overcome the limitations of image datasets, and improve the training effect of the model.

[0196] In this embodiment, the process of replacing at least one target sub-image region in a specified sample image with at least one candidate sub-image region is described. After determining the target sub-image regions in the specified sample image, based on the registration relationship between the specified sample image and the candidate sample image, a pairing relationship between n target sub-image regions and n candidate sub-image regions is determined, that is, there is a one-to-one correspondence between the n target sub-image regions and the n candidate sub-image regions. When obtaining the enhanced sample image, the i-th target sub-image region is replaced with the i-th candidate sub-image region, and the replacement between the n target sub-image regions and the n candidate sub-image regions is iteratively completed. Since n is a positive integer, compared with the specified sample image, the enhanced sample image is a similar image obtained after replacing at least one target sub-image region. The enhanced sample image can be used to assist in the training of the model, enabling the model to learn more patterns of different image background information, better identify the main subject of the image, and improve the training effect of the model.

[0197] In an optional embodiment, the above-described sample image generation method is applied to image data augmentation scenarios in deep learning. Illustratively, when there is insufficient sample image data available for training, or when the annotation cost of sample image data is high, the above-described sample image generation method can augment the limited sample image data to obtain enhanced sample images.

[0198] Optionally, deep learning-based image data augmentation scenarios include: image classification, object detection, and image segmentation. When using an analysis model to analyze image data, it is usually necessary to train the model with a large amount of sample image data so that the model can learn the patterns in the sample image data and analyze image data related to the sample image data.

[0199] However, the amount of sample image data is usually small, and it is typically presented as follows: Figure 9The distribution pattern of the long-tail data shown indicates that there are more head data (910) and fewer tail data (920). When tail data (920) is used as sample image data to train the analysis model, the limited number of sample images will greatly affect the training effect of the model, that is, the effect of small sample learning is poor. In addition, the head data (910) stores the sample set of the category with sufficient samples, while the tail data (920) stores the sample set of the category with scarce samples. When cross-category analysis of sample image data is required, it is usually necessary to use sample image data from multiple categories to train the analysis model. When tail data (920) is used as sample image data, the limited number of tail data (920) will also make it difficult for the analysis model to learn the differences between head data (910) and tail data (920), resulting in poor cross-domain learning.

[0200] Optionally, the above sample image generation method is implemented in the following four parts: (i) constructing a sample image set; (ii) constructing enhanced samples; (iii) constructing a probability distribution; and (iv) data augmentation.

[0201] (I) Constructing a sample image set

[0202] Optionally, by constructing a sample image set, a rich selection space can be provided for the specified sample images to be augmented.

[0203] For illustrative purposes, the sample images in the sample image set are a collection of unprocessed images from nature. Optionally, when acquiring sample images, it is not necessary to obtain their sample labels; that is, the sample image set can be implemented as an unlabeled dataset.

[0204] Optionally, sample images from natural scenes mostly conform to the following... Figure 9 The long-tail distribution is shown. Optionally, when using everyday images as specified sample images for sample augmentation, the set of sample images corresponding to the head data 910 is used as the sample image set; or, in few-shot incremental learning, i.e., when using scarce images as specified sample images for sample augmentation, the set of sample images corresponding to the tail data 920 is used as the sample image set, etc.; or, any dataset from a relatively similar domain is selected as the sample dataset, such as collected images, standard datasets on the Internet (e.g., ImageNet dataset), etc.

[0205] (II) Constructing Enhanced Samples

[0206] Optionally, after obtaining the sample dataset for enhancement, assuming that a sample image A needs to be enhanced, that is, the sample image A is used as the specified sample image to be enhanced, and the sample image A is transformed by the sample image generation method to obtain multiple transformed images related to the sample image A.

[0207] Optionally, sample images from the sample image set constructed above can be used to perform sample augmentation on sample image A. For example, multiple samples can be arbitrarily selected from the sample image set, denoted as B1, B2, B3, ...; then, sample image A can be divided into N*N squares horizontally and vertically, such as... Figure 10 As shown, after dividing the sample image A1010, a total of N are obtained. 2 1020 small squares.

[0208] Optionally, each small square is referred to as a point according to its horizontal and vertical coordinate positions. Then, when performing sample enhancement on sample image A using the sample image generation method, the enhancement process of sample image A is completed by replacing these points.

[0209] (III) Construction of probability distribution

[0210] Optionally, a standardized negative logarithmic composite bivariate Gaussian distribution or a two-dimensional Dirichlet distribution with equal probability distribution functions can be used for N. 2 Sampling is performed on small squares.

[0211] Indicative, such as Figure 6 As shown, a standardized binary Gaussian distribution composed of negative logarithms is used as the probability distribution function, which is visually represented as a bowl-shaped distribution.

[0212] Optionally, based on this probability distribution function, for N 2 N corresponding to each small square 2 Each point is sampled according to probability, resulting in several selected points after sampling. Illustratively, the above N... 2 Each small square corresponds to an N*N space, which is a two-dimensional matrix. The elements corresponding to the selected points are marked as 1, and the elements corresponding to the unselected points are marked as 0. This two-dimensional matrix is ​​called a mask.

[0213]

[0214] in, It is used to indicate a two-dimensional normal distribution; (x1,x2) is used to indicate any point; p(x1,x2) is used to indicate the probability that point (x1,x2) is selected; Used to indicate the center coordinates of the mask, where The x-axis is... The vertical axis is denoted by .

[0215] Schematic, using the above probability distribution formula, we obtain as follows: Figure 6 The diagram shows the probability distribution function of a bivariate Gaussian distribution.

[0216] Optionally, the above probability distribution construction process makes it easier to sample the surrounding areas of the specified sample image and the candidate sample image, thereby enabling the process of replacing different backgrounds for the specified sample image while protecting the main body of the specified sample image. Furthermore, the probability distribution function diagram obtained using the above probability distribution construction process also ensures that the central part has a certain low probability of being sampled, thus increasing the flexibility and diversity of sample augmentation.

[0217] (iv) Regional replacement

[0218] Optionally, after obtaining the mask representing the probability distribution, the mask M is applied to the specified sample image A and other candidate sample images B1, B2, B3, etc., used for the sample augmentation process. Here, M indicates the number of small squares marked as 1.

[0219] To illustrate, let's take a specified sample image A and a candidate sample image B1 as an example of sample enhancement. When blending the specified sample image A and the candidate sample image B1, pixel-level blending is performed on the two images. Since the specified sample image A and the candidate sample image B1 are divided into N*N small squares, and the mask M is also of N*N dimensions, the pixel-level blending process is performed using the following formula.

[0220]

[0221]

[0222] Where M indicates the small square marked as 1 (the selected target sub-graph region); (1-M) indicates the small square marked as 0 (the unselected sub-graph region); x A Used to indicate the x-coordinate of a specified sample image A; x B Used to indicate the x-coordinate of candidate sample image B; Used to indicate the x-axis of the enhanced sample image; Used to indicate the ordinate of the enhanced sample image; ⊙ is used to indicate the element-wise multiplication.

[0223] Optionally, a specified sample image A has a corresponding sample label a. After obtaining the enhanced sample image, the label of the enhanced sample image is the sample label a corresponding to the specified sample image A. That is, compared with related technologies, the idea of ​​the sample image generation method provided in this application is to enhance the image background, so there is no need to use label smoothing, nor is there a need for the model to associate the adjusted and generated enhanced sample image with the candidate sample image used for sample enhancement.

[0224] It is worth noting that the above are merely illustrative examples, and the embodiments of this application are not limited thereto.

[0225] In an optional embodiment, the above-described method for generating sample images can be applied to various tasks involving image data, such as image classification, object detection, and image segmentation.

[0226] Optionally, when the above sample image generation method is applied to an image classification task, since the image classification data is relatively simple single-subject data, fewer target sub-image regions can be selected from the specified sample image (e.g., the sub-image regions are sparsely divided) and a sample enhancement process can be performed; or, when the above sample image generation method is applied to an image segmentation task, since the training set data for image segmentation generally has multiple subjects, more target sub-image regions can be selected from the specified sample image (e.g., the sub-image regions are densely divided) and a sample enhancement process can be performed.

[0227] Indicative, such as Figure 11 As shown, the above sample image generation method is used as the data augmentation module 1110. In the data augmentation module 1110, candidate sample images 1111 are obtained from the sample image set. The above sample augmentation process is performed on the specified sample image 1112 using the candidate sample images 1111 to obtain the augmented sample image 1113. That is, a specified sample image is expanded into multiple augmented sample images by the sample image generation method. Then, the model training process is performed using the expanded augmented sample images.

[0228] To illustrate, multiple enhanced sample images 1113 are input into the feature extraction module 1120 and passed through a fully-connected layer to obtain better feature results.

[0229] In other words, the data augmentation module 1110, determined using the aforementioned sample image generation method, is inserted as a separate module into the model's input. This allows the sample image generation method to be applied to numerous models, thus enabling its integration with various image processing methods. For instance, when combined with methods such as object detection and image segmentation, inserting the data augmentation module 1110 obtained using the sample image generation method into the model's input achieves improved performance without altering other parts of the method.

[0230] It is worth noting that the above are merely illustrative examples, and different probability distribution functions can be selected according to different circumstances to achieve the best results. This application does not limit this.

[0231] In summary, since the probability distribution conditions are determined based on the distribution patterns of the main subject in the image, they not only better protect the image information of the main subject, but also expand other image regions in the specified sample image besides the main subject using candidate sub-image regions, resulting in a large number of enhanced sample images. By enhancing the sample images, the diversity of the specified sample images is improved. When training the model, the obtained enhanced sample images can be used to discover more image distribution patterns, overcome the limitations of image datasets, and improve the training effect of the model.

[0232] In this embodiment, a sample image is enhanced using a sample image generation method to enrich the image background while protecting the main body of the sample image from damage. Instead of training the model using only a small number of sample images, the enhanced sample images are used to train the model more comprehensively, resulting in stronger robustness and better performance in image classification, object detection, and image segmentation scenarios. The algorithm for generating the sample images described above is simple, easy to operate, plug-and-play, and requires no additional training auxiliary networks, making it a low-cost enhancement method.

[0233] Figure 12 This is a structural block diagram of a sample image generation apparatus provided in an exemplary embodiment of this application, such as... Figure 12 As shown, the device includes the following parts:

[0234] The acquisition module 1210 is used to acquire a specified sample image and a candidate sample image, wherein the specified sample image is an image to be sampled and enhanced by the candidate sample image;

[0235] The partitioning module 1220 is used to partition the specified sample image into multiple sub-image regions in the specified sample image.

[0236] The determining module 1230 is used to determine at least one target sub-image region that meets the probability requirements from the plurality of sub-image regions based on probability distribution conditions, as the sub-image region to be enhanced, wherein the probability distribution conditions are determined based on the distribution pattern of the image subject in the image;

[0237] The registration module 1240 is used to determine at least one candidate sub-image region that matches the at least one target sub-image region from the candidate sample image based on the registration relationship between the specified sample image and the candidate sample image;

[0238] Application module 1250 is used to apply the at least one candidate sub-image region to the region location of the at least one target sub-image region in the specified sample image to obtain an enhanced sample image, wherein the enhanced sample image is a sample image generated after adjusting the specified sample image.

[0239] In an optional embodiment, the registration module 1240 is further configured to determine at least one target sub-image region that meets the probability requirements from the plurality of sub-image regions based on a two-dimensional normal distribution condition; wherein, the first distribution probability of the first sub-image region in the specified sample image is higher than the second distribution probability of the second sub-image region, and the first distance between the first sub-image region and the center point of the specified sample image is greater than the second distance between the second sub-image region and the center point of the specified sample image.

[0240] In an optional embodiment, the registration module 1240 is further configured to determine the distribution probability corresponding to the multiple sub-image regions based on the two-dimensional normal distribution condition and the distance between the multiple sub-image regions and the center point of the specified sample image; obtain the sub-image regions in the multiple sub-image regions whose distribution probability is higher than the probability threshold; and determine the at least one target sub-image region from the sub-image regions whose distribution probability is higher than the probability threshold.

[0241] In an optional embodiment, the application module 1250 is further configured to replace at least one target sub-image region in the specified sample image with at least one candidate sub-image region to obtain the enhanced sample image.

[0242] In an optional embodiment, the specified sample image includes n target sub-image regions, and the candidate sample image includes n candidate sub-image regions, where n is a positive integer;

[0243] The application module 1250 is further configured to determine the pairing relationship between the n target sub-image regions and the n candidate sub-image regions based on the registration relationship between the specified sample image and the candidate sample image, wherein the i-th target sub-image region is paired with the i-th candidate sub-image region, 0 < i ≤ n, and i is an integer; replace the i-th target sub-image region with the i-th candidate sub-image region, and iteratively complete the replacement between the n target sub-image regions and the n candidate sub-image regions to obtain the enhanced sample image.

[0244] In an optional embodiment, the specified sample image is labeled with a specified sample tag;

[0245] The application module 1250 is further configured to apply the at least one candidate sub-image region to the region location of the at least one target sub-image region in the specified sample image, and use the specified sample label as the sample label corresponding to the enhanced sample image.

[0246] In an optional embodiment, the application module 1250 is further configured to input the enhanced sample image into an image recognition model, wherein the image recognition model is a recognition model to be trained; obtain the image prediction result output by the image recognition model; determine a loss value based on the difference between the image prediction result and the specified sample label corresponding to the enhanced sample image; and train the image recognition model with the loss value to obtain a target image recognition model, wherein the target image recognition model is used to perform image recognition on the image to be recognized.

[0247] In an optional embodiment, the probability distribution condition is a condition determined by multiple sample region images, which are pre-collected image data;

[0248] The registration module 1240 is further configured to perform region recognition on the plurality of sample region images through a region recognition model, determine the image subject region corresponding to the plurality of sample region images respectively, the image subject region being used to indicate the image region where the image subject is located in the sample region image; and comprehensively analyze the regional positions of the plurality of image subject regions in the corresponding sample region images to determine the probability distribution conditions.

[0249] In an optional embodiment, the plurality of image subject regions include a first image subject region and a second image subject region, wherein the first image subject region is used to indicate the image subject region corresponding to the first sample region image, and the second image subject region is used to indicate the image subject region corresponding to the second sample region image.

[0250] The registration module 1240 is further configured to determine the first region position of the first image subject region in the first sample region image, and the second region position of the second image subject region in the second sample region image; by comprehensively analyzing the first region position and the second region position, the distribution pattern of the subject region is determined, and the distribution pattern of the subject region is used as the probability distribution condition.

[0251] In an optional embodiment, the registration module 1240 is further configured to register the candidate sample image onto the designated sample image, using the designated sample image as the target registration image, and determine the registration relationship between the designated sample image and the candidate sample image.

[0252] In summary, since the probability distribution conditions are determined based on the distribution patterns of the main subject in the image, the aforementioned sample image generation device can not only better protect the image information of the main subject, but also expand other image regions in the specified sample image besides the main subject using candidate sub-image regions, obtaining a large number of enhanced sample images. By enhancing the sample images, the diversity of the specified sample images is improved. When training the model, the obtained enhanced sample images can be used to discover more image distribution patterns, overcome the limitations of image datasets, and improve the training effect of the model.

[0253] It should be noted that the sample image generation apparatus provided in the above embodiments is only an example of the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the sample image generation apparatus and the sample image generation method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.

[0254] Figure 13 This illustration shows a schematic diagram of a server provided in an exemplary embodiment of this application. The server 1300 includes a Central Processing Unit (CPU) 1301, a system memory 1304 including Random Access Memory (RAM) 1302 and Read Only Memory (ROM) 1303, and a system bus 1305 connecting the system memory 1304 and the CPU 1301. The server 1300 also includes a mass storage device 1306 for storing an operating system 1313, application programs 1314, and other program modules 1315.

[0255] Mass storage device 1306 is connected to central processing unit 1301 via a mass storage controller (not shown) connected to system bus 1305. Mass storage device 1306 and its associated computer-readable media provide non-volatile storage for server 1300. That is, mass storage device 1306 may include computer-readable media (not shown) such as hard disk or compact disc read-only memory (CD-ROM) drives.

[0256] Without loss of generality, computer-readable media can include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented using any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include RAM, ROM, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other solid-state storage technologies, CD-ROM, digital versatile disc (DVD) or other optical storage, magnetic tape cassettes, magnetic tape, disk storage, or other magnetic storage devices. Of course, those skilled in the art will recognize that computer storage media are not limited to the above-mentioned types. The system memory 1304 and mass storage device 1306 described above can be collectively referred to as memory.

[0257] According to various embodiments of this application, server 1300 can also be connected to a remote computer on a network, such as the Internet. That is, server 1300 can be connected to network 1312 via network interface unit 1311 connected to system bus 1305, or it can also use network interface unit 1311 to connect to other types of networks or remote computer systems (not shown).

[0258] The aforementioned memory also includes one or more programs, which are stored in the memory and configured to be executed by the CPU.

[0259] Embodiments of this application also provide a computer device, which includes a processor and a memory. The memory stores at least one instruction, at least one program, code set, or instruction set. The at least one instruction, at least one program, code set, or instruction set is loaded and executed by the processor to implement the sample image generation method provided in the above-described method embodiments.

[0260] Embodiments of this application also provide a computer-readable storage medium storing at least one instruction, at least one program, code set, or instruction set, wherein the at least one instruction, at least one program, code set, or instruction set is loaded and executed by a processor to implement the sample image generation method provided in the above-described method embodiments.

[0261] Embodiments of this application also provide a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform any of the sample image generation methods described in the above embodiments.

[0262] Optionally, the computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), solid-state drives (SSDs), or optical discs, etc. The random access memory may include resistive random access memory (ReRAM) and dynamic random access memory (DRAM). The sequence numbers of the embodiments in this application are merely descriptive and do not represent the superiority or inferiority of the embodiments.

[0263] Those skilled in the art will understand that all or part of the steps of the above embodiments can be implemented by hardware or by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as a read-only memory, a disk, or an optical disk.

[0264] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for generating a sample image, characterized in that, The method includes: Acquire a specified sample image and a candidate sample image, wherein the specified sample image is the image to be enhanced using the candidate sample image; The specified sample image is divided into regions to obtain multiple sub-image regions in the specified sample image; Based on probability distribution conditions, at least one target sub-image region that meets the probability requirements is determined from the plurality of sub-image regions as the sub-image region to be enhanced, wherein the probability distribution conditions are determined based on the distribution pattern of the image subject in the image; Based on the registration relationship between the specified sample image and the candidate sample image, at least one candidate sub-image region that matches the at least one target sub-image region is determined from the candidate sample image; The at least one candidate sub-image region is applied to the region location of the at least one target sub-image region in the specified sample image to obtain an enhanced sample image, wherein the enhanced sample image is a sample image generated after adjusting the specified sample image.

2. The method according to claim 1, characterized in that, The step of determining at least one target sub-graph region that meets the probability requirements from the plurality of sub-graph regions based on probability distribution conditions includes: Based on the two-dimensional normal distribution condition, at least one target sub-graph region that meets the probability requirements is determined from the plurality of sub-graph regions; Wherein, the first distribution probability of the first sub-image region in the specified sample image is higher than the second distribution probability of the second sub-image region, and the first distance between the first sub-image region and the center point of the specified sample image is greater than the second distance between the second sub-image region and the center point of the specified sample image.

3. The method according to claim 2, characterized in that, The determination of at least one target sub-map region that meets the probability requirements from the plurality of sub-map regions based on the two-dimensional normal distribution condition includes: Based on the two-dimensional normal distribution condition and the distance between the multiple sub-image regions and the center point of the specified sample image, the distribution probability corresponding to the multiple sub-image regions is determined respectively; Identify the sub-graph regions whose distribution probability is higher than a probability threshold among the plurality of sub-graph regions; The at least one target sub-graph region is determined from sub-graph regions whose distribution probability is higher than the probability threshold.

4. The method according to any one of claims 1 to 3, characterized in that, The step of applying the at least one candidate sub-image region to the region location of the at least one target sub-image region in the specified sample image to obtain an enhanced sample image includes: The enhanced sample image is obtained by replacing at least one target sub-image region in the specified sample image with at least one candidate sub-image region.

5. The method according to claim 4, characterized in that, The specified sample image includes n target sub-image regions, and the candidate sample image includes n candidate sub-image regions, where n is a positive integer; The step of replacing at least one target sub-image region in the specified sample image with at least one candidate sub-image region to obtain the enhanced sample image includes: Based on the registration relationship between the specified sample image and the candidate sample image, the pairing relationship between the n target sub-image regions and the n candidate sub-image regions is determined, wherein the i-th target sub-image region is paired with the i-th candidate sub-image region, 0 < i ≤ n, and i is an integer; The i-th target sub-image region is replaced with the i-th candidate sub-image region, and the replacement between the n target sub-image regions and the n candidate sub-image regions is iterated to obtain the enhanced sample image.

6. The method according to any one of claims 1 to 3, characterized in that, The specified sample images are labeled with specified sample tags; The method further includes: In the specified sample image, the at least one candidate sub-image region is applied to the region location of the at least one target sub-image region, and the specified sample label is used as the sample label corresponding to the enhanced sample image.

7. The method according to claim 6, characterized in that, After applying the at least one candidate sub-image region to the region location of the at least one target sub-image region in the specified sample image to obtain the enhanced sample image, the method further includes: The enhanced sample image is input into the image recognition model, which is the recognition model to be trained. Obtain the image prediction result output by the image recognition model; The loss value is determined based on the difference between the image prediction result and the specified sample label corresponding to the enhanced sample image; The image recognition model is trained using the loss value to obtain a target image recognition model, which is used to perform image recognition on the image to be recognized.

8. The method according to any one of claims 1 to 3, characterized in that, The probability distribution condition is determined by multiple sample region images, which are image data collected in advance. The method further includes: The region recognition model is used to perform region recognition on the multiple sample region images to determine the image subject region corresponding to each of the multiple sample region images. The image subject region is used to indicate the image region where the image subject is located in the sample region image. By comprehensively analyzing the regional locations of multiple image subject regions in the corresponding sample region images, the probability distribution conditions are determined.

9. The method according to claim 8, characterized in that, The plurality of image subject regions include a first image subject region and a second image subject region. The first image subject region is used to indicate the image subject region corresponding to the first sample region image, and the second image subject region is used to indicate the image subject region corresponding to the second sample region image. The comprehensive analysis of the regional locations of multiple image subject regions in the corresponding sample region image to determine the probability distribution conditions includes: Determine the first region position of the first image subject region in the first sample region image, and the second region position of the second image subject region in the second sample region image; By comprehensively analyzing the locations of the first and second regions, the distribution pattern of the main regions is determined, and the distribution pattern of the main regions is used as the probability distribution condition.

10. The method according to any one of claims 1 to 3, characterized in that, The method further includes: Using the specified sample image as the target registration image, the candidate sample image is registered onto the specified sample image to determine the registration relationship between the specified sample image and the candidate sample image.

11. A sample image generation apparatus, characterized in that, The device includes: The acquisition module is used to acquire a specified sample image and a candidate sample image, wherein the specified sample image is the image to be sampled and enhanced using the candidate sample image; The segmentation module is used to segment the specified sample image into regions to obtain multiple sub-image regions in the specified sample image. The determination module is used to determine at least one target sub-image region that meets the probability requirements from the plurality of sub-image regions based on probability distribution conditions, as the sub-image region to be enhanced, wherein the probability distribution conditions are determined based on the distribution pattern of the image subject in the image; A registration module is used to determine at least one candidate sub-image region that matches the at least one target sub-image region from the candidate sample image based on the registration relationship between the specified sample image and the candidate sample image; An application module is used to apply the at least one candidate sub-image region to the region location of the at least one target sub-image region in the specified sample image to obtain an enhanced sample image, wherein the enhanced sample image is a sample image generated after adjusting the specified sample image.

12. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing at least one instruction, which is loaded and executed by the processor to implement the method for generating sample images as described in any one of claims 1 to 10.

13. A computer-readable storage medium, characterized in that, The storage medium stores at least one instruction, which is loaded and executed by a processor to implement the method for generating sample images as described in any one of claims 1 to 10.

14. A computer program product, characterized in that, It includes computer instructions that, when executed by a processor, implement the method for generating sample images as described in any one of claims 1 to 10.

Citation Information

Patent Citations

  • X-ray sample image generation method, X-ray sample image generation equipment and storage device

    CN111242905A

  • Sample image increment method, image detection model training method and image detection method

    CN112949767A