Image augmentation method and apparatus, electronic device, and storage medium

CN115797727BActive Publication Date: 2026-09-11IFLYTEK SOUTH CHINA ARTIFICIAL INTELLIGENCE RES INST GUANGZHOU CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211678417.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-26
Publication Date
2026-09-11
Estimated Expiration
2042-12-26

AI Technical Summary

Technical Problem

[0004]本发明提供一种图像增广方法、装置、电子设备和存储介质,用以解决现有技术中无法增广得到新视角下的增广图像的缺陷,实现高效的图像增广

Benefits of technology

[0042] The image augmentation method, apparatus, electronic device, and storage medium provided by this invention can train an image augmentation model based on sample images from multiple viewpoints. This model represents a three-dimensional volume model, i.e., it reconstructs the three-dimensional volume model corresponding to the sample images from multiple viewpoints through inverse rendering. Therefore, by only determining the image shooting direction information and voxel position information corresponding to the target viewpoint, voxel sampling can be performed on the three-dimensional volume model represented by the image augmentation model based on the image shooting direction information and voxel position information to obtain the color and transparency information of multiple voxels. Then, based on each color and transparency information, an augmented image corresponding to the target viewpoint is generated, thereby augmenting the image to a new viewpoint. This effectively augments the sample data in terms of viewpoint, achieving efficient image augmentation. When training a model based on viewpoint-augmented sample data, this improves the model training effect and enhances the robustness of the machine learning model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115797727B_ABST
    Figure CN115797727B_ABST
Patent Text Reader

Abstract

The application provides an image augmentation method, device, electronic equipment and storage medium. The method comprises: determining image shooting direction information corresponding to a target view angle and voxel position information, the voxel position information comprising the spatial positions of a plurality of voxels; performing voxel sampling on a three-dimensional volume model represented by an image augmentation model based on the image shooting direction information and the voxel position information to obtain color information and transparency information of the plurality of voxels; and generating an augmented image corresponding to the target view angle based on the color information and the transparency information. The method, device, electronic equipment and storage medium provided by the application train an image augmentation model based on sample images of a plurality of view angles, so that the image augmentation model represents a three-dimensional volume model, so that only the image shooting direction information corresponding to the target view angle and the voxel position information need to be determined to obtain an augmented image under a new view angle, the sample data is effectively augmented in the view angle, and efficient image augmentation is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to an image augmentation method, apparatus, electronic device, and storage medium. Background Technology

[0002] With the rapid development of artificial intelligence technology, the application scenarios of various machine learning models are becoming increasingly widespread. Some of these models require massive amounts of image or video data as sample data for training, testing, and validation. Therefore, image augmentation is necessary for both image and video data.

[0003] Currently, image augmentation can be performed using geometric transformations such as horizontal flipping, random scaling, center cropping, and translation. It can also be achieved through pixel transformations such as noise reduction, color transformation, brightness adjustment, contrast adjustment, and saturation transformation. Furthermore, it can be done through content editing methods such as foreground segmentation, background replacement, random erasure, and cross-image manipulation. However, existing technologies cannot augment images from new perspectives, resulting in ineffective augmentation of sample data in terms of viewpoint. This leads to poor model training performance and reduced robustness of machine learning models. Summary of the Invention

[0004] This invention provides an image augmentation method, apparatus, electronic device, and storage medium to overcome the shortcomings of existing technologies that cannot augment images from new perspectives, thereby achieving efficient image augmentation.

[0005] This invention provides an image augmentation method, comprising:

[0006] Determine the image shooting direction information and voxel position information corresponding to the target viewpoint, wherein the voxel position information includes the spatial position of multiple voxels;

[0007] Based on the image shooting direction information and the voxel position information, voxel sampling is performed on the three-dimensional volume model represented by the image augmentation model to obtain the color information and transparency information of the multiple voxels.

[0008] Based on the color information and the transparency information, an augmented image corresponding to the target viewpoint is generated;

[0009] The image augmentation model is trained based on sample images from multiple perspectives, and the sample images from these multiple perspectives all depict the same subject.

[0010] According to an image augmentation method provided by the present invention, the image augmentation model is trained based on the following steps:

[0011] Based on the sample images from the multiple viewpoints, the shooting direction information and voxel position information of the sample images corresponding to each viewpoint are determined, and the voxel position information includes the spatial position of multiple voxels.

[0012] The shooting direction information of each sample image and the voxel position information of each sample image are input into the image augmentation model so that the image augmentation model represents the three-dimensional volume model based on the shooting direction information of each sample image and the voxel position information of each sample image;

[0013] Determine the current viewpoint of the current training round, determine the sample image shooting direction information corresponding to the current viewpoint from the sample image shooting direction information of each sample image, and determine the sample voxel position information corresponding to the current viewpoint from the sample voxel position information of each sample voxel.

[0014] Based on the shooting direction information of the sample image corresponding to the current viewpoint and the voxel position information of the sample image corresponding to the current viewpoint, the image augmentation model is trained to reconstruct the three-dimensional volume model.

[0015] Return to the step of determining the current perspective of the current training round, until the current training round is the last training round.

[0016] According to an image augmentation method provided by the present invention, the step of training the image augmentation model based on the shooting direction information of the sample image corresponding to the current viewpoint and the voxel position information of the sample image corresponding to the current viewpoint includes:

[0017] Based on the shooting direction information of the sample image corresponding to the current viewpoint and the position information of the sample voxels corresponding to the current viewpoint, voxel sampling is performed on the three-dimensional volume model to obtain the sample color information and sample transparency information of the multiple sample voxels corresponding to the current viewpoint.

[0018] Based on the color information and transparency information of each sample, an augmented image of the sample corresponding to the current viewpoint is generated;

[0019] The image augmentation model is trained based on the image difference between the sample image from the current viewpoint and the augmented sample image.

[0020] According to an image augmentation method provided by the present invention, the image difference degree is determined based on the color difference degree of multiple pixels;

[0021] The color difference of any pixel is determined based on the difference between a first color value and a second color value, wherein the first color value is the color value of any pixel in the sample image of the current viewpoint, and the second color value is the color value of any pixel in the sample augmented image.

[0022] According to an image augmentation method provided by the present invention, the step of inputting the shooting direction information of each sample image and the voxel position information of each sample image into the image augmentation model, so that the image augmentation model represents the three-dimensional volume model based on the shooting direction information of each sample image and the voxel position information of each sample image, includes:

[0023] The shooting direction information of each sample image and the position information of each sample voxel are encoded to obtain multiple position encoding vectors;

[0024] The plurality of location encoding vectors are input into the image augmentation model so that the image augmentation model represents the three-dimensional volume model based on the plurality of location encoding vectors.

[0025] According to an image augmentation method provided by the present invention, the step of performing voxel sampling on a three-dimensional volume model represented by an image augmentation model based on the image shooting direction information and the voxel position information to obtain color information and transparency information of the plurality of voxels includes:

[0026] Based on the image shooting direction information and the voxel position information, voxel sampling is performed on the three-dimensional volume model corresponding to the current frame to obtain the color information and transparency information of the multiple voxels;

[0027] The process of generating an augmented image corresponding to the target viewpoint based on the color information and the transparency information further includes:

[0028] Based on the augmented images in multiple frames, the augmented video is determined;

[0029] The image augmentation model is trained based on sample videos from multiple perspectives, and each sample video from any perspective includes multiple frames of sample images.

[0030] According to an image augmentation method provided by the present invention, the image augmentation model is trained based on the following steps:

[0031] Based on the video feature extraction model, feature extraction is performed on the sample videos from the multiple perspectives to obtain the video features corresponding to the multiple perspectives;

[0032] Based on the image augmentation model, each of the video features is encoded to obtain the first feature vector corresponding to the multiple viewpoints;

[0033] Based on the feature encoding model, each of the first feature vectors is encoded to obtain the second feature vectors corresponding to the multiple viewpoints;

[0034] The image augmentation model, the video feature extraction model, and the feature encoding model are trained based on the similarity of each of the second feature vectors across different viewpoints.

[0035] The present invention also provides an image augmentation device, comprising:

[0036] The determination module is used to determine the image shooting direction information and voxel position information corresponding to the target viewpoint, wherein the voxel position information includes the spatial position of multiple voxels;

[0037] The sampling module is used to perform voxel sampling on the three-dimensional volume model represented by the image augmentation model based on the image shooting direction information and the voxel position information, so as to obtain the color information and transparency information of the multiple voxels.

[0038] The generation module is used to generate an augmented image corresponding to the target viewpoint based on the color information and the transparency information.

[0039] The image augmentation model is trained based on sample images from multiple perspectives, and the sample images from these multiple perspectives all depict the same subject.

[0040] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the image augmentation method as described above.

[0041] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the image augmentation method as described above.

[0042] The image augmentation method, apparatus, electronic device, and storage medium provided by this invention can train an image augmentation model based on sample images from multiple viewpoints. This model represents a three-dimensional volume model, i.e., it reconstructs the three-dimensional volume model corresponding to the sample images from multiple viewpoints through inverse rendering. Therefore, by only determining the image shooting direction information and voxel position information corresponding to the target viewpoint, voxel sampling can be performed on the three-dimensional volume model represented by the image augmentation model based on the image shooting direction information and voxel position information to obtain the color and transparency information of multiple voxels. Then, based on each color and transparency information, an augmented image corresponding to the target viewpoint is generated, thereby augmenting the image to a new viewpoint. This effectively augments the sample data in terms of viewpoint, achieving efficient image augmentation. When training a model based on viewpoint-augmented sample data, this improves the model training effect and enhances the robustness of the machine learning model. Attached Figure Description

[0043] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0044] Figure 1 This is one of the flowcharts illustrating the image augmentation method provided by the present invention;

[0045] Figure 2 A second schematic flowchart of the image augmentation method provided by the present invention;

[0046] Figure 3 The third schematic flowchart of the image augmentation method provided by the present invention;

[0047] Figure 4 The fourth schematic flowchart of the image augmentation method provided by the present invention;

[0048] Figure 5 A schematic diagram of the structure of the image augmentation device provided by the present invention;

[0049] Figure 6 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0050] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0051] With the rapid development of artificial intelligence technology, the application scenarios of various machine learning models are becoming increasingly widespread. Some of these models require massive amounts of image or video data as sample data for training, testing, and validation. Therefore, image augmentation is necessary for both image and video data.

[0052] For example, to improve the accuracy of sign language recognition models, massive amounts of sign language data are needed as sample data for training, testing, and validation. For application scenarios of sign language recognition, considering that the images captured by sign language recognition devices of the subject can be from various perspectives, the sample data also needs to include images from multiple perspectives. However, the traditional method involves investing a significant amount of manpower and time to collect sign language data from various perspectives, which is extremely time-consuming and labor-intensive.

[0053] Currently, image augmentation can be performed using geometric transformations such as horizontal flipping, random scaling, center cropping, and translation. It can also be achieved through pixel transformations such as noise reduction, color transformation, brightness adjustment, contrast adjustment, and saturation transformation. Furthermore, it can be done through content editing methods such as foreground segmentation, background replacement, random erasure, and cross-image manipulation. However, existing technologies cannot augment images from new perspectives, resulting in ineffective augmentation of sample data in terms of viewpoint. This leads to poor model training performance and reduced robustness of machine learning models.

[0054] To facilitate understanding, we will use the application scenarios of sign language data augmentation as an example. Geometric transformation methods directly perform geometric transformations on the original image, which can simulate changes in the geometric shape of sign language movements to some extent. However, they cannot simulate sign language movements in new perspective scenarios. Pixel transformation methods directly change pixel values ​​on the original image, which can simulate changes in the color appearance of sign language movements to some extent. However, they cannot simulate changes in geometric shape and image content, and it is difficult to control the degree of pixel transformation, thus failing to generate realistic sign language data. Most importantly, they cannot simulate sign language movements in new perspective scenarios. Content editing methods extract salient foreground sign language movements from the original image, eliminate background interference, and combine them to generate new sign language data. This can simulate changes in the content of sign language images or sequences of sign language images to some extent. However, they cannot simulate changes in geometric shape and color appearance, and images or videos augmented using content editing methods often have significant artifacts, resulting in uneven content or discontinuous sequences. Most importantly, they cannot simulate sign language movements in new perspective scenarios. Based on the above, the existing sign language data augmentation can alleviate the insufficiency of sign language data to some extent, but it still cannot simulate sign language data in new perspective scenarios.

[0055] To address the above problems, the present invention proposes the following embodiments. Figure 1 This is one of the flowcharts illustrating the image augmentation method provided by the present invention, such as... Figure 1 As shown, the image augmentation method includes:

[0056] Step 110: Determine the image shooting direction information and voxel position information corresponding to the target viewpoint. The voxel position information includes the spatial positions of multiple voxels.

[0057] Here, the target viewpoint is the viewpoint that needs to be augmented, that is, the augmented image or video that needs to be obtained from the target viewpoint.

[0058] Here, the image capture direction information refers to the acquisition direction corresponding to the augmented image or augmented video of the target viewpoint. This image capture direction information can be used to characterize the viewing direction of the capturing device, that is, to characterize the target viewpoint. In one embodiment, the image capture direction information includes a horizontal angle and a vertical angle.

[0059] Here, the voxel position information refers to the voxel position corresponding to the augmented image or augmented video of the target viewpoint. This spatial position is used to characterize the position of each voxel. In one embodiment, the spatial position includes the X-axis coordinate position, the Y-axis coordinate position, and the Z-axis coordinate position, that is, the spatial position of any voxel can be characterized by (x, y, z).

[0060] Step 120: Based on the image shooting direction information and the voxel position information, voxel sampling is performed on the three-dimensional volume model represented by the image augmentation model to obtain the color information and transparency information of the multiple voxels.

[0061] Here, the image augmentation model is a neural network model, thus performing image augmentation based on neural rendering. The specific structure of this image augmentation model can be set according to actual needs, such as MLP (Multi-Layer Perceptron), convolutional neural network, Transformer, etc. This image augmentation model is used to implicitly represent a 3D volume model. This 3D volume model is obtained through a training phase, which optimizes the image augmentation model.

[0062] Here, color information is used to characterize the color value of a voxel. In one embodiment, the color information includes RGB color, but it may also include color information of other color types.

[0063] Here, transparency information is used to characterize the transparency of voxels. This transparency information is also known as volumetric density.

[0064] Specifically, based on the spatial positions of multiple voxels, voxel information of the 3D volumetric model is sampled along the camera viewing direction ray indicated by the image capture direction information to obtain the color and transparency information of these voxels. In other words, based on the spatial positions of multiple voxels, voxel sampling is performed on multiple voxels in the 3D volumetric model along the projection line indicated by the image capture direction information to obtain the color and transparency information of these voxels.

[0065] In one embodiment, when augmentation is needed to obtain an augmented video, i.e., when the augmented video includes multiple frames of augmented images, the image augmentation model represents multiple three-dimensional volumetric models, with each frame corresponding to one three-dimensional volumetric model. Based on the image shooting direction information and voxel position information, voxel sampling is performed on the multiple three-dimensional volumetric models represented by the image augmentation model to obtain the color and transparency information of multiple voxels corresponding to multiple frames. In other words, the image augmentation model represents a sequence of three-dimensional volumetric models, which includes multiple three-dimensional volumetric models.

[0066] Step 130: Based on the color information and the transparency information, generate an augmented image corresponding to the target viewpoint.

[0067] Specifically, based on the color information and transparency information of multiple voxels, an augmented image corresponding to the target viewpoint is rendered, that is, an augmented image of the new viewpoint scene is generated. Then, based on multiple consecutive augmented images, an augmented video of the new viewpoint scene can be generated, such as sign language action sequence data, that is, sign language action video.

[0068] In one embodiment, an augmented image corresponding to the target viewpoint is rendered using the RayMatching method based on the color information and transparency information of multiple voxels.

[0069] The image augmentation model is trained based on sample images from multiple perspectives, and the sample images from these multiple perspectives all depict the same subject.

[0070] Here, the sample images can be acquired using inexpensive RGB imaging devices, such as CMOS (Complementary Metal Oxide Semiconductor) cameras, thus enabling multi-view sample image acquisition without relying on specific depth cameras or 3D data acquisition equipment. Of course, other imaging devices can also be used, and this embodiment of the invention does not specifically limit the acquisition of such images.

[0071] It should be noted that the sample images from multiple perspectives are of the same subject, that is, sample images of the same subject from different perspectives are collected. This ensures that a highly accurate image augmentation model is trained based on sample images from multiple perspectives, and ultimately a highly accurate augmented image is obtained.

[0072] In one embodiment, when it is necessary to augment the video to obtain an augmented video, that is, when the augmented video includes multiple frames of augmented images, the image augmentation model is trained based on sample videos from multiple perspectives, and the sample video from any perspective includes multiple frames of sample images.

[0073] For example, if the augmented video is a sign language action video, then the sample video is also a sign language action video. This sign language action video includes multiple frames of sign language images depicting continuous sign language actions; that is, the sample images are sign language images. Based on this, the sample images represent the posture information of the sign language actions, which may include, but is not limited to, at least one of the following: hand shape, hand position, limb shape, limb position, facial expression, facial position, etc. Therefore, the information represented by each frame of the sample images may include, but is not limited to, hand movements, facial movements, limb movements, etc. Based on this, sample images from multiple perspectives correspond to the same sign language action; that is, sample videos from multiple perspectives correspond to the same sign language action sequence.

[0074] It is understandable that generating an augmented image corresponding to the target viewpoint, thereby augmenting the image to a new viewpoint, effectively augments the sample data in terms of viewpoint. This can improve the training effect of the sign language recognition model when training the model based on the viewpoint-augmented sample data, and thus improve the robustness of the sign language recognition model.

[0075] The image augmentation method provided in this invention can train an image augmentation model based on sample images from multiple viewpoints. This model represents a three-dimensional volume model, i.e., it reconstructs the three-dimensional volume model corresponding to the sample images from multiple viewpoints through inverse rendering. Therefore, by only determining the image shooting direction information and voxel position information corresponding to the target viewpoint, voxel sampling can be performed on the three-dimensional volume model represented by the image augmentation model based on the image shooting direction information and voxel position information to obtain the color and transparency information of multiple voxels. Then, based on each color and transparency information, an augmented image corresponding to the target viewpoint is generated, thereby augmenting the image to obtain an augmented image under the new viewpoint. This effectively augments the sample data in terms of viewpoint, achieving efficient image augmentation. When training the model based on the sample data augmented in viewpoint, the model training effect can be improved, thereby improving the robustness of the machine learning model.

[0076] Based on the above embodiments, Figure 2 This is a second schematic flowchart of the image augmentation method provided by the present invention, as shown below. Figure 2 As shown, the image augmentation model is trained based on the following steps:

[0077] Step 210: Based on the sample images from the multiple viewpoints, determine the sample image shooting direction information and sample voxel position information corresponding to each viewpoint. The sample voxel position information includes the spatial position of multiple sample voxels.

[0078] Here, the sample image shooting direction information refers to the acquisition direction of the sample image or sample video corresponding to any viewpoint. This sample image shooting direction information can be used to characterize the viewing direction of the shooting device. In one embodiment, the sample image shooting direction information includes a horizontal angle and a vertical angle. One viewpoint corresponds to one sample image shooting direction information.

[0079] Here, the sample voxel position information refers to the voxel position corresponding to the sample image or sample video at any given viewpoint. This spatial position is used to characterize the position of each sample voxel. In one embodiment, the spatial position includes X-axis coordinates, Y-axis coordinates, and Z-axis coordinates; that is, (x, y, z) can be used to characterize the spatial position of any sample voxel. One viewpoint corresponds to one sample voxel position information.

[0080] Specifically, the physical and geometric properties of sample images from multiple perspectives are calculated, and the shooting direction information and sample voxel position information of the sample images corresponding to each perspective in the scene are estimated based on these physical and geometric properties.

[0081] In one embodiment, when augmentation is required to obtain augmented video, the image augmentation model is trained based on sample videos from multiple perspectives. Based on this, the physical and geometric properties of the sample videos from multiple perspectives are calculated, and the shooting direction information and sample voxel position information of the sample images corresponding to each perspective in the scene are estimated based on the physical and geometric properties.

[0082] In one specific embodiment, the COLMAP (motion structure and multi-view stereo computing software) tool is used to estimate the shooting direction information and voxel position information of the sample images corresponding to each viewpoint in the scene based on sample images from multiple viewpoints. Further, the COLMAP tool is used to estimate the shooting direction information and voxel position information of the sample images corresponding to each viewpoint in the scene based on sample videos from multiple viewpoints.

[0083] Step 220: Input the shooting direction information of each sample image and the voxel position information of each sample image into the image augmentation model, so that the image augmentation model represents the three-dimensional volume model based on the shooting direction information of each sample image and the voxel position information of each sample image.

[0084] Here, the image augmentation model is used to reconstruct a three-dimensional volume model based on the shooting direction information and voxel position information of each sample image. In other words, the image augmentation model is used to implicitly represent the three-dimensional volume model from two-dimensional sample images from multiple different perspectives.

[0085] In one embodiment, when augmentation is needed to obtain augmented video, the image augmentation model represents multiple three-dimensional volumetric models. That is, it represents multiple three-dimensional volumetric models based on the shooting direction information and voxel position information of each sample image, thus representing a sequence of three-dimensional volumetric models. In other words, the image augmentation model is used to reconstruct the sequence of three-dimensional volumetric models, for example, to reconstruct the sequence of three-dimensional volumetric models corresponding to a sign language movement sequence. This means it implicitly represents the sequence of three-dimensional volumetric models of the sign language movement sequence from two-dimensional sample images from multiple different perspectives.

[0086] Step 230: Determine the current viewpoint of the current training round, determine the sample image shooting direction information corresponding to the current viewpoint from the sample image shooting direction information of each sample image, and determine the sample voxel position information corresponding to the current viewpoint from the sample voxel position information of each sample voxel.

[0087] It should be noted that the number of training epochs for the image augmentation model is the same as the number of sample images from multiple viewpoints. That is, the model is first trained based on sample images from one viewpoint, and then trained again based on sample images from the next viewpoint. Furthermore, the model is first trained based on sample videos from one viewpoint, and then trained again based on sample videos from the next viewpoint.

[0088] Here, the current viewpoint can be one of the various viewpoints, as long as the current viewpoint is different in each training round.

[0089] Step 240: Based on the shooting direction information of the sample image corresponding to the current viewpoint and the position information of the sample voxel corresponding to the current viewpoint, train the image augmentation model to reconstruct the three-dimensional volume model.

[0090] It should be noted that training the image augmentation model allows for the continuous reconstruction and optimization of the 3D volumetric model. In other words, the image augmentation model learns scene representations based on sample images from multiple viewpoints to obtain the 3D volumetric model.

[0091] In one embodiment, when augmented video is needed, the image augmentation model is trained based on the shooting direction information of the sample images corresponding to the current viewpoint and the voxel position information of the sample images corresponding to the current viewpoint to reconstruct multiple three-dimensional volumetric models. In other words, the image augmentation model performs scene representation learning based on sample videos from multiple viewpoints to obtain a sequence of three-dimensional volumetric models. For example, based on sign language action videos from multiple viewpoints, scene representation learning of the sign language action sequence is performed to obtain a sequence of three-dimensional volumetric models.

[0092] Step 250: Return to the step of determining the current view of the current training round until the current training round is the last training round.

[0093] Specifically, for each viewpoint sample image, steps 230 and 240 are performed to continuously optimize the three-dimensional volume model until the training is completed for all viewpoint sample images.

[0094] The image augmentation method provided in this invention determines the shooting direction information and voxel position information of the sample images corresponding to each viewpoint based on sample images from multiple perspectives. This information is then input into an image augmentation model, enabling the model to represent a three-dimensional volume model based on the shooting direction and voxel position information. This allows the generation of augmented images corresponding to new viewpoints based on the three-dimensional volume model. Simultaneously, the image augmentation model is trained on sample images from each viewpoint to continuously reconstruct the three-dimensional volume model, improving its accuracy and consequently the accuracy of the augmented images. This effectively augments the sample data across viewpoints, achieving efficient image augmentation. Furthermore, when training the model based on viewpoint-augmented sample data, the model training effect can be further improved, thereby enhancing the robustness of the machine learning model.

[0095] Based on any of the above embodiments Figure 3 This is the third flowchart illustrating the image augmentation method provided by the present invention, as shown below. Figure 3 As shown, step 240 above includes:

[0096] Step 241: Based on the sampling direction information of the sample image corresponding to the current viewpoint and the sample voxel position information corresponding to the current viewpoint, voxel sampling is performed on the three-dimensional volume model to obtain the sample color information and sample transparency information of the multiple sample voxels corresponding to the current viewpoint.

[0097] Here, sample color information is used to characterize the color value of the sample voxels. In one embodiment, the sample color information includes RGB color, but it may also include color information of other color types.

[0098] Here, sample transparency information is used to characterize the transparency of the sample voxels. This sample transparency information is also called volumetric density.

[0099] Specifically, based on the spatial positions of multiple sample voxels corresponding to the current viewpoint, voxel information on the camera viewing direction ray indicated by the sample image capture direction information of the 3D volumetric model is sampled to obtain the sample color information and sample transparency information of these multiple sample voxels. In other words, based on the spatial positions of multiple sample voxels corresponding to the current viewpoint, voxel sampling is performed on multiple sample voxels in the 3D volumetric model along the projection line indicated by the sample image capture direction information to obtain the sample color information and sample transparency information of these multiple sample voxels.

[0100] In one embodiment, when augmentation is needed to obtain augmented video, the image augmentation model represents multiple three-dimensional volume models, with each frame corresponding to one three-dimensional volume model. Based on the shooting direction information of the sample image corresponding to the current viewpoint and the voxel position information of the sample image corresponding to the current viewpoint, voxel sampling is performed on the multiple three-dimensional volume models to obtain the sample color information and sample transparency information of the multiple sample voxels corresponding to the current viewpoint and multiple frames.

[0101] Step 242: Based on the color information and transparency information of each sample, generate an augmented image of the sample corresponding to the current viewpoint.

[0102] Specifically, based on the sample color information and sample transparency information of multiple sample voxels corresponding to the current viewpoint, a sample augmented image corresponding to the current viewpoint is rendered. Then, based on multiple consecutive sample augmented images, a sample augmented video of the current viewpoint scene can be generated, such as sign language action sequence data, i.e., sign language action video.

[0103] In one embodiment, an augmented image of the sample corresponding to the current viewpoint is rendered using the Ray Matching method, based on the sample color information and sample transparency information of multiple sample voxels corresponding to the current viewpoint.

[0104] Step 243: Train the image augmentation model based on the image difference between the sample image at the current viewpoint and the sample augmented image.

[0105] Here, image dissimilarity is used to characterize the degree of difference between the sample image and the augmented image. It can be determined based on the similarity between the sample image and the augmented image, i.e., the higher the similarity, the smaller the image dissimilarity. It can also be determined based on the mean square error of each pixel value between the sample image and the augmented image, i.e., the mean of the offset of each pixel value. Of course, it can also be determined in other ways, which will not be elaborated here.

[0106] In other words, the loss function of the image augmentation model includes a rendering loss function, the value of which is determined based on the image difference between the sample image and the augmented sample image. For example, the rendering loss function is shown below:

[0107] L1 = ||G1-G2||;

[0108] In the formula, ||G1-G2|| represents the image difference degree, G1 represents the sample image, and G2 represents the sample augmented image.

[0109] In one embodiment, when augmentation is required to obtain an augmented video, the image augmentation model is trained based on the video difference between the sample video and the augmented sample video from the current viewpoint. This video difference is determined based on multiple image differences, each of which is based on the difference between the target sample image in the sample video from the current viewpoint and the corresponding target sample augmented image in the augmented sample video. The correspondence between the target sample image and the target sample augmented image is that they have the same number of frames.

[0110] It is understandable that by performing steps 241, 242 and 243 above from multiple perspectives, the final augmented image can be made consistent with the input sample images from multiple perspectives in terms of attributes such as shape or appearance. That is, the three-dimensional volume model can be reconstructed based on the image differences corresponding to multiple perspectives.

[0111] The image augmentation method provided in this invention trains an image augmentation model based on the image difference between the sample image at the current viewpoint and the augmented sample image. This ensures that the final augmented image is consistent with the input sample image in terms of attributes such as shape or appearance. In other words, a three-dimensional volume model is reconstructed based on the image difference to further improve the accuracy of the three-dimensional volume model, thereby further improving the accuracy of the augmented image. This effectively augments the sample data at the viewpoint, achieving efficient image augmentation. When training the model based on the sample data augmented at the viewpoint, the model training effect can be further improved, thereby further enhancing the robustness of the machine learning model.

[0112] Based on any of the above embodiments, in this method, the image difference degree is determined based on the color difference degree of multiple pixels; the color difference degree of any pixel is determined based on the difference degree of a first color value and a second color value, wherein the first color value is the color value of any pixel in the sample image of the current viewpoint, and the second color value is the color value of any pixel in the sample augmented image.

[0113] Here, the number of pixels is determined based on the number of pixels in the sample image at the current viewpoint. The correspondence between the first color value and the second color value is that the pixel positions are the same. The first color value and the second color value can be RGB pixel values, or of course, other types of color values.

[0114] Here, color difference is used to characterize the degree of difference between the first color value and the second color value. It can be determined based on the similarity between the first color value and the second color value, that is, the higher the similarity, the smaller the color difference. Of course, it can also be determined in other ways, which will not be elaborated here.

[0115] In one embodiment, the mean square error of the color values ​​of each pixel in the sample image and the sample augmented image can be used to determine the mean value of the offset of the color value of each pixel.

[0116] The image augmentation method provided in this invention determines the image difference between the sample image and the augmented sample image based on the difference in color values ​​of their corresponding pixels. Therefore, based on the image difference between the sample image and the augmented sample image at the current viewpoint, the image augmentation model is trained to ensure that the final augmented image is consistent with the input sample image in terms of color, appearance, and other attributes. This involves reconstructing a three-dimensional volume model based on the image difference, further improving the accuracy of the three-dimensional volume model, and consequently, the accuracy of the augmented image. This effectively augments the sample data from the viewpoint, achieving efficient image augmentation. Furthermore, when training the model based on the viewpoint-augmented sample data, the model training effect can be further improved, thereby enhancing the robustness of the machine learning model.

[0117] Based on any of the above embodiments, in this method, step 220 includes:

[0118] The shooting direction information of each sample image and the position information of each sample voxel are encoded to obtain multiple position encoding vectors;

[0119] The plurality of location encoding vectors are input into the image augmentation model so that the image augmentation model represents the three-dimensional volume model based on the plurality of location encoding vectors.

[0120] In one embodiment, the shooting direction information and voxel position information of the sample image corresponding to any viewpoint are position encoded to obtain the position encoding vector corresponding to that viewpoint.

[0121] In another embodiment, the shooting direction information and voxel position information of the sample image corresponding to any viewpoint are respectively position encoded to obtain two position encoding vectors corresponding to that viewpoint.

[0122] In one specific embodiment, higher harmonic functions are used to perform position encoding on the image capture direction information and voxel position information of each sample image, resulting in multiple position encoding vectors. Of course, other trigonometric functions can also be used for position encoding, and this embodiment of the invention does not specifically limit this.

[0123] In one embodiment, when it is necessary to augment the video to obtain an augmented image, the image augmentation model represents multiple three-dimensional volume models, that is, multiple three-dimensional volume models are represented based on multiple position encoding vectors, that is, a sequence of three-dimensional volume models.

[0124] The image augmentation method provided in this invention determines the shooting direction information and voxel position information of the sample images corresponding to each viewpoint based on sample images from multiple perspectives. It then performs position encoding on the shooting direction information and voxel position information of each sample image to obtain multiple position encoding vectors. These multiple position encoding vectors are then input into an image augmentation model, enabling the model to represent a higher-frequency three-dimensional volume model based on these multiple position encoding vectors. This results in better rendering of locations with high-frequency color and appearance changes, further improving the accuracy of the three-dimensional volume model and consequently the accuracy of the augmented image. This effectively augments the sample data from different perspectives, achieving efficient image augmentation. Furthermore, when training a model based on sample data augmented from different perspectives, this method can further improve the model training effect and enhance the robustness of the machine learning model.

[0125] Based on any of the above embodiments, in this method, step 120 includes:

[0126] Based on the image shooting direction information and the voxel position information, voxel sampling is performed on the three-dimensional volume model corresponding to the current frame to obtain the color information and transparency information of the multiple voxels.

[0127] Considering that augmentation may be needed to obtain augmented video, i.e., when the augmented video includes multiple frames of augmented images, the image augmentation model represents multiple three-dimensional volume models, with each frame corresponding to a three-dimensional volume model. Therefore, based on the image shooting direction information and voxel position information, voxel sampling is performed on the three-dimensional volume model corresponding to the current frame to obtain the color information and transparency information of multiple voxels corresponding to the current frame.

[0128] Accordingly, after step 130 above, the method further includes

[0129] Based on the augmented images in multiple frames, the augmented video is determined.

[0130] Specifically, steps 120 and 130 are performed for each frame to obtain augmented images corresponding to multiple target viewpoints. Then, based on the frame times of the multiple augmented images, the multiple augmented images are aggregated into an augmented video.

[0131] The number of augmented images can be the same as the number of three-dimensional volume models represented by the image augmentation model; or the number of augmented images can be greater than the number of three-dimensional volume models represented by the image augmentation model, meaning that the three-dimensional volume models corresponding to multiple adjacent frames can be the same.

[0132] The image augmentation model is trained based on sample videos from multiple perspectives, and each sample video from any perspective includes multiple frames of sample images.

[0133] Here, the sample video can be acquired using an inexpensive RGB imaging device, such as a CMOS camera, thus enabling multi-view sample video acquisition without relying on a specific depth camera or 3D data acquisition equipment. Of course, it can also be acquired using other imaging devices; this embodiment of the invention does not specifically limit its acquisition.

[0134] It should be noted that the sample videos from multiple perspectives all depict the same subject; that is, sample videos of the same subject are collected from different angles. This ensures that a highly accurate image augmentation model is trained based on sample videos from multiple perspectives, ultimately resulting in a highly accurate augmented video. Furthermore, the sample videos from multiple perspectives correspond to the same action sequence, further ensuring that a highly accurate image augmentation model is trained based on sample videos from multiple perspectives, ultimately resulting in a highly accurate augmented video.

[0135] In one embodiment, the augmented video is a sign language action video, and the sample video is also a sign language action video. This sign language action video includes multiple frames of sign language images depicting continuous sign language actions; that is, the sample images are sign language images. Based on this, sample videos from multiple perspectives correspond to the same sign language action sequence.

[0136] It is understandable that generating augmented videos corresponding to the target viewpoint, thereby augmenting the data to obtain augmented videos from the new viewpoint, effectively augments the sample data in terms of viewpoint. This can improve the training effect of the sign language recognition model when training the model based on the augmented sample data, and thus improve the robustness of the sign language recognition model.

[0137] The image augmentation method provided in this invention can train an image augmentation model based on sample videos from multiple viewpoints. This model represents multiple three-dimensional volume models, i.e., it reconstructs the three-dimensional volume models corresponding to the sample videos from multiple viewpoints through inverse rendering. Therefore, by only determining the image shooting direction information and voxel position information corresponding to the target viewpoint, voxel sampling can be performed on the multiple three-dimensional volume models represented by the image augmentation model based on the image shooting direction information and voxel position information, thereby generating the augmented video corresponding to the target viewpoint. This results in an augmented video from a new viewpoint, effectively augmenting the sample data in terms of viewpoint, achieving efficient image augmentation. This improves the model training effect when training the model based on the sample data augmented in viewpoint, thereby enhancing the robustness of the machine learning model.

[0138] Based on any of the above embodiments, in this method, the image augmentation model is trained based on the following steps:

[0139] Based on the video feature extraction model, feature extraction is performed on the sample videos from the multiple perspectives to obtain the video features corresponding to the multiple perspectives;

[0140] Based on the image augmentation model, each of the video features is encoded to obtain the first feature vector corresponding to the multiple viewpoints;

[0141] Based on the feature encoding model, each of the first feature vectors is encoded to obtain the second feature vectors corresponding to the multiple viewpoints;

[0142] The image augmentation model, the video feature extraction model, and the feature encoding model are trained based on the similarity of each of the second feature vectors across different viewpoints.

[0143] Here, the video feature extraction model is used to extract sequence features from sample videos to obtain video features. The specific structure of this video feature extraction model can be set according to actual needs, for example, a CNN (Convolutional Neural Networks) network model.

[0144] Here, the image augmentation model is used to further encode the video features to obtain a first feature vector. This first feature vector is used to characterize the sequence features of the sample video.

[0145] Here, the feature encoding model is used to further encode the first feature vector to obtain a second feature vector. This second feature vector is used to characterize the sequence features of the sample video. The specific structure of this feature encoding model can be set according to actual needs, for example, an RNN (Recurrent Neural Network) model, to extract sequence features.

[0146] It should be noted that model training is performed based on the similarity of each second feature vector across different perspectives, thus achieving self-supervised training (self-supervised learning).

[0147] Here, any similarity is the similarity between the second feature vector corresponding to the first viewpoint and the second feature vector corresponding to the second viewpoint, where the first viewpoint and the second viewpoint are different viewpoints.

[0148] In other words, the model is trained using a sequence self-supervised loss function to constrain the consistency of action sequences from multiple perspectives. The loss value of this sequence self-supervised loss function is determined based on the similarity scores. For example, the sequence self-supervised loss function is shown below:

[0149]

[0150] In the formula, similarity(fi) i ` fj j `) represents the similarity between the second feature vector corresponding to the i-th viewpoint and the second feature vector corresponding to the j-th viewpoint, fi i ` f represents the second feature vector corresponding to the i-th viewpoint. j ` Let N represent the second feature vector corresponding to the j-th viewpoint, and N represent the total number of viewpoints.

[0151] In one embodiment, the image augmentation model can be trained and optimized together with the rendering loss function described above, based on the sequence self-supervised loss function.

[0152] In one embodiment, the sample video is a sign language action video. Based on the similarity of each second feature vector across different viewpoints, the image augmentation model is trained to ensure that the final augmented video has consistency in sign language action sequences across multiple viewpoints. In other words, the final augmented video from the new viewpoint has consistency in appearance across multiple views, and the video content and action sequences are coherent and consistent. This results in a high-fidelity sign language action video, further improving the accuracy of the augmented video. This effectively augments the sample data in terms of viewpoints, achieving efficient video augmentation. When training a sign language recognition model based on viewpoint-augmented sample data, this further improves the training effect of the sign language recognition model, thereby further enhancing the robustness of the sign language recognition model.

[0153] The image augmentation method provided in this invention further encodes sample videos corresponding to each viewpoint using a video feature extraction model, an image augmentation model, and a feature encoding model to obtain multiple second feature vectors. Based on the similarity of these second feature vectors across different viewpoints, the image augmentation model is trained to ensure that the final augmented video exhibits consistency in action sequences across multiple viewpoints. In other words, the final augmented video from the new viewpoint possesses consistency in appearance across multiple views, and the video content and action sequences exhibit coherence and consistency. This results in high-fidelity action videos, further improving the accuracy of the augmented videos and effectively augmenting sample data across viewpoints. This achieves efficient video augmentation, further enhancing the training effect of the model when training it using sample data augmented from viewpoints, thereby improving the robustness of the machine learning model.

[0154] To facilitate understanding of the above embodiments, a specific embodiment will be used as an example for illustration. Figure 4As shown, assuming augmentation is needed to obtain augmented videos, sample videos from multiple perspectives are required, including sample videos from the first, second, and third perspectives. Based on these sample videos, the shooting direction information and voxel position information of the sample images corresponding to each perspective are determined. Position encoding is performed on the shooting direction information and voxel position information of each sample image to obtain multiple position encoding vectors. These vectors are then input into the image augmentation model, enabling the model to represent multiple 3D volumetric models based on these vectors. Based on the shooting direction information and voxel position information of the sample images, voxel sampling is performed on each of the multiple 3D volumetric models to obtain sample color information and sample voxel position information. Based on the color and transparency information of each sample, this process generates augmented sample videos. Then, based on the difference between the sample videos and the augmented sample videos, the image augmentation model is trained. Simultaneously, based on a video feature extraction model, features are extracted from sample videos at multiple viewpoints to obtain video features corresponding to each viewpoint. These video features are then encoded using the image augmentation model to obtain first feature vectors corresponding to multiple viewpoints. Finally, based on a feature encoding model, each first feature vector is encoded to obtain second feature vectors corresponding to multiple viewpoints. The image augmentation model, video feature extraction model, and feature encoding model are trained based on the similarity of these second feature vectors across different viewpoints. After training the image augmentation model, the image shooting direction and voxel position information corresponding to the target viewpoint are determined. Based on these information, voxel sampling is performed on the 3D volume model represented by the image augmentation model to obtain color and transparency information for multiple voxels. Based on this color and transparency information, an augmented image or video corresponding to the target viewpoint is generated.

[0155] In practical applications, based on the above embodiments, neural rendering technology is used to generate new perspective image or video data from image or video data from multiple different viewpoints. This effectively alleviates the contradiction between high sample data acquisition costs, limited sample data volume, and the heavy reliance of machine learning model performance on data volume. For example, through 3D volumetric model reconstruction and neural rendering of sign language action sequences, high-fidelity sign language video data in new perspective scenes can be generated from sample videos from multiple different viewpoints. This enriches sign language scenarios and expands the number of sign language samples, thereby improving the robustness of deep learning-based sign language action recognition and translation models, promoting the construction of scenario-rich sign language databases, and advancing sign language teaching.

[0156] The image augmentation apparatus provided by the present invention will now be described. The image augmentation apparatus described below can be referred to in correspondence with the image augmentation method described above.

[0157] Figure 5This is a schematic diagram of the image augmentation device provided by the present invention, as shown below. Figure 5 As shown, the image augmentation device includes:

[0158] The determination module 510 is used to determine the image shooting direction information and voxel position information corresponding to the target viewpoint, wherein the voxel position information includes the spatial position of multiple voxels.

[0159] The sampling module 520 is used to perform voxel sampling on the three-dimensional volume model represented by the image augmentation model based on the image shooting direction information and the voxel position information, so as to obtain the color information and transparency information of the multiple voxels.

[0160] The generation module 530 is used to generate an augmented image corresponding to the target viewpoint based on the color information and the transparency information.

[0161] The image augmentation model is trained based on sample images from multiple perspectives, and the sample images from these multiple perspectives all depict the same subject.

[0162] The image augmentation device provided in this invention can train an image augmentation model based on sample images from multiple viewpoints. This model represents a three-dimensional volume model, i.e., it reconstructs the three-dimensional volume model corresponding to the sample images from multiple viewpoints through inverse rendering. Therefore, by only determining the image shooting direction information and voxel position information corresponding to the target viewpoint, voxel sampling can be performed on the three-dimensional volume model represented by the image augmentation model based on the image shooting direction information and voxel position information to obtain the color information and transparency information of multiple voxels. Then, based on each color information and each transparency information, an augmented image corresponding to the target viewpoint is generated, thereby augmenting the image to obtain an augmented image under the new viewpoint. This effectively augments the sample data in terms of viewpoint, achieving efficient image augmentation. When training the model based on the sample data augmented in viewpoint, the model training effect can be improved, thereby improving the robustness of the machine learning model.

[0163] Based on any of the above embodiments, the device further includes a model training module, which includes:

[0164] The first information determination unit is used to determine the shooting direction information and voxel position information of the sample images corresponding to each viewpoint based on the sample images from the multiple viewpoints. The voxel position information includes the spatial position of multiple voxels.

[0165] An information input unit is used to input the shooting direction information of each sample image and the position information of each sample voxel into the image augmentation model, so that the image augmentation model can characterize the three-dimensional volume model based on the shooting direction information of each sample image and the position information of each sample voxel;

[0166] The second information determination unit is used to determine the current viewpoint of the current training round, determine the sample image shooting direction information corresponding to the current viewpoint from the sample image shooting direction information, and determine the sample voxel position information corresponding to the current viewpoint from the sample voxel position information.

[0167] The model training unit is used to train the image augmentation model based on the shooting direction information of the sample image corresponding to the current viewpoint and the voxel position information of the sample image corresponding to the current viewpoint, so as to reconstruct the three-dimensional volume model.

[0168] The step return unit is used to return the step of determining the current perspective of the current training round until the current training round is the last training round.

[0169] Based on any of the above embodiments, the model training unit is further used for:

[0170] Based on the shooting direction information of the sample image corresponding to the current viewpoint and the position information of the sample voxels corresponding to the current viewpoint, voxel sampling is performed on the three-dimensional volume model to obtain the sample color information and sample transparency information of the multiple sample voxels corresponding to the current viewpoint.

[0171] Based on the color information and transparency information of each sample, an augmented image of the sample corresponding to the current viewpoint is generated;

[0172] The image augmentation model is trained based on the image difference between the sample image from the current viewpoint and the augmented sample image.

[0173] Based on any of the above embodiments, the image difference degree is determined based on the color difference degree of multiple pixels;

[0174] The color difference of any pixel is determined based on the difference between a first color value and a second color value, wherein the first color value is the color value of any pixel in the sample image of the current viewpoint, and the second color value is the color value of any pixel in the sample augmented image.

[0175] Based on any of the above embodiments, the information input unit is further configured to:

[0176] The shooting direction information of each sample image and the position information of each sample voxel are encoded to obtain multiple position encoding vectors;

[0177] The plurality of location encoding vectors are input into the image augmentation model so that the image augmentation model represents the three-dimensional volume model based on the plurality of location encoding vectors.

[0178] Based on any of the above embodiments, the sampling module 520 is further used for:

[0179] Based on the image shooting direction information and the voxel position information, voxel sampling is performed on the three-dimensional volume model corresponding to the current frame to obtain the color information and transparency information of the multiple voxels.

[0180] The device also includes a video determination module, which is used for:

[0181] Based on the augmented images in multiple frames, the augmented video is determined;

[0182] The image augmentation model is trained based on sample videos from multiple perspectives, and each sample video from any perspective includes multiple frames of sample images.

[0183] Based on any of the above embodiments, the device further includes a model training module, which includes:

[0184] The feature extraction unit is used to extract features from the sample videos from the multiple perspectives based on the video feature extraction model, so as to obtain the video features corresponding to the multiple perspectives.

[0185] The feature encoding unit is used to encode each of the video features based on the image augmentation model to obtain the first feature vectors corresponding to the multiple viewpoints;

[0186] The feature encoding unit is also used to encode each of the first feature vectors based on the feature encoding model to obtain the second feature vectors corresponding to the multiple viewpoints;

[0187] The model training unit is also used to train the image augmentation model, the video feature extraction model, and the feature encoding model based on the similarity of each of the second feature vectors across different viewpoints.

[0188] Figure 6 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 6As shown, the electronic device may include a processor 610, a communication interface 620, a memory 630, and a communication bus 640, wherein the processor 610, the communication interface 620, and the memory 630 communicate with each other via the communication bus 640. The processor 610 can call logical instructions in the memory 630 to execute an image augmentation method. This method includes: determining image shooting direction information and voxel position information corresponding to the target viewpoint, wherein the voxel position information includes the spatial positions of multiple voxels; performing voxel sampling on a three-dimensional volume model represented by an image augmentation model based on the image shooting direction information and the voxel position information to obtain color information and transparency information of the multiple voxels; and generating an augmented image corresponding to the target viewpoint based on each of the color information and each of the transparency information; wherein the image augmentation model is trained based on sample images from multiple viewpoints, and the sample images from the multiple viewpoints depict the same shooting object.

[0189] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0190] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the image augmentation method provided by the above methods. The method includes: determining image shooting direction information and voxel position information corresponding to a target viewpoint, wherein the voxel position information includes the spatial positions of multiple voxels; performing voxel sampling on a three-dimensional volume model represented by an image augmentation model based on the image shooting direction information and the voxel position information to obtain color information and transparency information of the multiple voxels; and generating an augmented image corresponding to the target viewpoint based on the color information and the transparency information; wherein the image augmentation model is trained based on sample images from multiple viewpoints, and the sample images from the multiple viewpoints have the same shooting object.

[0191] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the image augmentation method provided by the above methods. The method includes: determining image shooting direction information and voxel position information corresponding to a target viewpoint, wherein the voxel position information includes the spatial positions of multiple voxels; performing voxel sampling on a three-dimensional volume model represented by an image augmentation model based on the image shooting direction information and the voxel position information to obtain color information and transparency information of the multiple voxels; and generating an augmented image corresponding to the target viewpoint based on each of the color information and each of the transparency information; wherein the image augmentation model is trained based on sample images from multiple viewpoints, and the sample images from the multiple viewpoints depict the same shooting object.

[0192] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0193] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0194] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An image augmentation method, characterized in that, include: Determine the image shooting direction information and voxel position information corresponding to the target viewpoint, wherein the voxel position information includes the spatial position of multiple voxels; Based on the image capture direction information and the voxel position information, voxel sampling is performed on the three-dimensional volume model represented by the image augmentation model to obtain the color and transparency information of the multiple voxels; the image augmentation model is used to implicitly represent the three-dimensional volume model. Based on the color information and the transparency information, an augmented image corresponding to the target viewpoint is generated; The image augmentation model is trained based on sample images from multiple perspectives, and the sample images from multiple perspectives all depict the same subject. The image augmentation model is trained based on the following steps: Based on the sample images from the multiple viewpoints, the shooting direction information and voxel position information of the sample images corresponding to each viewpoint are determined, and the voxel position information includes the spatial position of multiple voxels. The shooting direction information of each sample image and the position information of each sample voxel are input into the image augmentation model so that the image augmentation model represents the three-dimensional volume model based on the shooting direction information of each sample image and the position information of each sample voxel. Determine the current viewpoint of the current training round, determine the sample image shooting direction information corresponding to the current viewpoint from the sample image shooting direction information of each sample image, and determine the sample voxel position information corresponding to the current viewpoint from the sample voxel position information of each sample voxel. Based on the shooting direction information of the sample image corresponding to the current viewpoint and the voxel position information of the sample image corresponding to the current viewpoint, the image augmentation model is trained to reconstruct the three-dimensional volume model. Return to the step of determining the current perspective of the current training round, until the current training round is the last training round; The step of inputting the shooting direction information of each sample image and the voxel position information of each sample image into the image augmentation model, so that the image augmentation model represents the three-dimensional volume model based on the shooting direction information of each sample image and the voxel position information of each sample image, includes: Using higher harmonic functions, position encoding is performed on the shooting direction information of each sample image and the position information of each sample voxel to obtain multiple position encoding vectors; The plurality of location encoding vectors are input into the image augmentation model so that the image augmentation model represents the three-dimensional volume model based on the plurality of location encoding vectors.

2. The image augmentation method according to claim 1, characterized in that, The step of training the image augmentation model based on the shooting direction information of the sample image corresponding to the current viewpoint and the voxel position information of the sample image corresponding to the current viewpoint includes: Based on the shooting direction information of the sample image corresponding to the current viewpoint and the position information of the sample voxels corresponding to the current viewpoint, voxel sampling is performed on the three-dimensional volume model to obtain the sample color information and sample transparency information of the multiple sample voxels corresponding to the current viewpoint. Based on the color information and transparency information of each sample, an augmented image of the sample corresponding to the current viewpoint is generated; The image augmentation model is trained based on the image difference between the sample image from the current viewpoint and the augmented sample image.

3. The image augmentation method according to claim 2, characterized in that, The image difference is determined based on the color difference of multiple pixels; The color difference of any pixel is determined based on the difference between a first color value and a second color value, wherein the first color value is the color value of any pixel in the sample image of the current viewpoint, and the second color value is the color value of any pixel in the sample augmented image.

4. The image augmentation method according to any one of claims 1 to 3, characterized in that, Based on the image capture direction information and the voxel position information, voxel sampling is performed on the three-dimensional volume model represented by the image augmentation model to obtain the color and transparency information of the multiple voxels, including: Based on the image shooting direction information and the voxel position information, voxel sampling is performed on the three-dimensional volume model corresponding to the current frame to obtain the color information and transparency information of the multiple voxels; The process of generating an augmented image corresponding to the target viewpoint based on the color information and the transparency information further includes: Based on the augmented images in multiple frames, the augmented video is determined; The image augmentation model is trained based on sample videos from multiple perspectives, and each sample video from any perspective includes multiple frames of sample images.

5. The image augmentation method according to claim 4, characterized in that, The image augmentation model is trained based on the following steps: Based on the video feature extraction model, feature extraction is performed on the sample videos from the multiple perspectives to obtain the video features corresponding to the multiple perspectives; Based on the image augmentation model, each of the video features is encoded to obtain the first feature vector corresponding to the multiple viewpoints; Based on the feature encoding model, each of the first feature vectors is encoded to obtain the second feature vectors corresponding to the multiple viewpoints; The image augmentation model, the video feature extraction model, and the feature encoding model are trained based on the similarity of each of the second feature vectors across different viewpoints.

6. An image augmentation device, characterized in that, include: The determination module is used to determine the image shooting direction information and voxel position information corresponding to the target viewpoint, wherein the voxel position information includes the spatial position of multiple voxels; The sampling module is used to perform voxel sampling on the three-dimensional volume model represented by the image augmentation model based on the image shooting direction information and the voxel position information, so as to obtain the color information and transparency information of the multiple voxels. The image augmentation model is used to implicitly represent the three-dimensional volume model; The generation module is used to generate an augmented image corresponding to the target viewpoint based on the color information and the transparency information. The image augmentation model is trained based on sample images from multiple perspectives, and the sample images from multiple perspectives all depict the same subject. The device further includes a model training module, which includes: The first information determination unit is used to determine the shooting direction information and voxel position information of the sample images corresponding to each viewpoint based on the sample images from the multiple viewpoints. The voxel position information includes the spatial position of multiple voxels. An information input unit is used to input the shooting direction information of each sample image and the voxel position information of each sample image into the image augmentation model, so that the image augmentation model can characterize the three-dimensional volume model based on the shooting direction information of each sample image and the voxel position information of each sample image; The second information determination unit is used to determine the current viewpoint of the current training round, determine the sample image shooting direction information corresponding to the current viewpoint from the sample image shooting direction information of each sample image, and determine the sample voxel position information corresponding to the current viewpoint from the sample voxel position information of each sample voxel. The model training unit is used to train the image augmentation model based on the shooting direction information of the sample image corresponding to the current viewpoint and the voxel position information of the sample image corresponding to the current viewpoint, so as to reconstruct the three-dimensional volume model. The step return unit is used to return the step of determining the current perspective of the current training round until the current training round is the last training round; The step of inputting the shooting direction information of each sample image and the voxel position information of each sample image into the image augmentation model, so that the image augmentation model represents the three-dimensional volume model based on the shooting direction information of each sample image and the voxel position information of each sample image, includes: Using higher harmonic functions, position encoding is performed on the shooting direction information of each sample image and the position information of each sample voxel to obtain multiple position encoding vectors; The plurality of location encoding vectors are input into the image augmentation model so that the image augmentation model represents the three-dimensional volume model based on the plurality of location encoding vectors.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the image augmentation method as described in any one of claims 1 to 5.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the image augmentation method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • View angle image generation method and device, electronic equipment and storage medium

    CN114549731A