Image generation method, apparatus, device, and medium

By acquiring user and request information and adjusting the image generation model based on a preset material library, visual training images that meet user preferences are generated. This solves the problems of low efficiency and monotonous content in visual training image material generation, thereby improving user experience and training effectiveness.

CN118037893BActive Publication Date: 2026-02-17SHENZHEN QINGNIAORU ELECTRONICS CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410207052.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-02-26
Publication Date
2026-02-17
Estimated Expiration
2044-02-26

AI Technical Summary

Technical Problem

The current visual training image material generation efficiency is low and the material is insufficient, resulting in monotonous training content, poor user experience, and negatively impacting training effectiveness.

Method used

By acquiring user personal information and request information, and determining the background image set and 3D object image set based on a preset material library, the image generation model is fine-tuned to generate visual training images that meet user preferences.

Benefits of technology

It improves the efficiency and richness of visual training image generation, making the images more in line with user preferences and enhancing training effectiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118037893B_ABST
    Figure CN118037893B_ABST
Patent Text Reader

Abstract

The application discloses an image generation method, device, equipment and medium, wherein the method comprises: obtaining user personal information and user request information, and determining a background image set and a three-dimensional object image set in a preset material library based on the personal information; fine-tuning a preset image generation model according to the background image set and the three-dimensional object image set to obtain a target image generation model; and generating a visual training image according to the user request information, the target image generation model and preset motion parameters. The image for visual training can be generated according to the user's preference, thereby avoiding the problem of poor visual training effect caused by the single and repeated image for visual training.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of visual training, and in particular to an image generation method, device, equipment and medium. BACKGROUND

[0002] Visual training is to play a three-dimensional video containing moving objects, so that the user's eyes watch two different perspective pictures of a three-dimensional scene with a certain parallax. The eyeballs of the two eyes constantly move together and apart with the changes of the three-dimensional dynamic image, so as to achieve the effect of improving various visual abilities and preventing myopia.

[0003] Visual training often needs to be carried out for a long time for months or even years, and the time for daily visual training needs to reach twenty minutes. Therefore, a large number of different image video training image materials need to be played in the process of visual training. In the prior art, the training image materials are manually designed and produced by special staff, which has high time and fund costs and low production efficiency. As a result, there is a lack of training materials in the training material library, and the playing content is single and repetitive. At the same time, the users of visual training are mainly minors and / or children. In the process of visual training, the training image materials are repetitive and single, and the users of visual training are prone to have a tired and repulsive psychology in the mind, which causes a poor experience of visual training and further affects the effect of visual training. SUMMARY

[0004] Therefore, it is necessary to propose an image generation method, device, equipment and medium in view of the technical problem that the production efficiency of training image materials is low, the training image materials are insufficient, the training image materials are single, the experience of visual training is poor, and the effect of visual training is not good in the prior art.

[0005] In a first aspect, an image generation method is provided, and the method comprises:

[0006] obtaining target personal information of a user and user request information, and determining a background image set and a three-dimensional object image set in a preset material library based on the target personal information;

[0007] fine-tuning a preset image generation model according to the background image set and the three-dimensional object image set, to obtain a target image generation model;

[0008] generating a visual training image according to the user request information, the target image generation model and a preset motion parameter.

[0009] Optionally, before the step of determining the background image set and the three-dimensional object image set in the preset material library based on the target personal information, the method further comprises:

[0010] statistical pre-selected user for each background image in the preset material library of the first score data, and according to the pre-selected personal information of the pre-selected user and the first score data to build the first corresponding rule, the first corresponding rule is used to determine the background image set in the preset material library based on the target personal information;

[0011] The second score data of the pre-selected user for each three-dimensional object image in the preset material library is recorded, and the second corresponding rule is constructed according to the pre-selected personal information and the second score data, and the second corresponding rule is used to determine the three-dimensional object image set in the preset material library based on the target personal information.

[0012] Optionally, the step of determining the background image set and the three-dimensional object image set in the preset material library based on the target personal information comprises:

[0013] The similarity data of the target personal information and the pre-selected personal information is calculated to obtain a similarity data set, and the similarity data with the highest similarity value in the similarity data set is selected as the target similarity data;

[0014] The pre-selected personal information corresponding to the target similarity data is used as the similar personal information corresponding to the target personal information;

[0015] According to the first corresponding rule, the first target score data corresponding to the similar personal information is obtained in the first score data, the background image corresponding to the first target score data is used as the target background image, and the background image set is obtained;

[0016] According to the second corresponding rule, the second target score data corresponding to the similar personal information is obtained in the second score data, the three-dimensional object image corresponding to the second target score data is used as the target three-dimensional object image, and the three-dimensional object image set is obtained.

[0017] Optionally, the step of fine-tuning the preset image generation model according to the background image set and the three-dimensional object image set to obtain the target image generation model comprises:

[0018] According to the background image set, the first image generation model in the preset image generation model is trained to obtain a first target image generation model, the first image generation model comprises a two-dimensional encoder copy, a first preset convolutional layer and a preset two-dimensional image generation model, the two-dimensional encoder copy is the same as the encoder in the preset two-dimensional image generation model, and the two-dimensional encoder copy is connected to the preset two-dimensional image generation model through the first preset convolutional layer;

[0019] According to the three-dimensional object image set, a second image generation model in the preset image generation model is trained to obtain a second target image generation model, the second image generation model comprising a three-dimensional encoder copy, a second preset convolutional layer and a preset three-dimensional image generation model, the three-dimensional encoder copy being the same as an encoder in the preset three-dimensional image generation model, and the three-dimensional encoder copy being connected to the preset three-dimensional image generation model through the second preset convolutional layer.

[0020] Based on the first target image generation model and the second target image generation model, a target image generation model is obtained.

[0021] Optionally, the step of generating a visual training image according to the user request information, the target image generation model and preset motion parameters comprises:

[0022] The user request information is input into a first target image generation model in the target image generation model to generate a two-dimensional image corresponding to the user request information to obtain a target background image.

[0023] The user request information is input into a second target image generation model in the target image generation model to generate a three-dimensional image corresponding to the user request information to obtain a target three-dimensional object image.

[0024] According to the preset motion parameters, the target background image and the target three-dimensional object image, a visual training image is generated.

[0025] Optionally, the step of generating a visual training image according to the preset motion parameters, the target background image and the target three-dimensional object image comprises:

[0026] Based on the preset motion parameters, relative position data of the target background image and the target three-dimensional object image is determined.

[0027] Based on the relative position data, the target background image and the target three-dimensional object image are fused to generate a visual training image.

[0028] Optionally, the step of generating a visual training image by fusing the target background image and the target three-dimensional object image based on the relative position data comprises:

[0029] Based on the relative position data, an occluded area of the target background image occluded by the target three-dimensional object image is determined.

[0030] The occluded area is replaced by the target three-dimensional object image to generate a visual training image.

[0031] In a second aspect, an image generation apparatus is provided, and the apparatus comprises:

[0032] a data collection module configured to collect target personal information of a user and user request information, and determine a background image set and a three-dimensional object image set in a preset material library based on the target personal information;

[0033] a model training module configured to fine-tune a preset image generation model based on the background image set and the three-dimensional object image set, to obtain a target image generation model;

[0034] an image generation module configured to generate a visual training image based on the user request information, the target image generation model, and preset motion parameters.

[0035] In a third aspect, a computer device is provided, which comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the steps of the image generation method when executing the computer program.

[0036] In a fourth aspect, a computer readable storage medium is provided, which stores a computer program, and the computer program implements the steps of the image generation method when executed by a processor.

[0037] The present application collects target personal information of a user and user request information, and determines a background image set and a three-dimensional object image set in a preset material library based on the target personal information; fine-tunes a preset image generation model based on the background image set and the three-dimensional object image set, to obtain a target image generation model; and generates a visual training image based on the user request information, the target image generation model, and preset motion parameters. By determining a background image set and a three-dimensional object image set in a preset material library based on the target personal information, the preferences of the user are predicted, so that the generated visual training image is more in line with the preferences or habits of the user. Meanwhile, the preset image generation model is used to generate the visual training image, which improves the generation efficiency of the visual training image compared to the process of manually producing the visual training image, and thus makes the visual training image more abundant, solves the problem of insufficient visual training images, and takes into account that the images generated by the preset image generation model under unrestricted conditions may not be suitable for visual training. Therefore, the preset image generation model is fine-tuned based on the determination of the background image set and the three-dimensional object image set in the preset material library based on the target personal information, which ensures the generation efficiency of the visual training image while ensuring the quality of the visual training image. BRIEF DESCRIPTION OF DRAWINGS

[0038] In order to make the technical solutions in the embodiments of the present application or the prior art clearer, the accompanying drawings needed in the embodiments or prior art description will be briefly introduced. Obviously, the accompanying drawings in the following description only aim to some embodiments of the present application, and for those skilled in the art, other drawings can be obtained without creative effort based on these drawings.

[0039] Wherein:

[0040] Figure 1 An application environment diagram of the image generation method in an embodiment;

[0041] Figure 2 A flow chart of the image generation method in an embodiment;

[0042] Figure 3 A structural block diagram of the image generation device in an embodiment;

[0043] Figure 4 A structural block diagram of the computer device in an embodiment;

[0044] Figure 5 A structural block diagram of the computer device in another embodiment. DETAILED DESCRIPTION

[0045] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative effort belong to the scope of protection of the present application.

[0046] The image generation method provided by the embodiments of the present application can be applied in an application environment such as Figure 1 , wherein the mobile terminal 110 communicates with the server 120.

[0047] The server 120 is configured to: acquire target personal information of a user and user request information received by the mobile terminal 110, and determine a background image set and a three-dimensional object image set in a preset material library based on the target personal information; fine-tune a preset image generation model according to the background image set and the three-dimensional object image set, to obtain a target image generation model; generate a visual training image according to the user request information, the target image generation model and preset motion parameters, and feed back the visual training image to the mobile terminal 110, or feed back the visual training image to a specific display device. By determining the background image set and the three-dimensional object image set in the preset material library based on the target personal information, the preference of the user is predicted, so that the generated visual training image is more in line with the preference or habit of the user. At the same time, the preset image generation model is used to generate the visual training image, which improves the generation efficiency of the visual training image compared with the process of manually producing the visual training image, so that the visual training image is more abundant, and the problem of insufficient visual training images is solved. Considering that the images generated by the preset image generation model under unrestricted conditions may not be suitable for visual training, the preset image generation model is fine-tuned based on the target personal information in the preset material library to determine the background image set and the three-dimensional object image set, which ensures the generation efficiency of the visual training image while ensuring the quality of the visual training image.

[0048] To reduce the computing pressure of the server 120 and the data transmission pressure between the server 120 and the mobile terminal 110, the mobile terminal 110 can also be configured to: acquire target personal information of a user and user request information, and determine a background image set and a three-dimensional object image set in a preset material library based on the target personal information; fine-tune a preset image generation model according to the background image set and the three-dimensional object image set, to obtain a target image generation model; and generate a visual training image according to the user request information, the target image generation model and preset motion parameters.

[0049] The application will be described in detail below through specific embodiments.

[0050] Please refer to Figure 2 as shown, Figure 2 A flowchart of an image generation method provided by an embodiment of the application includes the following steps:

[0051] S101, acquire target personal information of a user and user request information, and determine a background image set and a three-dimensional object image set in a preset material library based on the target personal information;

[0052] Exemplarily, the target personal information is basic information of the user, such as age, gender, height, weight, hobby, region, and the like; the user request information is request information for visual training, such as request information for indicating to start visual training; and the preset material library is a database including different background images and different three-dimensional object images, the background images being, for example, lawn images, forest images, sky images, and the like, and the three-dimensional object images being images of objects moving in the background images, such as ball images, flower images, mushroom images, and the like.

[0053] Exemplarily, in the process of constructing the preset material library, images can be randomly selected from different data sources, such as different open source image libraries, to construct the preset material library, or the preset material library can be constructed in a targeted manner according to the age, gender, height, weight, hobby, region, and the like of different users, for example, according to the age of different users, animation images, car images, ball sport images, and the like are selected from different data sources.

[0054] Exemplarily, when the preset material library is constructed in a targeted manner according to the age, gender, height, weight, hobby, region, and the like of different users, it can be avoided that the preset material library is too large, and at the same time, the applicable range of the preset material library is ensured.

[0055] Exemplarily, in order to ensure the security of the personal information of the user, the user request information can be acquired first, and the user request information is identified, when it is identified that the user request information is used for visual training, the target personal information of the user is acquired, and thus the risk of leakage of the personal information of the user is reduced.

[0056] S102, fine-tuning a preset image generation model according to the set of background images and the set of three-dimensional object images to obtain a target image generation model;

[0057] Exemplarily, the preset image generation model for image generation includes a model that can generate a two-dimensional image, such as a text-to-image generation model (Stable Diffusion) based on a latent diffusion model (Latent Diffusion Models), the preset image generation model also includes a model that can generate a three-dimensional image, such as a three-dimensional image generation model (Stable 3D), and since the preset image generation model needs to output a single image, that is, the generated two-dimensional image is taken as a background image, and the generated three-dimensional image is taken as an object image in the background image, the preset image generation model also includes a fusion model for image fusion of the two-dimensional image and the three-dimensional image.

[0058] Exemplarily, the preset image generation model includes a model for generating a two-dimensional image and a model for generating a three-dimensional image, and the two-dimensional image and the three-dimensional image generated by the preset image generation model do not take into account the preferences of the user, i.e., are not targeted, so that the generated two-dimensional image and the generated three-dimensional model can not well arouse the interest of the user. Therefore, the preset image generation model is fine-tuned based on the target personal information to determine a background image set and a three-dimensional object image set in a preset material library, so that the preset image generation model learns the preferences of different users in the dimensions of image form, image content, and image style.

[0059] S103, generating a visual training image according to the user request information, the target image generation model, and preset motion parameters.

[0060] Exemplarily, a two-dimensional image and a three-dimensional image that meet the preferences of the user are generated according to the user request information and the target image generation model, and then the two-dimensional image and the three-dimensional image that meet the preferences of the user are fused according to preset motion parameters to obtain a visual training image.

[0061] Exemplarily, the preset motion parameters include a motion trajectory equation and motion parameters of an object image (three-dimensional image) in a background image (two-dimensional image), and then the position of the object image (three-dimensional image) in the background image (two-dimensional image) at a certain time can be determined according to the motion trajectory equation and the motion parameters, the image fusion is realized, and then the visual training image is generated.

[0062] The target personal information and the user request information of the user are obtained, and a background image set and a three-dimensional object image set are determined in a preset material library based on the target personal information. The preset image generation model is fine-tuned based on the background image set and the three-dimensional object image set to obtain a target image generation model. A visual training image is generated according to the user request information, the target image generation model, and preset motion parameters. The preferences of the user are predicted by determining the background image set and the three-dimensional object image set in the preset material library based on the target personal information, so that the generated visual training image is more in line with the preferences or habits of the user. At the same time, the preset image generation model is used to generate the visual training image, which improves the generation efficiency of the visual training image compared with the process of manually making the visual training image, and makes the visual training image more abundant, solves the problem of insufficient visual training images, and considers that the images generated by the preset image generation model under unrestricted conditions can not be suitable for visual training. Therefore, the preset image generation model is fine-tuned based on the target personal information to determine the background image set and the three-dimensional object image set in the preset material library, which ensures the generation efficiency of the visual training image while ensuring the quality of the visual training image.

[0063] In a possible implementation, before the step of determining the set of background images and the set of three-dimensional object images from the preset material library based on the target personal information, the method further includes:

[0064] collecting first score data of the preselected users on each background image in the preset material library, and constructing a first corresponding rule based on preselected personal information of the preselected users and the first score data, the first corresponding rule being used to determine the set of background images from the preset material library based on the target personal information;

[0065] collecting second score data of the preselected users on each three-dimensional object image in the preset material library, and constructing a second corresponding rule based on the preselected personal information and the second score data, the second corresponding rule being used to determine the set of three-dimensional object images from the preset material library based on the target personal information.

[0066] For example, different respondents (preselected users) are invited to manually score each background image in the preset material library, and the age, gender, height, weight, hobby, region, and other personal information (preselected personal information) of each respondent and the score (first score data) of each respondent on the degree of love for each image element are recorded. The higher the score given by a respondent to a certain image material, the higher the degree of personal subjective love for the material. The set of background images is determined from the preset material library according to a pre-set first screening percentage (first corresponding rule).

[0067] Similarly, different respondents are invited to manually score each three-dimensional object image in the preset material library, and the age, gender, height, weight, hobby, region, and other personal information (preselected personal information) of each respondent and the score (second score data) of each respondent on the degree of love for each image element are recorded. The higher the score given by a respondent to a certain image material, the higher the degree of personal subjective love for the material. The set of three-dimensional object images is determined from the preset material library according to a pre-set second screening percentage (second corresponding rule).

[0068] In a possible implementation, the step of determining the set of background images and the set of three-dimensional object images from the preset material library based on the target personal information includes:

[0069] calculating similarity data of the target personal information and the preselected personal information to obtain a set of similarity data, and selecting, from the set of similarity data, similarity data with the highest similarity value as target similarity data;

[0070] taking the preselected personal information corresponding to the target similarity data as similar personal information corresponding to the target personal information.

[0071] According to the first corresponding rule, first target score data corresponding to the similar personal information is obtained in the first score data, a background image corresponding to the first target score data is taken as a target background image, and the background image set is obtained;

[0072] According to the second corresponding rule, second target score data corresponding to the similar personal information is obtained in the second score data, a three-dimensional object image corresponding to the second target score data is taken as a target three-dimensional object image, and the three-dimensional object image set is obtained.

[0073] For example, N pieces of data such as age, gender, height, weight, hobby, and region are encoded into an N-dimensional vector, and each piece of data in the N pieces of data can be quantified into a value between 0 and 1. The target personal information vector can be represented as W = [w1w2w3...wn], and the preselected personal information of each respondent is also represented as a vector W' = [w'1w'2w'3...w'n] of the same format. Then, the target personal information vector is compared with the preselected personal information vector of each respondent for similarity, for example, the cosine distance is used to compare the target personal information vector and the preselected personal information vector, that is, W = [w1w2w3...wn] and W' = [w'1w'2w'3...w'n] are compared. The similarity result is N N N N That is, the dot product of the two vectors is divided by the product of the norms of the two vectors. The closer the similarity result is to 1, the closer the information of the two people is. The closer the similarity result is to 0, the more different the information of the two people is.

[0074] For example, the similarity data with the highest similarity value in the similarity data set can be selected as the target similarity data, or a preset number of similarity data with the highest similarity value in the similarity data set can be selected as the target similarity data. When a preset number of similarity data with the highest similarity value in the similarity data set is selected, the average score values of the preset number of respondents for the like degree of each background image are sorted from high to low to obtain the first score data, wherein the preset number of respondents are the preset number of respondents with the highest similarity value to the target personal information.

[0075] In a possible implementation, the step of fine-tuning the preset image generation model according to the background image set and the three-dimensional object image set to obtain the target image generation model comprises:

[0076] ​​​​According to the background image set, a first image generation model in the preset image generation model is trained to obtain a first target image generation model, the first image generation model includes a two-dimensional encoder copy, a first preset convolutional layer, and a preset two-dimensional image generation model, the two-dimensional encoder copy is the same as an encoder in the preset two-dimensional image generation model, and the two-dimensional encoder copy is connected to the preset two-dimensional image generation model through the first preset convolutional layer;

[0077] According to the three-dimensional object image set, a second image generation model in the preset image generation model is trained to obtain a second target image generation model, the second image generation model includes a three-dimensional encoder copy, a second preset convolutional layer, and a preset three-dimensional image generation model, the three-dimensional encoder copy is the same as an encoder in the preset three-dimensional image generation model, and the three-dimensional encoder copy is connected to the preset three-dimensional image generation model through the second preset convolutional layer;

[0078] Based on the first target image generation model and the second target image generation model, a target image generation model is obtained.

[0079] For example, for the encoder in the model for generating two-dimensional images (such as Stable Diffusion), a copy is made as a copy of the model for generating two-dimensional images (two-dimensional encoder copy), the copy is connected to the model for generating two-dimensional images through a convolutional layer (first preset convolutional layer) with an initial weight value of 0 to obtain a first image generation model, the background image set filtered according to the target personal information input by the user is used to train the first image generation model, and after weight parameter iterative optimization, a first target image generation model is obtained.

[0080] Similarly, for three-dimensional object images, a similar method is used to copy the encoder in the model for generating three-dimensional images (such as Stable 3D) as a copy of the model for generating three-dimensional images (three-dimensional encoder copy), the three-dimensional encoder copy is connected to the model for generating three-dimensional images through a convolutional layer (second preset convolutional layer) with an initial weight value of 0 to obtain a second image generation model, the three-dimensional object image set filtered according to the target personal information input by the user is used to train the second image generation model, and after weight parameter iterative optimization, a second target image generation model is obtained, wherein the data format processed by the model for generating three-dimensional images is a three-dimensional point cloud format.

[0081] In a possible implementation, the step of generating a visual training image according to the user request information, the target image generation model, and preset motion parameters includes:

[0082] input the user request information into a first target image generation model in the target image generation model, generate a two-dimensional image corresponding to the user request information, and obtain a target background image;

[0083] input the user request information into a second target image generation model in the target image generation model, generate a three-dimensional image corresponding to the user request information, and obtain a target three-dimensional object image;

[0084] generate a visual training image according to the preset motion parameter, the target background image, and the target three-dimensional object image.

[0085] For example, the user request information is input into the target image generation model, and the user request information is identified to distinguish first information in the user request information for a background image and second information in the user request information for an object image. For example, the user request information is "a ball rolling on a lawn", and it is identified that "a ball" is the second information for the object image and "a lawn" is the first information for the background image.

[0086] For example, since the target image generation model has learned user preference data, such as image style, a target three-dimensional object image corresponding to the "ball" that meets the user's preference and a target background image corresponding to the "lawn" that meets the user's preference are generated.

[0087] In a possible implementation, the step of generating a visual training image according to the preset motion parameter, the target background image, and the target three-dimensional object image includes:

[0088] determining relative position data of the target background image and the target three-dimensional object image based on the preset motion parameter;

[0089] fusing the target background image and the target three-dimensional object image based on the relative position data to generate a visual training image.

[0090] For example, since the visual training image is video data, i.e., video frame images at different time points need to be generated, the relative position data of the target background image and the target three-dimensional object image at different time points is determined based on the preset motion parameter, and the target background image and the target three-dimensional object image are fused based on the relative position data to generate video frame images at different time points. All video frame images are arranged in chronological order to obtain video data as a visual training image.

[0091] In one possible implementation, the step of fusing the target background image and the target 3D object image based on the relative position data to generate a visual training image includes:

[0092] Based on the relative position data, determine the occlusion area where the target background image is occluded by the target 3D object image;

[0093] The occluded area is replaced with the target 3D object image to generate a visual training image.

[0094] For example, if the background is considered as a fixed background, that is, the two-dimensional background image (target background image) in each frame of the video remains unchanged, the relative position data (x, y, x) of the three-dimensional object (target three-dimensional object image) at time point t can be determined using a given motion trajectory equation and parameters (preset motion parameters). t ,y t , z t ), where x t Let y be the displacement of the 3D object along the x-axis at time t. t Let z be the displacement of the 3D object along the y-axis at time t. t Let t be the displacement of the 3D object along the z-axis at time t. Place each 3D object model at its corresponding position in 3D space and remove the occluded parts from the 2D background image.

[0095] Please see Figure 3 As shown, in one embodiment, an image generation apparatus is provided, the apparatus comprising:

[0096] The data acquisition module 201 is used to acquire the user's target personal information and user request information, and determine the background image set and the three-dimensional object image set in the preset material library based on the target personal information;

[0097] Model training module 202 is used to fine-tune the preset image generation model based on the background image set and the three-dimensional object image set to obtain the target image generation model;

[0098] The image generation module 203 is used to generate visual training images based on the user request information, the target image generation model, and preset motion parameters.

[0099] In one possible implementation, the data acquisition module 201 is used for:

[0100] The system collects first rating data from pre-selected users for each background image in the preset material library, and constructs a first correspondence rule based on the pre-selected personal information of the pre-selected users and the first rating data. The first correspondence rule is used to determine the set of background images in the preset material library based on the target personal information.

[0101] record the second score data of the preselected user on each three-dimensional object image in the preset material library, and construct a second corresponding rule according to the preselected personal information and the second score data, the second corresponding rule being used to determine the set of three-dimensional object images in the preset material library based on the target personal information.

[0102] In a possible implementation, the data collection module 201 is configured to:

[0103] calculate the similarity data of the target personal information and the preselected personal information to obtain a set of similarity data, and select the similarity data with the highest similarity value in the set of similarity data as target similarity data;

[0104] select the preselected personal information corresponding to the target similarity data as the similar personal information corresponding to the target personal information;

[0105] obtain the first target score data corresponding to the similar personal information in the first score data according to the first corresponding rule, and obtain the set of background images by taking the background image corresponding to the first target score data as the target background image;

[0106] obtain the second target score data corresponding to the similar personal information in the second score data according to the second corresponding rule, and obtain the set of three-dimensional object images by taking the three-dimensional object image corresponding to the second target score data as the target three-dimensional object image.

[0107] In a possible implementation, the model training module 202 is configured to:

[0108] train the first image generation model in the preset image generation model according to the set of background images to obtain a first target image generation model, the first image generation model comprising a two-dimensional encoder copy, a first preset convolutional layer, and a preset two-dimensional image generation model, the two-dimensional encoder copy being the same as an encoder in the preset two-dimensional image generation model, and the two-dimensional encoder copy being connected to the preset two-dimensional image generation model through the first preset convolutional layer;

[0109] train the second image generation model in the preset image generation model according to the set of three-dimensional object images to obtain a second target image generation model, the second image generation model comprising a three-dimensional encoder copy, a second preset convolutional layer, and a preset three-dimensional image generation model, the three-dimensional encoder copy being the same as an encoder in the preset three-dimensional image generation model, and the three-dimensional encoder copy being connected to the preset two-dimensional image generation model through the second preset convolutional layer;

[0110] The target image generation model is generated based on the first target image generation model and the second target image generation model.

[0111] In a possible implementation, the image generation module 203 is configured to:

[0112] The user request information is input into a first target image generation model in the target image generation model, and generation of a two-dimensional image corresponding to the user request information is performed to obtain a target background image.

[0113] The user request information is input into a second target image generation model in the target image generation model, and generation of a three-dimensional image corresponding to the user request information is performed to obtain a target three-dimensional object image.

[0114] A visual training image is generated based on the preset motion parameter, the target background image, and the target three-dimensional object image.

[0115] In a possible implementation, the image generation module 203 is configured to:

[0116] Relative position data of the target background image and the target three-dimensional object image is determined based on the preset motion parameter.

[0117] The target background image and the target three-dimensional object image are fused based on the relative position data to generate a visual training image.

[0118] In a possible implementation, the image generation module 203 is configured to:

[0119] An occlusion area of the target background image occluded by the target three-dimensional object image is determined based on the relative position data.

[0120] The occlusion area is replaced with the target three-dimensional object image to generate a visual training image.

[0121] The application obtains target personal information and user request information of a user, determines a background image set and a three-dimensional object image set in a preset material library based on the target personal information, fine-tunes a preset image generation model according to the background image set and the three-dimensional object image set, and obtains a target image generation model, and generates a visual training image according to the user request information, the target image generation model and preset motion parameters. The background image set and the three-dimensional object image set are determined in the preset material library based on the target personal information, the preference of the user is predicted, the generated visual training image is more in line with the preference or habit of the user, meanwhile, the preset image generation model is used to generate the visual training image, compared with the process of manually making the visual training image, the generation efficiency of the visual training image is improved, and the visual training image is more abundant, the problem of insufficient visual training images is solved, and considering that the images generated by the preset image generation model under unlimited conditions may not be suitable for visual training, the preset image generation model is fine-tuned based on the determination of the background image set and the three-dimensional object image set in the preset material library based on the target personal information, so that the generation efficiency of the visual training image is ensured while the quality of the visual training image is ensured.

[0122] In one embodiment, a computer device, which can be a server, is provided, and an internal structure diagram thereof can be as shown in Figure 4 The computer device includes a processor, a memory, a network interface and a database connected through a system bus. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is configured to communicate with an external client through a network connection. The computer program is executed by the processor to implement functions or steps of a server side of an image generation method.

[0123] In one embodiment, a computer device, which can be a client, is provided, and an internal structure diagram thereof can be as shown in Figure 5 The computer device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is configured to communicate with an external server through a network connection. The computer program is executed by the processor to implement functions or steps of a client side of an image generation method.

[0124] In one embodiment, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, the processor implements the following steps when executing the computer program: obtaining target personal information of a user and user request information, and determining a background image set and a three-dimensional object image set in a preset material library based on the target personal information; fine-tuning a preset image generation model according to the background image set and the three-dimensional object image set to obtain a target image generation model; and generating a visual training image according to the user request information, the target image generation model, and preset motion parameters.

[0125] The present application obtains target personal information of a user and user request information, and determines a background image set and a three-dimensional object image set in a preset material library based on the target personal information; fine-tunes a preset image generation model according to the background image set and the three-dimensional object image set to obtain a target image generation model; and generates a visual training image according to the user request information, the target image generation model, and preset motion parameters. By determining a background image set and a three-dimensional object image set in a preset material library based on the target personal information, the preferences of the user are predicted, so that the generated visual training image is more in line with the preferences or habits of the user. At the same time, the preset image generation model is used to generate the visual training image, which improves the generation efficiency of the visual training image compared to the process of manually producing the visual training image, and thus makes the visual training image more abundant, solves the problem of insufficient visual training images, and takes into account that the images generated by the preset image generation model under unrestricted conditions may not be suitable for visual training. Therefore, the preset image generation model is fine-tuned based on the determination of the background image set and the three-dimensional object image set in the preset material library based on the target personal information, which ensures the generation efficiency of the visual training image while ensuring the quality of the visual training image.

[0126] In one embodiment, a computer readable storage medium is provided, which stores a computer program, the computer program is executed by a processor to implement the following steps: obtaining target personal information of a user and user request information, and determining a background image set and a three-dimensional object image set in a preset material library based on the target personal information; fine-tuning a preset image generation model according to the background image set and the three-dimensional object image set to obtain a target image generation model; and generating a visual training image according to the user request information, the target image generation model, and preset motion parameters.

[0127] The application obtains target personal information and user request information of a user, determines a background image set and a three-dimensional object image set in a preset material library based on the target personal information, fine-tunes a preset image generation model according to the background image set and the three-dimensional object image set, obtains a target image generation model, and generates a visual training image according to the user request information, the target image generation model and preset motion parameters. The background image set and the three-dimensional object image set are determined in the preset material library based on the target personal information, the preference of the user is predicted, the visual training image generated is more in line with the preference or habit of the user, meanwhile, the preset image generation model is used to generate the visual training image, compared with the process of manually making the visual training image, the generation efficiency of the visual training image is improved, and the visual training image is more abundant, the problem of insufficient visual training images is solved, and considering that the images generated by the preset image generation model under unrestricted conditions may not be suitable for visual training, the preset image generation model is fine-tuned based on the determination of the background image set and the three-dimensional object image set in the preset material library based on the target personal information, the generation efficiency of the visual training image is ensured, and the quality of the visual training image is ensured.

[0128] It should be noted that the functions or steps that the computer readable storage medium or the computer device can implement correspond to the related descriptions of the server side and the client side in the foregoing method embodiments, and will not be described again here to avoid repetition.

[0129] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when executed, can include the processes of the above-mentioned embodiment methods. Any reference to memory, storage, database or other medium used in the embodiments provided by the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0130] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is exemplified, and in actual application, the above-mentioned functions can be completed by different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.

[0131] The above-mentioned embodiments are only used to illustrate the technical solutions of the present application, but not limit it. Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features. The modification or replacement does not make the essence of the corresponding technical solution deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.

Claims

1. An image generation method characterized by, The method comprises: obtaining target personal information of a user and user request information, and determining a background image set and a three-dimensional object image set in a preset material library based on the target personal information; fine-tuning a preset image generation model according to the background image set and the three-dimensional object image set to obtain a target image generation model; The step of fine-tuning the preset image generation model according to the background image set and the three-dimensional object image set to obtain the target image generation model comprises: training a first image generation model in the preset image generation model according to the background image set to obtain a first target image generation model, wherein the first image generation model comprises a two-dimensional encoder copy, a first preset convolutional layer, and a preset two-dimensional image generation model, the two-dimensional encoder copy is the same as an encoder in the preset two-dimensional image generation model, and the two-dimensional encoder copy is connected to the preset two-dimensional image generation model through the first preset convolutional layer; training a second image generation model in the preset image generation model according to the three-dimensional object image set to obtain a second target image generation model, wherein the second image generation model comprises a three-dimensional encoder copy, a second preset convolutional layer, and a preset three-dimensional image generation model, the three-dimensional encoder copy is the same as an encoder in the preset three-dimensional image generation model, and the three-dimensional encoder copy is connected to the preset three-dimensional image generation model through the second preset convolutional layer; obtaining the target image generation model based on the first target image generation model and the second target image generation model; generating a visual training image according to the user request information, the target image generation model, and preset motion parameters.

2. The image generation method according to claim 1, characterized by, Before the step of determining the background image set and the three-dimensional object image set in the preset material library based on the target personal information, the method further comprises: statistically obtaining first score data of each background image in the preset material library by a preselected user, and constructing a first corresponding rule based on preselected personal information of the preselected user and the first score data, wherein the first corresponding rule is used to determine the background image set in the preset material library based on the target personal information; statistically obtaining second score data of each three-dimensional object image in the preset material library by the preselected user, and constructing a second corresponding rule based on the preselected personal information and the second score data, wherein the second corresponding rule is used to determine the three-dimensional object image set in the preset material library based on the target personal information.

3. The image generation method of claim 2, wherein, The step of determining the background image set and the three-dimensional object image set in the preset material library based on the target personal information comprises: calculating similarity data of the target personal information and the preselected personal information to obtain a similarity data set, and selecting similarity data with the highest similarity value in the similarity data set as target similarity data; taking the preselected personal information corresponding to the target similarity data as similar personal information corresponding to the target personal information. According to the first corresponding rule, first target score data corresponding to the similar personal information is obtained in the first score data, a background image corresponding to the first target score data is taken as a target background image, and the background image set is obtained; According to the second corresponding rule, second target score data corresponding to the similar personal information is obtained in the second score data, a three-dimensional object image corresponding to the second target score data is taken as a target three-dimensional object image, and the three-dimensional object image set is obtained.

4. The image generation method of claim 1, wherein, The step of generating a visual training image according to the user request information, the target image generation model, and a preset motion parameter includes: inputting the user request information into a first target image generation model in the target image generation model to generate a two-dimensional image corresponding to the user request information, and obtaining a target background image; inputting the user request information into a second target image generation model in the target image generation model to generate a three-dimensional image corresponding to the user request information, and obtaining a target three-dimensional object image; generating a visual training image according to the preset motion parameter, the target background image, and the target three-dimensional object image.

5. The image generation method of claim 4, wherein, The step of generating a visual training image according to the preset motion parameter, the target background image, and the target three-dimensional object image includes: determining relative position data of the target background image and the target three-dimensional object image based on the preset motion parameter; fusing the target background image and the target three-dimensional object image based on the relative position data to generate a visual training image.

6. The image generation method of claim 5, wherein, The step of fusing the target background image and the target three-dimensional object image based on the relative position data to generate a visual training image includes: determining an occlusion area of the target background image occluded by the target three-dimensional object image based on the relative position data; replacing the occlusion area with the target three-dimensional object image to generate a visual training image.

7. An image generation apparatus characterized by comprising: The device includes: a data acquisition module configured to obtain target personal information of a user and user request information, and determine a background image set and a three-dimensional object image set in a preset material library based on the target personal information; The model training module is configured to fine-tune a preset image generation model according to the background image set and the three-dimensional object image set to obtain a target image generation model. The step of fine-tuning the preset image generation model according to the background image set and the three-dimensional object image set to obtain the target image generation model includes: training a first image generation model in the preset image generation model according to the background image set to obtain a first target image generation model. The first image generation model includes a two-dimensional encoder copy, a first preset convolutional layer, and a preset two-dimensional image generation model. The two-dimensional encoder copy is the same as an encoder in the preset two-dimensional image generation model, and the two-dimensional encoder copy is connected to the preset two-dimensional image generation model through the first preset convolutional layer. Training a second image generation model in the preset image generation model according to the three-dimensional object image set to obtain a second target image generation model. The second image generation model includes a three-dimensional encoder copy, a second preset convolutional layer, and a preset three-dimensional image generation model. The three-dimensional encoder copy is the same as an encoder in the preset three-dimensional image generation model, and the three-dimensional encoder copy is connected to the preset three-dimensional image generation model through the second preset convolutional layer. The target image generation model is obtained based on the first target image generation model and the second target image generation model. The image generation module is configured to generate a visual training image according to the user request information, the target image generation model, and preset motion parameters.

8. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The computer program is executed by the processor to implement the steps of the image generation method according to any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 8. The computer program is executed by the processor to implement the steps of the image generation method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Method and system for visual function training

    CN110292515A

  • Image processing method, device and equipment and computer readable storage medium

    CN116704221A