Method, device, electronic device and storage medium for generating virtual image
By generating a 3D virtual image and utilizing mask images and three-dimensional modeling information, the problem of simple guidance signs on the navigation interface is solved, personalization and fun are enhanced, and the user experience is improved.
Patent Information
- Application Number
- CN202411719882.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-27
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2044-11-27
AI Technical Summary
In existing electronic navigation applications, guidance signs are too simple and easily confused with other signs on the navigation interface, which affects the user experience and makes it difficult to meet the user's personalized needs.
By generating a 3D virtual image, using mask images and 3D modeling information, combined with style description information, a virtual image that meets the user's personalized needs is generated, thereby improving the fun and user experience of the navigation interface.
It realizes personalized guidance signs, improves user experience and the aesthetics of the navigation interface, reduces user input pressure, and enhances user participation and fun.
Smart Images

Figure CN119559312B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, particularly computer vision, deep learning, and large models, and can be applied to scenarios such as AIGC-based content generation based on artificial intelligence. Specifically, it relates to a method, device, electronic device, and storage medium for generating a virtual image. Background Art
[0002] AIGC (Artificial Intelligence Generated Content) is a technology based on artificial intelligence techniques such as generative adversarial networks and large pre-trained models. It uses learning and recognition from existing data to generate relevant content with appropriate generalization capabilities. For example, it can be used to generate personalized guide signs. Summary of the Invention
[0003] The present disclosure provides a method, device, electronic device, and storage medium for generating a virtual image.
[0004] According to one aspect of the present disclosure, a method for generating a virtual image is provided, comprising: in response to received style description information for a target vehicle, obtaining a mask image and three-dimensional modeling information for the target vehicle; each mask area in the mask image represents a position for adding a material image in an initial texture image for the target vehicle; matching the material image with the style description information; extracting from the three-dimensional modeling information the initial texture image and a mapping relationship between each two-dimensional coordinate in the initial texture image and each three-dimensional coordinate in the virtual image; generating a target texture image by processing the style description information, the initial texture image and the mask image; and generating a target virtual image by processing the target texture image based on the mapping relationship.
[0005] According to another aspect of the present disclosure, a device for generating a virtual image is provided, including: an acquisition module, a first extraction module, a first generation module, and a second generation module.
[0006] An acquisition module is configured to acquire, in response to received style description information for a target vehicle, a mask image and three-dimensional modeling information for the target vehicle; each mask region in the mask image represents a location for adding a material image to an initial texture image for the target vehicle; and the material image is matched with the style description information.
[0007] The first extraction module is used to extract the initial texture image and the mapping relationship between each two-dimensional coordinate in the initial texture image and each three-dimensional coordinate in the virtual image from the three-dimensional modeling information.
[0008] The first generation module is used to generate a target texture image by processing style description information, an initial texture image and a mask image.
[0009] The second generating module is used to generate a target virtual image by processing the target texture image based on the mapping relationship.
[0010] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described above.
[0011] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to execute the method described above.
[0012] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein the computer program implements the method described above when executed by a processor.
[0013] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.
[0015] Figure 1 Schematically illustrates an exemplary system architecture to which the method and apparatus for generating a virtual image according to an embodiment of the present disclosure may be applied;
[0016] Figure 2 The flowchart of the method for generating a virtual image according to an embodiment of the present disclosure is schematically shown;
[0017] Figure 3A Schematically illustrates a schematic diagram of generating a target virtual image at a predetermined position based on a single material image according to an embodiment of the present disclosure;
[0018] Figure 3B Schematically illustrates a schematic diagram of generating a target virtual image based on matching mapping positions of multiple material images according to another embodiment of the present disclosure;
[0019] Figure 3C Schematically illustrates a schematic diagram of generating a target virtual image based on matching mapping positions of a custom image and multiple material images according to an embodiment of the present disclosure;
[0020] Figure 4AA schematic diagram of modifying an image uploaded by a user based on modification requirements to generate a customized image according to an embodiment of the present disclosure is schematically shown;
[0021] Figure 4B Schematically illustrates a schematic diagram of modifying an image uploaded by a user to generate a custom image based on modification requirements according to another embodiment of the present disclosure;
[0022] Figure 5 A block diagram schematically shows a virtual image generating device according to an embodiment of the present disclosure; and
[0023] Figure 6 A block diagram schematically shows an electronic device suitable for implementing a method for generating a virtual image according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0024] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0025] Electronic navigation applications have made travel more convenient for users. As the number of users increases, different users have developed personalized requirements for navigational signs within the navigation interface. Currently, when using electronic maps for navigation, the user's location is typically indicated by a dot or arrow on the navigation interface. These signs are overly simple and similar in shape to other navigational signs, which can easily cause visual confusion and affect the user experience. Therefore, there is an urgent need for navigational signs that can meet the personalized needs of different users to enhance the user experience.
[0026] Therefore, an embodiment of the present invention provides a method for generating a virtual image for a navigation map. By generating a 3D virtual image, it not only meets the user's personalized needs for guidance signs in the navigation interface, but also increases the fun of electronic navigation applications and improves user experience.
[0027] Figure 1 An exemplary system architecture to which the method and apparatus for generating a virtual image according to an embodiment of the present disclosure can be applied is schematically shown.
[0028] It should be noted that Figure 1The examples shown are merely examples of system architectures to which the embodiments of the present disclosure may be applied, to help those skilled in the art understand the technical content of the present disclosure. This does not mean that the embodiments of the present disclosure cannot be applied to other devices, systems, environments, or scenarios. For example, in another embodiment, an exemplary system architecture to which the method and apparatus for generating a virtual image may be applied may include a terminal device, but the terminal device may implement the method and apparatus for generating a virtual image provided by the embodiments of the present disclosure without interacting with a server.
[0029] like Figure 1 As shown, the system architecture 100 according to this embodiment may include a terminal device 101, a network 102, and a server 103. The network 102 is used to provide a medium for a communication link between the terminal device 101 and the server 103. The network 102 may include various connection types, such as wired and / or wireless communication links, etc.
[0030] A user may use a terminal device 101 to interact with a server 103 via a network 102 to receive or send messages, etc. Various communication client applications may be installed on the terminal device 101, such as electronic navigation applications, knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software (for example only).
[0031] The terminal devices 101 , 102 , and 103 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, and desktop computers.
[0032] Server 103 may be a server that provides various services, such as a background management server (for example only) that supports content browsed by a user using terminal device 101. The background management server may analyze and process received data such as user requests, and feed back processing results (e.g., web pages, information, or data obtained or generated based on user requests) to the terminal device.
[0033] It should be noted that the method for generating a virtual image provided in the embodiment of the present disclosure can generally be executed by the terminal device 101. Accordingly, the apparatus for generating a virtual image provided in the embodiment of the present disclosure can also be provided in the terminal device 101.
[0034] Alternatively, the virtual image generation method provided by the embodiment of the present disclosure may also be generally executed by the server 103. Accordingly, the virtual image generation device provided by the embodiment of the present disclosure may generally be set in the server 103. The virtual image generation method provided by the embodiment of the present disclosure may also be executed by a server or server cluster that is different from the server 103 and can communicate with the terminal device 101 and / or the server 103. Accordingly, the virtual image generation device provided by the embodiment of the present disclosure may also be set in a server or server cluster that is different from the server 103 and can communicate with the terminal device 101 and / or the server 103.
[0035] For example, a user may input style description information to the terminal device 101 via voice or text, such as "Please generate a cute 3D car logo for the model of XX" 110. The terminal device 101 may send the style description information to the server 103. The server 103 may obtain a mask image and a material image based on the style description information, generate a target virtual image 120 based on the method provided in an embodiment of the present invention, and send the target virtual image 120 to the terminal device 101.
[0036] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is merely illustrative. Any number of terminal devices, networks and servers may be provided as required.
[0037] In the technical solution disclosed herein, the collection, storage, use, processing, transmission, provision, disclosure and application of user personal information involved comply with the provisions of relevant laws and regulations, take necessary confidentiality measures, and do not violate public order and good morals.
[0038] In the technical solution of the present disclosure, the user's authorization or consent is obtained before obtaining or collecting the user's personal information. The images used in the examples are all from public image data sources.
[0039] Figure 2 The flowchart of the method for generating a virtual image according to an embodiment of the present disclosure is schematically shown.
[0040] like Figure 2 As shown, the method 200 includes operations S210 to S240.
[0041] In operation S210 , in response to the received style description information for the target vehicle, a mask image and three-dimensional modeling information for the target vehicle are acquired.
[0042] In operation S220 , an initial texture image and a mapping relationship between each two-dimensional coordinate in the initial texture image and each three-dimensional coordinate in the avatar are extracted from the three-dimensional modeling information.
[0043] In operation S230 , a target texture image is generated by processing the style description information, the initial texture image, and the mask image.
[0044] In operation S240 , a target avatar is generated by processing the target texture image based on the mapping relationship.
[0045] According to an embodiment of the present disclosure, the style description information can be voice information or text information. The content of the style description information includes, but is not limited to, the appearance color and overall style of the target avatar of the target vehicle. For example, the style may be simple, cute, or black.
[0046] According to an embodiment of the present disclosure, the target vehicle can be the vehicle currently being driven by the user, or can be a dream vehicle that the user wants to drive. The target virtual image can be used to be displayed in the navigation map.
[0047] According to an embodiment of the present disclosure, a material library containing material images of various styles can be pre-configured for material images, and material images matching the style description information can be determined from the material library. For example, the style description information can be cute, and the matching material images can be various cartoon characters.
[0048] In some embodiments, a plurality of material images matching the style description information may be recommended to the user first, and the final material image may be determined through interaction with the user.
[0049] In some embodiments, material images may be recommended to the user based on the user's historical preferences in combination with the style description information, thereby further improving the user's interactive experience.
[0050] According to the embodiments of the present disclosure, the initial texture image can be style-modified through a simple style description, which can reduce the user's input pressure and further improve the user experience.
[0051] According to embodiments of the present disclosure, a target texture image can be generated by invoking a generative model to process style description information, an initial texture image, and a mask image. For example, the generative model can be a generative adversarial network or a diffusion model. This is not specifically limited in the present disclosure.
[0052] According to an embodiment of the present disclosure, each mask region in the mask image represents a location for adding a material image to an initial texture image for a target vehicle.
[0053] According to an embodiment of the present disclosure, the 3D modeling information may include an initial texture image and 3D mesh information used to construct a 3D avatar of the target vehicle. The 3D mesh information represents the mapping relationship between each 2D coordinate in the initial texture image and each 3D coordinate in the avatar.
[0054] In related examples, multiple 2D images of a target vehicle are captured from multiple perspectives, and then a 3D virtual image is reconstructed from these 2D images. This approach results in inconsistent 3D textures in the 3D virtual image due to variations in camera pose during the generation of the 2D images from each perspective. For example, the camera poses of the same object may differ between the front and left views.
[0055] However, because the mask image in this embodiment of the present invention reserves space for adding source images, the initial mapping relationship between the 2D coordinates in the target texture image generated after adding the source images to the initial texture image and the 3D coordinates in the avatar remains consistent. Therefore, based on this mapping relationship, the texture of the target avatar generated by processing the target texture image is 3D consistent.
[0056] At the same time, by adding material images that match the style description information to the initial texture image, it not only meets the user's personalized needs, but also reduces the user's input pressure, further improving the user experience.
[0057] Reference below Figure 3A to Figure 3C 、 Figure 4A-4B , combined with specific embodiments Figure 2 The method shown is further explained.
[0058] Figure 3A The diagram schematically shows a method of generating a target virtual image at a predetermined position based on a single material image according to an embodiment of the present disclosure.
[0059] like Figure 3A As shown, in embodiment 300A, the style description information for the target vehicle input by the user may be "Please generate a cute 3D car logo of the XX model" 301. "Style: cute style" 302 and "Car model: XX" 303 may be extracted from the style description information.
[0060] According to an embodiment of the present disclosure, generating a target texture image by processing style description information, an initial texture image, and a mask image may include the following operations: determining a target material image based on the style description information; determining a target position for adding the target material image in the initial texture image based on a mask area in the mask image; and adding the target material image to the target position in the initial texture image by calling a generation model to generate a target texture image.
[0061] For example, first, based on “style: cute style” 302 , a material image 304A may be acquired.
[0062] According to an embodiment of the present disclosure, determining the target material image based on the style description information may include the following operations: extracting a first keyword from the style description information; and determining the target material image from the candidate material images based on the matching degree between the first keyword and the subject word of the candidate material image.
[0063] For example, the first keyword may be "cute", and the subject words of the candidate material images may be "cartoon", "cute animals", etc. The target material image may represent an image whose subject word of the candidate material image matches the first keyword more than a predetermined matching threshold.
[0064] Next, based on "Vehicle Type: XX" 303, an initial texture image 305A and a mask image 306A for vehicle type XX are obtained. Initial texture image 305A may include texture images of various vehicle components, such as the body, windows, and wheels, as well as hand images used to generate the virtual driver. The row / column position corresponding to the marker "TTTT" in mask image 306A indicates the location in initial texture image 305A where material image 304A is added. For example, this location may correspond to the location marked by the dashed box in the initial texture image.
[0065] Next, the diffusion model may be called to add the material image 304A to the initial texture image 305A according to the position indicated in the mask image 306A to generate a target texture image 307A.
[0066] Finally, based on the mapping relationship 308 between the two-dimensional coordinates in the initial texture image and the three-dimensional coordinates in the avatar, the target texture image 307A is mapped to the initial three-dimensional grid to generate a target avatar 309A.
[0067] According to the embodiments of the present disclosure, by screening material images based on the degree of matching between style keywords and subject words of the material images, the matching efficiency of the material images can be improved, thereby improving the generation efficiency of the target virtual image.
[0068] According to an embodiment of the present disclosure, a target texture image is generated under the guidance of a target position in a mask image by calling a generation model. Since the initial texture image reserves an addition position for the material image at the target position corresponding to the mask image, the initial mapping relationship between each two-dimensional coordinate in the target texture image and each three-dimensional coordinate in the virtual image still satisfies the initial mapping relationship, thereby improving the 3D texture consistency of the target texture image.
[0069] In some embodiments, the user may not only want to add a virtual character to the driving position, but may also want to decorate the exterior parts of the vehicle, such as the hood, windows, wheels, etc.
[0070] Therefore, for scenes with multiple material images and multiple mask areas, the correspondence between each target material image and each mask area can be generated based on the attributes of each target material image and the attributes of each mask area; and based on the correspondence, the target positions of each target material image can be determined.
[0071] According to an embodiment of the present disclosure, the attribute of each target material image can be the image size, and the attribute of each mask area can be the mask area size. For example, the region matching degree can be calculated based on the image size of each target material image and the size of each mask area. The region matching degree can be a percentage of the image area and the mask area area, or it can be the difference between the image area and the mask area area. The embodiment of the present disclosure does not specifically limit the calculation method of the region matching degree, as long as it can characterize the difference between the image area and the mask area area. The size of the target material image that matches the mask area can be the material image that is closest to the size of the mask area.
[0072] According to embodiments of the present disclosure, the attribute of each target material image can also be the location where the image is added. For example, for a material image of an animal, the location where the image is added can be pre-configured in the material library as the hood. Then, the mask area corresponding to the hood location in the mask image can be determined as corresponding to the animal material image.
[0073] Figure 3B The diagram schematically shows a method of generating a target virtual image based on matching map positions of multiple material images according to another embodiment of the present disclosure.
[0074] like Figure 3B As shown, in embodiment 300B, the style description information input by the user is the same as that in embodiment 300A. The difference is that the candidate material images 310 obtained based on "Style: Cute" 302 may include seven images, and the target material images 304B determined by the user through interaction with the user, for example, by displaying them to the user through a visual interface and selecting them through voice input or click operation, may include a tiger image, a unicorn image, and a baby elephant image.
[0075] In this embodiment 300B, the mask image 306B indicates the locations where each image in the target material image 304B is added to the initial texture image 305B. For example, the mask area marked "MM" is used to add a tiger image, the mask area marked "Y" is used to add a unicorn image, and the mask area marked "P" is used to add an elephant image.
[0076] The generative model is invoked to process target material image 304B, initial texture image 305B, and mask image 306B, generating target texture image 307B, which has the target material images added to the corresponding mask regions. Finally, based on mapping relationship 308, target texture image 307B is mapped onto the initial 3D mesh, generating target avatar 309B.
[0077] According to the embodiments of the present disclosure, by matching the attributes of the mask area with the attributes of each target material image, various virtual images can be generated flexibly and variably, and the fun can be increased through interaction with the user, further improving the user experience.
[0078] Furthermore, embodiments of the present invention can also generate target avatars based on user-defined source images. For example, based on the content of each target source image, mask regions corresponding to each target source image are determined according to a predetermined correspondence; and based on the attributes of each mask region, the attributes of each target source image are adjusted accordingly.
[0079] Figure 3C The diagram schematically shows a method of generating a target virtual image based on matching mapping positions of a custom image and multiple material images according to an embodiment of the present disclosure.
[0080] like Figure 3C As shown, in this embodiment 300C, the style description information is the same as in embodiments 300A and 300B, and the obtained initial texture image 305B and mask image 306B are the same as in embodiment 300B. The difference is that in addition to the pre-configured images in the material library, the candidate material images 320 also include user-defined images. User-defined images can be generated by the user through beautification and modification using any image processing software. User-defined images can be stored in the material library or uploaded by the user during user interaction.
[0081] According to an embodiment of the present disclosure, the content of each target material image can be a user-defined image, such as a head image. The predetermined correspondence can represent the correspondence between the material content and the mask region, for example, the head image corresponds to the mask region in the mask image representing the driver's position. For example, the mask region labeled "MM" in mask image 306B.
[0082] Because the size or orientation of the user-defined image may not match the mask area, the properties of the user-defined image can be adjusted based on the properties of each mask area. For example, the size of the user-defined image can be adjusted to be slightly smaller than the size of the mask area used to represent the driver's position. The orientation of the user-defined image can be adjusted to be the same as the orientation of the mask area used to represent the driver's position. For example, the positions of key points such as hair, ears, and nose can be finely identified in the mask area. The orientation and / or size of the user-defined image can be adjusted so that the key points of the user-defined image are in the same positions as the key points of the mask area.
[0083] The generation model is invoked to process target material image 304C, initial texture image 305B, and mask image 306B, generating target texture image 307C, which has the target material images added to the corresponding mask regions. Finally, based on mapping relationship 308, target texture image 307C is mapped onto the initial 3D mesh, generating target avatar 309C.
[0084] According to the embodiments of the present disclosure, the properties of each target material image are adjusted based on the mask area properties, thereby improving the consistency of each key position in the target texture image and the reserved position in the initial texture image, thereby improving the consistency and aesthetics of the 3D texture of the target virtual image and enhancing the user viewing experience.
[0085] For application scenarios where users want to generate 3D car logos through customized images, an embodiment of the present invention can perform semantic segmentation on the initial image to generate a first image of the target part; and in response to the received modification requirement information for the initial image, generate a material image by processing the first image based on the modification requirement information.
[0086] According to embodiments of the present disclosure, the initial image can be a user's own photo or a photo of the user's pet. The target location can be user-defined or determined based on the location the user wants to add to the initial texture image. For example, if the driver's head is added to the initial texture image, the target location can be the head. If the wheel hub is added to the initial texture image, the target location can be the pet's foot.
[0087] For example, an image segmentation model can be used to implement semantic segmentation of the initial image. Alternatively, key point detection can be performed on the initial image before semantic segmentation to determine the locations of key points of target parts, such as hair, ears, and nose. The image segmentation model can be any model capable of implementing image segmentation, and this disclosure does not impose specific limitations on this.
[0088] Figure 4AThe diagram schematically shows a process of modifying an image uploaded by a user based on modification requirements to generate a customized image according to an embodiment of the present disclosure.
[0089] like Figure 4A As shown, the image segmentation model is used to perform semantic segmentation on the initial image 401 to obtain a semantic segmentation map 402 , and the head image 403 is extracted from the initial image 401 according to the semantic segmentation map 402 .
[0090] According to embodiments of the present disclosure, the modification request information can include a modification style, such as "cool style," or can be specific to the part or item being modified, such as adding a cat-ear-shaped hairpin to the hair. The modification request information "Please add cartoon modification" 404 and the head image 403 are input into the diffusion model to generate a material image 405 decorated with cartoon accessories.
[0091] For example, for a target modification style, a first modification material image and a first position for adding the first modification material image to the first image may be determined according to the target modification style; and the first material image may be generated by adding the first modification material image to the first position.
[0092] According to embodiments of the present disclosure, the modification material images and addition locations corresponding to the target modification style can be pre-configured. The addition locations can be pixel coordinates. For example, for a "cute" modification style, the modification material images can be cartoon ornaments such as "red cheeks" or "little red flowers," and the addition locations can be "face," "hair," or "ears."
[0093] Finally, the first modified material image may be added to the first image at the first position using the diffusion model to generate a first material image with a "red cheek" and a "little red flower" on the hair.
[0094] According to the embodiments of the present disclosure, an initial image is modified based on a modification style, which reduces the input pressure of a user and improves the user experience for users who prefer efficiency.
[0095] For example: for the modification requirements of a part to be modified and a modification object for the part to be modified, the second position of the part to be modified can be extracted from the first image; according to the modification style, the second modification material image corresponding to the modification object is determined; and the second material image is generated by adding the second modification material image at the second position.
[0096] For example, the part to be modified could be "hair," and the modification object for that part could be "hat." The second location could be the pixel coordinates of the "hair" area in the first image. Based on a modification style, such as "cute," the second modified source image could be "a furry hat with bunny ears."
[0097] Finally, the second modified material image may be added to the first image at the second position using the diffusion model to generate a second material image with a "furry hat with rabbit ears" on the hair.
[0098] According to the embodiments of the present disclosure, the initial image is modified based on the position to be modified and the specific modifying object, which increases user participation and improves the user experience for users who prefer to pursue fun and aesthetics.
[0099] For users who require both efficiency and fun, we can first provide them with candidate materials based on the modification style, and then perform secondary modification on the candidate materials based on the interaction with the user.
[0100] In actual applications, when a user is immersed in the interactive selection of various material images, a scenario may occur where the modification style of the user-customized image conflicts with the modification style of the target virtual image.
[0101] Figure 4B The following schematically illustrates a schematic diagram of modifying an image uploaded by a user based on modification requirements to generate a customized image according to another embodiment of the present disclosure.
[0102] like Figure 4B As shown, in embodiment 400B, the process of extracting a head image 403 through semantic segmentation based on the initial image 401 is the same as in embodiment 400A. In this embodiment 400B, style conflict detection is added between the modification style request "Please add cartoon modification" 404 and the style description information "Please generate a cute 3D car logo for the XX model."
[0103] By executing S410, it is determined whether there is a conflict between the modification style requirement and the style description information. For example, a first keyword may be extracted from the style description information; a second keyword may be extracted from the target modification style; and in response to a match between the first keyword and the second keyword being less than a predetermined threshold, a warning message indicating a style conflict is generated.
[0104] For example, the first keyword may be "cute style" and the second keyword may be "cartoon." Since cartoon images generally represent a cute style, it can be determined that there is no style conflict. The diffusion model can be called to process the modification request information "Please add cartoon modification" 404 and the head image 403 to generate a material image 405 decorated with cartoon accessories.
[0105] When it is determined by executing S410 that there is a conflict between the modification style requirement and the style description information, a style conflict warning 406 may be sent to the user.
[0106] For example, the first keyword in the style description is "cute style," and the second keyword extracted from the modification request information is "bat accessories." Bat accessories typically represent a "dark style," which doesn't match the "cute style." Therefore, a style conflict warning can be sent to the user to prompt them to change their modification request information. The style conflict warning information can be presented through a visual interface or voice announcement, which is not specifically limited in this embodiment.
[0107] According to the embodiments of the present disclosure, through style conflict verification, the interaction with the user can be increased during the process of generating the target virtual image. At the same time, the probability of generating virtual images with very different styles is reduced, thereby improving user satisfaction.
[0108] Figure 5 A block diagram schematically shows a device for generating a virtual image according to an embodiment of the present disclosure.
[0109] like Figure 5 As shown, the apparatus 500 may include an acquisition module 510 , a first extraction module 520 , a first generation module 530 and a second generation module 540 .
[0110] The acquisition module 510 is used to obtain a mask image and three-dimensional modeling information for the target vehicle in response to the received style description information for the target vehicle; each mask area in the mask image represents a location for adding a material image to the initial texture image for the target vehicle; the material image is matched with the style description information.
[0111] The first extraction module 520 is configured to extract, from the three-dimensional modeling information, an initial texture image and a mapping relationship between each two-dimensional coordinate in the initial texture image and each three-dimensional coordinate in the virtual image.
[0112] The first generating module 530 is configured to generate a target texture image by processing the style description information, the initial texture image, and the mask image.
[0113] The second generating module 540 is configured to generate a target virtual image by processing the target texture image based on the mapping relationship.
[0114] According to an embodiment of the present disclosure, the first generating module includes: a first determining submodule, a second determining submodule and a first generating submodule.
[0115] The first determination submodule is configured to determine a target material image based on the style description information. The second determination submodule is configured to determine a target location within the initial texture image for adding the target material image based on a mask region within the mask image. The first generation submodule is configured to generate a target texture image by adding the target material image to the target location within the initial texture image by invoking a generation model.
[0116] According to an embodiment of the present disclosure, the target material image includes multiple images; the mask area includes multiple areas; and the second determination submodule includes: a first generation unit and a first determination unit.
[0117] The first generating unit is configured to generate a correspondence between each target material image and each mask area according to the attributes of each target material image and the attributes of each mask area. The first determining unit is configured to determine each target position of each target material image based on the correspondence.
[0118] According to an embodiment of the present disclosure, the target material image includes multiple images; the mask area includes multiple areas; the first generation module further includes: a third determination submodule and an adjustment submodule.
[0119] The third determination submodule is configured to determine, based on the content of each target material image and a predetermined correspondence relationship, each mask region corresponding to each target material image; wherein the predetermined correspondence relationship represents the correspondence between the material content and the mask region. The adjustment submodule is configured to adjust the attributes of each target material image accordingly based on the attributes of each mask region.
[0120] According to an embodiment of the present disclosure, the first determination submodule includes: an extraction unit and a second determination unit. The extraction unit is configured to extract a first keyword from the style description information; and the second determination unit is configured to determine a target material image from the candidate material images based on a degree of match between the first keyword and a subject word of the candidate material images.
[0121] According to an embodiment of the present disclosure, the apparatus further includes a segmentation module and a processing module. The segmentation module is configured to perform semantic segmentation on the initial image to generate a first image of the target area. The processing module is configured to generate a material image by processing the first image based on received modification requirement information for the initial image.
[0122] According to an embodiment of the present disclosure, modification requirement information includes a target modification style; and a processing module includes a fourth determination submodule and a second generation submodule. The fourth determination submodule is configured to determine, based on the target modification style, a first modification material image and a first position for adding the first modification material image to the first image. The second generation submodule is configured to generate the first material image by adding the first modification material image to the first position.
[0123] According to an embodiment of the present disclosure, the requirement information also includes a part to be modified and a modification target for the part to be modified; and the processing module further includes an extraction submodule and a fifth determination submodule. The extraction submodule is configured to extract a second position of the part to be modified from the first image. The fifth determination submodule is configured to determine a second modification material image corresponding to the modification target based on the modification style. The third generation submodule is configured to generate a second material image by adding the second modification material image to the second position.
[0124] According to an embodiment of the present disclosure, the device further includes: a second extraction module, a third extraction module and a third generation module.
[0125] The second extraction module is configured to extract a first keyword from the style description information. The third extraction module is configured to extract a second keyword from the target modification style. The third generation module is configured to generate warning information indicating a style conflict in response to a match between the first keyword and the second keyword being less than a predetermined threshold.
[0126] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0127] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described above.
[0128] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute the method described above.
[0129] According to an embodiment of the present disclosure, a computer program product includes a computer program, and when the computer program is executed by a processor, the computer program implements the method described above.
[0130] Figure 6 A schematic block diagram of an example electronic device 600 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0131] like Figure 6 As shown, device 600 includes a computing unit 601, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. RAM 603 may also store various programs and data required for the operation of device 600. Computing unit 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to bus 604.
[0132] Various components in device 600 are connected to I / O interface 605, including an input unit 606, such as a keyboard, mouse, etc.; an output unit 607, such as various types of displays, speakers, etc.; a storage unit 608, such as a magnetic disk, optical disk, etc.; and a communication unit 609, such as a network card, modem, wireless communication transceiver, etc. The communication unit 609 allows device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0133] The computing unit 601 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as the avatar generation method. For example, in some embodiments, the avatar generation method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the avatar generation method described above can be performed. Alternatively, in other embodiments, the computing unit 601 can be configured to perform the avatar generation method through any other suitable means (e.g., via firmware).
[0134] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-a-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0135] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0136] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0137] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0138] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0139] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.
[0140] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.
[0141] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A method for generating a virtual image, comprising: In response to the received style description information for the target vehicle, obtaining a mask image and three-dimensional modeling information for the target vehicle; Each mask region in the mask image represents a location for adding a material image to the initial texture image for the target vehicle; the material image matches the style description information; extracting the initial texture image and the mapping relationship between each two-dimensional coordinate in the initial texture image and each three-dimensional coordinate in the virtual image from the three-dimensional modeling information; Determining a target material image according to the style description information; determining a target position for adding the target material image in the initial texture image based on a mask area in the mask image; as well as By calling the generation model, adding the target material image to the target position in the initial texture image to generate a target texture image; as well as Based on the mapping relationship, a target virtual image is generated by processing the target texture image.
2. The method according to claim 1, wherein The target material images include multiple ones; the mask areas include multiple ones; The step of determining a target position for adding the target material image to the initial texture image based on a mask area in the mask image includes: generating a correspondence between each target material image and each mask area according to the attributes of each target material image and the attributes of each mask area; and Based on the corresponding relationship, each target position of each target material image is determined.
3. The method according to claim 1, wherein The target material images include a plurality of images; the mask areas include a plurality of regions; and the method further includes: Determining, based on the content of each target material image, each mask region corresponding to each target material image according to a predetermined correspondence relationship; wherein the predetermined correspondence relationship represents a correspondence between the material content and the mask region; According to the attributes of each mask area, the attributes of each target material image are correspondingly adjusted.
4. The method according to claim 1, wherein The step of determining a target material image according to the style description information includes: Extracting a first keyword from the style description information; and The target material image is determined from the candidate material images based on a matching degree between the first keyword and a subject word of the candidate material image.
5. The method according to claim 1, wherein The method further comprises: Performing semantic segmentation on the initial image to generate a first image of the target area; and In response to the received modification requirement information for the initial image, the material image is generated by processing the first image based on the modification requirement information.
6. The method according to claim 5, wherein: The modification requirement information includes a target modification style; The step of generating the material image by processing the first image based on the modification requirement information received for the initial image in response to the modification requirement information includes: determining, according to the target modification style, a first modification material image and a first position for adding the first modification material image to the first image; and A first material image is generated by adding the first modified material image to the first position.
7. The method according to claim 6, wherein: The modification requirement information also includes the part to be modified and the modification object for the part to be modified; The step of generating the material image by processing the first image based on the modification requirement information received for the initial image in response to the modification requirement information includes: extracting a second position of the part to be modified from the first image; determining a second modified material image corresponding to the modified object according to the modification style; A second material image is generated by adding the second modified material image at the second position.
8. The method according to claim 6, further comprising: extracting a first keyword from the style description information; extracting a second keyword from the target modification style; In response to the matching degree between the first keyword and the second keyword being less than a predetermined threshold, warning information for indicating a style conflict is generated.
9. The method according to any one of claims 1 to 8, wherein: The target virtual image is used for display in the navigation map.
10. A device for generating a virtual image, comprising: an acquisition module, configured to acquire a mask image and three-dimensional modeling information for the target vehicle in response to the received style description information for the target vehicle; Each mask region in the mask image represents a location for adding a material image to the initial texture image for the target vehicle; the material image matches the style description information; a first extraction module, configured to extract, from the three-dimensional modeling information, the initial texture image and a mapping relationship between each two-dimensional coordinate in the initial texture image and each three-dimensional coordinate in the virtual image; A first generating module, the first generating module comprising: a first determining submodule, configured to determine a target material image according to the style description information; a second determining submodule, configured to determine a target position for adding the target material image in the initial texture image based on a mask area in the mask image; and A first generating submodule is configured to generate a target texture image by adding the target material image to the target position in the initial texture image by calling a generation model; and The second generating module is configured to generate a target virtual image by processing the target texture image based on the mapping relationship.
11. The device according to claim 10, wherein The target material images include multiple images; the mask areas include multiple areas; the second determination submodule includes: a first generating unit, configured to generate a correspondence between each of the target material images and each of the mask regions according to an attribute of each of the target material images and an attribute of each of the mask regions; and The first determining unit determines each target position of each target material image based on the corresponding relationship.
12. The device according to claim 10, wherein The target material images include multiple images; the mask areas include multiple images; and the first generating module further includes: a third determining submodule, configured to determine, based on the content of each target material image and in accordance with a predetermined correspondence, each mask region corresponding to each target material image; wherein the predetermined correspondence represents a correspondence between the material content and the mask region; The adjustment submodule is configured to adjust the attributes of each target material image according to the attributes of each mask area.
13. The device according to claim 10, wherein The first determining submodule includes: an extraction unit, configured to extract a first keyword from the style description information; and The second determining unit is configured to determine the target material image from the candidate material images based on a matching degree between the first keyword and a subject word of the candidate material image.
14. The device according to claim 10, wherein The device further comprises: a segmentation module, configured to perform semantic segmentation on the initial image to generate a first image of the target part; and The processing module is configured to generate the material image by processing the first image based on the received modification requirement information for the initial image.
15. The device according to claim 14, wherein The modification requirement information includes a target modification style; The processing module includes: a fourth determining submodule, configured to determine, according to the target modification style, a first modification material image and a first position for adding the first modification material image to the first image; as well as The second generating submodule is configured to generate a first material image by adding the first modified material image to the first position.
16. The device according to claim 15, wherein The demand information also includes the part to be modified and the modification object for the part to be modified; The processing module further includes: an extraction submodule, configured to extract a second position of the part to be modified from the first image; a fifth determining submodule, configured to determine a second modified material image corresponding to the modified object according to the modification style; The third generating submodule is configured to generate a second material image by adding the second modified material image at the second position.
17. The apparatus according to claim 16, further comprising: a second extraction module, configured to extract a first keyword from the style description information; a third extraction module, configured to extract a second keyword from the target modification style; The third generating module is configured to generate warning information for indicating a style conflict in response to a matching degree between the first keyword and the second keyword being less than a predetermined threshold.
18. The device according to any one of claims 10 to 17, wherein: The target virtual image is used for display in the navigation map.
19. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 9.
20. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 9.
21. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Virtual reality display method and device and storage medium
CN114092670A
Three-dimensional marking method and device for vehicle
CN116051832A