Virtual object generation method and device based on large model, equipment and intelligent agent
By planning and processing the video shooting trajectory of the initial image and processing it, three-dimensional virtual objects are generated, which solves the problems of contour deformation, texture color difference and detail loss in the generation of three-dimensional virtual objects in the prior art, and achieves high-precision and efficient three-dimensional virtual object generation.
Patent Information
- Application Number
- CN202510322753.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-18
- Publication Date
- 2025-06-24
AI Technical Summary
Existing three-dimensional virtual objects generated based on visual big models have problems such as contour deformation, texture chromatic aberration and loss of important contour details.
By planning the video shooting trajectory of the initial image related to the object, multiple shooting trajectory points are obtained, and the initial image and multiple shooting trajectory points are processed using a large model to generate a generated image sequence for the object, and an image is generated based on at least one frame in the generated image sequence to generate a three-dimensional virtual object representing the object.
It realizes that without real image acquisition, image acquisition of objects is simulated under multi-view conditions, and a three-dimensional virtual object that clearly characterizes the object shape, texture and other information is generated, reducing the impact of occlusions and contour distortion in the initial image, and improving the generation accuracy and efficiency of three-dimensional virtual objects.
Smart Images

Figure CN120198591A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technologies, particularly to technical fields such as computer vision, deep learning, large models, etc., and can be applied to scenarios such as AIGC (Artificial Intelligence Generative Content), digital twin, and metaverse. Background Art
[0002] With the rapid development of artificial intelligence technologies, in scenarios such as animation production, video production, and metaverse construction, three-dimensional virtual objects such as building bodies and vehicles in a three-dimensional space can be generated by using a visual large model to improve the production efficiency of virtual scenarios. Summary of the Invention
[0003] The present disclosure provides a method, apparatus, device, and agent for generating virtual objects based on a large model.
[0004] According to one aspect of the present disclosure, a method for generating a virtual object based on a large model is provided, including: performing video shooting trajectory planning based on an initial image related to an object to obtain a plurality of shooting trajectory points, where the shooting trajectory points represent the movement trajectories required for an image acquisition device to acquire images of the object for the initial image; using the large model to process the initial image and the plurality of shooting trajectory points to obtain a sequence of generated images for the object; and generating a three-dimensional virtual object representing the object based on at least one frame of the generated image sequence.
[0005] According to another aspect of the present disclosure, a device for generating a virtual object based on a large model is provided, including: a shooting trajectory point obtaining module for performing video shooting trajectory planning based on an initial image related to an object to obtain a plurality of shooting trajectory points, where the shooting trajectory points represent the movement trajectories required for an image acquisition device to acquire images of the object for the initial image; a generated image sequence obtaining module for using the large model to process the initial image and the plurality of shooting trajectory points to obtain a sequence of generated images for the object; and a three-dimensional virtual object generating module for generating a three-dimensional virtual object representing the object based on at least one frame of the generated image sequence.
[0006] According to another aspect of the present disclosure, an agent of artificial intelligence is provided, including: an input module for receiving input information; a processing module for determining a target task based on the input information received by the input module, determining a large model based on the target task, and obtaining output information by calling the large model to execute the method provided in the embodiments of the present disclosure; and an output module for outputting the output information obtained by the processing module.
[0007] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method provided according to the embodiments of the present disclosure.
[0008] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method provided according to the embodiments of the present disclosure.
[0009] According to another aspect of the present disclosure, there is provided a computer program product, including a computer program, where the computer program, when executed by a processor, implements the method provided according to the embodiments of the present disclosure.
[0010] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0012] Figure 1 Schematically shows an exemplary system architecture to which the method and apparatus for generating virtual objects based on a large model according to the embodiments of the present disclosure can be applied;
[0013] Figure 2 Schematically shows a flowchart of the method for generating virtual objects based on a large model according to the embodiments of the present disclosure;
[0014] Figure 3 Schematically shows an application scenario diagram of the method for generating virtual objects based on a large model according to the embodiments of the present disclosure;
[0015] Figure 4 Schematically shows a schematic diagram of the principle of the method for generating virtual objects based on a large model according to the embodiments of the present disclosure;
[0016] Figure 5 Schematically shows a schematic diagram of the principle of the method for generating virtual objects based on a large model according to another embodiment of the present disclosure;
[0017] Figure 6 Schematically shows a schematic diagram of the principle of the method for generating virtual objects based on a large model according to still another embodiment of the present disclosure;
[0018] Figure 7Schematically shown is a block diagram of a virtual object generation device based on a large model according to an embodiment of the present disclosure;
[0019] Figure 8 Schematically shown is a structural block diagram of an intelligent agent of artificial intelligence according to an embodiment of the present disclosure; and
[0020] Figure 9 Shown is a schematic block diagram of an example electronic device that can be used to implement the method for generating a virtual object based on a large model according to an embodiment of the present disclosure. Detailed implementation manners
[0021] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0022] In the technical solution of the present disclosure, the acquisition, storage, and application of the user's personal information involved all comply with the provisions of relevant laws and regulations, necessary confidentiality measures are taken, and public order and good customs are not violated.
[0023] The inventors found that with the rapid development of artificial intelligence technology, the image generation ability of image generation large models can be used to generate three-dimensional virtual objects, thereby improving the work efficiency of various production scenarios such as animation game production, video production, and metaverse scene construction. For example, a visual large model can be used to process street view images captured by an in-vehicle camera to generate three-dimensional virtual objects of buildings beside the street. However, the three-dimensional virtual objects generated based on image generation models such as visual large models have defects such as contour deformation, color difference in textures, and loss of important contour details.
[0024] Embodiments of the present disclosure provide a method, device, equipment, and intelligent agent for generating a virtual object based on a large model. The method for generating a virtual object based on a large model includes: performing video shooting trajectory planning based on an initial image related to an object to obtain a plurality of shooting trajectory points, where the shooting trajectory points represent the movement trajectory required for an image acquisition device to acquire images of the object for the initial image; using a large model to process the initial image and the plurality of shooting trajectory points to obtain a generated image sequence for the object; and generating a three-dimensional virtual object representing the object based on at least one frame of the generated image sequence.
[0025] According to an embodiment of the present disclosure, by performing video shooting trajectory planning on an initial image related to an object, the shooting trajectory points that an image acquisition device should move to convert the image acquisition perspective of the object are planned. Furthermore, a large model is used to process multiple shooting trajectory points and the initial image, so as to simulate a generated image sequence obtained by the image acquisition device for image acquisition of the object under multi-perspective conditions without using the image acquisition device for actual shooting, so that the generated image sequence can more clearly represent information such as the shape and texture of the object by converting the shooting perspective, reduce the occlusion effect of occluders on the object in the initial image, or reduce the influence of the display effect caused by contour distortion in the initial image. Thus, a three-dimensional virtual object representing the object can be quickly and accurately generated based on the accuracy of the object representation and relatively rich image information in at least one frame of the generated image sequence, realizing relatively accurate three-dimensional virtual object generation based on less time cost and equipment cost, and further improving the construction efficiency and accuracy of virtual scenes such as the metaverse and digital twins.
[0026] Figure 1 Schematically shows an exemplary system architecture to which the virtual object generation method and apparatus based on a large model according to an embodiment of the present disclosure can be applied.
[0027] It should be noted that Figure 1 The shown is only an example of the system architecture to which the embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios. For example, in another embodiment, the exemplary system architecture to which the virtual object generation method and apparatus based on a large model can be applied may include a terminal device, but the terminal device can implement the virtual object generation method and apparatus based on a large model provided by the embodiments of the present disclosure without interacting with the server.
[0028] As Figure 1 shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for a communication link between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired and / or wireless communication links, etc.
[0029] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as knowledge reading applications, web browser applications, search applications, instant messaging tools, email clients, and / or social platform software, etc. (only as examples).
[0030] The terminal devices 101, 102, and 103 can be various electronic devices with a display screen and supporting web browsing, including but not limited to smartphones, tablets, laptop computers, desktop computers, and so on.
[0031] The server 105 can be a server that provides various services. For example, it can be a background management server (only for example) that supports the content browsed by users using the terminal devices 101, 102, and 103. The background management server can analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0032] The server 105 can also be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, solving the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server 105 can also be a server of a distributed system, or a server combined with a blockchain.
[0033] It should be noted that the method for generating virtual objects based on a large model provided by the embodiments of the present disclosure can generally be executed by the server 105. Correspondingly, the device for generating virtual objects based on a large model provided by the embodiments of the present disclosure can generally be set in the server 105. The method for generating virtual objects based on a large model provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103 and / or the server 105. Correspondingly, the device for generating virtual objects based on a large model provided by the embodiments of the present disclosure can also be set in a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103 and / or the server 105.
[0034] For example, when a user is reading an e - book online, the terminal devices 101, 102, and 103 can obtain the target content in the e - book pointed to by the user's line of sight, and then send the obtained target content to the server 105. The server 105 analyzes the target content to determine the characteristic information of the target content; predicts the content that the user is interested in according to the characteristic information of the target content; and extracts the content that the user is interested in. Or a server or a server cluster capable of communicating with the terminal devices 101, 102, 103 and / or the server 105 analyzes the target content and finally realizes extracting the content that the user is interested in.
[0035] It should be understood, Figure 1The numbers of the terminal devices, networks, and servers in [it] are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, networks, and servers.
[0036] Figure 2 Schematically shows a flowchart of a large model-based virtual object generation method according to an embodiment of the present disclosure.
[0037] As Figure 2 shown, the large model-based virtual object generation method includes operations S210 to S240.
[0038] In operation S210, based on an initial image related to an object, video shooting trajectory planning is performed to obtain a plurality of shooting trajectory points.
[0039] In operation S220, the initial image and the plurality of shooting trajectory points are processed using a large model to obtain a sequence of generated images for the object.
[0040] In operation S230, based on at least one frame of the generated image sequence, a three-dimensional virtual object representing the object is generated.
[0041] According to an embodiment of the present disclosure, the object can be a real object or a real creature such as a building, a vehicle, a mechanical device, a person, a pet, etc. in the real world, and the initial image can be obtained by image acquisition of the object using an image acquisition device.
[0042] For example, the initial image can be a street view image obtained by an in-vehicle image acquisition device for image acquisition of a building object by the roadside.
[0043] It should be noted that the data collected in the embodiments of the present disclosure, including but not limited to the initial image and other data, are all obtained under the condition of obtaining the authorization of relevant users or institutions. And before collecting the data, the purpose of the data is clearly informed and necessary encryption measures are taken to avoid data leakage.
[0044] According to an embodiment of the present disclosure, the shooting trajectory points represent the movement trajectory required for the image acquisition device to acquire an image of the object for the initial image. When the image acquisition device acquires an image of the object, there may be image display defects such as a large area of the object area in the initial image being blocked by an occluder or the contour of the object area in the initial image being distorted due to reasons such as the acquisition perspective being close to the horizontal plane and the image acquisition device being close to the object. The plurality of shooting trajectory points can represent that the image acquisition device changes the perspective of image acquisition of the object by moving the shooting position, so that the image acquisition device can change the image acquisition angle or distance for the object at at least one shooting trajectory point.
[0045] According to an embodiment of the present disclosure, the shooting trajectory points may be coordinate points in a preset space, such as real coordinate points in a world coordinate system. Alternatively, the shooting trajectory points may also be coordinate points in other types of preset three-dimensional space coordinate systems. For example, the shooting trajectory points may also be virtual coordinate points in a preset virtual coordinate system.
[0046] According to an embodiment of the present disclosure, video shooting trajectory planning based on an initial image may include performing coordinate system conversion on the actual shooting position based on the camera parameters of the image acquisition device for the initial image, converting the actual shooting position where the initial image is taken to a preset coordinate system in a preset space, thereby obtaining the initial shooting position of the image acquisition device in the preset space, and performing shooting trajectory planning by moving the initial shooting position in the preset coordinate system to obtain multiple shooting trajectory points in the preset space. Alternatively, multiple preset spaces may also be spaces constructed based on the world coordinate system, and the shooting trajectory points may be coordinate points in the world coordinate system. The embodiments of the present disclosure do not limit the specific type of the coordinate system for the shooting trajectory points, and those skilled in the art may select according to actual needs.
[0047] It should be understood that the actual shooting position of the initial image may be a coordinate in the world coordinate system. The actual shooting position may be determined based on the positioning information of the image acquisition device, or may also be processed by a camera calibration method for the camera parameters and the object coordinate points of the object in the world coordinate system to obtain the actual shooting position. The embodiments of the present disclosure do not limit the specific manner of determining the actual shooting position of the initial image.
[0048] According to an embodiment of the present disclosure, multiple shooting trajectory points may be multiple consecutive coordinate points in the shooting trajectory, or multiple shooting trajectory points may also be coordinate points spaced apart in the shooting trajectory. The embodiments of the present disclosure do not limit the specific coordinate type of the multiple shooting trajectory points, as long as it can enable the image acquisition device to convert the shooting perspective to perform image acquisition on the object. The image acquisition device moves from the initial shooting position in the preset space to multiple shooting trajectory points to perform image acquisition on the object through different shooting perspectives, which can reduce the area of the object region in the collected image blocked by the occluder and reduce the shape distortion of the lines in the collected image, thereby improving the display effect of the three-dimensional virtual object and the quality of the three-dimensional image.
[0049] According to an embodiment of the present disclosure, a large model may refer to a deep learning model with a large number of model parameters. A large model usually contains hundreds of millions, tens of billions, hundreds of billions, trillions or even more than one quadrillion model parameters. The large model may include visual large models or multimodal large models with video generation capabilities such as Video Diffusion Models algorithm and CogVideo algorithm. The large model involved in the embodiments of the present disclosure may be a general large model, or may also be an expert large model obtained after fine-tuning based on a specified inference task. The embodiments of the present disclosure do not limit this. The large model can be used to process information in any modality such as images, texts, videos, and audios.
[0050] According to an embodiment of the present disclosure, using a large model to process an initial image and multiple shooting trajectory points to obtain a generated image sequence for the object may include generating prompt information for controlling the large model to understand the video generation task based on the multiple shooting trajectory points. The large model can execute an image generation task based on the prompt information and image acquisition parameters such as the camera pose of the image acquisition device for acquiring the initial image, so that the large model can simulate image acquisition of an object in the real world by moving the image acquisition device according to the moving positions of the multiple shooting trajectory points, obtain the generated images corresponding to the respective multiple shooting trajectory points, and thus obtain the generated image sequence. In this way, under the condition of not using the image acquisition device to acquire images of the object, the large model can be used to simulate the image acquisition device moving from the initial shooting position in the preset space to multiple shooting trajectory points for image acquisition, and then simulate image acquisition of the object from different shooting perspectives by using the image acquisition device, so as to reduce the area of the object region in the generated image blocked by the occluder and reduce the shape distortion of the lines in the generated image, thereby improving the display effect and three-dimensional image quality for generating the three-dimensional virtual object.
[0051] According to an embodiment of the present disclosure, generating a three-dimensional virtual object representing an object based on at least one frame of generated image may include using a three-dimensional visual large model to process at least one frame of generated image to obtain a three-dimensional virtual object in a preset three-dimensional space. For example, a visual large model based on algorithms such as the Diffusion algorithm with three-dimensional image generation capabilities can be used to process the generated image to obtain a three-dimensional virtual object.
[0052] In one example, a visual large model can be used to process the last frame of the generated image sequence to obtain a three-dimensional virtual object. The three-dimensional virtual object can be, for example, a three-dimensional image representing a building.
[0053] In one example, a vision large model can be utilized to process multiple generated images in a generated image sequence to obtain a three-dimensional virtual object. This enables the vision large model to more accurately generate a three-dimensional virtual object based on the image acquisition perspectives of object regions in different generated images for objects in the real world.
[0054] According to an embodiment of the present disclosure, based on an initial image related to an object, video shooting trajectory planning is performed to obtain multiple shooting trajectory points, including: detecting the shooting position of the initial image to obtain an initial shooting position in a preset space; in response to a position movement request for the initial shooting position, moving the initial shooting position based on the movement trajectory information carried in the position movement request to obtain multiple shooting trajectory points in the preset space.
[0055] According to an embodiment of the present disclosure, detecting the shooting position of the initial image may include performing coordinate system conversion on the actual acquisition position where the image acquisition device acquires the initial image in the real world based on camera parameters such as the extrinsic parameters of the camera and the camera pose of the image acquisition device to obtain the initial shooting position in a preset coordinate system for the preset space.
[0056] In one example, the preset coordinate system may be a virtual space coordinate system constructed by a computing device such as a server. Conversion matrices between each coordinate system can be obtained based on relevant parameters such as the virtual space coordinate system, the world coordinate system, and the camera coordinate system of the image acquisition device. The actual acquisition position is subjected to coordinate system conversion through the conversion matrix to obtain the initial shooting position in the preset space.
[0057] According to an embodiment of the present disclosure, the position movement request may be generated by a user through an interaction operation, and the position movement request may be used to move the initial shooting position in the preset space to generate multiple shooting trajectory points in the movement trajectory.
[0058] According to an embodiment of the present disclosure, the movement trajectory information carried in the position movement request may include at least one of a movement trajectory direction and a trajectory point position. The movement direction may be a direction pointed from the initial shooting position as the starting position in the preset space, or may also be a direction pointed from any shooting trajectory point as the starting position in the preset space. The trajectory point position may be any coordinate point position in the preset space. The trajectory point position may be used to indicate the coordinates of the shooting trajectory point, or may be used to generate a smooth shooting trajectory by processing the movement direction and the trajectory point position based on a preset normalization algorithm or other algorithms for generating trajectories, and multiple shooting trajectory points are determined from the shooting trajectory.
[0059] According to an embodiment of the present disclosure, the virtual object generation method based on a large model may further include: displaying the initial shooting position in an interaction interface; and determining a position movement request in response to an interaction operation for the initial shooting position.
[0060] According to an embodiment of the present disclosure, the interactive interface can be used to display display information such as lines, coordinate points, textures, etc. in a preset space. The initial shooting position can be displayed in the interactive interface in the form of coordinate points.
[0061] According to an embodiment of the present disclosure, the interactive operation for the initial shooting position can be of any type. For example, it can be an information input operation, a voice input operation, a drag operation, etc. The embodiments of the present disclosure do not limit the specific type of the interactive operation, as long as it can generate movement trajectory information.
[0062] Figure 3 Fig. schematically shows an application scenario diagram of a virtual object generation method based on a large model according to an embodiment of the present disclosure.
[0063] As Figure 3 shown, the initial image T301 and the initial shooting position N301 in the preset space are shown in the interactive interface 300. The image detection result includes an object area T311 and an occlusion area T312. The initial image T301 can represent the initial shooting position N301 of the image acquisition device in the preset coordinate system of the preset space, and is obtained by image acquisition of the houses and trees by the roadside based on the first perspective V301. The user can determine the first trajectory point position, the second trajectory point position, and the third trajectory point position based on the interactive operation. At the same time, the first movement direction, the second movement direction, and the third movement direction starting from the initial shooting position N301, the first trajectory point position, and the second trajectory point position can also be determined based on the interactive operation. The server can use the first trajectory point position, the second trajectory point position, and the third trajectory point position, as well as the first movement direction, the second movement direction, and the third movement direction carried in the position movement request. Based on the movement trajectory information carried in the position movement request, the initial shooting position N301 is moved in the preset coordinate system of the preset space to generate a plurality of shooting trajectory points. The plurality of shooting trajectory points can include a first trajectory point N311, a second trajectory point N312, and a third trajectory point N313. At the same time, the first trajectory segment B311, the first trajectory point N311, the second trajectory segment B312, the second trajectory point N312, the third trajectory segment B313, and the third trajectory point N313 in the preset space can be displayed in the interactive interface to display the simulated shooting positions and shooting perspectives corresponding to the respective shooting trajectory points in the generated image, so as to facilitate the user to correct the shooting trajectory points.
[0064] It should be noted that Figure 3The first trajectory point N311, the second trajectory point N312, and the third trajectory point N313 shown in [Figure 0] can be generated by moving the initial shooting position N301 along the positive Y-axis, the negative X-axis, and the negative Z-axis of the preset coordinate system. By means of multiple shooting trajectory points, it is possible to simulate the way that the image acquisition device approaches the house in the preset space and raises the shooting angle for shooting, so that the generated image representing the shooting along the second viewing angle V302 based on the third trajectory point N313 can reduce the occluded image area of the occluding object, the tree, and display a larger area of the house image area. In this way, a high-quality three-dimensional virtual house object can be generated at least based on the generated image corresponding to the third trajectory point N313, improving the display accuracy and display effect of the three-dimensional virtual object for the house building in the real world.
[0065] According to an embodiment of the present disclosure, based on an initial image related to an object, video shooting trajectory planning is performed to obtain multiple shooting trajectory points, including: performing target detection on the occluding object area of the initial image to obtain an occluding object detection result; performing shooting position planning based on the occluding object detection result to obtain a first target position in the preset space; and determining multiple shooting trajectory points based on the first target position.
[0066] According to an embodiment of the present disclosure, the occluding object area represents an occluding object of the object. For example, it can be an image area representing the area of the object occluding the object in the initial image. The occluding object detection result can be information such as the position, size, and shape of the occluding object area in the initial image. The occluding object detection result can be a detection box representing the occluding object obtained after processing the initial image based on a target detection algorithm.
[0067] According to an embodiment of the present disclosure, the first target position is used to reduce the occluded area of the object in the generated image. Representing the first target position of the image acquisition device in the preset space for image acquisition of the object can reduce the occluded area of the occluding object area for the object area in the acquired image. Performing shooting position planning based on the occlusion detection result can include processing the object area and the occluding object detection result in the initial image based on a trained neural network model to obtain information such as the direction, distance, and position that the image acquisition device should move to avoid the occluding object for shooting the object, and then performing coordinate transformation according to the information such as the direction, distance, and position to be moved to determine the first target position in the preset space. The first target position can include one or more.
[0068] According to an embodiment of the present disclosure, determining a plurality of shooting trajectory points based on a first target position may include performing trajectory planning according to the positional relationship between an initial shooting position and the first target position in a preset space to obtain a movement trajectory from the initial shooting position to the first target position, and any one or more trajectory points in the movement trajectory may be used as shooting trajectory points. Alternatively, a plurality of first target positions may also be determined as a plurality of shooting trajectory points. The embodiment of the present disclosure does not limit the specific setting method for determining a plurality of shooting trajectory points based on the first target position, as long as it can represent that the target acquisition device performs image acquisition on an object based on at least one of the plurality of shooting trajectory points, and the area of the object occluded in the image can be reduced.
[0069] According to an embodiment of the present disclosure, performing shooting position planning based on an occlusion detection result to obtain a first target position in a preset space includes: determining an initial positional relationship between an occlusion area and an object area representing an object in an initial image based on the occlusion position; determining a target movement direction based on the initial positional relationship; and determining the first target position based on the target movement direction and the initial shooting position for the initial image.
[0070] According to an embodiment of the present disclosure, the initial positional relationship may represent information such as the distance and direction between a specified area point of the occlusion area and a specified area point of the object area in an image coordinate system or other specified coordinate systems, which characterizes the position where the occlusion in the initial image occludes the object. The target movement direction characterizes the direction in which the image acquisition device in the preset space should move. It can be understood that the target movement direction indicates the direction in which the image acquisition device in the preset space should move away from the initial shooting position in order to avoid occlusion by an obstacle for image acquisition.
[0071] According to an embodiment of the present disclosure, determining the first target position based on the target movement direction and the initial shooting position for the initial image may include processing the target movement direction and the initial shooting position based on a deep learning algorithm to implement video shooting trajectory planning for the image acquisition device in the preset space, so as to obtain one or more trajectory points in the video shooting trajectory as the first target position. It should be noted that the target movement direction may include one or more, and each target movement direction may correspond to one or more first target positions. A plurality of shooting trajectory points may be determined from the first target positions corresponding to any target movement direction.
[0072] Figure 4 Schematically shows a schematic diagram of the principle of a virtual object generation method based on a large model according to an embodiment of the present disclosure.
[0073] As Figure 4As shown, the initial image T410 may include an object region T411 and an occluder region T412. The initial positional relationship between the occluder center position of the occluder region T412 and the object center position of the object region T411 in the initial image T410 may indicate that the lower right region of the object region T411 is occluded by the occluder region T412. Based on the initial positional relationship, the target movement direction 401 in the preset space may have a positive Y-axis component and a negative X-axis component. By using a trajectory planning algorithm to process the target movement direction 401 and the initial shooting position N401, the video shooting movement trajectory B410 in the preset space can be obtained. In this way, any trajectory point in the video shooting movement trajectory B410 can be used as a shooting trajectory point. Among them, the first shooting trajectory point N411 may represent a shooting perspective of the image acquisition device that is higher than the initial shooting position N401 and avoids the occlusion of the house by the trees to acquire an image of the house. In this way, at least based on a vision large model, a plurality of sequentially arranged shooting trajectory points in the video shooting movement trajectory B410 can be processed to generate a first generated image of the house being shot by the first shooting trajectory point N411 of the simulated image acquisition device in the preset space. Thus, a three-dimensional virtual object can be generated based on the first generated image, and a three-dimensional image that can accurately represent important information such as the contour, shape, texture, window size, and number of windows of the house can be obtained, improving the generation accuracy of the three-dimensional virtual object.
[0074] According to an embodiment of the present disclosure, based on an initial image related to an object, video shooting trajectory planning is performed, and obtaining a plurality of shooting trajectory points includes: performing distortion detection on a reference object region of the initial image to obtain a distortion detection result; performing shooting position planning based on the distortion detection result to obtain a second target position in a preset space; and determining a plurality of shooting trajectory points based on the second target position.
[0075] According to an embodiment of the present disclosure, the distortion detection result may represent the degree of line contour distortion of the reference object region in the initial image. The reference object region may be an image region in the initial image that represents any reference object. The reference object may include reference objects with a preset line shape such as lane lines and street lamp poles. Or, the reference object may also be an occluded object or an occluder object. The embodiment of the present disclosure does not limit the setting manner of the reference object. The neural network algorithm may be used to calculate the difference in the shape contour represented by the reference object region to obtain the line contour difference information between the contour of the reference object represented by the reference object region and the actual contour shape of the reference object, and the distortion detection result may be determined based on the line contour difference information.
[0076] According to an embodiment of the present disclosure, the distortion detection result may characterize any distortion type such as barrel distortion, pincushion distortion, etc., but is not limited thereto. The distortion detection result may also represent the distortion degree of the line contour of the reference object area in the initial image.
[0077] According to an embodiment of the present disclosure, the second target position is used to reduce the distortion degree of the reference object in the generated image. Based on the distortion type or distortion degree characterized by the distortion detection result, the initial shooting position of the image acquisition device can be moved to reduce the distortion degree of the reference object in the image acquired by the image acquisition device in the preset space. For example, when the distortion detection result indicates that barrel distortion occurs in the reference object area, the object can be image-acquired at a second target position away from the object relative to the initial shooting position in the preset space, so that the barrel distortion degree of the acquired image is reduced. For another example, when the distortion detection result indicates that pincushion distortion occurs in the reference object area, the object can be image-acquired at a second target position closer to the object relative to the initial shooting position in the preset space, so that the pincushion distortion degree of the acquired image is reduced.
[0078] According to an embodiment of the present disclosure, the distance between the second target position and the initial shooting position in the preset space can be determined based on the distortion type and distortion degree characterized by the distortion detection result. For example, the distance between the second target position and the initial shooting position is in a proportional relationship with the distortion degree, so as to more accurately reduce the influence of this distortion type in the generated image based on the second target position.
[0079] According to an embodiment of the present disclosure, the number of second target positions can be one or more. Determining a plurality of shooting trajectory points based on the second target position may include determining shooting trajectory points from the connection lines from the initial shooting position to one or more second target positions, or may also include performing normalization processing based on the initial shooting position and a plurality of second target positions to obtain a video shooting trajectory in the preset space, and determining a plurality of shooting trajectory points from the video shooting trajectory. The plurality of shooting trajectory points determined based on the second target position can instruct the visual large model for generating video data to move along the shooting trajectory for reducing the distortion degree to generate a sequence of generated images with reduced image distortion degree. In this way, high-quality three-dimensional virtual objects can be generated based on the generated images with lower distortion degree, avoiding abnormal situations such as line distortion in the three-dimensional virtual objects.
[0080] Figure 5 Schematically shows a schematic diagram of the principle of a virtual object generation method based on a large model according to another embodiment of the present disclosure.
[0081] As Figure 5As shown, the initial image T510 may include an object region T511, and the reference object represented by the object region T511 may be a house object. By detecting the initial image T510, it can be obtained that the distortion detection result indicates that the object region T511 has pincushion distortion, and the degree of pincushion distortion is the first-level distortion degree. By determining the distance corresponding to the first-level distortion degree and the direction approaching the object from the initial shooting position N501, the second target position N511 can be determined. In addition, the initial shooting position N501 can also be moved to the trajectory points in the second video shooting trajectory B510 of the second target position N511 as other second target positions in addition to the second target position N511. In this way, by determining multiple second target positions as multiple shooting trajectory points, a vision large model with video generation capabilities can be controlled to generate generated images corresponding to each shooting trajectory point, obtaining a sequence of generated images, so that any frame of the generated images in the sequence of generated images can reduce the pincushion distortion degree of the object region, and further improve the image quality of the three-dimensional virtual object image generated based on the generated images and the display accuracy by improving the image quality of the generated images.
[0082] According to an embodiment of the present disclosure, based on an initial image related to an object, performing video shooting trajectory planning to obtain multiple shooting trajectory points may further include: performing target detection on the occluder region of the initial image to obtain an occluder detection result; performing shooting position planning based on the occluder detection result to obtain a first target position in a preset space; performing distortion detection on the reference object region of the initial image to obtain a distortion detection result; performing shooting position planning based on the distortion detection result to obtain a second target position in the preset space; and determining multiple shooting trajectory points based on the first target position and the second target position.
[0083] According to an embodiment of the present disclosure, determining multiple shooting trajectory points based on the first target position and the second target position may include performing trajectory fitting on the first target position and the second target position in the preset space to obtain a fitted trajectory, and determining multiple shooting trajectory points from the fitted trajectory. However, it is not limited thereto. Multiple shooting trajectory points may also be determined from the first target position and the second target position based on other methods. For example, determining multiple shooting trajectory points from multiple first target positions and multiple second target positions. The embodiment of the present disclosure does not limit the specific method for determining multiple shooting trajectory points, as long as the first target position and the second target position are used.
[0084] According to an embodiment of the present disclosure, by determining a plurality of shooting trajectory points based on a first target position and a second target position, it is possible to control a vision large model with a video generation function to simulate image acquisition of an object through the perspectives and camera poses of image acquisition devices corresponding to the plurality of shooting trajectory points using the plurality of shooting trajectory points and an initial image, so that the generated image sequence can simulate a generated video obtained by video shooting of the object in accordance with the camera movement mode characterized by the plurality of shooting trajectory points, reduce the degree of image distortion of the generated images in the generated image sequence, and reduce the occlusion area of the occlusion object on the object area, thereby improving the image quality of the generated images. In this way, a multi-modal large model with the ability to generate three-dimensional objects can generate a three-dimensional virtual object that can more accurately represent attributes such as the contour and texture of the object by processing at least one frame of the generated image.
[0085] According to an embodiment of the present disclosure, using a large model to process an initial image and a plurality of shooting trajectory points to obtain a generated image sequence of a recorded object includes: using a position feature extraction network of the large model to process N shooting trajectory points to obtain the position features of each of the N shooting trajectory points; using an image feature extraction network of the large model to process the initial image to obtain an initial image feature; using an image generation network of the large model to process the initial image feature and the first position feature of the first shooting trajectory point to obtain a first generated image; using the image generation network to process the (n - 1)th image feature of the (n - 1)th generated image in the generated image sequence and the nth position feature of the nth shooting trajectory point to obtain the nth generated image; and in the case of N = n, obtaining a generated image sequence based on the N generated images.
[0086] According to an embodiment of the present disclosure, N is an integer greater than 1, N ≥ n > 1, and the (n - 1)th image feature is determined according to the (n - 1)th generated image.
[0087] According to an embodiment of the present disclosure, the position feature extraction network and the image feature extraction network can be determined based on any type of deep learning algorithm. For example, the position feature extraction network and the image feature extraction network can be constructed based on the encoder layer of the attention network algorithm. The image generation network can be constructed based on a large model with video generation ability. For example, the image generation network can be constructed based on the Transformer algorithm.
[0088] Figure 6 Schematically shows a schematic diagram of the principle of a virtual object generation method based on a large model according to another embodiment of the present disclosure.
[0089] As Figure 6As shown in the figure, a large model with video generation capabilities may include a position feature extraction network M611, an image feature extraction network M612, and an image generation network M613. The video shooting trajectory B610 carrying N shooting trajectory points is input into the position feature extraction network M611, and the first position feature, the second position feature... up to the Nth position feature corresponding to each of the N shooting trajectory points are output. The first position feature, the second position feature... up to the Nth position feature respectively correspond to the first shooting trajectory point, the second shooting trajectory point... up to the Nth shooting trajectory point arranged in sequence in the video shooting trajectory B610. The initial image T601 is input into the image feature extraction network M612, and the initial image feature is output. The initial image feature and the first position feature are input into the image generation network, and the first generated image is output. The first generated image may represent an image obtained by the image acquisition device capturing an object at the first trajectory point position in the preset coordinate system.
[0090] As Figure 6 shown in the figure, the hidden feature of the first generated image obtained by the image generation network M613 is used as the first image feature. The first image feature and the second position feature are input into the image generation network M613, and the second generated image is output. The second generated image may represent an image obtained by the image acquisition device capturing an object at the second trajectory point position in the preset coordinate system. The hidden feature of the (N - 1)th generated image obtained by the image generation network M613 is used as the (N - 1)th image feature. The (N - 1)th image feature and the Nth position feature are input into the image generation network M613, and the Nth generated image is output. The Nth generated image may represent an image obtained by the image acquisition device capturing an object at the Nth trajectory point position in the preset coordinate system. Based on the first generated image to the Nth generated image arranged in sequence as a generated image sequence, the generated image sequence can be used as a generated video obtained by the image acquisition device shooting an object according to multiple shooting trajectory points in the preset space. The Nth generated image in the generated image sequence may have a higher shooting angle to avoid the image information of the object area being blocked by the occlusion area. At the same time, the Nth generated image can simulate the image acquisition device shooting the object at a more appropriate distance, thereby reducing the degree of image distortion. Inputting the Nth generated image into the multimodal large model can output a three-dimensional virtual object representing the object, so as to realize the generation of three-dimensional objects for objects such as buildings and bridges in the street view, and improve the display accuracy of the three-dimensional virtual object for the object.
[0091] It should be noted that the image generation network can also input a prompt (prompt) for other prompt information such as a conversion matrix, so that the image generation network can execute the video generation function based on the control of the prompt and generate a generated image representing the object being shot at the shooting trajectory point more accurately.
[0092] Figure 7 A block diagram of a virtual object generation device based on a large model according to an embodiment of the present disclosure is schematically shown.
[0093] As Figure 7 shown, the virtual object generation device 700 based on a large model includes a shooting trajectory point acquisition module 710, a generated image sequence acquisition module 720, and a three-dimensional virtual object generation module 730.
[0094] The shooting trajectory point acquisition module 710 is configured to perform video shooting trajectory planning based on an initial image related to an object to obtain a plurality of shooting trajectory points, and the shooting trajectory points represent the movement trajectory required for an image acquisition device to acquire an image of the object for the initial image.
[0095] The generated image sequence acquisition module 720 is configured to use a large model to process the initial image and the plurality of shooting trajectory points to obtain a generated image sequence for the object.
[0096] The three-dimensional virtual object generation module 730 is configured to generate a three-dimensional virtual object representing the object based on at least one frame of the generated image sequence.
[0097] According to an embodiment of the present disclosure, the shooting trajectory point acquisition module 710 includes: an occlusion detection result acquisition unit, a first target position acquisition unit, and a first determination unit.
[0098] The occlusion detection result acquisition unit is configured to perform target detection on the occlusion area of the initial image to obtain an occlusion detection result, and the occlusion area represents an occluder of the occluded object.
[0099] The first target position acquisition unit is configured to perform shooting position planning based on the occlusion detection result to obtain a first target position in a preset space, where the first target position is used to reduce the occluded area of the object in the generated image.
[0100] The first determination unit is configured to determine a plurality of shooting trajectory points based on the first target position.
[0101] According to an embodiment of the present disclosure, the occlusion detection result includes the occlusion position of the occlusion area in the initial image.
[0102] According to an embodiment of the present disclosure, the first target position acquisition unit includes: a first determination subunit, a second determination subunit, and a third determination subunit.
[0103] The first determination subunit is configured to determine an initial position relationship between the occlusion area in the initial image and the object area representing the object based on the occlusion position.
[0104] A second determination subunit, configured to determine a target movement direction based on an initial position relationship, where the target movement direction represents the direction in which an image acquisition device in a preset space should move.
[0105] A third determination subunit, configured to determine a first target position based on the target movement direction and an initial shooting position of an initial image, where the initial shooting position represents a position in a preset space.
[0106] According to an embodiment of the present disclosure, the shooting trajectory point obtaining module 710 includes: a distortion detection result obtaining unit, a second target position obtaining unit, and a second determination unit.
[0107] The distortion detection result obtaining unit is configured to perform distortion detection on a reference object area of an initial image to obtain a distortion detection result.
[0108] The second target position obtaining unit is configured to perform shooting position planning based on the distortion detection result to obtain a second target position in a preset space, where the second target position is used to reduce the distortion degree of a reference object in a generated image.
[0109] The second determination unit is configured to determine a plurality of shooting trajectory points based on the second target position.
[0110] According to an embodiment of the present disclosure, the shooting trajectory point obtaining module 710 includes: an occlusion detection result obtaining unit, a first target position determination unit, a distortion detection result obtaining unit, a second target position obtaining unit, and a third determination unit.
[0111] The occlusion detection result obtaining unit is configured to perform target detection on an occlusion area of an initial image to obtain an occlusion detection result, where the occlusion area represents an occluder of an occlusion object.
[0112] The first target position determination unit is configured to perform shooting position planning based on the occlusion detection result to obtain a first target position in a preset space, where the first target position is used to reduce the occluded area of an object in a generated image.
[0113] The distortion detection result obtaining unit is configured to perform distortion detection on a reference object area of an initial image to obtain a distortion detection result.
[0114] The second target position obtaining unit is configured to perform shooting position planning based on the distortion detection result to obtain a second target position in a preset space, where the second target position is used to reduce the distortion degree of a reference object in a generated image.
[0115] The third determination unit is configured to determine a plurality of shooting trajectory points based on the first target position and the second target position.
[0116] According to an embodiment of the present disclosure, the shooting trajectory point acquisition module 710 includes: an initial shooting position acquisition unit and a shooting trajectory point acquisition unit.
[0117] The initial shooting position acquisition unit is configured to detect the shooting position of the initial image to obtain the initial shooting position in the preset space.
[0118] The shooting trajectory point acquisition unit is configured to, in response to a position movement request for the initial shooting position, move the initial shooting position based on the movement trajectory information carried in the position movement request to obtain a plurality of shooting trajectory points in the preset space, where the movement trajectory information includes at least one of a movement trajectory direction and a trajectory point position.
[0119] According to an embodiment of the present disclosure, the virtual object generation device 700 based on a large model further includes: a display module and a request determination module.
[0120] The display module is configured to display the initial shooting position in the interaction interface.
[0121] The request determination module is configured to determine a position movement request in response to an interaction operation for the initial shooting position.
[0122] According to an embodiment of the present disclosure, the generated image sequence acquisition module 720 includes: a position feature acquisition unit, an initial image feature acquisition unit, a first generated image acquisition unit, an nth generated image acquisition unit, and a generated image sequence acquisition unit.
[0123] The position feature acquisition unit is configured to process N shooting trajectory points by using the position feature extraction network of the large model to obtain the position features of each of the N shooting trajectory points, where N is an integer greater than 1.
[0124] The initial image feature acquisition unit is configured to process the initial image by using the image feature extraction network of the large model to obtain the initial image feature.
[0125] The first generated image acquisition unit is configured to process the initial image feature and the first position feature of the first shooting trajectory point by using the image generation network of the large model to obtain the first generated image.
[0126] The nth generated image acquisition unit is configured to process the (n - 1)th image feature of the (n - 1)th generated image in the generated image sequence and the nth position feature of the nth shooting trajectory point by using the image generation network to obtain the nth generated image, where N ≥ n > 1, and the (n - 1)th image feature is determined according to the (n - 1)th generated image.
[0127] The generated image sequence acquisition unit is configured to, when N = n, obtain a generated image sequence based on the N generated images.
[0128] Figure 8 Schematically shows a structural block diagram of an agent of artificial intelligence according to an embodiment of the present disclosure.
[0129] In an embodiment of the present disclosure, as Figure 8 shown, the AI agent 800 may include an input module 810, a processing module 820, and an output module 830.
[0130] The input module 810 is configured to receive input information.
[0131] The processing module 820 is configured to determine a target task based on the input information received by the input module, determine a large model based on the target task, execute an interaction method based on the large model provided according to an embodiment of the present disclosure by invoking the large language model, or obtain output information by invoking the large model to execute a virtual object generation method based on the large model provided according to an embodiment of the present disclosure.
[0132] The output module 830 is configured to output the output information obtained by the processing module.
[0133] According to an embodiment of the present disclosure, the input module 810 is responsible for receiving or sensing information such as queries, requests, instructions, signals, or data from the outside world (such as users or the external environment), and converting it into a format that the AI agent 800 can understand and process. The input module 810 is the primary link for the AI agent 800 to interact with the outside world, enabling the AI agent 800 to efficiently and accurately obtain necessary "sensory" information from the outside world and respond to this information.
[0134] In an example, the input module 810 may input the initial image, shooting trajectory points, generated images, etc. described above.
[0135] In an example, the processing module 820 is the core support for the AI agent 800's ability to handle complex tasks. The processing module 820 may execute the virtual object generation method based on the large model described above.
[0136] In an example, the performance of the processing module 820 may be closely related to the large model on which the AI agent 800 is based. To fully utilize the capabilities of the large model, the internal structure of the processing module 820 can be designed to be highly configurable and extensible to cope with various different types of tasks and requirements in real scenarios.
[0137] In an example, after the AI agent 800 obtains the initial image and multiple shooting trajectory points, the processing module 820 may use a vision large model to process the initial image and multiple shooting trajectory points to obtain a generated image sequence, call a three-dimensional image generation model to process at least one frame of the generated image to obtain a three-dimensional virtual object, and transmit the three-dimensional virtual object to the output module 830.
[0138] It can be understood that although large language models have excellent language understanding and generation capabilities, like humans, the tasks they can solve without any tools are very limited. When the AI agent 800 is given the ability to call tools, tasks such as performing mathematical operations with the help of a calculator, conducting data analysis with the help of Python, and obtaining weather forecasts with the help of a search engine can be achieved.
[0139] In the example, the output module 830 can output the generated image sequence or three-dimensional virtual object described above.
[0140] The AI agent 800 according to the embodiments of the present disclosure can simply and effectively improve the degree of intelligence, and enhance flexibility and versatility.
[0141] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0142] According to an embodiment of the present disclosure, an electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method as described above.
[0143] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the method as described above.
[0144] According to an embodiment of the present disclosure, a computer program product includes a computer program, and the computer program implements the method as described above when executed by a processor.
[0145] Figure 9 A schematic block diagram of an example electronic device that can be used to implement the large model-based virtual object generation method according to the embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described herein and / or claimed.
[0146] As Figure 9As shown, device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0147] Multiple components in the device 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, an optical disc, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0148] The computing unit 901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 901 executes the various methods and processes described above, such as the method for generating virtual objects based on a large model. For example, in some embodiments, the method for generating virtual objects based on a large model can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the method for generating virtual objects based on a large model described above can be executed. Alternatively, in other embodiments, the computing unit 901 can be configured to execute the method for generating virtual objects based on a large model in any other appropriate manner (e.g., by means of firmware).
[0149] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGA), application specific integrated circuits (ASIC), application specific standard products (ASSP), systems on chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.
[0150] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine, or entirely on the remote machine or server.
[0151] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0152] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0153] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.
[0154] A computer system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client - server relationship is created by computer programs running on the respective computers and having a client - server relationship with each other. The server can be a cloud server, a server of a distributed system, or a server incorporating blockchain.
[0155] It should be understood that the various forms of the processes shown above can be reordered, added to, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.
[0156] The above - mentioned specific implementation manners do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of this disclosure shall be included within the protection scope of this disclosure.
Claims
1. A method for generating a virtual object based on a large model, comprising: Based on an initial image related to the object, video shooting trajectory planning is performed to obtain a plurality of shooting trajectory points, wherein the shooting trajectory points represent movement trajectories required for an image acquisition device to acquire an image of the object with respect to the initial image; Processing the initial image and the plurality of shooting trajectory points using a large model to obtain a generated image sequence for the object; as well as A three-dimensional virtual object representing the object is generated based on at least one frame of the generated image sequence.
2. The method according to claim 1, wherein: The video shooting trajectory planning is performed based on the initial image related to the object to obtain multiple shooting trajectory points, including: Performing target detection on an occluder region of the initial image to obtain an occluder detection result, wherein the occluder region represents an occluder that occludes the object; Performing shooting position planning based on the obstruction detection result to obtain a first target position in a preset space, wherein the first target position is used to reduce the obstructed area of the object in the generated image; and A plurality of shooting trajectory points are determined based on the first target position.
3. The method according to claim 2, wherein: The occluder detection result includes the occluder position of the occluder area in the initial image; The step of performing shooting position planning based on the obstruction detection result to obtain a first target position in a preset space includes: Based on the position of the occluder, determining an initial positional relationship between the occluder region and an object region representing the object in the initial image; Determining a target moving direction based on the initial position relationship, the target moving direction representing a direction in which the image acquisition device in the preset space should move; and The first target position is determined based on the target moving direction and an initial shooting position for the initial image, wherein the initial shooting position represents a position in the preset space.
4. The method according to claim 1, wherein: The video shooting trajectory planning is performed based on the initial image related to the object to obtain multiple shooting trajectory points, including: Performing distortion detection on a reference object region of the initial image to obtain a distortion detection result; Performing shooting position planning based on the distortion detection result to obtain a second target position in a preset space, wherein the second target position is used to reduce the degree of distortion of the reference object in the generated image; and A plurality of shooting trajectory points are determined based on the second target position.
5. The method according to claim 1, wherein: The video shooting trajectory planning is performed based on the initial image related to the object to obtain multiple shooting trajectory points, including: Performing target detection on an occluder region of the initial image to obtain an occluder detection result, wherein the occluder region represents an occluder that occludes the object; Performing shooting position planning based on the occlusion detection result to obtain a first target position in a preset space, wherein the first target position is used to reduce the occluded area of the object in the generated image; Performing distortion detection on a reference object region of the initial image to obtain a distortion detection result; Performing shooting position planning based on the distortion detection result to obtain a second target position in the preset space, wherein the second target position is used to reduce the degree of distortion of the reference object in the generated image; A plurality of shooting trajectory points are determined based on the first target position and the second target position.
6. The method according to claim 1, wherein: The video shooting trajectory planning is performed based on the initial image related to the object to obtain multiple shooting trajectory points, including: Performing shooting position detection on the initial image to obtain an initial shooting position in a preset space; In response to a position movement request for the initial shooting position, the initial shooting position is moved based on the movement trajectory information carried in the position movement request to obtain a plurality of shooting trajectory points in the preset space, wherein the movement trajectory information includes at least one of a movement trajectory direction and a trajectory point position.
7. The method according to claim 6, further comprising: Displaying the initial shooting position in the interactive interface; as well as In response to an interactive operation for the initial shooting position, the position movement request is determined.
8. The method according to claim 1, wherein: The using of the large model to process the initial image and the plurality of shooting trajectory points to obtain a generated image sequence for the object comprises: Processing the N shooting trajectory points using the position feature extraction network of the large model to obtain the position features of each of the N shooting trajectory points, where N is an integer greater than 1; Processing the initial image using the image feature extraction network of the large model to obtain initial image features; Using the image generation network of the large model to process the initial image features and the first position features of the first shooting trajectory point, to obtain a first generated image; Processing the n-1th image feature of the n-1th generated image in the generated image sequence and the n-th position feature of the n-th shooting trajectory point using the image generation network to obtain an n-th generated image, wherein N≥n>1, and the n-1th image feature is determined based on the n-1th generated image; and When N=n, the generated image sequence is obtained based on N generated images.
9. A virtual object generation device based on a large model, comprising: A shooting trajectory point acquisition module, used to plan a video shooting trajectory based on an initial image related to the object, and obtain a plurality of shooting trajectory points, wherein the shooting trajectory points represent the movement trajectory required for the image acquisition device to acquire an image of the object with respect to the initial image; A generated image sequence acquisition module, used for processing the initial image and a plurality of the shooting trajectory points using a large model to obtain a generated image sequence for the object; as well as The three-dimensional virtual object generation module is used to generate an image based on at least one frame in the generated image sequence to generate a three-dimensional virtual object representing the object.
10. An artificial intelligence agent, comprising: An input module, used for receiving input information; a processing module, configured to determine a target task based on the input information received by the input module, determine a large model based on the target task, and obtain output information by calling the large model to execute the method according to any one of claims 1 to 8; An output module is used to output the output information obtained by the processing module.
11. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 8.
12. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 8.
13. A computer program product, comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 4.
Citation Information
Cited By
Grinding and polishing track self-adaptive planning method and device based on intelligent vision
CN120645132A