Intelligent Agent Three-Dimensional Virtual Scene Generation Optimization Method, Device, Equipment and Medium

The multi-modal alignment framework automates the generation of high-quality 3D virtual scenes by retrieving optimal asset data and performing pose alignment and physical optimization, addressing the limitations of existing labor-intensive and library-dependent 3D scene construction methods.

CN120014209BActive Publication Date: 2025-07-15BEIJING INSTITUTE FOR GENERAL ARTIFICIAL INTELLIGENCE
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510474387.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-15
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

In the prior art, the generation of intelligent body three-dimensional virtual scenes relies on manual annotation, and the scalability is limited. The application scope is limited. The lack of an automated and efficient three-dimensional scene creation framework, making it difficult to achieve large-scale, high-quality, and simulateable three-dimensional scene construction.

Method used

By generating candidate asset data that recognizes the target scene image, the optimal asset data is retrieved using preset multimodal alignment rules, object pose alignment and physical optimization are performed, and efficient, real and high-quality three-dimensional virtual scenes are constructed.

Benefits of technology

It has realized the automated construction of large-scale, simulateable three-dimensional scenarios, reduce dependence on specific asset libraries, improve the environmental fidelity and interactive reliability of embodied intelligent research, and reduce manual intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014209B_ABST
    Figure CN120014209B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides an optimization method for generating an intelligent agent's three-dimensional virtual scene, which can be applied to the field of artificial intelligence technology. The optimization method for generating an intelligent agent's three-dimensional virtual scene includes: generating candidate asset data corresponding to a target object in a recognized target scene image; retrieving optimal asset data corresponding to the target object based on a preset multi-modal alignment rule, through the text description information, mask segmentation image, and candidate asset data corresponding to the target object; performing object pose alignment processing in a preset virtual scene according to the optimal asset data to generate a target virtual scene; and performing scene physics optimization for the target object on the target virtual scene to complete the optimization process of the virtual scene. An embodiment of the present invention also provides an optimization device, device, storage medium, and program product for generating an intelligent agent's three-dimensional virtual scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, specifically to the field of image processing technology, and more specifically to a method, device, equipment, medium and product for optimizing the generation of a three-dimensional virtual scene of an intelligent agent. Background Art

[0002] Artificial Intelligence (AI for short) is an important driving force for the new round of scientific and technological revolution and industrial transformation. It is a new key technical science that studies, develops theories, methods, technologies and application systems for simulating, extending and expanding human intelligence. As an important part of intelligent science, artificial intelligence attempts to understand the essence of intelligence and produce a new intelligent machine (i.e., an intelligent agent) that can respond in a way similar to human intelligence.

[0003] In the field of artificial intelligence, computer vision can use visual sensing devices (such as cameras, etc.) to replace the human eye to identify, track and detect targets, etc., and further through graphics processing, achieve the "seeing" function similar to that of the human eye. Among them, three-dimensional scene datasets play a core role in the field of computer vision, especially playing a crucial role in scene understanding and interaction tasks. At present, the construction of higher-quality three-dimensional indoor scene datasets is mainly achieved through two methods. One is to directly design the scene in a simulation environment, and the other is to directly obtain scene data using high-resolution devices during the scanning process. Although these datasets have greatly promoted the research of embodied intelligence, especially making important progress in tasks such as reasoning, navigation and operation. However, the construction of high-quality three-dimensional scenes still relies on a large amount of manual work. With the increasing importance of the scale of three-dimensional scene datasets, the existing construction of high-quality three-dimensional scenes still has limitations in scalability due to highly relying on manual annotation, limited application scope due to relying on specific asset datasets, and lack of an automated and efficient three-dimensional scene creation framework, making it difficult to achieve large-scale, efficient and high-quality three-dimensional scene construction. Summary of the Invention

[0004] In view of at least one of the above problems, embodiments of the present invention provide a method, device, equipment, medium and product for optimizing the generation of a three-dimensional virtual scene of an intelligent agent, so as to be able to provide an automated and efficient three-dimensional scene creation framework, realize the automated construction of large-scale three-dimensional scenes that are efficient, realistic, high-quality, simulatable and diverse, reduce the dependence on a specific asset library through an extensible method, and thus more efficiently achieve automated scene creation.

[0005] An aspect of an embodiment of the present invention provides a method for optimizing the generation of an intelligent agent three-dimensional virtual scene, which includes: generating candidate asset data corresponding to a target object in a recognized target scene image; retrieving the optimal asset data corresponding to the target object based on a preset multimodal alignment rule, using the text description information, mask segmentation image, and candidate asset data corresponding to the target object; performing object pose alignment processing on the preset virtual scene according to the optimal asset data to generate a target virtual scene; and performing scene physical optimization corresponding to the target object on the target virtual scene to complete the generation and optimization process of the three-dimensional virtual scene.

[0006] According to an embodiment of the present invention, before generating the candidate asset data corresponding to the target object in the recognized target scene image, it further includes: extracting a set of target perspective images of the target object in the preset virtual scene; and extracting the target scene image from the set of target perspective images.

[0007] According to an embodiment of the present invention, in generating the candidate asset data corresponding to the target object in the recognized target scene image, it includes: obtaining the text description information and mask segmentation image corresponding to the target object according to the target scene image; and generating candidate asset data according to the text description information and mask segmentation image.

[0008] According to an embodiment of the present invention, in obtaining the text description information and mask segmentation image corresponding to the target object according to the target scene image, it includes: extracting and complementing the mask segmentation image of the target scene image; and generating the text description information corresponding to the mask segmentation image through a preset language model.

[0009] According to an embodiment of the present invention, in generating candidate asset data according to the text description information and mask segmentation image, it includes: generating first candidate data according to the text description information; generating second candidate data according to the mask segmentation image; and retrieving third candidate data according to the text description information; wherein the candidate asset data includes the first candidate data, the second candidate data, and the third candidate data.

[0010] According to an embodiment of the present invention, in retrieving the optimal asset data corresponding to the target object based on a preset multimodal alignment rule, using the text description information, mask segmentation image, and candidate asset data corresponding to the target object, it includes: extracting the text features of the text description information, the image features of the mask segmentation image, and the point cloud features of the candidate asset data; generating the corresponding first matching information between the text features and the point cloud features and the corresponding second matching information between the image features and the point cloud features; generating a target loss function based on the first matching information and the second matching information according to the third matching information corresponding to the point cloud features and a preset best matching vector; generating an asset retrieval function by combining a preset auxiliary loss function and the target loss function; and retrieving the optimal asset data from the candidate asset data according to the asset retrieval function.

[0011] According to an embodiment of the present invention, when performing object pose alignment processing in a preset virtual scene based on optimal asset data to generate a target virtual scene, it includes: in the preset virtual scene, translating the center position of the virtual object corresponding to the optimal asset data to coincide with the real scene center position of the target object; in the preset virtual scene, performing asset scaling on the virtual object so that the longest side of the bounding box of the virtual object matches the real longest side of the target object; and in the preset virtual scene, performing rotation around a preset central axis of the virtual object at preset interval angles to complete the object pose alignment processing and generate the target virtual scene.

[0012] According to an embodiment of the present invention, when performing scene physical optimization for the target object on the target virtual scene, it includes: performing physical constraint optimization for spatial position relationships on the target virtual scene according to the hierarchical scene data corresponding to the target object; adding target physical attributes to the target object of the target virtual scene after physical constraint optimization through a physical simulation environment to complete the scene physical optimization.

[0013] Another aspect of the embodiments of the present invention provides a generation and optimization device for an intelligent agent three-dimensional virtual scene, which includes a data generation module, an asset retrieval module, a pose alignment module, and a physical optimization module. The data generation module is used to generate candidate asset data corresponding to the target object in the recognized target scene image; the asset retrieval module is used to retrieve the optimal asset data corresponding to the target object based on a preset multimodal alignment rule through the text description information, mask segmentation image, and candidate asset data corresponding to the target object; the pose alignment module is used to perform object pose alignment processing in a preset virtual scene based on the optimal asset data to generate the target virtual scene; and the physical optimization module is used to perform scene physical optimization for the target object on the target virtual scene to complete the generation and optimization processing of the three-dimensional virtual scene.

[0014] Another aspect of the embodiments of the present invention provides an electronic device, which includes one or more processors and a memory. The memory is used to store one or more programs. When the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the above-mentioned generation and optimization method for the intelligent agent three-dimensional virtual scene.

[0015] Another aspect of the embodiments of the present invention provides a computer-readable storage medium, on which executable instructions are stored. When the instructions are executed by a processor, the processor is caused to execute the above-mentioned generation and optimization method for the intelligent agent three-dimensional virtual scene.

[0016] Another aspect of the embodiments of the present invention provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the above-mentioned generation and optimization method for the intelligent agent three-dimensional virtual scene.

[0017] The method for optimizing the generation of the agent's three-dimensional virtual scene provided by the embodiments of the present invention can at least partially solve the problem of low intelligence level (such as low quality, low efficiency, and high distortion rate) existing in the process of constructing the agent's three-dimensional scene in the related art, and thus can at least achieve one of the following technical effects:

[0018] Based on the above method for optimizing the generation of the agent's three-dimensional virtual scene according to the embodiments of the present invention, by proposing an algorithm framework for automatically constructing a virtual copy of the real-world three-dimensional scanned scene, a large-scale, simulable, and high-quality three-dimensional scene dataset is constructed. This three-dimensional scene dataset can provide rich data support for embodied intelligence research by replacing the objects in the real-world three-dimensional scan with high-quality three-dimensional assets from multiple sources, promoting more realistic environment simulation and interaction. In addition, the above algorithm framework first uses a powerful multi-modal alignment model to select the most suitable replacement object as the optimal asset data from the candidate three-dimensional asset data, and precisely aligns its position, size, and orientation. On this basis, by further introducing physical simulation for scene optimization, it is ensured that the objects in the virtual scene conform to physical laws (such as stability, collision detection, etc.), thereby enhancing the authenticity and practicality of the virtual copy.

[0019] Therefore, based on the above method for optimizing the generation of the agent's three-dimensional virtual scene according to the embodiments of the present invention, it is possible to automatically construct a highly authentic and simulable three-dimensional scene copy, improve the environmental fidelity and interaction reliability of embodied intelligence research, while reducing manual intervention and achieving large-scale, efficient, and high-quality three-dimensional scene generation and optimization.

[0020] It should be understood that the above general description and the following specific embodiments are only exemplary and explanatory, and they do not limit the scope of what the present invention intends to claim. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Through the following description of the embodiments of the present invention with reference to the drawings, the above content and other objects, features, and advantages of the present invention will become clearer. In the drawings:

[0022] Figure 1 Schematically shows an application scenario diagram of the method, device, equipment, medium, and program product for optimizing the generation of the agent's three-dimensional virtual scene according to the embodiments of the present invention;

[0023] Figure 2 Schematically shows a flowchart of the method for optimizing the generation of the agent's three-dimensional virtual scene according to the embodiments of the present invention;

[0024] Figure 3A Schematically shows an application scenario diagram corresponding to the data collection stage of the method for optimizing the generation of the agent's three-dimensional virtual scene according to the embodiments of the present invention;

[0025] Figure 3B Schematically shows an application scenario diagram corresponding to the data annotation stage of the method for optimizing the generation of an agent three-dimensional virtual scene according to an embodiment of the present invention;

[0026] Figure 3C Schematically shows an application scenario diagram corresponding to the scene optimization stage of the method for optimizing the generation of an agent three-dimensional virtual scene according to an embodiment of the present invention;

[0027] Figure 4 Schematically shows an application scenario diagram for retrieving optimal asset data in the data annotation stage corresponding to the method for optimizing the generation of an agent three-dimensional virtual scene according to an embodiment of the present invention;

[0028] Figure 5 Schematically shows a structural block diagram of an apparatus for optimizing the generation of an agent three-dimensional virtual scene according to an embodiment of the present invention; and

[0029] Figure 6 Schematically shows a block diagram of an electronic device suitable for implementing the method for optimizing the generation of an agent three-dimensional virtual scene according to an embodiment of the present invention.

[0030] The above-mentioned drawings are a part of the description of the embodiments of the present invention, which illustrate the exemplary embodiments of the present invention. The accompanying drawings, together with the description of the specification, are used to explain the principles of the embodiments of the present invention. It should be understood that the above general description of the drawings and the following detailed description are only exemplary and explanatory, and do not limit the scope of the present invention that is intended to be claimed. Detailed Embodiments

[0031] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer and more understandable, the spirit of the content disclosed by the present invention will be clearly described below with reference to the drawings and in detail. After any person skilled in the art understands the embodiments of the content of the present invention, the techniques taught by the content of the present invention can be changed and modified, which does not deviate from the spirit and scope of the content of the present invention.

[0032] The schematic embodiments of the present invention and their descriptions are used to explain the present invention, but do not limit the present invention. Additionally, elements / components using the same or similar reference numerals in the drawings and embodiments are used to represent the same or similar parts.

[0033] Regarding the use of "first", "second",... etc. in the present invention, it does not particularly refer to the meaning of order or sequence, nor is it used to limit the present invention. It is only used to distinguish elements or operations described with the same technical terms.

[0034] Regarding the directional terms used in the present invention, such as: up, down, left, right, front or back, etc., they are only references to the directions in the attached drawings. Therefore, the directional terms used are for illustration and not for limiting this creation.

[0035] Regarding the terms "comprising", "including", "having", "containing", etc. used in the present invention, they are all open-ended terms, that is, they are meant to include but not be limited to.

[0036] Regarding the "and / or" used in the present invention, it includes any one or all combinations of the described things.

[0037] Regarding "a plurality of" in the present invention, it includes "two" and "more than two"; regarding "a plurality of groups" in the present invention, it includes "two groups" and "more than two groups".

[0038] Regarding the terms "substantially", "about", etc. used in the present invention, they are used to modify any quantity or error that can vary slightly, but these slight variations or errors do not change their nature. Generally, the range of such slight variations or errors modified by such terms can be 20% in some embodiments, 10% in some embodiments, 5% in some embodiments or other values. Those skilled in the art should understand that the aforementioned values can be adjusted according to actual needs and are not limited thereto.

[0039] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0040] In cases where expressions such as "at least one of A, B, and C, etc." are used, generally, it should be interpreted according to the meaning that a person skilled in the art would usually understand this expression (for example, "a system having at least one of A, B, and C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.). In cases where expressions such as "at least one of A, B, or C, etc." are used, generally, it should be interpreted according to the meaning that a person skilled in the art would usually understand this expression (for example, "a system having at least one of A, B, or C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.). A person skilled in the art should also understand that substantially any disjunctive conjunction and / or phrase representing two or more alternative items, whether in the specification, claims, or drawings, should be understood to give the possibility of including one of these items, either of these items, or both items. For example, the phrase "A or B" should be understood to include the possibility of "A" or "B", or "A and B".

[0041] Creating realistic and diverse simulatable 3D scenes from real data is a long-term challenge. In the prior art, there are ways to try to determine the pose of an object by matching the object in an image with a 3D model through key points; other prior arts consider using a single RGB image to optimize the size, position, orientation, and appearance of a 3D object. However, although these prior arts aim to improve the scene understanding ability, it is still difficult to generate realistic 3D objects in complex environments, lacking the robustness and generalization ability required in embodied intelligence research.

[0042] To address the challenges of object modeling in 3D scenes, there is another solution in the prior art. Specifically, it is possible to construct multiple large-scale datasets and provide detailed 3D asset matching annotations. However, these datasets still face the problem of limited asset types. For example, Scan2CAD matches the scan data in ScanNet with 3D CAD models in ShapeNet, but due to the need for a large amount of manual intervention in the adjustment, selection, or design of 3D assets, the scalability is limited, especially when it comes to objects with hinge structures. Therefore, these challenges further highlight the necessity of an automated scene creation framework. There is also another solution in the prior art. This solution (such as ACDC) uses a base model for object matching, but performs poorly in complex and realistic scenes and highly depends on existing asset datasets.

[0043] In summary, there are at least the following technical problems in the prior art for the generation of the 3D virtual scene of an agent that urgently need to be solved:

[0044] (1) Highly dependent on manual annotation with limited scalability: Existing solutions (such as Scan2CAD) require a large amount of manual intervention to adjust, select, or design 3D objects when matching scanned data with 3D assets. Especially when dealing with objects having hinge structures, it is difficult to scale up.

[0045] (2) Dependence on specific asset datasets limits the application scope: Existing methods (such as ACDC) rely heavily on fixed 3D asset libraries, and the coverage of these asset libraries is limited, resulting in large deviations when matching real scanned scenes and making it difficult to apply to more complex and realistic scenes.

[0046] (3) Lack of an automated and efficient 3D scene creation framework: Currently, the construction of 3D scene datasets mainly relies on manual design or high-resolution scanning. There is still a lack of a general and automated framework to replace objects in real scans and optimize their geometric properties to achieve high-quality and simulatable 3D scene reconstruction.

[0047] In view of at least one of the above problems, embodiments of the present invention provide a method, apparatus, device, medium, and product for generating and optimizing an agent-based 3D virtual scene, thereby being able to provide an automated and efficient 3D scene creation framework, realizing the automated construction of efficient, realistic, high-quality, simulatable, and diverse large-scale 3D scenes, reducing the dependence on specific asset libraries through scalable methods, and thus more efficiently realizing automated scene creation.

[0048] One aspect of an embodiment of the present invention provides a method for generating and optimizing an agent-based 3D virtual scene, including: generating candidate asset data corresponding to a target object in a recognized target scene image; retrieving optimal asset data corresponding to the target object based on a preset multimodal alignment rule through the text description information, mask segmentation image, and candidate asset data corresponding to the target object; performing object pose alignment processing in a preset virtual scene according to the optimal asset data to generate a target virtual scene; and performing scene physical optimization corresponding to the target object on the target virtual scene to complete the generation and optimization processing of the 3D virtual scene.

[0049] Figure 1 Schematically shows an application scenario diagram of the method, apparatus, device, medium, and program product for generating and optimizing an agent-based 3D virtual scene according to an embodiment of the present invention.

[0050] As Figure 1As shown, the application scenario 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0051] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only for examples).

[0052] The terminal devices 101, 102, 103 may be various electronic devices with a display screen and supporting web browsing, including but not limited to smartphones, tablets, laptop portable computers, and desktop computers, etc.

[0053] The server 105 may be a server providing various services, such as a background management server that supports the websites browsed by users using the terminal devices 101, 102, 103 (only for examples). The background management server may analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.

[0054] It should be noted that the method for generating and optimizing the intelligent agent three-dimensional virtual scene provided by the embodiments of the present invention can generally be executed by the server 105. Correspondingly, the device for generating and optimizing the intelligent agent three-dimensional virtual scene provided by the embodiments of the present invention can generally be set in the server 105. The method for generating and optimizing the intelligent agent three-dimensional virtual scene provided by the embodiments of the present invention can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103 and / or the server 105. Correspondingly, the device for generating and optimizing the intelligent agent three-dimensional virtual scene provided by the embodiments of the present invention can also be set in a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103 and / or the server 105.

[0055] It should be understood that Figure 1 the numbers of terminal devices, networks, and servers in

[0056] are merely illustrative. According to the implementation requirements, there may be any number of terminal devices, networks, and servers. Figure 1 which will be based on the Figures 2 to 4A method for optimizing the generation of an agent's three-dimensional virtual scene in an open embodiment is described in detail.

[0057] As Figure 2 shown, one aspect of an embodiment of the present invention provides a method for optimizing the generation of an agent's three-dimensional virtual scene, which includes operations S201 to S204.

[0058] In operation S201, candidate asset data corresponding to the target object in the recognized target scene image is generated;

[0059] In operation S202, based on a preset multimodal alignment rule, the optimal asset data corresponding to the target object is retrieved through the text description information, mask segmentation image, and candidate asset data corresponding to the target object;

[0060] In operation S203, object pose alignment processing is performed in a preset virtual scene according to the optimal asset data to generate a target virtual scene;

[0061] In operation S204, scene physical optimization corresponding to the target object is performed on the target virtual scene to complete the generation and optimization process of the three-dimensional virtual scene.

[0062] The agent can be the execution subject of the above-mentioned method for optimizing the generation of an agent's three-dimensional virtual scene in an embodiment of the present invention, or the execution party controlled by the method for optimizing the generation of an agent's three-dimensional virtual scene. Specifically, it can be a humanoid intelligent robot or other AI devices, which usually have actuators capable of completing specific action tasks, such as scanning and detecting the surrounding environment or scene through visual sensors such as cameras.

[0063] The three-dimensional virtual scene can generally be a three-dimensional space scene restoration of some real-space scenes provided by a virtual setting system such as a computer simulation system (such as a simulator Simulator). Specifically, for example, the agent establishes a virtual three-dimensional space scene by means of three-dimensional construction of the scanned living room environment data.

[0064] The target scene image is an image of the detection perspectives of the target object at different angles in the target scene, and can include two-dimensional images of the target object and its surrounding objects. The target object can be one or some specific objects targeted in the generation process of the three-dimensional virtual scene in the embodiments of the present invention. Usually, asset data collection, annotation, and optimization are performed on the target object so that the target object can achieve the best match in the virtual scene. In addition, the candidate asset data can be the basic three-dimensional scene data for annotation and optimization of the target object in the three-dimensional virtual scene, and can be generated through the target scene image. Specifically, it can be the three-dimensional virtual model of the corresponding target object in the three-dimensional virtual scene. The candidate asset data can be obtained by conversion through the target scene image, or data matching the target object can be queried from the existing asset library through the target scene image.

[0065] The preset multimodal alignment rule can be a rule for retrieving the optimal asset data for the above-mentioned candidate asset data. The preset multimodal alignment rule defines the rule for achieving the best annotation of the three-dimensional virtual model of each target object in the three-dimensional virtual scene. The text description information can be the text expression information in natural language of the target object corresponding to the above-mentioned target scene image, and specifically can involve text information such as the color, material, shape, and mass of the target object. The mask segmentation image can be the segmentation image of the target object corresponding to the above-mentioned target scene image, and specifically can involve the target object image segmented relative to the object background image. The optimal asset data can be the three-dimensional object model selected from the candidate asset data that satisfies the best placement relationship of the target object in the three-dimensional virtual scene, where the best placement relationship can be a relationship related to the placement features such as the position, attributes (color, material, shape, etc.) of the real target object in the real environment.

[0066] The preset virtual scene can be a simulated space generated by scanning and detecting the real space or real scene (i.e., the target scene) where the target object is located, and can include other relevant scene objects including the target object. For example, the preset virtual scene is constructed according to the scanning and detection data of the target scene. Therefore, the preset virtual scene can be understood as the construction basis of the three-dimensional virtual scene. By performing object pose alignment processing (such as position, size, and orientation, etc.) on the optimal asset data, the three-dimensional object model of the target object corresponding to the optimal asset data is adjusted in pose to be consistent with the pose of the real target object in the real scene, thereby completing the construction of the target virtual scene. Therefore, the target virtual scene can be a virtual scene copy formed after performing object pose alignment processing on the preset virtual scene for the target object.

[0067] The main purpose of physical scene optimization is to process the content in the target virtual scene that conflicts with the physical laws in the real scene. For example, there are situations where the three-dimensional object model of the target object shows physical interference (such as the object exceeding the scene boundary) in the target virtual scene, which violates physical laws.

[0068] Therefore, based on the above-mentioned intelligent agent three-dimensional virtual scene generation and optimization method of the embodiments of the present invention, a large-scale and simulatable three-dimensional scene dataset can be constructed. This three-dimensional scene dataset can replace the target objects in the real-world three-dimensional scan with high-quality three-dimensional assets from various sources, providing rich data support for embodied intelligence research and promoting more realistic environment simulation and interaction. In addition, based on the above method of the embodiments of the present invention, an algorithm framework for automatically constructing a virtual copy of the real-world three-dimensional scan scene can be provided. Among them, this framework can utilize a powerful multi-modal alignment model to select the most suitable replacement object from the candidate three-dimensional candidate assets and perform precise pose alignment on its position, size, orientation, etc. Subsequently, physical simulation is introduced for scene optimization to ensure that the objects in the three-dimensional virtual scene conform to physical laws (such as stability, collision detection, etc.), thereby improving the authenticity and practicality of the virtual copy.

[0069] In summary, the above-mentioned intelligent agent three-dimensional virtual scene generation and optimization method of the embodiments of the present invention can automatically construct a three-dimensional scene copy with high authenticity and simulability, improve the environmental fidelity and interaction reliability of embodied intelligence research, reduce manual intervention at the same time, and achieve large-scale and efficient three-dimensional scene generation and optimization.

[0070] To enable those skilled in the art to have a clearer understanding of the above-mentioned intelligent agent three-dimensional virtual scene generation and optimization method of the embodiments of the present invention, the following Figures 3A to 4 description is further provided.

[0071] As Figures 3A to 4 shown, in the embodiments of the present invention, the above-mentioned three-dimensional virtual scene generation and optimization method can be divided into three aspects: data collection 301 (as Figure 3A shown, that is, Collection), data annotation 302 (as Figure 3B shown, that is, Annotation), and scene optimization 303 (as Figure 3C shown, that is, Optimization) to realize the construction process of the three-dimensional scene dataset.

[0072] For each scanned object, the goal of the three-dimensional virtual scene generation and optimization method of the embodiments of the present invention is to find diverse and high-quality three-dimensional assets as replacement candidates, ensure that these assets match the original object as much as possible, and have good simulation performance.

[0073] AsFigures 3A to 4 As shown, before generating the candidate asset data corresponding to the target object in the recognized target scene image in operation S201 according to an embodiment of the present invention, it further includes:

[0074] Extracting a set of target perspective images of the target object in a preset virtual scene;

[0075] Extracting the target scene image from the set of target perspective images.

[0076] As Figure 3A shown, the intelligent agent detects and recognizes the selected target space through a visual sensing device (such as a camera), and uses the image data obtained by the detection and recognition to build three-dimensional scene data to form a preset virtual scene 311 (Scene PCD) based on point cloud data, as Figure 3A shown. By judging the three-dimensional assets with poor quality in the preset virtual scene 311, the corresponding target object can be determined as an alternative candidate target for the three-dimensional asset. Among them, the so-called three-dimensional asset can be understood as the three-dimensional model data of the corresponding object.

[0077] Based on the determined target object, in the preset virtual scene 311, for the position of the target object in the preset virtual scene, multiple perspectives can be selected simultaneously to obtain images of the target object in the preset virtual scene. As Figure 3A shown, the set of target perspective images 312 (Multiview Images) is a set of selected images of the target object in the preset virtual scene from the above different perspectives.

[0078] By preprocessing each target perspective image in the set of target perspective images, based on the preprocessing results of the target perspective images, the target perspective image with the best image quality can be selected as the target scene image 313 (Best-viewselection), as Figure 3A shown.

[0079] Therefore, by means of constructing the preset virtual scene, it is possible to confirm the target object that does not meet the virtual scene image quality, and thereby select the perspective image corresponding to the target object, thus laying a better data foundation for subsequent data annotation and physical optimization, and also being able to reduce the amount of data processing, realize the annotation and optimization of specific target objects, and thus be able to accelerate the generation efficiency of the three-dimensional virtual scene.

[0080] Each target perspective image in the target perspective image can be sharpened through techniques such as NAFNet (Nonlinear Activation Free Network for Image Restoration), thereby ensuring better image quality of the image and more accurate subsequent processing.

[0081] In addition, depth image technology can be further used to process the depth map of each target perspective image to extract the target scene image 313 (Best-view selection). Among them, the target scene image 313 can be a two-dimensional image with the least occlusion and the clearest view of the target object based on the target perspective image.

[0082] Therefore, a high-quality target object image with the least occlusion and the clearest view of the target object in the above target perspective image set can be obtained.

[0083] As Figures 3A to 4 shown, according to an embodiment of the present invention, in the operation S201 of generating candidate asset data corresponding to the target object in the recognized target scene image, it includes:

[0084] Obtain the text description information and mask segmentation image corresponding to the target object according to the target scene image;

[0085] Generate candidate asset data according to the text description information and the mask segmentation image.

[0086] The text description information can be the text expression information in natural language of the target object corresponding to the above target scene image, and specifically can involve text information such as the color, material, shape, and quality of the target object. The mask segmentation image can be the segmentation image of the target object corresponding to the above target scene image, and specifically can involve the target object image segmented relative to the object background image. The candidate asset data (Asset Candidates Creation) can be the basic three-dimensional scene data for annotating and optimizing the target object in the three-dimensional virtual scene, and can be generated through the target scene image. Specifically, it can be the three-dimensional virtual model of the corresponding target object in the three-dimensional virtual scene.

[0087] Thereby, a large-scale and simulable three-dimensional scene dataset can be constructed. This dataset can replace the objects in the real-world three-dimensional scan with high-quality three-dimensional assets from multiple sources, providing rich data support for embodied intelligence research and promoting more realistic environment simulation and interaction.

[0088] As Figures 3A to 4 shown, according to an embodiment of the present invention, in obtaining the text description information and mask segmentation image corresponding to the target object according to the target scene image, it includes:

[0089] Extract and complete the masked segmentation image of the target scene image; and

[0090] Generate text description information corresponding to the masked segmentation image through a preset language model.

[0091] As Figure 3A shown, image segmentation techniques such as SAM (Segment Anything Model) can be used to perform image segmentation processing around the target object on the target scene image, thereby generating a two-dimensional mask image of the target object, where the two-dimensional mask image can be a two-dimensional image obtained by segmenting and excluding the background information of the target object in the target scene image.

[0092] As Figure 3A shown, further, image completion techniques such as SD (Stable Diffusion) can be used to complete the incomplete target object in the above two-dimensional mask image, and the above-mentioned masked segmentation image 315 can be generated. Therefore, the masked segmentation image can be a two-dimensional image of the target object with a complete structure.

[0093] The preset language model 314 can be a preset language model with image-text conversion capabilities (such as Large Language Model, abbreviated as LLM), specifically GPT-4v. Taking the above masked segmentation image as the input data of the preset language model, the preset language model can automatically generate detailed descriptive natural language text covering the texture, color, and physical properties of the target object as the text description information 316. For example, if the target object is a wooden bar stool, the corresponding text description information 316 may include the following content:

[0094] “A wooden bar stool

[0095] Color: Brown

[0096] Texture: Wooden

[0097] Shape: Strip

[0098] Rigid body

[0099] Mass: 5.0 kg”.

[0100] Therefore, through the above preset language model, rich semantic description texts can be generated for each scanned target object, making the target object more concrete and significantly improving the subsequent processing efficiency.

[0101] As Figures 3A to 4As shown, according to an embodiment of the present invention, in generating candidate asset data based on text description information and mask segmentation images, it includes:

[0102] Generating first candidate data according to the text description information;

[0103] Generating second candidate data according to the mask segmentation image; and

[0104] Retrieving third candidate data according to the text description information;

[0105] Wherein the candidate asset data includes the first candidate data, the second candidate data, and the third candidate data.

[0106] For the above target object, more candidate asset data can be generated based on the above text description information and mask segmentation image. For example, a three-dimensional object model of the target object can be constructed based on advanced three-dimensional object modeling techniques. Additionally, the target model data can be retrieved from the existing three-dimensional object model database according to the target object.

[0107] As Figure 3A shown, the target three-dimensional model data of the target object can be generated using models such as Shape-E through the text description information, directly generating the three-dimensional object mesh of the target object, constituting the first candidate data 371, that is, text-to-3D generation.

[0108] In addition, the three-dimensional object model data of the target object can be formed using TripoSR, InstantMesh, and Michelangelo, etc. respectively through the mask segmentation image, constituting the second candidate data 372, that is, image-to-3D generation.

[0109] Moreover, using the above text description information as the retrieval information, existing three-dimensional object model data that conforms to the text description information can be retrieved from the candidate asset database, and the retrieved model data is used as the third candidate data 373, that is, text-to-3D retrieval. Specifically, different methods such as Uni3D and ULIP (Unified Representation of Language, Images, and Point) can be used to retrieve assets that match the text description information of the target object from the candidate asset library respectively. Among them, the candidate asset library can be a large-scale three-dimensional asset database, such as Objaverse.

[0110] Finally, as Figure 3AAs shown, in order to further improve the quality and realism of the generated 3D object mesh, texture optimization technology (such as Paint3D, etc.) is further applied to perform texture optimization processing 318 on the above candidate asset data to enhance the color fidelity and surface details of the generated assets and make them closer to real-world objects.

[0111] Therefore, a 3D asset retrieval scheme based on multimodal alignment can be provided, and a large-scale and simulable 3D scene dataset is constructed. By replacing the objects in real-world 3D scans with high-quality 3D assets from multiple sources, this dataset provides rich data support for embodied intelligence research and promotes more realistic environment simulation and interaction.

[0112] Such as Figure 3B As shown, after obtaining the candidate asset data of the 3D asset as candidates, data annotation 302 (Annotation) can be further implemented, and the optimal asset data is selected from these candidate asset data based on comprehensive judgments such as geometric similarity (such as shape) and visual appearance similarity (such as material and texture).

[0113] Such as Figures 3A to 4 As shown, according to an embodiment of the present invention, in operation S202, based on a preset multimodal alignment rule, by using the text description information, mask segmentation image, and candidate asset data corresponding to the target object to retrieve the optimal asset data corresponding to the target object, it includes:

[0114] Extract the text features of the text description information, the image features of the mask segmentation image, and the point cloud features of the candidate asset data;

[0115] Generate the corresponding first matching information between the text features and the point cloud features and the corresponding second matching information between the image features and the point cloud features;

[0116] Generate a target loss function based on the first matching information and the second matching information according to the third matching information corresponding to the point cloud features and a preset optimal matching vector;

[0117] Combine a preset auxiliary loss function and the target loss function to generate an asset retrieval function;

[0118] Retrieve the optimal asset data from the candidate asset data according to the asset retrieval function.

[0119] Such as Figure 4 As shown, in order to create a realistic and diverse simulable 3D scene from real data and have scalability, a complete automated virtual scene creation framework can be constructed as a powerful multimodal alignment model.

[0120] The model framework can generate optimal asset data based on various encoders and multi-layer perceptrons, etc. As Figure 4 shown, the text features of the text description information 316 are extracted through the frozen text encoder 401 (Text Encoder); the image features of the mask segmentation image 315 are extracted through the frozen image encoder 402 (Image Encoder); the point cloud features of the candidate asset data 317 are extracted through the point cloud encoder 403 (Point Cloud Encoder).

[0121] Specifically, as Figure 4 shown, through the multi-modal alignment model, the best asset candidate, i.e., the optimal asset data, is retrieved from the data set of the candidate asset data. For the target object i in the preset virtual scene, a quadruple < I i , T i , P i , y i > can be constructed, where I i can be the mask segmentation image 315, T i can be the text description information 316, P i ={ P 1 i , ¨¨¨, P L i} can be L a set of potential candidate point clouds (corresponding to the above candidate asset data 317), y i can be a one-hot vector used to represent the best match.

[0122] Based on the design of the above quadruple, the learning of optimal asset retrieval is achieved through the multi-modal contrast model.

[0123] First, the image features of the mask segmentation image and the text features of the text description information are respectively extracted using the frozen image encoder and text encoder h I i and text features h T i . Then, a learnable three-dimensional encoder Ɛ P is adopted to P k i ∈P i Extract point cloud features h P i,k = Ɛ P ( P k i )。Calculate the matching information (such as matching score) between each candidate asset data and the corresponding mask segmentation image or text description information through the following formula 1:

[0124]

[0125] Therefore, the corresponding first matching information between the generated text features and point cloud features is q T i , realizing text feature alignment 411 (i.e., Text-PCD Aligned); the corresponding second matching information between the generated image features and point cloud features is q I i , realizing image feature alignment 412 (i.e., Image-PCD Aligned).

[0126] In addition, as Figure 4 shown, it is also possible to input the above { h P i,l} L l=1 into a learnable multi-layer perceptron 404 (Multilayer Perceptron, abbreviated as MLP) to directly calculate the matching score corresponding to its third matching information from the point cloud q P i , to prevent the situation where no image or text is available.

[0127] Therefore, based on the above first matching information q T i , second matching information q I i and third matching information q P i and the above preset optimal matching vector y i , the following formula 2 can be constructed for the target loss function:

[0128]

[0129] where is the standard deviation. Using the above objective loss function can achieve better supervised learning of the model, ensuring the accurate and efficient acquisition of the optimal asset data.

[0130] To better align the point cloud features with the image or text features in different scenarios and object instances, a new candidate set P i ′ is further created to add additional supervision signals, which can be composed of the original best candidates and candidates randomly sampled from different scenarios. Calculate similar matching information (such as matching scores) following Equation 1 q T i ′ and q I i ′ to form the preset auxiliary loss function shown in Equation 3 below:

[0131]

[0132] Therefore, the final asset retrieval function can be expressed as Equation 4 below as the learning objective:

[0133]

[0134] In, the asset retrieval function shown in Equation 4 above can be used as the basis for retrieving the optimal asset data and also defines the above preset multimodal alignment rules. Therefore, the optimal asset data can be the best candidate asset for a 3D virtual scene that can ensure fit with the original real object scene through the pose and maximize the restoration of the real scene.

[0135] In summary, the above method for optimizing the generation of a 3D virtual scene in the embodiments of the present invention mainly uses a model architecture based on image, text, and point cloud information to select the most suitable 3D assets to replace the target scanned object, and through multimodal feature learning, it can significantly improve the accuracy of matching, specifically as Figure 4As shown in the figure. Among them, for the target object in the scene, by constructing a quadruple containing an image, a text description, multiple candidate point clouds, and an optimal matching annotation, the corresponding features are extracted using frozen image and text encoders, and the features of the candidate point clouds are extracted using a learnable 3D encoder. Further, in order to calculate matching information such as matching scores, multi-modal contrast learning is adopted to calculate the similarity between the point cloud features and the image and text features. At the same time, a method based on the MLP network is introduced to directly predict the matching score from the point cloud features to handle the situation where there is a lack of images or texts. In addition, in order to further improve the alignment effect in different scenes and object instances, the candidate asset set is expanded, and an auxiliary loss is introduced to enhance the matching ability between the point cloud and the image and text using additional supervision signals. Finally, the optimization objective of the model consists of the main matching loss and the auxiliary loss, thereby improving the accuracy and generalization ability of asset retrieval.

[0136] As Figures 3A to 4 shown, according to an embodiment of the present invention, in operation S203, performing object pose alignment processing in a preset virtual scene according to the optimal asset data to generate a target virtual scene, including:

[0137] In the preset virtual scene, translating the center position of the virtual object corresponding to the optimal asset data to coincide with the real scene center position of the target object;

[0138] In the preset virtual scene, performing asset scaling on the virtual object so that the longest side of the bounding box of the virtual object matches the longest side of the target object in reality; and

[0139] In the preset virtual scene, rotating around a preset central axis of the virtual object at a preset interval angle to complete the object pose alignment processing and generate a target virtual scene.

[0140] As Figure 3B shown, during the data annotation 302 process, a heuristic-based asset placement process can be adopted to accurately align the pose information such as the position, size, and orientation of the 3D object model corresponding to the selected optimal asset data in the preset virtual scene with the pose information of the scanned target object in the real scene, and align the best retrieved asset to this preset virtual scene. Among them, the real scene is the real space environment detected by the agent through scanning.

[0141] Specifically, as Figure 3BThe shown annotation interface can translate the center of the best retrieved asset to the center of the real scanned object. Additionally, the asset can be scaled so that the longest side of the asset bounding box matches the longest side of the real scanned object. Moreover, the asset is rotated around a preset central axis at a preset interval angle such as 30 degrees to find the minimum rotation angle, thereby achieving the best alignment with the real scanned object. Herein, the virtual object is a three-dimensional object model corresponding to the above optimal asset data in the preset virtual scene; the real scanned object is the target object in the real scene.

[0142] Therefore, through the above process of automatically aligning the complete object pose of the optimal asset data in the preset virtual scene, ultimately, three-dimensional asset retrieval and replacement based on multi-modal alignment can be achieved. The best candidate asset is placed into the three-dimensional scene, and its orientation and size are adjusted according to the actual situation to ensure its fit (Best Match) with the original object and maximize the restoration of the real scene.

[0143] It can be seen that through the above algorithm framework for automatically constructing a virtual copy of the real-world three-dimensional scanning scene and using a powerful multi-modal alignment model to select the most suitable replacement object from the candidate three-dimensional assets, precise alignment of its position, size, and orientation can be achieved.

[0144] As Figures 3A to 4 shown, according to an embodiment of the present invention, in the operation S204 of performing scene physical optimization of the target object on the target virtual scene, it includes:

[0145] Performing physical constraint optimization on the target virtual scene for the spatial position relationship according to the hierarchical scene data corresponding to the target object;

[0146] Adding target physical attributes to the target object of the target virtual scene after physical constraint optimization through a physical simulation environment to complete the scene physical optimization.

[0147] The hierarchical scene data can be the stored data of a three-dimensional hierarchical scene graph constructed based on scene point cloud technology, which can be specifically presented in the form of table information to reflect all spatial relationship information of the target object in the scene, such as support, embedding, and containment, etc. Among them, the hierarchical scene data can be used for physical constraint optimization (PhysicalOptimization) of the target object, specifically using the above spatial relationship information as optimization constraint conditions. Among them, physical constraint optimization is the constraint on the physical spatial relationship between the target object and the virtual scene, aiming to make the spatial relationship of the target object in the virtual scene more conform to the spatial physical relationship of the real scene.

[0148] AsFigure 3C As shown, virtual objects existing in a preset virtual scene will actually have various spatial relationships with the virtual scene. In many cases, these spatial relationships will show various states that do not conform to spatial physical relationships, such as the target object exceeding the scene range of the virtual scene (i.e., out-of-bounds Outside), the target object being suspended in the virtual scene without contacting other objects (i.e., floating Floating), or the target object having at least partial overlap with other objects (i.e., collision Collision), etc.

[0149] To achieve the above-mentioned optimization of physical constraints for the target object, optimization sampling techniques such as Markov Chain Monte Carlo (MCMC for short) can be selected for optimization (as Figure 3C shown in SSG & MCMC optimization 331), for example, fine-tuning the positional relationship of the target object in the horizontal direction in the preset virtual scene, specifically for optimizations such as penetration and touch. Therefore, an optimal trade-off can be achieved between the scene graph information provided by the hierarchical scene data and physical conflicts (such as collisions), guiding the optimization process to adjust the final position of the target object to make it meet physical rationality.

[0150] Furthermore, import the three-dimensional virtual scene optimized by the above physical constraints into a physical simulator such as Blender (as Figure 3C shown in simulator 332), and add target physical properties (such as material type, mass, etc. information) to the corresponding target object through the physical simulation environment of this physical simulator to enhance the physical authenticity of the reconstructed scene and make it more suitable for simulation and interaction tasks.

[0151] In summary, for the traditional optimization scheme that directly uses gradient-based methods to achieve complex constraint optimization, the embodiments of the present invention achieve lower computational costs through the method of global optimization of virtual scenes based on physical constraints, and can further ensure the physical rationality of objects in the scene; moreover, by introducing physical simulation for scene optimization, it can better ensure that the objects in the virtual scene conform to physical laws (such as stability, collision detection, etc.), thereby enhancing the authenticity and practicality of the virtual copy.

[0152] In summary, the above method for generating and optimizing the intelligent agent three-dimensional virtual scene according to the embodiments of the present invention can automatically construct a highly authentic and simulable three-dimensional scene copy, improve the environmental fidelity and interaction reliability of embodied intelligence research, while reducing manual intervention, and achieve large-scale, efficient, and high-quality three-dimensional scene generation and optimization.

[0153] Based on the above method for generating and optimizing the intelligent agent three-dimensional virtual scene, the present invention also provides a device for generating and optimizing the intelligent agent three-dimensional virtual scene. The following will be combined withFigure 5 Describe the device in detail.

[0154] Figure 5 The structural block diagram of the generation optimization device for the intelligent agent three-dimensional virtual scene according to an embodiment of the present invention is schematically shown.

[0155] As Figure 5 shown, the generation optimization device 500 for the intelligent agent three-dimensional virtual scene in this embodiment includes a data generation module 510, an asset retrieval module 520, an attitude alignment module 530, and a physical optimization module 540.

[0156] The data generation module 510 is used to generate candidate asset data corresponding to the target object in the recognized target scene image. In one embodiment, the data generation module 510 can be used to perform the operation S201 described above, which will not be elaborated here.

[0157] The asset retrieval module 520 is used to retrieve the optimal asset data corresponding to the target object based on the preset multimodal alignment rule, through the text description information, mask segmentation image, and candidate asset data corresponding to the target object. In one embodiment, the asset retrieval module 520 can be used to perform the operation S202 described above, which will not be elaborated here.

[0158] The attitude alignment module 530 is used to perform object attitude alignment processing in the preset virtual scene according to the optimal asset data to generate the target virtual scene. In one embodiment, the attitude alignment module 530 can be used to perform the operation S203 described above, which will not be elaborated here.

[0159] The physical optimization module 540 is used to perform scene physical optimization of the target object on the target virtual scene to complete the optimization processing of the virtual scene. In one embodiment, the physical optimization module 540 can be used to perform the operation S204 described above, which will not be elaborated here.

[0160] According to an embodiment of the present invention, any plurality of modules among the data generation module 510, the asset retrieval module 520, the pose alignment module 530, and the physical optimization module 540 may be combined and implemented in one module, or any one of them may be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules may be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the data generation module 510, the asset retrieval module 520, the pose alignment module 530, and the physical optimization module 540 may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or may be implemented by any other reasonable means such as hardware or firmware by integrating or packaging circuits, or may be implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, at least one of the data generation module 510, the asset retrieval module 520, the pose alignment module 530, and the physical optimization module 540 may be at least partially implemented as a computer program module, and when the computer program module is run, corresponding functions may be executed.

[0161] Figure 6 FIG. schematically shows a block diagram of an electronic device suitable for implementing the method for generating and optimizing an agent three-dimensional virtual scene according to an embodiment of the present invention.

[0162] The above-mentioned electronic device provided by the embodiment of the present invention includes one or more processors and a memory, and the memory is used to store one or more programs. Wherein, when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the above-mentioned method for generating and optimizing an agent three-dimensional virtual scene.

[0163] As Figure 6 shown, the electronic device 600 according to an embodiment of the present invention includes a processor 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage part 608 into a random access memory (RAM) 603. The processor 601 may include, for example, a general microprocessor (such as a CPU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (such as an application specific integrated circuit (ASIC)), etc. The processor 601 may also include on-board memory for caching purposes. The processor 601 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.

[0164] In the RAM 603, various programs and data required for the operation of the electronic device 600 are stored. The processor 601, the ROM 602, and the RAM 603 are connected to each other via the bus 604. The processor 601 performs various operations of the method flow according to the embodiments of the present invention by executing the programs in the ROM 602 and / or the RAM 603. It should be noted that the programs may also be stored in one or more memories other than the ROM 602 and the RAM 603. The processor 601 may also perform various operations of the method flow according to the embodiments of the present invention by executing the programs stored in the one or more memories.

[0165] According to an embodiment of the present invention, the electronic device 600 may further include an input / output (I / O) interface 605, and the input / output (I / O) interface 605 is also connected to the bus 604. The electronic device 600 may further include one or more of the following components connected to the I / O interface 605: an input portion 606 including a keyboard, a mouse, etc.; an output portion 607 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 608 including a hard disk, etc.; and a communication portion 609 including a network interface card such as a LAN card, a modem, etc. The communication portion 609 performs communication processing via a network such as the Internet. The drive 610 is also connected to the I / O interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as needed so that a computer program read therefrom can be installed into the storage portion 608 as needed.

[0166] The present invention also provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, cause the processor to execute the above-mentioned method for optimizing the generation of the intelligent agent three-dimensional virtual scene.

[0167] Among them, the computer-readable storage medium may be included in the device / apparatus / system described in the above embodiments; or it may exist separately without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of the present invention is implemented.

[0168] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present invention, the computer-readable storage medium may include one or more memories other than the above-described ROM 602 and / or RAM 603 and / or ROM 602 and RAM 603.

[0169] An embodiment of the present invention further includes a computer program product, which includes a computer program that, when executed by a processor, implements the above-mentioned method for optimizing the generation of the intelligent agent three-dimensional virtual scene.

[0170] Wherein, the computer program contains program codes for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program codes are used to enable the computer system to implement the method provided by the embodiment of the present invention.

[0171] When the computer program is executed by the processor 601, it executes the above-mentioned functions defined in the system / apparatus of the embodiment of the present invention. According to an embodiment of the present invention, the above-described systems, apparatuses, modules, units, etc. can be implemented by computer program modules.

[0172] In one embodiment, the computer program can rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program can also be transmitted and distributed in the form of a signal on a network medium, and is downloaded and installed through the communication part 609, and / or installed from the removable medium 611. The program codes contained in the computer program can be transmitted by any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0173] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 609, and / or installed from the removable medium 611. When the computer program is executed by the processor 601, it executes the above-mentioned functions defined in the system of the embodiment of the present invention. According to an embodiment of the present invention, the above-described systems, devices, apparatuses, modules, units, etc. can be implemented by computer program modules.

[0174] According to embodiments of the present invention, program code for executing the computer programs provided by the embodiments of the present invention can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. The programming languages include, but are not limited to, such as Java, C++, Python, the "C" language, or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., by using an Internet service provider to connect through the Internet).

[0175] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutively represented blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, and the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0176] In addition, all actions of obtaining information, signals, or data in the present invention are carried out on the premise of complying with the corresponding data protection laws, regulations, and policies of the country where it is located, and obtaining the authorization given by the owner of the corresponding device.

[0177] Those skilled in the art can understand that the features described in various embodiments and / or claims of the present invention can be combined or / and combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features described in various embodiments and / or claims of the present invention can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.

[0178] The embodiments of the present invention have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although the embodiments have been described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. The scope of the present invention is defined by the appended claims and their equivalents. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and these substitutions and modifications should fall within the scope of the present invention.

Claims

1. An optimization method for generating an agent virtual scene, characterized in that, Including: Generating candidate asset data corresponding to the target object in the recognized target scene image; Based on the preset multimodal alignment rules, retrieving the optimal asset data corresponding to the target object through the text description information, mask segmentation image, and the candidate asset data generated by the preset language model corresponding to the target object, including: extracting the text features of the text description information, the image features of the mask segmentation image, and the point cloud features of the candidate asset data; generating the first matching information corresponding between the text features and the point cloud features and the second matching information corresponding between the image features and the point cloud features; generating a target loss function based on the first matching information and the second matching information according to the third matching information corresponding to the point cloud features and the preset optimal matching vector; combining the preset auxiliary loss function and the target loss function to generate an asset retrieval function; retrieving the optimal asset data from the candidate asset data according to the asset retrieval function; Performing object pose alignment processing on the preset virtual scene according to the optimal asset data to generate a target virtual scene; and Performing scene physical optimization corresponding to the target object on the target virtual scene to complete the generation and optimization processing of the three-dimensional virtual scene; Among them, the preset multimodal alignment rules are the rules for retrieving the optimal asset data for the candidate asset data, and define the rules for the three-dimensional virtual model of each target object to achieve the best annotation in the three-dimensional virtual scene; the optimal asset data is the three-dimensional object model selected from the candidate asset data that satisfies the best placement relationship of the target object in the three-dimensional virtual scene, where the best placement relationship is the correlation with the placement characteristics of the real target object in the real environment.

2. The method according to claim 1, wherein Before generating the candidate asset data corresponding to the target object in the recognized target scene image, it further includes: Extracting the target perspective image set of the target object in the preset virtual scene; Extracting the target scene image from the target perspective image set.

3. The method according to claim 1, wherein Among the generating the candidate asset data corresponding to the target object in the recognized target scene image, it includes: Obtaining the text description information and the mask segmentation image corresponding to the target object according to the target scene image; Generating the candidate asset data according to the text description information and the mask segmentation image.

4. The method according to claim 3, characterized in that Among the obtaining the text description information and the mask segmentation image corresponding to the target object according to the target scene image, it includes: Extracting and complementing the mask segmentation image of the target scene image; and Generating the text description information corresponding to the mask segmentation image through a preset language model.

5. The method according to claim 3, wherein Among the generating the candidate asset data according to the text description information and the mask segmentation image, it includes: Generating the first candidate data according to the text description information; Generating the second candidate data according to the mask segmentation image; and Retrieving the third candidate data according to the text description information; Wherein the candidate asset data includes the first candidate data, the second candidate data, and the third candidate data.

6. The method according to claim 1, characterized in that, Among the performing object pose alignment processing on the preset virtual scene according to the optimal asset data to generate a target virtual scene, it includes: In the preset virtual scene, translate the center position of the virtual object corresponding to the optimal asset data to coincide with the real scene center position of the target object; In the preset virtual scene, perform asset scaling on the virtual object so that the longest side of the bounding box of the virtual object matches the real longest side of the target object; and In the preset virtual scene, perform rotation around the preset central axis of the virtual object at preset interval angles to complete the object pose alignment process and generate a target virtual scene.

7. The method according to claim 6, characterized in that In the scene physical optimization of the target virtual scene corresponding to the target object, it includes: Perform physical constraint optimization for spatial position relationships on the target virtual scene according to the hierarchical scene data corresponding to the target object; Add target physical attributes to the target object of the target virtual scene after the physical constraint optimization through a physical simulation environment to complete the scene physical optimization.

8. An intelligent agent three-dimensional virtual scene generation and optimization device, characterized in that, It includes: A data generation module for generating candidate asset data corresponding to the target object in the recognized target scene image; An asset retrieval module for retrieving the optimal asset data corresponding to the target object based on a preset multimodal alignment rule, through the text description information, mask segmentation image, and the candidate asset data generated by the preset language model corresponding to the target object, including: extracting the text features of the text description information, the image features of the mask segmentation image, and the point cloud features of the candidate asset data; generating the corresponding first matching information between the text features and the point cloud features and the corresponding second matching information between the image features and the point cloud features; generating a target loss function based on the third matching information corresponding to the point cloud features and the preset best matching vector; combining the preset auxiliary loss function and the target loss function to generate an asset retrieval function; retrieving the optimal asset data from the candidate asset data according to the asset retrieval function; A pose alignment module for performing object pose alignment processing in the preset virtual scene according to the optimal asset data to generate a target virtual scene; and A physical optimization module for performing scene physical optimization of the target virtual scene corresponding to the target object to complete the generation and optimization process of the three-dimensional virtual scene; Among them, the preset multimodal alignment rule is a rule for retrieving the optimal asset data for the candidate asset data, which defines the rule for the three-dimensional virtual model of each target object to achieve the best annotation in the three-dimensional virtual scene; the optimal asset data is a three-dimensional object model selected from the candidate asset data that satisfies the best placement relationship of the target object in the three-dimensional virtual scene, where the best placement relationship is the relevant relationship with the placement characteristics of the real target object in the real environment.

9. An electronic device, including: One or more processors; A memory for storing one or more programs, Wherein, when the one or more programs are executed by the one or more processors, the one or more processors execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having executable instructions stored thereon, which when executed by a processor cause the processor to perform the method according to any one of claims 1 to 7.

11. A computer program product comprising a computer program, which when executed by a processor implements the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Virtual scene generation method and system

    CN119166236A