Generation optimization method and device of intelligent body three-dimensional virtual scene, equipment and medium

Through the intelligent body three-dimensional virtual scene generation optimization method, multimodal alignment rules and physical optimization technology are used to solve the problem that three-dimensional scene construction in the existing technology relies on manual and specific asset data sets, and realize the automated construction of high-quality and simulateable three-dimensional scenes, which improves the environmental fidelity of embodied intelligent research.

CN120014209AActive Publication Date: 2025-05-16BEIJING INSTITUTE FOR GENERAL ARTIFICIAL INTELLIGENCE

Patent Information

Application Number
CN202510474387.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-05-16
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

When building high-quality three-dimensional indoor scene datasets, the existing technology relies on a large amount of manual work, resulting in limited scalability, relying on specific asset datasets, lacking automated efficient three-dimensional scene creation frameworks, making it difficult to achieve large-scale, efficient and high-quality three-dimensional scene construction.

Method used

It provides a method for generating and optimization of the three-dimensional virtual scene of an intelligent body. By generating and identifying the target object candidate asset data in the target scene image, searching the optimal asset data based on multimodal alignment rules, object posture alignment and scene physical optimization are carried out, and efficient, real and high-quality three-dimensional scenes are automatically constructed.

Benefits of technology

It realizes automated construction of high-reality and simulateable three-dimensional scene copies, improves the environmental fidelity and interactive reliability of embodied intelligent research, reduces manual intervention, and realizes large-scale and efficient three-dimensional scene generation and optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014209A_ABST
    Figure CN120014209A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an intelligent body three-dimensional virtual scene generation optimization method which can be applied to the technical field of artificial intelligence. The intelligent agent three-dimensional virtual scene generation optimization method comprises the steps of generating candidate asset data corresponding to a target object in an identified target scene image; based on a preset multi-modal alignment rule, searching optimal asset data corresponding to the target object through the text description information corresponding to the target object, the mask segmentation image and the candidate asset data; performing object attitude alignment processing in a preset virtual scene according to the optimal asset data to generate a target virtual scene; and performing scene physical optimization of the corresponding target object on the target virtual scene to complete optimization processing of the virtual scene. The embodiment of the invention further provides a device and equipment for generating and optimizing the three-dimensional virtual scene of the intelligent agent, a storage medium and a program product.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, specifically to the field of image processing technology, and more specifically to a method, device, equipment, medium and product for optimizing the generation of a three-dimensional virtual scene of an intelligent body. Background Art

[0002] Artificial Intelligence (AI) is an important driving force for the new round of scientific and technological revolution and industrial transformation. It is a new key technical science that studies and develops theories, methods, technologies and application systems for simulating, extending and expanding human intelligence. As an important part of intelligent science, artificial intelligence attempts to understand the essence of intelligence and produce a new intelligent machine (i.e., intelligent agent) that can respond in a similar way to human intelligence.

[0003] In the field of artificial intelligence, computer vision can use visual sensing devices (such as cameras, etc.) to replace human eyes to identify, track and detect targets, and further realize the "seeing" function similar to human eyes through graphics processing. Among them, three-dimensional scene datasets occupy a core position in the field of computer vision, especially in scene understanding and interaction tasks. At present, the construction of higher-quality three-dimensional indoor scene datasets is mainly achieved in two ways. One is to design the scene directly in the simulation environment, and the other is to directly obtain scene data using high-resolution equipment during the scanning process. Although these datasets have greatly promoted the research of embodied intelligence, especially in tasks such as reasoning, navigation and operation, they have made important progress. However, the construction of high-quality three-dimensional scenes still relies on a lot of manual work. As the importance of the scale of three-dimensional scene datasets becomes increasingly prominent, the construction of existing high-quality three-dimensional scenes still has the problems of limited scalability due to high dependence on manual annotation, limited application scope due to dependence on specific asset datasets, and lack of automated and efficient three-dimensional scene creation framework, making it difficult to achieve large-scale, efficient and high-quality three-dimensional scene construction. Summary of the invention

[0004] In view of at least one of the above problems, an embodiment of the present invention provides a method, device, equipment, medium and product for optimizing the generation of intelligent three-dimensional virtual scenes, thereby providing an automatic and efficient three-dimensional scene creation framework, realizing the automatic construction of efficient, realistic, high-quality, simulatable and diverse large-scale three-dimensional scenes, reducing dependence on specific asset libraries through an extensible method, thereby more efficiently realizing automated scene creation.

[0005] One aspect of an embodiment of the present invention provides a method for optimizing the generation of a three-dimensional virtual scene of an intelligent body, which includes: generating candidate asset data corresponding to a target object in a recognized target scene image; retrieving optimal asset data corresponding to the target object based on preset multimodal alignment rules through text description information, mask segmentation image and candidate asset data corresponding to the target object; performing object posture alignment processing in a preset virtual scene according to the optimal asset data to generate a target virtual scene; and performing scene physics optimization corresponding to the target object on the target virtual scene to complete the generation optimization processing of the three-dimensional virtual scene.

[0006] According to an embodiment of the present invention, before generating candidate asset data corresponding to the target object in the identified target scene image, it also includes: extracting a target perspective image set of the target object in a preset virtual scene; and extracting a target scene image from the target perspective image set.

[0007] According to one embodiment of the present invention, in generating candidate asset data corresponding to a target object in a recognized target scene image, it includes: acquiring text description information and a mask segmentation image corresponding to the target object based on the target scene image; and generating the candidate asset data based on the text description information and the mask segmentation image.

[0008] According to one embodiment of the present invention, obtaining text description information and a mask segmentation image corresponding to a target object based on a target scene image includes: extracting and completing the mask segmentation image of the target scene image; and generating text description information corresponding to the mask segmentation image through a preset language model.

[0009] According to one embodiment of the present invention, generating candidate asset data based on text description information and a mask segmented image includes: generating first candidate data based on the text description information; generating second candidate data based on the mask segmented image; and retrieving third candidate data based on the text description information; wherein the candidate asset data includes the first candidate data, the second candidate data and the third candidate data.

[0010] According to one embodiment of the present invention, in retrieving the optimal asset data corresponding to the target object based on a preset multimodal alignment rule through the text description information, mask segmentation image and candidate asset data corresponding to the target object, it includes: extracting text features of the text description information, image features of the mask segmentation image and point cloud features of the candidate asset data; generating first matching information corresponding to the text features and the point cloud features and second matching information corresponding to the image features and the point cloud features; generating a target loss function based on the first matching information and the second matching information according to the third matching information corresponding to the point cloud features and the preset best matching vector; generating an asset retrieval function in combination with the preset auxiliary loss function and the target loss function; and retrieving the optimal asset data from the candidate asset data according to the asset retrieval function.

[0011] According to one embodiment of the present invention, in performing object posture alignment processing in a preset virtual scene according to optimal asset data to generate a target virtual scene, it includes: in the preset virtual scene, translating the center position of the virtual object corresponding to the optimal asset data to coincide with the center position of the real scene of the target object; in the preset virtual scene, performing asset scaling on the virtual object so that the longest side of the bounding box of the virtual object matches the real longest side of the target object; and in the preset virtual scene, performing rotation around the preset central axis of the virtual object at a preset interval angle to complete the object posture alignment processing and generate the target virtual scene.

[0012] According to one embodiment of the present invention, in performing scene physics optimization of a target object corresponding to a target virtual scene, it includes: performing physical constraint optimization on the target virtual scene for a spatial position relationship according to hierarchical scene data corresponding to the target object; adding target physical properties to the target object of the target virtual scene after physical constraint optimization through a physical simulation environment to complete the scene physics optimization.

[0013] Another aspect of the embodiments of the present invention provides a generation and optimization device for an intelligent three-dimensional virtual scene, which includes a data generation module, an asset retrieval module, a posture alignment module and a physical optimization module. The data generation module is used to generate candidate asset data corresponding to a target object in a recognized target scene image; the asset retrieval module is used to retrieve the optimal asset data corresponding to the target object based on a preset multimodal alignment rule through text description information corresponding to the target object, a mask segmentation image and candidate asset data; the posture alignment module is used to perform object posture alignment processing in a preset virtual scene according to the optimal asset data to generate a target virtual scene; and the physical optimization module is used to perform scene physical optimization corresponding to the target object on the target virtual scene to complete the generation and optimization processing of the three-dimensional virtual scene.

[0014] Another aspect of an embodiment of the present invention provides an electronic device, comprising one or more processors and a memory, the memory being used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors execute the above-mentioned method for optimizing the generation of a three-dimensional virtual scene of an intelligent body.

[0015] Another aspect of an embodiment of the present invention provides a computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, causes the processor to execute the above-mentioned method for optimizing the generation of a three-dimensional virtual scene of an intelligent body.

[0016] Another aspect of an embodiment of the present invention provides a computer program product, including a computer program, which, when executed by a processor, implements the above-mentioned method for optimizing the generation of a three-dimensional virtual scene of an intelligent body.

[0017] The method for optimizing the generation of an intelligent three-dimensional virtual scene provided by the embodiment of the present invention can at least partially solve the problem of low intelligence level (such as low quality, low efficiency and high distortion rate) in the process of constructing an intelligent three-dimensional scene in the related art, and thus can achieve at least one of the following technical effects: Based on the above-mentioned intelligent three-dimensional virtual scene generation optimization method of the embodiment of the present invention, a large-scale, simulatable high-quality three-dimensional scene dataset is constructed by proposing an algorithm framework for automatically constructing a virtual copy of a real-world three-dimensional scanned scene. The three-dimensional scene dataset can replace objects in real-world three-dimensional scans with high-quality three-dimensional assets from a variety of sources, providing rich data support for embodied intelligence research and promoting more realistic environmental simulation and interaction. In addition, the above-mentioned algorithm framework first uses a powerful multimodal alignment model to select the most suitable replacement object from the candidate three-dimensional asset data as the optimal asset data, and accurately aligns its position, size and orientation. On the above basis, by further introducing physical simulation for scene optimization, it is ensured that the objects in the virtual scene conform to physical laws (such as stability, collision detection, etc.), thereby improving the authenticity and practicality of the virtual copy.

[0018] Therefore, based on the generation and optimization method of the intelligent body three-dimensional virtual scene of the above-mentioned embodiment of the present invention, a highly realistic and simulatable three-dimensional scene copy can be automatically constructed, thereby improving the environmental realism and interaction reliability of embodied intelligence research, while reducing human intervention, and realizing large-scale, efficient, high-quality three-dimensional scene generation and optimization.

[0019] It should be understood that the above general description and the following detailed description are merely exemplary and illustrative and are not intended to limit the scope of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The above contents and other objects, features and advantages of the present invention will become more apparent through the following description of the embodiments of the present invention with reference to the accompanying drawings, in which: Figure 1 A schematic diagram of an application scenario of a method, device, equipment, medium and program product for optimizing generation of a three-dimensional virtual scene of an intelligent body according to an embodiment of the present invention is shown; Figure 2 A flowchart of a method for optimizing the generation of a three-dimensional virtual scene of an intelligent agent according to an embodiment of the present invention is schematically shown; Figure 3A Schematically showing an application scenario diagram corresponding to the data collection phase of the method for optimizing the generation of a three-dimensional virtual scene of an intelligent agent according to an embodiment of the present invention; Figure 3BSchematically showing an application scenario diagram corresponding to the data annotation stage of the method for optimizing the generation of a three-dimensional virtual scene of an intelligent agent according to an embodiment of the present invention; Figure 3C The application scene diagram corresponding to the scene optimization stage of the method for optimizing the generation of the intelligent three-dimensional virtual scene according to the embodiment of the present invention is schematically shown; Figure 4 A schematic diagram of an application scenario of retrieving optimal asset data in a data annotation phase corresponding to a method for optimizing the generation of a three-dimensional virtual scene of an intelligent agent according to an embodiment of the present invention is shown; Figure 5 A schematic diagram of a structure of a device for optimizing generation of a three-dimensional virtual scene of an intelligent agent according to an embodiment of the present invention is shown; and Figure 6 A block diagram of an electronic device suitable for implementing a method for optimizing generation of a three-dimensional virtual scene of an intelligent agent according to an embodiment of the present invention is schematically shown.

[0021] The above-mentioned drawings are part of the specification of the embodiments of the present invention, which illustrate exemplary embodiments of the present invention. The attached drawings and the description of the specification are used together to illustrate the principles of the embodiments of the present invention. It should be understood that the above general description of the drawings and the following specific implementations are only exemplary and illustrative, and they cannot limit the scope of the present invention. DETAILED DESCRIPTION

[0022] In order to make the objectives, technical solutions and advantages of the embodiments of the present invention more clearly understood, the spirit of the contents disclosed by the present invention will be clearly explained with the accompanying drawings and detailed descriptions below. After understanding the embodiments of the contents of the present invention, any technician in the relevant technical field can change and modify the techniques taught by the contents of the present invention without departing from the spirit and scope of the contents of the present invention.

[0023] The exemplary embodiments and descriptions of the present invention are used to explain the present invention, but are not intended to limit the present invention. In addition, elements / components with the same or similar reference numerals used in the drawings and embodiments are used to represent the same or similar parts.

[0024] The terms “first”, “second”, etc. used in the present invention do not particularly refer to an order or sequence, nor are they used to limit the present invention. They are only used to distinguish elements or operations described with the same technical terms.

[0025] The directional terms used in the present invention, such as up, down, left, right, front or back, etc., are only used to refer to the directions of the drawings. Therefore, the directional terms used are used to illustrate and not to limit the present invention.

[0026] The words “include,” “including,” “have,” “contain,” etc. used in the present invention are open-ended terms, meaning including but not limited to.

[0027] The term "and / or" used in the present invention includes any or all combinations of the items mentioned.

[0028] Regarding the present invention, "plurality" includes "two" and "more than two"; regarding the present invention, "plurality of groups" includes "two groups" and "more than two groups".

[0029] The terms "substantially" and "approximately" used in the present invention are used to modify any quantity or error that may vary slightly, but these slight changes or errors do not change their essence. Generally speaking, the range of slight changes or errors modified by such terms may be 20% in some embodiments, 10% in some embodiments, 5% in some embodiments, or other values. Those skilled in the art should understand that the aforementioned values ​​can be adjusted according to actual needs and are not limited thereto.

[0030] All terms (including technical and scientific terms) used herein have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0031] In the case of using expressions such as "at least one of A, B, and C, etc.", it should generally be interpreted in accordance with the meaning of the expression generally understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but not be limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or a system having A, B, C, etc.). In the case of using expressions such as "at least one of A, B, or C, etc.", it should generally be interpreted in accordance with the meaning of the expression generally understood by those skilled in the art (for example, "a system having at least one of A, B, or C" should include but not be limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or a system having A, B, C, etc.). Those skilled in the art should also understand that any transitional conjunctions and / or phrases that substantially represent two or more optional items, whether in the specification, claims, or drawings, should be understood to give the possibility of including one of these items, either of these items, or both of these items. For example, the phrase "A or B" should be understood to include the possibilities of "A" or "B", or "A and B".

[0032] Creating realistic and diverse simulatable 3D scenes from real data is a long-standing challenge. Some existing technologies attempt to match objects in images with 3D models by key points to determine the pose of the objects; other existing technologies consider optimizing the size, position, orientation, and appearance of 3D objects using a single RGB image. However, although these existing technologies aim to improve scene understanding capabilities, they still have difficulty generating realistic 3D objects in complex environments and lack the robustness and generalization capabilities required for embodied intelligence research.

[0033] In order to meet the challenges of object modeling in three-dimensional scenes, there is another solution in the prior art. Specifically, it can be achieved by constructing multiple large-scale datasets and providing detailed three-dimensional asset matching annotations. However, these datasets still face the problem of limited asset types. For example, Scan2CAD matches scan data in ScanNet with three-dimensional CAD models in ShapeNet, but the adjustment, selection or design of three-dimensional assets requires a lot of manual intervention, so the scalability is limited, especially when it comes to objects with hinged structures. This limitation is more obvious. Therefore, these challenges further highlight the necessity of an automated scene creation framework. There is another solution in the prior art. Although this solution (such as ACDC) uses a basic model for object matching, it performs poorly in complex and realistic scenes and is highly dependent on existing asset datasets.

[0034] In summary, in the prior art, there are at least the following technical problems to be solved for the generation of intelligent three-dimensional virtual scenes: (1) High reliance on manual annotation and limited scalability: Existing solutions (such as Scan2CAD) require a lot of manual intervention to adjust, select or design 3D objects when matching scan data with 3D assets. This is especially difficult to scale when dealing with objects with hinged structures.

[0035] (2) Dependence on specific asset datasets limits the scope of application: Existing methods (such as ACDC) rely heavily on fixed 3D asset libraries, which have limited coverage, resulting in large deviations when matching real scanned scenes, making it difficult to apply to more complex and realistic scenes.

[0036] (3) Lack of an automated and efficient 3D scene creation framework: Currently, the construction of 3D scene datasets mainly relies on manual design or high-resolution scanning. There is still a lack of a general, automated framework to replace objects in real scans and optimize their geometric properties to achieve high-quality, simulatable 3D scene reconstruction.

[0037] In view of at least one of the above problems, an embodiment of the present invention provides a method, device, equipment, medium and product for optimizing the generation of intelligent three-dimensional virtual scenes, thereby providing an automatic and efficient three-dimensional scene creation framework, realizing the automatic construction of efficient, realistic, high-quality, simulatable and diverse large-scale three-dimensional scenes, reducing dependence on specific asset libraries through an extensible method, thereby more efficiently realizing automated scene creation.

[0038] One aspect of an embodiment of the present invention provides a method for optimizing the generation of a three-dimensional virtual scene of an intelligent body, which includes: generating candidate asset data corresponding to a target object in a recognized target scene image; retrieving optimal asset data corresponding to the target object based on preset multimodal alignment rules through text description information, mask segmentation image and candidate asset data corresponding to the target object; performing object posture alignment processing in a preset virtual scene according to the optimal asset data to generate a target virtual scene; and performing scene physics optimization corresponding to the target object on the target virtual scene to complete the generation optimization processing of the three-dimensional virtual scene.

[0039] Figure 1 The application scenario diagram schematically shows the generation optimization method, device, equipment, medium and program product of the intelligent three-dimensional virtual scene according to the embodiment of the present invention.

[0040] like Figure 1 As shown, the application scenario 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104 and a server 105. The network 104 is used to provide a medium for a communication link between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links or optical fiber cables, etc.

[0041] Users can use terminal devices 101, 102, 103 to interact with server 105 through network 104 to receive or send messages, etc. Various communication client applications can be installed on terminal devices 101, 102, 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only examples).

[0042] The terminal devices 101 , 102 , and 103 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop computers, and desktop computers.

[0043] The server 105 may be a server that provides various services, such as a background management server (only as an example) that provides support for websites browsed by users using the terminal devices 101, 102, and 103. The background management server may analyze and process the received data such as user requests, and feed back the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal device.

[0044] It should be noted that the generation optimization method of the intelligent three-dimensional virtual scene provided in the embodiment of the present invention can generally be executed by the server 105. Accordingly, the generation optimization device of the intelligent three-dimensional virtual scene provided in the embodiment of the present invention can generally be set in the server 105. The generation optimization method of the intelligent three-dimensional virtual scene provided in the embodiment of the present invention can also be executed by a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105. Correspondingly, the generation optimization device of the intelligent three-dimensional virtual scene provided in the embodiment of the present invention can also be set in a server or server cluster that is different from the server 105 and can communicate with the terminal devices 101, 102, 103 and / or the server 105.

[0045] It should be understood that Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration. Any number of terminal devices, networks and servers may be provided according to implementation requirements.

[0046] The following will be based on Figure 1 The scene described by Figure 2~Figure 4 The generation and optimization method of the intelligent agent three-dimensional virtual scene of the disclosed embodiment is described in detail.

[0047] like Figure 2 As shown, one aspect of an embodiment of the present invention provides a method for optimizing the generation of a three-dimensional virtual scene of an intelligent body, which includes operations S201 to S204.

[0048] In operation S201, candidate asset data corresponding to a target object in a recognized target scene image is generated; In operation S202, based on a preset multimodal alignment rule, optimal asset data corresponding to the target object is retrieved through text description information corresponding to the target object, the mask segmentation image, and the candidate asset data; In operation S203, object posture alignment processing is performed in a preset virtual scene according to the optimal asset data to generate a target virtual scene; In operation S204, scene physics optimization corresponding to the target object is performed on the target virtual scene to complete the generation optimization process of the three-dimensional virtual scene.

[0049] The intelligent agent can be the executor of the above-mentioned intelligent agent three-dimensional virtual scene generation optimization method of the embodiment of the present invention, or it can be an executor controlled by the intelligent agent three-dimensional virtual scene generation optimization method. Specifically, it can be a humanoid intelligent robot or other AI device, which usually has an actuator that can complete specific action tasks, such as scanning and detecting the environment or scene through visual sensors such as cameras.

[0050] The three-dimensional virtual scene can usually be a three-dimensional space scene restoration for certain real space scenes provided by a virtual setting system such as a computer simulation system (such as a simulator). Specifically, the intelligent agent uses the scanned living room environment data to build a virtual three-dimensional space scene through three-dimensional construction and other methods.

[0051] The target scene image is an image of the detection perspective of the target object at different angles in the target scene, and may include a two-dimensional image of the target object and its surrounding objects. The target object may be one or more specific objects targeted in the generation process of the three-dimensional virtual scene in the embodiment of the present invention. Usually, asset data is collected, labeled and optimized for the target object so that the target object can be best matched in the virtual scene. In addition, the candidate asset data may be basic three-dimensional scene data that is labeled and optimized for the target object in the three-dimensional virtual scene, which may be generated through the target scene image, and may specifically be a three-dimensional virtual model of the corresponding target object in the three-dimensional virtual scene. The candidate asset data may be obtained by converting the target scene image, or may be obtained by querying the target scene image for data matching the target object in the existing asset library.

[0052] The preset multimodal alignment rule may be a rule for retrieving the optimal asset data for the candidate asset data, and the preset multimodal alignment rule defines the rule for achieving the best annotation of the three-dimensional virtual model of each target object in the three-dimensional virtual scene. The text description information may be the text expression information of the target object corresponding to the target scene image in natural language, and may specifically involve text information such as the color, material, shape, and quality of the target object. The mask segmentation image may be a segmented image of the target object corresponding to the target scene image, and may specifically involve an image of the target object segmented relative to the object background image. The optimal asset data may be a three-dimensional object model selected from the candidate asset data that satisfies the optimal placement relationship of the target object in the three-dimensional virtual scene, wherein the optimal placement relationship may be a relationship related to the placement characteristics such as the position, attributes (color, material, shape, etc.) of the real target object in the real environment.

[0053] The preset virtual scene can be a simulated space generated by scanning and detecting the real space or real scene (i.e., the target scene) where the target object is located, which may include other related scene objects including the target object. For example, the preset virtual scene is generated by constructing a virtual scene based on the scanning and detection data of the target scene. Therefore, the preset virtual scene can be understood as the basis for constructing the three-dimensional virtual scene. By performing object posture alignment processing (such as position, size, and orientation, etc.) on the optimal asset data, the three-dimensional object model of the target object corresponding to the optimal asset data is adjusted through posture to be consistent with the posture of the real target object in the real scene, thereby completing the construction of the target virtual scene. Therefore, the target virtual scene can be a copy of the virtual scene formed after the preset virtual scene is subjected to object posture alignment processing of the target object.

[0054] The purpose of physical scene optimization is mainly to deal with the content in the target virtual scene that conflicts with the physical laws in the real scene. For example, the three-dimensional object model of the target object has physical interference in the target virtual scene (such as the object exceeds the scene boundary) and other violations of physical laws.

[0055] Therefore, based on the above-mentioned intelligent three-dimensional virtual scene generation optimization method of the embodiment of the present invention, a large-scale, simulatable three-dimensional scene dataset can be constructed. The three-dimensional scene dataset can replace the target objects in the real-world three-dimensional scan by using high-quality three-dimensional assets from multiple sources, providing rich data support for embodied intelligence research and promoting more realistic environmental simulation and interaction. In addition, based on the above-mentioned method of the embodiment of the present invention, an algorithm framework for automatically constructing a virtual copy of a real-world three-dimensional scan scene can be provided. Among them, the framework can use a powerful multimodal alignment model to select the most suitable replacement object from the candidate three-dimensional candidate assets, and accurately align its position, size, orientation, etc. Subsequently, physical simulation is introduced to optimize the scene to ensure that the objects in the three-dimensional virtual scene conform to physical laws (such as stability, collision detection, etc.), thereby improving the authenticity and practicality of the virtual copy.

[0056] In summary, the above-mentioned intelligent body three-dimensional virtual scene generation and optimization method of the embodiment of the present invention can automatically construct highly realistic and simulatable three-dimensional scene copies, improve the environmental realism and interaction reliability of embodied intelligence research, while reducing human intervention, and realize large-scale and efficient three-dimensional scene generation and optimization.

[0057] In order to enable those skilled in the art to have a clearer understanding of the above-mentioned intelligent three-dimensional virtual scene generation optimization method of the embodiment of the present invention, the following is further provided: Figure 3A-Figure 4 Description.

[0058] like Figure 3A-Figure 4As shown, in the embodiment of the present invention, the above-mentioned three-dimensional virtual scene generation optimization method can be divided into data collection 301 (such as Figure 3A As shown, that is, Collection), data annotation 302 (such as Figure 3B As shown, namely Annotation) and scene optimization 303 (such as Figure 3C As shown in Figure 1, the three aspects of optimization are used to realize the construction process of the three-dimensional scene dataset.

[0059] For each scanned object, the goal of the method for optimizing the generation of a three-dimensional virtual scene in an embodiment of the present invention is to find diverse and high-quality three-dimensional assets as replacement candidates, ensuring that these assets match the original object as closely as possible and have good simulation performance.

[0060] like Figure 3A-Figure 4 As shown, according to an embodiment of the present invention, before generating candidate asset data corresponding to the target object in the identified target scene image in operation S201, the method further includes: Extracting a target perspective image set of a target object in a preset virtual scene; Extract the target scene image from the target view image set.

[0061] like Figure 3A As shown, the intelligent agent detects and identifies the selected target space through a visual sensor device (such as a camera), and uses the image data obtained by the detection and identification to build three-dimensional scene data to form a preset virtual scene 311 (Scene PCD) based on point cloud data, such as Figure 3A By judging the poor quality three-dimensional assets in the preset virtual scene 311, the corresponding target object can be determined as a candidate target for replacing the three-dimensional asset. The so-called three-dimensional asset can be understood as the three-dimensional model data of the corresponding object.

[0062] Based on the determined target object, in the preset virtual scene 311, multiple viewing angles can be selected to obtain images of the target object in the preset virtual scene according to the position of the target object in the preset virtual scene. Figure 3A As shown, the target view image set 312 (Multiview Images) is a set of selected images of the target object in the preset virtual scene under the above different view angles.

[0063] By preprocessing each target view image in the target view image set, based on the preprocessing result of the target view image, the target view image with the best image quality can be selected as the target scene image 313 (Best-view selection), such as Figure 3A shown.

[0064] Therefore, by constructing a preset virtual scene, it is possible to confirm the target object that does not meet the image quality of the virtual scene, and thereby select the corresponding perspective image of the target object, thereby laying a better data foundation for subsequent data annotation and physical optimization. It can also reduce the amount of data processing, achieve annotation and optimization of specific target objects, and thus speed up the generation efficiency of three-dimensional virtual scenes.

[0065] Each target perspective image in the target perspective image can be sharpened by a technique such as NAFNet (Nonlinear Activation Free Network for Image Restoration), thereby ensuring better image quality for the image and more accurate subsequent processing.

[0066] In addition, the depth image technology can be further used to process the depth image of each target perspective image to extract the target scene image 313 (Best-view selection). The target scene image 313 can be a two-dimensional image with the least occlusion and the clearest target object based on the target perspective image.

[0067] Therefore, a high-quality target object image with the least occlusion of the target object and the clearest image in the above target perspective image set can be obtained.

[0068] like Figure 3A-Figure 4 As shown, according to an embodiment of the present invention, in operation S201, the candidate asset data corresponding to the target object in the identified target scene image is generated, including: Obtain text description information and mask segmentation image corresponding to the target object according to the target scene image; Generate candidate asset data based on text description information and mask segmentation images.

[0069] The text description information may be the textual expression information of the target object corresponding to the target scene image in natural language, and may specifically involve textual information such as the color, material, shape, and quality of the target object. The mask segmentation image may be the segmentation image of the target object corresponding to the target scene image, and may specifically involve the target object image segmented relative to the object background image. The candidate asset data (Asset Candidates Creation) may be the basic three-dimensional scene data for annotating and optimizing the target object in the three-dimensional virtual scene, and may be generated through the target scene image, and may specifically be the three-dimensional virtual model of the corresponding target object in the three-dimensional virtual scene.

[0070] This allows us to build a large-scale, simulatable 3D scene dataset, which can provide rich data support for embodied intelligence research by replacing objects in real-world 3D scans with high-quality 3D assets from a variety of sources, and promote more realistic environment simulation and interaction.

[0071] like Figure 3A-Figure 4 As shown, according to an embodiment of the present invention, in acquiring text description information corresponding to a target object and a mask segmentation image according to a target scene image, the method includes: extracting and completing a mask segmentation image of a target scene image; and Generate text description information corresponding to the mask segmentation image through the preset language model.

[0072] like Figure 3A As shown, the target scene image can be segmented around the target object with the help of image segmentation techniques such as SAM (Segment Anything Model), so as to generate a two-dimensional mask image of the target object, wherein the two-dimensional mask image can be a two-dimensional image that excludes background information of the target object in the target scene image through segmentation.

[0073] like Figure 3A As shown, further, by using image completion techniques such as SD (Stable Diffusion) to complete the image of the incomplete target object in the above two-dimensional mask image, the above mask segmentation image 315 can be generated. Therefore, the mask segmentation image can be a two-dimensional image with a complete structure of the target object.

[0074] The preset language model 314 may be a preset language model (such as Large Language Model, LLM for short) with image-text conversion capability, such as GPT-4v. The above mask segmented image is used as input data of the preset language model, and the preset language model can automatically generate a detailed description of the texture, color and physical properties of the target object as text description information 316. For example, if the target object is a wooden bar stool, the corresponding text description information 316 may include the following content: “A wooden bar stool Color: Brown Texture: Wooden Shape: Strip Rigid body Mass: 5.0 kg".

[0075] Therefore, through the above preset language model, rich semantic description text can be generated for each scanned target object, making the target object more concrete and significantly improving subsequent processing efficiency.

[0076] like Figure 3A-Figure 4 As shown, according to an embodiment of the present invention, in generating candidate asset data according to text description information and mask segmentation image, it includes: Generate first candidate data according to the text description information; generating second candidate data according to the mask segmentation image; and Retrieving third candidate data according to the text description information; The candidate asset data includes first candidate data, second candidate data and third candidate data.

[0077] For the above target object, more candidate asset data can be generated based on the above text description information and mask segmentation image. For example, a 3D object model of the target object can be constructed based on advanced 3D object modeling technology, and target model data can also be retrieved from an existing 3D object model database based on the target object.

[0078] like Figure 3A As shown, the target 3D model data of the target object can be generated by using the Shape-E model through the text description information, and the 3D object mesh of the target object can be directly generated to form the first candidate data 371, namely, Text-to-3D generation.

[0079] In addition, the mask segmentation image can also be used to use TripoSR, InstantMesh and Michelangelo to respectively form 3D object model data of the target object to constitute the second candidate data 372, that is, Image-to-3D generation.

[0080] Moreover, by using the above text description information as retrieval information, existing 3D object model data that matches the text description information can be retrieved in the candidate asset database, and the retrieved model data can be used as the third candidate data 373, that is, text-to-3D retrieval. Specifically, assets that match the text description information of the target object can be retrieved from the candidate asset library through different methods such as Uni3D and ULIP (Unified Representation of Language, Images, and Point). The candidate asset library can be a large-scale 3D asset database, such as Objaverse.

[0081] Finally, if Figure 3A As shown, in order to further improve the quality and realism of the generated three-dimensional object mesh, texture optimization technology (such as Paint3D, etc.) is further applied to perform texture optimization 318 processing on the above-mentioned candidate asset data to enhance the color fidelity and surface details of the generated assets, making them closer to real-world objects.

[0082] Therefore, a 3D asset retrieval scheme based on multimodal alignment can be provided, and a large-scale, simulatable 3D scene dataset can be constructed. This dataset provides rich data support for embodied intelligence research by replacing objects in real-world 3D scans with high-quality 3D assets from multiple sources, and promotes more realistic environment simulation and interaction.

[0083] like Figure 3B As shown, after obtaining the candidate asset data of the three-dimensional asset as the candidate items, data annotation 302 (Annotation) can be further implemented to select the best asset data from these candidate asset data based on comprehensive judgments such as geometric similarity (such as shape) and visual appearance similarity (such as material and texture).

[0084] like Figure 3A-Figure 4 As shown, according to an embodiment of the present invention, in operation S202, based on a preset multimodal alignment rule, the optimal asset data corresponding to the target object is retrieved through the text description information corresponding to the target object, the mask segmentation image and the candidate asset data, including: Extract text features of text description information, image features of mask segmentation images, and point cloud features of candidate asset data; Generate first matching information corresponding to the text feature and the point cloud feature and second matching information corresponding to the image feature and the point cloud feature; Generate a target loss function based on the first matching information and the second matching information according to the third matching information corresponding to the point cloud feature and the preset best matching vector; Generate an asset retrieval function by combining the preset auxiliary loss function and the target loss function; The optimal asset data is retrieved from the candidate asset data according to the asset retrieval function.

[0085] like Figure 4 As shown, in order to create realistic and diverse simulatable 3D scenes from real data with scalability, a complete automated virtual scene creation framework can be constructed as a powerful multimodal alignment model.

[0086] This model framework can realize the generation of optimal asset data based on various encoders and multi-layer perceptrons. Figure 4As shown, the text features of the text description information 316 are extracted through the frozen text encoder 401 (Text Encoder); the image features of the mask segmentation image 315 are extracted through the frozen image encoder 402 (Image Encoder); and the point cloud features of the candidate asset data 317 are extracted through the point cloud encoder 403 (Point Cloud Encoder).

[0087] Specifically, Figure 4 As shown in FIG. 1 , the best asset candidate, i.e., the optimal asset data, is retrieved from the data set of candidate asset data through the multimodal alignment model. For the target object i in the preset virtual scene, a quadruple < I i , T i , P i , y i >, where I i It can be a mask segmentation image 315, T i It can be text description information 316, P i ={ P 1 i ,¨¨¨, P L i} can be L A set of potential candidate point clouds (corresponding to the candidate asset data 317 above), y i Can be a one-hot vector representing the best match.

[0088] Based on the above quadruple design, the optimal asset retrieval is learned through a multimodal comparison model.

[0089] First, a frozen image encoder and a text encoder are used to extract the image features of the mask segmentation image and the text description information respectively. h I i and text features h T i Next, a learnable 3D encoder is used Ɛ P , for each candidate point cloud P k i ∈ P i Extract point cloud features h Pi,k = Ɛ P ( P k i ). The matching information (such as matching score) between each candidate asset data and the corresponding mask segmentation image or text description information is calculated by the following formula 1:

[0090] Therefore, the first matching information corresponding to the text feature and the point cloud feature is generated as q T i , realize text feature alignment 411 (i.e., Text-PCD Aligned); generate the second matching information corresponding to the image feature and the point cloud feature as q I i , achieving image feature alignment 412 (ie, Image-PCD Aligned).

[0091] In addition, if Figure 4 As shown, you can also replace the above { h P i,l} L l=1 Input a learnable multilayer perceptron 404 (Multilayer Perceptron, MLP for short) to directly calculate the matching score corresponding to the third matching information from the point cloud q P i , in case no image or text is available.

[0092] Therefore, based on the first matching information q T i , Second matching information q I i And the third matching information q P i and the above preset best matching vector y i , we can construct the target loss function as shown in Formula 2 below:

[0093] Where, is the standard deviation. The above objective loss function can be used to achieve better supervised learning of the model and ensure accurate and efficient acquisition of optimal asset data.

[0094] In order to better align point cloud features with image or text features in different scenes and object instances, we further create a new candidate set P i ′ to add additional supervisory signals. The candidate set can consist of the original best candidate and candidates randomly sampled from different scenes. Similar matching information (such as matching score) is calculated according to formula 1 q T i 'and q I i ′, forming the preset auxiliary loss function shown in the following formula 3:

[0095] Therefore, the final asset retrieval function as the learning objective can be expressed as follows:

[0096] In the above formula 4, the asset retrieval function can be used as the retrieval basis for the optimal asset data, and also defines the above preset multimodal alignment rule. Therefore, the optimal asset data can be the best candidate asset of the three-dimensional virtual scene that can ensure the fit with the original real object scene through the posture and restore the real scene to the greatest extent.

[0097] In summary, the above-mentioned generation optimization method of the three-dimensional virtual scene in the embodiment of the present invention mainly uses the model architecture based on image, text and point cloud information to select the most suitable three-dimensional asset to replace the target scanned object. It can significantly improve the matching accuracy through multimodal feature learning. Figure 4 As shown in the figure. For the target object in the scene, a quadruple containing an image, a text description, multiple candidate point clouds and the best matching annotation is constructed, and the corresponding features are extracted using a frozen image and text encoder, and the features of the candidate point cloud are extracted through a learnable three-dimensional encoder. Furthermore, in order to calculate matching information such as matching scores, multimodal contrastive learning is used to calculate the similarity between point cloud features and image and text features. At the same time, a method based on an MLP network is introduced to directly predict matching scores from point cloud features to cope with the lack of images or text. In addition, in order to further improve the alignment effect in different scenes and object instances, the candidate asset set is expanded and auxiliary losses are introduced to enhance the matching ability between point clouds and images and texts using additional supervisory signals. Finally, the optimization objective of the model is composed of the main matching loss and the auxiliary loss, thereby improving the accuracy and generalization ability of asset retrieval.

[0098] like Figure 3A-Figure 4As shown, according to an embodiment of the present invention, in operation S203, object posture alignment processing is performed in a preset virtual scene according to the optimal asset data to generate a target virtual scene, including: In the preset virtual scene, the center position of the virtual object corresponding to the optimal asset data is translated to coincide with the center position of the real scene of the target object; In the preset virtual scene, performing asset scaling on the virtual object so that the longest side of the bounding box of the virtual object matches the real longest side of the target object; and In the preset virtual scene, rotation is performed around the preset central axis of the virtual object at a preset interval angle to complete the object posture alignment process and generate the target virtual scene.

[0099] like Figure 3B As shown, in the data annotation 302 process, a heuristic-based asset placement process can be used to accurately align the position, size, orientation and other posture information of the three-dimensional object model corresponding to the selected optimal asset data in the preset virtual scene with the posture information of the scanned target object in the real scene, and align the best retrieved asset to the preset virtual scene. The real scene is the real space environment scanned and detected by the intelligent body.

[0100] Specifically, Figure 3B The Annotation Interface shown can translate the center of the best retrieved asset to the center of the real scanned object; in addition, the asset can be scaled so that the longest side of the asset bounding box matches the longest side of the real scanned object. Moreover, the asset is rotated around the preset central axis of the asset at a preset interval angle such as 30 degrees to find the minimum rotation angle, thereby achieving the best alignment with the real scanned object. Among them, the virtual object is the three-dimensional object model corresponding to the above-mentioned optimal asset data in the preset virtual scene; the real scanned object is the target object in the real scene.

[0101] Therefore, through the above-mentioned process of automatic alignment of the complete object posture of the optimal asset data in the preset virtual scene, it is ultimately possible to achieve three-dimensional asset retrieval and replacement based on multimodal alignment, place the best candidate asset in the three-dimensional scene, and adjust its direction and size according to actual conditions to ensure that it fits the original object (Best Match) and restores the real scene to the greatest extent.

[0102] It can be seen that through the above-mentioned algorithm framework for automatically constructing a virtual copy of the real-world 3D scan scene, and utilizing a powerful multimodal alignment model, the most suitable replacement object can be selected from the candidate 3D assets to achieve precise alignment of their position, size and orientation.

[0103] like Figure 3A-Figure 4 As shown, according to an embodiment of the present invention, in operation S204, performing scene physics optimization of a target object on a target virtual scene includes: Performing physical constraint optimization on the spatial position relationship of the target virtual scene according to the hierarchical scene data corresponding to the target object; Through the physical simulation environment, target physical properties are added to the target objects in the target virtual scene after physical constraint optimization to complete the scene physics optimization.

[0104] Hierarchical scene data can be the storage data of a three-dimensional hierarchical scene graph constructed based on scene point cloud technology, which can be presented in the form of table information to reflect all spatial relationship information of the target object in the scene, such as support, embedding and inclusion. Among them, hierarchical scene data can be used for physical constraint optimization (PhysicalOptimization) of the target object, and the above spatial relationship information can be used as optimization constraint conditions. Among them, physical constraint optimization is a constraint on the physical spatial relationship between the target object and the virtual scene, aiming to make the spatial relationship of the target object in the virtual scene more consistent with the spatial physical relationship of the real scene.

[0105] like Figure 3C As shown, the virtual objects existing in the preset virtual scene will actually have various spatial relationships with the virtual scene. In many cases, these spatial relationships will appear in various states that do not conform to the spatial physical relationship, such as the target object exceeds the scene range of the virtual scene (i.e., Outside), the target object is suspended in the virtual scene and has no contact with other objects (i.e., Floating), or the target object has at least partial overlap with other objects (i.e., Collision), etc.

[0106] To achieve the above physical constraint optimization for the target object, an optimization sampling technique such as Markov Chain Monte Carlo (MCMC) can be selected for optimization (e.g. Figure 3C The SSG & MCMC optimization 331 shown in FIG. 3 is used to fine-tune the horizontal position relationship of the target object in the preset virtual scene, such as optimization of penetration, contact, etc. Therefore, an optimization trade-off can be achieved between the scene graph information provided by the hierarchical scene data and physical conflicts (such as collisions), guiding the optimization process to adjust the final position of the target object to meet physical rationality.

[0107] Furthermore, the three-dimensional virtual scene optimized by the above physical constraints is imported into a physical simulator such as Blender (e.g. Figure 3CThe simulator 332 shown in the figure adds target physical properties (such as material type, quality and other information) to the corresponding target object through the physical simulation environment of the physical simulator to enhance the physical reality of the reconstructed scene and make it more suitable for simulation and interactive tasks.

[0108] In summary, compared with the traditional optimization scheme that directly uses gradient-based methods to achieve complex constraint optimization, the embodiments of the present invention achieve lower computing costs through a method of global optimization of virtual scenes based on physical constraints, and can further ensure the physical rationality of objects in the scene; and by introducing physical simulation for scene optimization, it can better ensure that objects in the virtual scene comply with physical laws (such as stability, collision detection, etc.), thereby improving the authenticity and practicality of virtual copies.

[0109] In summary, the above-mentioned intelligent body three-dimensional virtual scene generation and optimization method of the embodiment of the present invention can automatically construct highly realistic and simulatable three-dimensional scene copies, improve the environmental realism and interaction reliability of embodied intelligence research, and at the same time reduce human intervention, thereby achieving large-scale, efficient, and high-quality three-dimensional scene generation and optimization.

[0110] Based on the above-mentioned intelligent three-dimensional virtual scene generation optimization method, the present invention also provides an intelligent three-dimensional virtual scene generation optimization device. Figure 5 The device is described in detail.

[0111] Figure 5 The structure block diagram of the generation and optimization device of the intelligent three-dimensional virtual scene according to an embodiment of the present invention is schematically shown.

[0112] like Figure 5 As shown, the generation and optimization device 500 of the intelligent body three-dimensional virtual scene of this embodiment includes a data generation module 510, an asset retrieval module 520, a posture alignment module 530 and a physical optimization module 540.

[0113] The data generation module 510 is used to generate candidate asset data corresponding to the target object in the identified target scene image. In one embodiment, the data generation module 510 can be used to perform the operation S201 described above, which will not be described in detail here.

[0114] The asset retrieval module 520 is used to retrieve the optimal asset data corresponding to the target object based on the preset multimodal alignment rules through the text description information corresponding to the target object, the mask segmentation image and the candidate asset data. In one embodiment, the asset retrieval module 520 can be used to perform the operation S202 described above, which will not be repeated here.

[0115] The posture alignment module 530 is used to perform object posture alignment processing in a preset virtual scene according to the optimal asset data to generate a target virtual scene. In one embodiment, the posture alignment module 530 can be used to perform the operation S203 described above, which will not be repeated here.

[0116] The physical optimization module 540 is used to perform scene physical optimization corresponding to the target object on the target virtual scene to complete the optimization process of the virtual scene. In one embodiment, the physical optimization module 540 can be used to perform the operation S204 described above, which will not be repeated here.

[0117] According to an embodiment of the present invention, any multiple modules of the data generation module 510, the asset retrieval module 520, the posture alignment module 530 and the physical optimization module 540 can be combined into one module for implementation, or any one of the modules can be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module. According to an embodiment of the present invention, at least one of the data generation module 510, the asset retrieval module 520, the posture alignment module 530 and the physical optimization module 540 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application specific integrated circuit (ASIC), or can be implemented by hardware or firmware such as any other reasonable way of integrating or packaging the circuit, or implemented in any one of the three implementation methods of software, hardware and firmware or in a suitable combination of any of them. Alternatively, at least one of the data generation module 510, the asset retrieval module 520, the posture alignment module 530, and the physical optimization module 540 may be at least partially implemented as a computer program module, which may perform corresponding functions when executed.

[0118] Figure 6 A block diagram of an electronic device suitable for implementing a method for optimizing generation of a three-dimensional virtual scene of an intelligent agent according to an embodiment of the present invention is schematically shown.

[0119] The electronic device provided by an embodiment of the present invention includes one or more processors and a memory, the memory being used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors execute the above-mentioned method for optimizing the generation of a three-dimensional virtual scene of an intelligent body.

[0120] like Figure 6As shown, the electronic device 600 according to an embodiment of the present invention includes a processor 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage part 608 to a random access memory (RAM) 603. The processor 601 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 601 may also include an onboard memory for caching purposes. The processor 601 may include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present invention.

[0121] In RAM 603, various programs and data required for the operation of electronic device 600 are stored. Processor 601, ROM 602 and RAM 603 are connected to each other via bus 604. Processor 601 performs various operations of the method flow according to the embodiment of the present invention by executing the program in ROM 602 and / or RAM 603. It should be noted that the program can also be stored in one or more memories other than ROM 602 and RAM 603. Processor 601 can also perform various operations of the method flow according to the embodiment of the present invention by executing the program stored in the one or more memories.

[0122] According to an embodiment of the present invention, the electronic device 600 may further include an input / output (I / O) interface 605, which is also connected to the bus 604. The electronic device 600 may further include one or more of the following components connected to the I / O interface 605: an input portion 606 including a keyboard, a mouse, etc.; an output portion 607 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage portion 608 including a hard disk, etc.; and a communication portion 609 including a network interface card such as a LAN card, a modem, etc. The communication portion 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as needed. A removable medium 611, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 610 as needed, so that a computer program read therefrom is installed into the storage portion 608 as needed.

[0123] The present invention also provides a computer-readable storage medium on which executable instructions are stored. When the instructions are executed by a processor, the processor executes the above-mentioned method for optimizing the generation of a three-dimensional virtual scene of an intelligent body.

[0124] The computer-readable storage medium may be included in the device / apparatus / system described in the above embodiment; or it may exist independently without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the method according to the embodiment of the present invention is implemented.

[0125] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, the computer-readable storage medium may be any tangible medium containing or storing a program, which may be used by or in combination with an instruction execution system, an apparatus or a device. For example, according to an embodiment of the present invention, the computer-readable storage medium may include the ROM 602 and / or RAM 603 described above and / or one or more memories other than ROM 602 and RAM 603.

[0126] An embodiment of the present invention also includes a computer program product, which includes a computer program, and when the computer program is executed by a processor, the above-mentioned intelligent body three-dimensional virtual scene generation optimization method is implemented.

[0127] The computer program includes program codes for executing the method shown in the flow chart. When the computer program product is run in a computer system, the program codes are used to enable the computer system to implement the method provided by the embodiment of the present invention.

[0128] The computer program executes the above functions defined in the system / device of the embodiment of the present invention when it is executed by the processor 601. According to the embodiment of the present invention, the system, device, module, unit, etc. described above can be implemented by a computer program module.

[0129] In one embodiment, the computer program may rely on tangible storage media such as optical storage devices, magnetic storage devices, etc. In another embodiment, the computer program may also be transmitted and distributed in the form of signals on a network medium, and downloaded and installed through the communication part 609, and / or installed from a removable medium 611. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0130] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 609, and / or installed from the removable medium 611. When the computer program is executed by the processor 601, the above functions defined in the system of the embodiment of the present invention are performed. According to the embodiment of the present invention, the system, device, means, module, unit, etc. described above can be implemented by a computer program module.

[0131] According to an embodiment of the present invention, the program code for executing the computer program provided by the embodiment of the present invention can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level process and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, Java, C++, python, "C" language or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on the remote computing device, or entirely on the remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., using an Internet service provider to connect through the Internet).

[0132] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present invention. In this regard, each box in the flow chart or block diagram can represent a module, a program segment, or a part of a code, and the above-mentioned module, program segment, or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flow chart, and the combination of the boxes in the block diagram or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0133] In addition, all actions of acquiring information, signals or data in the present invention are carried out in compliance with the corresponding data protection laws, regulations and policies of the country where they are located, and with the authorization given by the owner of the corresponding device.

[0134] It will be appreciated by those skilled in the art that the features described in the various embodiments and / or claims of the present invention may be combined and / or combined in various ways, even if such combinations and / or combinations are not explicitly described in the present invention. In particular, the features described in the various embodiments and / or claims of the present invention may be combined and / or combined in various ways without departing from the spirit and teachings of the present invention. All of these combinations and / or combinations fall within the scope of the present invention.

[0135] The embodiments of the present invention are described above. However, these embodiments are only for the purpose of illustration, and are not intended to limit the scope of the present invention. Although each embodiment is described above, it does not mean that the measures in each embodiment cannot be used in combination. The scope of the present invention is defined by the attached claims and their equivalents. Without departing from the scope of the present invention, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present invention.

Claims

1. A method for optimizing the generation of an intelligent agent virtual scene, characterized in that: include: Generate candidate asset data corresponding to the target object in the identified target scene image; Based on a preset multimodal alignment rule, the optimal asset data corresponding to the target object is retrieved through the text description information corresponding to the target object, the mask segmentation image and the candidate asset data; Performing object posture alignment processing in a preset virtual scene according to the optimal asset data to generate a target virtual scene; as well as The target virtual scene is subjected to scene physical optimization corresponding to the target object, thereby completing the generation optimization process of the three-dimensional virtual scene.

2. The method according to claim 1, characterized in that Before generating candidate asset data corresponding to the target object in the identified target scene image, the method further includes: Extracting a target perspective image set of the target object in the preset virtual scene; The target scene image is extracted from the target perspective image set.

3. The method according to claim 1, characterized in that: The candidate asset data corresponding to the target object in the generated and identified target scene image includes: Acquire text description information and a mask segmentation image corresponding to the target object according to the target scene image; The candidate asset data is generated according to the text description information and the mask segmentation image.

4. The method according to claim 3, characterized in that In the step of acquiring text description information and mask segmentation image corresponding to the target object according to the target scene image, the step includes: Extracting and completing the mask segmentation image of the target scene image; and The text description information corresponding to the mask segmented image is generated by a preset language model.

5. The method according to claim 3, characterized in that: The step of generating the candidate asset data according to the text description information and the mask segmentation image includes: Generate first candidate data according to the text description information; generating second candidate data according to the mask segmentation image; and Retrieving the third candidate data according to the text description information; The candidate asset data includes the first candidate data, the second candidate data and the third candidate data.

6. The method according to claim 1, characterized in that In the retrieving the optimal asset data corresponding to the target object based on the preset multimodal alignment rule through the text description information corresponding to the target object, the mask segmentation image and the candidate asset data, the method includes: extracting text features of the text description information, image features of the mask segmentation image, and point cloud features of the candidate asset data; Generate first matching information corresponding to the text feature and the point cloud feature and second matching information corresponding to the image feature and the point cloud feature; Generate a target loss function based on the first matching information and the second matching information according to the third matching information corresponding to the point cloud feature and the preset best matching vector; Generate an asset retrieval function by combining a preset auxiliary loss function and the target loss function; The optimal asset data is retrieved from the candidate asset data according to the asset retrieval function.

7. The method according to claim 1, characterized in that In performing object posture alignment processing in a preset virtual scene according to the optimal asset data to generate a target virtual scene, the method includes: In the preset virtual scene, the center position of the virtual object corresponding to the optimal asset data is translated to coincide with the center position of the real scene of the target object; In the preset virtual scene, performing asset scaling on the virtual object so that the longest side of the bounding box of the virtual object matches the real longest side of the target object; and In the preset virtual scene, rotation is performed around the preset central axis of the virtual object at preset interval angles to complete the object posture alignment process and generate a target virtual scene.

8. The method according to claim 7, characterized in that In performing scene physics optimization corresponding to the target object on the target virtual scene, the method includes: Performing physical constraint optimization for spatial position relationship on the target virtual scene according to hierarchical scene data corresponding to the target object; The physical optimization of the scene is completed by adding target physical properties to the target object of the target virtual scene after the physical constraint optimization through the physical simulation environment.

9. A device for optimizing the generation of a three-dimensional virtual scene of an intelligent agent, characterized in that: include: A data generation module, used to generate candidate asset data corresponding to a target object in the identified target scene image; An asset retrieval module, configured to retrieve optimal asset data corresponding to the target object through text description information corresponding to the target object, the mask segmentation image and the candidate asset data based on a preset multimodal alignment rule; A posture alignment module, used to perform object posture alignment processing in a preset virtual scene according to the optimal asset data to generate a target virtual scene; as well as The physical optimization module is used to perform scene physical optimization corresponding to the target object on the target virtual scene to complete the generation optimization processing of the three-dimensional virtual scene.

10. An electronic device comprising: one or more processors; a memory for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors execute the method according to any one of claims 1 to 8.

11. A computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, causes the processor to execute the method according to any one of claims 1 to 8.

12. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Railway signal indoor equipment maintenance system based on augmented reality

    CN113989462A

  • Virtual content generation method and device, electronic equipment and storage medium

    CN114904270A

  • Three-dimensional virtual scene generation method and device, equipment, storage medium and product

    CN118279448A

  • Virtual scene generation method and system

    CN119166236A

  • Multi-modal digital twinning scene building method and device based on image semantic fusion

    CN119251687A

Cited By

  • Intelligent agent autonomous evolutionary algorithm for business scene and application system of intelligent agent autonomous evolutionary algorithm

    CN120689523A

  • Three-dimensional scene generation method and device, electronic equipment and storage medium

    CN120747362A

  • Method and device for determining alignment area in virtual and real light alignment, virtual and real light alignment method and device and medium

    CN121544671A