Modeling method of entity object and related device
By obtaining the content of operation instructions through human-computer interaction, the problem of task understanding ambiguity in modeling of neural networks and deep learning models is solved, and the accurate generation of virtual models is achieved.
Patent Information
- Application Number
- CN202410430950.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-10
- Publication Date
- 2025-10-17
AI Technical Summary
Existing neural network models and deep learning models may have ambiguous understanding of tasks when modeling physical objects, resulting in the execution of incorrect tasks and the generation of incorrect virtual models.
The first operation instruction content, including the first description information and the first position information, is obtained through human-computer interaction, indicating the intention and operation position of the interactive object to perform the creation operation, so as to accurately understand the task during modeling and generate a correct virtual model.
The generation of virtual models is controlled through human-computer interaction, which avoids the execution of incorrect tasks and ensures the accuracy of the generated virtual models.
Smart Images

Figure CN120807760A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computer, in particular to an entity object modeling method and related device. BACKGROUND
[0002] With the rapid development of today's digital technology, an entity object in the real world can be accurately converted into a virtual model of the virtual world, so as to dynamically simulate, monitor, analyze and control the entity object through digital means. For example, a virtual model of an entity object can be generated through digital twinning, and the entity object may, for example, be an industrial product, a city, a village, a vehicle, etc.
[0003] At present, an entity object can be modeled through a machine learning model to obtain a virtual model of the entity object. However, in this way, the neural network model may understand ambiguities for the task to be performed, thereby causing the task to be performed incorrectly and obtaining an incorrect virtual model. SUMMARY
[0004] To solve the above technical problems, the present application provides an entity object modeling method and related device to avoid performing incorrect tasks and obtaining incorrect virtual models.
[0005] The present application embodiment discloses the following technical solutions:
[0006] In one aspect, the present application embodiment provides an entity object modeling method, which comprises:
[0007] An image to be processed including an entity object to be modeled is obtained, and first operation instruction content is obtained, the first operation instruction content including first description information and first position information, the first description information being used to indicate an intention of an interactive object to perform a creation operation, and the first position information being used to indicate a position of the entity object to be modeled in the image to be processed to which the creation operation is directed;
[0008] According to the position of the entity object to be modeled in the image to be processed indicated by the first position information and the creation operation indicated by the first description information, the entity object to be modeled is modeled to obtain a virtual model of the entity object to be modeled.
[0009] In one aspect, the present application embodiment provides an entity object modeling device, which comprises an acquisition unit and a modeling unit:
[0010] The acquisition unit is configured to acquire a to-be-processed image including a to-be-modeled entity object, and acquire first operation instruction content, the first operation instruction content including first description information and first position information, the first description information being used to indicate an intention of an interactive object to perform a creation operation, and the first position information being used to indicate a position of the to-be-modeled entity object in the to-be-processed image to which the creation operation is directed;
[0011] The modeling unit is configured to model the to-be-modeled entity object according to the to-be-processed image, a position of the to-be-modeled entity object in the to-be-processed image indicated by the first position information, and the creation operation indicated by the first description information, to obtain a virtual model of the to-be-modeled entity object.
[0012] In an aspect, an embodiment of the present application provides a computer device, the computer device comprising a processor and a memory:
[0013] The memory is configured to store a computer program and transmit the computer program to the processor.
[0014] The processor is configured to execute the method according to the instructions in the computer program.
[0015] In an aspect, an embodiment of the present application provides a computer readable storage medium, the computer readable storage medium being configured to store a computer program, the computer program causing a processor to execute the method according to any one of the preceding aspects when the computer program is executed by the processor.
[0016] In an aspect, an embodiment of the present application provides a computer program product, comprising a computer program, the computer program being executed by a processor to implement the method according to any one of the preceding aspects.
[0017] It can be seen from the technical solution that in the embodiment, the generation of the virtual model is controlled by human-computer interaction when the entity object to be modeled is modeled. Specifically, the image to be processed including the entity object to be modeled and the first operation instruction content are acquired, and the first operation instruction content includes the first description information and the first position information, so that the task to be performed on the image to be processed can be accurately understood based on the first operation instruction. The first description information is used to indicate the intention of the interactive object to perform the creation operation, and the first position information is used to indicate the position of the entity object to be modeled in the image to be processed. Therefore, the first description information and the first position information can reflect the real operation intention of the interactive object. Thus, the entity object to be modeled can be modeled according to the position of the entity object to be modeled in the image to be processed indicated by the first position information and the creation operation indicated by the first description information based on the image to be processed, and the virtual model of the entity object to be modeled is obtained. In the modeling process, the virtual model can be generated based on the image to be processed under the control of the real operation intention reflected by the first operation instruction content, so that the wrong task can be avoided. BRIEF DESCRIPTION OF DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0019] Figure 1 An application scenario architecture diagram of an entity object modeling method provided by the embodiment of the present application;
[0020] Figure 2 A flowchart of an entity object modeling method provided by the embodiment of the present application;
[0021] Figure 3 An acquisition method of first description information provided by the embodiment of the present application;
[0022] Figure 4 Another acquisition method of first description information provided by the embodiment of the present application;
[0023] Figure 5 An acquisition method of first position information provided by the embodiment of the present application;
[0024] Figure 6 A schematic diagram of a method for acquiring an image to be processed provided in an embodiment of the present application;
[0025] Figure 7 This is a diagram of the overall architecture of a multimodal large model provided in an embodiment of the present application;
[0026] Figure 8 A schematic diagram of a secondary type of semantic element provided in an embodiment of the present application;
[0027] Figure 9 A schematic diagram of displaying a virtual model of a physical object to be modeled provided in an embodiment of the present application;
[0028] Figure 10 A schematic diagram of editing a virtual model provided in an embodiment of the present application;
[0029] Figure 11 A schematic diagram of a city twin provided in an embodiment of the present application;
[0030] Figure 12 An architectural diagram of city twinning using a multimodal large model provided in an embodiment of the present application;
[0031] Figure 13 A schematic diagram of the overall process of city twinning provided in an embodiment of the present application;
[0032] Figure 14 A schematic diagram of a multimodal large model in a city twin scenario provided in an embodiment of the present application;
[0033] Figure 15 A schematic structural diagram of a device for modeling a physical object provided in an embodiment of the present application;
[0034] Figure 16 A structural diagram of a terminal provided in an embodiment of the present application;
[0035] Figure 17 A structural diagram of a server provided in an embodiment of the present application. DETAILED DESCRIPTION
[0036] The embodiments of the present application are described below with reference to the accompanying drawings.
[0037] In related technologies, physical objects in the real world can be converted into virtual models in the virtual world. Physical objects can be physical objects in the real world, and virtual models can be high-fidelity digital mappings that provide multi-dimensional, multi-temporal and multi-space scales for physical objects.
[0038] Currently, an entity object can be modeled by a machine learning model to obtain a virtual model of the entity object. The process of converting an entity object into a virtual model can be applied in a variety of different scenarios. For example, a neural network model can be used to construct a city in the real world as a virtual object based on digital twinning. Analyzing the digital virtual object is more accurate and efficient than directly analyzing the entity object. Digital twinning is a process of digitizing the real world, similar to generating an electronic version of a twin scene. Digital twinning can include processes such as semantic information extraction, model instantiation, model merging, etc.
[0039] However, whether it is to generate a virtual model based on a city in the real world or to edit an already generated virtual model, it is implemented by using the same neural network model. The neural network model has the ability to perform a variety of different tasks, which can cause the neural network model to understand ambiguities about the task that needs to be performed, and thus cause the wrong task to be performed, resulting in an incorrect virtual model.
[0040] For another example, a deep learning computed tomography (CT) image model can be used to generate a virtual object based on a CT image. The CT image can be a CT medical image, and a virtual object can be generated based on the CT medical image. The virtual object can be applied to medical processes such as disease detection, preoperative precise simulation planning, intraoperative navigation, and postoperative quantitative evaluation.
[0041] However, whether it is to generate a virtual model based on a CT medical image or to segment different parts of an already generated virtual model, it is implemented by using a deep learning computed tomography image model. The deep learning computed tomography image model has the ability to perform a variety of different tasks. When performing a task, the computed tomography image model can understand ambiguities about the task that needs to be performed, and thus cause the wrong task to be performed, resulting in an incorrect virtual model.
[0042] To solve the above technical problems, an embodiment of the present application provides a method for modeling a physical object. This method directly provides operation instructions when modeling the physical object, allowing the neural network model to correctly understand the task to be performed, avoid executing the wrong task, and ultimately obtain a correct virtual model. The method for modeling a physical object provided in an embodiment of the present application controls the generation of the virtual model through human-computer interaction when modeling the physical object to be modeled. Human-computer interaction refers to the process of information exchange between a person and a computer using a certain dialogue language and a certain interactive method to complete a specific task. When modeling the physical object to be modeled, a first operation instruction content can be obtained. The first operation instruction content includes first description information and first position information. The first description information is used to indicate the intention of the interactive object to perform the creation operation, and the first position information is used to indicate the position of the physical object to be modeled in the image to be processed for the creation operation. When modeling the physical object, the first description information and the first position information are combined based on the image to be processed corresponding to the physical object. The virtual model is generated by controlling the intention of the creation operation reflected by the first description information and the position of the physical object to be modeled in the image to be processed reflected by the first position information, thereby avoiding executing the wrong task and obtaining an incorrect virtual model.
[0043] It should be noted that the entity object modeling method provided in the embodiments of this application can be applied to various scenarios for modeling entity objects, including but not limited to cloud technology, artificial intelligence, smart transportation, assisted driving, etc. In these scenarios, specific modeling scenarios based on digital twins and modeling based on sensor data may be involved. The embodiments of this application will be introduced using the modeling scenario based on digital twins as an example. Modeling scenarios based on digital twins can include scenarios such as city twins and medical model generation. City twins can be large-scale digital twins at the city level, and city twins can be the process of generating virtual models based on real-world cities.
[0044] The entity object modeling method provided in the embodiments of the present application can be executed by a computer device, which can be, for example, a server or a terminal. The server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. Terminals include but are not limited to smartphones, computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, etc.
[0045] like Figure 1 As shown, Figure 1 An application scenario architecture diagram of a modeling method for an entity object is shown, and the application scenario is introduced as a computer device being a server.
[0046] The application scenario can include a server 100 and a terminal 200. The server 100 can be a stand-alone physical server, a server cluster composed of multiple physical servers, or a distributed system, or a cloud server providing cloud computing services. The terminal 200 includes, but is not limited to, a smart phone, a computer, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, an aircraft, and the like. The terminal 200 can install a modeling application, which can model an entity object to obtain a virtual model through a modeling function. The server 100 can provide modeling services for the modeling application on the terminal 200 to obtain the virtual model. The server 100 and the terminal 200 can be directly or indirectly connected through wired or wireless communication, which is not limited in the present application. For example, the server 100 and the terminal 200 can be connected through a network, which can be a wired or wireless network.
[0047] The terminal 200 is used to interact with an interactive object. The interactive object can open the modeling application on the terminal 200, send first operation instruction content to the server 100 through the terminal 200, and the server 100 receives the first operation instruction content. The server 100 can generate a virtual model based on the first operation instruction content and the to-be-processed image, and return the virtual model to the terminal 200, so as to display the virtual model to the interactive object on the terminal 200.
[0048] Specifically, the server 100 can obtain a to-be-processed image including a to-be-modeled entity object, and obtain first operation instruction content. The first operation instruction content includes first description information and first position information. The first description information is used to indicate the intention of the interactive object to perform a creation operation, and the first position information is used to indicate the position of the to-be-modeled entity object in the to-be-processed image. The to-be-modeled entity object can be an entity object waiting to be modeled, such as a city, a village, an industrial product, a vehicle, and the like. The to-be-processed image can be an image waiting to be processed including the to-be-modeled entity object. The to-be-processed image can be obtained by image acquisition on the entity object including the to-be-modeled entity object.
[0049] The operation instruction can be an instruction for the server 100 to correctly perform a task, and in some cases can reflect an instruction to perform an operation on a certain entity object. In the embodiments of the present application, the operation instruction can be an instruction to perform a creation operation on a certain entity object, an instruction to perform an editing operation on a certain entity object, and the like. The operation instruction can include description information and position information, the description information can be an intention to perform an operation, and the position information can indicate the position of the entity object on which the operation is performed. The description information can be, for example, an instruction to perform a creation operation or an instruction to perform an editing operation. The operation instruction content can reflect the content of the operation instruction, and different operation instruction contents can correspond to different operation instructions, for example, the first operation instruction content. The first operation instruction content can be content for indicating the intention of the interactive object to perform a creation operation on the entity object to be modeled, the first operation instruction content can include first description information and first position information, and the first operation instruction content can be sent by the terminal 200 to the server 100. Since the first description information in the first operation instruction content indicates the intention of the interactive object to perform a creation operation, and the first position information indicates the position of the entity object to be modeled in the to-be-processed image on which the creation operation is performed, the server 100 can accurately understand the task required to be performed on the to-be-processed image based on the first operation instruction content, thereby avoiding performing an incorrect task.
[0050] The server 100 models the entity object to be modeled according to the to-be-processed image, the position of the entity object to be modeled in the to-be-processed image indicated by the first position information, and the creation operation indicated by the first description information, to obtain a virtual model of the entity object to be modeled. Since the first operation instruction content including the first description information and the first position information is used when modeling the entity object to be modeled, the first description information indicates the intention of the interactive object to perform a creation operation, and the first position information indicates the position of the entity object to be modeled in the to-be-processed image on which the creation operation is performed, the server 100 can more accurately understand the task required to be performed on the to-be-processed image based on the first operation instruction content. That is, the virtual model is generated under the control of the intention reflected by the first operation instruction content, thereby avoiding performing an incorrect task.
[0051] After generating the virtual model, the server 100 can return the virtual model to the terminal 200, so that the terminal 200 can perform operations such as dynamic simulation, monitoring, analysis, and control on the virtual model.
[0052] Since the embodiment of the application controls the generation of the virtual model through human-computer interaction when modeling the entity object to be modeled. The first operation instruction content used when modeling the entity object indicates the intention of the interactive object to perform a creation operation on the entity object to be modeled. The first description information in the first operation instruction content indicates the intention to perform the creation operation, and the first position information indicates the position of the entity object to be modeled in the to-be-processed image to which the creation operation is directed. Generating the virtual model under the control of the real operation intention reflected by the first operation instruction content can avoid performing incorrect tasks. By inputting the first operation instruction content of the interactive object through human-computer interaction, the task required to be performed on the to-be-processed image can be more accurately understood based on the first operation instruction content, thereby avoiding incorrect tasks and obtaining incorrect virtual models.
[0053] It should be noted that the method provided by the embodiment of the application can involve artificial intelligence technology. Artificial intelligence (AI) is the use of digital computers or digital computer-controlled machines to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science, which aims to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines, so that machines have the functions of perception, reasoning and decision-making.
[0054] Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, both hardware and software technologies. Artificial intelligence basic technologies generally include, such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-training model technology, operation / interaction system, mechatronics, etc. Among them, the pre-training model is also called large model, basic model, which can be widely used in downstream tasks of various directions of artificial intelligence after fine-tuning. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology and machine learning / deep learning, etc.
[0055] It can be understood that the modeling method of the entity object provided in the embodiments of the present application can involve natural language processing. Natural language processing (NLP) is an important direction in the field of computer science and the field of artificial intelligence. Natural language processing is a science integrating linguistics, computer science and mathematics. Therefore, the research in this field will involve natural language, i.e. the language used in daily life, so it has a close relationship with the research of linguistics. Natural language processing technology usually includes text processing, semantic understanding, machine translation, robot question and answer, knowledge graph and the like. In the process of modeling the entity object, the embodiments of the present application can use text processing, semantic understanding and the like.
[0056] In the modeling of the entity object, machine learning can also be involved. Machine learning (ML) is a multi-field interdisciplinary subject, involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and the like. It is a subject specially studying how a computer simulates or realizes the learning behavior of human beings to obtain new knowledge or skills, and reorganizes the existing knowledge structure to constantly improve the performance. Machine learning is the core of artificial intelligence and is the fundamental approach to making computers intelligent, and its application is widespread in various fields of artificial intelligence. Machine learning and deep learning usually include artificial neural network, belief network, reinforcement learning, transfer learning, inductive learning, rule teaching learning and the like. In the embodiments of the present application, machine learning can be used to train a model for extracting semantic information and a model for feature coding.
[0057] The pre-training model (PTM), also known as cornerstone model, large model, refers to a deep neural network (DNN) with large parameters, which is trained on a large amount of unlabeled data. The PTM extracts common features from the data using the function approximation capability of the large parameter DNN. Through fine tuning, parameter-efficient fine-tuning (PEFT), prompt tuning and other techniques, the PTM is suitable for downstream tasks. Therefore, the pre-training model can achieve ideal results in few-shot or zero-shot scenarios. PTM can be divided into language model, visual model, speech model and multi-modal model according to the data modality processed, wherein the language model can include word representation model (Embeddings from Language Models, ELMO), pre-training language representation model (Bidirectional Encoder Representation from Transformers, BERT) or natural language generation model (Generative Pre-training Transformer, GPT), wherein the visual model can include image segmentation model (Swin-Transformer, ST), image classification model (Vision Transformer, ViT) or visual model (Vision MoE, V-MOE), wherein the speech model can include text-to-speech model (Voice Allocation Language and Expression, VALL-E), wherein the multi-modal model can include visual language multi-modal model (Vision-and-Language Bidirectional Encoder Representation from Transformers, ViBERT), image text multi-modal model (Contrastive Language-Image Pre-training, CLIP), few-shot visual language model (Flamingo) or multi-modal intelligent model (Generalist-Agent, Gato), wherein the multi-modal model refers to a model that establishes feature representation of two or more data modalities. The pre-training model is an important tool for outputting artificial intelligence generated content (AIGC), and can also be used as a general interface connecting multiple specific task models.
[0058] It should be noted that in the specific embodiments of the present application, user information and other related data may be involved throughout the process. When the above embodiments of the present application are applied to specific products or technologies, the individual consent or individual license of the user needs to be obtained, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards of the country and region.
[0059] Next, taking a computer device as an example, the modeling method of an entity object provided by the embodiments of the present application will be introduced in combination with the accompanying drawings. Referring to FIG. 1, Figure 2 , Figure 2 a flowchart of a modeling method of an entity object is shown, and the method comprises:
[0060] S201, obtaining a to-be-processed image comprising a to-be-modeled entity object, and obtaining first operation instruction content.
[0061] When the interactive object uses an application with a modeling function, the interactive object can trigger a modeling request for the to-be-modeled entity object through the terminal. The server can model the to-be-modeled entity object in response to the modeling request for the to-be-modeled entity object sent by the terminal. When modeling the to-be-modeled entity object, the server can obtain a to-be-processed image comprising the to-be-modeled entity object based on the modeling request for the to-be-modeled entity object, and obtain first operation instruction content.
[0062] The first operation instruction content comprises first description information and first position information, the first description information is used to indicate the intention of the interactive object to perform a creation operation, and the first position information is used to indicate the position of the to-be-modeled entity object in the to-be-processed image to which the creation operation is directed. The description information can be natural language description information, and the first description information can be description information for describing the creation operation, which can be text information or voice information, etc. The position information can be information for indicating a position, which can be referred to as Prompt information in the embodiments of the present application. The Prompt information can be prompt information input by the interactive object in the process of interacting with the terminal, such as a point or region information in a picture. In order to adapt to different application scenarios, the position information can be represented in different forms, such as longitude and latitude, coordinates in a three-dimensional coordinate system, pixel points in the to-be-processed image, etc. The first position information can be used to indicate the position of the to-be-modeled entity object in the to-be-processed image to which the creation operation is directed.
[0063] The entity object to be modeled can be an entity object to be modeled, the image to be processed can be an image to be processed including the entity object to be modeled, and the image to be processed can be obtained by image acquisition on the entity object including the entity object to be modeled. It can be understood that the type of the entity object to be modeled can be different in different scenarios. For example, in the scenario of city twinning, the entity object to be modeled can be a residential area or a street in the city, the geographic entity can be an entity object with a certain geographic location, and the geographic entity can be a city, a village, or a street. In the scenario of generating a virtual object by a deep learning computer tomography image model, the entity object to be modeled can be a part of a physical entity in the computer tomography image, the physical entity can be an entity object with a certain physical structure, and the physical entity can be a vehicle or a device.
[0064] In different application scenarios, the type of the image to be processed can be different. For example, in the scenario of city twinning, the entity object to be modeled can be a city, and the image to be processed can be a satellite image including the entity object to be modeled. The satellite image can be an image fed back to a user by a satellite to reflect the real appearance of the earth's surface. In the scenario of generating a virtual object by a deep learning computer tomography image model, the image to be processed can be a computer tomography image including the entity object to be modeled. In the scenario of modeling a vehicle, the image to be processed can be a photo including the vehicle.
[0065] In the scenario of city twinning, the satellite image including the entity object to be modeled can be obtained in various ways, such as through various map software.
[0066] The operation instruction can be an instruction for controlling the server to correctly perform a task. In the embodiments of the present application, the first operation instruction content can be the content of the intention of the interactive object to perform a creation operation on the entity object to be modeled, and the first operation instruction content includes first description information and first position information. The interactive object can be an object capable of interacting with the terminal, and the interactive object can be a user, for example. The first description information can be information in the form of voice, text, etc. The first position information can be longitude and latitude, coordinates in a three-dimensional coordinate system, a pixel point in the image to be processed, etc.
[0067] In different scenarios, the first description information is used to indicate the intention of the interactive object to perform the creation operation, and the first position information is used to indicate the position of the entity object to be modeled in the to-be-processed image. For example, in a scenario of constructing a city in the real world into a virtual object based on digital twinning by using a neural network model, the first description information can indicate the intention of the interactive object to perform the creation operation, and the first position information can indicate the position of a certain residential area in the to-be-processed image. In a scenario of generating a virtual object by using a deep learning computerized tomography image model, the first description information can indicate the intention of the interactive object to perform the creation operation, and the first position information can indicate the position of a certain entity object in the computerized tomography image.
[0068] In the scenario of city twinning, the first description information can be in the form of text “reconstruct the twinned scene in the specified area” or in the form of voice “reconstruct the twinned scene in the specified area”, and the first position information can be “longitude 20° to 21°, latitude -4° to -5°”. The first position information can indicate the longitude and latitude coordinates of the specified area, and thus determine the specified area, so as to model the specified area as the entity object to be modeled.
[0069] It can be understood that the first description information and the first position information can be input by the interactive object through one interaction or through multiple interactions. For example, the interactive object can first input the first description information through one interaction, and then input the first position information through another interaction. By inputting the first description information and the first position information through multiple interactions, the interactive object can flexibly select an interaction mode for inputting different information.
[0070] The server can obtain the first description information and the first position information in multiple ways, so that there are many ways to obtain the first operation instruction content. In one possible implementation, the server can directly obtain the first description information input by the interactive object. The first description information can be text information or voice information. For example, the interactive object can input the first description information “reconstruct the twinned scene in the specified area” by text, and the server can directly obtain the first description information. For another example, the interactive object can input the first description information “reconstruct the vehicle in the photo” by voice, and the server can also directly obtain the first description information.
[0071] In another possible implementation, the server can obtain the first description information in response to the triggering of the creation operation control by the interactive object. That is, when the interactive object triggers the creation operation control, it indicates that the interactive object wants to perform the creation operation, and thus the first description information indicating the creation operation is generated.
[0072] Figure 3 A schematic diagram of a first description information acquisition method provided by an embodiment of the present application,Figure 3 The corresponding scenario is city twinning, Figure 3 The satellite image is above the city, and the interactive object performs human-computer interaction with the terminal through voice. The interactive object inputs "create a virtual model of the specified area" to the terminal through voice, and "create a virtual model of the specified area" is the first description information. The terminal can send the obtained "create a virtual model of the specified area" to the server, so that the server uses the description information to model.
[0073] In addition to inputting the first operation instruction content through voice, the interactive object can also input the first operation instruction content through text, and of course, editing operations such as adding, deleting, modifying, and canceling can be performed on the input text.
[0074] Figure 4 Another first description information acquisition method provided by the embodiment of the application is shown in the schematic diagram, Figure 4 The corresponding scenario is modeling a vehicle, Figure 4 The photo of the vehicle is above, Figure 4 The creation operation control in the figure includes a virtual button of "create a virtual model", the interactive object triggers the virtual button of "create a virtual model", and the server acquires the first description information "create a virtual model" in response to the triggering of the virtual button of "create a virtual model" by the interactive object. Figure 4 In a possible implementation, the server can directly acquire the first location information input by the interactive object. The first location information can correspond to one entity object to be modeled, or can correspond to multiple entity objects to be modeled. Taking the first location information as an example, the interactive object can input the first location information "longitude 20° to 21°, latitude -4° to -5°", and the server can directly acquire the first location information, which corresponds to the entity objects to be modeled between "longitude 20° to 21°, latitude -4° to -5°". For another example, the interactive object can input the first location information "longitude 20°, latitude -4°", and the first location information corresponds to the entity objects to be modeled at "longitude 20°, latitude -4°".
[0075] In another possible implementation, the server can acquire the first location information in response to the selection operation of the interactive object on the entity object to be modeled in the image to be processed. For example, in the scenario of city twinning, the interactive object selects a certain area of the image to be processed, and the server acquires the first location information of the area in response to the selection operation. The selection operation can be clicking or framing, etc.
[0076]
[0077] Another first location information acquisition method provided by the embodiment of the application is shown in the schematic diagram, Figure 5 Figure 5 The corresponding scenario is city twinning, Figure 5 The satellite map above, Figure 5 The various tools for selecting the entity object to be modeled are below, the interactive object selects the entity object to be modeled in the satellite map through the dashed rectangular frame tool, and the server acquires the first position information in response to the selection operation.
[0078] It should be noted that the first description information and the first position information in the first operation instruction content can be acquired simultaneously or sequentially. In a possible implementation, the server can acquire the first description information and the first position information simultaneously. For example, in the scenario of city twinning, the interactive object wants to perform twinning reconstruction on a region, and the interactive object can input the first description information "construct a twinning reconstruction scene of the selected region" and simultaneously input the longitude and latitude of the region. The server simultaneously acquires the first description information and the longitude and latitude input by the interactive object.
[0079] The embodiments of the present application provide various information in the first operation instruction content, and the interactive object can use a corresponding human-computer interaction mode. The interactive object can input the first description information in the form of text or voice, and input the first position information by clicking or frame selection. The interactive object only needs to use a simple method to input the first description information and the first position information through human-computer interaction, and the operation is very convenient. The human-computer interaction mode of the interactive object is very friendly.
[0080] The embodiments of the present application provide various ways to acquire the image to be processed. In a possible implementation, the server can directly use the original entity object image as the image to be processed. In some cases, the original entity object image can also include redundant entity objects, which are entity objects that do not need to be modeled. In order to more accurately model the entity object to be modeled and obtain a virtual model that meets the requirements of the interactive object, in another possible implementation, the server can first acquire the original entity object image, which is an image that has not been processed, and the entity object to be modeled is part of the entity objects included in the original entity image. Then, the server can determine the image corresponding to the first position information in the original entity object image as the image to be processed.
[0081] Figure 6 A schematic diagram of a method for acquiring an image to be processed according to an embodiment of the present application, Figure 6The corresponding scenario is city twinning. The leftmost satellite map is an original entity object image. Part of the entity objects included in the original entity image are to-be-modeled entity objects. In the middle satellite map, the first position information is obtained. The position represented by the first position information is the area shown by the dashed box, that is, the entity objects in the area are to-be-modeled entity objects. Therefore, in order to remove redundant entity objects, the server can determine the image corresponding to the first position information in the original entity object image as a to-be-processed image.
[0082] The embodiment of the present application determines the to-be-processed image through the first position information, thereby eliminating redundant entity objects and improving the modeling speed of the to-be-modeled entity objects.
[0083] S202, according to the to-be-processed image, the position of the to-be-modeled entity object in the to-be-processed image indicated by the first position information, and the creation operation indicated by the first description information, the to-be-modeled entity object is modeled to obtain a virtual model of the to-be-modeled entity object.
[0084] The server can model the to-be-modeled entity object according to the to-be-processed image, the position of the to-be-modeled entity object in the to-be-processed image indicated by the first position information, and the creation operation indicated by the first description information, to obtain a virtual model of the to-be-modeled entity object. The virtual model is generated under the control of the intention reflected by the first operation instruction content.
[0085] In the modeling process, in order to enable the server to better understand the to-be-processed image, the first description information and the first position information, the server can process the to-be-processed image and the first operation instruction content. In a possible implementation manner, the implementation manner of S202 can be that the server extracts semantic information from the to-be-processed image to obtain image semantic features, extracts semantic information from the first description information to obtain first operation semantic features, and encodes the first position information to obtain first position features. Then, the server models the to-be-modeled entity object according to the image semantic features, the first operation semantic features and the first position features to obtain a virtual model of the to-be-modeled entity object.
[0086] The semantic information extraction can be a process of extracting semantic information of different pixels in the to-be-processed image. The image semantic features obtained through the semantic information extraction can reflect the semantic information corresponding to different pixels in the to-be-processed image, so as to obtain related information such as deep meaning and interpretation of the to-be-modeled entity object, attributes of the to-be-modeled entity object, or components included in the to-be-modeled entity object and the relationship between the components, thereby providing the most direct semantic support for modeling. In the scenario of city twinning, the to-be-processed image is a satellite image, and the semantic information extraction on the satellite image can obtain image semantic features. The image semantic features can reflect semantic elements included in the satellite image, for example, buildings, vegetation, terrain, roads, and the like. Specifically, the satellite image includes buildings and roads, and the image semantic features obtained through the semantic information extraction on the satellite image can represent the buildings and roads in the satellite image.
[0087] After obtaining the first operation instruction content, the server can process the first operation instruction content to obtain the first operation instruction feature, so as to convert the first operation instruction content into a feature space for subsequent processing. Since the first operation instruction feature includes the first description information and the first position information, the server actually performs semantic information extraction on the first description information to obtain the first operation semantic feature, and performs feature encoding on the first position information to obtain the first position feature. That is, the first operation instruction feature includes the first operation semantic feature and the first position feature.
[0088] The order of the semantic information extraction on the to-be-processed image, the semantic information extraction on the first description information, and the feature encoding on the first position information is not limited in the embodiments of the present application. For example, the semantic information extraction on the to-be-processed image and the first description information can be performed simultaneously, or can be performed in a certain order.
[0089] In the embodiments of the present application, the server can perform semantic information extraction on the to-be-processed image in multiple ways to obtain image semantic features, and the manner of the semantic information extraction is not limited in the embodiments of the present application. In one possible implementation manner, the server can perform semantic information extraction on the to-be-processed image by using an algorithm capable of obtaining image semantic features.
[0090] In another possible implementation, the server can perform semantic information extraction on the to-be-processed image by using an image processing model to obtain image semantic features. The image processing model can be a model for processing images. In different scenarios, the image processed by the image processing model can be different. For example, when modeling a city, the to-be-processed image can be a satellite image, and the image processing model can be a trained image processing model that can process satellite images. When modeling a vehicle, the to-be-processed image can be a picture taken of the vehicle, and the image processing model can be a trained image processing model that can process pictures.
[0091] In the embodiment of the present application, the image processing model is used to perform semantic information extraction on the to-be-processed image. By using the image processing model, more accurate image semantic features can be obtained, and the virtual model generated based on the image semantic features has less difference from the to-be-modeled entity object.
[0092] In the embodiment of the present application, the server can process the first operation instruction content in multiple ways to obtain the first operation instruction feature, and the manner of feature encoding is not limited in the embodiment of the present application. In one possible implementation, the server can process the first operation instruction content by using an instruction processing model to obtain the first operation instruction feature. The instruction processing model can be a model for processing instructions, and the first operation instruction feature can reflect the intention of the interactive object to perform a creation operation on the to-be-modeled entity object.
[0093] It should be understood that the first operation instruction content can include first description information and first position information, and the processing of the first operation instruction content by the server is the processing of the first description information and the first position information. Specifically, the server can perform semantic information extraction on the first description information to obtain a first operation semantic feature, and perform feature encoding on the first position information to obtain a first position feature. The first operation semantic feature can reflect the semantics of the description class information of the creation operation, and the first position feature can reflect the position of the to-be-modeled entity object in the to-be-processed image.
[0094] In the embodiment of the present application, the server can perform semantic information extraction on the first description information in multiple ways to obtain the first operation semantic feature, and the manner of semantic information extraction is not limited in the embodiment of the present application. In one possible implementation, the server can perform semantic information extraction on the first description information by using an algorithm capable of performing semantic information extraction. In another possible implementation, the server can perform semantic information extraction on the first description information by using a description information processing model to obtain the first operation semantic feature.
[0095] Similarly, the server can encode the first location information by various manners to obtain the first location feature, and embodiments of the present application do not limit the manner of feature encoding. In a possible implementation manner, the server can encode the first location information by using a feature encoding algorithm. In another possible implementation manner, the server can encode the first location information by using a location information processing model to obtain the first location feature.
[0096] The various manners of extracting semantic information from the first description information and the various manners of encoding the first location information can be used in different combinations. When the first description information is extracted by using the description information processing model and the first location information is encoded by using the location information processing model, the instruction processing model can include the description information processing model and the location information processing model.
[0097] In embodiments of the present application, different kinds of information in the first operation instruction content are processed by corresponding manners, the semantic information of the first description information is extracted, and the first location information is encoded. The first operation semantic feature and the first location feature that are more consistent with the creation intention of the interactive object can be obtained by processing by the corresponding processing manner, and a virtual model that is more consistent with the creation intention of the interactive object can be obtained when modeling.
[0098] In a possible implementation manner, the implementation manner of S202 can be that the server fuses the image semantic feature, the first operation semantic feature, and the first location feature to obtain first fused features. The first fused features include the type of the operation that the interactive object wants to perform, the position of the entity object to be modeled for performing the creation operation, and the image semantic feature. Then, the server models the entity object to be modeled based on the first fused features to obtain the virtual model of the entity object to be modeled. The above method obtains the first fused features by feature fusion. The first fused features can represent the association and interaction between different features. By modeling by using the first fused features, the consistency of the virtual model as a whole and the accuracy of details are improved.
[0099] In embodiments of the present application, the server can model the entity object to be modeled by various manners to obtain the virtual model of the entity object to be modeled, and the present application does not limit the manner of modeling the entity object to be modeled. In a possible implementation manner, the server can model the entity object to be modeled by using a three-dimensional generation model according to the image semantic feature, the first operation semantic feature, and the first location feature to obtain the virtual model of the entity object to be modeled. The three-dimensional generation model can be a neural network model for generating a three-dimensional virtual model.
[0100] The three-dimensional generation model can better understand the intention of the interactive object to the creation operation of the entity object to be modeled in the first operation semantic feature and the first position feature during modeling, so that the three-dimensional generation model can correctly perform the task and obtain the virtual model.
[0101] In the city twin scenario, the three-dimensional generation model can model the entity object to be modeled according to the image semantic feature, the first operation semantic feature and the first position feature. The modeling process includes generation of three-dimensional buildings, generation of scene terrain, generation of vegetation and generation of roads. The three-dimensional generation model can generate virtual models of three-dimensional buildings, scene terrain, vegetation and roads respectively, and finally combine the above virtual models to obtain the virtual model of the entity object to be modeled.
[0102] In the image processing model, the description information processing model, the position information processing model and the three-dimensional generation model in the embodiments of the present application can be large models. At this time, the image processing model can be an image processing large model, the description information processing model can be a description information processing large model, the position information processing model can be a position information processing large model, and the image processing model, the description information processing model, the position information processing model and the three-dimensional generation model can be applied to the modeling process of the entity object to be modeled independently. In the process of modeling the entity object to be modeled, the above large models are used to process the data of the corresponding modal respectively.
[0103] In order to further speed up the modeling speed, the server can use a multi-modal large model to model the entity object to be modeled. The multi-modal large model can include a plurality of different modal artificial intelligence algorithms, such as natural language processing, picture recognition segmentation, style transfer, text-to-image, image-to-three-dimensional model and other related artificial intelligence algorithms. The multi-modal large model used in the embodiments of the present application at least includes an image processing large model, a description information processing model, a position information processing model and a three-dimensional generation large model. The image processing large model is used to process the image to be processed, and the image processing model is used to extract semantic information from the image to be processed to obtain image semantic features. The description information processing model is used to extract semantic information from the first description information, and the description information processing model is used to extract semantic information from the first description information to obtain the first operation semantic feature. The position information processing model is used to encode the first position information, and the position information processing model is used to encode the first position information to obtain the first position feature. The three-dimensional generation model can be responsible for generating a virtual model by integrating a plurality of different features, that is, modeling the entity object to be modeled by the three-dimensional generation model according to the image semantic feature, the first operation semantic feature and the first position feature, to obtain the virtual model of the entity object to be modeled.
[0104] Figure 7 An overall architecture diagram of a multi-modal large model provided for an embodiment of the present application. Figure 7 The large rectangle of the multi-modal large model, the image processing model, the description information processing model, the position information processing model, and the three-dimensional generation model are components of the multi-modal large model. The multi-modal large model can be a large model that processes multiple modalities at the same time. The to-be-processed image and the first operation instruction content are inputs of the multi-modal large model. The first operation instruction content includes first description information and first position information. Specifically, the to-be-processed image is an input of the image processing model in the multi-modal large model. The first description information in the first operation instruction content is an input of the description information processing model in the multi-modal large model. The first position information in the first operation instruction content is an input of the position information processing model in the multi-modal large model. The three-dimensional generation model in the multi-modal large model uses image semantic features, first operation semantic features, and first position features to model the to-be-modeled entity object to obtain a virtual model of the to-be-modeled entity object. When the multi-modal large model is used to model the to-be-modeled entity object, the image processing model, the description information processing model, the position information processing model, and the three-dimensional generation model in the multi-modal large model work in coordination. The multi-modal large model can process data of different modalities to generate a virtual model. The modality corresponding to the to-be-processed image is an image. The modalities corresponding to the description information processing model and the position information processing model are natural languages.
[0105] Unlike traditional models, the multi-modal large model can process data of multiple modalities, has stronger semantic information extraction and feature encoding capabilities, and can more quickly process the to-be-processed image and the first operation instruction content through the multi-modal large model to obtain more accurate image semantic features and first operation instruction features.
[0106] It should be noted that the training and use of the multi-modal large model have certain requirements for hardware devices. Generally, it is recommended that the graphics processing unit (GPU) memory of the hardware device used to train and use the multi-modal large model be greater than or equal to 16 GB.
[0107] It can be seen from the technical solution that, in the embodiment, the generation of the virtual model is controlled by human-computer interaction when the entity object to be modeled is modeled. Specifically, the image to be processed including the entity object to be modeled and the first operation instruction content are acquired, and the first operation instruction content includes first description information and first position information, so that the task to be performed on the image to be processed can be accurately understood based on the first operation instruction. The first description information is used to indicate the intention of the interactive object to perform a creation operation, and the first position information is used to indicate the position of the entity object to be modeled in the image to be processed. Therefore, the first description information and the first position information can reflect the real operation intention of the interactive object. Thus, the entity object to be modeled can be modeled according to the position of the entity object to be modeled in the image to be processed indicated by the first position information and the creation operation indicated by the first description information based on the image to be processed, and the virtual model of the entity object to be modeled is obtained. In the modeling process, the virtual model can be generated based on the image to be processed under the control of the real operation intention reflected by the first operation instruction content, thereby avoiding the execution of an incorrect task. The first operation instruction content of the interactive object is input by human-computer interaction in the embodiment, so that the task to be performed on the image to be processed can be more accurately understood based on the first operation instruction content, thereby avoiding the execution of an incorrect task and obtaining an incorrect virtual model.
[0108] The entity object to be modeled can include semantic elements, which can represent the semantics of the entity object to be modeled in the image to be processed. In order to accurately model the entity object to be modeled, the server needs to distinguish the semantic elements corresponding to the entity object to be modeled, and the semantic elements corresponding to the entity object to be modeled can be different in different scenarios. For example, in the scenario of city twinning, the entity object to be modeled is a city, which can include elements such as buildings, roads, and vegetation. In order to accurately model the city, the server needs to distinguish the semantic elements of each part in the image to be processed. For another example, in the scenario of modeling a vehicle, the entity object to be modeled is a vehicle, which can have attributes such as driving type, vehicle type, and color. In order to accurately model the vehicle, the server also needs to distinguish the semantic elements of each part in the image to be processed. In order to enrich the semantic elements corresponding to the entity object to be modeled and model the entity object to be modeled in more detail to obtain a virtual model that meets the requirements of the interactive object, the semantic elements can be classified in more detail in the embodiment, so that an image processing model capable of classifying semantic elements in more detail can be trained.
[0109] Specifically, the server can obtain a plurality of sample images and a semantic label of each sample image in the plurality of sample images, the semantic label being used to indicate a secondary type of a semantic element in the sample image, semantic elements of a same primary type being classified into a plurality of secondary types. The server trains an initial network model by using the plurality of sample images and the respective semantic labels corresponding to the plurality of sample images to obtain the image processing model.
[0110] The primary type can be obtained by classifying entity objects according to unique characteristics and functions of the entity objects in a macroscopic manner, and the secondary type can be obtained by classifying entity objects of a certain primary type based on a certain attribute of the entity objects. For example, trees are a primary type, and large trees, small trees, dry trees, arbors, and shrubs are a plurality of secondary types of the same primary type. For another example, building materials are a primary type, and steel, bricks, and concrete are a plurality of secondary types of the same primary type.
[0111] Figure 8 A schematic diagram of a secondary type of a semantic element provided by an embodiment of the present application, Figure 8 The semantic element of the sample image in the upper left corner is a building, and the semantic element of the sample image in the upper right corner is a villa. Houses are a primary type, and buildings and villas are secondary types of the houses. The semantic element of the sample image in the lower left corner is a crossroad, and the semantic element of the sample image in the lower right corner is a single-lane road. Roads are a primary type, and crossroads and single-lane roads are secondary types of the roads.
[0112] An embodiment of the present application trains an image processing model by using sample images and semantic labels indicating secondary types of semantic elements in the sample images. The image processing model can more finely distinguish types of the semantic elements, thereby increasing richness of semantic information and enriching modeling effects of an entity object to be modeled.
[0113] Since modeling is performed by the server, an interactive object can hardly intuitively understand a virtual model obtained by the server. Therefore, the virtual model can be visually displayed. Visualization is a tool for intuitively displaying data, and the server can visually display the virtual model of the entity object to be modeled. The present application does not limit a manner in which the server displays the virtual model of the entity object to be modeled. For example, the server can display the virtual model of the entity object to be modeled by using an external display, or the server can send the virtual model of the entity object to be modeled to a terminal to display the virtual model of the entity object to be modeled by the terminal. The virtual model can be displayed in various forms, and the present application does not limit a display form of the virtual model. For example, the virtual model can be displayed by using a webpage or by using Unreal Engine (UE).
[0114] Figure 9A display schematic of a virtual model of a to-be-modeled entity object is provided for an embodiment of the present application, Figure 9 The virtual model in the display schematic is displayed in a terminal, and the virtual model displayed in the terminal is sent to the terminal by a server, Figure 9 The virtual model visualized in the display schematic is a building.
[0115] Through the display of the virtual model of the to-be-modeled entity object, the interactive object can intuitively see the virtual model, thereby bringing the most intuitive experience to the interactive object.
[0116] After obtaining the virtual model of the to-be-modeled entity object, the server can store the virtual model in the form of a file for subsequent application. The virtual model has a wide range of application scenarios, and the requirements for the output format of the virtual model are different in different application scenarios. Based on this, in a possible implementation manner, the server can first determine the file output type according to the application scenario of the virtual model. After determining the file output type, the server exports the model file of the virtual model according to the file output type.
[0117] For example, the application scenario of the virtual model is game development, the server determines the file output type as a game scenario file according to the application scenario of the virtual model, and the server can export the virtual model as a UE resource, which can be applied to game development.
[0118] For another example, the application scenario of the virtual model is a city traffic scenario, the server determines the file output type as a traffic scenario file according to the application scenario of the virtual model, and the server can export the virtual model as a traffic scenario file, which can be applied to city traffic management, smart travel and travel, and automatic driving simulation in multiple directions.
[0119] The above method, after obtaining the virtual model, in order to apply the virtual model, the server can store the virtual model in the form of a file, in different scenarios, the server can export the file form required by the file application scenario corresponding to the virtual model, so that the exported model file is more suitable for the corresponding application scenario.
[0120] After the virtual model of the entity object to be modeled is visualized, the interactive object can intuitively see the effect of the virtual model, so as to determine whether the virtual model meets the requirements of the interactive object. In order to make the virtual model more meet the requirements of the interactive object, the interactive object can further edit the virtual model of the entity object to be modeled according to whether the interactive object is satisfied with the effect of the virtual model, so as to adjust the virtual model as needed. The server can first obtain second operation instruction content for the target region in the virtual model, the second operation instruction content including second description information and second position information, the second description information being used to indicate the intention of the interactive object to perform a target editing operation, and the second position information being used to indicate the position of the target region in the virtual model to which the target editing operation is directed. Then, the server can edit the target region according to the position of the target region in the virtual model indicated by the second position information and the target editing operation indicated by the second description information, to obtain an edited virtual model.
[0121] The target region is a region in the virtual model waiting for a subsequent editing operation to be performed. In the scenario of city twinning, the target region can be a virtual model corresponding to a specific building, vegetation, road or terrain in the virtual model.
[0122] Similarly to the implementation manner of S202, in a possible implementation manner, the manner in which the server obtains the edited virtual model can be that the server can perform semantic information extraction on the second description information to obtain second operation semantic features, and perform feature encoding on the second position information to obtain second position features. Then, the target region in the virtual model is edited according to the image semantic features, the second operation semantic features and the second position features, to obtain an edited virtual model, the edited virtual model being generated under the control of the intention reflected by the second description information and the second position information.
[0123] The server can obtain the second description information and the second position information in various manners. The manner of obtaining the second description information is similar to that of obtaining the first description information, and the manner of obtaining the second position information is similar to that of obtaining the first position information, which will not be described herein again.
[0124] The interactive object may have various requirements for the virtual model. In order to make the virtual model meet the various requirements of the interactive object, the second operation instruction content may correspond to the various requirements of the interactive object. For example, in the scenario of urban twins, the interactive object requires that the building in the virtual model be a building, but the building displayed in the virtual model is a bungalow. At this time, the interactive object can indicate the replacement of the architectural style in the virtual model through the second description information in the second operation instruction content. In the same scenario, in addition to requirements for the style of the building, the interactive object also has requirements for the height of the building. The interactive object requires that the building in the virtual model cannot be higher than fifty meters, but the virtual model displays a building higher than fifty meters. At this time, the interactive object can indicate the editing of the building higher than fifty meters in the virtual model through the second description information in the second operation instruction content.
[0125] Figure 10 A schematic diagram of editing a virtual model provided in an embodiment of the present application, wherein the area in the dotted box of the left virtual model is the target area. The interactive object inputs a second operation instruction content including second descriptive information and second position information, wherein the second descriptive information is "style switch to office building", and the second position information is the latitude and longitude of the target area. The architectural style in the left virtual model is a bungalow, and the second operation instruction content indicates the intention of the interactive object to perform a secondary type modification on the modeled entity object. The server switches the building in the target area from the secondary type "bungalow" to the secondary type "office building" based on the image semantic features and the second operation instruction features.
[0126] The server edits the target area in the virtual model according to the image semantic feature, the second operation semantic feature and the second position feature by fusing the image semantic feature, the second operation semantic feature and the second position feature to obtain a second fused feature, where the second fused feature includes the intention of the interactive object to perform the target editing operation. Afterwards, the server can edit the target area in the virtual model based on the second fused feature to obtain the edited virtual model.
[0127] In the embodiment of the application, when the virtual model of the entity object to be modeled is further edited, the target editing operation performed on the target region is controlled in a man-machine interactive manner. Specifically, the second operation instruction content can be acquired, the second operation instruction content including second description information and second position information, the second description information being used to indicate the intention of the interactive object to perform the editing operation, and the second position information being used to indicate the position of the target region in the virtual model to which the editing operation is directed. Thus, the target editing operation can be accurately performed on the virtual model based on the second description information and the second position information, and the edited virtual model is obtained. In the embodiment of the application, the virtual model can be directly edited as needed, and the editing process of the virtual model does not need a long intermediate chain. The entity object to be modeled, the virtual model and the modified virtual model have a direct and intuitive association. The generation process of the virtual model of the entity object to be modeled and the editing process of the virtual model are very simple and fast. For the interactive object, only simple man-machine interaction is needed to achieve the editing of the virtual model, which is very friendly to the interactive object.
[0128] The above embodiment details the modeling method of the entity object. One main application scenario of modeling the entity object is city twinning. City twinning has a wide application prospect, including city traffic management, smart travel and travel, automatic driving simulation and industrial simulation and the like. However, the city twinning in the related art often relies on a plurality of different data, such as map vector data, building height data, road line and polygon region of vegetation and water and the like. The sources of these data are generally different, for example, the map vector data can be from a map, and the building height data can be from a database. The related art relies very much on a plurality of different data when modeling, for example, relies on the building height data when modeling the building, and relies on the map vector data when modeling the road. In order to reduce the problem of data dependence, the scenario of city twinning is taken as an example and described in detail in combination with the embodiment of the application.
[0129] In the scenario of city twinning, the entity object to be modeled is a geographical entity, and the image to be processed is a satellite map. The server can obtain the element model of each type of semantic element from the image to be processed under the control of the intention reflected by the first position information and the first description information. The element model can be a virtual model corresponding to the semantic element, and each type of semantic element can provide direct semantic support for city twinning. Then, the server can combine the element models of each type of semantic element to obtain the virtual model of the entity object to be modeled.
[0130] The type of the semantic element can be a first-level type or a second-level type. The semantic elements in the city twin scenario include first-level semantic elements such as building elements, vegetation elements, road elements, and terrain elements, and the vegetation elements can further include second-level semantic elements such as large trees, small trees, dry trees, shrubs, and the like.
[0131] Figure 11 A city twin schematic diagram provided by an embodiment of the present application, Figure 11 The left image is a satellite image, which can be used as a to-be-processed image. The satellite image includes semantic elements of types such as jungle, lake, bridge, and house. The server can obtain element models of semantic elements of various types according to the to-be-processed image under the control of the intention of the creation operation of the interactive object reflected by the first position information and the first description information. Then, the server combines the element models of semantic elements of various types to obtain a virtual model of the to-be-modeled entity object. Figure 11 The right side is a virtual model corresponding to the left satellite image.
[0132] In the city twin scenario, the embodiment of the present application only needs to input a satellite image and first operation instruction content to generate a virtual model, and does not depend on map vector data, building height data, road lines, and polygon areas of vegetation and water areas, thereby greatly reducing the time for obtaining data and reducing the cost. Since map vector data is not required during modeling, city twinning can be performed even in remote areas where map vector data is relatively scarce, thereby improving the application range of city twinning. In the city twin scenario, the embodiment of the present application does not need to process different data respectively, the process is very simple, and does not involve the parsing and merging of multiple data formats, thereby reducing the cost.
[0133] Figure 12 An architecture diagram for city twinning using a multi-modal large model provided by an embodiment of the present application, in which Figure 12 The first operation instruction content includes Prompt information and first description information, and the Prompt information is the first position information. The Prompt information, the first description information, and the satellite image are inputs of the multi-modal large model. The multi-modal large model generates a virtual model based on the Prompt information, the first description information, and the satellite image, and then visualizes the virtual model. Finally, the virtual model is exported according to the application scenario of city twinning.
[0134] Under the corresponding architecture, city twinning can be performed through the overall flowchart of city twinning provided by the embodiment of the present application, Figure 12 Figure 13 Figure 13 An overall flow diagram of city twinning provided by an embodiment of the present application is shown. In the city twinning scenario, the image to be processed is a satellite image, and the server extracts semantic information from the satellite image to obtain image semantic features. The description information represents the intention of the interactive object to perform an operation, and the Prompt information is location information. The server extracts semantic information from the description information to obtain operation semantic features, and encodes the features of the Prompt information to obtain location features. The overall flow of city twinning can include a creation phase and an editing phase of a virtual model. In the creation phase, the description information is first description information, and the Prompt information is first location information. The description information and the Prompt information constitute first operation instruction content. In the editing phase, the description information is second description information, and the Prompt information is second location information. The description information and the Prompt information constitute second operation instruction content. The server fuses the image semantic features, the operation semantic features, and the location features to obtain fused features. In the creation phase, the fused features are first fused features, and in the editing phase, the fused features are second fused features. The server models based on the fused features to obtain a virtual model, and then visualizes the virtual model. The interactive object can choose whether to modify the virtual model based on the visualization. If modified, the process of the editing phase is performed, and if not modified, the virtual model is exported.
[0135] In Figure 13 The overall flow provided uses a multi-modal large model, Figure 14 The multi-modal large model can be used in the overall flow, Figure 14 An illustration of a multi-modal large model in a city twinning scenario provided by an embodiment of the present application is shown. In the multi-modal large model shown, Figure 14 In the multi-modal large model shown, the satellite image, the first description information, and the Prompt information are inputs to the multi-modal large model, and the Prompt information is first location information. The image processing large model extracts semantic information from the satellite image to obtain image semantic features. The description information processing large model extracts semantic information from the first description information to obtain first operation semantic features. The Prompt information processing large model encodes the features of the Prompt information to obtain first location features. The Prompt information processing large model can be a location information processing large model. The multi-modal large model fuses the image semantic features, the first operation semantic features, and the first location features to obtain first fused features, and the cross-shaped circle in the figure represents feature fusion. The first fused features are inputs to a three-dimensional generation model, which generates a virtual model using the first fused features and visualizes the virtual model.
[0136] It should be noted that the implementation manners provided by the present application in the above aspects can be further combined to provide more implementation manners.
[0137] Based on Figure 2 The modeling method of the entity object provided by the corresponding embodiment also provides a modeling device 1500 of an entity object. Referring to Figure 15 As shown in the figure, the barrage display device 1500 includes an acquisition unit 1501 and a modeling unit 1502:
[0138] The acquisition unit 1501 is configured to acquire a to-be-processed image including a to-be-modeled entity object, and acquire first operation instruction content, the first operation instruction content including first description information and first position information, the first description information being used to indicate an intention of an interactive object to perform a creation operation, and the first position information being used to indicate a position of the to-be-modeled entity object in the to-be-processed image to which the creation operation is directed.
[0139] The modeling unit 1502 is configured to model the to-be-modeled entity object according to the position of the to-be-modeled entity object in the to-be-processed image indicated by the first position information and the creation operation indicated by the first description information, and obtain a virtual model of the to-be-modeled entity object.
[0140] In a possible implementation, the acquisition unit is specifically configured to:
[0141] acquire the first description information input by the interactive object, or acquire the first description information in response to a trigger of a creation operation control by the interactive object;
[0142] acquire the first position information input by the interactive object, or acquire the first position information in response to a selection operation of the to-be-modeled entity object by the interactive object.
[0143] In a possible implementation, the acquisition unit is specifically configured to:
[0144] acquire an original entity object image, and the to-be-modeled entity object is a partial entity object included in the original entity object image;
[0145] determine an image corresponding to the first position information in the original entity object image as the to-be-processed image.
[0146] In a possible implementation, the modeling unit is specifically configured to:
[0147] perform semantic information extraction on the to-be-processed image to obtain image semantic features, perform semantic information extraction on the first description information to obtain first operation semantic features, and perform feature encoding on the first position information to obtain first position features;
[0148] According to the image semantic feature, the first operation semantic feature, and the first position feature, the modeling unit models the entity object to be modeled to obtain a virtual model of the entity object to be modeled.
[0149] In a possible implementation, the modeling unit is specifically configured to:
[0150] The image processing model is used to extract semantic information of the image to be processed to obtain the image semantic feature.
[0151] The description information processing model is used to extract semantic information of the first description information to obtain the first operation semantic feature.
[0152] The position information processing model is used to encode the first position information to obtain the first position feature.
[0153] According to the image semantic feature, the first operation semantic feature, and the first position feature, the modeling unit models the entity object to be modeled by using a three-dimensional generation model to obtain a virtual model of the entity object to be modeled.
[0154] In a possible implementation, the entity object to be modeled includes semantic elements, and the apparatus further includes a training unit, which is specifically configured to:
[0155] Obtain a plurality of sample images and semantic labels of each sample image in the plurality of sample images, the semantic label being used to indicate a secondary type of a semantic element in the sample image, and semantic elements of a same primary type being divided into a plurality of secondary types.
[0156] The initial network model is trained by using the plurality of sample images and the semantic labels corresponding to each sample image to obtain the image processing model.
[0157] In a possible implementation, the modeling unit is specifically configured to:
[0158] The image semantic feature, the first operation semantic feature, and the first position feature are fused to obtain a first fused feature.
[0159] The modeling unit models the entity object to be modeled based on the first fused feature to obtain a virtual model of the entity object to be modeled.
[0160] In a possible implementation, after the modeling unit models the to-be-modeled entity object according to the position of the to-be-modeled entity object in the to-be-processed image indicated by the first position information and the creation operation indicated by the first description information, the apparatus further includes a display unit, in particular configured to:
[0161] display the virtual model of the to-be-modeled entity object.
[0162] In a possible implementation, the apparatus further includes an editing unit, in particular configured to:
[0163] obtain second operation instruction content for a target region in the virtual model, the second operation instruction content including second description information and second position information, the second description information being used to indicate an intention of the interactive object to perform a target editing operation, and the second position information being used to indicate a position of the target region in the virtual model to which the target editing operation is directed;
[0164] edit the target region according to the position of the target region in the virtual model indicated by the second position information and the target editing operation indicated by the second description information, to obtain an edited virtual model.
[0165] In a possible implementation, the apparatus further includes an output unit, in particular configured to:
[0166] determine a file output type according to an application scenario of the virtual model;
[0167] export a model file of the virtual model according to the file output type.
[0168] In a possible implementation, the to-be-modeled entity object is a geographic entity, the to-be-processed image is a satellite image, the image semantic feature is used to indicate a type of a semantic element included in the geographic entity, and the modeling unit is specifically configured to:
[0169] obtain an element model of each type of semantic element according to the to-be-processed image under control of the intention reflected by the first position information and the first description information;
[0170] combine the element models of the each type of semantic element to obtain the virtual model of the to-be-modeled entity object.
[0171] It can be seen from the technical solution that, in the embodiment, the generation of the virtual model is controlled by human-computer interaction when the entity object to be modeled is modeled. Specifically, the image to be processed including the entity object to be modeled and the first operation instruction content are acquired, and the first operation instruction content includes the first description information and the first position information, so that the task to be performed on the image to be processed can be accurately understood based on the first operation instruction. The first description information is used to indicate the intention of the interactive object to perform the creation operation, and the first position information is used to indicate the position of the entity object to be modeled in the image to be processed. Therefore, the first description information and the first position information can reflect the real operation intention of the interactive object. Thus, the entity object to be modeled can be modeled according to the position of the entity object to be modeled in the image to be processed indicated by the first position information and the creation operation indicated by the first description information based on the image to be processed, and the virtual model of the entity object to be modeled is obtained. In the modeling process, the virtual model can be generated based on the image to be processed under the control of the real operation intention reflected by the first operation instruction content, so as to avoid performing the wrong task.
[0172] The embodiment of the application further provides a computer device which can execute the modeling method of the entity object. Figure 16 The structure diagram of the terminal is shown. Figure 16 In the embodiment, the terminal is taken as a smart phone as an example.
[0173] Referring to Figure 16 , the smart phone includes the radio frequency (RF) circuit 1610, the memory 1620, the input unit 1630, the display unit 1640, the sensor 1650, the audio circuit 1660, the wireless fidelity (WiFi) module 1670, the processor 1680, and the power supply 1690, and the like. The input unit 1630 can include the touch panel 1631 and other input devices 1632, and the display unit 1640 can include the display panel 1641. The audio circuit 1660 can include the speaker 1661 and the microphone 1662. It can be understood that Figure 16 The structure of the smart phone shown in the embodiment is not limited to the smart phone, and can include more or fewer components than the diagram, or combine some components, or different component arrangement.
[0174] The memory 1620 can be used to store software programs and modules. The processor 1680 executes the various functional applications and data processing of the smartphone by running the software programs and modules stored in the memory 1620. The memory 1620 may mainly include a program storage area and a data storage area. The program storage area may store an operating system, at least one application required for a function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created based on the use of the smartphone (such as audio data, a phone book, etc.). In addition, the memory 1620 may include a high-speed random access memory and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage device.
[0175] Processor 1680 is the control center of the smartphone, connecting all components of the smartphone using various interfaces and circuits. It executes software programs and / or modules stored in memory 1620 and accesses data stored in memory 1620 to perform various smartphone functions and process data. Optionally, processor 1680 may include one or more processing units. Preferably, processor 1680 integrates an application processor and a modem processor. The application processor primarily handles the operating system, user interface, and application programs, while the modem processor primarily handles wireless communications. It is understood that the modem processor may not be integrated into processor 1680.
[0176] In this embodiment, the processor 1680 in the smartphone can execute the entity object modeling method provided in each embodiment of the present application.
[0177] The computer device provided in the embodiment of the present application may also be a server, see Figure 17 As shown, Figure 17 The structural diagram of the server 1700 provided in the embodiment of the present application, the server 1700 may have relatively large differences due to different configurations or performances, and may include one or more processors, such as a central processing unit (CPU) 1722, and a memory 1732, one or more storage media 1730 (such as one or more massive storage devices) storing application programs 1742 or data 1744. Among them, the memory 1732 and the storage medium 1730 can be temporary storage or persistent storage. The program stored in the storage medium 1730 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Furthermore, the central processing unit 1722 can be configured to communicate with the storage medium 1730 to execute a series of instruction operations in the storage medium 1730 on the server 1700.
[0178] The server 1700 can also include one or more power supplies 1726, one or more wired or wireless network interfaces 1750, one or more input / output interfaces 1758, and / or one or more operating systems 1741, such as Windows Server TM , Mac OS X TM , Unix TM , Linux TM , FreeBSD TM , etc.
[0179] In this embodiment, the central processing unit 1722 in the server 1700 can execute the modeling method of the solid object provided in the embodiments of the present application.
[0180] According to an aspect of the present application, a computer readable storage medium is provided, which is used to store a computer program, the computer program being used to execute the modeling method of the solid object provided in the embodiments.
[0181] According to an aspect of the present application, a computer program product is provided, which includes a computer program stored in a computer readable storage medium. A processor of a computer device reads the computer program from the computer readable storage medium, and the processor executes the computer program, so that the computer device executes the method provided in the various optional implementation manners of the above embodiments.
[0182] The descriptions of the corresponding flow or structure of each of the above figures are each focused on, and the parts not described in detail in a certain flow or structure can be referred to the related descriptions of other flows or structures.
[0183] The terms "first", "second", "third", "fourth" and the like in the description of the present application and in the above figures (if any) are used to distinguish similar objects, and do not necessarily have to be used to describe a particular order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0184] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other manners. For example, the described device embodiments are merely schematic. The division of the units is merely a logical function division. There can be another division manner for the actual implementation, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electrical, mechanical or other forms.
[0185] The units described as separated components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed on a plurality of network units. In actual implementation, some or all of the units can be selected according to the actual needs to achieve the purposes of the embodiments.
[0186] In addition, each functional unit in the embodiments of the present application can be integrated in a processing unit, or each unit can exist physically as a separate unit, or two or more units can be integrated in one unit. The integrated unit can be implemented in the form of hardware, or in the form of software functional units.
[0187] When the integrated unit is implemented in the form of software functional units and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on such an understanding, the technical solutions of the present application essentially, or the part that contributes to the prior art, or all or a part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions for causing a computer device (which can be a terminal, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the embodiments of the present application. The foregoing storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, and various other media that can store computer programs.
[0188] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program with a predetermined function, and works together with other related parts to achieve a predetermined target, and can be implemented entirely or partially by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, one processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of an integral module or unit that includes the functions of the module or unit.
[0189] The above-described embodiments are merely used to illustrate the technical solutions of the present application, but not limit the present application; although the present application has been described in detail with reference to the foregoing embodiments, it should be understood by those of ordinary skill in the art that the technical solutions recorded in the foregoing embodiments can be modified, or some technical features thereof can be replaced equivalently; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for modeling a physical object, characterized in that: The method comprises: Acquiring an image to be processed that includes a physical object to be modeled, and acquiring first operation instruction content, wherein the first operation instruction content includes first description information and first position information, wherein the first description information is used to indicate the intention of the interactive object to perform a creation operation, and the first position information is used to indicate the position of the physical object to be modeled targeted by the creation operation in the image to be processed; According to the image to be processed, the entity object to be modeled is modeled according to the position of the entity object to be modeled in the image to be processed indicated by the first position information and the creation operation indicated by the first description information to obtain a virtual model of the entity object to be modeled.
2. The method according to claim 1, characterized in that The obtaining of the first operation instruction content includes: Acquire the first description information input by the interactive object, or acquire the first description information in response to the interactive object triggering a create operation control; The first position information input by the interactive object is obtained, or the first position information is obtained in response to a selection operation of the interactive object on the entity object to be modeled.
3. The method according to claim 1, characterized in that The step of obtaining an image to be processed including an entity object to be modeled comprises: Acquire an original entity object image, wherein the entity object to be modeled is a partial entity object included in the original entity image; An image in the original physical object image corresponding to the first position information is determined as the image to be processed.
4. The method according to claim 1, wherein The step of modeling the entity object to be modeled according to the image to be processed, according to the position of the entity object to be modeled in the image to be processed indicated by the first position information and the creation operation indicated by the first description information, to obtain a virtual model of the entity object to be modeled, including: Extracting semantic information from the image to be processed to obtain image semantic features, extracting semantic information from the first description information to obtain first operational semantic features, and encoding the first position information to obtain first position features; The entity object to be modeled is modeled according to the image semantic feature, the first operational semantic feature, and the first position feature to obtain a virtual model of the entity object to be modeled.
5. The method according to claim 4, characterized in that The step of extracting semantic information from the image to be processed to obtain image semantic features includes: Extracting semantic information from the image to be processed using an image processing model to obtain semantic features of the image; The extracting semantic information from the first description information to obtain a first operational semantic feature includes: extracting semantic information from the first description information using a description information processing model to obtain the first operational semantic feature; The performing feature encoding on the first position information to obtain a first position feature includes: Performing feature encoding on the first position information using a position information processing model to obtain the first position feature; The step of modeling the entity object to be modeled according to the image semantic feature, the first operational semantic feature, and the first position feature to obtain a virtual model of the entity object to be modeled includes: The entity object to be modeled is modeled by using a three-dimensional generative model according to the image semantic feature, the first operational semantic feature, and the first position feature to obtain a virtual model of the entity object to be modeled.
6. The method according to claim 5, characterized in that The entity object to be modeled includes semantic elements, and the method further includes: Acquire a plurality of sample images and a semantic label of each of the plurality of sample images, wherein the semantic label is used to indicate a secondary type of a semantic element in the sample image, and semantic elements of the same primary type are divided into multiple secondary types; The initial network model is trained using the multiple sample images and the semantic labels corresponding to each sample image to obtain the image processing model.
7. The method according to claim 4, characterized in that The step of modeling the entity object to be modeled according to the image semantic feature, the first operational semantic feature, and the first position feature to obtain a virtual model of the entity object to be modeled includes: Performing feature fusion on the image semantic feature, the first operational semantic feature, and the first position feature to obtain a first fused feature; The entity object to be modeled is modeled based on the first fused features to obtain a virtual model of the entity object to be modeled.
8. The method according to any one of claims 1 to 7, characterized in that After modeling the entity object to be modeled based on the image to be processed, according to the position of the entity object to be modeled in the image to be processed indicated by the first position information and the creation operation indicated by the first description information, to obtain a virtual model of the entity object to be modeled, the method further includes: A virtual model of the entity object to be modeled is displayed.
9. The method according to claim 8, characterized in that The method further comprises: Obtaining a second operation instruction content for a target area in the virtual model, the second operation instruction content including second description information and second position information, the second description information being used to indicate the intention of the interactive object to perform a target editing operation, and the second position information being used to indicate a position of the target area targeted by the target editing operation in the virtual model; According to the position of the target area in the virtual model indicated by the second position information and the target editing operation indicated by the second description information, the target area is edited to obtain an edited virtual model.
10. The method according to any one of claims 1 to 7, characterized in that The method further comprises: Determining a file output type according to an application scenario of the virtual model; Export the model file of the virtual model according to the file output type.
11. The method according to any one of claims 1 to 7, characterized in that: The entity object to be modeled is a geographic entity, the image to be processed is a satellite image, the image semantic features are used to indicate the type of semantic elements included in the geographic entity, and modeling the entity object to be modeled according to the image to be processed, according to the position of the entity object to be modeled in the image to be processed indicated by the first position information and the creation operation indicated by the first description information, to obtain a virtual model of the entity object to be modeled, including: Under the control of the intention reflected by the first position information and the first description information, obtaining element models of various types of semantic elements according to the image to be processed; The element models of the semantic elements of various types are combined to obtain a virtual model of the entity object to be modeled.
12. A modeling device for a physical object, characterized in that: The device comprises an acquisition unit and a modeling unit: The acquisition unit is configured to acquire an image to be processed including a physical object to be modeled, and acquire first operation instruction content, wherein the first operation instruction content includes first description information and first position information, wherein the first description information is used to indicate the intention of the interactive object to perform a creation operation, and the first position information is used to indicate the position of the physical object to be modeled targeted by the creation operation in the image to be processed; The modeling unit is used to model the entity object to be modeled according to the image to be processed, according to the position of the entity object to be modeled in the image to be processed indicated by the first position information and the creation operation indicated by the first description information, so as to obtain a virtual model of the entity object to be modeled.
13. A computer device, characterized in that: The computer device includes a processor and a memory: The memory is used to store a computer program and transmit the computer program to the processor; The processor is configured to execute the method according to any one of claims 1 to 11 according to instructions in the computer program.
14. A computer-readable storage medium, characterized in that The computer-readable storage medium is used to store a computer program, which, when executed by a processor, causes the processor to perform the method according to any one of claims 1 to 11.
15. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 11 is implemented.