Human-computer interaction methods, devices, media and electronic equipment

By using a multimodal large model to infer and segment the digital 3D model, the problem of users having difficulty understanding the digital display of cultural heritage is solved, and a more intuitive interactive learning experience and enhanced immersion are achieved.

CN119225530BActive Publication Date: 2025-10-28PEKING UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411208770.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-30
Publication Date
2025-10-28
Estimated Expiration
2044-08-30

AI Technical Summary

Technical Problem

When faced with digital displays of cultural heritage, users often struggle to effectively connect abstract textual explanations with specific cultural details, impacting user experience and learning outcomes.

Method used

The digital 3D model is inferred through a multimodal large model to obtain the segmented areas related to the introductory text, and the introductory text is aligned and segmented with 2D images from multiple perspectives. In addition, a three-dimensional mask is generated in combination with text instruction prompts to display the segmented areas and play audio.

Benefits of technology

It enhances the user's sense of immersion and learning depth, significantly improves the user experience and educational outcomes, and achieves an effective connection between abstract text and specific cultural details.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119225530B_ABST
    Figure CN119225530B_ABST
Patent Text Reader

Abstract

This application discloses a human-computer interaction method, device, medium, and electronic device. The method includes: responding to a user's selection instruction for a target object, displaying a digital 3D model of the target object; when the display time of the digital 3D model reaches a preset time, obtaining introductory text for the target object; based on the introductory text, performing reasoning on the digital 3D model using a multimodal large model to obtain segmented regions related to the introductory text, wherein the multimodal large model is used to align and segment the introductory text with 2D images from multiple perspectives, the 2D images from multiple perspectives being rendered from different predefined viewpoints of the digital 3D model; playing audio corresponding to the introductory text, and displaying the segmented regions. Therefore, by using the embodiments of this application, when users encounter digital displays of cultural heritage, they can effectively connect abstract textual explanations with specific cultural details, thereby improving the user experience and learning effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of mixed reality and human-computer interaction technology, and in particular to a human-computer interaction method, device, medium and electronic device. Background Technology

[0002] With the rapid development of information technology, human-computer interaction technology has made significant progress, especially in the fields of 3D display and virtual reality. These technologies have provided entirely new ways for users to interact with digital content, greatly enriching application scenarios in multiple fields such as education, entertainment, healthcare, and tourism. In particular, the digital preservation and dissemination of cultural heritage has become a field of great interest.

[0003] Currently, digital technology allows people to appreciate and engage with cultural heritage in an immersive way. This not only provides new avenues for the protection and transmission of cultural heritage but also offers the public a richer and more intuitive experience. However, because cultural heritage typically contains profound historical and cultural knowledge, its complexity often exceeds the comprehension of ordinary users without a professional background. This makes it difficult for users to effectively connect abstract textual explanations with specific cultural details when faced with digital displays of cultural heritage, thus affecting their experience and learning outcomes. Summary of the Invention

[0004] This application provides a human-computer interaction method, apparatus, medium, and electronic device. To provide a basic understanding of some aspects of the disclosed embodiments, a brief summary is given below. This summary is not intended as a general description, nor is it intended to identify key / important components or describe the scope of protection of these embodiments. Its sole purpose is to present some concepts in a simple form as a prelude to the detailed description that follows.

[0005] In a first aspect, embodiments of this application provide a human-computer interaction method, the method comprising:

[0006] In response to the user's selection command for the target object, display a digital 3D model of the target object;

[0007] Once the display time of the digital 3D model reaches the preset time, obtain the descriptive text of the target object;

[0008] Based on the introductory text, the digital 3D model is inferred through a multimodal large model to obtain the segmentation region related to the introductory text. The multimodal large model is used to align and segment the introductory text with 2D images from multiple perspectives. The 2D images from multiple perspectives are obtained by rendering different predefined viewpoints of the digital 3D model.

[0009] Play the audio corresponding to the introductory text and display the segmented area.

[0010] Optionally, the method also includes:

[0011] The device receives attitude information sent by the sensing device. The attitude information is obtained by the sensing device when the user operates the pre-marked 3D prototype of the target object and is calculated based on the spatial information of the sensing device. The 3D prototype is a physical object manufactured based on the digital 3D model.

[0012] Based on the pose information, update the display orientation of the digital 3D model to keep the pose of the digital 3D model consistent with the pose of the 3D prototype.

[0013] Optionally, based on the introductory text, inference is performed on the digital 3D model using a multimodal large model to obtain segmented regions related to the introductory text, including:

[0014] Render different predefined viewpoints of the digital 3D model to generate 2D images from multiple perspectives;

[0015] Retrieve text instruction prompts present in the introductory text;

[0016] Input 2D images from multiple perspectives and text instructions into a multimodal large model, and output multiple target 2D regions described by the text instructions;

[0017] Merge multiple target 2D regions into a 3D mask;

[0018] Preprocess the 3D mask to obtain segmented regions related to the introductory text.

[0019] Optionally, the multimodal large model includes an alignment module and a segmentation module;

[0020] Inputting 2D images from multiple perspectives and text instructions into a multimodal large model, the output model describes multiple target 2D regions, including:

[0021] The alignment module receives 2D images and text prompts from multiple perspectives, extracts visual features from the 2D images, extracts keywords and phrases from the text prompts, uses deep learning technology to align the visual features and text features, and outputs multiple aligned feature maps.

[0022] The segmentation module receives multiple feature maps and 2D images from multiple viewpoints from the alignment module. Based on the multiple feature maps, it performs region segmentation on the 2D image from each viewpoint and outputs multiple target 2D regions described by text instructions.

[0023] Optionally, the 3D mask is preprocessed to obtain segmented regions related to the introductory text, including:

[0024] Refine the 3D mask;

[0025] Eliminate discontinuities caused by viewpoint changes in refined 3D masks;

[0026] Filter out low-confidence areas in the 3D mask to eliminate discontinuities caused by viewpoint changes;

[0027] The 3D mask, which filters out low-confidence areas, is projected onto the digital 3D model to obtain the segmented regions related to the introductory text.

[0028] Optionally, the refinement uses a Gaussian geodesic reweighting algorithm, the elimination of discontinuities caused by viewpoint changes uses visibility smoothing technology, the filtering of low-confidence areas uses a global filtering strategy, and the audio is obtained by converting the introductory text using the Text2Speech algorithm.

[0029] Optionally, a digital 3D model of the target object may be generated by following these steps:

[0030] Collect object information of the target object, including geometric data, image information, and on-site measurement information;

[0031] Based on geometric data, image information, and on-site measurement information, point cloud data of the target object is constructed.

[0032] The point cloud data is converted into a network model to obtain the initial 3D structure;

[0033] The initial 3D structure is modeled in three dimensions and textured to obtain a digital 3D model of the target object.

[0034] Secondly, embodiments of this application provide a human-computer interaction device, the device comprising:

[0035] The display module is used to display a digital 3D model of the target object in response to a selection command for the target object;

[0036] The acquisition module is used to acquire the descriptive text of the target object when the display time of the digital 3D model reaches the preset time.

[0037] The reasoning module is used to reason about the digital 3D model based on the introductory text using a multimodal large model to obtain the segmented regions related to the introductory text. The multimodal large model is used to align and segment the introductory text with 2D images from multiple perspectives, which are obtained by rendering different predefined viewpoints of the digital 3D model.

[0038] The playback and display module is used to play the audio corresponding to the introductory text and display the segmented area.

[0039] Thirdly, embodiments of this application provide a computer storage medium storing multiple instructions adapted for loading and execution of the above-described method steps by a processor.

[0040] Fourthly, embodiments of this application provide an electronic device that may include: a processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and executed by the above-described method steps.

[0041] The technical solutions provided in this application embodiment may include the following beneficial effects:

[0042] In this embodiment, based on the introductory text, a segmented region related to the digital 3D model is obtained by inferring from the digital 3D model using a multimodal large model. This multimodal large model is used to align and segment the introductory text with 2D images from multiple perspectives, which are obtained by rendering different predefined viewpoints of the digital 3D model. This process involves accurately aligning and segmenting the introductory text with a set of 2D images rendered from the digital 3D model at multiple predefined viewpoints. This effectively links abstract textual explanations with specific cultural details, greatly enhancing user immersion and learning depth, thereby significantly improving user experience and educational effectiveness.

[0043] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and do not limit this application. Attached Figure Description

[0044] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0045] Figure 1 This is a flowchart illustrating a human-computer interaction method provided in an embodiment of this application;

[0046] Figure 2 This is a schematic diagram of a three-dimensional mask blending process provided in this application;

[0047] Figure 3 This application provides a schematic diagram of a segmented region related to the introductory text;

[0048] Figure 4 This is a schematic diagram of a three-dimensional prototype generation process provided in an embodiment of this application;

[0049] Figure 5 This is an interactive schematic diagram provided by an embodiment of the present application, illustrating how a user interacts with a three-dimensional prototype of a target object.

[0050] Figure 6This is a schematic diagram of the overall architecture of human-computer interaction provided in an embodiment of this application;

[0051] Figure 7 This is a schematic diagram of the structure of a human-computer interaction device provided in this application;

[0052] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0053] The following description and accompanying drawings fully illustrate specific embodiments of this application to enable those skilled in the art to practice them.

[0054] It should be understood that the described embodiments are merely some, not all, of the embodiments in this application. All other embodiments obtained by those skilled in the art based on the embodiments in this application without inventive effort are within the scope of protection of this application.

[0055] In the following description, when referring to the accompanying drawings, the same numbers in different drawings denote the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0056] In the description of this application, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Those skilled in the art can understand the specific meaning of the above terms in this application based on the specific circumstances. Furthermore, in the description of this application, unless otherwise stated, "multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship.

[0057] This application provides a human-computer interaction method, apparatus, medium, and electronic device to solve the problems existing in the aforementioned related technologies. In the embodiments of this application, based on the introductory text, a segmented region related to the introductory text can be obtained by reasoning about a digital 3D model using a multimodal large model. This multimodal large model is used to align and segment the introductory text with 2D images from multiple perspectives, which are obtained by rendering different predefined viewpoints of the digital 3D model. This process involves accurately aligning and segmenting the introductory text with a set of 2D images rendered from the digital 3D model at multiple predefined viewpoints. This method effectively links abstract textual explanations with specific cultural details, greatly enhancing the user's immersion and learning depth, thereby significantly improving the user experience and educational effectiveness. Exemplary embodiments are described in detail below.

[0058] The following will be combined with the appendix Figure 1 -Appendix Figure 6 This application provides a detailed description of the human-computer interaction method provided in its embodiments. This method can be implemented using a computer program and can run on human-computer interaction devices based on the von Neumann architecture. The computer program can be integrated into an application or run as a standalone utility application.

[0059] See Figure 1 This is a flowchart illustrating a human-computer interaction method provided in an embodiment of this application. Figure 1 As shown, the method in this application embodiment may include the following steps:

[0060] S101, in response to the user's selection instruction for the target object, displays a digital 3D model of the target object;

[0061] In this context, "user" refers to an individual interacting with the system or device, and could be a learner of cultural heritage. The target object is the specific cultural heritage that the user selects or focuses on. Selection instructions are generated by the user's selection or interaction with the target object; this can be through clicks, touch, voice commands, or other forms of user input. A digital 3D model is a virtual three-dimensional object created using digital technology.

[0062] In some embodiments, users select cultural heritage items of interest, such as a piece of porcelain, by tapping a touchscreen, using a mouse, or issuing selection commands via a voice recognition system through a graphical user interface (GUI). The system receives the user's selection command and translates it into a query for the corresponding cultural heritage item in the database. Based on the user's selection, the system retrieves a digital 3D model associated with the selected cultural heritage item. Using 3D rendering technology, the retrieved digital 3D model is rendered and displayed on the screen at high resolution and with realism.

[0063] The database contains a pre-built mapping between identifiers and digital 3D models.

[0064] In this embodiment of the application, a digital 3D model of the target object is generated according to the following steps: collecting object information of the target object, including geometric data, image information and on-site measurement information; constructing point cloud data of the target object based on the geometric data, image information and on-site measurement information; converting the point cloud data into a network model to obtain an initial 3D structure; and performing 3D modeling and texture mapping on the initial 3D structure to obtain a digital 3D model of the target object.

[0065] Furthermore, after obtaining the digital 3D model of each target object, an identifier for each target object can be created, and then the identifier of each target object and the digital 3D model of each target object can be stored to obtain the pre-built mapping relationship between the identifier and the digital 3D model.

[0066] S102, when the display time of the digital 3D model reaches the preset time, obtain the introductory text of the target object;

[0067] The display duration refers to the length of time the digital 3D model is displayed on the user interface. The preset duration is a pre-set time limit used to determine whether to retrieve the introductory text of the target object. The introductory text is textual information provided about the background, historical significance, and artistic value of the target object.

[0068] In some embodiments, when a user selects a cultural heritage object, the system displays the corresponding digital 3D model and starts a timer. The system continuously monitors the duration of the currently displayed 3D model. Once the preset duration is reached, the timer triggers an event. After the event is triggered, the system automatically retrieves the descriptive text associated with the displayed cultural heritage object from the database.

[0069] In this embodiment of the application, the timer continues even if the display time of the digital 3D model has not reached the preset time.

[0070] S103, Based on the introductory text, the digital 3D model is inferred through a multimodal large model to obtain the segmentation region related to the introductory text. The multimodal large model is used to align and segment the introductory text with 2D images from multiple perspectives. The 2D images from multiple perspectives are obtained by rendering different predefined viewpoints of the digital 3D model.

[0071] In 3D modeling, a segmented region is a specific part divided from the overall image or model, and these parts are associated with specific information or features. A 2D image is a two-dimensional image, that is, an image on a plane, which only contains width and height information compared to a 3D model. A viewpoint is the specific position and direction when observing or capturing an image; different viewpoints can provide different visual information. A predefined viewpoint is a specific position and direction set in advance during the 3D modeling and rendering process for observing or rendering a 3D model.

[0072] In some embodiments of this application, the specific process of obtaining segmented regions related to the introductory text by reasoning about a digital 3D model using a multimodal large model based on the introductory text includes: rendering different predefined viewpoints of the digital 3D model to generate 2D images from multiple perspectives; obtaining textual instruction prompts present in the introductory text; inputting the 2D images from multiple perspectives and the textual instruction prompts into the multimodal large model to output multiple target 2D regions described by the textual instruction prompts; fusing the multiple target 2D regions into a 3D mask; and preprocessing the 3D mask to obtain segmented regions related to the introductory text.

[0073] Text prompts are instructions or hints contained in the introductory text, used to guide users in learning relevant knowledge about the target object. A 3D mask merges multiple 2D segmented regions to form a 3D mask that can be applied to a 3D model, used to identify and distinguish different parts of the model.

[0074] In this embodiment, by rendering 2D images from multiple predefined viewpoints using a digital 3D model, and combining this with instruction prompts in the introductory text, a multimodal large model is used for deep analysis and reasoning. This allows for the accurate identification and segmentation of 2D target regions closely related to the text description. These regions are then further fused to form a 3D mask, which is then preprocessed to ultimately obtain segmented regions that precisely correspond to the content of the introductory text. This approach not only enhances the user's understanding and perception of the 3D model but also provides a more intuitive and interactive learning experience by combining abstract text with concrete visual elements. It is particularly suitable for the digital display and education of cultural heritage, significantly improving the efficiency of information delivery.

[0075] The multimodal large model includes an alignment module and a segmentation module.

[0076] In some embodiments of this application, the specific process of inputting 2D images from multiple perspectives and text prompts into a multimodal large model and outputting multiple target 2D regions described by the text prompts includes: an alignment module receiving 2D images from multiple perspectives and text prompts, extracting visual features from the 2D images, extracting keywords and phrases from the text prompts, aligning the visual features and text features using deep learning technology, and outputting multiple aligned feature maps; and a segmentation module receiving the multiple feature maps output by the alignment module and 2D images from multiple perspectives, performing region segmentation on the 2D images from each perspective based on the multiple feature maps, and outputting multiple target 2D regions described by the text prompts.

[0077] In some embodiments of this application, the specific process of preprocessing the 3D mask to obtain the segmented region related to the introductory text includes: refining the 3D mask; eliminating discontinuities caused by viewpoint changes in the refined 3D mask; filtering out low-confidence regions from the 3D mask with eliminated discontinuities caused by viewpoint changes; and projecting the 3D mask with the low-confidence regions filtered out onto a digital 3D model to obtain the segmented region related to the introductory text.

[0078] Specifically, the Gaussian geodesic reweighting algorithm is used to refine the process, visibility smoothing technology is used to eliminate discontinuities caused by changes in viewpoint, a global filtering strategy is used to filter out low-confidence areas, and the audio is obtained by converting the introductory text using the Text2Speech algorithm.

[0079] For example Figure 2 As shown, the target object is a humanoid cultural heritage. First, the digital 3D model of this humanoid cultural heritage is rendered from different predefined viewpoints to generate multi-view images. Then, the 2D images from multiple viewpoints, along with text prompts, are input into a multimodal large model, outputting multiple target 2D regions described by the text prompts. These multiple target 2D regions are then blended into a 3D mask using a fusion module. The 3D mask is preprocessed and projected onto the digital 3D model of the humanoid cultural heritage to obtain, for example... Figure 3 The segmented region related to the introductory text.

[0080] S104, play the audio corresponding to the introductory text and display the segmented area.

[0081] In some embodiments of this application, after obtaining the segmented region related to the introductory text, the introductory text can be converted into audio using the Text2Speech algorithm, and the audio can be played, as well as the segmented region can be displayed.

[0082] In this embodiment, based on the introductory text, a segmented region related to the digital 3D model is obtained by inferring from the digital 3D model using a multimodal large model. This multimodal large model is used to align and segment the introductory text with 2D images from multiple perspectives, which are obtained by rendering different predefined viewpoints of the digital 3D model. This process involves accurately aligning and segmenting the introductory text with a set of 2D images rendered from the digital 3D model at multiple predefined viewpoints. This effectively links abstract textual explanations with specific cultural details, greatly enhancing user immersion and learning depth, thereby significantly improving user experience and educational effectiveness.

[0083] In this embodiment, after displaying the digital 3D model of the target object, the user can manipulate the pre-marked 3D prototype of the target object to control the display direction of the digital 3D model; or after playing the audio corresponding to the introductory text and displaying the segmented area, the user can manipulate the pre-marked 3D prototype of the target object to control the display direction of the digital 3D model. The user's manipulation of the pre-marked 3D prototype of the target object can be determined by the user, and the system can control the display direction of the digital 3D model at different times based on the user's operation.

[0084] Specifically, the system receives posture information sent by the sensing device. The posture information is obtained by the sensing device when the user operates the pre-marked 3D prototype of the target object and is calculated based on the spatial information of the sensing device. The 3D prototype is a physical object manufactured based on the digital 3D model. According to the posture information, the display orientation of the digital 3D model is updated to keep the posture of the digital 3D model consistent with the posture of the 3D prototype.

[0085] Sensing devices refer to devices capable of detecting and identifying target objects in the environment, such as cameras and motion trackers. Posture information is data describing the orientation and position of an object in space, including rotation angle, tilt, and translation. Physical objects refer to real-world entities, as opposed to digital models. Display orientation is the direction in which the digital 3D model faces on the screen or in a virtual reality environment.

[0086] The 3D prototype is created using a 3D printer and a digital 3D model within a 3D mesh. For example... Figure 4 As shown, a three-dimensional basic prototype is obtained by printing a digital 3D model in a three-dimensional mesh using 3D printing technology, and then the three-dimensional basic prototype is colored to obtain a three-dimensional prototype.

[0087] After obtaining the 3D prototype, it is attached to an acrylic plate with ARUco markings at its four corners. Each marking has a unique structure and serves as a key input to the pose estimation algorithm. The camera captures the spatial information of these markings, and then the pose estimation algorithm processes this information to determine the orientation of the acrylic plate, and thus the orientation of the prototype, obtaining the pose information of the 3D prototype. The camera finally sends the pose information to the system.

[0088] For example Figure 5 As shown, the user manipulates a 3D prototype of the target object. The camera captures the pose information of the 3D prototype and sends it to the system in the holographic cabin. Based on the pose information, the system updates the display orientation of the digital 3D model to ensure that the pose of the digital 3D model matches that of the 3D prototype. Finally, the user can observe the updated digital 3D model displayed in the holographic cabin.

[0089] In this embodiment, by receiving posture information sent by a sensing device, the system can accurately capture the spatial dynamics of a user interacting with a pre-labeled 3D prototype. This information is derived from real-time monitoring and calculation of user behavior by the sensing device, ensuring synchronization between the digital 3D model and the actual object. Updating the display orientation of the digital 3D model to match the posture of the 3D prototype not only enhances the intuitiveness and accuracy of user interaction but also provides an immersive experience, allowing users to interact with virtual content more naturally. This method has significant advantages in multiple fields such as education, exhibitions, and design verification. Particularly in the digital display of cultural heritage, it allows users to explore and learn in a completely new way, greatly improving user experience and educational effectiveness.

[0090] For example Figure 6 As shown, Figure 6 This application provides a schematic diagram of the overall architecture of human-computer interaction. By reconstructing the cultural heritage in three dimensions, a digital 3D model of the cultural heritage can be obtained. This digital 3D model can be input into a holographic cabin for display. At the same time, the text description of the cultural heritage and the digital 3D model are input into a multimodal large model for processing, which can obtain segmented regions related to the description text. These segmented regions related to the description text are sent to the holographic cabin for display. Simultaneously, the text description is converted into audio for playback. Users can control the display angle of the digital 3D model of the cultural heritage by operating the 3D printed prototype.

[0091] In this embodiment, based on the introductory text, a segmented region related to the digital 3D model is obtained by inferring from the digital 3D model using a multimodal large model. This multimodal large model is used to align and segment the introductory text with 2D images from multiple perspectives, which are obtained by rendering different predefined viewpoints of the digital 3D model. This process involves accurately aligning and segmenting the introductory text with a set of 2D images rendered from the digital 3D model at multiple predefined viewpoints. This effectively links abstract textual explanations with specific cultural details, greatly enhancing user immersion and learning depth, thereby significantly improving user experience and educational effectiveness.

[0092] The following are device embodiments of the present application, which can be used to implement the method embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application.

[0093] See Figure 7 This illustration shows a structural schematic diagram of a human-computer interaction device provided in an exemplary embodiment of this application. This human-computer interaction device can be implemented as all or part of an electronic device through software, hardware, or a combination of both. The device 1 includes a display module 10, an acquisition module 20, an inference module 30, and a playback and display module 40.

[0094] Display module 10 is used to display a digital 3D model of the target object in response to a selection command for the target object;

[0095] The acquisition module 20 is used to acquire the description text of the target object when the display time of the digital 3D model reaches the preset time.

[0096] The reasoning module 30 is used to reason about the digital 3D model based on the introductory text using a multimodal large model to obtain the segmented regions related to the introductory text. The multimodal large model is used to align and segment the introductory text with 2D images from multiple perspectives. The 2D images from multiple perspectives are obtained by rendering different predefined viewpoints of the digital 3D model.

[0097] The playback and display module 40 is used to play the audio corresponding to the introductory text and display the segmented area.

[0098] It should be noted that the human-computer interaction device provided in the above embodiments is only illustrated by the division of the above functional modules when executing the human-computer interaction method. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the human-computer interaction device and the human-computer interaction method embodiments provided in the above embodiments belong to the same concept, and the implementation process is detailed in the method embodiments, which will not be repeated here.

[0099] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0100] In this embodiment, based on the introductory text, a segmented region related to the digital 3D model is obtained by inferring from the digital 3D model using a multimodal large model. This multimodal large model is used to align and segment the introductory text with 2D images from multiple perspectives, which are obtained by rendering different predefined viewpoints of the digital 3D model. This process involves accurately aligning and segmenting the introductory text with a set of 2D images rendered from the digital 3D model at multiple predefined viewpoints. This effectively links abstract textual explanations with specific cultural details, greatly enhancing user immersion and learning depth, thereby significantly improving user experience and educational effectiveness.

[0101] This application also provides a computer-readable medium having program instructions stored thereon, which, when executed by a processor, implement the human-computer interaction methods provided in the above-described method embodiments.

[0102] This application also provides a computer program product containing instructions that, when run on a computer, causes the computer to execute the human-computer interaction methods of the various method embodiments described above.

[0103] See Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 8 As shown, the electronic device 1000 may include: at least one processor 1001, at least one network interface 1004, a user interface 1003, a memory 1005, and at least one communication bus 1002.

[0104] The communication bus 1002 is used to realize the connection and communication between these components.

[0105] The user interface 1003 may include a display screen and a camera. Optionally, the user interface 1003 may also include a standard wired interface and a wireless interface.

[0106] The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wi-Fi interface).

[0107] The processor 1001 may include one or more processing cores. The processor 1001 connects to various parts within the electronic device 1000 using various interfaces and lines. It executes various functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 1005, and by calling data stored in the memory 1005. Optionally, the processor 1001 may be implemented using at least one hardware form selected from Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), and Programmable Logic Array (PLA). The processor 1001 may integrate one or more of the following: a Central Processing Unit (CPU), a Graphics Processing Unit (GPU), and a modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content to be displayed on the screen; and the modem handles wireless communication. It is understood that the modem may also be implemented as a separate chip, without being integrated into the processor 1001.

[0108] The memory 1005 may include random access memory (RAM) or read-only memory. Optionally, the memory 1005 may include a non-transitory computer-readable storage medium. The memory 1005 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 1005 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-described method embodiments, etc.; the data storage area may store data involved in the above-described method embodiments, etc. Optionally, the memory 1005 may also be at least one storage system located remotely from the aforementioned processor 1001. Figure 8 As shown, the memory 1005, which serves as a computer storage medium, may include an operating system, a network communication module, a user interface module, and human-computer interaction applications.

[0109] exist Figure 8In the illustrated electronic device 1000, the user interface 1003 is mainly used to provide an input interface for the user and to acquire user input data; while the processor 1001 can be used to call the human-computer interaction application stored in the memory 1005 and specifically perform the following operations:

[0110] In response to the user's selection command for the target object, display a digital 3D model of the target object;

[0111] Once the display time of the digital 3D model reaches the preset time, obtain the descriptive text of the target object;

[0112] Based on the introductory text, the digital 3D model is inferred through a multimodal large model to obtain the segmentation region related to the introductory text. The multimodal large model is used to align and segment the introductory text with 2D images from multiple perspectives. The 2D images from multiple perspectives are obtained by rendering different predefined viewpoints of the digital 3D model.

[0113] Play the audio corresponding to the introductory text and display the segmented area.

[0114] In one embodiment, the processor 1001 also performs the following operations:

[0115] The device receives attitude information sent by the sensing device. The attitude information is obtained by the sensing device when the user operates the pre-marked 3D prototype of the target object and is calculated based on the spatial information of the sensing device. The 3D prototype is a physical object manufactured based on the digital 3D model.

[0116] Based on the pose information, update the display orientation of the digital 3D model to keep the pose of the digital 3D model consistent with the pose of the 3D prototype.

[0117] In one embodiment, when the processor 1001 performs reasoning on the digital 3D model based on the introductory text using a multimodal large model to obtain the segmented region related to the introductory text, it specifically performs the following operations:

[0118] Render different predefined viewpoints of the digital 3D model to generate 2D images from multiple perspectives;

[0119] Retrieve text instruction prompts present in the introductory text;

[0120] Input 2D images from multiple perspectives and text instructions into a multimodal large model, and output multiple target 2D regions described by the text instructions;

[0121] Merge multiple target 2D regions into a 3D mask;

[0122] Preprocess the 3D mask to obtain segmented regions related to the introductory text.

[0123] In one embodiment, when the processor 1001 inputs 2D images from multiple perspectives and text instructions into a multimodal large model and outputs multiple target 2D regions described by the text instructions, it specifically performs the following operations:

[0124] The alignment module receives 2D images and text prompts from multiple perspectives, extracts visual features from the 2D images, extracts keywords and phrases from the text prompts, uses deep learning technology to align the visual features and text features, and outputs multiple aligned feature maps.

[0125] The segmentation module receives multiple feature maps and 2D images from multiple viewpoints from the alignment module. Based on the multiple feature maps, it performs region segmentation on the 2D image from each viewpoint and outputs multiple target 2D regions described by text instructions.

[0126] In one embodiment, when processor 1001 performs preprocessing of the 3D mask to obtain segmented regions related to the introductory text, it specifically performs the following operations:

[0127] Refine the 3D mask;

[0128] Eliminate discontinuities caused by viewpoint changes in refined 3D masks;

[0129] Filter out low-confidence areas in the 3D mask to eliminate discontinuities caused by viewpoint changes;

[0130] The 3D mask, which filters out low-confidence areas, is projected onto the digital 3D model to obtain the segmented regions related to the introductory text.

[0131] In one embodiment, when the processor 1001 executes the process of generating a digital 3D model of the target object, it specifically performs the following operations:

[0132] Collect object information of the target object, including geometric data, image information, and on-site measurement information;

[0133] Based on geometric data, image information, and on-site measurement information, point cloud data of the target object is constructed.

[0134] The point cloud data is converted into a network model to obtain the initial 3D structure;

[0135] The initial 3D structure is modeled in three dimensions and textured to obtain a digital 3D model of the target object.

[0136] In this embodiment, based on the introductory text, a segmented region related to the digital 3D model is obtained by inferring from the digital 3D model using a multimodal large model. This multimodal large model is used to align and segment the introductory text with 2D images from multiple perspectives, which are obtained by rendering different predefined viewpoints of the digital 3D model. This process involves accurately aligning and segmenting the introductory text with a set of 2D images rendered from the digital 3D model at multiple predefined viewpoints. This effectively links abstract textual explanations with specific cultural details, greatly enhancing user immersion and learning depth, thereby significantly improving user experience and educational effectiveness.

[0137] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The human-computer interaction program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium for the human-computer interaction program can be a magnetic disk, optical disk, read-only memory, or random access memory, etc.

[0138] The above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of this application. Therefore, any equivalent variations made in accordance with the claims of this application shall still fall within the scope of this application.

Claims

1. A human-computer interaction method, characterized in that, The method includes: In response to a user's selection command for a target object, a digital 3D model of the target object is displayed; the digital 3D model is a virtual three-dimensional object created through digital technology. If the display time of the digital 3D model reaches the preset time, obtain the descriptive text of the target object; Based on the introductory text, the digital 3D model is inferred using a multimodal large model to obtain segmentation regions related to the introductory text. The multimodal large model is used to align and segment the introductory text with 2D images from multiple perspectives. The 2D images from multiple perspectives are obtained by rendering different predefined viewpoints of the digital 3D model. The step of reasoning about the digital 3D model using a multimodal large model based on the introductory text to obtain segmented regions related to the introductory text includes: The digital 3D model is rendered from different predefined viewpoints to generate 2D images from multiple perspectives. Obtain the text instruction prompts present in the introductory text; The 2D images from multiple perspectives and the text instructions are input into a multimodal large model, and the multiple target 2D regions described by the text instructions are output. The multiple target 2D regions are merged into a 3D mask; Preprocess the 3D mask to obtain segmented regions related to the introductory text; Play the audio corresponding to the introductory text and display the segmented area.

2. The method according to claim 1, characterized in that, The method further includes: The device receives posture information sent by a sensing device. The posture information is obtained by the sensing device when the user operates the pre-marked three-dimensional prototype of the target object and is calculated based on the perceived spatial information. The three-dimensional prototype is a physical object manufactured based on the digital 3D model. Based on the posture information, the display orientation of the digital 3D model is updated to ensure that the posture of the digital 3D model is consistent with the posture of the three-dimensional prototype.

3. The method according to claim 1, characterized in that, The multimodal large model includes an alignment module and a segmentation module; The process of inputting the 2D images from multiple perspectives and the text prompts into a multimodal large model, and outputting multiple target 2D regions described by the text prompts, includes: The alignment module receives 2D images from multiple perspectives and the text instruction prompts, extracts visual features from the 2D images, extracts keywords and phrases from the text instruction prompts, uses deep learning technology to align the visual features and text features, and outputs multiple aligned feature maps. The segmentation module receives multiple feature maps output by the alignment module and 2D images from multiple viewpoints. Based on the multiple feature maps, it performs region segmentation on the 2D image from each viewpoint and outputs multiple target 2D regions described by the text instruction prompt.

4. The method according to claim 1, characterized in that, The preprocessing of the 3D mask yields segmented regions related to the introductory text, including: Refine the 3D mask; The refined 3D mask eliminates discontinuities caused by viewpoint changes; The 3D mask used to eliminate discontinuities caused by viewpoint changes filters out low-confidence areas. The 3D mask, which filters out low-confidence regions, is projected onto the digital 3D model to obtain the segmented region related to the introductory text.

5. The method according to claim 4, characterized in that, The refinement process employs a Gaussian geodesic reweighting algorithm; the elimination of discontinuities caused by viewpoint changes utilizes visibility smoothing technology; the filtering out of low-confidence regions employs a global filtering strategy; and the audio is obtained by converting the introductory text using the Text2Speech algorithm.

6. The method according to claim 1, comprising the following steps to generate a digital 3D model of the target object: Collect object information of the target object, including geometric data, image information, and on-site measurement information; Based on the geometric data, image information, and on-site measurement information, point cloud data of the target object is constructed; The point cloud data is converted into a network model to obtain the initial 3D structure; The initial 3D structure is modeled in three dimensions and textured to obtain a digital 3D model of the target object.

7. A human-computer interaction device implemented using the method according to any one of claims 1-6, characterized in that, The device includes: The display module is used to display a digital 3D model of the target object in response to a selection command for the target object; The acquisition module is used to acquire the description text of the target object when the display time of the digital 3D model reaches a preset time. The inference module is used to infer the digital 3D model based on the introductory text using a multimodal large model to obtain segmented regions related to the introductory text. The multimodal large model is used to align and segment the introductory text with 2D images from multiple perspectives. The 2D images from multiple perspectives are obtained by rendering different predefined viewpoints of the digital 3D model. The playback and display module is used to play the audio corresponding to the introductory text and display the segmented area.

8. A computer storage medium, characterized in that, The computer storage medium stores a plurality of instructions adapted for loading by a processor and executing the method as described in any one of claims 1-6.

9. An electronic device, characterized in that, include: A processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and executed as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Information processing method based on an augmented reality technology and electronic equipment

    CN109917918A

  • Three-dimensional model segmentation method and device

    CN117422848A

  • Directive 3D instance segmentation method based on chain type perception

    CN117593527A