Scene semantic graph generation method based on mixed reality interaction
By combining mixed reality interaction and scene semantic graph generation models, the problem of poor adaptability of 3D point clouds on edge devices is solved, the accuracy and robustness of scene semantic graphs are improved, the generation steps are simplified, and the user experience is enhanced.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIHANG UNIV
- Filing Date
- 2026-01-30
- Publication Date
- 2026-05-15
AI Technical Summary
Existing technologies have poor adaptability to 3D point clouds when generating scene semantic maps in edge device environments, resulting in low accuracy and robustness of the generated scene semantic maps. Furthermore, the generation process is cumbersome and leads to a poor user experience.
By using mixed reality interaction, scene point cloud and user interaction stroke information are obtained. A pre-trained scene semantic graph generation model is used to generate intent type information. Scene semantic graphs are generated through feature similarity and correction processing to reduce the impact of noise and occlusion point cloud data and simplify the generation process.
It improves the adaptability of 3D point clouds and the accuracy and robustness of generating scene semantic maps, simplifies the generation process, and enhances the user experience.
Smart Images

Figure CN122049219A_ABST
Abstract
Description
Technical Field
[0001] The embodiments disclosed herein relate to the field of computer technology, and more specifically to a method for generating scene semantic graphs based on mixed reality interaction. Background Technology
[0002] Scene semantic graphs, as a structured representation that simultaneously encodes object categories and their spatial relationships within a scene, have become an important intermediate representation supporting 3D scene cognition and reasoning. Existing methods for generating 3D scene semantic graphs mainly employ reasoning mechanisms based on graph neural networks, generating scene semantic graphs from 3D point clouds through joint prediction of nodes and edges using image, point cloud local features, and spatial neighborhood information.
[0003] However, in practice, the following problems often occur when using existing methods to generate scene semantic graphs: When generating scene semantic maps solely from 3D point clouds in an edge device environment, the 3D point clouds contain a lot of coarse, noisy, and highly occluded point cloud data, resulting in poor adaptability of the 3D point clouds. This leads to low accuracy and poor robustness of the generated scene semantic maps for the corresponding edge devices. Furthermore, it is necessary to display and label a large number of nodes or edges corresponding to the 3D point clouds to generate scene semantic maps, making the process cumbersome and time-consuming, which in turn leads to a poor user experience.
[0004] The information disclosed in this background section is only intended to enhance the understanding of the background of the inventive concept, and therefore may contain information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0005] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.
[0006] Some embodiments of this disclosure propose a method, apparatus, electronic device, and computer-readable medium for generating scene semantic graphs based on mixed reality interaction to solve one or more of the technical problems mentioned in the background section above.
[0007] In a first aspect, some embodiments of this disclosure propose a scene semantic graph generation method based on mixed reality interaction. The method includes: acquiring scene point clouds and user interaction stroke information corresponding to a target scene; generating intent type information corresponding to the user interaction stroke information based on the user interaction stroke information using a first sub-model included in a pre-trained scene semantic graph generation model; and performing the following steps using a second sub-model included in the pre-trained scene semantic graph generation model: generating a set of object 3D point information corresponding to the scene point cloud based on the scene point cloud; and generating the user interaction stroke information based on the intent type information. The alignment feature information is obtained; the feature similarity information that meets the preset threshold condition in the feature similarity information set between the above user interaction stroke information and the above object 3D point information set is determined as a feature similarity information group; based on the above alignment feature information, the above feature similarity information group, the first preset injection information, the second preset injection information, and each object 3D point in the above object 3D point information set associated with the above user interaction stroke information, the above object 3D point information set is corrected to obtain a changed object 3D point information set; based on the above changed object 3D point information set, a scene semantic map corresponding to the above target scene is generated.
[0008] Secondly, some embodiments of this disclosure provide a scene semantic graph generation apparatus, including an acquisition unit configured to acquire scene point clouds and user interaction stroke information corresponding to a target scene; a generation unit configured to generate intent type information corresponding to the user interaction stroke information based on the user interaction stroke information using a first sub-model included in a pre-trained scene semantic graph generation model; and an execution unit configured to execute the following steps using a second sub-model included in the pre-trained scene semantic graph generation model: generating a set of object 3D point information corresponding to the scene point cloud based on the scene point cloud; and generating intent type information corresponding to the user interaction stroke information based on the intent type information. Alignment feature information of user interaction stroke information; determine each feature similarity information that meets the preset threshold condition in the feature similarity information set between the above user interaction stroke information and the above object 3D point information set as a feature similarity information group; according to the above alignment feature information, the above feature similarity information group, the first preset injection information, the second preset injection information, and each object 3D point in the above object 3D point information set associated with the above user interaction stroke information, perform correction processing on the above object 3D point information set to obtain a modified object 3D point information set; generate a scene semantic map corresponding to the above target scene based on the above modified object 3D point information set.
[0009] Thirdly, some embodiments of this disclosure provide an electronic device, including: one or more processors; and a storage device having one or more programs stored thereon, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation of the first aspect above.
[0010] Fourthly, some embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the method described in any of the implementations of the first or second aspect.
[0011] The above-described embodiments of this disclosure have the following beneficial effects: Through the scene semantic graph generation method based on mixed reality interaction of some embodiments of this disclosure, the adaptability of 3D point clouds can be improved, thereby improving the accuracy and robustness of the generated scene semantic graph corresponding to the edge device, and the steps of generating scene semantic graphs can be simplified, thereby shortening the time spent generating scene semantic graphs and improving the user experience. Specifically, the poor adaptability of 3D point clouds, the low accuracy and robustness of the generated scene semantic maps for edge devices, and the cumbersome steps involved in generating scene semantic maps all contribute to the poor user experience. Specifically, when generating scene semantic maps from 3D point clouds in an edge device environment, the 3D point clouds contain a large amount of coarse, noisy, and highly occluded point cloud data, resulting in poor adaptability and consequently, low accuracy and robustness of the generated scene semantic maps for edge devices. Furthermore, the need to explicitly label a large number of nodes or edges corresponding to the 3D point clouds to generate scene semantic maps makes the process cumbersome and time-consuming, further contributing to the poor user experience. Therefore, some embodiments of the mixed reality-based scene semantic map generation method disclosed in this invention first acquire the scene point cloud of the corresponding target scene and user interaction stroke information. This allows the acquisition of user interaction strokes and related scene point clouds. Then, using the first sub-model of the pre-trained scene semantic graph generation model, intent type information corresponding to the aforementioned user interaction stroke information is generated based on the user interaction stroke information. Thus, the intent type corresponding to the user interaction stroke can be obtained. Next, using the second sub-model of the pre-trained scene semantic graph generation model, the following steps are performed: First, based on the aforementioned scene point cloud, a set of 3D object point information corresponding to the aforementioned scene point cloud is generated. Thus, a set of pre-segmented object nodes and the features of nodes and edges can be obtained. Second, based on the aforementioned intent type information, alignment feature information corresponding to the aforementioned user interaction stroke information is generated. Thus, the alignment features from the user interaction strokes to the 3D point cloud space can be obtained. Then, each feature similarity information in the feature similarity information set between the aforementioned user interaction stroke information and the aforementioned set of 3D object point information that meets a preset threshold condition is determined as a feature similarity information group. Thus, the similarity between the user interaction stroke alignment features and node features can be determined, obtaining the node feature values that meet the threshold condition and the corresponding similarity set. Subsequently, based on the alignment feature information, the feature similarity information group, the first preset injection information, the second preset injection information, and the three-dimensional points of the objects associated with the user interaction stroke information in the three-dimensional point information set of the objects, the three-dimensional point information set of the objects is corrected to obtain the modified three-dimensional point information set of the objects.Therefore, by jointly injecting strong and weak guiding features into the nodes and edges included in the aforementioned 3D point information set of the object, the aforementioned 3D point information set can be updated to obtain an updated 3D point information set of the object. Finally, based on the aforementioned modified 3D point information set of the object, a scene semantic map corresponding to the aforementioned target scene is generated. Thus, a semantic map of the target scene can be obtained. Because the scene semantic map is not generated solely from the 3D point cloud, but rather from the scene point cloud of the target scene and user interaction stroke information, the scene point cloud of the target scene can be corrected using user interaction stroke information. This reduces the coarse, noisy, and highly occluded point cloud data contained in the 3D point cloud, thereby improving the adaptability of the 3D point cloud and enhancing the accuracy and robustness of the generated scene semantic map corresponding to the edge device. Furthermore, because user interaction stroke information is introduced, the intent type can be generated from the introduced user interaction stroke information, thereby generating the scene semantic map without needing to explicitly label nodes or edges during correction. This simplifies the steps of generating the scene semantic map, shortens the time required, and ultimately improves the user's interactive experience. Furthermore, because the 3D point information set of objects in the target scene is corrected using the aforementioned user interaction stroke information, the accuracy of the generated scene semantic map can be improved by generating a scene semantic map based on the corrected 3D point information set. This, in turn, enhances the stability and generalization ability of the generated scene semantic map. It also improves the adaptability of the 3D point cloud, thereby increasing the accuracy and robustness of the generated scene semantic map for the corresponding edge device. Additionally, it simplifies the steps involved in generating the scene semantic map, reducing the time required and ultimately improving the user experience. Attached Figure Description
[0012] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.
[0013] Figure 1 This is a flowchart of some embodiments of the scene semantic graph generation method based on mixed reality interaction according to the present disclosure; Figure 2 This is a schematic diagram of the structure of some embodiments of the scene semantic graph generation device based on mixed reality interaction according to the present disclosure; Figure 3 This is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed Implementation
[0014] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0015] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.
[0016] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0017] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0018] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0019] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.
[0020] Figure 1 A flow 100 of some embodiments of the scene semantic graph generation method based on mixed reality interaction according to this disclosure is shown. The scene semantic graph generation method based on mixed reality interaction includes the following steps: Step 101: Obtain the scene point cloud and user interaction stroke information of the corresponding target scene.
[0021] In some embodiments, the execution entity (e.g., a computing device) of the scene semantic graph generation method based on mixed reality interaction can acquire scene point clouds and user interaction stroke information of the corresponding target scene. The target scene can represent any indoor three-dimensional space. The scene point cloud can represent discrete points of objects or scene surface morphology within the target scene. The user interaction stroke information can represent the trajectory drawn by the user using an input device in the target scene. The drawn trajectory consists of various three-dimensional coordinate points. The user interaction stroke information may include a 100Hz sampling point column. The sampling point column may include various three-dimensional coordinate points. The specific type of the input device is not limited here. For example, the input device can be a touchscreen. In practice, the execution entity can obtain the scene point cloud corresponding to the target scene through laser scanning and call the API interface of the input device to obtain the user interaction stroke information of the corresponding input device.
[0022] Step 102: The first sub-model of the pre-trained scene semantic graph generation model generates intent type information corresponding to the user interaction stroke information based on the user interaction stroke information.
[0023] In some embodiments, the aforementioned execution entity can generate intent type information corresponding to the aforementioned user interaction stroke information based on the first sub-model included in the pre-trained scene semantic graph generation model. The intent type information can characterize the type of the aforementioned user interaction stroke information. The intent type information can characterize stroke type 0, stroke type 1, or stroke type 2. Stroke type 0 can characterize the type of the aforementioned interaction stroke information as guided by homogeneous objects. The type of guided homogeneous objects can characterize the aforementioned user interaction stroke information as being used to classify objects in the aforementioned target scene that have the same physical or digital attributes. Stroke type 1 can characterize the type of the aforementioned interaction stroke information as guided by relation reconstruction. The type of guided relation reconstruction can characterize the aforementioned user interaction stroke information as adding, deleting, or modifying edge relationships between various nodes (objects) in the aforementioned target scene through line segments. The edge relationships can characterize the relative position, contact relationship (whether they are in contact), or topological relationship between various nodes (objects) in space. Stroke type 2 can characterize the type of the aforementioned interaction stroke information as a type of ambiguity disambiguation annotation. The aforementioned disambiguation annotation type can characterize the user interaction stroke information as fine-grained feature adjustment of target objects in the aforementioned target scene through closed trajectories. This fine-grained feature adjustment can characterize updating the node attributes of target objects in the aforementioned target scene. For example, fine-grained feature adjustment can correct the node attributes (e.g., label classification) of target objects. The aforementioned target objects can characterize objects in the aforementioned target scene. It should be noted that the objects in each object in the aforementioned target scene can be the nodes included in the aforementioned scene point cloud after coarse segmentation. The aforementioned first sub-model can include a temporal encoder, a spatial encoder, and an intent type information generation layer. The aforementioned temporal encoder can take user interaction stroke information as input and temporal dynamic features as output. The aforementioned temporal encoder can include a Long Short-Term Memory (LSTM) network layer, an attention mechanism layer, and a context vector generation layer. The aforementioned spatial encoder can take user interaction stroke information as input and spatial structure features as output. The aforementioned intent type information generation layer can take spatial structure features and temporal dynamic features as input and intent type information from user interaction stroke information as output.
[0024] In some optional implementations of certain embodiments, the aforementioned execution entity may generate intent type information corresponding to the aforementioned user interaction stroke information by using the first sub-model included in the pre-trained scene semantic graph generation model through the following steps: The first step involves generating temporal dynamic features and spatial structural features corresponding to the aforementioned user interaction stroke information. The temporal dynamic features characterize the change patterns, behavioral patterns, or state evolution processes of the user interaction stroke information over time. These temporal dynamic features can include temporal sequence features, motion speed variation features, and pause segment features. Both the temporal dynamic features and the spatial structural features can be 16-dimensional. The temporal sequence features characterize the change patterns of the user interaction stroke information over time. The motion speed variation features characterize the speed variation characteristics of the user interaction stroke information. The pause segment features characterize the pause characteristics of the user interaction stroke information. For example, the pause segment feature can characterize the pause duration of the user interaction stroke information during the drawing process. The spatial structural features characterize the geometric shape, spatial positioning attributes, and topological behavior representation characteristics of the user interaction stroke information in three-dimensional space. The geometric shape characterizes the number of three-dimensional coordinate points, total trajectory length, average curvature, and bounding volume of the user interaction stroke information. The aforementioned bounding volume can represent the spatial size occupied by a bounding box that encloses all the three-dimensional coordinate points of the aforementioned user-interactive stroke information. The aforementioned spatial positioning attributes can represent the centroid position, starting point position (coordinates of the starting point of the aforementioned user-interactive stroke information), and ending point position (coordinates of the ending point of the aforementioned user-interactive stroke information) of the aforementioned user-interactive stroke information. The aforementioned topological behavior representation can represent the cumulative directional change of the aforementioned user-interactive stroke information. The aforementioned cumulative directional change can represent the total degree of change in the trajectory direction of the aforementioned user-interactive stroke information along the entire trajectory; its value is the sum of the directional changes between adjacent trajectory segments. The aforementioned directional change can represent the angle between two adjacent trajectory directions.
[0025] The second step is to generate intent type information corresponding to the aforementioned user interaction stroke information based on the aforementioned time dynamic features and spatial structure features.
[0026] In some optional implementations of certain embodiments, the first sub-model described above can generate temporal dynamic features and spatial structural features corresponding to the user interaction stroke information by following the following steps: First, the aforementioned execution entity can input the aforementioned user interaction stroke information into the aforementioned time encoder to obtain the time dynamic features. The execution steps of the aforementioned time encoder include: The first step is to determine the corresponding state feature sequence information based on the aforementioned user interaction stroke information. This state feature sequence information can characterize the hidden state sequence of the aforementioned user interaction stroke information. In practice, the aforementioned time encoder can input the aforementioned user interaction stroke information into a Long Short-Term Memory (LSTM) network layer to obtain the hidden state sequence.
[0027] The second step involves multiplying the preset weight matrix and the preset weight vector to determine the weight matrix. The preset weight vector represents the degree of attention the temporal encoder pays to different positions in the input hidden state sequence. The preset weight matrix represents the transformation matrix of the hidden state sequence from the Long Short-Term Memory (LSTM) network layer to the attention mechanism layer. Both the preset weight vector and the preset weight matrix represent the weight vector and weight matrix obtained after modifying the randomly generated weight vector and weight matrix during the training of the scene semantic graph generation model; they are fixed, learnable parameters in the trained scene semantic graph generation model. The specific values of the preset weight matrix and the preset weight vector are not limited here.
[0028] The third step is to normalize the weight matrix to obtain a normalized weight matrix. In practice, the time encoder described above can use the softmax function to normalize the weight matrix to obtain the normalized weight matrix.
[0029] The fourth step involves updating the normalized weight matrix to obtain the changed weight matrix. This changed weight matrix represents the transpose of the normalized weight matrix. In practice, the time encoder can use the transpose of the normalized weight matrix as the changed weight matrix.
[0030] Fifth, based on the aforementioned state feature sequence information and the aforementioned change weight matrix, determine the temporal dynamic features corresponding to the aforementioned user interaction stroke information. In practice, firstly, the aforementioned time encoder can multiply the aforementioned state feature sequence information with the aforementioned change weight matrix to obtain the temporal dynamic features corresponding to the aforementioned user interaction stroke information.
[0031] Then, the aforementioned execution entity can input the aforementioned user interaction stroke information into the aforementioned spatial encoder to obtain spatial structural features. The execution steps of the aforementioned spatial encoder include: The first step involves feature extraction processing of the aforementioned user interaction stroke information to obtain interactive stroke feature information. This interactive stroke feature information characterizes the geometric shape, spatial positioning attributes, and topological behavior of the user interaction stroke information in three-dimensional space. Specifically, it includes the number of sampling points in the user interaction stroke information, the trajectory length of the user interaction stroke information, the average curvature of the user interaction stroke information, the enclosing volume of the minimum closed volume formed by the sampling points in space, the centroid position representing the average position of all coordinate points in the trajectory, the start and end positions of the first and last three-dimensional coordinate points of the trajectory, and the cumulative direction change vector representing the accumulated value of the direction change of each line segment in the trajectory. In practice, the spatial encoder can first determine the interactive stroke feature information corresponding to the aforementioned interactive stroke information using the NumPy / SciPy library in Python.
[0032] The second step is to determine the spatial structure features corresponding to the aforementioned user interaction stroke information based on the above-mentioned interactive stroke feature information. In practice, the execution entity can directly concatenate the number of points, total trajectory length, average curvature, bounding volume, centroid position, start and end point positions, and cumulative direction change vector included in the above-mentioned interactive stroke feature information using the above-mentioned spatial encoder to obtain the spatial structure features.
[0033] In some optional implementations of certain embodiments, the first sub-model described above can generate intent type information corresponding to the user interaction stroke information by following the steps described above based on the aforementioned temporal dynamic features and spatial structural features: First, the aforementioned executing entity can input the aforementioned temporal dynamic features and spatial structural features into the aforementioned intent type information generation layer to obtain intent type information. The execution steps of the aforementioned intent type information generation layer include: The first step involves concatenating the aforementioned temporal dynamic features and spatial structural features to obtain the interactive stroke feature vector corresponding to the aforementioned user interaction stroke information. This interactive stroke feature vector represents the feature vector obtained by concatenating the spatial structural features and temporal dynamic features of the aforementioned interactive stroke information. In practice, the aforementioned intent type information generation layer can directly concatenate the aforementioned temporal dynamic features and spatial structural features to obtain the interactive stroke feature vector corresponding to the aforementioned user interaction stroke information.
[0034] The second step involves classifying the aforementioned interactive stroke feature vectors to obtain the interactive stroke matrix information corresponding to the user's interactive stroke information. This interactive stroke matrix information can represent a three-dimensional vector. It can include various numerical values. These values represent the scores for different intent types of the user's interactive stroke information; for example, the intent type information can be stroke type 0. In practice, the intent type information generation layer can use a multilayer perceptron classifier to classify the interactive stroke feature vectors and obtain the interactive stroke matrix information.
[0035] The third step is to generate intent type information corresponding to the aforementioned user interaction stroke information based on the above-mentioned interaction stroke matrix information.
[0036] In some optional implementations of certain embodiments, the intent type information generation layer can generate intent type information corresponding to the user interaction stroke information based on the interaction stroke matrix information through the following steps: The first step is to normalize the aforementioned interactive stroke matrix information to obtain normalized interactive stroke matrix information. This normalized interactive stroke matrix information can include individual interactive stroke values. Each interactive stroke value represents the probability of the type represented by the interactive stroke matrix information. The sum of all the interactive stroke values in the normalized interactive stroke matrix information is 1. For example, if the interactive stroke value representing the user's interactive stroke information as type 0 is 0.3, it can represent that the probability of the user's interactive stroke information being type 0 is 0.3. The normalized interactive stroke matrix information can be [0.3, 0.5, 0.2]. In practice, the intent type information generation layer can use the softmax function to normalize the interactive stroke matrix information to obtain the interactive stroke matrix information. Each interactive stroke value in the normalized interactive stroke matrix information corresponds to an intent type.
[0037] The second step involves determining the target interactive stroke values from the normalized interactive stroke matrix information that meet preset selection criteria. These preset selection criteria can be that the interactive stroke value is the largest among the normalized interactive stroke matrix information. In practice, the intent type information generation layer can determine the largest interactive stroke value included in the normalized interactive stroke matrix information as the target interactive stroke value.
[0038] The third step is to determine the intent type information corresponding to the above target interaction stroke values as the intent type information corresponding to the above user interaction stroke information.
[0039] Step 103: Generate the second sub-model included in the model using the pre-trained scene semantic graph, and perform the following steps: Step 1031: Generate a set of 3D point information of objects corresponding to the scene point cloud based on the scene point cloud.
[0040] In some embodiments, the execution entity can generate a set of 3D point information of objects corresponding to the scene point cloud based on the scene point cloud. The 3D point information of the objects in the set of 3D point information can characterize the features of the target object in 3D space. The 3D point information of the objects in the set of 3D point information can include geometric and physical attributes, semantic and category information, and representations of relationships between objects. The geometric and physical attributes can characterize the geometric features of the target object (e.g., centroid, bounding box, direction of normal, surface curvature). The semantic and category information can characterize the category of the target object (e.g., chair). The representations of relationships between objects can characterize the relationship type of the target object in the 3D space (e.g., the target object is adjacent to other objects). In practice, the execution entity can obtain a scene point cloud cluster by coarsely segmenting the scene point cloud, and then extract the features of node and edge relationships in the scene point cloud through the obj_encoder layer and the rel_encoder layer to obtain the set of 3D point information of objects corresponding to the scene point cloud. The coarse segmentation can be Euclidean fusion. The aforementioned `obj_encoder` and `rel_encoder` layers can each include one ReLU activation function and three convolutional layers. The second sub-model can include a guided injection layer and a scene semantic graph generation layer. The guided injection layer can take the aforementioned intent type information, temporal dynamic features, and spatial structure features as input, and output a set of 3D point information of the changed object. The scene semantic graph generation layer can take the aforementioned set of 3D point information of the changed object as input and output a scene semantic graph. The scene semantic graph generation layer can include a graph neural network and two decoders. The graph neural network can include GraphEdgeAttenNetworkLayers layers for multi-layer message passing of graph structure data. The GraphEdgeAttenNetworkLayers layers include two MSG_FAN layers and one dropout layer. The MSG_FAN layer can process graph data and update node or edge features. The dropout layer can prevent overfitting. The two decoders can each include three fully connected layers, one dropout layer, and one ReLU activation function layer.
[0041] Step 1032: Generate alignment feature information corresponding to the user interaction stroke information based on the intent type information.
[0042] In some embodiments, the execution entity can generate alignment feature information corresponding to the user interaction stroke information based on the intent type information. The alignment feature information can characterize the feature vector of the user interaction stroke information aligned to the three-dimensional space. In practice, the guidance injection layer can input the intent type information, the temporal dynamic features, the spatial structure features, and the interaction stroke feature vector into a first preset function to obtain the alignment feature information corresponding to the user interaction stroke information.
[0043] As an example, the first preset function mentioned above can be: .
[0044] in, It can represent intent type information. It can characterize the alignment feature information corresponding to the above user interaction stroke information. This can be used to characterize the Dirac function. For example, when t=0, =1, =0, =0; when t=1 =1, =0, =0, when t=2 =1, =0, =0. , These represent the projectors corresponding to stroke types 0 and 1, respectively. Interactive stroke feature vectors can represent the above-mentioned user interaction stroke information. This can characterize the features of the aforementioned user interaction stroke information and the point cloud context within the enclosing region. The aforementioned enclosing region can characterize the area enclosed by the aforementioned user interaction stroke information. This can characterize a projector corresponding to two types of strokes. Among them, the above... and the above Each layer can include a first linear layer, an activation function, and a second linear layer. The input to the first linear layer can be 32-dimensional, and the output can be 64-dimensional. The input to the second linear layer can be 64-dimensional, and the output can be 256-dimensional. The activation function can be the ReLU function. It may include fully connected layers, activation functions, spatial encoders, classifiers, and a preset function information layer. The fully connected layers may include the first linear layer and the second linear layer mentioned above. The classifier may include three fully connected layers, one dropout layer, and one ReLU activation function layer. The preset function information layer may be an identity mapping.
[0045] Step 1033: Determine each feature similarity information that meets the preset threshold condition in the feature similarity information set between the user interaction stroke information and the object's three-dimensional point information set as a feature similarity information group.
[0046] In some embodiments, the executing entity can determine each feature similarity information that meets a preset threshold condition in the feature similarity information set between the user interaction stroke information and the object 3D point information set as a feature similarity information group. The feature similarity information in the feature similarity information group can characterize the correspondence between the features and similarities of the object 3D point information in the object 3D point information set. For example, the feature similarity information in the feature similarity information group can be: "Feature: xx, Similarity: 70%". The similarity can characterize the similarity between the features of the object 3D point information and the alignment feature information. The feature similarity information in the feature similarity information group can include the features and similarities of the object 3D point information in the object 3D point information set. The preset threshold condition can be that the similarity included in the feature similarity information set is greater than 70%. In practice, firstly, for each object 3D point information in the object 3D point information set, the executing entity can determine the similarity of the object 3D point information by the cosine similarity between the object 3D point information and the alignment feature information. Then, the similarity and the corresponding feature are determined as feature similarity information. Finally, the feature similarity information groups whose similarity is greater than a preset similarity are identified as feature similarity information groups. The preset similarity can be 70%.
[0047] In the process of adopting technical solutions to address the aforementioned technical problems, the following technical problem often arises: Correcting only local areas (partial 3D point information within the scene) results in poor consistency of the generated scene semantic map. This necessitates frequent maintenance and updates to improve the accuracy of the scene semantic map, leading to high computational resource consumption and a poor user experience with an inaccurate scene semantic map. The conventional solution to this second technical problem is to detect errors in local areas and correct the erroneous content. However, considering the shortcomings of detecting and correcting errors in local areas, and leveraging the advantages of our organization in correcting local areas within the scene, we have decided to adopt the following solution: Step 1034: Based on the alignment feature information, feature similarity information group, first preset injection information, second preset injection information, and the object three-dimensional point information set associated with the user interaction stroke information, the object three-dimensional point information set is corrected to obtain the changed object three-dimensional point information set.
[0048] In some embodiments, the execution entity can perform correction processing on the object 3D point information set based on the alignment feature information, the feature similarity information group, the first preset injection information, the second preset injection information, and the object 3D point information set associated with the user interaction stroke information, to obtain a modified object 3D point information set. The object 3D points associated with the user interaction stroke information may include not only nodes (representing the object's 3D points) but also edge relationships. The object 3D points associated with the user interaction stroke information can represent 3D points or edge relationships connected to the user interaction stroke information, nodes within the area enclosed by the user interaction stroke information, or nodes near the user interaction stroke information. The distance between nodes near the user interaction stroke information and nodes represented by the user interaction stroke information is not limited; for example, the distance can be 5mm. The first preset injection information can represent the injection intensity of the alignment feature information. The first preset injection information can be 35%. The second preset injection information can represent the injection intensity of the alignment feature information. The second preset injection information can be 15%. The guided injection layer can perform the following steps: The first step is to determine the target feature vector by multiplying the alignment feature information and the first preset injection information.
[0049] The second step involves determining a set of fused feature vectors based on the target feature vector and the feature vectors representing the 3D points of each object associated with the user-interactive stroke information. The fused feature vector in this set represents the vector obtained by adding the target feature vector to the feature vector representing a 3D point of an object associated with the user-interactive stroke information. The fused feature vectors in this set can have a one-to-one correspondence with the 3D points of the objects. In practice, for each 3D point, the executing entity can add the target feature vector to the feature vector representing the 3D point to obtain the fused feature vector. Finally, the resulting fused feature vectors are defined as the set of fused feature vectors.
[0050] The third step is to determine the target value by summing the preset value and the first preset injected information. The preset value can represent a pre-defined value. The preset value can be 1. The target value can be used to improve the stability of the final result range.
[0051] The fourth step involves correcting the 3D points of each object associated with the user-interactive stroke information based on the aforementioned fused feature vector set and target value, thereby obtaining a first set of 3D object point information. The first object 3D point information in this set represents the 3D object points obtained after correcting the 3D object points associated with the user-interactive stroke information. In practice, firstly, for each 3D object point associated with the user-interactive stroke information, the following steps are performed: the ratio of the fused feature vector corresponding to the 3D object point to the target value is determined as the corrected 3D object point, which is then used as the first object 3D point information. Then, the obtained first object 3D point information is used to form the first object 3D point information set.
[0052] It should be noted that the aforementioned first preset injection information can be dynamically adjusted based on the aforementioned intent type information, thereby allowing for dynamic reduction. This reduces over-correction caused by abnormal situations such as user mislabeling, thereby enhancing the robustness and flexibility of the pre-trained scene semantic graph generation model.
[0053] Fifth, based on the aforementioned second preset injection information, the aforementioned alignment feature information, and the aforementioned target value, the aforementioned feature similarity information group is fused to obtain a second object 3D point information set. The second object 3D point information in the aforementioned second object 3D point information set can represent the object 3D points obtained by fusing each feature included in the aforementioned feature similarity information group with the corresponding object 3D points in the aforementioned object 3D point information set. In practice, the aforementioned guided injection layer can input the features of each node or edge included in the aforementioned feature similarity information group, the aforementioned second preset injection information, and the aforementioned alignment feature information into a third preset function to obtain the second object 3D point information set.
[0054] As an example, the third preset function mentioned above can be: .
[0055] in, This can characterize the aforementioned second preset injection information. It can characterize the features of a node or edge of the three-dimensional point information of an object included in the above set of three-dimensional point information of an object. It can characterize the corresponding features in the above feature similarity information group. similarity, It can represent the corresponding The second object's three-dimensional point information. The index can be characterized as The alignment features of the strokes.
[0056] Step 6: Based on the aforementioned first set of 3D object point information and the aforementioned second set of 3D object point information, update the aforementioned set of 3D object point information to obtain a modified set of 3D object point information. This modified set of 3D object point information includes the updated features of nodes or edges, as well as the features of other nodes or edges that have not been updated. In practice, the aforementioned guidance injection layer can replace the 3D object point information corresponding to the first set of 3D object point information in the aforementioned set of 3D object point information with the first set of 3D object point information, and replace the 3D object point information corresponding to the second set of 3D object point information in the aforementioned set of 3D object point information with the second set of 3D object point information, to obtain the modified set of 3D object point information. The modified 3D object point information in this modified set of 3D object point information can represent either the updated or unupdated 3D object point information. The updated 3D object point information can represent either the first 3D object point information in the first set of 3D object point information or the second 3D object point information in the second set of 3D object point information. The unupdated 3D object point information can represent 3D object point information in the aforementioned set of 3D object point information that does not meet the preset update conditions. The above-mentioned preset update conditions can be that the object's 3D point information is not associated with the object's 3D points of the user interaction stroke information and the similarity of the object's 3D point information does not meet the above-mentioned preset threshold conditions.
[0057] The above technical solution, as an inventive point of this disclosure, solves technical problem two: "The generated scene semantic graph has poor consistency, leading to frequent maintenance and updates to improve its accuracy, resulting in high computational resource consumption and a poor user experience." The reasons for this poor consistency, frequent maintenance and updates, high computational resource consumption, and poor user experience are as follows: Correcting only local areas (partial 3D point information) in the scene results in poor consistency, necessitating frequent maintenance and updates, leading to high computational resource consumption and a poor user experience. Solving these factors improves the consistency of the generated scene semantic graph, reduces the frequency of maintenance and updates, reduces computational resource consumption, and enhances the user experience. To achieve this effect, the disclosed method for generating scene semantic graphs based on mixed reality interaction firstly corrects the 3D point information of each object in the object 3D point information set that is associated with the aforementioned user interaction stroke information. Then, it determines the similarity between the user interaction stroke information and the object 3D point information. Finally, it fuses the similarity regions based on the similarity. This improves the consistency of the generated scene semantic graph, reduces the frequency of maintenance and updates, reduces computational resource consumption, and enhances the user experience.
[0058] Step 1035: Generate a scene semantic map of the corresponding target scene based on the set of 3D point information of the changed object.
[0059] In some embodiments, the execution entity can generate a scene semantic map corresponding to the target scene based on the set of 3D point information of the changed object. In practice, the execution entity can input the set of 3D point information of the changed object into the scene semantic map generation layer to obtain a scene semantic map corresponding to the target scene.
[0060] Optionally, after step 1035, the executing entity may further label the node semantic label information and relation semantic label information corresponding to the scene semantic graph into a three-dimensional space. The node semantic label information can represent the node semantic labels of the scene semantic graph. For example, a node semantic label can be "table" or "sofa". The relation semantic label information can represent the relation semantic labels of the scene semantic graph. For example, an edge semantic label can be "standing on" or "attached to". In practice, the executing entity can use a labeling algorithm based on 3D detection / segmentation to label the node semantic label information and relation semantic label information corresponding to the scene semantic graph into a three-dimensional space.
[0061] In addressing the aforementioned technical problems using technical solutions, the following third technical issue often arises: Training objectives are often singular and isolated, typically relying solely on core loss functions such as node or relation classification. This leads to a lack of comprehensive constraints on model behavior, resulting in semantic maps where different parts of the same object are predicted as different categories and the precise spatial relationship between user strokes and point clouds is ignored. Ultimately, this results in insufficient stability and generalization ability of the model during complex inference processes. Furthermore, multi-stage training increases redundant computation and storage consumption. A conventional solution to this third technical problem is to introduce multi-task auxiliary loss to jointly optimize node and relation classification. Considering the advantages of our institution in refining scene regions using multi-task auxiliary loss for joint optimization of node and relation classification, we have decided to adopt the following solution: The first step is to acquire a dataset, which includes 3D point cloud data and simulated user interaction stroke information. The 3D point cloud data represents a set of discrete points representing objects or surface features within a scene. This 3D point cloud data is used to train the pre-trained scene semantic graph generation model. The simulated user interaction stroke information is used to train the pre-trained scene semantic graph generation model. This simulated user interaction stroke information may include a set of stroke sampling points, intent type information, target object index, relation edge index, expected semantic label, number of sampling points, and trajectory length. The 3D point cloud data may include semantic labels for real nodes and edges. The set of stroke sampling points represents a series of discrete points collected along the stroke trajectory. The target object index represents a unique identifier or index of the object associated with the user interaction stroke information. The relation edge index represents a unique identifier or index for identifying and accessing specific relation edges. The expected semantic label represents the semantic result that the user expects to trigger through the user interaction stroke information. This semantic result can be the result after using the user interaction stroke information. For example, semantic results could be the result of repairing a target object using user-interacted stroke information.
[0062] The second step involves performing the following training steps based on the dataset: The first sub-step involves inputting 3D point cloud data of at least one scene from the dataset and user interaction stroke information into an initial neural network to obtain the corresponding intent type information, semantic labels for nodes and edges (edge relationships) of the scene semantic graph for each user interaction stroke in the aforementioned at least one scene. It should be noted that the structure of the initial neural network is the same as the structure of the scene semantic graph generation model, and will not be repeated here.
[0063] The second sub-step generates a loss value based on the intent type information corresponding to each user interaction stroke, the semantic labels of nodes and edges in the scene semantic graph, the semantic labels of real nodes and edges in the scene semantic graph, and the real intent type corresponding to each user interaction stroke. This loss value can include stroke-guided intent classification loss, type 2 stroke cue autoencoder classification loss, node classification loss, relation classification loss, a first loss value, a third loss value, a fourth loss value, and a fifth loss value. The stroke-guided intent classification loss can be the loss value used to predict the intent type information of each stroke, and can be generated using the cross-entropy loss function to supervise the consistency between the predicted intent type information and the real intent type. The type 2 stroke cue autoencoder classification loss can represent the loss value generated using the cross-entropy loss function, used for supervision to improve the distinguishability of the enclosing region and the reasoning assistance capability. The node classification loss can be generated using the standard cross-entropy loss function, a loss value used to supervise the relationship between the predicted category (the semantic label of the node in the predicted scene semantic graph) and its real label (the semantic label of the node in the real scene semantic graph) of each object node, to optimize node classification accuracy. The aforementioned relation classification loss characterizes the edge relation prediction used for connecting node pairs in the scene semantic graph. The loss value generated using the cross-entropy loss function improves the consistency between the relation categories output by the supervised inference network (the semantic labels of the predicted edges in the scene semantic graph) and the true relation labels (the semantic labels of the actual edges in the scene semantic graph), thus enhancing the accuracy of edge semantic inference. In practice, the aforementioned execution entity can generate a loss value based on the intent type information corresponding to each user interaction stroke, the semantic labels of nodes and edges in the scene semantic graph, the semantic labels of the actual nodes and edges in the scene semantic graph, and the actual intent type corresponding to each user interaction stroke information through the following steps: The first step involves the execution entity inputting the guiding features of user interaction stroke information and the features of nodes or edges into the first loss function to obtain the first loss value. As an example, the first loss function can be: .
[0064] in, It can characterize the alignment features of stroke type 0. It can characterize the alignment features of two types of strokes. It can characterize the alignment features of one type of stroke. , An index that can represent stroke information from user interaction. , , These can be used to represent the index of a node. It can represent nodes in user interaction stroke information With nodes The index of the edge between them. Based on the representation index The characteristics of the nodes. It can represent nodes and nodes The characteristics of the edges formed between them. The function can represent the mapping relationship between the node index and the index of the aforementioned user-interactive stroke information, or the mapping relationship between the node index pair and the index of the aforementioned user-interactive stroke information. It can represent the number of nodes. It can represent the number of node pairs, that is, the number of edges. It can represent the first loss value.
[0065] The second step is to determine the classification loss value by summing the above stroke guidance intention classification loss, the above type 2 stroke prompt autoencoder classification loss, the node classification loss and the relation classification loss.
[0066] The third step involves inputting the classification loss value, the preset loss parameter, and the first loss value into the second loss function to obtain the second loss value. The preset loss parameter can be a pre-defined parameter, such as 1.0.
[0067] As an example, the second loss function can be: .
[0068] in, The loss value can be one of the above classifications. The above-mentioned preset loss parameters can be used. It can be the second loss value mentioned above.
[0069] The fourth step involves inputting the semantic labels of the nodes and edges in the aforementioned scene semantic graph into the third loss function to obtain the third loss value. As an example, the third loss function can be: , .
[0070] in, For predicted nodes Semantic tags. For predicted nodes Semantic tags. For predicted nodes and nodes The semantic labels of the edges between them. This can be any 3D point cloud traversed in 3D space by stroke type 0. (The above...) This can be achieved through point cloud detection methods, which detect the geometric contact relationship between the stroke path and the point cloud. For example, the detection method can be an octree / KD tree. , Each can be a neighborhood set near the starting and ending points of a stroke type, typically composed of local segments extracted within a fixed spatial fitting radius centered on the endpoint. , They can respectively characterize the first The, the The weighting factor of each node is defined as the distance function between the center point of the segment and the corresponding stroke endpoint. The weights can be assigned to the mutual influence between two nodes. It is one of the predefined relationship types in the scene semantic graph, indicating that two objects belong to the same whole.
[0071] Fifth, the aforementioned execution entity can input the predicted semantic labels of nodes and edges into the fourth loss function to obtain the fourth loss value. As an example, the fourth loss function can be: .
[0072] in, It can represent the number of nodes in the neighborhood set (e.g., nodes within a radius of 5mm centered at the start or end point of stroke type 1) near the start point (e.g., 5mm) of stroke type 1. It can represent the number of each node in the neighborhood set near the endpoint. It can represent the index of a node when traversing the neighborhood set near the starting point. It can represent the index of a node when traversing the neighborhood set near the endpoint. This can represent the foreground group corresponding to the two stroke types. The foreground group can represent a set of high-response nodes near the center of the prompt area. These high-response nodes can be nodes with a confidence level greater than or equal to 0.9 belonging to the object the user intends to point to. The prompt area can represent the enclosed region of the two stroke types mentioned above. This can represent background groups corresponding to two types of strokes. These background groups can represent elements in the prompt area that... The set of nodes obtained by sampling with Gaussian probability decay, using a first preset value as the radius and taking a point as the origin. Here, the specific value of the first preset value is not limited. For example, the first preset value can be 5. . This can represent the number of nodes in the aforementioned foreground group. It can represent the number of nodes in the background group mentioned above. Nodes that can be characterized for prediction and nodes The semantic labels of the edges between them. The index can be characterized as The characteristics of the nodes. The index can be characterized as The characteristics of the nodes. It can characterize the weights of a normal distribution. The index of the background group mentioned above can be... The three-dimensional coordinates of the node. It can represent the center coordinates of the above-mentioned prompt area. This can represent the radius of the aforementioned indicated area. It can characterize the mean of a normal distribution. =2 , It can characterize the standard deviation of a normal distribution. It can characterize the response strength of the "same part" relation label in the false true prediction results. An indicator vector that can represent the "same part" relationship label. It can characterize tiny constants that avoid numerical instability.
[0073] In the sixth step, the aforementioned execution entity can input the predicted semantic labels of the nodes and edges into the fifth loss function to obtain the fifth loss value. As an example, the fifth loss function can be: .
[0074] in, It can represent a vertically upward unit vector in the Cartesian coordinate system. It can represent a group of vertical plane objects. The vertical plane objects in the above group of vertical plane objects can represent objects (nodes) that are perpendicular to the plane. It can represent a group of horizontal plane objects. The horizontal plane objects in the above group can represent objects that are horizontal to a plane. The index can be any of the above-mentioned vertical plane object groups or the above-mentioned horizontal plane object groups. The unit normal vector along the principal direction of the object. It can characterize the object's pair and The edge relationships between them. It can represent the total number of edge relationships. It can represent objects A collection of three-dimensional point clouds. It can represent objects A collection of three-dimensional point clouds. , They can represent objects respectively and objects The surrounding box. Can characterize , The crossover ratio of the areas obtained by projecting onto the xoy plane. It can represent an indicator vector, which takes a value of 1 when the relation label is a hierarchical relation (such as embedded or supporting), and 0 otherwise. It is used to constrain the vertical projection relationship in the hierarchical structure.
[0075] Step 7: Input the third, fourth, and fifth loss values mentioned above into the sixth loss function to obtain the loss value. As an example, the sixth pre-loss function can be: .
[0076] in, The weighting coefficients can represent the fifth loss value mentioned above. It can characterize the fifth loss value. This can characterize the third loss value mentioned above. It can characterize the fourth loss value. It can represent the sixth loss value.
[0077] The third sub-step involves determining, based on the loss value, whether the initial neural network has reached a preset optimization objective. This optimization objective can be defined as minimizing the loss values obtained in a preset number of rounds of adjustments to the initial neural network. The preset number of rounds can be 1000.
[0078] The fourth sub-step is to determine the initial neural network as the trained scene semantic graph generation model in response to the determination that the initial neural network has achieved the above optimization objective.
[0079] Optionally, the aforementioned execution entity may also, in response to determining that the initial neural network has not achieved the aforementioned optimization objective, adjust the network parameters of the initial neural network, and use a dataset composed of unused data, use the adjusted initial neural network as the initial neural network, and execute the aforementioned training steps again.
[0080] The above-described technical solution, as an inventive point of this disclosure, solves technical problem three: "The generated scene semantic graph is prone to having different parts of the same object predicted as different categories and ignoring the precise spatial relationship between user strokes and point clouds, resulting in insufficient stability and generalization ability of the model in complex reasoning processes, and increasing redundant computation and storage consumption." The reasons why the generated scene semantic graph is prone to having different parts of the same object predicted as different categories and ignoring the precise spatial relationship between user strokes and point clouds, leading to insufficient stability and generalization ability of the model in complex reasoning processes, are as follows: The training objective is singular and isolated, usually relying only on core loss functions such as node or relationship classification. This results in a lack of comprehensive constraints on the model's behavior. Therefore, the generated semantic graph will have different parts of the same object predicted as different categories and ignore the precise spatial relationship between user strokes and point clouds, ultimately leading to insufficient stability and generalization ability of the model in complex reasoning processes. Furthermore, the use of multi-stage training increases redundant computation and storage consumption. If the above factors are addressed, the consistency of predictions for parts of the same object can be determined, and the spatial relationship between strokes and point clouds can be accurately captured. This improves the stability and generalization ability of the model in complex reasoning, while reducing the consumption of redundant computation and storage resources. To achieve this effect, the scene semantic graph generation method based on mixed reality interaction disclosed in this paper adopts a joint optimization strategy of core loss and intervention loss. The core loss includes stroke guidance intent classification loss, type 2 stroke prompt autoencoder classification loss, node classification loss, relation classification loss, and a first loss value, which is used to ensure the basic accuracy of reasoning. The intervention loss includes local consistency constraints (third loss value), error rejection constraints (fourth loss value), and geometric sensitivity constraints (fifth loss value), which are used to further improve the stability and overall consistency of reasoning results in local guided regions. Through the synergistic optimization of core loss and intervention loss, the accuracy, stability, and robustness of the reasoning process are improved. Thus, the consistency of predictions for parts of the same object can be determined, and the spatial relationship between strokes and point clouds can be accurately captured, thereby improving the stability and generalization ability of the model in complex reasoning, while reducing the consumption of redundant computation and storage resources.
[0081] The above-described embodiments of this disclosure have the following beneficial effects: Through the scene semantic graph generation method based on mixed reality interaction of some embodiments of this disclosure, the adaptability of 3D point clouds can be improved, thereby improving the accuracy and robustness of the generated scene semantic graph corresponding to the edge device, and the steps of generating scene semantic graphs can be simplified, thereby shortening the time spent generating scene semantic graphs and improving the user experience. Specifically, the poor adaptability of 3D point clouds, the low accuracy and robustness of the generated scene semantic maps for edge devices, and the cumbersome steps involved in generating scene semantic maps all contribute to the poor user experience. Specifically, when generating scene semantic maps from 3D point clouds in an edge device environment, the 3D point clouds contain a large amount of coarse, noisy, and highly occluded point cloud data, resulting in poor adaptability and consequently, low accuracy and robustness of the generated scene semantic maps for edge devices. Furthermore, the need to explicitly label a large number of nodes or edges corresponding to the 3D point clouds to generate scene semantic maps makes the process cumbersome and time-consuming, further contributing to the poor user experience. Therefore, some embodiments of the mixed reality-based scene semantic map generation method disclosed in this invention first acquire the scene point cloud of the corresponding target scene and user interaction stroke information. This allows the acquisition of user interaction strokes and related scene point clouds. Then, using the first sub-model of the pre-trained scene semantic graph generation model, intent type information corresponding to the aforementioned user interaction stroke information is generated based on the user interaction stroke information. Thus, the intent type corresponding to the user interaction stroke can be obtained. Next, using the second sub-model of the pre-trained scene semantic graph generation model, the following steps are performed: First, based on the aforementioned scene point cloud, a set of 3D object point information corresponding to the aforementioned scene point cloud is generated. Thus, a set of pre-segmented object nodes and the features of nodes and edges can be obtained. Second, based on the aforementioned intent type information, alignment feature information corresponding to the aforementioned user interaction stroke information is generated. Thus, the alignment features from the user interaction strokes to the 3D point cloud space can be obtained. Then, each feature similarity information in the feature similarity information set between the aforementioned user interaction stroke information and the aforementioned set of 3D object point information that meets a preset threshold condition is determined as a feature similarity information group. Thus, the similarity between the user interaction stroke alignment features and node features can be determined, obtaining the node feature values that meet the threshold condition and the corresponding similarity set. Subsequently, based on the alignment feature information, the feature similarity information group, the first preset injection information, the second preset injection information, and the three-dimensional points of the objects associated with the user interaction stroke information in the three-dimensional point information set of the objects, the three-dimensional point information set of the objects is corrected to obtain the modified three-dimensional point information set of the objects.Therefore, by jointly injecting strong and weak guiding features into the nodes and edges included in the aforementioned 3D point information set of the object, the aforementioned 3D point information set can be updated to obtain an updated 3D point information set of the object. Finally, based on the aforementioned modified 3D point information set of the object, a scene semantic map corresponding to the aforementioned target scene is generated. Thus, a semantic map of the target scene can be obtained. Because the scene semantic map is not generated solely from the 3D point cloud, but rather from the scene point cloud of the target scene and user interaction stroke information, the scene point cloud of the target scene can be corrected using user interaction stroke information. This reduces the coarse, noisy, and highly occluded point cloud data contained in the 3D point cloud, thereby improving the adaptability of the 3D point cloud and enhancing the accuracy and robustness of the generated scene semantic map corresponding to the edge device. Furthermore, because user interaction stroke information is introduced, the intent type can be generated from the introduced user interaction stroke information, thereby generating the scene semantic map without needing to explicitly label nodes or edges during correction. This simplifies the steps of generating the scene semantic map, shortens the time required, and ultimately improves the user's interactive experience. Furthermore, because the 3D point information set of objects in the target scene is corrected using the aforementioned user interaction stroke information, the accuracy of the generated scene semantic map can be improved by generating a scene semantic map based on the corrected 3D point information set. This, in turn, enhances the stability and generalization ability of the generated scene semantic map. It also improves the adaptability of the 3D point cloud, thereby increasing the accuracy and robustness of the generated scene semantic map for the corresponding edge device. Additionally, it simplifies the steps involved in generating the scene semantic map, reducing the time required and ultimately improving the user experience.
[0082] Further reference Figure 2 As an implementation of the methods shown in the above figures, this disclosure provides some embodiments of a scene semantic graph generation method based on mixed reality interaction. These device embodiments are similar to... Figure 1 Corresponding to the method embodiments shown, the device can be specifically applied to various electronic devices.
[0083] like Figure 2As shown, a scene semantic graph generation device 200 based on mixed reality interaction in some embodiments includes: an acquisition unit 201, a generation unit 202, and an execution unit 203. The acquisition unit 201 is configured to acquire scene point cloud and user interaction stroke information corresponding to a target scene; the generation unit 202 is configured to generate intent type information corresponding to the user interaction stroke information based on the user interaction stroke information using a first sub-model included in a pre-trained scene semantic graph generation model; the execution unit 203 is configured to perform the following steps using a second sub-model included in the pre-trained scene semantic graph generation model: generating a set of object 3D point information corresponding to the scene point cloud based on the scene point cloud; and generating a set of object 3D point information corresponding to the user interaction stroke information based on the intent type information. Alignment feature information; Each feature similarity information that meets a preset threshold condition in the feature similarity information set between the aforementioned user interaction stroke information and the aforementioned object 3D point information set is determined as a feature similarity information group; Based on the aforementioned alignment feature information, the aforementioned feature similarity information group, the first preset injection information, the second preset injection information, and each object 3D point in the aforementioned object 3D point information set associated with the aforementioned user interaction stroke information, the aforementioned object 3D point information set is corrected to obtain a modified object 3D point information set; Based on the aforementioned modified object 3D point information set, a scene semantic map corresponding to the aforementioned target scene is generated.
[0084] It is understandable that the units and references described in the mixed reality interaction-based scene semantic graph generation device 200 are... Figure 1 The steps in the described method correspond to each other. Therefore, the operations, features, and beneficial effects described above for the method also apply to the positioning information generation device 200 and the units contained therein, and will not be repeated here.
[0085] The following is for reference. Figure 3 It shows a schematic diagram of the structure of an electronic device 300 (e.g., a computing device) suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.
[0086] like Figure 3As shown, the electronic device 300 may include a processing unit 301 (e.g., a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. The RAM 303 also stores various programs and data required for the operation of the electronic device 300. The processing unit 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0087] Typically, the following devices can be connected to I / O interface 305: input devices 306 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, mixed reality controllers, etc.; output devices 307 including, for example, liquid crystal displays (LCDs), speakers, vibrators, head-mounted mixed reality displays, etc.; storage devices 308 including, for example, magnetic tapes, hard disks, etc.; and communication devices 309. Communication device 309 allows electronic device 300 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 3 An electronic device 300 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 3 Each box shown can represent a device or multiple devices as needed.
[0088] In particular, according to some embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 309, or installed from storage device 308, or installed from ROM 302. When the computer program is executed by processing device 301, it performs the functions defined in the methods of some embodiments of this disclosure.
[0089] It should be noted that, in some embodiments of this disclosure, the computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0090] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0091] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: acquire scene point cloud and user interaction stroke information corresponding to the target scene; generate intent type information corresponding to the aforementioned user interaction stroke information based on the aforementioned user interaction stroke information using a first sub-model included in a pre-trained scene semantic graph generation model; and perform the following steps using a second sub-model included in the pre-trained scene semantic graph generation model: generate a set of object 3D point information corresponding to the aforementioned scene point cloud based on the aforementioned scene point cloud; and generate information corresponding to the aforementioned user interaction stroke information based on the aforementioned intent type information. Alignment feature information of stroke information; determine each feature similarity information that meets the preset threshold condition in the feature similarity information set between the above user interaction stroke information and the above object 3D point information set as a feature similarity information group; according to the above alignment feature information, the above feature similarity information group, the first preset injection information, the second preset injection information, and each object 3D point in the above object 3D point information set associated with the above user interaction stroke information, perform correction processing on the above object 3D point information set to obtain a modified object 3D point information set; generate a scene semantic map corresponding to the above target scene based on the modified object 3D point information set.
[0092] Computer program code for performing operations of some embodiments of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0093] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0094] The units described in some embodiments of this disclosure can be implemented in software or hardware. The described units can also be housed in a processor; for example, a processor may be described as including an acquisition unit, a generation unit, and an execution unit. The names of these units do not necessarily limit the unit itself; for example, an acquisition unit may also be described as "a unit that acquires scene point clouds and user interaction stroke information corresponding to a target scene."
[0095] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.
[0096] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.
Claims
1. A method for generating scene semantic graphs based on mixed reality interaction, comprising: Obtain the scene point cloud and user interaction stroke information for the corresponding target scene; The first sub-model of the pre-trained scene semantic graph generation model generates intent type information corresponding to the user interaction stroke information based on the user interaction stroke information. The second sub-model, which is generated from a pre-trained scene semantic graph, is performed by executing the following steps: Based on the scene point cloud, generate a set of three-dimensional point information of objects corresponding to the scene point cloud; Based on the intent type information, generate alignment feature information corresponding to the user interaction stroke information; Each feature similarity information that meets a preset threshold condition in the feature similarity information set between the user interaction stroke information and the object's three-dimensional point information set is determined as a feature similarity information group; Based on the alignment feature information, the feature similarity information group, the first preset injection information, the second preset injection information, and each object 3D point in the object 3D point information set that is associated with the user interaction stroke information, the object 3D point information set is corrected to obtain the changed object 3D point information set. Based on the set of 3D point information of the changed object, a scene semantic map corresponding to the target scene is generated.
2. The method according to claim 1, wherein, The method further includes: The node semantic label information and relation semantic label information corresponding to the scene semantic graph are labeled to the three-dimensional space corresponding to the target scene.
3. The method according to claim 1, wherein, The first sub-model of the pre-trained scene semantic graph generation model generates intent type information corresponding to the user interaction stroke information based on the user interaction stroke information, including: Based on the user interaction stroke information, generate the corresponding temporal dynamic features and spatial structural features of the user interaction stroke information; Based on the time dynamic features and the spatial structure features, intent type information corresponding to the user interaction stroke information is generated.
4. The method according to claim 3, wherein, The step of generating temporal dynamic features and spatial structural features corresponding to the user interaction stroke information based on the user interaction stroke information includes: Based on the user interaction stroke information, determine the state feature sequence information corresponding to the user interaction stroke information; The product of the preset weight matrix and the preset weight vector is used to determine the weight matrix; The weight matrix is normalized to obtain the normalized weight matrix; The normalized weight matrix is updated to obtain the modified weight matrix; Based on the state feature sequence information and the change weight matrix, determine the time dynamic features corresponding to the user interaction stroke information; The user interaction stroke information is subjected to feature extraction processing to obtain interactive stroke feature information; Based on the interactive stroke feature information, determine the spatial structure features corresponding to the user interactive stroke information.
5. The method according to claim 3, wherein, The step of generating intent type information corresponding to the user interaction stroke information based on the temporal dynamic features and the spatial structural features includes: The time dynamic features and the spatial structure features are concatenated to obtain the interactive stroke feature vector corresponding to the user interaction stroke information. The interactive stroke feature vector is classified to obtain interactive stroke matrix information corresponding to the user interactive stroke information, wherein the interactive stroke matrix information includes various types of numerical values. Based on the interactive stroke matrix information, intent type information corresponding to the user interactive stroke information is generated.
6. The method according to claim 5, wherein, The step of generating intent type information corresponding to the user interaction stroke information based on the interaction stroke matrix information includes: The interactive stroke matrix information is normalized to obtain normalized interactive stroke matrix information. The interactive stroke values that meet the preset selection conditions in the normalized interactive stroke matrix information are determined as the target interactive stroke values. The intent type information corresponding to the target interactive stroke value is determined as the intent type information corresponding to the user interactive stroke information.
7. A scene semantic graph generation device based on mixed reality interaction, comprising: The acquisition unit is configured to acquire scene point cloud and user interaction stroke information of the corresponding target scene; The generation unit is configured to generate the first sub-model of the pre-trained scene semantic graph generation model, and generate intent type information corresponding to the user interaction stroke information based on the user interaction stroke information. The execution unit is configured to generate a second sub-model including the pre-trained scene semantic map and perform the following steps: generate a set of object 3D point information corresponding to the scene point cloud based on the scene point cloud; generate alignment feature information corresponding to the user interaction stroke information based on the intent type information. Each feature similarity information that meets a preset threshold condition in the feature similarity information set between the user interaction stroke information and the object's three-dimensional point information set is determined as a feature similarity information group; Based on the alignment feature information, the feature similarity information group, the first preset injection information, the second preset injection information, and the object 3D point information set associated with the user interaction stroke information, the object 3D point information set is corrected to obtain a modified object 3D point information set; based on the modified object 3D point information set, a scene semantic map corresponding to the target scene is generated.
8. An electronic device, comprising: One or more processors; A storage device on which one or more programs are stored; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1 to 6.
9. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 6.