A synchronous mixed reality method with both immersion and interactive performance
By using the interleaving algorithm and instance segmentation model in synchronous mixed reality technology, the target information is selectively expressed, and the problem of difficulty in taking into account the immersion and interaction efficiency in the virtual and real fusion scenario is solved, and efficient virtual and real space fusion is achieved.
Patent Information
- Application Number
- CN202510309354.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2045-03-17
AI Technical Summary
The existing synchronous mixed reality technology is difficult to have both environmental immersion and system interaction efficiency in virtual and real fusion scenarios, resulting in limited application.
The target information is selectively expressed using the interleaving algorithm, and the visual information and pose information of physical instances are extracted through the instance segmentation model, and direct clue drawing and featured indirect clue drawing are performed in the virtual environment to realize the rendering and interaction of virtual and real fusion scenes.
It improves the target interaction accuracy, expands the rendering freedom of virtual and real fusion scenes, realizes the intelligent fusion of virtual and real spaces, taking into account environmental immersion and system interaction efficiency.
Smart Images

Figure CN119828897B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of virtual-real fusion and deep learning, and more specifically, to a synchronous hybrid reality method with both immersion and interaction efficiency. Background Art
[0002] Synchronized Mixed Reality (SMR) technology refers to a virtual-real fusion technology that uses an immersive head-mounted display device to construct a "synchronous space", matches virtual and real objects as "synchronous interaction objects" with common attributes through video perspective, and performs real-time expression of the synchronous interaction objects in the synchronous space to achieve "object fusion" and "environment synchronization".
[0003] The "synchronous space" refers to a unified space that has both the characteristics of a virtual space and the dimensions of a physical space. This space is constructed based on an immersive display device, and information is superimposed using video perspective. According to different information superimposition ratios, it can achieve intelligent changes from a pure virtual environment to a pure physical environment. The "synchronous interaction object" refers to an interaction object existing in the synchronous space, which is jointly composed of a physical entity in the physical space and a virtual agent in the virtual space. The physical entity has the functional attributes of the object itself (for example, a keyboard has an input function), and the virtual agent has both the visual presentation attribute of the physical object (the visual presentation of the keyboard) and the virtual information superimposition attribute (graphical drawing and functional redefinition of keyboard keys). The combination of the physical entity and the virtual agent makes the synchronous interaction object have both the characteristics of function penetration and information superimposition and function expansion, thereby expanding the functions of the virtual-real fusion system and reducing the construction difficulty of complex systems. "Object fusion" refers to extracting information from a physical target and then matching the physical entity and the virtual agent with a certain spatial transformation relationship to achieve their spatio-temporal synchronization. "Environment synchronization" refers to achieving real-time synchronization of the physical environment and the virtual environment through the conversion of the virtual space coordinate system and the physical space coordinate system, and at the same time achieving spatial synchronization of the characteristics (information such as position, attitude, deformation, temperature, etc.) of the synchronous interaction object through visual cue drawing or superimposition.
[0004] Traditional virtual-real fusion technologies include augmented virtuality technology with physical objects as the fusion object, such as Figure 1 shown; and augmented reality technology with virtual objects as the fusion object, such as Figure 2 shown; compared with traditional technologies, synchronous mixed reality technology performs virtual-real fusion with space as the object, not only including the real-time fusion of virtual objects and physical objects, but also including the feature fusion of physical space and virtual space, aiming to achieve Figure 3The "intelligent synchronization", "space synchronization", and "time synchronization" among the user, physical space, and virtual space as shown. "Intelligent synchronization" emphasizes the two-way synchronization mechanism between physical objects and virtual objects; different from the existing fusion methods with a constant virtual-real ratio, the synchronous mixed reality technology constructs a new immersive environment with equal virtual-real relationships in a unified space according to user needs, and can intelligently adjust the fusion ratio of virtual information and physical information, thereby realizing intelligent virtual-real fusion from a pure physical space to a pure virtual space. "Space synchronization" emphasizes the consistency between the virtual space and the physical space in terms of terrain and the shape structure of objects, that is, in the new immersive environment, there is a virtual proxy for each entity object in the physical space, and the two maintain spatial consistency so that what the user sees is what they get. "Time synchronization" emphasizes the system's perception and manipulation ability of the dynamic changes in the environment, requiring the system to monitor the changes in the visual information and pose information of physical objects in real time and synchronize them to the corresponding virtual objects.
[0005] The synchronous mixed reality technology can extract and fuse information in any environment, can meet different system functions, improves the freedom of virtual-real fusion while reducing the cost of virtual-real fusion, and has high application value. However, the current synchronous mixed reality technology mostly stays in the theoretical stage. Due to the inability to combine the environmental immersion and system interaction efficiency when rendering virtual-real fusion scenes, its development encounters bottlenecks and it is difficult to be widely applied. Summary of the Invention
[0006] To overcome the above-mentioned defects of the existing technology that it is impossible to combine the environmental immersion and system interaction efficiency when rendering virtual-real fusion scenes, the present invention provides a synchronous mixed reality method with both immersion and interaction efficiency, which realizes the selective expression of target information through the intersection-over-union algorithm, improves the target interaction accuracy on the basis of ensuring environmental immersion, expands the rendering freedom of virtual-real fusion scenes, and realizes the intelligent fusion of virtual and real spaces.
[0007] To solve the above technical problems, the technical solution of the present invention is as follows:
[0008] A synchronous mixed reality method with both immersion and interaction efficiency, comprising the following steps:
[0009] S1: Use a camera to obtain a physical scene image, and use a pre-trained instance segmentation model to extract the visual information, pose information, and segmentation bounding boxes of several physical instances from the physical scene image;
[0010] The physical instances include: active movable objects and passive movable objects;
[0011] S2: Load the virtual environment, perform direct clue drawing on the visual information of active movable objects, and perform feature-based indirect clue drawing on the visual information of passive movable objects; obtain virtual objects corresponding one by one to each physical instance;
[0012] S3: Obtain the pose information of each virtual object in the virtual environment, and perform virtual-real coordinate mapping with the pose information of the corresponding physical instance respectively to complete the rendering of the virtual-real fusion scene;
[0013] S4: Based on the rendered virtual-real fusion scene, calculate the intersection over union between the segmentation bounding box of the active movable object and the segmentation bounding box of each passive movable object, predict the user's interaction target according to the calculation result, and classify all passive movable objects into interaction targets to be interacted with and non-interaction targets;
[0014] S5: Redraw the visual information of the interaction target to be interacted with by means of direct clue drawing, obtain the secondary-drawn virtual object corresponding to the interaction target to be interacted with and perform visual expression; for the virtual objects corresponding to the active movable object and the non-interaction target, directly perform visual expression;
[0015] S6: Repeat steps S4 to S6 to complete the synchronous mixed reality of selective expression of target information.
[0016] Preferably, in the step S1, the pre-trained instance segmentation model is specifically: the YOLACT instance segmentation model;
[0017] The YOLACT instance segmentation model includes: the target detection network RetinaNet, and the mask coefficient prediction network and the fully convolutional neural network FCN set in parallel; the output of the target detection network RetinaNet is respectively connected to the inputs of the mask coefficient prediction network and the fully convolutional neural network FCN;
[0018] The output of the mask coefficient prediction network is processed by the non-maximum suppression algorithm NMS and then combined with the output of the fully convolutional neural network FCN to obtain the output of the YOLACT instance segmentation model.
[0019] Preferably, in the YOLACT instance segmentation model, the target detection network RetinaNet includes the RestNet-101 residual sub-network and the feature pyramid sub-network FPN connected in sequence;
[0020] The mask coefficient prediction network includes the classification sub-network, the box regression sub-network and the mask coefficient prediction sub-network set in parallel;
[0021] The fully convolutional neural network FCN is specifically the Protonet prototype mask generation sub-network;
[0022] The non-maximum suppression algorithm NMS is specifically the Fast NMS algorithm.
[0023] Preferably, after the output of the mask coefficient prediction network is processed by the non-maximum suppression algorithm NMS, a coefficient matrix is obtained, where n is the number of physical instances calculated by the non-maximum suppression algorithm NMS, and k is the mask coefficient of each physical instance;
[0024] The output of the fully convolutional neural network FCN is a prototype mask matrix , where h and w are the first and second dimensions of the prototype mask respectively, is the number of prototype masks for each physical instance;
[0025] The coefficient matrix and the prototype mask matrix are combined and calculated according to the following formula to obtain the output of the YOLACT instance segmentation model :
[0026]
[0027] where is the Sigmoid function.
[0028] Preferably, the loss function of the YOLACT instance segmentation model during the pre-training process includes: a classification loss function , a bounding box regression loss function and a mask coefficient prediction loss function ;
[0029] The classification loss function is specifically the Softmax cross-entropy loss function;
[0030] The bounding box regression loss function includes: a confidence loss and a position loss , the confidence loss is specifically the Softmax cross-entropy loss function, and the position loss is specifically the Smooth L1 loss function;
[0031] The mask coefficient prediction loss function is specifically the binary cross-entropy loss function.
[0032] Preferably, step S1 further includes: setting a category hash table, and several types of physical instance categories that can be fused into the virtual environment are stored in the category hash table;
[0033] Judge whether each of the extracted physical instances belongs to the physical instance category in the category hash table. If it belongs, directly execute step S2; if not, delete the visual information, pose information, and segmentation bounding box of the physical instance, and execute step S2.
[0034] Preferably, in steps S2 and S5, the direct clue drawing includes: drawing of the physical instance body and drawing of background information;
[0035] The drawing of the physical instance body is specifically: according to the visual information of the physical instance, draw the contour, texture, and color of the physical instance pixel by pixel in the virtual space, and obtain the corresponding virtual object;
[0036] The drawing of the background information is specifically: taking the pixels other than the pixels of the physical instance body as background pixels, setting the pixel values of the background pixels to 0, and completing the drawing of the background information.
[0037] Preferably, in step S2, the feature-based indirect clue drawing includes:
[0038] Obtain the size and the position coordinates of the center point of the physical instance according to the visual information of the physical instance, and perform feature replacement on the visual information of the physical instance by using a preset virtual feature map, and obtain the corresponding virtual object;
[0039] The size and the position coordinates of the center point of the preset virtual feature map are the same as those of the physical instance.
[0040] Preferably, in step S4, calculate the intersection over union between the segmentation bounding box A of the active movable object and the segmentation bounding box B of each passive movable object according to the following formula :
[0041]
[0042] where, is the side length of the intersection of the segmentation bounding box A and the segmentation bounding box B on the x axis; is the side length of the intersection of the segmentation bounding box A and the segmentation bounding box B on the y axis; is the coordinate of the lower left vertex of the segmentation bounding box A; is the coordinate of the upper right vertex of the segmentation bounding box A; is the coordinate of the lower left vertex of the segmentation bounding box B; is the coordinate of the upper right vertex of the segmentation bounding box B.
[0043] Preferably, in step S4, predicting the user's interaction target according to the calculation result includes:
[0044] Determine whether the intersection over union (IoU) between the segmentation bounding box of the active movable object and the segmentation bounding box of each passive movable object is greater than or equal to a preset threshold. If it is greater than or equal to the threshold, the category of the passive movable object is the target to be interacted with; if it is less than the threshold, the category of the passive movable object is a non-interaction target.
[0045] Compared with the prior art, the beneficial effects of the technical solution of the present invention are as follows:
[0046] The present invention provides a synchronous hybrid reality method with both immersion and interaction efficiency. First, a camera is used to obtain a physical scene image, and a pre-trained instance segmentation model is used to extract the visual information, pose information, and segmentation bounding boxes of several physical instances from the physical scene image; the physical instances include: active movable objects and passive movable objects; a virtual environment is loaded, and for the visual information of the active movable objects, direct cue rendering is performed; for the visual information of the passive movable objects, feature-based indirect cue rendering is performed; virtual objects corresponding to each physical instance are obtained; the pose information of each virtual object in the virtual environment is obtained and respectively mapped to the pose information of the corresponding physical instance for virtual-real coordinate mapping to complete the rendering of the virtual-real fusion scene; based on the rendered virtual-real fusion scene, the IoU between the segmentation bounding box of the active movable object and the segmentation bounding box of each passive movable object is calculated, and the user's interaction target is predicted according to the calculation result, and all passive movable objects are classified into targets to be interacted with and non-interaction targets; the visual information of the targets to be interacted with is redrawn by means of direct cue rendering, and the secondary-rendered virtual objects corresponding to the targets to be interacted with are obtained and visually expressed; for the virtual objects corresponding to the active movable objects and non-interaction targets, direct visual expression is performed; the above steps are repeated to complete the synchronous hybrid reality of selective expression of target information;
[0047] The present invention uses a deep learning instance segmentation algorithm to extract the information of physical instances in real time, predicts the user's interaction purpose based on the IoU algorithm after virtual-real fusion, and then realizes the selective expression of physical target information in the virtual environment, improves the target interaction accuracy on the basis of ensuring the environmental immersion, can balance the environmental performance and the system interaction efficiency, and at the same time expands the rendering freedom of the virtual-real fusion scene to realize the intelligent fusion of virtual and real spaces. Description of the Drawings
[0048] Figure 1 It is a schematic diagram of augmented virtual technology provided in the background art.
[0049] Figure 2 It is a schematic diagram of augmented reality technology provided in the background art.
[0050] Figure 3 Schematic diagram of the synchronous mixed reality technology provided in the background art.
[0051] Figure 4 Flowchart of a synchronous mixed reality method with both immersion and interaction efficiency provided in Embodiment 1.
[0052] Figure 5 Implementation architecture diagram of the synchronous mixed reality technology provided in Embodiment 1.
[0053] Figure 6 Structure diagram of the YOLACT instance segmentation model provided in Embodiment 2.
[0054] Figure 7 Structure diagram of the RetinaNet object detection network provided in Embodiment 2.
[0055] Figure 8 Structure diagram of the residual bottleneck block in the RestNet-101 residual subnetwork provided in Embodiment 2.
[0056] Figure 9 Structure diagram of the Protonet prototype mask generation subnetwork provided in Embodiment 2.
[0057] Figure 10 Structure diagram of the mask coefficient prediction network provided in Embodiment 2.
[0058] Figure 11 Schematic diagram of the intelligent extraction of target information after setting up the category hash table provided in Embodiment 2.
[0059] Figure 12 Schematic diagram of the direct clue drawing of the virtual-real fusion scene provided in Embodiment 2.
[0060] Figure 13 Schematic diagram of the indirect clue drawing of the virtual-real fusion scene features provided in Embodiment 2.
[0061] Figure 14 Schematic diagram of the graphical clue drawing of different target object information provided in Embodiment 2.
[0062] Figure 15 Schematic diagram of the side length representation of the intersection part of the two bounding boxes provided in Embodiment 2.
[0063] Figure 16 Schematic diagram of the selective expression of target information based on the anticipation of interaction purposes provided in Embodiment 2.
[0064] Figure 17 Schematic diagram of using the synchronous mixed reality hardware system for scene drawing provided in Embodiment 3. Detailed implementation manners
[0065] The accompanying drawings are only for illustrative purposes and should not be construed as limiting the present application;
[0066] To better illustrate this embodiment, some components in the accompanying drawings will be omitted, enlarged or reduced, and do not represent the dimensions of the actual product;
[0067] For those skilled in the art, it is understandable that some well-known structures and their descriptions in the accompanying drawings may be omitted.
[0068] The technical solutions of the present invention will be further described below with reference to the accompanying drawings and embodiments.
[0069] Embodiment 1
[0070] As Figure 4 shown, this embodiment provides a synchronous hybrid reality method with both immersion and interaction efficiency, including the following steps:
[0071] S1: Use a camera to obtain a physical scene image, and use a pre-trained instance segmentation model to extract the visual information, pose information, and segmentation bounding boxes of several physical instances from the physical scene image;
[0072] The physical instances include: active movable objects and passive movable objects;
[0073] S2: Load the virtual environment. For the visual information of the active movable objects, perform direct cue drawing; for the visual information of the passive movable objects, perform feature-based indirect cue drawing; obtain virtual objects corresponding to each physical instance;
[0074] S3: Obtain the pose information of each virtual object in the virtual environment, and perform virtual-real coordinate mapping with the pose information of the corresponding physical instance respectively to complete the rendering of the virtual-real fusion scene;
[0075] S4: Based on the rendered virtual-real fusion scene, calculate the intersection-over-union ratio between the segmentation bounding box of the active movable object and the segmentation bounding boxes of each passive movable object, predict the user's interaction target according to the calculation result, and classify all passive movable objects into interaction targets to be interacted with and non-interaction targets;
[0076] S5: Redraw the visual information of the interaction targets to be interacted with in the way of direct cue drawing, obtain the secondary-drawn virtual objects corresponding to the interaction targets to be interacted with and perform visual expression; for the virtual objects corresponding to the active movable objects and non-interaction targets, directly perform visual expression;
[0077] S6: Repeat steps S4 to S6 to complete the synchronous hybrid reality of selective expression of target information.
[0078] In the specific implementation process, such as Figure 5 shown in the implementation architecture diagram of synchronous mixed reality technology, which mainly includes four key steps: physical target information extraction, target feature drawing, virtual-real mapping, and target selective spatio-temporal synchronization. Among them, the target information extraction process extracts the visual information and pose information of the target in the physical scene through an image segmentation method; the target feature drawing realizes the visual feature synchronization of virtual and real objects by extracting the feature information of virtual objects and assigning them to physical targets; the virtual-real mapping transforms the physical space coordinate system and the virtual space coordinate system to realize target pose tracking; the target selective spatio-temporal synchronization selectively expresses specific target information in the virtual environment according to the user's needs;
[0079] In the physical target information extraction based on instance segmentation, first, the physical space information is captured by an RGB camera, and the category, bounding box, and mask prototype information of the physical instance are calculated using a deep learning instance segmentation algorithm, and its visual information and pose information are obtained;
[0080] To achieve the seamless integration of physical information in the virtual environment, it is necessary to classify the physical objects in the information extraction space before constructing the virtual-real fusion environment. According to the motion characteristics, interaction paradigms, and existence states of physical objects, the target objects in the physical space can be divided into Active Movable Objects (AMO), Passive Movable Objects (PMO), and Immovable Objects (IMO), and the priorities of the three types of information decrease in turn;
[0081] The first level is the active movable objects represented by people, animals, robots, etc. Such objects have subjective initiative, and their motion space and activity trajectories are determined by the moving objects themselves. They have the characteristics of high degrees of freedom of motion, large activity ranges, and uncontrollable existence states. In addition, such objects have obvious personalized characteristics, and the personalized characteristics have an important impact on interaction. In the virtual-real fusion scenario, the most typical representatives of such objects are the users themselves. They are both the objects of information extraction and the issuers and executors of interaction behaviors, with the highest interaction weight and information presentation level. In addition to users, the objects at this level also include environmental participants, robots that can move freely, pets, and other physical objects with active motion capabilities. The interaction weight and existence state of such objects in the physical space are uncontrollable. Therefore, to improve the interaction efficiency of the virtual-real fusion system and reduce environmental collisions, such objects have the highest information extraction priority when extracting information;
[0082] The second level is passive movable objects represented by interactive tools, such as objects with specific functions but uncertain motion states like keyboards, mobile phones, cups, etc.; such objects have the characteristics that their motion trajectories and activity ranges are controllable, their appearance and disappearance states in the physical space are stable, although they have personalized characteristics, information such as visual texture, color, geometric shape, etc. has little impact on interaction, so real-time rendering is not required, but real-time tracking is required;
[0083] The third level is immovable objects represented by walls and tables, etc. The spatial positions of such objects are generally fixed, with a lower interaction weight, and precise expression is not required. However, to avoid collisions, the movable space of the user needs to be defined according to the third-level objects;
[0084] In this embodiment, the extracted physical instances only retain active movable objects and passive movable objects. For the extracted immovable objects, their visual information will not be rendered subsequently;
[0085] During the target feature rendering process, first, a physical target label library needs to be constructed according to the target category. Secondly, a feature extraction method is used to calculate the virtual environment features and endow the virtual feature information to the physical target labels to achieve the feature presentation of the label library. Finally, the feature labels are used to replace the physical information to achieve the unity of the visual features of the physical target and the virtual environment visual features, ensuring the immersion of the virtual-real fusion environment;
[0086] During the virtual-real mapping process, the camera coordinate system of the virtual space is converted into the camera coordinate system of the physical space through spatial transformation to achieve the unity of the positions, directions, and distances of the virtual space and the physical space;
[0087] During the selective spatio-temporal synchronization of target information, first, according to the target depth occlusion relationship, the intersection over union (IoU) algorithm of the bounding box is used to predict the user's interaction purpose. Secondly, according to the user's interaction purpose, the per-pixel rendering method is used to directly express the visual information of the interaction target, and the feature icon rendering method is used to indirectly express the visual information of non-interaction objects. Finally, the selective expression of the target information in the physical space in the virtual environment is realized, improving the system interaction efficiency while ensuring the immersion of the virtual-real fusion environment.
[0088] Embodiment 2
[0089] This embodiment provides a synchronous mixed reality method with both immersion and interaction efficiency, including the following steps:
[0090] S1: Use a camera to obtain a physical scene image, and use a pre-trained instance segmentation model to extract the visual information, pose information, and segmentation bounding boxes of several physical instances from the physical scene image;
[0091] The physical instances include: active movable objects and passive movable objects;
[0092] S2: Load the virtual environment, perform direct clue drawing for the visual information of active movable objects, and perform feature-based indirect clue drawing for the visual information of passive movable objects; obtain virtual objects corresponding one by one to each physical instance;
[0093] S3: Obtain the pose information of each virtual object in the virtual environment, and perform virtual-real coordinate mapping with the pose information of the corresponding physical instance respectively to complete the rendering of the virtual-real fusion scene;
[0094] S4: Based on the rendered virtual-real fusion scene, calculate the intersection over union between the segmentation bounding box of the active movable object and the segmentation bounding box of each passive movable object, predict the user's interaction target according to the calculation result, and classify all passive movable objects into interaction targets to be interacted and non-interaction targets;
[0095] S5: Redraw the visual information of the interaction target to be interacted by means of direct clue drawing, obtain the secondary-drawn virtual object corresponding to the interaction target to be interacted and perform visual expression; for the virtual objects corresponding to the active movable object and the non-interaction target, directly perform visual expression;
[0096] S6: Repeat steps S4 - S6 to complete the synchronous mixed reality of selective expression of target information;
[0097] In the said step S1, the pre-trained instance segmentation model is specifically: YOLACT instance segmentation model;
[0098] The said YOLACT instance segmentation model includes: the target detection network RetinaNet, and the mask coefficient prediction network and the fully convolutional neural network FCN set in parallel; the output of the target detection network RetinaNet is respectively connected to the inputs of the mask coefficient prediction network and the fully convolutional neural network FCN;
[0099] After the output of the said mask coefficient prediction network is processed by the non-maximum suppression algorithm NMS, it is combined with the output of the said fully convolutional neural network FCN to obtain the output of the YOLACT instance segmentation model;
[0100] In the said YOLACT instance segmentation model, the target detection network RetinaNet includes the RestNet-101 residual sub-network and the feature pyramid sub-network FPN connected in sequence;
[0101] The said mask coefficient prediction network includes the classification sub-network, the box regression sub-network and the mask coefficient prediction sub-network set side by side;
[0102] The said fully convolutional neural network FCN is specifically the Protonet prototype mask generation sub-network;
[0103] The non-maximum suppression algorithm NMS is specifically the Fast NMS algorithm;
[0104] After the output of the mask coefficient prediction network is processed by the non-maximum suppression algorithm NMS, a coefficient matrix is obtained , where n is the number of physical instances calculated by the non-maximum suppression algorithm NMS, and k is the mask coefficient of each physical instance;
[0105] The output of the fully convolutional neural network FCN is a prototype mask matrix , where h and w are the first and second dimensions of the prototype mask respectively, is the number of prototype masks for each physical instance;
[0106] The coefficient matrix and the prototype mask matrix are combined and calculated according to the following formula to obtain the output of the YOLACT instance segmentation model :
[0107]
[0108] where is the Sigmoid function;
[0109] The loss function of the YOLACT instance segmentation model during the pre-training process includes: a classification loss function , a bounding box regression loss function and a mask coefficient prediction loss function ;
[0110] The classification loss function is specifically the Softmax cross-entropy loss function;
[0111] The bounding box regression loss function includes: a confidence loss and a location loss , the confidence loss is specifically the Softmax cross-entropy loss function, and the location loss is specifically the Smooth L1 loss function;
[0112] The mask coefficient prediction loss function is specifically the binary cross-entropy loss function;
[0113] Step S1 further includes: setting a category hash table, and several categories of physical instances that can be integrated into the virtual environment are stored in the category hash table;
[0114] Judge whether each of the extracted physical instances belongs to the physical instance category in the category hash table. If it belongs, directly execute step S2; if not, delete the visual information, pose information, and segmentation bounding box of the physical instance, and execute step S2;
[0115] In the above-mentioned step S2 and step S5, the direct clue drawing includes: physical instance ontology drawing and background information drawing;
[0116] The physical instance ontology drawing is specifically: according to the visual information of the physical instance, draw the contour, texture, and color of the physical instance pixel by pixel in the virtual space, and obtain the corresponding virtual object;
[0117] The background information drawing is specifically: take the pixels other than the physical instance ontology pixels as background pixels, and set the pixel values of the background pixels to 0 to complete the background information drawing;
[0118] In the above-mentioned step S2, the feature-based indirect clue drawing includes:
[0119] Obtain the size and the center point position coordinates of the physical instance according to the visual information of the physical instance, and use a preset virtual feature map to perform feature replacement on the visual information of the physical instance, and obtain the corresponding virtual object;
[0120] The size and the center point position coordinates of the preset virtual feature map are the same as those of the physical instance;
[0121] In the above-mentioned step S4, calculate the intersection over union between the segmentation bounding box A of the active movable object and the segmentation bounding box B of each passive movable object according to the following formula :
[0122]
[0123] where, is the side length of the intersection of the segmentation bounding box A and the segmentation bounding box B on the x axis; is the side length of the intersection of the segmentation bounding box A and the segmentation bounding box B on the y axis; is the lower left vertex coordinate of the segmentation bounding box A; is the upper right vertex coordinate of the segmentation bounding box A; is the lower left vertex coordinate of the segmentation bounding box B; is the upper right vertex coordinate of the segmentation bounding box B;
[0124] In the above-mentioned step S4, predicting the user's interaction target according to the calculation result includes:
[0125] Determine whether the intersection over union (IoU) between the segmentation bounding box of the active movable object and the segmentation bounding box of each passive movable object is greater than or equal to a preset threshold. If it is greater than or equal to the threshold, the category of the passive movable object is the target to be interacted with; if it is less than the threshold, the category of the passive movable object is a non-interaction target.
[0126] In the specific implementation process, the realization of the synchronous mixed reality technology mainly includes four key steps: physical target information extraction, target feature mapping, virtual-real mapping, and target selective spatio-temporal synchronization. Among them, the target information extraction process extracts the visual information and pose information of the target in the physical scene through an image segmentation method; the target feature mapping realizes the visual feature synchronization of virtual and real objects by extracting the feature information of virtual objects and assigning them to physical targets; the virtual-real mapping transforms the physical space coordinate system and the virtual space coordinate system to achieve target pose tracking; the target selective spatio-temporal synchronization selectively expresses specific target information in the virtual environment according to the user's needs.
[0127] In the physical target information extraction based on instance segmentation, first, the physical space information is captured by an RGB camera, and the category, bounding box, and mask prototype information of the physical instance are calculated using a deep learning instance segmentation algorithm, and its visual information and pose information are obtained. Second, the classification information is judged through a classification hash table, and the physical target information is selectively extracted.
[0128] In this embodiment, only the active movable objects and passive movable objects are retained among the extracted physical instances. For the extracted immovable objects, their pixel values are directly set to 0.
[0129] The segmentation algorithm in this embodiment is specifically the YOLACT instance segmentation model. The YOLACT model is a "first-order" instance segmentation model, which improves the algorithm speed by abandoning the feature localization step and adopting a two-way parallel method. The YOLACT network structure is as Figure 6 shown. When an image is input into the YOLACT network, first, the image is feature-extracted by a target detection network (RetinaNet) to obtain a set of feature layer images. Second, the feature layer images are processed in a two-way parallel manner. One branch obtains the mask, bounding box, and category output of the target through a mask coefficient prediction network and a non-maximum suppression algorithm (NMS). The other branch predicts the target prototype through a fully convolutional neural network (FCN) to obtain several prototype masks. Finally, the mask information of each instance is obtained by linearly combining the mask prototype and mask coefficient of each instance, and the output result of the image instance segmentation is obtained through bounding box cropping and thresholding.
[0130] The YOLACT network uses a first-order RetinaNet object detection network to extract the feature information of target objects in the physical space. This network consists of a residual sub-network (RestNet-101) and a feature pyramid sub-network (Feature Pyramid Network, FPN). Its network framework is as Figure 7 shown;
[0131] The YOLACT network selects a residual bottleneck block as Figure 8 shown to construct a convolutional network with a depth level of 101 layers, so as to extract more complex image features; during the feature extraction process, the layers with unchanged feature map size during the convolution process of the RestNet-101 network are used as a stage, and the output result of the last layer of each stage is extracted as the feature pyramid layer. The outputs of the conv2, conv3, conv4, and conv5 stages are extracted, corresponding to the C2, C3, C4, and C5 layers of the feature pyramid. The convolutional structure of each layer is shown in Table 1;
[0132] Table 1 Convolutional layer structures of the RestNet-101 residual sub-network
[0133]
[0134] The feature pyramid sub-network FPN consists of three parts: a bottom-up convolutional neural network, a top-down feature sampling, and a lateral connection. Among them, the bottom-up convolutional network is the RestNet-101 network, which can effectively extract the coarse-grained features of the image, and as the number of layers of the convolutional network increases, the high-level semantic information of the target becomes more abundant; however, as the convolution deepens, the spatial resolution of the image decreases and more spatial information is lost. The top-down sampling process uses an interpolation method to extract feature information and obtains a feature map with weak high-level semantic information but rich spatial information; thereafter, the feature map with high semantic information and low spatial information obtained by bottom-up sampling is laterally connected with the feature map with low semantic information and high spatial information obtained by top-down sampling, and a 3×3 convolutional kernel is used to convolve the connection result, so as to eliminate the aliasing effect of upsampling and finally obtain a group of feature maps with both high spatial resolution and strong semantic information.
[0135] In the RetinaNet object detection network, first, a set of feature maps {C1, C2, C3, C4, C5} is obtained through bottom-up sampling. To reduce the computational complexity and improve the algorithm speed, the C1 feature layer with too low sampling stride is discarded. Secondly, according to the Feature Pyramid Network, the {C2, C3, C4, C5} feature layers are iteratively processed layer by layer to obtain a feature pyramid {P2, P3, P4, P5} that has both high spatial resolution and strong semantic information. After that, to further improve the algorithm speed, the P2 feature layer with too high resolution is discarded. At the same time, to improve the algorithm accuracy, a strided convolution with a stride of 2 is performed on the P5 feature layer to obtain the P6 feature layer, and the striding operation is continued on the P6 feature layer to obtain the P7 feature layer and construct a feature pyramid containing the {P3, P4, P5, P6, P7} feature maps. In this feature pyramid, the deeper the feature layer, the higher the resolution of the spatial information, and the higher the feature layer, the richer the semantic information. Finally, the target feature information is extracted by operating on the feature layers {P3, P4, P5, P6, P7};
[0136] The goal of the prototype generation network is to obtain a prototype mask with as high a resolution as possible and as accurate target information as possible. Since the P3 layer in the feature pyramid generated by the RetinaNet object detection network has the highest resolution information, the P3 layer is used to generate the prototype mask; the prototype mask generation network is called the Protonet network, which is constructed based on the Fully Convolutional Network (FCN) and activated by the ReLU function; the structure of the Protonet network is as Figure 9 shown. In the process of obtaining the prototype template, first, the input P3 feature map is subjected to 4 identical convolution operations using a 3×3 convolutional kernel; secondly, a 2×2 transposed convolutional kernel is used to perform a transposed convolution operation on the convolution-processed image to expand the image size and improve the accuracy of small target segmentation; finally, a 1×1 convolution operation is performed on the upsampled image to obtain an image with K channels, and each channel of the image can be regarded as a prototype mask;
[0137] The mask coefficient prediction network is used to predict the mask coefficient, and an anchor detection model is mostly used for coefficient prediction. In the mask prediction process, first, corresponding anchors need to be generated for each feature layer; in the YOLCAT algorithm, translation-invariant anchors are used to traverse the feature map to extract feature information; as Figure 10 shown, it includes a classification sub-network, a box regression sub-network, and a mask coefficient prediction sub-network arranged in parallel;
[0138] In the pre-training stage of the network, the corresponding loss function of the classification sub-network is: classification loss function , which is calculated using the Softmax cross-entropy loss:
[0139]
[0140]
[0141] Among them, is the true label, is the output value of the i-th node of the classification sub-network, is the output value and the probability value in C categories;
[0142] The bounding box regression loss function of the bounding box regression sub-network includes: confidence loss and location loss , and the two are weighted and added:
[0143]
[0144] Among them, x is the matching degree between the A-th prior box and the B-th bounding box in the p-th category; c is the predicted value of the class confidence; l is the predicted value of the prior box position; g is the bounding box position parameter; N is the number of prior box samples matched to the bounding box; i j ; is the proportionality coefficient for adjusting the ratio between and , and is generally set to 1;
[0145] Confidence loss adopts the Softmax cross-entropy loss function,
[0146]
[0147]
[0148] Among them, represents the positive sample set, represents the negative sample set, p represents the p-th category, represents the probability of the output value of the i-th node in the p-th category;
[0149] Location loss is calculated using the Smooth L1 loss:
[0150]
[0151]
[0152] Among them, is the predicted value of the i-th prior box predicting the corresponding bounding box; The encoding value for matching the i-th prior box with the j-th ground truth bounding box; In the function Indicates the difference between the ground truth and the predicted value;
[0153] The mask coefficient prediction sub-network processes the feature map pixel by pixel, predicting a mask coefficient for each pixel i in the feature map. The mask coefficient prediction loss function Is the binary cross-entropy loss function:
[0154]
[0155]
[0156] The non-maximum suppression algorithm NMS in this embodiment is specifically the Fast NMS algorithm, which can reduce the impact of the confidence score sorting on the algorithm running speed. It does not require sorting all the confidence output results. Only n results with the highest scores need to be extracted and the IoU coefficient of this result is calculated;
[0157] In the YOLCAT instance segmentation network, the feature layer extracted by the RetinaNet object detection network is divided into two branches for parallel processing to improve the algorithm running speed; the prototype mask network processes the P3 feature layer through the Protonet network to obtain the prototype mask; the coefficient prediction network predicts the mask category, bounding box, and mask coefficient through the prediction head network, and uses the non-maximum suppression algorithm to obtain the local optimal solution of the mask coefficient; thereafter, a simple linear combination of the prototype mask and the mask coefficient can obtain the final mask information; the mask information synthesis operation can be regarded as a combined operation of basic matrix multiplication and the Sigmoid function, and the operation formula is as follows:
[0158] The coefficient matrix And the prototype mask matrix Perform a combined operation according to the following formula to obtain the output of the YOLACT instance segmentation model :
[0159]
[0160] Among them, Is the Sigmoid function;
[0161] When constructing the virtual-real fusion environment, first, the physical space information shown in (a) in Figure 11 needs to be captured by an RGB camera; secondly, the YOLCAT model is used to perform real-time segmentation on the physical space information to obtain the physical object category, bounding box, and mask prototype information shown in (b) in Figure 11 ; finally, the information extraction method is used to obtain as shown in Figure 11The target visual feature information shown in (c); however, in real application scenarios, the physical environment contains a lot of useless information. Therefore, before performing virtual-real fusion, it is necessary to process the spatial information segmented by the YOLCAT network to remove redundant information; in this embodiment, a classification hash table is constructed to intelligently extract the target information;
[0162] According to the system functions and user requirements, store the categories of target objects that need to be fused into the virtual environment in the hash table, and extract specific target information through category gating judgment; in the actual application process, the categories in the list can be changed according to different requirements of the system, so as to change the fused target information and realize the intelligent extraction of target information;
[0163] The information extraction method combining the YOLCAT instance segmentation network and the classification hash table can selectively extract the visual information of the target objects in the physical space and fuse it into the virtual environment to achieve virtual-real fusion. However, since most physical environments and virtual environments are quite different, directly fusing the extracted information into the virtual environment will destroy the environmental immersion; therefore, in order to ensure the visual consistency of virtual information and physical information and maintain the immersion of the virtual-real fusion environment, target feature rendering is required;
[0164] During the target feature rendering process, first, a physical target label library needs to be constructed according to the target category. Secondly, a feature extraction method is used to calculate the virtual environment features and endow the virtual feature information to the physical target labels to achieve the feature rendering of the label library. Finally, the feature labels are used to replace the physical information to achieve the unity of the visual features of the physical target and the virtual environment, and ensure the immersion of the virtual-real fusion environment;
[0165] The movement range and activity trajectory of active movable objects are uncontrollable, but their personalized features are obvious and the personalized visual information is of great significance to interaction. Therefore, in the virtual environment, a direct image presentation method is used for information expression; as Figure 12 shown, the direct clue rendering method performs per-pixel rendering of the virtual agent according to the features such as the contour, texture, and color of the physical object; the physical information rendering includes not only the physical target rendering but also the background information rendering. In the direct clue rendering scenario, the per-pixel rendering method is used to render the visual feature information of the physical target, and the redundant information is eliminated by setting the pixel values of the background or irrelevant information to 0, so as to realize the real-time rendering of the target information in the virtual environment;
[0166] The direct clue rendering method can present the personalized feature information of the target, can react to the changes of the target visual features in real time, and helps to improve the interaction accuracy of the system. However, limited by information extraction, the accuracy of the information rendered by this method is slightly lower, and when the visual features of the physical target are too different from the virtual environment, this method will destroy the immersion of the virtual-real fusion environment;
[0167] Passive movable objects have a controllable movement trajectory and a relatively large number, but their personalized visual features do not have specific significance. Therefore, in a virtual environment, they can be replaced by characterized virtual substitutes to maintain the immersion of the environment while ensuring the interaction efficiency. Indirect cue drawing replaces the target visual information according to the size of the target object and the position coordinates of the center point, and realizes the expansion of the target function or the visualization unity of virtual and real information while retaining some features of the physical object. The indirect cue drawing process is as follows Figure 13 As shown, during the information drawing process, instead of drawing the target information pixel by pixel, a pre-constructed virtual proxy is used to indirectly express the physical object information. The features such as texture, color, size, and shape of the virtual proxy may be different from those of the physical object. This drawing method increases the consistency of virtual and real information and has a relatively low impact on the immersion of the virtual-real fusion environment. However, due to the loss of the target visual information, the interaction accuracy of the system is also relatively low.
[0168] Immovable objects are fixed in position and have a low interaction frequency. After matching the virtual space and the physical space, they can be not drawn.
[0169] As Figure 14 shown, this is a schematic diagram of the graphical cue drawing of different target object information in this embodiment. For the active movable object (the user's hand), it is directly drawn, and for the passive movable object (the cup), it is indirectly drawn.
[0170] During the virtual-real mapping process, the camera coordinate system in the virtual space is converted to the camera coordinate system in the physical space through space transformation to achieve the unity of position, direction, and distance in the virtual space and the physical space, and complete the rendering of the virtual-real fusion scene.
[0171] For the virtual-real fusion scene, the ideal state is that when there is no need to interact with the object, the object information is presented in the form of characterized icons to ensure the immersion of the virtual environment; when it is necessary to interact with such an object, the drawing mode of the target information switches back to direct drawing to achieve the accurate expression of physical information. To achieve this goal, this embodiment uses the intersection over union (IoU) algorithm to predict the user's interaction purpose and presents the visual information of the physical target through a selective expression method.
[0172] As Figure 15 shown, the intersection over union between the segmentation bounding box A of the active movable object and the segmentation bounding box B of each passive movable object is calculated according to the following formula :
[0173]
[0174] where is the side length of the intersection of the segmentation bounding box A and the segmentation bounding box B on the x axis. is the side length of the intersection of the segmentation bounding box A and the segmentation bounding box B on the y axis; is the coordinate of the lower left vertex of the segmentation bounding box A; is the coordinate of the upper right vertex of the segmentation bounding box A; is the coordinate of the lower left vertex of the segmentation bounding box B; is the coordinate of the upper right vertex of the segmentation bounding box B; The larger the value, the higher the degree of overlap;
[0175] As Figure 16 shown, in the process of target information selective spatio-temporal synchronization, first, according to the target depth occlusion relationship, the intersection over union (IoU) algorithm is used to predict the user's interaction purpose. When the user does not need to interact with the object, the method of drawing feature-based icons is used to express physical information to ensure the immersion of the virtual-real fusion environment; when the user needs to interact with the target, the interaction object is converted from a feature-based icon to a physical icon, while other non-interaction objects still maintain the feature-based icon presentation, so as to balance the conflict between the immersion of the virtual environment and the interaction efficiency;
[0176] After the physical scene information is segmented by YOLCAT, in addition to obtaining the instance and instance classification labels, the instance bounding box information is also obtained; in the process of information selective expression, first, based on the bounding box information output by the instance segmentation network, the above formula is used to calculate the intersection over union between instances and ; secondly, a threshold T is set, and the calculated is compared with the threshold T. If ≥T, it indicates that the two bounding boxes intersect, and it is predicted that this instance is the user's interaction target; thereafter, the direct cue drawing method is used to express the target visual information, and the feature-based icon drawing method is used to indirectly express the non-interaction target visual information to achieve the selective expression of the target;
[0177] In the process of calculating the intersection over union, the first depth level object is used as the base bounding box, and the intersection over union between the bounding box of the first depth level object and the bounding box of the second depth level object is calculated; in Figure 16 , the first depth level object is the user's hand information. By calculating the intersection over union between the user's hand bounding box and the bounding boxes of objects such as cups, mice, and keyboards, different scene drawing methods are selected; in the figure, the user's hand is on the keyboard, and the value of the keyboard bounding box and the hand bounding box is greater than the threshold T. At this time, the keyboard is presented in real time; while the water cup bounding box and the hand bounding box do not overlap or the overlapping part is too small, and the between the two is less than the threshold, then the feature-based icon drawing method is used; this method of information selective expression based on interaction purpose prediction can improve the system interaction efficiency while maintaining the immersion of the virtual environment.
[0178] Example 3
[0179] This embodiment provides a synchronous mixed reality hardware system based on the synchronous mixed reality method with both immersion and interaction efficiency described in Embodiment 1 or 2.
[0180] In the specific implementation process, the synchronous mixed reality hardware system performs virtual-real fusion in units of space, which can be divided into small-range specific space virtual-real fusion and large-range movable space virtual-real fusion according to the fusion space range; the two types of space virtual-real fusion have different requirements for the information capture field of view. When performing small-range and close-range physical object fusion, this type of space generally includes immovable objects and passive movable objects, with lower requirements for the field of view but higher requirements for information extraction resolution, so a high-resolution monocular camera is used for construction; when performing large-range and long-distance physical object fusion, this type of space includes three types of interactive characteristic objects, with higher requirements for the field of view, so a binocular camera is used for construction;
[0181] The two system structures are similar, both including an information capture and calculation unit, an information calculation and rendering unit, and an information visualization presentation unit. Among them, the information capture unit completes the real-time capture of physical space information, mainly a binocular or monocular camera; the information calculation and rendering unit is used to complete the real-time segmentation, extraction, and rendering of space information using the method in Embodiment 1 or 2, and is processed by a desktop computer; the information visualization presentation unit is used to present the real-time information of the virtual-real fusion scene, mainly a head-mounted display device;
[0182] As Figure 17 shown, it is a schematic diagram of using the synchronous mixed reality hardware system for scene drawing. As Figure 17 shown in (a) below, the physical environment contains a table and multiple ceramic cups with different textures and colors. The hardware used by the user is a monocular camera and an HTC VIVE head-mounted display device; the virtual environment is as Figure 17 shown in (b) below, and a table corresponding to the physical space is drawn to place the extracted cup models; Figure 17 Shown in (c) below represents the synchronous mixed reality scene drawn by direct cues, Figure 17In figure (d), it shows the characterization of the image rendering scenario and the selective presentation effect of the target information; it can be seen that the system can successfully achieve the selective expression of the target information; in addition, after experimental verification using this system, it shows that the physical information selective expression method proposed in Embodiment 1 or 2 can effectively solve the conflict between the environmental presence and the system interaction efficiency in the virtual-real fusion scenario rendering, can improve the virtual environment immersion by 15.3%, increase the interaction speed by 14.6% and reduce the number of collisions by 34.3%. On the basis of ensuring the environmental immersion, it improves the target interaction accuracy, can balance the environmental performance and the system interaction efficiency, and at the same time expands the rendering freedom of the virtual-real fusion scenario, realizing the intelligent fusion of the virtual and real spaces.
[0183] The same or similar reference numerals correspond to the same or similar components;
[0184] The terms used to describe the positional relationship in the drawings are only for illustrative purposes and should not be construed as a limitation of this application;
[0185] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention, rather than limitations on the implementation manners of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made on the basis of the above description. It is not necessary and impossible to enumerate all the implementation manners here. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included in the protection scope of the claims of the present invention.
Claims
1. A synchronous mixed reality method with both immersion and interactive performance, characterized in that: The following steps are involved: S1: Use a camera to acquire a physical scene image, and use a pre-trained instance segmentation model to extract visual information, pose information, and segmentation bounding boxes of several physical instances from the physical scene image; The physical instances include: active movable objects and passive movable objects; S2: Load the virtual environment, draw direct clues for the visual information of the active movable object, draw characteristic indirect clues for the visual information of the passive movable object, and obtain the virtual object corresponding to each physical instance; The direct clue drawing includes: physical instance ontology drawing and background information drawing; The physical instance ontology drawing specifically includes: drawing the outline, texture and color of the physical instance pixel by pixel in the virtual space according to the visual information of the physical instance, and obtaining the corresponding virtual object; The background information drawing is specifically as follows: taking pixels other than the pixels of the physical instance as background pixels, setting the pixel values of the background pixels to 0, and completing the background information drawing; The characterized indirect clue drawing includes: The size of the physical instance and the coordinates of the center point position are obtained according to the visual information of the physical instance, and the visual information of the physical instance is replaced with a preset virtual feature map to obtain a corresponding virtual object; The size and center point position coordinates of the preset virtual feature map are the same as those of the physical instance; S3: Obtain the position and posture information of each virtual object in the virtual environment, and map the virtual and real coordinates with the position and posture information of the corresponding physical instance to complete the rendering of the virtual and real fusion scene; S4: Based on the rendered virtual-reality fusion scene, calculate the intersection-over-union ratio between the segmentation bounding box of the active movable object and the segmentation bounding box of each passive movable object, predict the user's interaction target according to the calculation result, and classify all passive movable objects into targets to be interacted and non-interaction targets; S5: redrawing the visual information of the target to be interacted with by direct clue drawing, obtaining the secondary drawn virtual object corresponding to the target to be interacted with and performing visual expression; for the virtual objects corresponding to the active movable objects and non-interactive targets, performing visual expression directly; S6: Repeat steps S4 to S5 to complete the synchronous mixed reality of selectively expressing the target information.
2. A synchronous mixed reality method with both immersion and interactive performance according to claim 1, characterized in that: In step S1, the pre-trained instance segmentation model is specifically: a YOLACT instance segmentation model; The YOLACT instance segmentation model includes: a target detection network RetinaNet, and a mask coefficient prediction network and a fully convolutional neural network FCN set in parallel; the output of the target detection network RetinaNet is connected to the input of the mask coefficient prediction network and the fully convolutional neural network FCN respectively; The output of the mask coefficient prediction network is processed by the non-maximum suppression algorithm NMS, and then combined with the output of the fully convolutional neural network FCN to obtain the output of the YOLACT instance segmentation model.
3. A synchronous mixed reality method with both immersion and interactive performance according to claim 2, characterized in that: In the YOLACT instance segmentation model, the target detection network RetinaNet includes a RestNet-101 residual subnetwork and a feature pyramid subnetwork FPN connected in sequence; The mask coefficient prediction network includes a classification subnetwork, a frame regression subnetwork and a mask coefficient prediction subnetwork arranged in parallel; The fully convolutional neural network FCN is specifically a Protonet prototype mask generation subnetwork; The non-maximum suppression algorithm NMS is specifically a Fast NMS algorithm.
4. A synchronous mixed reality method with both immersion and interactive performance according to claim 2, characterized in that: The output of the mask coefficient prediction network is processed by the non-maximum suppression algorithm NMS to obtain The coefficient matrix of , where n is the number of physical instances calculated by the non-maximum suppression algorithm NMS, and k is the mask coefficient of each physical instance; The output of the fully convolutional neural network FCN is The prototype mask matrix , where h and w are the first and second sizes of the prototype mask, respectively. The number of prototype masks for each physical instance; The coefficient matrix and the prototype mask matrix Perform a combined operation according to the following formula to obtain the output of the YOLACT instance segmentation model : in, is the Sigmoid function.
5. A synchronous mixed reality method with both immersion and interactive performance according to claim 2, characterized in that: The loss functions of the YOLACT instance segmentation model in the pre-training process include: classification loss function , box regression loss function And the mask coefficient prediction loss function ; The classification loss function Specifically, it is the Softmax cross entropy loss function; The box regression loss function Includes: Confidence loss and position loss , the confidence loss Specifically, it is the Softmax cross entropy loss function, the position loss Specifically, it is the Smooth L1 loss function; The mask coefficient prediction loss function Specifically, it is a binary cross entropy loss function.
6. A synchronous mixed reality method with both immersion and interactive performance according to claim 1, characterized in that: The step S1 further includes: setting a category hash table, wherein the category hash table stores several categories of physical instances that can be integrated into the virtual environment; Determine whether each extracted physical instance belongs to the physical instance category in the category hash table. If so, directly execute step S2; if not, delete the visual information, posture information and segmentation bounding box of the physical instance and execute step S2.
7. A synchronous mixed reality method with both immersion and interactive performance according to claim 1, characterized in that: In step S4, the intersection-and-union ratio between the segmentation bounding box A of the active movable object and the segmentation bounding box B of each passive movable object is calculated according to the following formula: : in, The intersection of segmentation bounding box A and segmentation bounding box B is x The length of the side on the axis; The intersection of segmentation bounding box A and segmentation bounding box B is y The length of the side on the axis; is the coordinate of the lower left corner vertex of the segmentation bounding box A; is the coordinate of the upper right corner vertex of the segmentation bounding box A; is the coordinate of the lower left corner vertex of the segmentation bounding box B; is the coordinate of the upper right corner vertex of the segmentation bounding box B.
8. The synchronous mixed reality method with both immersion and interactive performance according to claim 1, characterized in that: In step S4, predicting the user's interaction target according to the calculation result includes: It is determined whether the intersection-and-union ratio between the segmentation bounding box of the active movable object and the segmentation bounding box of each passive movable object is greater than or equal to a preset threshold. If so, the category of the passive movable object is a target to be interacted with; if less than, the category of the passive movable object is a non-interactive target.
Citation Information
Patent Citations
Non-rigid object virtual and real shielding method and system based on convolutional neural network
CN117475117A
Multi-source virtual-real fusion office system construction method oriented to long-time immersion
CN117809003A