A method for constructing a visual and audio scene in an offline mode based on online master-slave communication
By collecting and reconstructing information about real locations and stakeholders in online interactions, a digital human model is generated, which solves the problem of lack of participation in online interactions, enables offline perception experiences for stakeholders, and improves communication efficiency.
Patent Information
- Application Number
- CN202510357638.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2026-02-24
- Estimated Expiration
- 2045-03-25
AI Technical Summary
Existing online master-slave communication methods cannot provide an experience comparable to offline participation. The master cannot perceive the slave's participation status and body movements in real time, and the slave cannot perceive the master's gaze and movement, which affects the effectiveness of communication.
By collecting information on real-world activity locations and the master and slave entities, and using mixed reality technology to perform 3D reconstruction on a central server to generate a master and slave digital human model, and then displaying the virtual scene on mixed reality glasses and terminal devices, the patented new technology has been applied to the field of virtual reality technology, specifically involving a method for constructing an offline sensory and audiovisual scene for online master-slave interaction.
Both the subject and the follower can have the same participatory and perceptual experience as offline. The subject can perceive the follower's state by turning their head and moving around, and the follower can perceive the subject's gaze and movement, thus improving the efficiency and effectiveness of communication.
Smart Images

Figure CN120318470B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of virtual reality technology, specifically relating to a method for constructing an offline sensory audiovisual scene for online master-slave interaction. Background Technology
[0002] Human activities are inseparable from mutual exchange and communication. In teaching and training, communication activities involve a subject and a follower, known as master-slave communication. In master-slave communication, the subject conveys content to the follower through explanation, demonstration, and questioning, while the follower receives the content through observation and listening. Unlike non-master-slave communication, master-slave communication typically occurs between the subject and the follower. During master-slave communication, the subject perceives the follower's participation by observing or approaching them, thus adjusting the pace, method, and content of the communication accordingly, and even drawing the follower's attention. Conversely, the follower perceives (or anticipates) whether they are being watched or noticed by the subject based on eye contact or proximity. Online communication is currently a common method for master-slave communication, offering advantages over offline methods such as freedom from time and space constraints and lower costs. However, existing online methods often rely on online meeting software (such as Tencent Meeting), which does not provide the same level of engagement as offline communication. Due to the physical limitations of display screens, existing online methods cannot clearly display videos (containing facial expressions) of participants along with the event content on the same screen. This forces participants to scroll through screens to perceive the participant's engagement beforehand, which undoubtedly interrupts their train of thought, affects their enthusiasm for understanding the participant's engagement, and ultimately impacts the effectiveness of the communication activity. Furthermore, in existing online methods, participants cannot perceive non-verbal cues such as eye contact or physical actions like approaching from the event location, making it difficult to transition from a potentially detached state to a focused one. Summary of the Invention
[0003] The purpose of this invention is to address the shortcomings of existing technologies and propose a method for constructing offline audiovisual scenes with online master-slave communication.
[0004] To achieve the above objectives, the present invention is implemented through the following technical solution: a method for constructing an offline sensory audiovisual scene with online master-slave communication, wherein the offline sensory audiovisual scene construction includes the following entities: a real activity venue, a master, mixed reality glasses, several slave entities, several slave terminal entities, and a central server; the method includes the following steps;
[0005] Step S1: Location information acquisition: The location information includes location view image information and location point cloud information. The location point cloud information includes the global location point cloud and the initial location multi-view point cloud.
[0006] Step S2: Master-slave information collection: The master-slave information includes the position and head posture of the master in the actual activity location, the direction of the master's eye gaze, the facial information of the slave, and the voice information of the master and slave;
[0007] The collected location information and master-slave information are transmitted to the central server via the network. The central server then completes the 3D reconstruction of the master and slave in the map coordinate system, the generation of the audiovisual scene from the master's perspective, and the generation of the audiovisual scene from the slave's perspective.
[0008] Step S3: 3D reconstruction of master and slave bodies in map coordinate system: including the construction of slave digital human in world coordinate system and the construction of master and slave digital human in map coordinate system. The coordinate system of the global point cloud of the location is used as the world coordinate system, and the coordinate system of the multi-view point cloud of the initial location is used as the map coordinate system.
[0009] Step S3.1, Constructing a volumetric digital human in the world coordinate system
[0010] Step S3.1.1: Extract the occupant from the global point cloud of the venue through human interaction. The occupant is a three-dimensional space with length, width and height that is assumed to be occupied by the subject sitting in the real activity venue.
[0011] Step S3.1.2: Based on the spatial range of the occupier and its position in the world coordinate system, establish standard digital human models for all slave bodies respectively;
[0012] Step S3.1.3: Personalize the head model of the standard digital human according to the data types supported by the slave terminal and the collected slave facial information;
[0013] Step S3.2, Construction of master-slave digital human in map coordinate system
[0014] Step S3.2.1: Establish a mathematical model of the transformation relationship between the world coordinate system and the map coordinate system based on the matching relationship between the initial location multi-view point cloud and the location global point cloud;
[0015] Step S3.2.2: Based on the obtained mathematical model, transform the human body digital human model in the world coordinate system to the human body digital human model in the map coordinate system;
[0016] Step S3.2.3: Establish a digital human model of the subject based on the subject's characteristics;
[0017] Step S4: Generating the main perspective audiovisual scene
[0018] The subject's position and head posture in the real activity venue are used to project perspective images of the slave digital human model to the left and right eyes respectively, generating two slave RGB images. The two generated slave RGB images are output to the left and right eye displays of the mixed reality glasses and superimposed on the scene of the real activity venue observed by the subject. In order to give the subject the feeling of being present in person, the slave's voice is corrected according to its distance from the subject in the map coordinate system, and then the corrected slave voice is output to the sound device of the mixed reality glasses.
[0019] Step S5: Generating an audiovisual scene from a volumetric perspective
[0020] Using the center of the eyes of the slave digital human in map coordinates as the viewpoint and the center of the head of the main digital human in map coordinates as the observation point, the slave's line of sight is established. Based on the viewpoint and line of sight, perspective projection is performed on the main digital human model in map coordinates and displayed on the slave terminal. In order to give the slave a sense of being present offline, virtual footsteps of the main body are set, and the virtual footsteps received by each slave are corrected according to the distance between the main and slave in map coordinates.
[0021] Preferably, the location perspective image information is an RGB image and a depth image of the real activity location acquired by the subject through mixed reality glasses; the location global point cloud is a 3D point cloud acquired by the subject through a 3D point cloud acquisition device scanning the entire real activity location before the activity begins; and the initial position location multi-view point cloud is a 3D point cloud of the real activity location acquired by the subject wearing mixed reality glasses, standing at a certain position in the real activity location, and rotating their body.
[0022] Preferably, the position and head posture of the subject in the real activity location are obtained by using the location perspective image information through the algorithm embedded in the mixed reality glasses; the direction of the subject's eye gaze is obtained by the mixed reality glasses; the facial information of the slave is obtained by the slave terminal; and the voice information of the subject and slave is obtained from the mixed reality glasses and the slave terminal, respectively.
[0023] Preferably, step S3.1.3, which involves personalizing the head model of the standard digital human based on the data types supported by the slave terminal and the collected slave facial information, specifically includes the following types:
[0024] (1) When the data type supported by the slave terminal is slave facial expression video stream, the array human model head is replaced with a cuboid, and the front or all faces of the cuboid are covered with the texture of the slave facial video.
[0025] (2) When the data type supported by the slave terminal is slave facial key feature points, the facial parameters of the digital human model are changed by the slave facial key feature points, and the face of the digital human model is textured using the slave's facial photo.
[0026] (3) When the data type supported by the slave terminal is the slave face 3D point cloud, a surface is constructed using the slave face 3D point cloud, and then the face photo texture of the slave is used to obtain the face of the digital human model.
[0027] Preferably, the mathematical model for establishing the transformation relationship between the world coordinate system and the map coordinate system based on the matching relationship between the initial location multi-view point cloud and the global point cloud of the location, as described in step S3.2.1, is as follows:
[0028] The mathematical model expression is:
[0029] x g =Ax v0 +ε
[0030] In the formula, x g The coordinates of a point in the world coordinate system, x v0 Let A be the coordinates of a point in the map coordinate system, A be the transformation matrix, and ε be the model transformation error.
[0031] The point set that matches the initial location location multi-view point cloud with the location global point cloud is extracted using a point cloud matching algorithm;
[0032] Substitute the coordinates of the matched point set into the mathematical model, and use the least squares method to calculate the coefficients in the mathematical model, thus obtaining the mathematical model of the transformation relationship between the world coordinate system and the map coordinate system.
[0033] Preferably, step S3.2.3, which involves establishing a digital human model of the subject based on the subject's features, includes the following steps: using a subject terminal device to collect the subject's facial geometric and texture features in advance, establishing a standard digital human model of the subject, and then combining the obtained facial photos of the subject to personalize the face of the standard digital human model.
[0034] Preferably, the establishment of the main standard digital human model specifically falls into the following types:
[0035] When the pre-collected facial geometric features are key feature points, these key feature points are used to change the facial parameters of the subject's standard digital human model, and at the same time, the subject's facial photo is used to perform texture mapping on the face of the subject's standard digital human model.
[0036] When the facial geometric features collected in advance are a 3D point cloud of the face, a surface is constructed using these point clouds, and then a main facial photograph is used for texture mapping to replace the face of the standard digital human model.
[0037] Preferably, in steps S4 and S5, the correction formula is:
[0038]
[0039] In the formula, represents the corrected volume, v represents the volume of the main body, and d represents the current straight-line distance between the slave and the main body.
[0040] Compared with existing methods, the subject and the follower in this invention can obtain a perceptual experience equivalent to offline participation, specifically including: 1) Similar to offline, the subject only needs to turn their head to visually perceive the follower's participation status or freely switch between perceiving the follower's participation status and viewing the communication content on the display screen; 2) Similar to offline, the subject only needs to walk into the follower to visually perceive the follower's participation status at close range, thereby guiding the follower to focus their attention in a timely manner; 3) Similar to offline, the subject only needs to keep their gaze on a follower for a period of time to make the follower perceive that the subject is paying attention to them, thereby guiding the follower to focus their attention; 4) Similar to offline, the subject can judge the approximate location of the speaker through spatial hearing, thereby quickly turning around and paying attention to the speaker; 5) Similar to offline, when the subject walks into the follower, the follower can perceive the subject's approach through sound or vision, thereby timely increasing their attention and entering a ready state in advance; 6) Similar to offline, the follower can visually perceive whether they are being watched, thereby timely increasing their attention and entering a ready state in advance. Attached Figure Description
[0041] Figure 1 This is a flowchart of a method for constructing an offline sensory audiovisual scene with online master-slave communication according to the present invention;
[0042] Figure 2 This is a schematic diagram illustrating the implementation framework of an online master-slave communication offline sensory audiovisual scene construction method of the present invention;
[0043] Figure 3 This is a schematic diagram of the seat occupant of the present invention;
[0044] Figure 4 This is a schematic diagram of the digital human construction method of the present invention;
[0045] Figure 5 This is a schematic diagram of the main visual scene generation of the present invention;
[0046] Figure 6 This is a schematic diagram of the generation of a volumetric visual scene according to the present invention. Detailed Implementation
[0047] The present invention will be further described below with reference to embodiments, but these embodiments are not intended to limit the scope of the invention.
[0048] Figure 2 This diagram illustrates the implementation framework of an online master-slave interactive offline audiovisual scene construction method. The method requires the following entities: a real activity venue 1, a master 2, a pair of mixed reality glasses 3, several slaves 4, several slave terminals 5 (including computers, mobile phones, etc.), and a central server 6. The real activity venue provides the master with a physical space for real-world movement and the necessary facilities for the activity, including a public display screen 7 and several tables and chairs 8 for the slaves to sit on. The mixed reality glasses are wearable glasses that integrate RGB image acquisition, depth image acquisition, eye gaze acquisition, binocular stereo display, voice pickup, and voice playback functions, and provide computing capabilities, such as Microsoft's HoloLens 2 glasses. Slave terminals are computing devices or systems with functions such as facial information collection, voice pickup, graphics display, and voice playback, such as computers with cameras, microphones, and speakers, smartphones, or computers with integrated Kinect motion-sensing devices. The central server is a high-performance computer connected to a network.
[0049] like Figure 1 As shown, the implementation of the method of the present invention consists of five steps: location information collection, master-slave information collection, master-slave 3D reconstruction in map coordinate system, generation of audiovisual scene from the perspective of the master, and generation of audiovisual scene from the perspective of the slave.
[0050] The location information collection includes the collection of location perspective image information and location point cloud information; the location perspective image is an RGB and depth image of the real activity location captured by the subject as they move around using mixed reality glasses; the mixed reality glasses are a pair of glasses that integrate RGB image acquisition, depth image acquisition, eye gaze acquisition, binocular stereo display, voice pickup and voice playback functions, and can be worn on the head and provide computing functions; the location point cloud information includes two types: global location point cloud and initial location multi-view point cloud; the global location point cloud is a 3D point cloud acquired by the subject before the activity begins by scanning the entire real activity location 1 using a 3D point cloud acquisition device (such as a portable 3D laser scanner); the real activity location is the physical space where the subject is located or walks during the communication activity and consists of the necessary facilities required for the activity; the necessary facilities required for the activity include a public display screen and several tables and chairs for the subject to sit on; the initial location multi-view point cloud is a 3D point cloud of the real activity location acquired by the subject after wearing and turning on the mixed reality glasses, standing at a certain position in the real activity location and rotating their body.
[0051] The acquisition of master-slave information includes the acquisition of the master's position and head posture in the real activity environment, the master's eye gaze direction, the slave's facial information, and the master-slave's voice information. The master's position and head posture in the real activity environment are obtained from the location view image through an algorithm embedded in the mixed reality glasses (e.g., calculated using the SLAM algorithm). The master's eye gaze direction is acquired by an eye tracker embedded in the mixed reality glasses. The slave's facial information consists of facial expression videos, key facial feature points, or 3D facial point clouds acquired through the slave terminal. The master's voice information is acquired through the mixed reality glasses, and the slave's voice information is acquired through a voice pickup device on the slave terminal. The slave terminal is a computing device or system with functions such as facial information acquisition, voice pickup, graphic display, and voice playback.
[0052] The collected location information and master / slave information are transmitted to the central server via the network. The central server then completes three steps: master / slave 3D reconstruction in the map coordinate system, generation of audiovisual scenes from the master's perspective, and generation of audiovisual scenes from the slave's perspective.
[0053] The 3D reconstruction of the master-slave body in the map coordinate system includes two steps: the construction of the slave body's personalized digital human in the world coordinate system and the construction of the master-slave body's digital human in the map coordinate system. The world coordinate system is the coordinate system in which the global point cloud of the location is located; the map coordinate system is the coordinate system in which the multi-view point cloud of the initial location is located.
[0054] The method for constructing a personalized digital human in the world coordinate system is as follows:
[0055] 1) Extract the occupant volume 9 in the world coordinate system and construct an occupant volume database; the occupant volume 9 is a three-dimensional space with length, width, and height assumed to be occupied by the subject sitting in a real activity space (e.g., Figure 3 As shown), it was extracted from the global point cloud of the venue through manual interaction; after extracting the occupant 9, a number was assigned to each occupant 9 so that it could be used to arrange seats for the occupants in the future; the numerous occupants and their numbers constitute the occupant database;
[0056] 2) Based on certain seating rules, select a placeholder in the placeholder database for each newly added follower. Then, based on the spatial range of the placeholder and its position in the world coordinate system, create a standard digital human model for each follower, starting with head model 10 (e.g., ...). Figure 4 ) and torso model 11 (such as Figure 4 Composed of [various elements]; certain seating arrangement rules allow for either random seating or seating according to the order of joining;
[0057] 3) Based on the data types supported by the slave terminal and the collected slave facial information, the head model of the standard digital human model is personalized. The specific method is as follows:
[0058] ① When the supported type is a facial expression video stream, the head of the standard digital human model is represented by a cuboid 12 (e.g., Figure 4 Instead, the front face (the same plane as the standard digital human model) or all faces of cuboid 12 are textured with the texture 13 from the body face video (such as...). Figure 4 );
[0059] ② When the supported type is facial key feature points, then these key feature points 14 (such as...) are used. Figure 4 This is used to change the facial parameters of the standard digital human model, while simultaneously using a facial photograph of the subject to texture map the face of the standard digital human model;
[0060] ③ When the supported type is a 3D point cloud of a face, a surface 15 (e.g., ...) is constructed in real time using these point clouds. Figure 4 Then, a facial photograph from the body is used to texture the face and replace the face of the standard digital human model.
[0061] The method for constructing a master-slave digital human in the map coordinate system is as follows:
[0062] 1) Establish the transformation relationship between the world coordinate system and the map coordinate system. The mathematical model is as follows:
[0063] x g =Ax v0 +ε (1)
[0064] In the formula, x g The coordinates of a point in the world coordinate system, x v0 Let A be the coordinates of a point in the map coordinate system, A be the transformation matrix, and ε be the model transformation error.
[0065] 2) Use point cloud matching algorithms to extract point sets that match the multi-view point clouds of the initial location with the global point cloud of the location;
[0066] 3) Substitute the coordinates of the matching point set into the above mathematical model, and then use the least squares method to calculate the transformation matrix A of equation (1);
[0067] 4) Using the transformation relationship of equation (1), the personalized digital human model in the world coordinate system is transformed to the map coordinate system (e.g., ...). Figure 2 (Issued No. 16);
[0068] 5) Based on the characteristics of the subject, establish a digital human model that closely resembles the subject (refer to the construction method of the subject's digital human model);
[0069] Specifically, a standard digital human model of the subject is established based on the geometric features (key facial feature points or three-dimensional point clouds of the face) and texture features (facial photos) of the subject's face collected in advance using devices such as similar physical terminals.
[0070] When the pre-collected facial geometric features are key feature points, these key feature points are used to change the facial parameters of the subject's standard digital human model, and at the same time, the subject's facial photo is used to perform texture mapping on the face of the subject's standard digital human model.
[0071] When the facial geometric features collected in advance are a 3D point cloud of the face, a surface is constructed using these point clouds, and then the main face photo is used for texture mapping to replace the face of the standard digital human model.
[0072] The subject wears mixed reality glasses, which calculate the subject's position in the real activity area in real time based on the collected location view image information and location point cloud information. Combined with the subject's position in the map coordinate system, the constructed digital human model of the subject is moved to the corresponding position; the orientation of the subject's digital human model's head is adjusted according to the collected subject's head posture data.
[0073] The method for generating the subject's perspective audiovisual scene is as follows:
[0074] 1) Based on the current position and head posture of the subject, perform left and right eye perspective projections on the personalized digital human model to generate two RGB images of the subject. Figure 5 (a) One of the volumetric RGB images is given; the two generated volumetric RGB images are output to the left and right eye displays of the mixed reality glasses respectively and superimposed on the scene of the real activity venue observed by the subject to generate a visual image of the subject participating offline. Figure 5 (b) Images of the actual activity locations observed by the subject; Figure 5 (c) is the image obtained by superimposing the RGB image of the subject with the image of the real event location, that is, the visual image that the subject finally obtains through the mixed reality glasses (due to the limitation of the two-dimensional graphic display, the original three-dimensional visual image is displayed as a two-dimensional visual image in this figure).
[0075] 2) Treat the sound of each slave as a point sound source, then correct the sound intensity of each slave according to the principle of sound attenuation over distance and the distance between the slave and the subject. Then output the corrected slave sound to the sound device of the subject's mixed reality glasses to generate an auditory experience for the subject to participate in the offline experience.
[0076]
[0077] In the formula, represents the corrected volume, v represents the volume of the main body, and d represents the current straight-line distance between the slave and the main body.
[0078] The method for generating audiovisual scenes from a volumetric perspective is as follows:
[0079] 1) Using the center of the eyes of the personalized digital human in the map coordinate system as the viewpoint and the center of the head of the main digital human model in the map coordinate system as the observation point, establish the visual observation direction of the human body as 16 (e.g., Figure 6 Based on the viewpoint and the direction of the line of sight, a perspective projection is performed on the main digital human model in the map coordinate system; the projected main digital human model is then displayed in a certain area of the slave terminal, so that the slave can perceive the current orientation of the main body while viewing the current communication content. Figure 6 (a) The screen displayed on a certain slave terminal includes the current communication content 17 and the main digital human model 18 displayed on the public display screen of the real activity venue; if the subject's eyes fall within the range of a certain slave occupier for more than 3 seconds at a certain moment, then a pair of constantly flashing gazes 19 are emitted from the eyes of the main digital human model of the slave terminal (e.g. Figure 6 So that the body can perceive the subject's gaze. Figure 6 (b) The screen displayed on another slave terminal includes the main digital human model 20. At this time, the main body is looking at the other slave, so the main digital human model from the front view is displayed on the other slave terminal.
[0080] 2) When the subject walks in a real activity area, a virtual footstep sound of the subject walking is set, and then the virtual footstep sound received by each slave is improved according to the point source attenuation principle of equation (2); the size of the footstep sound is related to the distance between the subject and each slave; the closer the subject is to the current slave, the louder the virtual footstep sound heard by the current slave, and vice versa; in this way, the slave can indirectly perceive the approach of the subject through hearing, so as to make timely preparations; when the subject approaches into the alert range of a certain slave, the event is expressed on the terminal of the slave with the help of visual or auditory elements, so that the slave can perceive the subject next to it.
[0081] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. However, the above description is merely a specific embodiment of the present invention, and the technical features of the present invention are not limited thereto. Any other embodiments derived by those skilled in the art without departing from the technical solution of the present invention should be covered within the patent scope of the present invention.
Claims
1. A method for constructing an offline immersive audiovisual scene with online master-slave interaction, wherein the offline immersive audiovisual scene construction includes the following entities: a real activity venue, a master, mixed reality glasses, several slave entities, several slave entity terminals, and a central server; characterized in that, The method includes the following steps; Step S1: Location information acquisition: The location information includes location view image information and location point cloud information. The location point cloud information includes the global location point cloud and the initial location multi-view point cloud. Step S2: Master-slave information collection: The master-slave information includes the position and head posture of the master in the actual activity location, the direction of the master's eye gaze, the facial information of the slave, and the voice information of the master and slave; The collected location information and master-slave information are transmitted to the central server via the network. The central server then completes the 3D reconstruction of the master and slave in the map coordinate system, the generation of the audiovisual scene from the master's perspective, and the generation of the audiovisual scene from the slave's perspective. Step S3: 3D reconstruction of master and slave bodies in map coordinate system: including the construction of slave digital human in world coordinate system and the construction of master and slave digital human in map coordinate system. The coordinate system of the global point cloud of the location is used as the world coordinate system, and the coordinate system of the multi-view point cloud of the initial location is used as the map coordinate system. Step S3.1, Constructing a volumetric digital human in the world coordinate system Step S3.1.1: Extract the occupant from the global point cloud of the venue through human interaction. The occupant is a three-dimensional space with length, width and height that is assumed to be occupied by the subject sitting in the real activity venue. Step S3.1.2: Based on the spatial range of the occupier and its position in the world coordinate system, establish standard digital human models for all slave bodies respectively; Step S3.1.3: Personalize the head model of the standard digital human according to the data types supported by the slave terminal and the collected slave facial information; Step S3.2, Construction of master-slave digital human in map coordinate system Step S3.2.1: Establish a mathematical model of the transformation relationship between the world coordinate system and the map coordinate system based on the matching relationship between the initial location multi-view point cloud and the location global point cloud; Step S3.2.2: Based on the obtained mathematical model, transform the human body digital human model in the world coordinate system to the human body digital human model in the map coordinate system; Step S3.2.3: Establish a digital human model of the subject based on the subject's characteristics; Step S4: Generating the main perspective audiovisual scene The subject's position and head posture in the real activity venue are used to project perspective images of the slave digital human model to the left and right eyes respectively, generating two slave RGB images. The two generated slave RGB images are output to the left and right eye displays of the mixed reality glasses and superimposed on the scene of the real activity venue observed by the subject. In order to give the subject the feeling of being present in person, the slave's voice is corrected according to its distance from the subject in the map coordinate system, and then the corrected slave voice is output to the sound device of the mixed reality glasses. Step S5: Generating an audiovisual scene from a volumetric perspective Using the center of the eyes of the slave digital human in map coordinates as the viewpoint and the center of the head of the main digital human in map coordinates as the observation point, the slave's line of sight is established. Based on the viewpoint and line of sight, perspective projection is performed on the main digital human model in map coordinates and displayed on the slave terminal. In order to give the slave a sense of being present offline, virtual footsteps of the main body are set, and the virtual footsteps received by each slave are corrected according to the distance between the main and slave in map coordinates.
2. The method for constructing an offline sensory-audio-visual scene based on online master-slave communication according to claim 1, characterized in that, The location perspective image information consists of RGB and depth images of the real activity location captured by the subject through mixed reality glasses; the global point cloud of the location consists of 3D point clouds acquired by the subject through a 3D point cloud acquisition device before the activity begins; and the initial position location multi-view point cloud consists of 3D point clouds of the real activity location acquired by the subject wearing mixed reality glasses, standing at a certain position in the real activity location, and rotating their body.
3. The method for constructing an offline sensory-audio-visual scene based on online master-slave communication according to claim 1, characterized in that, The position and head posture of the subject in the real activity location are obtained by using the location perspective image information through the algorithm embedded in the mixed reality glasses; the direction of the subject's eye gaze is obtained by the mixed reality glasses; the facial information of the slave is obtained by the slave terminal; the voice information of the subject and slave is obtained by the mixed reality glasses and the slave terminal respectively.
4. The method for constructing an offline sensory-audio-visual scene based on online master-slave communication according to claim 1, characterized in that, Step S3.1.3 describes the personalized expression of the standard digital human's head model based on the data types supported by the slave terminal and the collected slave facial information. Specifically, it includes the following types: (1) When the data type supported by the slave terminal is slave facial expression video stream, the array human model head is replaced with a cuboid, and the front or all faces of the cuboid are covered with the texture of the slave facial video. (2) When the data type supported by the slave terminal is slave facial key feature points, the facial parameters of the digital human model are changed by the slave facial key feature points, and the face of the digital human model is textured using the slave's facial photo. (3) When the data type supported by the slave terminal is the slave face 3D point cloud, a surface is constructed using the slave face 3D point cloud, and then the face photo texture of the slave is used to obtain the face of the digital human model.
5. The method for constructing an offline sensory-audio-visual scene based on online master-slave communication according to claim 1, characterized in that, Step S3.2.1, which establishes the mathematical model for the transformation relationship between the world coordinate system and the map coordinate system based on the matching relationship between the initial location multi-view point cloud and the global point cloud of the location, specifically includes: The mathematical model expression is: x g =Ax v0 +ε In the formula, x g The coordinates of a point in the world coordinate system, x v0 Let A be the coordinates of a point in the map coordinate system, A be the transformation matrix, and ε be the model transformation error. The point set that matches the initial location location multi-view point cloud with the location global point cloud is extracted using a point cloud matching algorithm; Substitute the coordinates of the matched point set into the mathematical model, and use the least squares method to calculate the coefficients in the mathematical model, thus obtaining the mathematical model of the transformation relationship between the world coordinate system and the map coordinate system.
6. The method for constructing an offline sensory-audio-visual scene based on online master-slave communication according to claim 1, characterized in that, Step S3.2.3, which involves establishing a digital human model of the subject based on the subject's characteristics, includes the following specific steps: By pre-collecting the facial geometric and texture features of the subject using a terminal device, a standard digital human model of the subject is established. Then, the face of the standard digital human model is personalized by combining the obtained facial photos of the subject.
7. The method for constructing an offline sensory-audio-visual scene based on online master-slave communication according to claim 6, characterized in that, The establishment of the main standard digital human model is specifically divided into the following types: When the pre-collected facial geometric features are key feature points, these key feature points are used to change the facial parameters of the subject's standard digital human model, and at the same time, the subject's facial photo is used to perform texture mapping on the face of the subject's standard digital human model. When the facial geometric features collected in advance are a 3D point cloud of the face, a surface is constructed using these point clouds, and then a main facial photograph is used for texture mapping to replace the face of the standard digital human model.
8. The method for constructing an offline sensory-audio-visual scene based on online master-slave communication according to claim 1, characterized in that, In steps S4 and S5, the corrected formula is: In the formula, represents the corrected volume, v represents the volume of the main body, and d represents the current straight-line distance between the slave and the main body.
Citation Information
Patent Citations
Digital twinborn scene intelligent generation method based on multi-modal visual identification
CN117456136A
Augmented reality-based remote guidance method and device, terminal, and storage medium
WO2019242262A1