Virtual Reality Interaction Method and Audio-Visual Device
The method optimizes virtual reality systems by constructing a virtual space model and audio-visual integration based on user interaction, enhancing immersion and reducing computational load.
Patent Information
- Application Number
- CN202411278199.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-12
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2044-09-12
AI Technical Summary
The existing virtual reality interaction technology has a computing resource burden under high immersion demand, resulting in the problem of system performance degradation.
By inputting image data and real-time face camera processing on virtual reality devices, combining head sensing data, virtual vision and audio areas are generated, regional audio optimization is performed, computing resource burden is reduced, and system performance is improved.
It realizes reducing the burden of computing resources under high immersion, improving system performance, and providing a customized, immersive audio experience.
Smart Images

Figure CN119248106B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of interactive data processing, and particularly to a virtual reality interaction method and an audiovisual device. Background Art
[0002] Modern virtual reality interaction technology is in a rapid development stage. Technological innovation, content expansion, market demand, and social acceptance have jointly promoted the progress of this field. The technical background of virtual reality interaction methods and audiovisual devices lies in creating a fully immersive experience where users can enter a virtual world and interact with it naturally. This generally involves using computer technology to generate a three-dimensional virtual environment and various input and output devices to achieve multi-sensory interaction. Through audiovisual devices, an immersive visual experience is provided, and spatial audio technology simulates a realistic sound environment, capable of creating a highly immersive environment, jointly constructing a multi-sensory integrated virtual world. At the same time, the interactivity is enhanced, and the natural and intuitive interaction method enables a more free exploration of the virtual world and the virtual construction space. However, if there is a need to improve the high immersion, there is a burden on the computing resources of the virtual space, resulting in a reduction in system performance. Summary of the Invention
[0003] Based on this, it is necessary for the present invention to provide a virtual reality interaction method and an audiovisual device to solve at least one of the above technical problems.
[0004] To achieve the above object, a virtual reality interaction method includes the following steps:
[0005] Step S1: Input image data into a virtual reality device to obtain spatial image data; construct a virtual space model from the spatial image data to obtain a virtual space model; construct a virtual audio space from the virtual space model to obtain a virtual audio space;
[0006] Step S2: Perform real-time face camera processing on the virtual reality device to obtain real-time camera stream data; detect a face image from the real-time camera stream data to obtain photographed face image data; generate a fixation point from the photographed face image data to obtain line-of-sight fixation point data; map the line-of-sight fixation point data pair to a coordinate system to obtain fixation point coordinate data;
[0007] Step S3: Obtain head sensing data from the virtual reality device to obtain head sensing data; calculate a deflected fixation point vector from the line-of-sight fixation point data based on the head sensing data to obtain deflected fixation point vector data; calculate a deflected coordinate from the fixation point coordinate data based on the deflected fixation point vector data to obtain a deflected fixation point coordinate;
[0008] Step S4: Based on the deflected fixation point coordinates, obtain the virtual visual area of the virtual space model to get virtual visual area data; based on the virtual visual area data, locate the fixation audio area in the virtual audio space to obtain the fixation virtual audio area; optimize the area audio for the virtual visual area data and the fixation virtual audio area to obtain area audio optimization data.
[0009] Through the initial setup and spatial construction of the virtual reality device, and through image data input, the system can construct a virtual space model and further create a corresponding virtual audio space. This can create a realistic virtual environment for users and lay a foundation for subsequent audio optimization; focus on capturing and analyzing the user's facial images and fixation points. Through real-time face photography and face image detection technology, the system can accurately obtain the user's facial data and the photographed face image. Then, the system generates line-of-sight fixation point data and determines the user's accurate fixation position through coordinate system mapping; utilize head sensing data to enhance the system's understanding of the user's fixation behavior. By combining head sensing data and line-of-sight fixation point data, the system can calculate the deflected fixation point vector, thereby accurately determining the deflected fixation point coordinates of the user in the virtual space; combines the virtual space model and the user's fixation behavior. By obtaining the deflected fixation point coordinates of the user, the system can extract the corresponding virtual visual area from the virtual space model and further locate the virtual audio area that the user is focusing on. Subsequently, the system optimizes the audio for these areas, thereby providing a customized and immersive audio experience. At the same time, by only optimizing the audio for the visual area to reduce the virtual space calculation amount, it is beneficial to reduce the computational resource burden and improve the system performance.
[0010] Preferably, step S1 includes the following steps:
[0011] Step S11: Input image data into the virtual reality device to obtain spatial image data;
[0012] Step S12: Mark the spatial point cloud data for the spatial image data to obtain spatial image point cloud data;
[0013] Step S13: Register the spatial image point cloud data to obtain spatial point cloud registration data;
[0014] Step S14: Extract spatial features from the spatial image point cloud data to obtain spatial point cloud feature data;
[0015] Step S15: Segment the spatial entity from the spatial point cloud feature data to obtain spatial entity data;
[0016] Step S16: Based on the spatial entity data and the spatial point cloud registration data, construct a virtual space model to obtain the virtual space model;
[0017] Step S17: Construct a virtual audio space for the virtual space model to obtain a virtual audio space.
[0018] The present invention obtains image information of the real world through a device, providing basic data for subsequent spatial audio construction, enabling the virtual audio space to correspond to the real environment and enhancing the immersion; converting two-dimensional images into three-dimensional point cloud data with spatial position information, enabling sounds to be placed at more precise spatial positions, such as walls, furniture, etc., enhancing the sense of sound localization and realism; through registration, stitching and fusing point cloud data obtained from different perspectives and at different times to construct a more complete and accurate environmental model, providing a more accurate simulation basis for the propagation of sound in space; extracting geometric and semantic features in the space, such as the shape, size, material of the room, etc., which can be used to simulate acoustic characteristics such as reflection, absorption and diffraction of sound in different spaces, making the virtual audio more realistic; dividing the space into different entity regions, such as walls, floors, ceilings, etc., and assigning different acoustic properties to them, making the propagation and reflection of sound in different regions more in line with the real situation and enhancing the spatial layering of sound; constructing a complete three-dimensional virtual space model, including spatial geometric information, semantic information and acoustic properties, providing a complete data basis for the construction of the virtual audio space; based on the constructed virtual space model, using an acoustic simulation algorithm to calculate the propagation path of sound in the space and the sound field distribution, and finally generating a realistic spatial audio experience that matches the virtual environment.
[0019] Preferably, step S17 includes the following steps:
[0020] Step S171: Set sound sources for the virtual space model to obtain sound source setting data;
[0021] Step S172: Aggregate the sound source setting data to obtain a sound source data set;
[0022] Step S173: Analyze the sound propagation path of the sound source data set based on the virtual space model to obtain sound propagation path data;
[0023] Step S174: Set the propagation path medium for the virtual space model based on the sound propagation path data to obtain propagation path medium data;
[0024] Step S175: Analyze the sound reflection of the propagation path medium data based on the sound source data set and the sound propagation path data to obtain sound reflection data;
[0025] Step S176: Analyze the sound obstacles of the propagation path medium data based on the sound source data set and the sound propagation path data to obtain sound obstacle data;
[0026] Step S177: Perform audio data spatial fitting on the virtual space model based on the sound reflection data and the sound obstacle data to obtain a virtual audio space.
[0027] In the present invention, virtual sound sources, such as human voices, music, etc., are placed at specific positions in the virtual space model, and their acoustic properties, such as volume, timbre, etc., are set to provide basic data for subsequent acoustic simulations; the data of multiple sound sources are integrated together to form a complete sound source data set, which facilitates subsequent unified sound propagation path analysis and audio rendering of multiple sound sources; according to the acoustic principle and the geometric information of the virtual space model, the sound propagation path from the sound source to the listener is calculated, including direct sound, reflected sound, diffracted sound, etc., to provide data support for simulating the real sound propagation effect; according to the sound propagation path, the medium properties of different regions on the path, such as air, walls, furniture, etc., are identified, and different acoustic parameters, such as absorption coefficient, reflection coefficient, etc., are assigned to more accurately simulate the propagation changes of sound in different media; analyze the reflection situation of sound when it encounters different medium surfaces on the propagation path, including reflection direction, reflection intensity, etc., to make the virtual sound more conform to the reflection law of sound in the real world and enhance the sense of space; analyze the occlusion and diffraction situations of sound when it encounters obstacles on the propagation path, such as sound bypassing walls, furniture, etc., to make the spatial occlusion effect of the virtual sound more realistic and credible; apply the sound reflection, obstacle and other data obtained in the previous steps to the virtual space model and perform final audio rendering to generate a complete virtual audio space containing sound sources, spatial information and acoustic effects, so that users can obtain an immersive auditory experience.
[0028] Preferably, step S2 includes the following steps:
[0029] Step S21: Perform real-time face camera processing on the virtual reality device to obtain real-time camera stream data;
[0030] Step S22: Perform face image detection on the real-time camera stream data to obtain photographed face image data;
[0031] Step S23: Extract the eye region from the photographed face image data to obtain photographed eye region data;
[0032] Step S24: Locate the pupil center of the photographed eye region data to obtain pupil center location data;
[0033] Step S25: Extract the iris contour from the photographed eye region data to obtain iris contour image data;
[0034] Step S26: Generate a fixation point based on the pupil center location data and the iris contour image data to obtain line-of-sight fixation point data;
[0035] Step S27: Perform coordinate system mapping on the line-of-sight fixation point data pairs to obtain fixation point coordinate data.
[0036] Through real-time face camera processing of the virtual reality device, the present invention can obtain the facial image data stream when the user wears the device, providing a basic data source for subsequent face detection and eye tracking, enabling the system to capture and analyze the user's facial information in real time; performing face image detection on the real-time camera stream data can accurately identify the position and range of the face in the image, and then separate the face region from the background, providing accurate input data for subsequent eye region extraction; eye region extraction is a key step for the entire system to analyze the user's line-of-sight focus. By accurately extracting the eye region in the photographed face image, other interference information in the image can be effectively removed, improving the accuracy of subsequent pupil positioning and iris recognition; pupil center positioning is an important step to determine the user's line-of-sight direction. This step analyzes the eye image data to identify the position of the pupil in the eye, providing basic reference data for line-of-sight tracking and focus judgment; iris contour extraction can provide important features for accurate iris recognition and subsequent line-of-sight tracking. By extracting the contour shape, texture, and other unique features of the iris, the user can be uniquely identified, and the accuracy of the entire system can be improved; combining the pupil center and iris contour information, through geometric calculation or deep learning models, the line-of-sight fixation point of the user in the three-dimensional space is calculated, that is, the direction and position that the user is observing; mapping the line-of-sight fixation point from the image coordinate system to the coordinate system of the virtual scene enables the virtual scene to perform real-time interaction and feedback according to the user's line of sight, such as fixation point selection, line-of-sight interaction, etc., improving the immersion and interactivity of the experience.
[0037] Preferably, step S26 includes the following steps:
[0038] Step S261: Extract iris feature points from the iris contour image data to obtain iris feature point data;
[0039] Step S262: Construct an eyeball model based on the pupil center positioning data and the iris feature point data to obtain an eyeball model;
[0040] Step S263: Obtain the eyeball optical axis direction data based on the pupil center positioning data for the eyeball model;
[0041] Step S264: Simulate the line-of-sight intersection point for the eyeball optical axis direction data to obtain line-of-sight intersection point data;
[0042] Step S265: Generate a line-of-sight fixation point for the line-of-sight intersection point data to obtain line-of-sight fixation point data.
[0043] Through this step, the present invention extracts representative feature points from the iris contour image, such as iris texture, color change, etc. These feature points can be used to more accurately fit the eyeball model and improve the accuracy of gaze estimation; The construction of the eyeball model is an important step in simulating the structure and movement of the human eye. By combining the pupil center localization data and the iris feature point data, an accurate three-dimensional eyeball model can be established. This model can simulate the shape, size and structure of the human eye and provide a basic framework for analyzing the optical axis and gaze movement of the eyeball; Obtaining the direction of the eyeball optical axis data is the key to understanding the gaze direction. Through the pupil center localization data, the direction vector of the eyeball optical axis can be calculated. This vector represents the direction from the pupil center to the exit direction of the eyeball model and provides an important reference for simulating the gaze intersection point and finally determining the gaze fixation point; The gazes of the left and right eyes will intersect at a point in space, that is, the gaze intersection point, which represents the specific position that the user is looking at in the three-dimensional space. By simulating the intersection process of the binocular gazes, the user's fixation point can be inferred more accurately; Taking the calculated gaze intersection point position as the user's gaze fixation point, this data can be used in various interactions and applications in the system, such as gaze selection, fixation point rendering, eye movement analysis, etc., to enhance the immersion and interactivity of the experience.
[0044] Preferably, step S3 includes the following steps:
[0045] Step S31: Obtain head sensing data for the virtual reality device to obtain head sensing data;
[0046] Step S32: Analyze the head posture data of the head sensing data to obtain head posture data;
[0047] Step S33: Analyze the head deflection characteristics of the head sensing data based on the head posture data to obtain head deflection characteristic data;
[0048] Step S34: Calculate the deflected fixation point vector for the gaze fixation point data based on the head deflection characteristic data to obtain deflected fixation point vector data;
[0049] Step S35: Calculate the deflected coordinates for the fixation point coordinate data based on the deflected fixation point vector data to obtain the deflected fixation point coordinates.
[0050] The head sensing data of the virtual reality device in the present invention is the basis for capturing the head movement and posture changes of the user. Through sensors built into the virtual reality device, such as gyroscopes, accelerometers, and magnetometers, the position, rotation, and movement information of the user's head in three-dimensional space can be obtained in real time, providing the necessary data support for subsequent analysis of head posture and deflection characteristics; the analysis of head posture data can help the system understand the current state of the user's head. By processing and analyzing the head sensing data, posture information such as the rotation, pitch, and deflection angles of the head in the virtual space can be obtained, providing a basic reference for subsequent deflection characteristic analysis and gaze tracking; the analysis of head deflection characteristics aims to identify and quantify the lateral deflection movement of the user's head. By further processing the head posture data, the deflection angle and speed of the head in the horizontal direction can be extracted, providing key input data for calculating the deflection gaze point vector and finally determining the deflection gaze point coordinates; according to the head deflection characteristics of the user, a deflection vector is calculated to correct the original gaze point direction to make it more in line with the actual gaze direction of the user; by combining the head sensing data of the virtual reality device and the deflection gaze point vector, the new gaze point coordinates generated by the user's head deflection in the virtual space can be calculated. This coordinate data can be corresponded to the objects in the virtual environment to achieve more accurate user interaction and a more natural human-computer interaction experience.
[0051] Preferably, step S33 includes the following steps:
[0052] Step S331: Set a reference posture for the head posture data to obtain head reference posture data;
[0053] Step S332: Calculate the head deflection angle based on the head reference posture data for the head sensing data to obtain the head deflection angle;
[0054] Step S333: Analyze the direction of the deflection vector for the head deflection angle to obtain deflection vector direction data;
[0055] Step S334: Calculate the magnitude of the deflection vector for the head deflection angle to obtain deflection vector magnitude data;
[0056] Step S335: Fit the characteristic data of the deflection vector direction data and the deflection vector magnitude data to obtain head deflection characteristic data.
[0057] By setting the head reference pose, the present invention can provide a reference basis for subsequent calculation and analysis of the deflection angle. By establishing a standard or initial head pose as a reference, the degree of deflection of the user's head can be calculated and quantified more accurately. This reference pose is usually defined as the pose when the user looks straight ahead, providing a neutral reference point for subsequent analysis; The calculation of the head deflection angle is the basis for understanding the lateral movement of the user's head. By comparing the head sensing data with the head reference pose data, the deflection angle of the user's head in the horizontal direction can be calculated. This angle represents the degree of deviation of the head from the reference pose, providing a quantitative index for subsequent deflection vector analysis and simulation; By analyzing the head deflection angle, the direction vector of the deflection can be determined. This direction vector indicates whether the head deflection is in the left, right, up, down or other specific directions, providing important input information for simulating and simulating the user's line-of-sight movement; By analyzing the head deflection angle, the magnitude of the deflection vector can be calculated, that is, the distance or angle size of the deflection. This magnitude data represents the intensity of the head deflection, providing parameters for simulating and analyzing the user's visual field range and head movement amplitude; Feature data fitting is the process of integrating the direction and magnitude data of the deflection vector. By mathematically fitting the direction data and magnitude data, a more complete head deflection feature model can be established. This model synthesizes the deflection direction and magnitude information, can more accurately represent the user's head deflection behavior, and provides reliable input data for subsequent line-of-sight tracking, virtual space interaction or other related applications.
[0058] Preferably, step S4 includes the following steps:
[0059] Step S41: Perform coordinate grid mapping on the virtual space model to obtain a virtual space coordinate grid;
[0060] Step S42: Based on the coordinates of the deflected fixation point, perform spatial positioning of the deflected fixation point on the virtual space coordinate grid to obtain spatial deflected fixation point data;
[0061] Step S43: Obtain virtual visual area data by acquiring the virtual visual area from the spatial deflected fixation point data;
[0062] Step S44: Based on the virtual visual area data, perform fixation audio area positioning on the virtual audio space to obtain a fixation virtual audio area;
[0063] Step S45: Optimize the area audio for the virtual visual area data and the fixation virtual audio area to obtain area audio optimization data.
[0064] The present invention can establish a coordinate framework for a virtual space by performing coordinate grid mapping on the virtual space model. By creating a three-dimensional coordinate grid system, a unique coordinate value can be assigned to each point in the virtual space. This provides a basic framework for deflected fixation point positioning, virtual vision acquisition, and audio optimization; the process of spatial positioning of the deflected fixation point is to map the coordinates of the deflected fixation point onto the virtual space coordinate grid. Through this step, the exact position of the user's deflected line-of-sight focus in the virtual space can be determined, providing a basic reference for subsequent virtual vision and audio region acquisition; by analyzing the deflected fixation point data, the region that enters the user's field of view in the virtual space, i.e., the virtual vision region, can be determined. This region will become the part of the virtual environment currently visible to the user, providing basic input for fixation audio region positioning and regional audio optimization; by analyzing the virtual vision region data, the virtual audio region currently fixated by the user can be determined. This step ensures that audio processing focuses on the virtual space region that the user is interested in or pays attention to, enhancing the relevance and accuracy of the audio experience; regional audio optimization aims to improve the quality of the user's audio experience in the virtual space. Through comprehensive analysis of the virtual vision region data and the fixation audio region data, the audio effect can be optimized. For example, enhancing the audio clarity within the virtual vision region, increasing the volume of the virtual audio region, adding ambient sound field effects, etc., to make the audio experience more realistic and immersive.
[0065] Preferably, step S45 includes the following steps:
[0066] Step S451: Analyze the regional acoustic characteristics of the fixated virtual audio region to obtain regional acoustic characteristic data;
[0067] Step S452: Based on the regional acoustic characteristic data, perform spatial region sound field simulation on the virtual audio space to obtain spatial region sound field data;
[0068] Step S453: Obtain the fixated visual path from the virtual vision region data;
[0069] Step S454: Identify the audio occluders on the fixated visual path to obtain path audio occluder data;
[0070] Step S455: Based on the path audio occluder data, perform regional audio simulation optimization on the spatial region sound field data to obtain regional audio optimization data.
[0071] The present invention aims to understand and quantify the sound characteristics of a virtual audio region through regional acoustic characteristic analysis. By analyzing the acoustic reflection, absorption, diffusion, and other acoustic characteristics of this region, regional acoustic characteristic data can be obtained. These data include material characteristics, boundary effects, acoustic environment, etc., providing basic parameters for subsequent spatial region sound field simulation; by applying acoustic principles and regional acoustic characteristic data, the propagation, reflection, and attenuation effects of sound in a specific virtual audio region can be simulated and predicted. This step helps to establish a realistic spatial audio model, providing a basic framework for final audio optimization; by analyzing virtual visual region data, the visual path from the user's deflected fixation point to other points in the virtual visual region can be determined. This path represents the movement trajectory of the user's line of sight in the virtual space, providing an important reference for subsequent audio occluder recognition and optimization; audio occluder recognition aims to identify obstacles in the virtual space that affect sound propagation and auditory experience. By analyzing the fixation visual path, audio occluders on the visual path can be identified, such as virtual walls, furniture, or other virtual objects. This step provides the necessary data for regional audio simulation optimization, ensuring that the propagation and reflection of sound conform to the virtual environment settings; through the comprehensive analysis of these two factors, the regional audio effect can be optimized. For example, simulating the attenuation, reflection, or diffraction of sound when encountering an occluder, and applying an appropriate acoustic model to achieve a more realistic audio experience. This optimization step ensures that the audio presentation in the virtual space is consistent with the actual physical environment, enhancing the immersion and spatial sense in virtual reality or mixed reality applications.
[0072] The present invention also provides an audiovisual device, including a head-mounted display for performing the virtual reality interaction method as described above. The head-mounted display of this audiovisual device includes:
[0073] A virtual audio space construction module for inputting image data into the virtual reality device to obtain spatial image data; constructing a virtual space model from the spatial image data to obtain a virtual space model; constructing a virtual audio space from the virtual space model to obtain a virtual audio space;
[0074] A fixation point coordinate generation module for performing real-time face camera processing on the virtual reality device to obtain real-time camera stream data; performing face image detection on the real-time camera stream data to obtain photographed face image data; generating a fixation point from the photographed face image data to obtain line-of-sight fixation point data; performing coordinate system mapping on the line-of-sight fixation point data pair to obtain fixation point coordinate data;
[0075] The deflection fixation point generation module is used to obtain head sensing data from a virtual reality device, and obtain head sensing data; calculate the deflection fixation point vector for the line-of-sight fixation point data based on the head sensing data to obtain deflection fixation point vector data; calculate the deflection coordinates for the fixation point coordinate data based on the deflection fixation point vector data to obtain the deflection fixation point coordinates.
[0076] The regional audio optimization module is used to obtain the virtual visual area data by acquiring the virtual visual area of the virtual space model based on the deflection fixation point coordinates; perform fixation audio area positioning on the virtual audio space based on the virtual visual area data to obtain the fixation virtual audio area; perform regional audio optimization on the virtual visual area data and the fixation virtual audio area to obtain the regional audio optimization data.
[0077] In summary, the present invention provides a virtual reality interaction method and an audiovisual device. The virtual reality interaction method and the audiovisual device are composed of a virtual audio space construction module, a fixation point coordinate generation module, a deflection fixation point generation module, and a regional audio optimization module, and can implement any virtual reality interaction method described in the present invention. It is used to jointly implement any virtual reality interaction method through the operations between the computer programs running on each module. The internal structure of the system cooperates with each other, which can greatly reduce repetitive work and manpower input, and can quickly and effectively provide a more accurate and efficient virtual reality interaction process, thereby simplifying the virtual reality interaction method and the audiovisual device. Description of the Drawings
[0078] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, purposes, and advantages of the present invention will become more obvious:
[0079] Figure 1 It is a schematic flowchart of the steps of a virtual reality interaction method of the present invention;
[0080] Figure 2 is Figure 1 a detailed schematic flowchart of step S4 in
[0081] Figure 3 is Figure 2 a detailed schematic flowchart of step S45 in Detailed Embodiment
[0082] The technical method of the present invention will be clearly and completely described below with reference to the drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.
[0083] In addition, the accompanying drawings are only schematic illustrations of the present invention and are not necessarily drawn to scale. The same reference numerals in the drawings denote the same or similar parts, and thus repeated descriptions thereof will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. The functional entities may be implemented in software form, or in one or more hardware modules or integrated circuits, or in different networks and / or processor methods and / or microcontroller methods.
[0084] It should be understood that although terms such as "first" and "second" may be used herein to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, the first unit may be referred to as the second unit, and similarly the second unit may be referred to as the first unit. The term "and / or" used herein includes any and all combinations of one or more of the listed associated items.
[0085] To achieve the above object, please refer to Figures 1 to 3 , the present invention provides a virtual reality interaction method, including the following steps:
[0086] Step S1: Input image data into the virtual reality device to obtain spatial image data; construct a virtual space model from the spatial image data to obtain a virtual space model; construct a virtual audio space from the virtual space model to obtain a virtual audio space;
[0087] Step S2: Perform real-time face camera processing on the virtual reality device to obtain real-time photography stream data; detect a face image from the real-time photography stream data to obtain photographed face image data; generate a fixation point from the photographed face image data to obtain line-of-sight fixation point data; map the line-of-sight fixation point data pair to a coordinate system to obtain fixation point coordinate data;
[0088] Step S3: Obtain head sensing data from the virtual reality device to obtain head sensing data; calculate a deflected fixation point vector from the line-of-sight fixation point data based on the head sensing data to obtain deflected fixation point vector data; calculate deflected coordinates from the fixation point coordinate data based on the deflected fixation point vector data to obtain deflected fixation point coordinates;
[0089] Step S4: Obtain a virtual visual area from the virtual space model based on the deflected fixation point coordinates to obtain virtual visual area data; locate a fixation audio area in the virtual audio space based on the virtual visual area data to obtain a fixation virtual audio area; optimize the area audio from the virtual visual area data and the fixation virtual audio area to obtain area audio optimization data.
[0090] In an embodiment of the present invention, please refer to Figure 1 As shown, it is a schematic flow chart of the steps of the virtual reality interaction method of the present invention. In this example, the virtual reality interaction method includes the following steps:
[0091] Step S1: Input image data into the virtual reality device to obtain spatial image data; construct a virtual space model from the spatial image data to obtain a virtual space model; construct a virtual audio space from the virtual space model to obtain a virtual audio space;
[0092] In an embodiment of the present invention, by using a virtual reality device to scan a target space, and collect color images and depth information. Set scanning parameters, scan the space, and perform real-time image stitching to generate spatial image data. By using point cloud processing software, convert the two-dimensional image into three-dimensional point cloud data, and perform operations such as noise reduction, feature extraction, and depth image conversion. Through feature matching and pose estimation, fuse the point cloud data from different perspectives into a complete spatial point cloud model. Based on the point cloud registration data, extract spatial feature information, including normal vectors, curvatures, geometric feature descriptions, and color features, etc. According to the point cloud feature data, segment the point cloud into different spatial entities and add semantic labels. According to the segmented spatial entity data, use 3D modeling software to construct a virtual space model. Based on the virtual space model, use virtual acoustics software to simulate the interaction of sound in the space and construct a realistic virtual audio space.
[0093] Step S2: Perform real-time face camera processing on the virtual reality device to obtain real-time photography stream data; perform face image detection on the real-time photography stream data to obtain photographed face image data; generate a gaze point from the photographed face image data to obtain gaze point data; perform coordinate system mapping on the gaze point data pair to obtain gaze point coordinate data;
[0094] In an embodiment of the present invention, by using the built-in camera of the virtual reality device to capture the user's face image in real time and decode it into digital image frames to form real-time photography stream data. For each frame of data, first perform face detection to locate the face area, then extract the eye area and perform pupil center positioning and iris contour extraction. Calculate the user's gaze point by combining the pupil center and iris information. Finally, map the gaze point from the image coordinate system to the screen coordinate system or three-dimensional space coordinate system of the virtual reality device to obtain the final gaze point coordinate data.
[0095] Step S3: Obtain head sensing data from the virtual reality device to obtain head sensing data; calculate a deflected gaze point vector from the gaze point data based on the head sensing data to obtain deflected gaze point vector data; calculate a deflected coordinate from the gaze point coordinate data based on the deflected gaze point vector data to obtain a deflected gaze point coordinate;
[0096] In the embodiment of the present invention, by using the high-precision sensors built in the virtual reality device, such as gyroscopes and accelerometers, the movement of the user's head is tracked to obtain information such as the head rotation angle and acceleration. The data is denoised and fused through algorithms such as Kalman filtering to obtain accurate head pose data. Then, the head pose change rate is calculated as the head deflection feature, and a mapping relationship is established to convert it into a line-of-sight deflection vector. Finally, the original line-of-sight fixation point coordinates and the deflection fixation point vector are added together to obtain the screen position coordinates that the user is about to fixate on.
[0097] Step S4: Based on the deflection fixation point coordinates, obtain the virtual visual area of the virtual space model to get the virtual visual area data; based on the virtual visual area data, locate the fixation audio area in the virtual audio space to obtain the fixation virtual audio area; optimize the area audio for the virtual visual area data and the fixation virtual audio area to obtain the area audio optimization data.
[0098] In the embodiment of the present invention, the virtual space model is loaded into the system and divided into a series of coordinate grids with unique identifiers. Then, the user's deflection fixation point coordinates are mapped to the virtual space coordinate grid to find the corresponding grid, and information such as its coordinate identifier is recorded as the spatial deflection fixation point data. Based on this data, the user's virtual visual area is determined, for example, a spherical area centered on the deflection fixation point with the user's field of view as the radius. Then, the corresponding fixation virtual audio area is located in the virtual audio space, for example, the projection area of the virtual visual area in the virtual audio space. Finally, according to the virtual visual area data and the fixation virtual audio area, the virtual audio is optimized, for example, improving the sound clarity and volume of the fixation area, and information such as the optimized audio parameters is recorded as the area audio optimization data.
[0099] Through the initial setup and spatial construction of the virtual reality device, and through the input of image data, the system can construct a virtual space model and further create a corresponding virtual audio space. This can create a realistic virtual environment for users and lay a foundation for subsequent audio optimization; focus on capturing and analyzing the user's facial images and fixation points. Through real-time face camera and face image detection technologies, the system can accurately obtain the user's facial data and the photographed face image. Then, the system generates line-of-sight fixation point data and determines the user's accurate fixation position through coordinate system mapping; utilize head sensing data to enhance the system's understanding of the user's fixation behavior. By combining head sensing data and line-of-sight fixation point data, the system can calculate the deflected fixation point vector to accurately determine the deflected fixation point coordinates of the user in the virtual space; combine the virtual space model and the user's fixation behavior. By obtaining the deflected fixation point coordinates of the user, the system can extract the corresponding virtual visual area from the virtual space model and further locate the virtual audio area that the user is focusing on. Subsequently, the system performs audio optimization on these areas to provide a customized and immersive audio experience. At the same time, by only performing audio optimization on the visual area to reduce the virtual space calculation amount, it is beneficial to reduce the computational resource burden and improve the system performance.
[0100] Preferably, step S1 includes the following steps:
[0101] Step S11: Input image data into the virtual reality device to obtain spatial image data;
[0102] In an embodiment of the present invention, by using a virtual reality device with a depth sensor to scan the target space, the device will simultaneously collect color images and depth information. The specific operation is to set scanning parameters, and according to the size and detail requirements of the target space, set parameters such as scanning resolution, scanning range, and frame rate; perform spatial scanning, hold the virtual reality device, and comprehensively scan the target space along a predetermined route, keeping the device moving steadily to avoid image blur or distortion; perform real-time image stitching; the algorithm built into the device will stitch the collected images and depth information in real time to generate preliminary spatial image data.
[0103] Step S12: Mark the spatial point cloud data of the spatial image data to obtain spatial image point cloud data;
[0104] In the embodiments of the present invention, by using the acquired spatial image data and combining with point cloud processing software, the two-dimensional image is converted into three-dimensional point cloud data and marked. The specific operations include depth image conversion. By using the depth value of each pixel in the depth image and combining with the camera parameters, each pixel is back-projected into the three-dimensional space to form point cloud data; downsampling of point cloud data. The original point cloud data is usually very large and needs to be downsampled to reduce the computational amount. Methods such as uniform downsampling and random downsampling can be used; denoising of point cloud data. Due to factors such as sensor errors, there will be noise points in the point cloud data, and methods such as statistical filtering and radius filtering need to be used to remove them.
[0105] Step S13: Perform point cloud data registration on the spatial image point cloud data to obtain spatial point cloud registration data;
[0106] In the embodiments of the present invention, the point cloud data from different perspectives is stitched into a complete spatial point cloud model. The specific operations include feature point extraction. Feature points such as corner points and edge points are extracted from the point cloud data from different perspectives. Commonly used feature point extraction algorithms include ISS, SIFT, NARF, etc.; feature matching. Using the geometric information of the feature points, such as normal vectors and curvatures, for feature matching to find the corresponding relationship between the point cloud data from different perspectives; pose estimation. According to the matched feature point pairs, calculate the rotation and translation transformation matrices between the point cloud data from different perspectives; point cloud fusion. Using the calculated pose transformation matrix, convert the point cloud data from different perspectives to the same coordinate system for fusion to obtain a complete spatial point cloud model, and finally obtain the complete spatial point cloud data after registration and fusion, that is, spatial point cloud registration data.
[0107] Step S14: Perform spatial feature extraction on the spatial image point cloud data to obtain spatial point cloud feature data;
[0108] In the embodiments of the present invention, based on the spatial point cloud registration data, spatial feature information is extracted to provide a data basis for subsequent entity segmentation and model construction. The specific operations include normal vector analysis. Calculate the normal vector of each point to describe the orientation information of the point cloud surface. Commonly used normal vector estimation methods include PCA, least squares method, etc.; curvature calculation. Calculate the curvature of each point to describe the degree of curvature of the point cloud surface. Curvature can be used to distinguish different shaped objects such as planes, spheres, and cylinders; geometric feature description. Using the information such as the normal vector and curvature of the point cloud, construct geometric feature descriptors of the point cloud, such as PFH, SHOT, FPFH, etc.; color feature extraction. If the point cloud data contains color information, color features such as color histograms and color moments can be extracted. Finally, point cloud data containing information such as normal vectors, curvatures, geometric feature descriptors, and color features is obtained, that is, spatial point cloud feature data.
[0109] Step S15: Perform spatial entity segmentation on the spatial point cloud feature data to obtain spatial entity data;
[0110] In the embodiments of the present invention, by using the point cloud feature data, clustering algorithms, region growing algorithms, etc. are used to segment the point cloud into different spatial entities, such as walls, floors, furniture, etc., and semantic labels are added to each entity. For example, using a region-growing-based segmentation algorithm, according to features such as the normal vector and curvature of the point cloud, the point cloud is segmented into different planar regions. Then, according to the semantic labels and spatial position relationships of the point cloud, the segmented planar regions are merged into different spatial entities, such as walls, floors, etc.
[0111] Step S16: Construct a virtual space model based on the spatial entity data and the spatial point cloud registration data to obtain a virtual space model;
[0112] In the embodiments of the present invention, by using three-dimensional modeling software, such as 3dsMax, Maya, Blender, etc., a virtual space model is constructed according to the segmented spatial entity data. Different modeling methods can be selected according to needs. For example, simple geometric bodies can be used to represent the spatial structure, or more refined models can be used to restore the real scene. For example, using Blender software, a corresponding three-dimensional model is created according to the geometric shape and size of the spatial entity.
[0113] Step S17: Construct a virtual audio space for the virtual space model to obtain a virtual audio space.
[0114] In the embodiments of the present invention, based on the constructed virtual space model, virtual acoustic software can be used to simulate the propagation and interaction of sound in space to construct a virtual audio space. The construction of a virtual audio space needs to consider physical phenomena such as sound reflection, diffraction, and absorption, as well as the auditory characteristics of the human ear. During specific operations, parameters such as the position, direction, and material properties of the sound source, as well as information such as the size, shape, and sound absorption coefficient of the room, need to be set. The virtual acoustic software will calculate the propagation path of sound in space and the sound field distribution according to these parameters to simulate a realistic auditory experience. For example, according to the geometric shape and material of the room, the sound absorption coefficients of objects such as walls, floors, and furniture can be set to simulate the reflection and absorption of sound on different material surfaces to construct an audio space.
[0115] The present invention acquires image information of the real world through a device, providing basic data for subsequent spatial audio construction, enabling the virtual audio space to correspond to the real environment and enhancing the immersion; converting two-dimensional images into three-dimensional point cloud data with spatial position information, enabling sounds to be placed at more precise spatial positions, such as walls, furniture, etc., enhancing the sense of sound localization and realism; through registration, stitching and fusing point cloud data acquired from different perspectives and at different times to construct a more complete and accurate environmental model, providing a more accurate simulation basis for the propagation of sound in space; extracting geometric and semantic features in the space, such as the shape, size, material of the room, etc., which can be used to simulate acoustic characteristics such as reflection, absorption, and diffraction of sound in different spaces, making the virtual audio more vivid; dividing the space into different entity regions, such as walls, floors, ceilings, etc., and assigning different acoustic properties to them, enabling the propagation and reflection of sound in different regions to be more in line with the real situation and enhancing the spatial layering of sound; constructing a complete three-dimensional virtual space model, including spatial geometric information, semantic information, and acoustic properties, providing a complete data basis for the construction of the virtual audio space; based on the constructed virtual space model, using acoustic simulation algorithms to calculate the propagation path and sound field distribution of sound in the space, and finally generating a realistic spatial audio experience that matches the virtual environment.
[0116] Preferably, step S17 includes the following steps:
[0117] Step S171: Set the sound source for the virtual space model to obtain sound source setting data;
[0118] In the embodiment of the present invention, according to the design requirements of the virtual audio space, parameters such as the position, direction, type, and volume of the sound source are set in the virtual space model, and a suitable audio file is selected to correspond to it. For example, in a virtual meeting room scenario, a sound source can be set at the position of each seat to simulate the speech of the participants; in a virtual concert hall scenario, a sound source can be set in the center of the stage to simulate the performance of musical instruments. The setting data of each sound source includes its three-dimensional coordinates in the virtual space, sound source type, such as point sound source, line sound source, surface sound source, directivity, initial volume, audio file path, and other information.
[0119] Step S172: Aggregate the sound source setting data to obtain a sound source data set;
[0120] In the embodiment of the present invention, by integrating the setting data of all sound sources to form a sound source data set for subsequent sound propagation path analysis and audio rendering, the sound source data set can be stored using data structures such as lists and arrays. Each element represents a sound source and contains all the setting data of that sound source.
[0121] Step S173: Perform acoustic propagation path analysis on the sound source dataset based on the virtual space model to obtain acoustic propagation path data;
[0122] In the embodiments of the present invention, by according to the geometric information of the sound source dataset and the virtual space model, using methods such as the ray tracing algorithm and the path tracing algorithm, calculate the propagation path of sound from the sound source to the listener. Acoustic propagation path analysis needs to consider phenomena such as sound reflection, diffraction, and transmission, as well as the influence of different materials on sound propagation. For example, using the ray tracing algorithm, emit multiple sound rays from each sound source, simulate the propagation of sound in the virtual space, and record information such as the propagation path, reflection times, and transmission coefficient of each sound ray.
[0123] Step S174: Perform propagation path medium setting on the virtual space model based on the acoustic propagation path data to obtain propagation path medium data;
[0124] In the embodiments of the present invention, by according to the acoustic propagation path data, identify different media passed through on the sound propagation path, and set corresponding acoustic parameters for them, such as the sound absorption coefficient, reflection coefficient, propagation speed, etc. For example, according to the results of ray tracing, the materials of walls, floors, furniture, etc. passed through by each sound ray can be determined, and their acoustic parameters can be set according to the material type.
[0125] Step S175: Perform sound reflection analysis on the propagation path medium data based on the sound source dataset and the acoustic propagation path data to obtain sound reflection data;
[0126] In the embodiments of the present invention, by according to the sound source dataset, the acoustic propagation path data, and the propagation path medium data, calculate the reflection phenomenon that occurs on the sound propagation path. Factors such as the reflection coefficient, reflection angle, and reflection energy attenuation of different materials need to be considered. For example, according to the intersection points of each sound ray with the surfaces of different media, calculate the propagation direction and energy attenuation of the reflected sound ray, and add the reflected sound ray to the acoustic propagation path data.
[0127] Step S176: Perform sound obstacle analysis on the propagation path medium data based on the sound source dataset and the acoustic propagation path data to obtain sound obstacle data;
[0128] In the embodiments of the present invention, by according to the sound source dataset, the acoustic propagation path data, and the propagation path medium data, analyze the obstacles encountered on the sound propagation path, and calculate the diffraction phenomenon of sound. Factors such as the size, shape, and material of the obstacles on the influence of sound diffraction need to be considered. For example, for the obstacles existing on the sound ray propagation path, the diffraction angle and energy attenuation of sound can be calculated according to their size and shape, and the diffracted sound ray can be added to the acoustic propagation path data.
[0129] Step S177: Perform audio data spatial fitting on the virtual space model based on the sound reflection data and the sound obstacle data to obtain a virtual audio space.
[0130] In the embodiment of the present invention, by performing audio data spatial fitting on the virtual space model according to the sound reflection data and the sound obstacle data, the calculated sound field information is integrated with the virtual space model to construct a virtual audio space. The specific operations include adding information such as sound sources, sound propagation paths, reflected sounds, and diffracted sounds to the virtual space model, and adjusting the sound energy according to factors such as distance attenuation and air absorption, and finally generating audio data containing spatial information.
[0131] In the present invention, virtual sound sources, such as human voices, music, etc., are placed at specific positions in the virtual space model, and their acoustic properties, such as volume, timbre, etc., are set to provide basic data for subsequent acoustic simulations; the data of multiple sound sources are integrated together to form a complete sound source data set, which is convenient for subsequent unified sound propagation path analysis and audio rendering of multiple sound sources; according to the acoustic principle and the geometric information of the virtual space model, the sound propagation path from the sound source to the listener is calculated, including direct sound, reflected sound, diffracted sound, etc., to provide data support for simulating the real sound propagation effect; according to the sound propagation path, the medium properties of different regions on the path are identified, such as air, walls, furniture, etc., and different acoustic parameters, such as absorption coefficient, reflection coefficient, etc., are assigned to more accurately simulate the propagation changes of sound in different media; analyze the reflection situation of sound when it encounters different medium surfaces on the propagation path, including reflection direction, reflection intensity, etc., to make the virtual sound more conform to the reflection law of sound in the real world and enhance the sense of space; analyze the occlusion and diffraction situations of sound when it encounters obstacles on the propagation path, such as sound bypassing walls, furniture, etc., to make the spatial shielding effect of the virtual sound more realistic and credible; apply the sound reflection, obstacle and other data obtained in the previous steps to the virtual space model and perform final audio rendering to generate a complete virtual audio space containing sound sources, spatial information and acoustic effects, so that users can obtain an immersive auditory experience.
[0132] Preferably, step S2 includes the following steps:
[0133] Step S21: Perform real-time face camera processing on the virtual reality device to obtain real-time camera stream data;
[0134] In the embodiments of the present invention, by using a camera integrated inside a virtual reality device, a face image of the user is captured in real time. The camera can be an infrared camera or a visible light camera, which is selected according to the application scenario and accuracy requirements. For example, an in-built infrared camera can be used for eye movement tracking to obtain higher-precision eye images. The image sequence collected by the camera is decoded into digital image frames to form real-time photographic stream data.
[0135] Step S22: Perform face image detection on the real-time photographic stream data to obtain photographic face image data;
[0136] In the embodiments of the present invention, by performing face detection on each frame of the real-time photographic stream data, the face area is identified and located. Methods such as Haar feature-based, HOG feature-based, and deep learning can be used for face detection. For example, the MTCNN model is used to perform face detection on the image frame to obtain the coordinate information of the face bounding box, and according to the bounding box coordinates, the face area is cropped from the image frame to obtain the photographic face image data.
[0137] Step S23: Extract the eye area from the photographic face image data to obtain photographic eye area data;
[0138] In the embodiments of the present invention, by further locating and extracting the eye area in the detected face image. Feature-based methods such as ASM and CLM, or deep learning-based methods such as EyeNet and GazeNet can be used. For example, using a pre-trained model provided by the library to detect the eye key points in the face image, including the positions of the eye corners and eyelids. According to the key point coordinates, the range of the eye area is determined, and the eye area image is cropped to obtain the photographic eye area data.
[0139] Step S24: Locate the pupil center for the photographic eye area data to obtain pupil center location data;
[0140] In the embodiments of the present invention, by performing pupil center location on the extracted eye area image. Methods based on circular detection such as Hough transform and RANSAC, or methods based on gray-scale features such as the Starburst method and the center offset method can be used. For example, using the center offset method based on gray-scale projection to calculate the horizontal and vertical gray-scale projection curves of the eye area image, finding the peak points of the projection curves as the candidate positions of the pupil center, and screening according to the gray-scale features of the pupil area to determine the final pupil center coordinates to obtain the pupil center location data.
[0141] Step S25: Extract the iris contour for the photographic eye area data to obtain iris contour image data;
[0142] In an embodiment of the present invention, the contour information of the iris is extracted from the eye region image. A method based on edge detection can be used, such as the Canny operator, Sobel operator, etc., or a method based on circle detection, such as the Hough transform, RANSAC, etc. For example, using the Hough transform method based on circle detection, circular edges are detected in the eye region image, a circular contour that conforms to the iris characteristics is found, and the edge pixel coordinates are extracted to obtain iris contour image data.
[0143] Step S26: Generate a fixation point from the pupil center localization data and the iris contour image data to obtain line-of-sight fixation point data;
[0144] In an embodiment of the present invention, the line of sight of the user is calculated by combining the pupil center localization data and the iris contour image data. A method based on geometric relationships can be used, such as the vector method, trigonometric function method, etc. For example, according to the pupil center coordinates and the center coordinates of the iris contour, the line-of-sight direction vector is calculated; then, according to the line-of-sight direction vector and the eyeball model, the intersection point of the line of sight and the screen of the virtual reality device is calculated as the line of sight fixation point of the user to obtain the line-of-sight fixation point data.
[0145] Step S27: Map the line-of-sight fixation point data pair to obtain fixation point coordinate data.
[0146] In an embodiment of the present invention, by mapping the line-of-sight fixation point data from the image coordinate system to the screen coordinate system or the three-dimensional space coordinate system of the virtual reality device, subsequent application processing is facilitated. For example, according to the screen resolution and field of view angle of the virtual reality device, the line-of-sight fixation point is mapped from the image coordinate system to the screen coordinate system to obtain the screen coordinates of the user's fixation point; or according to the pose information and camera parameters of the virtual reality device, the line-of-sight fixation point is mapped from the image coordinate system to the three-dimensional space coordinate system to obtain the three-dimensional coordinates of the user's fixation point in the virtual world, that is, the fixation point coordinate data.
[0147] Through real-time face camera processing of virtual reality devices, the present invention can obtain the facial image data stream when the user wears the device, providing a basic data source for subsequent face detection and eye tracking, enabling the system to capture and analyze the user's facial information in real time; performing face image detection on the real-time camera stream data can accurately identify the position and range of the face in the image, and then separate the face area from the background, providing accurate input data for subsequent eye area extraction; eye area extraction is a key step in the system to analyze the user's line of sight focus. By accurately extracting the eye area in the photographed face image, other interference information in the image can be effectively removed, improving the accuracy of subsequent pupil positioning and iris recognition; pupil center positioning is an important step in determining the user's line of sight direction. This step analyzes the eye image data to identify the position of the pupil in the eye, providing basic reference data for line of sight tracking and focus judgment; iris contour extraction can provide important features for accurate iris recognition and subsequent line of sight tracking. By extracting the contour shape, texture, and other unique features of the iris, the user can be uniquely identified, and the accuracy of the entire system can be improved; combining the pupil center and iris contour information, through geometric calculation or deep learning models, the line of sight fixation point of the user in the three-dimensional space is calculated, that is, the direction and position the user is observing; mapping the line of sight fixation point from the image coordinate system to the coordinate system of the virtual scene enables the virtual scene to perform real-time interaction and feedback according to the user's line of sight, such as fixation point selection, line of sight interaction, etc., improving the immersion and interactivity of the experience.
[0148] Preferably, step S26 includes the following steps:
[0149] Step S261: Extract iris feature points from the iris contour image data to obtain iris feature point data;
[0150] In the embodiment of the present invention, by using the iris contour image data, the feature points on the iris are extracted. These feature points can be the significant features of the iris texture, such as the edges and spots of the texture. Image processing techniques, such as edge detection and corner detection, can be used, or deep learning models can be used for feature extraction. For example, the Canny edge detection algorithm is used to extract the edge information in the iris contour image, and the edge points are used as iris feature points; or a feature point detection network based on CNN is used to extract iris feature points.
[0151] Step S262: Construct an eyeball model based on the pupil center positioning data and the iris feature point data to obtain an eyeball model;
[0152] In an embodiment of the present invention, a three-dimensional model of the eyeball is constructed based on pupil center localization data and iris feature point data. For example, a sphere model is used to approximately represent the eyeball, with the pupil center as the center of the sphere and the pupil radius as the radius of the sphere; or an ellipsoid model is used to more accurately describe the shape of the eyeball, and the parameters of the ellipsoid model are fitted using iris feature points.
[0153] Step S263: Obtain the eyeball optical axis data direction of the eyeball model based on the pupil center localization data to obtain the eyeball optical axis direction data;
[0154] In an embodiment of the present invention, for the sphere model, the vector from the pupil center to the center of the sphere can be directly calculated as the eyeball optical axis direction vector. For the ellipsoid model, the optical axis direction vector needs to be calculated based on the parameters of the ellipsoid model and the pupil center position. For example, the pupil center coordinates can be first transformed into the local coordinate system of the ellipsoid model, and then the direction vector of the optical axis in the local coordinate system can be calculated according to the equation of the ellipsoid model, and finally the direction vector is transformed back to the world coordinate system.
[0155] Step S264: Simulate the line-of-sight intersection point for the eyeball optical axis direction data to obtain the line-of-sight intersection point data;
[0156] In an embodiment of the present invention, after obtaining the eyeball optical axis direction vectors of both eyes, the two optical axis direction vectors can be extended, and the intersection point of the two lines can be calculated as the line-of-sight intersection point. It should be noted that due to factors such as measurement errors, the two optical axis direction vectors do not intersect. In this case, methods such as the least squares method can be used to calculate the optimal intersection point.
[0157] Step S265: Generate the line-of-sight fixation point for the line-of-sight intersection point data to obtain the line-of-sight fixation point data.
[0158] In an embodiment of the present invention, according to the screen parameters of the virtual reality device and the user's head pose information, the three-dimensional coordinates of the line-of-sight intersection point need to be transformed into the screen coordinate system of the virtual reality device. Then, the three-dimensional coordinates of the line-of-sight intersection point are projected onto the screen plane of the virtual reality device to obtain the two-dimensional coordinates of the line-of-sight fixation point on the screen.
[0159] In this invention, representative feature points are extracted from the iris contour image in this step, such as iris texture, color change, etc. These feature points can be used to more precisely fit the eyeball model and improve the accuracy of gaze estimation; The construction of the eyeball model is an important step in simulating the structure and movement of the human eye. By combining the pupil center localization data and the iris feature point data, an accurate three-dimensional eyeball model can be established. This model can simulate the shape, size, and structure of the human eye, providing a basic framework for analyzing the optical axis and gaze movement of the eyeball; Obtaining the direction of the eyeball optical axis data is the key to understanding the gaze direction. Through the pupil center localization data, the direction vector of the eyeball optical axis can be calculated. This vector represents the piercing direction from the pupil center to the eyeball model, providing an important reference for simulating the gaze intersection point and finally determining the gaze fixation point; The gazes of the left and right eyes will intersect at a point in space, that is, the gaze intersection point, which represents the specific position that the user is looking at in the three-dimensional space. By simulating the process of the intersection of the binocular gazes, the user's gaze fixation point can be inferred more accurately; Using the calculated gaze intersection point position as the user's gaze fixation point, this data can be used in various interactions and applications in the system, such as gaze selection, fixation point rendering, eye movement analysis, etc., to enhance the immersion and interactivity of the experience.
[0160] Preferably, step S3 includes the following steps:
[0161] Step S31: Obtain head sensing data from the virtual reality device to get the head sensing data;
[0162] In the embodiment of this invention, by equipping high-precision sensors, such as gyroscopes, accelerometers, etc., to track the user's head movement. By accessing the API interface provided by the device, the original data stream of the head sensor can be obtained in real time, including information such as head rotation angle, acceleration, timestamp, etc.
[0163] Step S32: Analyze the head pose data of the head sensing data to get the head pose data;
[0164] In the embodiment of this invention, algorithms such as Kalman filtering and complementary filtering are used to denoise and fuse the head sensing data. For example, the gyroscope data and the accelerometer data are fused to obtain a more accurate head pose estimation. At the same time, the head pose data can be converted into different representation forms according to needs, such as converting Euler angles into rotation matrices.
[0165] Step S33: Analyze the head deflection characteristics of the head sensing data based on the head pose data to get the head deflection characteristic data;
[0166] In the embodiments of the present invention, by calculating the change rate of the head pose data in the time series, such as the head rotation angular velocity, angular acceleration, etc., as the head deflection feature, it is also possible to analyze the direction and amplitude of the head rotation, such as determining whether the head turns left or right, and whether the rotation angle is large or small, etc., and use this information as the head deflection feature.
[0167] Step S34: Calculate the deflected fixation point vector based on the head deflection feature data for the line-of-sight fixation point data to obtain the deflected fixation point vector data;
[0168] In the embodiments of the present invention, by establishing a corresponding mapping relationship according to the type and magnitude of the head deflection feature, the head deflection feature is converted into a line-of-sight deflection vector. For example, according to the magnitude and direction of the head rotation angular velocity, a two-dimensional vector proportional thereto can be calculated to represent the deflection direction and distance of the line of sight on the screen plane.
[0169] Step S35: Calculate the deflected coordinates for the fixation point coordinate data based on the deflected fixation point vector data to obtain the deflected fixation point coordinates.
[0170] In the embodiments of the present invention, by adding the original line-of-sight fixation point coordinates and the deflected fixation point vector, a new two-dimensional coordinate is obtained, representing the screen position that the user is about to fixate on. It should be noted that the calculation of the deflected fixation point coordinates needs to consider the screen parameters of the virtual reality device, such as the screen resolution, field of view angle, etc.
[0171] The head sensing data of the virtual reality device in the present invention is the basis for capturing the head movement and posture changes of the user. Through sensors built into the virtual reality device, such as gyroscopes, accelerometers, and magnetometers, the position, rotation, and movement information of the user's head in three-dimensional space can be obtained in real time, providing the necessary data support for subsequent analysis of head posture and deflection characteristics; the analysis of head posture data can help the system understand the current state of the user's head. Through the processing and analysis of the head sensing data, posture information such as the rotation, pitch, and deflection angles of the head in the virtual space can be obtained, providing a basic reference for subsequent deflection characteristic analysis and gaze tracking; the analysis of head deflection characteristics aims to identify and quantify the lateral deflection movement of the user's head. Through further processing of the head posture data, the deflection angle and speed of the head in the horizontal direction can be extracted, providing key input data for calculating the deflection gaze point vector and finally determining the deflection gaze point coordinates; according to the head deflection characteristics of the user, a deflection vector is calculated to correct the original gaze point direction to make it more consistent with the actual gaze direction of the user; by combining the head sensing data of the virtual reality device and the deflection gaze point vector, the new gaze point coordinates generated by the user's head deflection in the virtual space can be calculated. This coordinate data can be corresponded to the objects in the virtual environment to achieve more accurate user interaction and a more natural human-computer interaction experience.
[0172] Preferably, step S33 includes the following steps:
[0173] Step S331: Set a reference posture for the head posture data to obtain the head reference posture data;
[0174] In the embodiment of the present invention, by defining a standard head posture as the reference posture, such as the head being upright and the eyes looking straight ahead. This posture can be adjusted according to the actual application scenario. Then, the sensor is used to obtain the head posture data of the current user, including the rotation angles of the head around the X-axis, Y-axis, and Z-axis. The obtained head posture data is compared with the pre-defined reference posture data, and the offset of the current head posture relative to the reference posture is calculated, and this offset is recorded as the head reference posture data.
[0175] Step S332: Calculate the head deflection angle based on the head reference posture data for the head sensing data to obtain the head deflection angle;
[0176] In an embodiment of the present invention, by obtaining the head sensing data of the user in real time, including the rotation angles of the head around the X-axis, Y-axis, and Z-axis. Compare the head sensing data obtained in real time with the head reference posture data, and calculate the deflection angles of the head in each direction. For example, if the reference posture data is (0°, 0°, 0°), and the current head sensing data is (10°, -5°, 3°), then the head deflection angle is (10°, -5°, 3°).
[0177] Step S333: Analyze the deflection vector direction of the head deflection angle to obtain the deflection vector direction data;
[0178] In an embodiment of the present invention, by converting the calculated head deflection angle into a deflection vector in three-dimensional space. This can be achieved by converting the deflection angle into a direction vector in the spherical coordinate system. The deflection vector direction data can be represented as a unit vector, such as (0.87, -0.5, 0.26), which is used to represent the direction of the head deflection.
[0179] Step S334: Calculate the deflection vector amplitude of the head deflection angle to obtain the deflection vector amplitude data;
[0180] In an embodiment of the present invention, by calculating the amplitude of the head deflection according to the calculated head deflection angle. For example, the Euclidean distance formula can be used to calculate the magnitude of the head deflection angle vector, and a scalar value is obtained to represent the degree of head deflection, such as 11.18°.
[0181] Step S335: Fit the deflection vector direction data and the deflection vector amplitude data to obtain the head deflection feature data.
[0182] In an embodiment of the present invention, by integrating the obtained deflection vector direction data and the deflection vector amplitude data, a multi-dimensional feature vector is formed. For example, a four-dimensional vector (0.87, -0.5, 0.26, 11.18) can be used to represent the current head deflection feature data, where the first three elements represent the direction and the last element represents the amplitude.
[0183] By setting the head reference pose, the present invention can provide a reference basis for subsequent calculation and analysis of the deflection angle. By establishing a standard or initial head pose as a reference, the degree of deflection of the user's head can be calculated and quantified more accurately. This reference pose is usually defined as the pose when the user looks straight ahead, providing a neutral reference point for subsequent analysis; The calculation of the head deflection angle is the basis for understanding the lateral movement of the user's head. By comparing the head sensing data with the head reference pose data, the deflection angle of the user's head in the horizontal direction can be calculated. This angle represents the degree of deviation of the head from the reference pose, providing a quantitative index for subsequent deflection vector analysis and simulation; By analyzing the head deflection angle, the direction vector of the deflection can be determined. This direction vector indicates whether the head deflection is in the left, right, up, down or other specific directions, providing important input information for simulating and simulating the user's line-of-sight movement; By analyzing the head deflection angle, the magnitude of the deflection vector can be calculated, that is, the distance or angle size of the deflection. This magnitude data represents the intensity of the head deflection, providing parameters for simulating and analyzing the user's visual field range and head movement amplitude; Feature data fitting is the process of integrating the direction and magnitude data of the deflection vector. By mathematically fitting the direction data and magnitude data, a more complete head deflection feature model can be established. This model combines the deflection direction and magnitude information, can more accurately represent the user's head deflection behavior, and provides reliable input data for subsequent line-of-sight tracking, virtual space interaction or other related applications.
[0184] Preferably, step S4 includes the following steps:
[0185] Step S41: Perform coordinate grid mapping on the virtual space model to obtain a virtual space coordinate grid;
[0186] Step S42: Based on the coordinates of the deflected fixation point, perform spatial positioning of the deflected fixation point on the virtual space coordinate grid to obtain spatial deflected fixation point data;
[0187] Step S43: Obtain virtual visual area data by acquiring the virtual visual area from the spatial deflected fixation point data;
[0188] Step S44: Based on the virtual visual area data, perform fixation audio area positioning on the virtual audio space to obtain a fixation virtual audio area;
[0189] Step S45: Optimize the area audio for the virtual visual area data and the fixation virtual audio area to obtain area audio optimization data.
[0190] As an embodiment of the present invention, refer to Figure 2 shown, for Figure 1 the detailed step flow diagram of step S4 in
[0191] Step S41: Perform coordinate grid mapping on the virtual space model to obtain a virtual space coordinate grid;
[0192] In the embodiment of the present invention, the virtual space model is loaded into the system. Then, according to the preset precision, the virtual space model is divided into a series of regular or irregular coordinate grids. Each grid corresponds to a specific area in the virtual space and is assigned a unique coordinate identifier. The density of the grids can be adjusted according to requirements. For example, for areas that require precise positioning, denser grids can be used.
[0193] Step S42: Perform deflection fixation point spatial positioning on the virtual space coordinate grid based on the deflection fixation point coordinates to obtain spatial deflection fixation point data;
[0194] In the embodiment of the present invention, by obtaining the deflection fixation point coordinates of the user, and then mapping the coordinates into the established virtual space coordinate grid, the grid containing the coordinates is found. This grid is the position where the deflection fixation point is located in the virtual space, and the coordinate identifier of the grid and other relevant information, such as the distance and direction from adjacent grids, are recorded as spatial deflection fixation point data.
[0195] Step S43: Obtain virtual visual area data by performing virtual visual area acquisition on the spatial deflection fixation point data;
[0196] In the embodiment of the present invention, the virtual visual area of the user is determined according to the obtained deflection fixation point spatial positioning information. This area can be simulated according to factors such as the user's field of view and deflection angle. For example, a spherical area centered on the deflection fixation point with the user's field of view as the radius can be defined as the virtual visual area. In addition, according to the actual application scenario, the shape and size of the virtual visual area can be adjusted. For example, simulate the field of view when the user wears a virtual reality device. Finally, all the coordinate grids and relevant information within the virtual visual area are recorded as virtual visual area data.
[0197] Step S44: Perform fixation audio area positioning on the virtual audio space based on the virtual visual area data to obtain a fixation virtual audio area;
[0198] In the embodiment of the present invention, the corresponding fixation audio area is located in the virtual audio space according to the obtained virtual visual area data. For example, the projection area of the virtual visual area in the virtual audio space can be defined as the fixation virtual audio area. This area contains the source of the sound that the user currently hears, as well as the target area for audio optimization.
[0199] Step S45: Perform regional audio optimization on the virtual visual area data and the fixation virtual audio area to obtain regional audio optimization data.
[0200] In an embodiment of the present invention, the virtual audio is optimized according to the obtained virtual visual area data and the gazed virtual audio area. For example, the sound clarity and volume within the gazed virtual audio area can be increased, the sound in other areas can be reduced, or the spatial orientation of the sound can be adjusted according to the deflected gaze point of the user, thereby enhancing the user's immersion and auditory experience. Finally, information such as the optimized audio parameters and area division is recorded as area audio optimization data.
[0201] The present invention can establish a coordinate framework for a virtual space by performing coordinate grid mapping on the virtual space model. By creating a three-dimensional coordinate grid system, a unique coordinate value can be assigned to each point in the virtual space. This provides a basic framework for deflected gaze point positioning, virtual vision acquisition, and audio optimization; the deflected gaze point spatial positioning is the process of mapping the coordinates of the deflected gaze point onto the virtual space coordinate grid. Through this step, the exact position of the user's deflected line of sight focus in the virtual space can be determined, providing a basic reference for subsequent virtual vision and audio area acquisition; by analyzing the deflected gaze point data, the area that enters the user's field of view in the virtual space, i.e., the virtual visual area, can be determined. This area will become the visible part of the virtual environment for the user currently, providing basic input for gazed audio area positioning and area audio optimization; by analyzing the virtual visual area data, the virtual audio area that the user is currently gazing at can be determined. This step ensures that the audio processing focuses on the virtual space area that the user is interested in or pays attention to, enhancing the relevance and accuracy of the audio experience; area audio optimization aims to improve the quality of the audio experience of the user in the virtual space. Through comprehensive analysis of the virtual visual area data and the gazed audio area data, the audio effect can be optimized. For example, enhancing the audio clarity within the virtual visual area, increasing the volume of the virtual audio area, adding ambient sound field effects, etc., to make the audio experience more realistic and immersive.
[0202] Preferably, step S45 includes the following steps:
[0203] Step S451: Analyze the regional acoustic characteristics of the gazed virtual audio area to obtain regional acoustic characteristic data;
[0204] Step S452: Based on the regional acoustic characteristic data, perform spatial area sound field simulation on the virtual audio space to obtain spatial area sound field data;
[0205] Step S453: Obtain the gazed visual path from the virtual visual area data;
[0206] Step S454: Identify the path audio occluders for the gazed visual path to obtain path audio occluder data;
[0207] Step S455: Based on the path audio occluder data, perform regional audio simulation optimization on the spatial region sound field data to obtain regional audio optimization data.
[0208] As an embodiment of the present invention, refer to Figure 3 shown in Figure 2 the detailed step flow diagram of step S45 in
[0209] Step S451: Analyze the regional acoustic characteristics of the gaze virtual audio region to obtain regional acoustic characteristic data;
[0210] In an embodiment of the present invention, through the three-dimensional model of the target virtual audio region. This model should include information such as the geometric shape, size, and material of the region. Then, select a suitable acoustic simulation software. For example, import the constructed three-dimensional model into the software. Set the sound source parameters according to the actual situation, such as the sound source type, sound pressure level, directivity, etc., and the air properties in the region, such as temperature, humidity, etc. Finally, run the acoustic simulation software for ray tracing and sound field calculation to obtain the acoustic parameters at different positions in the region, such as the reverberation time RT60, clarity index C50, speech intelligibility STI, etc. These acoustic parameters constitute the regional acoustic characteristic data for subsequent spatial region sound field simulation.
[0211] Step S452: Based on the regional acoustic characteristic data, perform spatial region sound field simulation on the virtual audio space to obtain spatial region sound field data;
[0212] In an embodiment of the present invention, through the regional acoustic characteristic data, combined with the overall structure and sound source distribution of the virtual audio space, use acoustic simulation software to perform sound field simulation on the entire virtual audio space. During the simulation, it is necessary to consider the sound propagation and attenuation between different regions, as well as the reflection, absorption, and scattering effects of walls, floors, objects, etc. on sound. The acoustic effects in the real environment can be simulated by setting different acoustic material properties and geometric structures. After the simulation is completed, data such as the sound pressure level, frequency response, and sound field distribution at each position in the virtual audio space can be obtained. These data constitute the spatial region sound field data.
[0213] Step S453: Obtain the gaze visual path for the virtual visual region data to obtain the gaze visual path;
[0214] In the embodiments of the present invention, by using the eye movement tracking function of an eye movement tracking device or a virtual reality head-mounted display, the eye movement trajectory data of the user within the virtual visual area is recorded. According to the eye movement trajectory data, the objects or regions gazed at by the user in the virtual scene are extracted, and adjacent fixation points are connected to form a continuous path, namely the fixation visual path. In order to improve the accuracy and stability of the path, a smoothing algorithm can be used to preprocess the eye movement trajectory data, such as moving average filtering, Gaussian filtering, etc.
[0215] Step S454: Perform audio occluder recognition on the fixation visual path to obtain path audio occluder data;
[0216] In the embodiments of the present invention, by according to the obtained fixation visual path and combining with the three-dimensional model data of the virtual scene, it is judged whether there are objects blocking the user's line of sight on the path. For each gazed object, the connection line between it and the user's viewpoint is calculated, and it is judged whether this connection line intersects with other objects. If an intersection occurs, it is considered that the object is blocked, and the information of the occluding object is recorded, including the object ID, the degree of occlusion, etc. Finally, the path audio occluder data is obtained.
[0217] Step S455: Perform regional audio simulation optimization on the spatial region sound field data based on the path audio occluder data to obtain regional audio optimization data.
[0218] In the embodiments of the present invention, by according to the obtained path audio occluder data, the obtained spatial region sound field data is optimized. For the occluded sound sources, parameters such as sound pressure level and frequency response are adjusted according to the degree of occlusion to simulate the attenuation and change of the sound after being blocked. For example, the sound transmission coefficient can be calculated according to the material and thickness of the occluding object, so as to adjust the volume and timbre of the occluded sound source. In this way, the sound in the virtual audio space can be made more consistent with the user's visual perception, and finally the regional audio optimization data is obtained.
[0219] The present invention aims to understand and quantify the sound characteristics of a virtual audio region through regional acoustic characteristic analysis. By analyzing the acoustic reflection, absorption, diffusion, and other acoustic characteristics of this region, regional acoustic characteristic data can be obtained. These data include material characteristics, boundary effects, acoustic environment, etc., providing basic parameters for subsequent spatial region sound field simulation; By applying acoustic principles and regional acoustic characteristic data, the propagation, reflection, and attenuation effects of sound in a specific virtual audio region can be simulated and predicted. This step helps to establish a realistic spatial audio model, providing a basic framework for final audio optimization; By analyzing virtual visual region data, the visual path from the user's deflected fixation point to other points in the virtual visual region can be determined. This path represents the movement trajectory of the user's line of sight in the virtual space, providing an important reference for subsequent audio occluder recognition and optimization; Audio occluder recognition aims to identify obstacles in the virtual space that affect sound propagation and auditory experience. By analyzing the fixation visual path, audio occluders on the visual path can be identified, such as virtual walls, furniture, or other virtual objects. This step provides the necessary data for regional audio simulation optimization, ensuring that the propagation and reflection of sound conform to the virtual environment settings; Through the comprehensive analysis of these two factors, the regional audio effect can be optimized. For example, simulating the attenuation, reflection, or diffraction of sound when encountering an occluder, and applying an appropriate acoustic model to achieve a more realistic audio experience. This optimization step ensures that the audio presentation in the virtual space is consistent with the actual physical environment, enhancing the immersion and spatial sense in virtual reality or mixed reality applications.
[0220] The present invention also provides an audiovisual device, including a head-mounted display for performing the virtual reality interaction method as described above. The head-mounted display of this audiovisual device includes:
[0221] A virtual audio space construction module for inputting image data into the virtual reality device to obtain spatial image data; constructing a virtual space model from the spatial image data to obtain a virtual space model; constructing a virtual audio space from the virtual space model to obtain a virtual audio space;
[0222] A fixation point coordinate generation module for performing real-time face camera processing on the virtual reality device to obtain real-time camera stream data; performing face image detection on the real-time camera stream data to obtain photographed face image data; generating a fixation point from the photographed face image data to obtain line-of-sight fixation point data; mapping the line-of-sight fixation point data pair to a coordinate system to obtain fixation point coordinate data;
[0223] A deflection fixation point generation module, configured to obtain head sensing data from a virtual reality device, and acquire the head sensing data; calculate a deflection fixation point vector for the line-of-sight fixation point data based on the head sensing data, and obtain deflection fixation point vector data; calculate a deflected coordinate for the fixation point coordinate data based on the deflection fixation point vector data, and obtain a deflected fixation point coordinate.
[0224] A regional audio optimization module, configured to obtain a virtual visual area of a virtual space model based on the deflected fixation point coordinates, and acquire virtual visual area data; perform fixation audio area localization on a virtual audio space based on the virtual visual area data, and obtain a fixation virtual audio area; perform regional audio optimization on the virtual visual area data and the fixation virtual audio area, and obtain regional audio optimization data.
[0225] In summary, the present invention provides a virtual reality interaction method and an audiovisual device. The virtual reality interaction method and the audiovisual device are composed of a virtual audio space construction module, a fixation point coordinate generation module, a deflection fixation point generation module, and a regional audio optimization module, and can implement any virtual reality interaction method described in the present invention. It is used to jointly realize any virtual reality interaction method through the operations between computer programs running on each module, and the internal structure of the system cooperates with each other. In this way, it can greatly reduce repetitive work and manpower investment, and can quickly and effectively provide a more accurate and efficient virtual reality interaction process, thereby simplifying the virtual reality interaction method and the audiovisual device.
[0226] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art. The general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A virtual reality interaction method, characterized in that It includes the following steps: Step S1: Input image data into the virtual reality device to obtain spatial image data; construct a virtual space model from the spatial image data; construct a virtual audio space from the virtual space model; Step S1 includes the following steps: Step S11: Input image data into the virtual reality device to obtain spatial image data; Step S12: Mark the spatial point cloud data of the spatial image data to obtain spatial image point cloud data; Step S13: Register the spatial image point cloud data to obtain spatial point cloud registration data; Step S14: Extract spatial features from the spatial image point cloud data to obtain spatial point cloud feature data; Step S15: Segment the spatial entities from the spatial point cloud feature data to obtain spatial entity data; Step S16: Construct a virtual space model based on the spatial entity data and the spatial point cloud registration data to obtain a virtual space model; Step S17: Construct a virtual audio space from the virtual space model. Step S17 includes the following steps: Step S171: Set the sound source of the virtual space model to obtain sound source setting data; Step S172: Aggregate the sound source data of the sound source setting data to obtain a sound source data set; Step S173: Analyze the sound propagation path of the sound source data set based on the virtual space model to obtain sound propagation path data; Step S174: Set the propagation path medium for the virtual space model based on the sound propagation path data to obtain propagation path medium data; Step S175: Analyze the sound reflection of the propagation path medium data based on the sound source data set and the sound propagation path data to obtain sound reflection data; Step S176: Analyze the sound obstacle of the propagation path medium data based on the sound source data set and the sound propagation path data to obtain sound obstacle data; Step S177: Fit the audio data space of the virtual space model based on the sound reflection data and the sound obstacle data to obtain a virtual audio space; Step S2: Perform real-time face camera processing on the virtual reality device to obtain real-time photography stream data; detect the face image from the real-time photography stream data to obtain photographed face image data; generate a fixation point from the photographed face image data to obtain line-of-sight fixation point data; map the line-of-sight fixation point data to a coordinate system to obtain fixation point coordinate data; Step S3: Obtain head sensing data from the virtual reality device; calculate the deflected fixation point vector from the line-of-sight fixation point data based on the head sensing data to obtain deflected fixation point vector data; calculate the deflected coordinates from the fixation point coordinate data based on the deflected fixation point vector data to obtain deflected fixation point coordinates; Step S4: Obtain virtual visual area data by acquiring the virtual visual area based on the deflected fixation point coordinates for the virtual space model; perform fixation audio area localization on the virtual audio space based on the virtual visual area data to obtain the fixation virtual audio area; perform area audio optimization on the virtual visual area data and the fixation virtual audio area to obtain area audio optimization data; Step S4 includes the following steps: Step S41: Perform coordinate grid mapping on the virtual space model to obtain the virtual space coordinate grid; Step S42: Perform deflected fixation point space localization on the virtual space coordinate grid based on the deflected fixation point coordinates to obtain space deflected fixation point data; Step S43: Obtain virtual visual area data by acquiring the virtual visual area for the space deflected fixation point data; Step S44: Perform fixation audio area localization on the virtual audio space based on the virtual visual area data to obtain the fixation virtual audio area; Step S45: Perform area audio optimization on the virtual visual area data and the fixation virtual audio area to obtain area audio optimization data; Step S45 includes the following steps: Step S451: Analyze the regional acoustic characteristics of the fixation virtual audio area to obtain regional acoustic characteristic data; Step S452: Perform spatial area sound field simulation on the virtual audio space based on the regional acoustic characteristic data to obtain spatial area sound field data; Step S453: Obtain the fixation visual path by acquiring the fixation visual path for the virtual visual area data; Step S454: Identify audio occluders for the fixation visual path to obtain path audio occluder data; Step S455: Perform regional audio simulation optimization on the spatial area sound field data based on the path audio occluder data to obtain area audio optimization data.
2. The virtual reality interaction method according to claim 1, wherein Step S2 includes the following steps: Step S21: Perform real-time face camera processing on the virtual reality device to obtain real-time photography stream data; Step S22: Perform face image detection on the real-time photography stream data to obtain photographed face image data; Step S23: Extract the eye region from the photographed face image data to obtain photographed eye region data; Step S24: Locate the pupil center for the photographed eye region data to obtain pupil center localization data; Step S25: Extract the iris contour from the photographed eye region data to obtain iris contour image data; Step S26: Generate a fixation point for the pupil center localization data and the iris contour image data to obtain line-of-sight fixation point data; Step S27: Perform coordinate system mapping on the line-of-sight fixation point data to obtain fixation point coordinate data.
3. The virtual reality interaction method according to claim 2, wherein Step S26 includes the following steps: Step S261: Extract iris feature points from the iris contour image data to obtain iris feature point data; Step S262: Construct an eyeball model for the pupil center localization data and the iris feature point data to obtain the eyeball model; Step S263: Obtain the eyeball optical axis direction data by acquiring the direction of the eyeball optical axis data for the eyeball model based on the pupil center localization data; Step S264: Simulate the line-of-sight intersection point for the eyeball optical axis direction data to obtain line-of-sight intersection point data; Step S265: Generate a gaze point from the line-of-sight intersection point data to obtain gaze point data.
4. The virtual reality interaction method according to claim 1, characterized in that Step S3 includes the following steps: Step S31: Obtain head sensing data by performing head sensing data acquisition on the virtual reality device to obtain head sensing data; Step S32: Analyze the head pose data of the head sensing data to obtain head pose data; Step S33: Analyze the head deflection characteristics of the head sensing data based on the head pose data to obtain head deflection characteristic data; Step S34: Calculate the deflection gaze point vector based on the head deflection characteristic data for the gaze point data to obtain deflection gaze point vector data; Step S35: Calculate the deflection coordinates for the gaze point coordinate data based on the deflection gaze point vector data to obtain the deflection gaze point coordinates.
5. The virtual reality interaction method according to claim 4, wherein Step S33 includes the following steps: Step S331: Set a reference pose for the head pose data to obtain head reference pose data; Step S332: Calculate the head deflection angle of the head sensing data based on the head reference pose data to obtain the head deflection angle; Step S333: Analyze the direction of the deflection vector for the head deflection angle to obtain deflection vector direction data; Step S334: Calculate the magnitude of the deflection vector for the head deflection angle to obtain deflection vector magnitude data; Step S335: Fit the feature data of the deflection vector direction data and the deflection vector magnitude data to obtain head deflection characteristic data.
6. An audiovisual device, characterized in that, It includes a head-mounted display for performing the virtual reality interaction method as described in claim 1. The head-mounted display of this audio-visual device includes: A virtual audio space construction module for inputting image data to the virtual reality device to obtain spatial image data; constructing a virtual space model from the spatial image data to obtain a virtual space model; constructing a virtual audio space from the virtual space model to obtain a virtual audio space; A gaze point coordinate generation module for performing real-time face camera processing on the virtual reality device to obtain real-time photography stream data; detecting a face image from the real-time photography stream data to obtain photographed face image data; generating a gaze point from the photographed face image data to obtain gaze point data; mapping the gaze point data to a coordinate system to obtain gaze point coordinate data; A deflected gaze point generation module for obtaining head sensing data by performing head sensing data acquisition on the virtual reality device; calculating a deflected gaze point vector based on the head sensing data for the gaze point data to obtain deflected gaze point vector data; Calculating deflected coordinates for the gaze point coordinate data based on the deflected gaze point vector data to obtain deflected gaze point coordinates; A regional audio optimization module for obtaining virtual visual region data by obtaining a virtual visual region from the virtual space model based on the deflected gaze point coordinates; Locating a gaze audio region in the virtual audio space based on the virtual visual region data to obtain a gazed virtual audio region; Performing regional audio optimization on the virtual visual region data and the gazed virtual audio region to obtain regional audio optimization data.
Citation Information
Patent Citations
Information processing device, information processing system, and information processing method
CN108885799A
Apparatus and associated methods for presentation of captured spatial audio content
CN111512371A
Modify audio based on physiological observations
CN113906368A
Sound effect rendering method and device, electronic equipment and readable storage medium
CN115412832A