Three-dimensional picture display method and system based on indoor space positioning
By using multi-camera systems and deep learning models in the digital twin exhibition hall, high-precision indoor space positioning and three-dimensional picture display are solved, and the problems of insufficient positioning accuracy and limited picture adaptation capabilities in the existing technology are solved, improving the user experience.
Patent Information
- Application Number
- CN202510143221.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-02-10
AI Technical Summary
The existing digital twin exhibition hall faces the problems of insufficient positioning accuracy, limited adaptability of multi-view picture, and poor real-time distortion correction results in complex interactive scenarios, resulting in a bottleneck in user experience.
By setting up two cameras at different angles, efficient detection and tracking of the head and facial information of the target user, enhancing the robustness of character position information collection in complex environments. Using coordinate matching and deep learning models, the target user's position information and pose information are generated, and the positioning server and the three-dimensional rendering engine are combined to correct the distortion of the three-dimensional picture in real time.
It improves the accuracy of indoor space positioning and the accuracy of three-dimensional picture display, enhances the user's visual experience and interaction effects, and solves the problems of insufficient positioning accuracy and limited picture adaptability.
Smart Images

Figure CN120147584A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical fields of indoor space positioning and computer vision processing, and particularly to a three-dimensional picture display method and system based on indoor space positioning. Background Art
[0002] As an interactive platform for the deep integration of the physical world and the virtual world, the digital twin exhibition hall provides users with an immersive and interactive exhibition experience by real-time mapping physical environment data to the virtual space. Its core depends on high-precision space positioning technology and the dynamic adaptation ability of three-dimensional pictures to ensure that users obtain distortion-free visual feedback during movement. However, there are still significant bottlenecks in the positioning accuracy and picture adaptation in multi-user dynamic scenarios in the existing technology.
[0003] Currently, indoor space positioning is mostly based on the binocular stereo disparity principle, and the three-dimensional coordinates of the target are calculated by simulating the human eye disparity with two cameras. However, traditional methods face two major challenges in complex exhibition hall environments: firstly, static calibration modes (such as checkerboard calibration) are difficult to adapt to multi-view and large-scale scenarios, resulting in the accumulation of camera parameter calibration errors; secondly, in multi-user scenarios, problems such as head occlusion and non-frontal orientation are likely to cause feature point matching failures and insufficient positioning robustness. In addition, existing systems usually rely on a single visual data source (such as head or face detection) and are difficult to balance accuracy and real-time performance in dynamic environments. Therefore, it is impossible to support the dynamic adjustment and optimization of the distortion of the three-dimensional pictures in the exhibition hall based on accurate space positioning, which greatly affects the display effect of the digital twin pictures in the exhibition hall.
[0004] In summary, the existing digital twin exhibition halls face core problems such as insufficient positioning accuracy, limited multi-view picture adaptation ability, and poor real-time distortion correction effect in complex interaction scenarios, and there is an urgent need for a technical solution that can fuse multi-source perception data and achieve high-precision dynamic positioning to break through the existing user experience bottleneck. Summary of the Invention
[0005] In view of the above technical problems, this application provides a three-dimensional picture display method and system based on indoor space positioning, which can accurately position users in the indoor environment while correcting the distortion in the three-dimensional pictures in real time, thereby improving the user's visual experience and interaction effect.
[0006] In a first aspect, an embodiment of this application provides a three-dimensional picture display method based on indoor space positioning, including:
[0007] Obtaining a plurality of head position pixel coordinates through a first camera and obtaining a plurality of face position pixel coordinates through a second camera, where the first camera is located at the top of the digital twin exhibition hall, and the second camera is located in front of the display screen of the digital twin exhibition hall;
[0008] Coordinate matching is performed based on matrix parameters, each of the head position pixel coordinates, each of the face position pixel coordinates, and face orientation information to generate position information of the target user. The position information includes three-dimensional coordinates and pose information. The matrix parameters are obtained by calibrating the first camera and the second camera;
[0009] Convert the position information to a preset Cartesian coordinate system to generate corresponding coordinate conversion position information;
[0010] Input the three-dimensional model data of the actual space environment and the coordinate conversion position information into a preset positioning server, so that the positioning server generates the best viewing point position of the target user;
[0011] Input the coordinate conversion position information into a preset three-dimensional rendering engine, so that the three-dimensional rendering engine generates an initial three-dimensional image;
[0012] According to the best viewing point position and the coordinate conversion position information, adjust the display position and display angle of the initial three-dimensional image through a position deviation correction algorithm, generate a final three-dimensional image and display it on the display screen.
[0013] The embodiment of the present application provides a three-dimensional image display method based on indoor space positioning. By setting two cameras at different angles, efficient detection and tracking of the head and face information of the target user are realized, the robustness of collecting the position information of people in a complex environment is enhanced, and the accuracy of subsequent indoor space positioning is improved; the position information of the target user is determined from multiple coordinates through coordinate matching, providing data support for subsequent three-dimensional image display according to the target user; according to the position information after coordinate conversion, the best viewing point position of the target user is generated by the positioning server, the initial three-dimensional image is generated by the three-dimensional rendering engine, and finally the best viewing point position is combined with the coordinate conversion position information of the target user, and the angle of the initial three-dimensional image is adjusted through the position deviation correction algorithm, generating a final three-dimensional image and displaying it on the display screen, realizing the correction of the distortion of the initial three-dimensional image, thereby enhancing the visual experience and interaction effect of the user in the digital twin exhibition hall environment.
[0014] Further, the obtaining of a plurality of head position pixel coordinates by the first camera and the obtaining of a plurality of face position pixel coordinates by the second camera include:
[0015] Obtain a first video stream and a second video stream through the first camera and the second camera respectively;
[0016] Input the first video stream into a preset head detection model, so that the head detection model detects the head information of the person in the first video stream, and then generates a plurality of head position pixel coordinates;
[0017] Input the second video stream into a preset face recognition model, so that the face recognition model can recognize the facial information of the people in the second video stream, and then generate a number of facial position pixel coordinates and facial orientation information;
[0018] Among them, both the head detection model and the face recognition model are deep learning models.
[0019] The embodiment of the present application provides a method for obtaining pixel coordinates. The first video stream and the second video stream are respectively obtained through different cameras, and then the corresponding deep learning models are combined to detect and recognize the first video stream and the second video stream, and a number of head position pixel coordinates, a number of facial position pixel coordinates and facial orientation information are respectively extracted, realizing the preliminary positioning of multiple users in the exhibition hall, providing a data basis for the generation and angle adjustment of subsequent three-dimensional pictures, making the generated three-dimensional pictures accurately match the viewing angles of users, and improving the visual experience and interaction effect of users.
[0020] Further, the coordinate matching according to the matrix parameters, each of the head position pixel coordinates, each of the facial position pixel coordinates and the facial orientation information to generate the position information of the target user includes:
[0021] Substitute each of the facial position pixel coordinates into a preset coordinate matching equation to obtain each corresponding camera epipolar line, and the coordinate matching equation is constructed according to the fundamental matrix in the matrix parameters;
[0022] According to the distances between each of the camera epipolar lines and each of the head position pixel coordinates, determine a number of matching head position pixel coordinates corresponding to each of the camera epipolar lines, and then perform coordinate matching on each of the facial position pixel coordinates and each of the matching head position pixel coordinates to generate each corresponding coordinate matching point pair;
[0023] Calculate each corresponding three-dimensional coordinate from each of the coordinate matching point pairs and the matrix parameters;
[0024] According to the corresponding relationship between each of the facial position pixel coordinates and the facial orientation information, determine each facial orientation information corresponding to each of the three-dimensional coordinates, and then generate each corresponding pose information;
[0025] Combine each of the three-dimensional coordinates and each corresponding pose information to generate each corresponding position information;
[0026] Determine the position information of the target user according to the distances between each of the position information and the second camera.
[0027] The embodiments of the present application provide a coordinate matching method. Considering that the field of view of the second camera in front of the screen is smaller than that of the second camera at the top, and the visitors may not face the screen or be blocked, resulting in the number of recognizable faces usually being less than the number of recognized faces at the top of the head. Therefore, the embodiments of the present application substitute the pixel coordinates of the facial position into the coordinate matching equation to obtain the corresponding camera epipolar lines for each. Then, according to each camera epipolar line, the corresponding pixel coordinates of the head position are matched. The specific matching rule is to match according to the distance between the camera epipolar line and each pixel coordinate of the head position. For example, for a certain camera epipolar line, the pixel coordinates of the head position with a distance less than the set threshold from the epipolar line are selected for matching, and then the corresponding coordinate matching point pairs are established. Then, according to each coordinate matching point pair and the matrix parameters, the corresponding three-dimensional coordinates are calculated for each, realizing the conversion from two-dimensional coordinates to three-dimensional coordinates, and generating the corresponding pose information in combination with the facial orientation information, determining the positions and facial orientations of each user in the digital twin exhibition hall. Further, when there are multiple position information, the target user is determined according to the distance between each position information and the second camera. For example, the user closest to the second camera is selected as the target user. This strategy preferentially meets the needs of visitors facing the display plane and closer to the screen, conforming to the typical standing habits of key visitors in the exhibition hall, thereby enhancing the visual experience and interaction effect of users in the digital twin exhibition hall environment.
[0028] In a possible implementation manner, the matrix parameters are obtained by calibrating the first camera and the second camera based on the red sphere calibration method, including:
[0029] Move the red sphere to different positions in the digital twin exhibition hall;
[0030] Each time the red sphere is moved, the first pixel coordinates and the second pixel coordinates are respectively obtained through the first camera and the second camera, and then the corresponding sphere coordinate matching point pairs are constructed;
[0031] The fundamental matrix between the first camera and the second camera is calculated according to a number of sphere coordinate matching point pairs;
[0032] The essential matrix between the first camera and the second camera is calculated according to the fundamental matrix and the camera intrinsic matrix of the first camera and the second camera;
[0033] Combining the fundamental matrix and the essential matrix, the matrix parameters are constructed and obtained.
[0034] An embodiment of the present application provides a calibration method based on a red sphere. By simulating the position of a human head, feature matching point pairs of a camera are dynamically generated, and then the camera is calibrated. Since the embodiment of the present application uses a first camera and a second camera that are perpendicular to each other and have a large viewing angle deviation, if the traditional checkerboard calibration method is used, a too large calibration board size is required and it is difficult to find a suitable pitch angle that simultaneously meets the viewing requirements of the two cameras. Therefore, a red sphere similar in size to a human head is selected as the calibration tool. This method makes the calibration process unaffected by the camera direction and the position of the sphere center can be measured accurately. At the same time, the reflective characteristics and bright red color of the sphere also facilitate accurate identification by the camera, improving the calibration accuracy of the camera, and further improving the accuracy of subsequent indoor space positioning and three-dimensional scene generation.
[0035] Further, the positioning server generates the optimal viewing point position of the target user, including:
[0036] According to the coordinate conversion position information, the three-dimensional model data, and the three-dimensional scene content currently displayed on the display screen, the optimal alignment point of the line of sight of the target user is calculated through a geometric optimization algorithm, and the geometric optimization algorithm is a minimum viewing point deviation algorithm or a maximum visible area algorithm;
[0037] Taking the optimal alignment point of the line of sight of the target user as the optimal viewing point position.
[0038] In a possible implementation manner, the three-dimensional rendering engine generates an initial three-dimensional scene, including:
[0039] Inputting the coordinate conversion position information into a preset three-dimensional video dynamic projection model, so that the three-dimensional video dynamic projection model calculates a projection matrix according to the coordinate conversion position information and the three-dimensional scene content currently displayed on the display screen;
[0040] Inputting the projection matrix into a preset three-dimensional video dynamic rendering adjustment model, so that the three-dimensional video dynamic rendering adjustment model renders the three-dimensional scene content according to the projection matrix to generate the initial three-dimensional scene.
[0041] The embodiment of the present application presets a three-dimensional video dynamic projection model and a three-dimensional video dynamic rendering adjustment model. Among them, the three-dimensional video dynamic projection model is used to calculate the video screen display angle suitable for the current viewer according to the coordinate conversion position information and the three-dimensional scene content, and generate a corresponding projection matrix (rendering parameters of a virtual camera); the three-dimensional video dynamic rendering adjustment model applies the projection matrix to the rendering engine to generate an initial three-dimensional scene, ensuring that when the user observes the screen from different positions, the three-dimensional content displayed on the screen can be dynamically adjusted according to the user's viewing angle.
[0042] In a possible implementation manner, when generating the initial three-dimensional picture, the undistort function in OpenCV is used to correct the image distortion of the initial three-dimensional picture.
[0043] Since the initial three-dimensional picture is generated according to the rendering parameters of the virtual camera, there will be certain degrees of image distortion and projection deformation problems. Therefore, in the embodiment of the present application, when generating the initial three-dimensional picture, the undistort function in OpenCV is further used to correct the image distortion of the initial three-dimensional picture, eliminate the above-mentioned image distortion problems, and improve the user's visual experience.
[0044] In a second aspect, correspondingly, the embodiment of the present application provides a three-dimensional picture display system based on indoor space positioning, including a coordinate acquisition module, a position information generation module, a coordinate conversion module, a positioning module, a rendering module, and a display module;
[0045] Among them, the coordinate acquisition module is used to acquire a plurality of head position pixel coordinates through a first camera and a plurality of face position pixel coordinates through a second camera, where the first camera is located at the top of the digital twin exhibition hall, and the second camera is located in front of the display screen of the digital twin exhibition hall;
[0046] The position information generation module is used to perform coordinate matching according to matrix parameters, each of the head position pixel coordinates, each of the face position pixel coordinates, and face orientation information, and generate the position information of the target user. The position information includes three-dimensional coordinates and attitude information, and the matrix parameters are obtained by calibrating the first camera and the second camera;
[0047] The coordinate conversion module is used to convert the position information into a preset Cartesian coordinate system to generate corresponding coordinate conversion position information;
[0048] The positioning module is used to input the three-dimensional model data of the actual space environment and the coordinate conversion position information into a preset positioning server, so that the positioning server generates the best viewing point position of the target user;
[0049] The rendering module is used to input the coordinate conversion position information into a preset three-dimensional rendering engine, so that the three-dimensional rendering engine generates an initial three-dimensional picture;
[0050] The display module is used to adjust the display position and display angle of the initial three-dimensional picture through a position deviation correction algorithm according to the best viewing point position and the coordinate conversion position information, generate a final three-dimensional picture, and display it on the display screen.
[0051] Further, the coordinate acquisition module includes a video stream acquisition unit, a head detection unit and a face recognition unit;
[0052] The video stream acquisition unit is used to acquire the first video stream and the second video stream through the first camera and the second camera respectively;
[0053] The head detection unit is used to input the first video stream into a preset head detection model, so that the head detection model detects the head information of the person in the first video stream, and then generates a number of head position pixel coordinates;
[0054] The face recognition unit is used to input the second video stream into a preset face recognition model so that the face recognition model recognizes the facial information of the person in the second video stream, and then generates a number of facial position pixel coordinates and facial orientation information;
[0055] Among them, the head detection model and the face recognition model are both deep learning models.
[0056] Furthermore, the position information generation module performs coordinate matching according to the matrix parameters, the pixel coordinates of each head position, the pixel coordinates of each face position and the face orientation information to generate the position information of the target user, including:
[0057] Substituting the pixel coordinates of each facial position into a preset coordinate matching equation to obtain each corresponding camera epipolar line, wherein the coordinate matching equation is constructed based on the basic matrix in the matrix parameters;
[0058] Determine a plurality of matching head position pixel coordinates corresponding to each camera epipolar line according to the distance between each camera epipolar line and each head position pixel coordinate, and then coordinate match each facial position pixel coordinate with each matching head position pixel coordinate to generate corresponding coordinate matching point pairs;
[0059] Match each of the coordinate point pairs and the matrix parameters to calculate and obtain the corresponding three-dimensional coordinates;
[0060] Determine the facial orientation information corresponding to each of the three-dimensional coordinates according to the correspondence between the pixel coordinates of each facial position and the facial orientation information, and then generate the corresponding posture information;
[0061] Combining each of the three-dimensional coordinates with the corresponding each of the posture information to generate corresponding each of the position information;
[0062] The location information of the target user is determined according to the distance between each piece of location information and the second camera. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] Figure 1 : A flow chart of a three-dimensional image display method based on indoor space positioning provided in an embodiment of the present application.
[0064] Figure 2 : Schematic diagram of video stream acquisition of the second camera in a three-dimensional image display method based on indoor space positioning provided in an embodiment of the present application.
[0065] Figure 3 : A schematic diagram of displaying a three-dimensional image in a three-dimensional image display method based on indoor space positioning provided in an embodiment of the present application.
[0066] Figure 4 : A detailed flow chart of a three-dimensional image display method based on indoor space positioning provided in an embodiment of the present application.
[0067] Figure 5 : A structural diagram of a three-dimensional image display system based on indoor space positioning provided in an embodiment of the present application. DETAILED DESCRIPTION
[0068] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0069] It should be noted that the step numbers in the text are only for the convenience of explanation of the specific embodiments and do not serve to limit the order in which the steps are executed. In the description of this application, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features.
[0070] Embodiment 1:
[0071] like Figure 1 As shown, the first embodiment provides a three-dimensional image display method based on indoor space positioning, including steps S1-S6:
[0072] Step S1, obtaining a number of head position pixel coordinates through a first camera, and obtaining a number of face position pixel coordinates through a second camera, wherein the first camera is located at the top of the digital twin exhibition hall, and the second camera is located in front of the display screen of the digital twin exhibition hall;
[0073] Step S2: Perform coordinate matching based on the matrix parameters, the pixel coordinates of each of the head positions, the pixel coordinates of each of the face positions, and the face orientation information to generate the position information of the target user, where the position information includes three-dimensional coordinates and pose information, and the matrix parameters are obtained by calibrating the first camera and the second camera;
[0074] Step S3: Convert the position information to a preset Cartesian coordinate system to generate corresponding coordinate conversion position information;
[0075] Step S4: Input the three-dimensional model data of the actual space environment and the coordinate conversion position information into a preset positioning server so that the positioning server generates the optimal viewing point position of the target user;
[0076] Step S5: Input the coordinate conversion position information into a preset three-dimensional rendering engine so that the three-dimensional rendering engine generates an initial three-dimensional image;
[0077] Step S6: According to the optimal viewing point position and the coordinate conversion position information, adjust the display position and display angle of the initial three-dimensional image through a position correction algorithm to generate a final three-dimensional image and display it on the display screen.
[0078] The embodiment of the present application provides a three-dimensional image display method based on indoor space positioning. By setting two cameras at different angles, efficient detection and tracking of the head and face information of the target user are realized, the robustness of collecting the position information of people in a complex environment is enhanced, and the accuracy of subsequent indoor space positioning is improved; the position information of the target user is determined from multiple coordinates through coordinate matching, providing data support for subsequent three-dimensional image display according to the target user; according to the position information after coordinate conversion, the optimal viewing point position of the target user is generated by the positioning server, the initial three-dimensional image is generated by the three-dimensional rendering engine, and finally, the optimal viewing point position is combined with the coordinate conversion position information of the target user, and the angle of the initial three-dimensional image is adjusted through the position correction algorithm to generate a final three-dimensional image and display it on the display screen, realizing the correction of the distortion of the initial three-dimensional image, thereby enhancing the visual experience and interaction effect of the user in the digital twin exhibition hall environment.
[0079] Further, in step S1, the obtaining a plurality of head position pixel coordinates through the first camera and obtaining a plurality of face position pixel coordinates through the second camera includes:
[0080] Obtain a first video stream and a second video stream through the first camera and the second camera respectively;
[0081] Input the first video stream into a preset head detection model, so that the head detection model detects the head information of the people in the first video stream, and then generates a number of head position pixel coordinates;
[0082] Input the second video stream into a preset face recognition model, so that the face recognition model recognizes the face information of the people in the second video stream, and then generates a number of face position pixel coordinates and face orientation information;
[0083] Among them, both the head detection model and the face recognition model are deep learning models.
[0084] The embodiment of the present application provides a method for obtaining pixel coordinates. The first video stream and the second video stream are respectively obtained through different cameras, and then the corresponding deep learning models are combined to detect and recognize the first video stream and the second video stream, and a number of head position pixel coordinates, a number of face position pixel coordinates and face orientation information are respectively extracted, so as to realize the preliminary positioning of multiple users in the exhibition hall, provide a data basis for the generation and angle adjustment of the subsequent three-dimensional picture, make the generated three-dimensional picture accurately match the viewing angle of the user, and improve the visual experience and interaction effect of the user.
[0085] In a preferred embodiment, the first camera and the second camera are high-resolution cameras, which are used to capture continuous RGB-D video streams of people. As Figure 2 shown, the second camera is collecting the video stream. The head detection model can specifically be the Fully Convolutional Head Detector model, and the face recognition model can specifically be the MediaPipe Face Mesh model.
[0086] Further, in step S2, the coordinate matching according to the matrix parameters, each of the head position pixel coordinates, each of the face position pixel coordinates and the face orientation information to generate the position information of the target user includes:
[0087] Substitute each of the face position pixel coordinates into a preset coordinate matching equation to obtain each corresponding camera epipolar line, and the coordinate matching equation is constructed according to the fundamental matrix in the matrix parameters;
[0088] According to the distance between each of the camera epipolar lines and each of the head position pixel coordinates, determine a number of matching head position pixel coordinates corresponding to each of the camera epipolar lines, and then perform coordinate matching on each of the face position pixel coordinates and each of the matching head position pixel coordinates to generate each corresponding coordinate matching point pair;
[0089] Calculate each corresponding three-dimensional coordinate by using each of the coordinate matching point pairs and the matrix parameters;
[0090] Determine the facial orientation information corresponding to each of the three-dimensional coordinates according to the corresponding relationship between each of the facial position pixel coordinates and the facial orientation information, and then generate each corresponding pose information;
[0091] Generate each corresponding position information by combining each of the three-dimensional coordinates and the corresponding pose information;
[0092] Determine the position information of the target user according to the distance between each position information and the second camera.
[0093] The embodiment of the present application provides a coordinate matching method. Considering that the field of view of the second camera in front of the screen is smaller than that of the second camera at the top, and the visitors may not face the screen or be blocked, resulting in the number of recognizable faces being usually less than the number of recognized faces at the top of the head. Therefore, the embodiment of the present application substitutes the facial position pixel coordinates into the coordinate matching equation to obtain each corresponding camera epipolar line. Then, according to each camera epipolar line, each corresponding head position pixel coordinate is matched. The specific matching rule is to match according to the distance between the camera epipolar line and each head position pixel coordinate. For example, for a certain camera epipolar line, the head position pixel coordinates with a distance less than the set threshold from the epipolar line are selected for matching, and then the corresponding coordinate matching point pairs are established. Then, each corresponding three-dimensional coordinate is calculated by using each coordinate matching point pair and the matrix parameters, realizing the conversion from two-dimensional coordinates to three-dimensional coordinates, and generating the corresponding pose information in combination with the facial orientation information, thereby determining the positions and facial orientations of each user in the digital twin exhibition hall. Further, when there are multiple position information, the target user is determined according to the distance between each position information and the second camera. For example, the user closest to the second camera is selected as the target user. This strategy preferentially meets the needs of visitors facing the display plane and close to the screen, conforms to the typical standing position habits of key visitors in the exhibition hall, and thus improves the visual experience and interaction effect of users in the digital twin exhibition hall environment.
[0094] In a preferred embodiment, the coordinate matching equation is:
[0095] x ′T Fx = 0
[0096] where x ′ and x are respectively the pixel point coordinates (represented by homogeneous coordinates) corresponding to the two cameras, and F is the fundamental matrix in the matrix parameters.
[0097] Further, to prevent screen image jitter, when the person moves, the embodiments of the present application update the position information of each user only when the changes in each three-dimensional coordinate exceed a preset threshold, ensuring the stability of the display effect of the display screen.
[0098] In a possible implementation manner, in step S2, the matrix parameters are obtained by calibrating the first camera and the second camera based on the red sphere calibration method, including:
[0099] Move the red sphere to different positions in the digital twin exhibition hall;
[0100] Each time the red sphere is moved, obtain the first pixel coordinates and the second pixel coordinates through the first camera and the second camera respectively, and then construct the corresponding sphere coordinate matching point pairs;
[0101] Calculate the fundamental matrix between the first camera and the second camera according to a number of sphere coordinate matching point pairs;
[0102] Calculate the essential matrix between the first camera and the second camera according to the fundamental matrix and the camera intrinsic matrix of the first camera and the second camera;
[0103] Combine the fundamental matrix and the essential matrix to construct the matrix parameters.
[0104] The embodiments of the present application provide a calibration method based on a red sphere. By simulating the position of the human head, feature matching point pairs of the camera are dynamically generated, and then the camera is calibrated. Since the embodiments of the present application use the first camera and the second camera that are perpendicular to each other and have a large viewing angle deviation, if the traditional checkerboard calibration method is used, a too large calibration board size is required and it is difficult to find a suitable pitch angle that simultaneously meets the viewing requirements of the two cameras. Therefore, a red sphere similar in size to the human head is selected as the calibration tool. This method makes the calibration process not affected by the camera direction and the measurement of the sphere center position is accurate. At the same time, the reflective characteristics and bright red color of the sphere are also convenient for the camera to accurately identify, improving the calibration accuracy of the camera, and further improving the accuracy of subsequent indoor space positioning and three-dimensional image generation.
[0105] In a preferred embodiment, the specific process of calibrating the first camera and the second camera is as follows:
[0106] Move the red sphere to different positions in the room. After each move, simultaneously capture the images in the first camera at the top and the second camera in front of the screen, record the pixel coordinates of the sphere, select the center of the sphere as the feature point, and after multiple acquisitions, obtain a number of feature matching point pairs. The matching point pairs (x, x′) satisfy the following relationship:
[0107] x′T Fx = 0
[0108] Then, calculate the fundamental matrix and the essential matrix respectively:
[0109] The fundamental matrix (F) describes the geometric relationship between two cameras and is used to map points in one camera image to the epipolar lines in the other camera image. The calculation steps are as follows:
[0110] (1) Construct a linear equation system: Use at least 8 pairs of feature matching points (x, x'), and construct a homogeneous linear equation: AF = 0, where A is a matrix constructed by point coordinates.
[0111] (2) Singular value decomposition (SVD): Perform SVD decomposition on matrix A and solve for F by minimizing ||AF||.
[0112] (3) Constraint enforcement: To ensure that the rank of F is 2, perform SVD decomposition on the solved matrix F and set the minimum singular value to zero.
[0113] The essential matrix (E) further describes the relative pose (rotation and translation) between two cameras and is calculated through the relationship between the camera's internal parameter matrix and the fundamental matrix:
[0114] (1) Use the fundamental matrix F obtained from the previous steps and the known camera internal parameter matrices K1 and K2 to calculate the essential matrix E:
[0115]
[0116] (2) Standardize E. Perform SVD decomposition on E and enforce its singular values to satisfy the constraint: the first two singular values are equal, and the third is 0.
[0117] (3) Recover the relative motion parameters from the essential matrix E. Using SVD decomposition, 4 sets of possible rotation matrix R and translation vector t solutions can be obtained:
[0118]
[0119] (4) Select the correct combination of R and t from the 4 sets of solutions by verifying the constraint that the reconstructed 3D points are in front of both cameras.
[0120] In a preferred embodiment, in step S3, the SLAM algorithm is used to convert the three-dimensional coordinates and pose of the target person into a Cartesian coordinate system with the center of the lower edge of the display screen as the origin, and generate the corresponding coordinate transformation position information.
[0121] Further, in step S4, the positioning server generates the best viewing point position of the target user, including:
[0122] According to the coordinate conversion position information, the three-dimensional model data, and the three-dimensional scene content currently displayed on the display screen, calculate the best alignment point of the line of sight of the target user through a geometric optimization algorithm, where the geometric optimization algorithm is a minimum viewing point deviation algorithm or a maximum visible area algorithm;
[0123] Take the best alignment point of the line of sight of the target user as the best viewing point position.
[0124] In a possible implementation manner, in step S5, the three-dimensional rendering engine generates an initial three-dimensional picture, including:
[0125] Input the coordinate conversion position information into a preset three-dimensional video dynamic projection model, so that the three-dimensional video dynamic projection model calculates a projection matrix according to the coordinate conversion position information and the three-dimensional scene content currently displayed on the display screen;
[0126] Input the projection matrix into a preset three-dimensional video dynamic rendering adjustment model, so that the three-dimensional video dynamic rendering adjustment model renders the three-dimensional scene content according to the projection matrix to generate the initial three-dimensional picture.
[0127] The embodiment of the present application presets a three-dimensional video dynamic projection model and a three-dimensional video dynamic rendering adjustment model. Among them, the three-dimensional video dynamic projection model is used to calculate the video picture display angle suitable for the current viewer according to the coordinate conversion position information and the three-dimensional scene content, and generate a corresponding projection matrix (rendering parameters of the virtual camera); the three-dimensional video dynamic rendering adjustment model applies the projection matrix to the rendering engine to generate an initial three-dimensional picture, ensuring that when the user observes the screen from different positions, the three-dimensional content displayed on the screen can be dynamically adjusted according to the user's perspective.
[0128] In a preferred embodiment, in step S6, combine the best viewing point position calculated by the positioning server with the actual position information of the target user for precise overlap, and adjust the initial three-dimensional picture through a position deviation correction algorithm to generate the final three-dimensional picture, as Figure 3 shown, to ensure that the user can obtain the best visual experience at any position.
[0129] In a possible implementation manner, when generating the initial three-dimensional picture, perform image distortion correction on the initial three-dimensional picture through the undistort function in OpenCV.
[0130] Since the initial three-dimensional (3D) image is generated according to the rendering parameters of a virtual camera, there will be a certain degree of image distortion and projection deformation problems. Therefore, when generating the initial 3D image in the embodiments of the present application, the undistort function in OpenCV is further used to correct the image distortion of the initial 3D image, eliminate the above-mentioned image distortion problems, and improve the user's visual experience.
[0131] In a preferred embodiment, the detailed step flowchart of the 3D image display method based on indoor space positioning is as Figure 4 shown. The facial position information and head position information of a person are collected and recognized by cameras at different positions, and then converted into the spatial position of the person's face and uploaded to a positioning server and a 3D rendering engine. Finally, the position of the initial 3D image generated in the 3D rendering engine is corrected in combination with the positioning information in the positioning server, and the final 3D image is generated and displayed on a display screen. While achieving precise positioning of the user in an indoor environment, the distortion in the 3D image is corrected in real time, thereby enhancing the user's visual experience and interaction effect. Compared with the prior art, in the embodiments of the present application, through a high-resolution camera combined with a deep learning model, efficient detection and tracking of head and facial information are achieved, and the robustness of collecting the position information of a person in a complex environment is enhanced; a calibration method based on a red sphere is adopted, feature points are matched using epipolar constraint, and through the SLAM algorithm, the precise mapping of the person's position in a standard Cartesian coordinate system is realized; a dynamic projection model and a rendering adjustment model based on viewpoint tracking are constructed, and the screen image is dynamically optimized according to the user's position and line-of-sight direction, reducing distortion and visual error; OpenCV is used to complete projection deformation and lens distortion correction to ensure the ideal visual effect of 3D image rendering and enhance the immersion and comfort.
[0132] Embodiment Two:
[0133] As Figure 5 shown, Embodiment Two provides a 3D image display system based on indoor space positioning, including a coordinate acquisition module 10, a position information generation module 20, a coordinate conversion module 30, a positioning module 40, a rendering module 50, and a display module 60;
[0134] Among them, the coordinate acquisition module 10 is used to obtain a plurality of pixel coordinates of the head position through a first camera and a plurality of pixel coordinates of the facial position through a second camera. Among them, the first camera is located at the top of the digital twin exhibition hall, and the second camera is located in front of the display screen of the digital twin exhibition hall;
[0135] The position information generation module 20 is configured to perform coordinate matching based on matrix parameters, each of the head position pixel coordinates, each of the facial position pixel coordinates, and facial orientation information to generate the position information of the target user, where the position information includes three-dimensional coordinates and pose information, and the matrix parameters are obtained by calibrating the first camera and the second camera;
[0136] The coordinate conversion module 30 is configured to convert the position information into a preset Cartesian coordinate system to generate corresponding coordinate conversion position information;
[0137] The positioning module 40 is configured to input the three-dimensional model data of the actual space environment and the coordinate conversion position information into a preset positioning server, so that the positioning server generates the best viewing point position of the target user;
[0138] The rendering module 50 is configured to input the coordinate conversion position information into a preset three-dimensional rendering engine, so that the three-dimensional rendering engine generates an initial three-dimensional image;
[0139] The display module 60 is configured to adjust the display position and display angle of the initial three-dimensional image through a position deviation correction algorithm according to the best viewing point position and the coordinate conversion position information, generate a final three-dimensional image and display it on the display screen.
[0140] Further, the coordinate acquisition module 10 includes a video stream acquisition unit, a head detection unit, and a face recognition unit;
[0141] Among them, the video stream acquisition unit is configured to acquire a first video stream and a second video stream through the first camera and the second camera respectively;
[0142] The head detection unit is configured to input the first video stream into a preset head detection model, so that the head detection model detects the head information of the person in the first video stream, and then generates a plurality of head position pixel coordinates;
[0143] The face recognition unit is configured to input the second video stream into a preset face recognition model, so that the face recognition model recognizes the face information of the person in the second video stream, and then generates a plurality of facial position pixel coordinates and facial orientation information;
[0144] Among them, both the head detection model and the face recognition model are deep learning models.
[0145] Further, the position information generation module 20 performs coordinate matching based on matrix parameters, each of the head position pixel coordinates, each of the facial position pixel coordinates, and facial orientation information to generate the position information of the target user, including:
[0146] Substitute each of the facial position pixel coordinates into a preset coordinate matching equation to obtain respective corresponding camera epipolar lines, where the coordinate matching equation is constructed based on the fundamental matrix in the matrix parameters;
[0147] Based on the distances between each of the camera epipolar lines and each of the head position pixel coordinates, determine a number of matching head position pixel coordinates corresponding to each of the camera epipolar lines, and then perform coordinate matching between each of the facial position pixel coordinates and each of the matching head position pixel coordinates to generate respective corresponding coordinate matching point pairs;
[0148] Calculate respective corresponding three-dimensional coordinates from each of the coordinate matching point pairs and the matrix parameters;
[0149] Based on the corresponding relationship between each of the facial position pixel coordinates and the facial orientation information, determine each of the facial orientation information corresponding to each of the three-dimensional coordinates, and then generate respective corresponding pose information;
[0150] Combine each of the three-dimensional coordinates and the corresponding pose information to generate respective corresponding position information;
[0151] Based on the distances between each of the position information and the second camera, determine the position information of the target user.
[0152] In a possible implementation manner, the matrix parameters are obtained by calibrating the first camera and the second camera based on the red sphere calibration method, including:
[0153] Move the red sphere to different positions in the digital twin exhibition hall;
[0154] Each time the red sphere is moved, obtain first pixel coordinates and second pixel coordinates through the first camera and the second camera respectively, and then construct a corresponding sphere coordinate matching point pair;
[0155] Calculate the fundamental matrix between the first camera and the second camera based on a number of sphere coordinate matching point pairs;
[0156] Calculate the essential matrix between the first camera and the second camera based on the fundamental matrix and the camera intrinsic matrix of the first camera and the second camera;
[0157] Combine the fundamental matrix and the essential matrix to construct the matrix parameters.
[0158] Furthermore, the positioning server generates the best viewing point position of the target user, including:
[0159] According to the position information after coordinate conversion, the three-dimensional model data, and the three-dimensional scene content currently displayed on the display screen, calculate the optimal alignment point of the target user's line of sight through a geometric optimization algorithm, where the geometric optimization algorithm is the minimum viewpoint deviation algorithm or the maximum visible area algorithm;
[0160] Take the optimal alignment point of the target user's line of sight as the optimal viewpoint position.
[0161] Further, the three-dimensional rendering engine generates an initial three-dimensional picture, including:
[0162] Input the position information after coordinate conversion into a preset three-dimensional video dynamic projection model, so that the three-dimensional video dynamic projection model calculates a projection matrix according to the position information after coordinate conversion and the three-dimensional scene content currently displayed on the display screen;
[0163] Input the projection matrix into a preset three-dimensional video dynamic rendering adjustment model, so that the three-dimensional video dynamic rendering adjustment model renders the three-dimensional scene content according to the projection matrix to generate the initial three-dimensional picture.
[0164] In a possible implementation manner, the three-dimensional picture display system further includes an image correction module, and the image correction module is used to perform image distortion correction on the initial three-dimensional picture through the undistort function in OpenCV when the initial three-dimensional picture is generated.
[0165] The embodiment of the present application provides a three-dimensional picture display system based on indoor space positioning. By setting two cameras at different angles, efficient detection and tracking of the head and face information of the target user are realized, the robustness of collecting the position information of people in a complex environment is enhanced, and the accuracy of subsequent indoor space positioning is improved; the position information of the target user is determined from multiple coordinates through coordinate matching, providing data support for subsequent three-dimensional picture display according to the target user; according to the position information after coordinate conversion, the optimal viewpoint position of the target user is generated by the positioning server, the initial three-dimensional picture is generated by the three-dimensional rendering engine, and finally, the optimal viewpoint position is combined with the position information of the target user after coordinate conversion, and the angle of the initial three-dimensional picture is adjusted through a position deviation correction algorithm to generate a final three-dimensional picture and display it on the display screen, realizing the correction of the distortion of the initial three-dimensional picture, thereby enhancing the visual experience and interaction effect of users in the digital twin exhibition hall environment.
[0166] The more detailed working principle and step flow of this embodiment can, but are not limited to, refer to the relevant records in Embodiment 1.
[0167] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present application. It should be understood that the above description is only for the specific embodiments of the present application and is not used to limit the protection scope of the present application. It is particularly pointed out that for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. A three-dimensional image display method based on indoor space positioning, characterized in that: include: Acquire a number of pixel coordinates of the head position through a first camera, and acquire a number of pixel coordinates of the face position through a second camera, wherein the first camera is located at the top of the digital twin exhibition hall, and the second camera is located in front of the display screen of the digital twin exhibition hall; Coordinate matching is performed according to matrix parameters, each of the head position pixel coordinates, each of the face position pixel coordinates and face orientation information to generate the position information of the target user, wherein the position information includes three-dimensional coordinates and posture information, and the matrix parameters are obtained by calibrating the first camera and the second camera; Convert the position information into a preset Cartesian coordinate system to generate corresponding coordinate conversion position information; Inputting the three-dimensional model data of the actual space environment and the coordinate conversion position information into a preset positioning server so that the positioning server generates the optimal viewpoint position of the target user; Inputting the coordinate conversion position information into a preset three-dimensional rendering engine so that the three-dimensional rendering engine generates an initial three-dimensional image; According to the optimal viewpoint position and the coordinate conversion position information, the display position and display angle of the initial three-dimensional image are adjusted by a position correction algorithm to generate a final three-dimensional image and display it on the display screen.
2. A three-dimensional image display method based on indoor space positioning as claimed in claim 1, characterized in that: The method of acquiring a plurality of pixel coordinates of head positions through the first camera and acquiring a plurality of pixel coordinates of face positions through the second camera includes: Acquire a first video stream and a second video stream through the first camera and the second camera respectively; Inputting the first video stream into a preset head detection model so that the head detection model detects the head information of the person in the first video stream, and then generates a number of head position pixel coordinates; Inputting the second video stream into a preset face recognition model so that the face recognition model recognizes the facial information of the person in the second video stream, and then generates a number of facial position pixel coordinates and facial orientation information; Among them, the head detection model and the face recognition model are both deep learning models.
3. A three-dimensional image display method based on indoor space positioning as claimed in claim 1, characterized in that: The step of performing coordinate matching according to the matrix parameters, the pixel coordinates of each head position, the pixel coordinates of each face position and the face orientation information to generate the position information of the target user includes: Substituting the pixel coordinates of each facial position into a preset coordinate matching equation to obtain each corresponding camera epipolar line, wherein the coordinate matching equation is constructed based on the basic matrix in the matrix parameters; Determine a plurality of matching head position pixel coordinates corresponding to each camera epipolar line according to the distance between each camera epipolar line and each head position pixel coordinate, and then coordinate match each facial position pixel coordinate with each matching head position pixel coordinate to generate corresponding coordinate matching point pairs; Match each of the coordinate point pairs and the matrix parameters to calculate and obtain the corresponding three-dimensional coordinates; Determine the facial orientation information corresponding to each of the three-dimensional coordinates according to the correspondence between the pixel coordinates of each facial position and the facial orientation information, and then generate the corresponding posture information; Combining each of the three-dimensional coordinates with the corresponding each of the posture information to generate corresponding each of the position information; The location information of the target user is determined according to the distance between each piece of location information and the second camera.
4. A three-dimensional image display method based on indoor space positioning as claimed in claim 1, characterized in that: The matrix parameters are obtained by calibrating the first camera and the second camera based on the red sphere calibration method, including: Move the red sphere to different locations in the digital twin exhibition hall; Each time the red sphere is moved, the first pixel coordinate and the second pixel coordinate are acquired through the first camera and the second camera, respectively, so as to construct a corresponding sphere coordinate matching point pair; Obtaining a basic matrix between the first camera and the second camera by calculation according to a plurality of spherical coordinate matching point pairs; Calculate and obtain an essential matrix between the first camera and the second camera according to the basic matrix and the camera intrinsic parameter matrices of the first camera and the second camera; The matrix parameters are constructed by combining the basic matrix and the essential matrix.
5. The three-dimensional image display method based on indoor space positioning according to claim 1, characterized in that: The positioning server generates the best viewpoint position of the target user, including: Calculate the best alignment point of the sight line of the target user by a geometric optimization algorithm according to the coordinate conversion position information, the three-dimensional model data and the three-dimensional scene content currently displayed on the display screen, wherein the geometric optimization algorithm is a viewpoint deviation minimization algorithm or a visual area maximization algorithm; The optimal alignment point of the target user's sight line is used as the optimal viewpoint position.
6. A three-dimensional image display method based on indoor space positioning as claimed in claim 1, characterized in that: The three-dimensional rendering engine generates an initial three-dimensional picture, including: Inputting the coordinate conversion position information into a preset three-dimensional video dynamic projection model, so that the three-dimensional video dynamic projection model calculates and obtains a projection matrix according to the coordinate conversion position information and the three-dimensional scene content currently displayed on the display screen; The projection matrix is input into a preset three-dimensional video dynamic rendering adjustment model, so that the three-dimensional video dynamic rendering adjustment model renders the three-dimensional scene content according to the projection matrix to generate the initial three-dimensional picture.
7. A three-dimensional image display method based on indoor space positioning according to any one of claims 1 to 6, characterized in that: When the initial three-dimensional image is generated, image distortion correction is performed on the initial three-dimensional image through the undistort function in OpenCV.
8. A three-dimensional image display system based on indoor space positioning, characterized in that: It includes a coordinate acquisition module, a position information generation module, a coordinate conversion module, a positioning module, a rendering module and a display module; The coordinate acquisition module is used to acquire a number of pixel coordinates of the head position through a first camera, and acquire a number of pixel coordinates of the face position through a second camera, wherein the first camera is located at the top of the digital twin exhibition hall, and the second camera is located in front of the display screen of the digital twin exhibition hall; The position information generation module is used to perform coordinate matching according to matrix parameters, each of the head position pixel coordinates, each of the face position pixel coordinates and face orientation information to generate the position information of the target user, wherein the position information includes three-dimensional coordinates and posture information, and the matrix parameters are obtained by calibrating the first camera and the second camera; The coordinate conversion module is used to convert the position information into a preset Cartesian coordinate system to generate corresponding coordinate conversion position information; The positioning module is used to input the three-dimensional model data of the actual space environment and the coordinate conversion position information into a preset positioning server, so that the positioning server generates the best viewpoint position of the target user; The rendering module is used to input the coordinate conversion position information into a preset three-dimensional rendering engine, so that the three-dimensional rendering engine generates an initial three-dimensional picture; The display module is used to adjust the display position and display angle of the initial three-dimensional image according to the optimal viewpoint position and the coordinate conversion position information through a position correction algorithm to generate a final three-dimensional image and display it on the display screen.
9. A three-dimensional image display system based on indoor space positioning as claimed in claim 8, characterized in that: The coordinate acquisition module includes a video stream acquisition unit, a head detection unit and a face recognition unit; The video stream acquisition unit is used to acquire the first video stream and the second video stream through the first camera and the second camera respectively; The head detection unit is used to input the first video stream into a preset head detection model, so that the head detection model detects the head information of the person in the first video stream, and then generates a number of head position pixel coordinates; The face recognition unit is used to input the second video stream into a preset face recognition model so that the face recognition model recognizes the facial information of the person in the second video stream, and then generates a number of facial position pixel coordinates and facial orientation information; Among them, the head detection model and the face recognition model are both deep learning models.
10. A three-dimensional image display system based on indoor space positioning as claimed in claim 8, characterized in that: The position information generation module performs coordinate matching according to matrix parameters, each of the head position pixel coordinates, each of the face position pixel coordinates and face orientation information to generate the position information of the target user, including: Substituting the pixel coordinates of each facial position into a preset coordinate matching equation to obtain each corresponding camera epipolar line, wherein the coordinate matching equation is constructed based on the basic matrix in the matrix parameters; Determine a plurality of matching head position pixel coordinates corresponding to each camera epipolar line according to the distance between each camera epipolar line and each head position pixel coordinate, and then coordinate match each facial position pixel coordinate with each matching head position pixel coordinate to generate corresponding coordinate matching point pairs; Match each of the coordinate point pairs and the matrix parameters to calculate and obtain the corresponding three-dimensional coordinates; Determine the facial orientation information corresponding to each of the three-dimensional coordinates according to the correspondence between the pixel coordinates of each facial position and the facial orientation information, and then generate the corresponding posture information; Combining each of the three-dimensional coordinates with the corresponding each of the posture information to generate corresponding each of the position information; The location information of the target user is determined according to the distance between each piece of location information and the second camera.
Citation Information
Patent Citations
Target positioning and tracking system and method based on video and three-dimensional spatial information registration fusion
CN106204656A
Digital twinborn scene-oriented real trajectory simulation and visual angle capture content recommendation method
CN115128965A
Indoor space positioning method and device based on AR application
CN115760985A
Method and system for displaying indoor positioning result in three-dimensional scene
CN115937487A
Positioning Enhancements to Localization Process for Three-Dimensional Visualization
US20190389600A1
Cited By
Focusing calibration algorithm and system for assembling camera lens by Monte Carlo tree search
CN121309800A
Monte carlo tree search in camera lens assembly focusing calibration algorithm and system
CN121309800B