A three-dimensional picture display method and system based on indoor space positioning

By using cameras at different angles and deep learning models to detect facial and head information in the digital twin exhibition hall, and combining this with the red sphere calibration method, high-precision user position recognition and 3D image distortion correction were achieved, enhancing the user's immersive interactive experience.

CN120147584BActive Publication Date: 2026-02-27GUANGDONG GUODI TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510143221.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2026-02-27
Estimated Expiration
2045-02-10

AI Technical Summary

Technical Problem

Existing digital twin exhibition halls suffer from insufficient positioning accuracy in multi-user dynamic scenarios, limited ability to adapt to multi-view images, and poor real-time distortion correction, making it impossible to achieve high-precision dynamic positioning and 3D image distortion adjustment.

Method used

The system uses two cameras at different angles to acquire the pixel coordinates of the head and face, detects and recognizes them using a deep learning model, and calibrates the cameras using the red sphere calibration method to generate the target user's location information. It then uses a 3D rendering engine and a positioning server to adjust the image, achieving efficient detection and tracking and correcting 3D image distortion.

Benefits of technology

It improves the visual experience and interactive effects for users in digital twin exhibition halls, enhances positioning accuracy and screen adaptation capabilities in complex environments, and ensures that users receive the best visual feedback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147584B_ABST
    Figure CN120147584B_ABST
Patent Text Reader

Abstract

The application provides a three-dimensional picture display method and system based on indoor space positioning, the method comprising: acquiring a plurality of head position pixel coordinates through a first camera and a plurality of face position pixel coordinates through a second camera; performing coordinate matching according to matrix parameters, the head position pixel coordinates, the face position pixel coordinates, and face orientation information to generate position information of a target user; converting the position information into a preset Cartesian coordinate system to generate corresponding coordinate conversion position information; generating an optimal viewpoint position of the target user through a preset positioning server according to the coordinate conversion position information; generating an initial three-dimensional picture through a three-dimensional rendering engine according to the coordinate conversion position information; adjusting the initial three-dimensional picture according to the optimal viewpoint position and the coordinate conversion position information to generate a final three-dimensional picture for display, thereby improving the visual experience and interaction effect of the user.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of indoor space positioning and computer vision processing, and in particular to a three-dimensional picture display method and system based on indoor space positioning. BACKGROUND

[0002] As an interactive platform for deep integration of the physical world and the virtual world, the digital twin exhibition hall provides users with immersive and interactive exhibition experiences by mapping physical environment data to virtual space in real time. Its core relies on high-precision space positioning technology and dynamic adaptation capability of three-dimensional pictures to ensure that users obtain undistorted visual feedback during movement. However, the existing technology still has significant bottlenecks in positioning accuracy and picture adaptation in multi-user dynamic scenarios.

[0003] At present, indoor space positioning is mostly based on the binocular stereo parallax principle, which calculates the target three-dimensional coordinates by simulating human eye parallax through double cameras. However, the traditional method faces two major challenges in complex exhibition hall environments: first, static calibration modes (such as chessboard calibration) are difficult to adapt to multi-angle and large-scale scenes, resulting in cumulative errors in camera parameter calibration; second, in multi-user scenarios, head occlusion, non-frontal orientation and other problems easily cause feature point matching failure, and the positioning robustness is insufficient. In addition, existing systems usually rely on a single visual data source (such as head or face detection), which is difficult to balance accuracy and real-time performance in dynamic environments. Therefore, it is impossible to support distortion dynamic adjustment and optimization of exhibition hall three-dimensional pictures based on accurate space positioning, which greatly affects the display effect of digital twin pictures in the exhibition hall.

[0004] In summary, the existing digital twin exhibition hall faces core problems such as insufficient positioning accuracy, limited multi-angle picture adaptation capability, and poor real-time distortion correction effect in complex interactive scenarios, and there is an urgent need for a technical solution that can integrate multi-source perception data and achieve high-precision dynamic positioning to break through the existing user experience bottleneck. SUMMARY

[0005] To solve the above technical problems, the present application provides a three-dimensional picture display method and system based on indoor space positioning, which realizes accurate positioning of users in indoor environments while correcting distortions in three-dimensional pictures in real time, thereby improving the visual experience and interactive effect of users.

[0006] In a first aspect, the embodiments of the present application provide a three-dimensional picture display method based on indoor space positioning, comprising:

[0007] Obtaining a plurality of head position pixel coordinates through a first camera and a plurality of face position pixel coordinates through a second camera, wherein the first camera is located at the top of the digital twin exhibition hall, and the second camera is located in front of the display screen of the digital twin exhibition hall;

[0008] The coordinate matching is performed according to the matrix parameter, the head position pixel coordinates, the face position pixel coordinates and the face orientation information, to generate position information of the target user, the position information including three-dimensional coordinates and attitude information, and the matrix parameter being obtained by calibrating the first camera and the second camera;

[0009] The position information is converted into a preset Cartesian coordinate system to generate corresponding coordinate conversion position information;

[0010] The three-dimensional model data of the actual space environment and the coordinate conversion position information are input into a preset positioning server, so that the positioning server generates an optimal viewpoint position of the target user;

[0011] The coordinate conversion position information is input into a preset three-dimensional rendering engine, so that the three-dimensional rendering engine generates an initial three-dimensional picture;

[0012] According to the optimal viewpoint position and the coordinate conversion position information, the display position and the display angle of the initial three-dimensional picture are adjusted through a position correction algorithm, to generate a final three-dimensional picture and display the final three-dimensional picture in the display screen.

[0013] The embodiment of the application provides a three-dimensional picture display method based on indoor space positioning, which realizes efficient detection and tracking of head and face information of a target user through two cameras with different angles, enhances the robustness of character position information collection in a complex environment, and improves the accuracy of subsequent indoor space positioning; the position information of the target user is determined from multiple coordinates through coordinate matching, which provides data support for subsequent display of a three-dimensional picture according to the target user; according to the position information after coordinate conversion, the positioning server generates an optimal viewpoint position of the target user, the three-dimensional rendering engine generates an initial three-dimensional picture, and finally the optimal viewpoint position and the coordinate conversion position information of the target user are combined, the initial three-dimensional picture is adjusted in angle through a position correction algorithm, a final three-dimensional picture is generated and displayed in the display screen, the distortion of the initial three-dimensional picture is corrected, and the visual experience and interaction effect of the user in the digital twin exhibition hall environment are improved.

[0014] Further, the obtaining of the head position pixel coordinates through the first camera and the face position pixel coordinates through the second camera comprises:

[0015] The first video stream and the second video stream are respectively obtained through the first camera and the second camera;

[0016] The first video stream is input into a preset head detection model, so that the head detection model detects the head information of the character in the first video stream, and then generates a plurality of head position pixel coordinates;

[0017] inputting the second video stream into a preset face recognition model, so that the face recognition model recognizes face information of a person in the second video stream, and further generates a plurality of face position pixel coordinates and face orientation information;

[0018] The head detection model and the face recognition model are both deep learning models.

[0019] The method provided by the embodiments of the present application can obtain pixel coordinates, and can obtain a first video stream and a second video stream through different cameras, and then detect and recognize the first video stream and the second video stream by combining corresponding deep learning models, extract a plurality of head position pixel coordinates, a plurality of face position pixel coordinates and face orientation information, realize preliminary positioning of a plurality of users in an exhibition hall, provide a data basis for subsequent generation and angle adjustment of a three-dimensional picture, and make the generated three-dimensional picture accurately match the viewing angle of a user, thereby improving the visual experience and interaction effect of the user.

[0020] Further, the coordinate matching according to the matrix parameters, the plurality of head position pixel coordinates, the plurality of face position pixel coordinates and the face orientation information to generate the position information of the target user comprises:

[0021] substitute the plurality of face position pixel coordinates into a preset coordinate matching equation to obtain a plurality of corresponding camera epipolar lines, wherein the coordinate matching equation is obtained according to a fundamental matrix in the matrix parameters;

[0022] determine a plurality of matching head position pixel coordinates corresponding to the plurality of camera epipolar lines according to distances between the plurality of camera epipolar lines and the plurality of head position pixel coordinates, and then perform coordinate matching on the plurality of face position pixel coordinates and the plurality of matching head position pixel coordinates to generate a plurality of corresponding coordinate matching point pairs;

[0023] calculate the plurality of three-dimensional coordinates corresponding to the plurality of coordinate matching point pairs and the matrix parameters;

[0024] determine the plurality of face orientation information corresponding to the plurality of three-dimensional coordinates according to a corresponding relationship between the plurality of face position pixel coordinates and the face orientation information, and then generate the plurality of pose information corresponding to the plurality of three-dimensional coordinates;

[0025] generate the plurality of position information corresponding to the plurality of three-dimensional coordinates and the plurality of pose information;

[0026] determine the position information of the target user according to distances between the plurality of position information and the second camera.

[0027] The embodiment of the present application provides a coordinate matching method, considering that the field of view range of the second camera in front of the screen is smaller than that of the second camera on the top, and the visitor may not face the screen or be blocked, resulting in that the number of identifiable faces is usually less than the number of overhead recognitions, therefore, the embodiment of the present application substitutes the face position pixel coordinates into a coordinate matching equation to obtain each corresponding camera epipolar line. Then, each corresponding head position pixel coordinate is matched according to each camera epipolar line, and the specific matching rule is to match the head position pixel coordinates with a distance less than a set threshold to the camera epipolar line, for example, for a certain camera epipolar line, the head position pixel coordinates with a distance less than a set threshold to the epipolar line are selected for matching, and then a corresponding coordinate matching point pair is established. Then, each corresponding three-dimensional coordinate is calculated and obtained according to each coordinate matching point pair and a matrix parameter, the conversion from two-dimensional coordinates to three-dimensional coordinates is realized, and corresponding posture information is generated in combination with the face orientation information, so as to determine the position and face orientation of each user in the digital twin exhibition hall. Further, when there are multiple position information, a target user is determined according to the distance between each position information and the second camera, for example, the user closest to the second camera is selected as the target user, this strategy can meet the needs of the visitors facing the display plane and being close to the screen, and conforms to the typical standing habit of the key visitors in the exhibition hall, so as to improve the visual experience and interaction effect of the user in the digital twin exhibition hall environment.

[0028] In a possible implementation manner, the matrix parameter is obtained based on a red sphere calibration method, and the first camera and the second camera are calibrated, comprising:

[0029] The red sphere is moved at different positions in the digital twin exhibition hall;

[0030] Each time the red sphere is moved, the first pixel coordinates and the second pixel coordinates are acquired by the first camera and the second camera respectively, and then a corresponding sphere coordinate matching point pair is constructed;

[0031] The basic matrix between the first camera and the second camera is calculated and obtained according to a plurality of sphere coordinate matching point pairs;

[0032] The essential matrix between the first camera and the second camera is calculated and obtained according to the basic matrix and the camera intrinsic parameter matrix of the first camera and the second camera;

[0033] The matrix parameter is constructed and obtained in combination with the basic matrix and the essential matrix.

[0034] The embodiment of the present application provides a calibration method based on a red sphere, generates feature matching point pairs of a camera dynamically by simulating the position of a human head, and then calibrates the camera. Since the embodiment of the present application adopts a first camera and a second camera which are perpendicular to each other and have a large angle of view deviation, if a traditional checkerboard calibration method is used, a large calibration board size is required and it is difficult to find a suitable pitch angle that meets the requirements of the two cameras, therefore, a red sphere with a size similar to that of a human head is selected as a calibration tool. This method makes the calibration process not affected by the direction of the camera and the measurement of the center of the sphere is accurate. Meanwhile, the reflective property of the sphere and the bright red color facilitate the accurate recognition of the camera, improve the calibration accuracy of the camera, and further improve the accuracy of subsequent indoor space positioning and three-dimensional picture generation.

[0035] Further, the positioning server generates the best viewpoint position of the target user, comprising:

[0036] According to the coordinate conversion position information, the three-dimensional model data and the three-dimensional scene content currently displayed by the display screen, a best alignment point of the line of sight of the target user is calculated by a geometric optimization algorithm, and the geometric optimization algorithm is a minimum viewpoint deviation algorithm or a maximum visible area algorithm.

[0037] The best alignment point of the line of sight of the target user is taken as the best viewpoint position.

[0038] In a possible implementation manner, the three-dimensional rendering engine generates an initial three-dimensional picture, comprising:

[0039] The coordinate conversion position information is input into a preset three-dimensional video dynamic projection model, so that the three-dimensional video dynamic projection model calculates and obtains a projection matrix according to the coordinate conversion position information and the three-dimensional scene content currently displayed by the display screen;

[0040] The projection matrix is input into a preset three-dimensional video dynamic rendering adjustment model, so that the three-dimensional video dynamic rendering adjustment model renders the three-dimensional scene content according to the projection matrix, and generates the initial three-dimensional picture.

[0041] The embodiment of the present application presets a three-dimensional video dynamic projection model and a three-dimensional video dynamic rendering adjustment model, wherein the three-dimensional video dynamic projection model is used to calculate the video picture display angle suitable for the current viewer according to the coordinate conversion position information and the three-dimensional scene content, and generate a corresponding projection matrix (rendering parameter of a virtual camera); the three-dimensional video dynamic rendering adjustment model applies the projection matrix to the rendering engine to generate an initial three-dimensional picture, so as to ensure that the three-dimensional content displayed on the screen can be dynamically adjusted according to the viewing angle of the user when the user observes the screen from different positions.

[0042] In a possible implementation manner, when the initial three-dimensional picture is generated, the initial three-dimensional picture is subjected to image distortion correction by means of an undistort function in OpenCV.

[0043] Since the initial three-dimensional picture is generated according to the rendering parameters of the virtual camera, there are problems of image distortion and projection deformation to a certain extent, therefore, the embodiment of the present application further subjects the initial three-dimensional picture to image distortion correction by means of an undistort function in OpenCV when the initial three-dimensional picture is generated, so as to eliminate the above-mentioned image distortion problem and improve the visual experience of the user.

[0044] In a second aspect, correspondingly, the embodiment of the present application provides a three-dimensional picture display system based on indoor space positioning, comprising a coordinate acquisition module, a position information generation module, a coordinate conversion module, a positioning module, a rendering module and a display module.

[0045] The coordinate acquisition module is configured to acquire a plurality of head position pixel coordinates through a first camera and a plurality of face position pixel coordinates through a second camera, wherein the first camera is located at the top of the digital twin exhibition hall, and the second camera is located in front of the display screen of the digital twin exhibition hall.

[0046] The position information generation module is configured to perform coordinate matching according to matrix parameters, the head position pixel coordinates, the face position pixel coordinates and face orientation information, to generate position information of a target user, wherein the position information comprises three-dimensional coordinates and attitude information, and the matrix parameters are obtained by calibrating the first camera and the second camera.

[0047] The coordinate conversion module is configured to convert the position information into a preset Cartesian coordinate system to generate corresponding coordinate conversion position information.

[0048] The positioning module is configured to input three-dimensional model data of an actual space environment and the coordinate conversion position information into a preset positioning server, so that the positioning server generates a best viewpoint position of the target user.

[0049] The rendering module is configured to input the coordinate conversion position information into a preset three-dimensional rendering engine, so that the three-dimensional rendering engine generates an initial three-dimensional picture.

[0050] The display module is configured to adjust the display position and display angle of the initial three-dimensional picture by means of a position correction algorithm according to the best viewpoint position and the coordinate conversion position information, to generate a final three-dimensional picture and display the final three-dimensional picture in the display screen.

[0051] Further, the coordinate acquisition module comprises a video stream acquisition unit, a head detection unit and a face recognition unit.

[0052] The video stream acquisition unit is configured to acquire a first video stream and a second video stream through the first camera and the second camera respectively.

[0053] The head detection unit is configured to input the first video stream into a preset head detection model, so that the head detection model detects the head information of a person in the first video stream, and further generates a plurality of head position pixel coordinates.

[0054] The face recognition unit is configured to input the second video stream into a preset face recognition model, so that the face recognition model recognizes the face information of a person in the second video stream, and further generates a plurality of face position pixel coordinates and face orientation information.

[0055] The head detection model and the face recognition model are both deep learning models.

[0056] Further, the position information generation module performs coordinate matching according to the matrix parameters, the plurality of head position pixel coordinates, the plurality of face position pixel coordinates and the face orientation information, to generate the position information of the target user, comprising:

[0057] Substitute each of the plurality of face position pixel coordinates into a preset coordinate matching equation to obtain a plurality of corresponding camera epipolar lines, wherein the coordinate matching equation is obtained according to a fundamental matrix in the matrix parameters;

[0058] According to the distance between each of the plurality of camera epipolar lines and each of the plurality of head position pixel coordinates, determine a plurality of matching head position pixel coordinates corresponding to each of the plurality of camera epipolar lines, and further perform coordinate matching on each of the plurality of face position pixel coordinates and each of the plurality of matching head position pixel coordinates to generate a plurality of corresponding coordinate matching point pairs;

[0059] Calculate each of the plurality of three-dimensional coordinates corresponding to each of the plurality of coordinate matching point pairs and the matrix parameters;

[0060] According to the corresponding relationship between each of the plurality of face position pixel coordinates and the face orientation information, determine each of the plurality of face orientation information corresponding to each of the plurality of three-dimensional coordinates, and further generate each of the plurality of posture information corresponding thereto;

[0061] Combine each of the plurality of three-dimensional coordinates and each of the plurality of posture information corresponding thereto to generate each of the plurality of position information corresponding thereto.

[0062] According to the distance between each of the plurality of position information and the second camera, determine the position information of the target user. BRIEF DESCRIPTION OF DRAWINGS

[0063] Figure 1 FIG. 1 is a flowchart of a three-dimensional picture display method based on indoor space positioning according to an embodiment of the present application.

[0064] Figure 2 FIG. 2 is a schematic diagram of video stream acquisition of a second camera in a three-dimensional picture display method based on indoor space positioning according to an embodiment of the present application.

[0065] Figure 3 FIG. 3 is a schematic diagram of display of a three-dimensional picture in a three-dimensional picture display method based on indoor space positioning according to an embodiment of the present application.

[0066] Figure 4 FIG. 4 is a detailed flowchart of a three-dimensional picture display method based on indoor space positioning according to an embodiment of the present application.

[0067] Figure 5 FIG. 5 is a structural schematic diagram of a three-dimensional picture display system based on indoor space positioning according to an embodiment of the present application. DETAILED DESCRIPTION

[0068] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.

[0069] It should be noted that the step numbers in the text are only for the convenience of explanation of the specific embodiments, and do not serve as the function of limiting the execution sequence of the steps. In the description of the present application, the terms "first", "second" are only for the purpose of description, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features limited by "first", "second" can explicitly or implicitly include one or more of the features.

[0070] Embodiment One:

[0071] As shown in Figure 1 Embodiment One provides a three-dimensional picture display method based on indoor space positioning, comprising steps S1-S6:

[0072] Step S1, acquiring a plurality of head position pixel coordinates through a first camera and a plurality of face position pixel coordinates through a second camera, wherein the first camera is located at the top of a digital twin exhibition hall, and the second camera is located in front of a display screen of the digital twin exhibition hall;

[0073] Step S2, coordinate matching is performed according to the matrix parameters, the respective head position pixel coordinates, the respective face position pixel coordinates and the face orientation information, to generate position information of the target user, the position information including three-dimensional coordinates and attitude information, the matrix parameters being obtained by calibrating the first camera and the second camera;

[0074] Step S3, the position information is converted into a preset Cartesian coordinate system to generate corresponding coordinate conversion position information;

[0075] Step S4, three-dimensional model data of an actual space environment and the coordinate conversion position information are input into a preset positioning server, so that the positioning server generates an optimal viewpoint position of the target user;

[0076] Step S5, the coordinate conversion position information is input into a preset three-dimensional rendering engine, so that the three-dimensional rendering engine generates an initial three-dimensional picture;

[0077] Step S6, according to the optimal viewpoint position and the coordinate conversion position information, the display position and the display angle of the initial three-dimensional picture are adjusted through a position rectification algorithm, to generate a final three-dimensional picture and display the final three-dimensional picture in the display screen.

[0078] The embodiment of the application provides a three-dimensional picture display method based on indoor space positioning. The two cameras with different angles are arranged to realize efficient detection and tracking of head and face information of a target user, to enhance the robustness of position information collection of a person in a complex environment, and to improve the accuracy of subsequent indoor space positioning. The position information of the target user is determined from multiple coordinates through coordinate matching, to provide data support for subsequent display of a three-dimensional picture according to the target user. According to the position information after coordinate conversion, the positioning server generates an optimal viewpoint position of the target user, the three-dimensional rendering engine generates an initial three-dimensional picture, and finally the optimal viewpoint position and the coordinate conversion position information of the target user are combined, the initial three-dimensional picture is adjusted in angle through a position rectification algorithm, a final three-dimensional picture is generated and displayed in the display screen, the distortion of the initial three-dimensional picture is corrected, and the visual experience and interaction effect of a user in a digital twin exhibition hall environment are improved.

[0079] Further, in step S1, the first camera is used to obtain a plurality of head position pixel coordinates, and the second camera is used to obtain a plurality of face position pixel coordinates, including:

[0080] The first camera and the second camera are respectively used to obtain a first video stream and a second video stream;

[0081] inputting the first video stream into a preset head detection model, so that the head detection model detects head information of a person in the first video stream, and generates a plurality of head position pixel coordinates;

[0082] inputting the second video stream into a preset face recognition model, so that the face recognition model recognizes face information of the person in the second video stream, and generates a plurality of face position pixel coordinates and face orientation information;

[0083] wherein the head detection model and the face recognition model are both deep learning models.

[0084] The embodiment of the present application provides a method for obtaining pixel coordinates, which acquires a first video stream and a second video stream through different cameras respectively, and then detects and recognizes the first video stream and the second video stream in combination with corresponding deep learning models, extracts a plurality of head position pixel coordinates, a plurality of face position pixel coordinates and face orientation information respectively, realizes preliminary positioning of a plurality of users in an exhibition hall, provides a data basis for subsequent generation and angle adjustment of a three-dimensional picture, makes the generated three-dimensional picture accurately match the viewing angle of the user, and improves the visual experience and interaction effect of the user.

[0085] In a preferred embodiment, the first camera and the second camera are high-resolution cameras for capturing continuous RGB-D video streams of a person, as shown in Figure 2 The head detection model can be a Fully Convolutional Head Detector model, and the face recognition model can be a MediaPipe Face Mesh model.

[0086] Further, in step S2, the coordinate matching is performed according to the matrix parameters, the plurality of head position pixel coordinates, the plurality of face position pixel coordinates and the face orientation information, to generate position information of a target user, including:

[0087] each of the face position pixel coordinates is substituted into a preset coordinate matching equation to obtain a plurality of corresponding camera epipolar lines, the coordinate matching equation being obtained according to a fundamental matrix in the matrix parameters;

[0088] a plurality of matching head position pixel coordinates corresponding to each of the camera epipolar lines are determined according to a distance between each of the camera epipolar lines and each of the head position pixel coordinates, and then the coordinate matching is performed between each of the face position pixel coordinates and each of the matching head position pixel coordinates to generate a plurality of corresponding coordinate matching point pairs;

[0089] each of the three-dimensional coordinates is calculated according to each of the coordinate matching point pairs and the matrix parameter;

[0090] According to the correspondence between each of the face position pixel coordinates and the face orientation information, each of the face orientation information corresponding to each of the three-dimensional coordinates is determined, and each of the posture information is further generated;

[0091] Each of the position information is generated in combination with each of the three-dimensional coordinates and each of the posture information;

[0092] The position information of the target user is determined according to the distance between each of the position information and the second camera.

[0093] The embodiment of the present application provides a coordinate matching method. Considering that the field of view of the second camera in front of the screen is smaller than that of the second camera on the top, and the visitor may not face the screen or be blocked, the number of identifiable faces is usually less than the number of top-identified faces. Therefore, the embodiment of the present application substitutes the face position pixel coordinates into the coordinate matching equation to obtain each corresponding camera epipolar line. Then, each corresponding head position pixel coordinate is matched according to each camera epipolar line. The specific matching rule is to match the head position pixel coordinates with a distance less than a set threshold from the camera epipolar line, for example, for a certain camera epipolar line, the head position pixel coordinates with a distance less than a set threshold from the epipolar line are selected for matching, and then the corresponding coordinate matching point pairs are established. Then, each of the corresponding three-dimensional coordinates is calculated according to each of the coordinate matching point pairs and the matrix parameter, the conversion from two-dimensional coordinates to three-dimensional coordinates is realized, and the corresponding posture information is generated in combination with the face orientation information, so as to determine the position and face orientation of each user in the digital twin exhibition hall. Further, when there are multiple position information, the target user is determined according to the distance between each of the position information and the second camera, for example, the user closest to the second camera is selected as the target user. This strategy gives priority to meeting the needs of visitors facing the display plane and being close to the screen, which conforms to the typical standing habits of key visitors in the exhibition hall, thereby improving the visual experience and interaction effect of users in the digital twin exhibition hall environment.

[0094] In a preferred embodiment, the coordinate matching equation is:

[0095] x ′T Fx=0

[0096] wherein x ′ and x are the pixel point coordinates (homogeneous coordinate representation) corresponding to the two cameras respectively, and F is the fundamental matrix in the matrix parameter.

[0097] Further, to prevent screen image jitter, when a person moves, the embodiment of the application only updates the position information of each user when each three-dimensional coordinate changes by more than a preset threshold, ensuring the stability of the display screen display effect.

[0098] In one possible implementation, in step S2, the matrix parameter is obtained by calibrating the first camera and the second camera based on a red sphere calibration method, comprising:

[0099] moving the red sphere at different positions in the digital twin exhibition hall;

[0100] acquiring first pixel coordinates and second pixel coordinates by the first camera and the second camera respectively each time the red sphere is moved, and then constructing corresponding sphere coordinate matching point pairs;

[0101] calculating the essential matrix between the first camera and the second camera according to a plurality of sphere coordinate matching point pairs;

[0102] calculating the essential matrix between the first camera and the second camera according to the intrinsic matrix of the camera of the first camera and the second camera;

[0103] combining the fundamental matrix and the essential matrix to construct the matrix parameter.

[0104] The embodiment of the application provides a calibration method based on a red sphere, which generates feature matching point pairs of a camera dynamically by simulating the position of a human head, and then calibrates the camera. Since the embodiment of the application uses a first camera and a second camera which are perpendicular to each other and have a large angle of view deviation, if a traditional checkerboard calibration method is used, a large calibration board size is required and it is difficult to find a suitable pitch angle that meets the requirements of the two cameras, therefore, a red sphere with a size similar to that of a human head is selected as a calibration tool. This method makes the calibration process not affected by the direction of the camera and the measurement of the center of the sphere accurate, and the reflective properties of the sphere and the bright red color also facilitate accurate recognition by the camera, improving the calibration accuracy of the camera and the accuracy of subsequent indoor space positioning and three-dimensional image generation.

[0105] In one preferred embodiment, the specific process of calibrating the first camera and the second camera is as follows:

[0106] move the red sphere at different positions in the room, capture the images in the first camera at the top and the second camera in front of the screen at the same time each time the red sphere is moved, record the pixel coordinates of the red sphere, select the center of the red sphere as a feature point, and obtain a plurality of feature matching point pairs after a plurality of acquisitions. The matching point pair (x, x') satisfies the following relationship:

[0107] x′T Fx = 0

[0108] Then, the fundamental matrix and the essential matrix are calculated respectively:

[0109] The fundamental matrix (F) describes the geometric relationship between two cameras, used to map a point in one camera image to the epipolar line in the other camera image. The calculation steps are as follows:

[0110] (1) Construct linear equations: using at least 8 pairs of feature matching points (x, x'), construct homogeneous linear equations: AF = 0, where A is a matrix constructed by point coordinates.

[0111] (2) Singular Value Decomposition (SVD): perform SVD decomposition on matrix A to solve F by minimizing ||AF||.

[0112] (3) Constraint enforcement: to ensure the rank of F is 2, perform SVD decomposition on the solved matrix F, and set the smallest singular value to zero.

[0113] The essential matrix (E) further describes the relative pose (rotation and translation) between two cameras, which is calculated by the relationship between the intrinsic matrix of the camera and the fundamental matrix:

[0114] (1) Calculate the essential matrix E using the fundamental matrix F obtained in the previous step and the known camera intrinsic matrices K1, K2:

[0115]

[0116] (2) Standardize E. Perform SVD decomposition on E and enforce its singular values to satisfy the constraint: the first two singular values are equal, and the third is 0.

[0117] (3) Recover the relative motion parameters from the essential matrix E. Using SVD decomposition, we can get 4 sets of possible rotation matrices R and translation vectors t solutions:

[0118]

[0119] (4) Select the correct R and t combination from the 4 sets of solutions by verifying the constraint that the reconstructed 3D points are in front of both cameras.

[0120] In a preferred embodiment, in step S3, using the SLAM algorithm, the three-dimensional coordinates and pose of the target person are converted to a Cartesian coordinate system with the lower edge of the display screen as the origin, and the corresponding coordinate conversion position information is generated.

[0121] Further, in step S4, the positioning server generates the optimal viewpoint position of the target user, including:

[0122] According to the coordinate conversion position information, the three-dimensional model data and the three-dimensional scene content currently displayed by the display screen, the optimal alignment point of the line of sight of the target user is calculated by a geometric optimization algorithm, which is a minimum viewpoint deviation algorithm or a maximum visible area algorithm;

[0123] The optimal alignment point of the line of sight of the target user is taken as the optimal viewpoint position.

[0124] In a possible implementation, in step S5, the three-dimensional rendering engine generates an initial three-dimensional picture, including:

[0125] The coordinate conversion position information is input into a preset three-dimensional video dynamic projection model, so that the three-dimensional video dynamic projection model calculates a projection matrix according to the coordinate conversion position information and the three-dimensional scene content currently displayed by the display screen;

[0126] The projection matrix is input into a preset three-dimensional video dynamic rendering adjustment model, so that the three-dimensional video dynamic rendering adjustment model renders the three-dimensional scene content according to the projection matrix to generate the initial three-dimensional picture.

[0127] The three-dimensional video dynamic projection model is used to calculate a video picture display angle suitable for a current viewer according to coordinate conversion position information and three-dimensional scene content, and generate a corresponding projection matrix (a rendering parameter of a virtual camera); and the three-dimensional video dynamic rendering adjustment model applies the projection matrix to a rendering engine to generate an initial three-dimensional picture, so as to ensure that three-dimensional content displayed on a screen can be dynamically adjusted according to a viewing angle of a user when the user observes the screen from different positions.

[0128] In a preferred embodiment, in step S6, the optimal viewpoint position calculated by the positioning server is combined with actual position information of the target user to perform accurate superposition, and the initial three-dimensional picture is adjusted by a position correction algorithm to generate a final three-dimensional picture, as shown in FIG. 6, so as to ensure that the user can obtain the best visual experience at any position. Figure 3

[0129] In a possible implementation, when the initial three-dimensional picture is generated, an undistort function in OpenCV is used to correct image distortion of the initial three-dimensional picture.

[0130] ​Since the initial three-dimensional picture is generated according to the rendering parameters of the virtual camera, there will be a certain degree of image distortion and projection deformation problem, therefore, the embodiment of the application further utilizes the undistort function in OpenCV to correct the image distortion of the initial three-dimensional picture when generating the initial three-dimensional picture, eliminates the above image distortion problem, and improves the visual experience of the user.

[0131] In a preferred embodiment, the detailed step flow chart of the three-dimensional picture display method based on indoor space positioning is as shown in Figure 4 The face position information and head position information of the person are collected and recognized by the cameras at different positions, and then converted into the spatial position of the face of the person, uploaded to the positioning server and the three-dimensional rendering engine, and finally the initial three-dimensional picture generated in the three-dimensional rendering engine is corrected in position by the positioning information in the positioning server to generate the final three-dimensional picture and display it in the display screen. While realizing the accurate positioning of the user in the indoor environment, the distortion in the three-dimensional picture is corrected in real time, thereby improving the visual experience and interaction effect of the user. Compared with the prior art, the embodiment of the application realizes efficient detection and tracking of head and face information by high-resolution cameras combined with deep learning models, enhances the robustness of the collection of person position information in complex environments, realizes accurate mapping of the position of the person in the standard Cartesian coordinate system by using the red sphere calibration method, matching feature points by epipolar constraint, and SLAM algorithm, constructs a dynamic projection model and rendering adjustment model based on viewpoint tracking, dynamically optimizes the screen picture according to the user position and line of sight direction, reduces distortion and visual error, and uses OpenCV to complete projection deformation and lens distortion correction, ensuring the ideal visual effect of three-dimensional picture rendering and enhancing the sense of immersion and comfort.

[0132] Embodiment two:

[0133] As shown in Figure 5 Embodiment two provides a three-dimensional picture display system based on indoor space positioning, which includes a coordinate acquisition module 10, a position information generation module 20, a coordinate conversion module 30, a positioning module 40, a rendering module 50, and a display module 60.

[0134] The coordinate acquisition module 10 is used to acquire a plurality of head position pixel coordinates through a first camera and a plurality of face position pixel coordinates through a second camera, wherein the first camera is located at the top of the digital twin exhibition hall, and the second camera is located in front of the display screen of the digital twin exhibition hall.

[0135] The position information generation module 20 is configured to perform coordinate matching according to the matrix parameter, the head position pixel coordinates, the face position pixel coordinates, and the face orientation information, to generate position information of the target user, the position information including three-dimensional coordinates and attitude information, and the matrix parameter being obtained by calibrating the first camera and the second camera.

[0136] The coordinate conversion module 30 is configured to convert the position information into a preset Cartesian coordinate system to generate corresponding coordinate conversion position information.

[0137] The positioning module 40 is configured to input three-dimensional model data of an actual space environment and the coordinate conversion position information into a preset positioning server, so that the positioning server generates an optimal viewpoint position of the target user.

[0138] The rendering module 50 is configured to input the coordinate conversion position information into a preset three-dimensional rendering engine, so that the three-dimensional rendering engine generates an initial three-dimensional picture.

[0139] The display module 60 is configured to adjust a display position and a display angle of the initial three-dimensional picture by a position correction algorithm according to the optimal viewpoint position and the coordinate conversion position information, to generate a final three-dimensional picture and display the final three-dimensional picture on the display screen.

[0140] Further, the coordinate acquisition module 10 includes a video stream acquisition unit, a head detection unit, and a face recognition unit.

[0141] The video stream acquisition unit is configured to acquire a first video stream and a second video stream through the first camera and the second camera, respectively.

[0142] The head detection unit is configured to input the first video stream into a preset head detection model, so that the head detection model detects human head information in the first video stream, and further generates a plurality of head position pixel coordinates.

[0143] The face recognition unit is configured to input the second video stream into a preset face recognition model, so that the face recognition model recognizes human face information in the second video stream, and further generates a plurality of face position pixel coordinates and face orientation information.

[0144] The head detection model and the face recognition model are both deep learning models.

[0145] Further, the position information generation module 20 performs coordinate matching according to the matrix parameter, the head position pixel coordinates, the face position pixel coordinates, and the face orientation information, to generate position information of the target user, including:

[0146] substituting each of the face position pixel coordinates into a preset coordinate matching equation to obtain each corresponding camera epipolar line, the coordinate matching equation being constructed according to a fundamental matrix in the matrix parameter;

[0147] determining a plurality of matching head position pixel coordinates corresponding to each of the camera epipolar lines according to a distance between each of the camera epipolar lines and each of the head position pixel coordinates, and then performing coordinate matching on each of the face position pixel coordinates and each of the matching head position pixel coordinates to generate each corresponding coordinate matching point pair;

[0148] calculating each of the three-dimensional coordinates corresponding to each of the coordinate matching point pairs and the matrix parameter;

[0149] determining each of the face orientation information corresponding to each of the three-dimensional coordinates according to a corresponding relationship between each of the face position pixel coordinates and the face orientation information, and then generating each of the pose information corresponding thereto;

[0150] generating each of the position information corresponding to each of the three-dimensional coordinates and each of the pose information corresponding thereto;

[0151] determining the position information of the target user according to a distance between each of the position information and the second camera;

[0152] In one possible implementation manner, the matrix parameter is obtained by calibrating the first camera and the second camera based on a red sphere calibration method, and includes:

[0153] moving the red sphere at different positions in the digital twin exhibition hall;

[0154] acquiring first pixel coordinates and second pixel coordinates by the first camera and the second camera respectively each time the red sphere is moved, and then constructing a corresponding sphere coordinate matching point pair;

[0155] calculating the fundamental matrix between the first camera and the second camera according to a plurality of sphere coordinate matching point pairs;

[0156] calculating the essential matrix between the first camera and the second camera according to the fundamental matrix and an intrinsic matrix of the first camera and the second camera;

[0157] constructing the matrix parameter according to the fundamental matrix and the essential matrix.

[0158] Further, the positioning server generates the best viewpoint position of the target user, including:

[0159] According to the coordinate conversion position information, the three-dimensional model data and the three-dimensional scene content currently displayed by the display screen, a best alignment point of the line of sight of the target user is calculated by a geometric optimization algorithm, the geometric optimization algorithm being a minimum viewpoint deviation algorithm or a maximum visible area algorithm.

[0160] The best alignment point of the line of sight of the target user is taken as the best viewpoint position.

[0161] Further, the three-dimensional rendering engine generates an initial three-dimensional picture, including:

[0162] The coordinate conversion position information is input to a preset three-dimensional video dynamic projection model, so that the three-dimensional video dynamic projection model calculates a projection matrix according to the coordinate conversion position information and the three-dimensional scene content currently displayed by the display screen;

[0163] The projection matrix is input to a preset three-dimensional video dynamic rendering adjustment model, so that the three-dimensional video dynamic rendering adjustment model renders the three-dimensional scene content according to the projection matrix to generate the initial three-dimensional picture.

[0164] In a possible implementation manner, the three-dimensional picture display system further includes an image correction module, which is configured to perform image distortion correction on the initial three-dimensional picture by using an undistort function in OpenCV when the initial three-dimensional picture is generated.

[0165] The embodiment of the present application provides a three-dimensional picture display system based on indoor space positioning. Efficient detection and tracking of head and face information of a target user are achieved by using two cameras with different angles, the robustness of collection of position information of a person in a complex environment is enhanced, and the accuracy of subsequent indoor space positioning is improved. The position information of the target user is determined from multiple coordinates through coordinate matching, which provides data support for subsequent display of a three-dimensional picture according to the target user. According to the position information after coordinate conversion, the best viewpoint position of the target user is generated by a positioning server, an initial three-dimensional picture is generated by a three-dimensional rendering engine, and finally the best viewpoint position and the coordinate conversion position information of the target user are combined, the initial three-dimensional picture is adjusted in angle by a position correction algorithm, a final three-dimensional picture is generated and displayed in the display screen, the distortion of the initial three-dimensional picture is corrected, and the visual experience and interaction effect of a user in a digital twin exhibition hall environment are improved.

[0166] The working principle and step flow of the embodiment are described in more detail, but are not limited to, the related description of the first embodiment.

[0167] The above-described specific embodiments, purposes, technical solutions and beneficial effects of the present application are further described in detail, and it should be understood that the above-described is only a specific embodiment of the present application and is not used to limit the protection scope of the present application. It is particularly pointed out that for those skilled in the art, any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A three-dimensional picture display method based on indoor space positioning, characterized by, The method comprises the following steps: obtaining a plurality of head position pixel coordinates through a first camera and a plurality of face position pixel coordinates through a second camera, wherein the first camera is located at the top of the digital twin exhibition hall, and the second camera is located in front of the display screen of the digital twin exhibition hall; performing coordinate matching according to matrix parameters, the head position pixel coordinates, the face position pixel coordinates, and face orientation information to generate position information of a target user, wherein the position information comprises three-dimensional coordinates and attitude information, and the matrix parameters are obtained by calibrating the first camera and the second camera; converting the position information into a preset Cartesian coordinate system to generate corresponding coordinate conversion position information; inputting three-dimensional model data of an actual space environment and the coordinate conversion position information into a preset positioning server to enable the positioning server to generate an optimal viewpoint position of the target user; inputting the coordinate conversion position information into a preset three-dimensional rendering engine to enable the three-dimensional rendering engine to generate an initial three-dimensional image; adjusting the display position and display angle of the initial three-dimensional image through a position correction algorithm according to the optimal viewpoint position and the coordinate conversion position information to generate a final three-dimensional image and display the final three-dimensional image on the display screen.

2. A three-dimensional picture presentation method based on indoor space positioning according to claim 1, characterized in that, The method of obtaining a plurality of head position pixel coordinates through a first camera and a plurality of face position pixel coordinates through a second camera comprises the following steps: obtaining a first video stream and a second video stream through the first camera and the second camera respectively; inputting the first video stream into a preset head detection model to enable the head detection model to detect the head information of a person in the first video stream and further generate a plurality of head position pixel coordinates; inputting the second video stream into a preset face recognition model to enable the face recognition model to recognize the face information of a person in the second video stream and further generate a plurality of face position pixel coordinates and face orientation information; wherein the head detection model and the face recognition model are both deep learning models.

3. The three-dimensional picture presentation method based on indoor space positioning according to claim 1, wherein, The method of performing coordinate matching according to matrix parameters, the head position pixel coordinates, the face position pixel coordinates, and face orientation information to generate position information of a target user comprises the following steps: substituting the face position pixel coordinates into a preset coordinate matching equation to obtain a plurality of corresponding camera epipolar lines, wherein the coordinate matching equation is obtained according to a fundamental matrix in the matrix parameters; determining a plurality of matching head position pixel coordinates corresponding to each of the camera epipolar lines according to the distance between each of the camera epipolar lines and each of the head position pixel coordinates, and further performing coordinate matching on each of the face position pixel coordinates and each of the matching head position pixel coordinates to generate a plurality of corresponding coordinate matching point pairs; calculating the coordinate matching point pairs and the matrix parameters to obtain a plurality of corresponding three-dimensional coordinates; determining a plurality of face orientation information corresponding to each of the three-dimensional coordinates according to the corresponding relationship between each of the face position pixel coordinates and the face orientation information, and further generating a plurality of corresponding attitude information. The corresponding position information is generated by combining each three-dimensional coordinate and corresponding each attitude information; The position information of the target user is determined according to the distance between each position information and the second camera.

4. The three-dimensional picture presentation method based on indoor space positioning according to claim 1, wherein, The matrix parameter is obtained by calibrating the first camera and the second camera based on a red sphere calibration method, including: The red sphere is moved at different positions in the digital twin exhibition hall; The first pixel coordinate and the second pixel coordinate are obtained by the first camera and the second camera respectively each time the red sphere is moved, and then the corresponding sphere coordinate matching point pair is constructed; The essential matrix between the first camera and the second camera is calculated according to a plurality of sphere coordinate matching point pairs; The intrinsic matrix between the first camera and the second camera is calculated according to the essential matrix and the camera intrinsic matrix of the first camera and the second camera; The matrix parameter is constructed and obtained by combining the essential matrix and the intrinsic matrix.

5. The three-dimensional picture presentation method based on indoor space positioning according to claim 1, wherein, The positioning server generates the best viewpoint position of the target user, including: The best alignment point of the line of sight of the target user is calculated by a geometric optimization algorithm according to the coordinate conversion position information, the three-dimensional model data and the three-dimensional scene content currently displayed by the display screen, and the geometric optimization algorithm is a minimum viewpoint deviation algorithm or a maximum visible area algorithm; The best alignment point of the line of sight of the target user is taken as the best viewpoint position.

6. The three-dimensional picture presentation method based on indoor space positioning according to claim 1, wherein, The three-dimensional rendering engine generates an initial three-dimensional picture, including: The coordinate conversion position information is input into a preset three-dimensional video dynamic projection model, so that the three-dimensional video dynamic projection model calculates a projection matrix according to the coordinate conversion position information and the three-dimensional scene content currently displayed by the display screen; The projection matrix is input into a preset three-dimensional video dynamic rendering adjustment model, so that the three-dimensional video dynamic rendering adjustment model renders the three-dimensional scene content according to the projection matrix to generate the initial three-dimensional picture.

7. A three-dimensional picture presentation method based on indoor space positioning according to any one of claims 1-6, characterized in that, When the initial three-dimensional picture is generated, the image distortion correction of the initial three-dimensional picture is performed by the undistort function in OpenCV.

8. A three-dimensional picture display system based on indoor space positioning, characterized by, The coordinate acquisition module, the position information generation module, the coordinate conversion module, the positioning module, the rendering module and the display module are included; The coordinate acquisition module is used to acquire a plurality of head position pixel coordinates by the first camera and a plurality of face position pixel coordinates by the second camera, wherein the first camera is located at the top of the digital twin exhibition hall, and the second camera is located in front of the display screen of the digital twin exhibition hall; The position information generation module is used to generate the position information of the target user by coordinate matching according to the matrix parameter, each head position pixel coordinate, each face position pixel coordinate and face orientation information, wherein the position information includes three-dimensional coordinates and attitude information, and the matrix parameter is obtained by calibrating the first camera and the second camera; The coordinate conversion module is used to convert the position information into a preset Cartesian coordinate system to generate corresponding coordinate conversion position information. The positioning module is configured to input the three-dimensional model data of the actual space environment and the coordinate conversion position information to a preset positioning server, so that the positioning server generates an optimal viewpoint position of the target user; The rendering module is configured to input the coordinate conversion position information to a preset three-dimensional rendering engine, so that the three-dimensional rendering engine generates an initial three-dimensional picture; The display module is configured to adjust the display position and display angle of the initial three-dimensional picture by a position correction algorithm according to the optimal viewpoint position and the coordinate conversion position information, generate a final three-dimensional picture, and display the final three-dimensional picture on the display screen.

9. A three-dimensional picture presentation system based on indoor space positioning as claimed in claim 8, characterized in that, The coordinate acquisition module includes a video stream acquisition unit, a head detection unit, and a face recognition unit; The video stream acquisition unit is configured to acquire a first video stream and a second video stream through the first camera and the second camera, respectively; The head detection unit is configured to input the first video stream into a preset head detection model, so that the head detection model detects the head information of a person in the first video stream, and further generates a plurality of head position pixel coordinates; The face recognition unit is configured to input the second video stream into a preset face recognition model, so that the face recognition model recognizes the face information of the person in the second video stream, and further generates a plurality of face position pixel coordinates and face orientation information; The head detection model and the face recognition model are both deep learning models.

10. A three-dimensional picture presentation system based on indoor space positioning as claimed in claim 8, characterized in that, The position information generation module performs coordinate matching according to the matrix parameters, the plurality of head position pixel coordinates, the plurality of face position pixel coordinates, and the face orientation information, to generate the position information of the target user, including: Substituting the plurality of face position pixel coordinates into a preset coordinate matching equation to obtain a plurality of corresponding camera epipolar lines, the coordinate matching equation being obtained according to a fundamental matrix in the matrix parameters; According to the distances between each of the camera epipolar lines and each of the head position pixel coordinates, a plurality of matching head position pixel coordinates corresponding to each of the camera epipolar lines are determined, and then the plurality of face position pixel coordinates and the plurality of matching head position pixel coordinates are subjected to coordinate matching to generate a plurality of corresponding coordinate matching point pairs; Each of the coordinate matching point pairs and the matrix parameters are used to calculate a plurality of corresponding three-dimensional coordinates; According to the corresponding relationship between each of the face position pixel coordinates and the face orientation information, each of the face orientation information corresponding to each of the three-dimensional coordinates is determined, and then each of the posture information corresponding to each of the three-dimensional coordinates is generated; Each of the position information corresponding to each of the three-dimensional coordinates and each of the posture information is generated; According to the distances between each of the position information and the second camera, the position information of the target user is determined.

Citation Information

Patent Citations

  • Target positioning and tracking system and method based on video and three-dimensional spatial information registration fusion

    CN106204656A

  • Digital twinborn scene-oriented real trajectory simulation and visual angle capture content recommendation method

    CN115128965A