Three-dimensional tour scene display method, display system, and display terminal
By producing three-dimensional tour scene videos on the display terminal and combining them with sensing devices, the problem of tour experience when there is a lack of tour guides at the tour venue is solved, and a better tour and exhibition experience is achieved.
Patent Information
- Application Number
- CN202210147182.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-17
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2042-02-17
AI Technical Summary
In tourist places such as museums, when the number of tourists is large, the terminal display screen cannot meet the tourists' needs for sightseeing and exhibition viewing, resulting in a poor sightseeing experience, especially in the absence of tour guides.
By producing a three-dimensional tour scene video on the display terminal and combining it with a sensing device to sense whether there are tourists and tour guides at the tour site, the playback of the three-dimensional tour scene video is intelligently controlled, including recording the tour guide's explanation video and converting it into a three-dimensional video to synthesize the three-dimensional tour scene.
It improves the tourists' tour and exhibition experience, allowing them to have an immersive three-dimensional tour experience even without a tour guide.
Smart Images

Figure CN114529673B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of display technology, and in particular relates to a display method and display system for a three-dimensional tour scene, and a display terminal. Background Art
[0002] At present, museums and other tourist attractions are usually equipped with large-size all-in-one machines, such as terminal display screens, which mainly play some pre-recorded two-dimensional video explanations in a single format. Based on this, in museums and other tourist attractions, if tourists want to have a deeper understanding of the detailed information of the museum's cultural relics, they still mainly rely on the detailed explanations of the venue's tour guides.
[0003] However, in museums and other tourist attractions, due to the small number of tour guides, when the number of visitors is large, relying solely on the terminal display screens in the museum is often unable to meet the tourists' high touring and exhibition needs, resulting in poor touring and exhibition experience for tourists in museums and other tourist attractions. Summary of the Invention
[0004] To address the aforementioned issues, the present invention provides a method, system, and display terminal for displaying a three-dimensional tour scene. This method provides visitors with a more immersive tour experience at the site, while also resolving the technical problem of enabling visitors to better explore and view exhibitions when there are no tour guides at the site, thereby enhancing their experience.
[0005] The present invention provides a method for displaying a three-dimensional tour scene, wherein the three-dimensional tour scene is a three-dimensional tour scene created based on a tour site, comprising:
[0006] A three-dimensional tour scene video is produced in advance on a display terminal; the three-dimensional tour scene video includes the three-dimensional tour scene and a tour guide's explanation of the tour site; the tour site is equipped with the display terminal and a sensing device;
[0007] The sensing device senses whether there are tourists at the tourist site and no guide, and determines whether the tourists need an explanation at the tourist site;
[0008] If yes, the display terminal is controlled to play the three-dimensional tour scene video.
[0009] Optionally, the display terminal is configured with a recording mode;
[0010] The step of producing a 3D tour scene video on a display terminal in advance includes:
[0011] Turning on the recording mode of the display terminal;
[0012] The display terminal records the video of the tour guide's explanation of the tour site;
[0013] The display terminal converts the explanation video of the tourist site into the three-dimensional tourist scene video.
[0014] Optionally, the display terminal is configured with a recording mode;
[0015] The step of producing a 3D tour scene video on a display terminal in advance includes:
[0016] The display terminal records the video of the tour site in real time;
[0017] The sensing device senses whether the tour guide is present at the tourist site;
[0018] If yes, then determine whether the lecturer is in the lecture state;
[0019] If yes, the display terminal intercepts the video of the tour guide's explanation from the tour site video, and converts the video of the tour guide's explanation into the three-dimensional tour scene video.
[0020] Optionally, the step of producing a 3D tour scene video on a display terminal in advance includes:
[0021] Establishing a three-dimensional model of the tourist site in advance, and storing the three-dimensional model in the display terminal;
[0022] Collecting the head portrait of the tour guide in advance, synthesizing the three-dimensional face of the tour guide, and storing the three-dimensional face in the display terminal;
[0023] The display terminal collects the explanation voice, explanation movements and expression information of the explainer in real time;
[0024] The display terminal combines the three-dimensional model of the tour site, the three-dimensional face of the tour guide, and the tour guide's explanation voice, explanation movements, and expression information during the explanation into the three-dimensional tour scene video.
[0025] Optionally, the step of producing a 3D tour scene video on a display terminal in advance includes:
[0026] Establishing a three-dimensional model of the tourist site in advance, and storing the three-dimensional model in the display terminal;
[0027] Collecting the head portrait of the tour guide in advance, synthesizing the three-dimensional face of the tour guide, and storing the three-dimensional face in the display terminal;
[0028] synthesizing explanation voice information according to the voice clip of the tour guide and the text information that the tour guide needs to explain about the tourist site;
[0029] The display terminal synthesizes the three-dimensional model of the tour site, the three-dimensional face of the tour guide, and the explanation voice information into the three-dimensional tour scene video.
[0030] Optionally, the sensing device senses whether the tour guide is present at the tourist site, including:
[0031] The sensing device captures the portraits of people in the tourist site;
[0032] Comparing the photographed head portrait with the head portrait of the tour guide stored in the display terminal;
[0033] If there is a head portrait among the photographed heads that is consistent with the head portrait of the tour guide, it is determined that the tour guide is at the tourist site.
[0034] Optionally, the determining whether the instructor is in an explanation state includes:
[0035] Determining whether the number of people in the tourist site other than the tour guide is greater than or equal to 1, and determining whether there is any audio related to the tour site in the audio of the tourist site;
[0036] If yes, it is determined that the lecturer is in the lecture state;
[0037] Alternatively, determining whether the stay time of the person at the tourist site is greater than or equal to a first set time;
[0038] If so, it is determined that the lecturer is in the lecture state.
[0039] Optionally, when the display terminal intercepts the video of the guide's explanation from the video of the tour site, the start time of the intercepted video is customized to be when the number of people in the tour site other than the guide is greater than or equal to 1 and the guide's explanation voice begins to appear in the video of the tour site;
[0040] The end time of the intercepted video is customized to the time when people disappear from the tour site and there is no sound.
[0041] Optionally, the step of establishing a three-dimensional model of the tourist site in advance includes:
[0042] Using a depth camera to photograph the tourist site to obtain a depth image of the tourist site;
[0043] Performing denoising on the depth image of the tourist site;
[0044] Estimating the depth camera pose and unifying the depth images taken by the depth camera at different poses;
[0045] The unified depth image is fused into the reconstructed three-dimensional model.
[0046] Optionally, the step of establishing a three-dimensional model of the tourist site in advance, after fusing the unified depth image into the reconstructed three-dimensional model, further comprises:
[0047] Color texture information is added to the reconstructed three-dimensional model.
[0048] Optionally, collecting the tour guide's head portrait in advance and synthesizing the tour guide's three-dimensional face includes:
[0049] Use camera equipment to obtain facial images from different viewpoints and build a general three-dimensional face mesh model;
[0050] Extracting facial feature points from the facial images of different viewpoints;
[0051] Calculating the point positions of the facial feature points in three-dimensional space, and deforming the general three-dimensional face mesh model based on the point positions to establish a geometric model of the face;
[0052] The texture image of the face is synthesized based on the face images of different viewpoints and texture mapping is performed to establish a three-dimensional face with a realistic feeling.
[0053] Optionally, synthesizing the explanation voice information based on the voice clips of the tour guide and the text information that the tour guide needs to explain about the tourist site includes:
[0054] Collecting the audio clips of the tour guide in advance;
[0055] Extracting the voice feature information of the narrator from the voice segment;
[0056] Extracting a text vector from the text information that the tour guide needs to explain about the tourist site;
[0057] Combining the sound feature information and the text vector into a speech spectrum;
[0058] The speech spectrum is converted into the explanation speech information.
[0059] Optionally, the sensing device senses whether there are tourists but no guide at the tourist site, including:
[0060] The sensing device collects facial information of people in the tourist site;
[0061] Comparing the facial information with the facial information of the tour guide stored in the display terminal;
[0062] If the comparison results are inconsistent, it is determined that there are tourists at the tourist site but no tour guide.
[0063] Optionally, the determining whether the tourist needs an explanation of the tourist site includes:
[0064] The sensing device identifies whether the tourist is looking at a target object in the tourist site;
[0065] and / or, the sensing device identifies whether the visitor stays in front of the target object for more than a second set time;
[0066] If at least one recognition result is yes, it is determined that the tourist needs an explanation of the tourist site.
[0067] The present invention also provides a three-dimensional tour scene display system, wherein the three-dimensional tour scene is a three-dimensional tour scene created based on a tour site, comprising:
[0068] A display terminal is placed in the tourist site;
[0069] A sensing device is configured in the tourist site; the sensing device is coupled to the display terminal;
[0070] The sensing device is used to sense whether there are tourists and no tour guide at the tourist site, and to determine whether the tourists need a tour guide at the tourist site;
[0071] The display terminal is used to prepare a three-dimensional tour scene video in advance; and is also used to play the three-dimensional tour scene video when the sensing judgment result of the sensing device is yes;
[0072] The three-dimensional tour scene video includes the three-dimensional tour scene and the tour guide's explanation of the tour site.
[0073] Optionally, the three-dimensional tour scene display system runs the above-mentioned three-dimensional tour scene display method;
[0074] The display terminal includes: a voice collector, a sound feature encoder, a text vector generator, a speech synthesizer and a vocoder;
[0075] The speech collector, the sound feature encoder, the speech synthesizer and the vocoder are connected in sequence; the text vector generator is connected to the speech synthesizer;
[0076] The voice collector is used to collect the voice clips of the tour guide;
[0077] The sound feature encoder is used to receive the voice segment of the lecturer and extract the voice feature information of the lecturer from the voice segment;
[0078] The text vector generator is configured to receive input text information that the tour guide needs to explain about the tourist site, and extract a text vector from the text information;
[0079] The speech synthesizer is configured to receive the sound feature information and the text vector, and synthesize the sound feature information and the text vector into a speech spectrum;
[0080] The vocoder is used to receive the speech spectrum and convert the speech spectrum into the explanation speech information.
[0081] The present invention also provides a display terminal, comprising the above-mentioned three-dimensional tour scene display system.
[0082] The beneficial effects of the present invention are as follows: the display method and display system of the three-dimensional tour scene provided by the present invention can enable tourists to have a better immersive tour experience in the tour site by producing a three-dimensional tour scene video on a display terminal and displaying it through different three-dimensional display modes that can be realized by the display terminal. At the same time, by sensing whether there is a tour guide in the tour site and whether the tourists need an explanation through the sensing device, the display terminal can be intelligently controlled to play the three-dimensional tour scene video in a timely manner, thereby solving the technical problem of how to enable tourists to better tour and view the exhibition when there is no tour guide in the tour site, and improving the tourists' tour and exhibition experience.
[0083] The display terminal provided by the present invention solves the technical problem of how to enable tourists to better tour and view exhibitions when there are no tour guides at the tour site, thereby improving the tourists' tour and exhibition viewing experience, by adopting the above-mentioned three-dimensional tour scene display system. BRIEF DESCRIPTION OF THE DRAWINGS
[0084] Figure 1 A flowchart of a method for displaying a three-dimensional tour scene provided by an embodiment of the present invention;
[0085] Figure 2 A flow chart for producing a 3D tour scene video on a display terminal in advance;
[0086] Figure 3 Another flow chart for producing a 3D tour scene video on a display terminal in advance;
[0087] Figure 4 Another flow chart for producing a 3D tour scene video on a display terminal in advance;
[0088] Figure 5 A flow chart for building a 3D model of the tour site in advance;
[0089] Figure 6 Flowchart for synthesizing the narrator's 3D face;
[0090] Figure 7 A flow chart for synthesizing explanation voice information based on the explanation voice clips and the text information that the explanation needs to explain about the tour site;
[0091] Figure 8 This is a block diagram showing the principle of synthesizing explanation voice information by a terminal in an embodiment of the present invention. DETAILED DESCRIPTION
[0092] In order to enable those skilled in the art to better understand the technical solution of the present invention, the following further describes in detail a three-dimensional tour scene display method, display system, and display terminal of the present invention in conjunction with the drawings and specific embodiments.
[0093] The terminal display screens of the above-mentioned museums and other tourist sites can only play two-dimensional video explanation content. The two-dimensional video can only be configured as pictures or two-dimensional video images of the museums and other tourist sites, and the explanation content is usually played in text form or concise voice form, which makes the explanation content relatively simple and concise. On the one hand, this cannot provide tourists with a better immersive tour experience, and on the other hand, the understanding of the tourist site is limited to a simple understanding.
[0094] The present invention aims at solving the above problems existing in current tourist sites such as museums and provides a method for displaying a three-dimensional tourist scene. The three-dimensional tourist scene is a three-dimensional tourist scene created based on the tourist site, such as Figure 1 Shown, including:
[0095] Step S01: Create a 3D tour scene video on a display terminal in advance. The 3D tour scene video includes the 3D tour scene and the tour guide's explanation of the tour site; the tour site is equipped with a display terminal and a sensing device.
[0096] Step S02: The sensing device senses whether there are tourists and no tour guide at the tourist site, and determines whether the tourists need a tour guide at the tourist site.
[0097] If yes, then step S03 is executed: controlling the display terminal to play the 3D tour scene video. If no, the display terminal will not play the 3D tour scene video.
[0098] The tourist venue can be a museum, exhibition hall, zoo, botanical garden, tourist attraction, scenic spot, or other tourist destinations. The display terminal can be a terminal display screen capable of displaying three-dimensional video (i.e., a 3D display screen). Visitors can wear 3D glasses to view the 3D video, thereby obtaining a three-dimensional visual experience. The display terminal can also be VR glasses (i.e., virtual reality head-mounted display devices). VR glasses use head-mounted display devices to block the user's vision and hearing from the outside world, guiding the user to experience a virtual environment. The display principle is that the left and right screens display the images for the left and right eyes respectively. The human eye receives this differential information and creates a three-dimensional perception in the mind. The display terminal can also implement AR display (i.e., augmented reality display), also known as mixed display. This uses computer technology to apply virtual information to the real world, superimposing the real environment and virtual objects in real time on the same screen or space. The sensing device can be a camera, video camera, or other device capable of capturing images or audio and video.
[0099] In this embodiment, by producing a three-dimensional tour scene video on the display terminal and displaying it through different three-dimensional display modes that the display terminal can achieve, tourists can have a better immersive tour experience in the tour site. At the same time, by using a sensing device to sense whether there is a tour guide at the tour site and whether the tourists need a tour guide, the display terminal can be intelligently controlled to play the three-dimensional tour scene video in a timely manner, thereby solving the technical problem of how to enable tourists to better tour and view the exhibition when there is no tour guide at the tour site, and improving the tourists' tour and exhibition experience.
[0100] Optionally, the display terminal is configured with a recording mode; a 3D tour scene video is produced on the display terminal in advance, such as Figure 2 Shown, including:
[0101] Step S101: Start the recording mode of the presentation terminal.
[0102] Step S102: The display terminal records the tour guide's explanation video of the tour site.
[0103] Step S103: the display terminal converts the explanation video of the tour site into a three-dimensional tour scene video.
[0104] Optionally, the recording mode of the presentation terminal can be enabled through a remote control, gesture recognition, or voice control.
[0105] Optionally, the display terminal is configured with a recording mode; a 3D tour scene video is produced on the display terminal in advance, such as Figure 2 Shown, including:
[0106] Step S101 ′: the display terminal records the video of the tour site in real time.
[0107] Step S102 ′: the sensing device senses whether there is a tour guide at the tourist site.
[0108] If yes, then execute step S103': determine whether the lecturer is in the lecture state.
[0109] If yes, step S104' is executed: the display terminal intercepts the video of the tour guide's explanation from the tour site video, and converts the video of the tour guide's explanation into a three-dimensional tour scene video.
[0110] Among them, the above two methods of producing three-dimensional tour scene videos on the display terminal in advance are autonomous explanation modes that the tour guide can choose to produce on the display terminal. The tour guide can choose any one of the autonomous explanation modes. After each autonomous explanation mode is selected, the three-dimensional tour scene video is produced in advance according to the above steps, and the three-dimensional tour scene video is stored in the display terminal for subsequent use.
[0111] Optionally, before the tour guide selects any of the autonomous tour modes, the method for displaying the three-dimensional tour scene may further include: Figure 2 As shown,
[0112] Step S100: The sensing device recognizes the face of a person who selects a mode on the display terminal, and determines whether the person who selects the mode on the display terminal is a tour guide.
[0113] If so, the tour guide can choose to enter the display terminal function management, and then enter the autonomous explanation mode selection function of the display terminal, and execute the above steps S101-step S103 or execute the above steps S101'-step S104' to realize the advance production of the three-dimensional tour scene video on the display terminal.
[0114] Optionally, in step S100, the sensing device may capture the face of the person who selects the mode on the display terminal and compare the face with the face of the tour guide stored in the display terminal. If the comparison is consistent, it is determined that the person who selects the mode on the display terminal is the tour guide.
[0115] Optionally, step S102': the sensing device senses whether there is a guide at the tourist site, including:
[0116] The sensing device captures the portraits of people in the tourist area;
[0117] Compare the captured headshot with the headshot of the tour guide stored in the display terminal;
[0118] If there is a portrait among the photographed heads that matches the guide's portrait, it is confirmed that there is a guide at the tour site.
[0119] Optionally, step S103′: determining whether the lecturer is in a lecture state, includes:
[0120] Determine whether the number of people other than the tour guide in the tour site is greater than or equal to 1, and determine whether there is any tour guide voice related to the tour site among the sounds in the tour site;
[0121] If yes, it is determined that the lecturer is in the lecture state;
[0122] Alternatively, determining whether the stay time of the person at the tourist site is greater than or equal to a first set time;
[0123] If so, it is determined that the lecturer is in the lecture state.
[0124] Alternatively, the system can determine whether the audio content of a visitor's visit contains relevant explanations of the site. For example, in a museum, the system can determine whether the audio content is related to key information such as the names and historical dates of the museum's artifacts. This determination can be achieved through natural language processing (NLP) technology. Natural language processing technology is a field that intersects computer science, artificial intelligence, and linguistics. Its goal is to enable computers to process or "understand" natural language to perform tasks such as language translation and question answering.
[0125] Optionally, the first set time may be a reasonable set time length such as 10 seconds or 20 seconds.
[0126] Optionally, step S104': when the display terminal intercepts the video of the tour guide's explanation from the video of the tour site, the start time of the intercepted video is customized to when the number of people in the tour site other than the tour guide is greater than or equal to 1, and the tour guide's explanation voice begins to appear in the video of the tour site; the end time of the intercepted video is customized to when the people in the tour site disappear and there is no sound.
[0127] Optionally, a 3D tour scene video is produced on the display terminal in advance, such as Figure 3 Shown, including:
[0128] Step S201: Create a three-dimensional model of the tourist site in advance and store the three-dimensional model in a display terminal;
[0129] Step S202: collecting the tour guide's head portrait in advance, synthesizing the tour guide's three-dimensional face, and storing the three-dimensional face in the display terminal;
[0130] Step S203: The presentation terminal collects the tour guide's voice, movements, and facial expressions in real time.
[0131] Step S204: the display terminal combines the three-dimensional model of the tour site, the three-dimensional face of the tour guide, and the tour guide's voice, movements, and facial expressions during the tour into a three-dimensional tour scene video.
[0132] Optionally, a 3D tour scene video is produced on the display terminal in advance, such as Figure 4 Shown, including:
[0133] Step S301: Create a three-dimensional model of the tourist site in advance and store the three-dimensional model in a display terminal;
[0134] Step S302: collecting the tour guide's head portrait in advance, synthesizing the tour guide's three-dimensional face, and storing the three-dimensional face in the display terminal;
[0135] Step S303: synthesizing explanation voice information based on the explanation voice segment of the tour guide and the text information that the tour guide needs to explain about the tour site;
[0136] Step S304: the display terminal synthesizes the three-dimensional model of the tour site, the three-dimensional face of the tour guide, and the explanation voice information into a three-dimensional tour scene video.
[0137] Among them, in step S201 and step S301, a three-dimensional model of the tourist site is established in advance, such as Figure 5 Shown, including:
[0138] Step S2301: Use a depth camera to shoot the tourist site to obtain a depth image of the tourist site.
[0139] In this step, a handheld depth camera can be used to scan the tourist site to obtain a depth image of the tourist site.
[0140] Step S2302: Denoising the depth image of the tourist site.
[0141] In this step, noise in the depth image is categorized into three types: missing depth, caused by factors such as images being too close or too far away, surface discontinuities, highlights, or shadows; incorrect depth, which indicates that the depth measurement has a certain degree of accuracy; and inconsistent depth, which indicates that the depth measured at the same point may be inconsistent over time. Bilateral filtering can be used to remove noise from the depth image.
[0142] After denoising, KinectFusion parsing begins real-time 3D reconstruction using an RGBD camera. The KinectFusion parsing algorithm consists of four steps: first, processing the acquired raw depth image to obtain the coordinates and normal vectors of the point cloud voxels; then, computing the current camera position and pose based on the point cloud of the current frame and the predicted point cloud from the previous frame; then, updating the TSDF values based on the camera position and pose, fusing the point clouds; and finally, estimating the surface based on the TSDF values. KinectFusion parsing uses downsampling to create a three-layer depth image pyramid for subsequent camera pose estimation. KinectFusion evenly divides a fixed-size space (e.g., 3m×3m×3m) into smaller blocks (e.g., 512×512×512). Each block is a voxel, storing the TSDF values and weights. The resulting 3D reconstruction is a linear interpolation of these voxels.
[0143] Step S2303: Estimate the depth camera pose and unify the depth images captured by the depth camera at different poses.
[0144] In this step, the typical approach is to find point correspondences and then estimate the transformation matrix. The camera pose typically refers to a six-degree-of-freedom transformation, represented by a rigid body transformation matrix T. ICP (Iterative Closest Point) is a very important algorithm for camera pose estimation. It is a common method for processing point clouds. By minimizing the difference between two point clouds, iteratively solves for the relative position of the cameras capturing them. There are different ways to describe the difference between point clouds, the most common being point-to-point and point-to-plane. KinectFusion parsing uses the point-to-plane approach, projecting point-to-point distances onto the normal vector. ICP is primarily used for 3D shape registration. By calculating the matching relationships between point clouds in adjacent frames and minimizing the Euclidean distance between point pairs, a rigid body transformation is calculated. However, this approach presents a problem: errors between adjacent frames accumulate during the scanning process, often referred to as cumulative error.
[0145] To eliminate the problem of cumulative error, a frame-to-model camera tracking method is adopted. That is, each time the entire reconstructed model of the current frame is registered, rather than being registered with the previous frame. This method reduces drift during camera tracking to a certain extent, but it does not completely solve the problem of cumulative error. That is, the error in camera pose estimation continues to accumulate over time. This accumulated drift can eventually prevent loop closure. Therefore, a global pose optimization method is proposed, and the concept of keyframes is introduced. Whenever the cumulative error is greater than a threshold, the current frame is selected as the keyframe for loop detection. The pose is jointly optimized using the pose graph and sparse bundle adjustment (BA). Depth information and color information are used for global pose estimation. In addition to the above-mentioned registration process, matching point pairs is also an important step in camera tracking. It can be divided into sparse and dense methods based on the number of points used: sparse methods only use feature points for matching, while dense methods use all points for matching.
[0146] Sparse point pair matching: SIFT, SURF, ORB. BundleFusion parsing uses SIFT for coarse registration, followed by dense methods for fine registration. Dense point pair matching: Traditional point pair matching methods are too time-consuming, while projection data association algorithms are faster.
[0147] The main process of step S2303 is to project the input 3D point coordinates onto the pixels of the target depth map based on the camera pose, and then take the 3D point corresponding to the nearest neighbor pixel as the target point corresponding to the input point. There are different calculation methods for measuring the distance between the input point and the target point, such as point-to-point and point-to-plane. Point-to-plane is to calculate the distance between the two points projected onto the normal vector, which converges faster. In addition to distance error, other distance measurement methods such as photometric error can also be introduced.
[0148] Step S2304: Fusing the unified depth image into the reconstructed 3D model.
[0149] In this step, the depth maps of the stereo representation are fused by taking the weighted sum of the TSDF of the current frame and the global TSDF. The SDF (Signed Distance Function) describes the distance from a point to a surface and is 0 on the surface, positive on one side of the surface, and negative on the other. The TSDF (Truncated SDF) only considers the SDF value within the neighborhood of the surface. If the maximum value of the neighborhood is the max truncation, the actual distance is divided by the max truncation value to achieve normalization. Therefore, the TSDF value is between -1 and +1.
[0150] Depth map fusion of patch representation: Each vertex, normal vector, and radius of the current frame must be integrated into the global model. This mainly involves three steps:
[0151] 1. Project the vertices in the current 3D model onto the image plane of the current frame camera and find matching point pairs;
[0152] 2. If a matching point pair is found, the most reliable point and the new point are weighted averaged; if no matching point pair is found, the new point is added to the global model as an unstable point;
[0153] 3. As more and more frames are processed, the global model will clean up outliers.
[0154] Optionally, a three-dimensional model of the tourist site is created in advance, and after step S2304: fusing the unified depth image into the reconstructed three-dimensional model, the following steps may be further performed:
[0155] Step S2305: Add color texture information to the reconstructed three-dimensional model.
[0156] In this step, the goal of offline texture reconstruction is to reconstruct high-quality, globally consistent textures for the 3D model using multi-view RGB images. Directly fusing different images can cause ghosting and over-smoothing, which are typically addressed through optimization algorithms. Common optimization objectives include color consistency, alignment of images with geometric features, and maximizing mutual information between projected images.
[0157] The goal of online texture reconstruction is to generate high-quality textures while scanning and reconstructing objects. The advantage is that users can observe the texture of the object in real time.
[0158] Optionally, in step S202 and step S302, the head portrait of the tour guide is collected in advance and the three-dimensional face of the tour guide is synthesized, such as Figure 6 Shown, including:
[0159] Step S3201: Use a camera to obtain facial images from different viewpoints and establish a general three-dimensional face mesh model.
[0160] In this step, the imaging device is a device such as a video camera or a still camera that can capture facial images.
[0161] Step S3202: extracting facial feature points from facial images of different viewpoints.
[0162] In this step, facial feature points such as the corners of the eyes, corners of the mouth, and the tip of the nose are used to identify features of the human face.
[0163] Step S3203: Calculate the point positions of facial feature points in three-dimensional space, and deform the general three-dimensional face mesh model based on the point positions to establish a geometric model of the face.
[0164] Step S3204: synthesizing a facial texture image based on facial images from different viewpoints and performing texture mapping to create a realistic three-dimensional face.
[0165] Optionally, Figure 4 In the method of making a three-dimensional tour scene video on a display terminal in advance, the tour guide can be a staff member of the tour site or other non-staff member, such as a singer, movie star or other person that tourists like, that is, the three-dimensional face can be the face of a singer, movie star or other person that tourists like.
[0166] Optionally, step S303: synthesize the explanation voice information according to the voice clip of the tour guide and the text information that the tour guide needs to explain about the tour site, such as Figure 7 Shown, including:
[0167] Step S3031: Collect the audio clips of the tour guide in advance.
[0168] In this step, the voice clips of the tour guide are collected by a voice collector such as a recording device.
[0169] Step S3032: extracting the voice feature information of the tour guide from the voice segment.
[0170] In this step, the voice feature information of the speaker is extracted through a voice feature encoder. The voice feature encoder is a voice feature extraction model obtained through pre-training.
[0171] Step S3033: extracting text vectors from the text information that the tour guide needs to explain about the tourist site.
[0172] In this step, text vectors are extracted using a text vector generator. The text vector generator is a text vector extraction model obtained through pre-training.
[0173] Step S3034: synthesize the sound feature information and the text vector into a speech spectrum.
[0174] In this step, the sound feature information and the text vector are synthesized into a speech spectrum by a speech synthesizer. The speech synthesizer is a speech synthesis model obtained by training the sound feature extraction model and the text vector extraction model.
[0175] Step S3035: Convert the speech spectrum into explanation speech information.
[0176] In this step, the speech spectrum is converted into explanation speech information through a vocoder. The vocoder is a speech spectrum conversion model obtained by training the speech synthesis model.
[0177] Optionally, Figure 4In the method for synthesizing explanation voice information based on the voice clips of the tour guide and the text information that the tour guide needs to explain about the tourist site, the tour guide can be a staff member of the tourist site or other non-staff member, such as a singer, movie star or other person that tourists like, that is, the explanation voice information can be the explanation voice information of the singer, movie star or other person that tourists like.
[0178] Optionally, in step S02, the sensing device senses whether there are tourists and no guides at the tourist site, including: the sensing device collects facial information of people at the tourist site;
[0179] Compare the facial information with the facial information of the tour guide stored in the display terminal;
[0180] If the comparison results are inconsistent, it is determined that there are tourists but no tour guide at the tour site.
[0181] Optionally, in step S02, determining whether the tourist needs a tour guide for the site includes:
[0182] The sensing device identifies whether the tourist is looking at the target object in the tourist area;
[0183] and / or, the sensing device identifies whether the visitor stays in front of the target object for more than a second set time;
[0184] If at least one recognition result is yes, it is determined that the tourist needs a tour guide for the site.
[0185] The second set time can be a reasonable set time length such as 30 seconds or 40 seconds.
[0186] The three-dimensional tour scene display method provided in this embodiment can provide tourists with a better immersive tour experience at the tour site by producing a three-dimensional tour scene video on a display terminal and displaying it through different three-dimensional display modes that can be realized by the display terminal. At the same time, by using a sensing device to sense whether there is a tour guide at the tour site and whether the tourists need a tour guide, the display terminal can be intelligently controlled to play the three-dimensional tour scene video in a timely manner, thereby solving the technical problem of how to enable tourists to better tour and view the exhibition when there is no tour guide at the tour site, thereby improving the tourists' tour and exhibition viewing experience.
[0187] An embodiment of the present invention also provides a three-dimensional tour scene display system, in which the three-dimensional tour scene is a three-dimensional tour scene created based on a tour site, and includes: a display terminal, which is placed in the tour site; a sensing device, which is configured in the tour site; the sensing device is coupled to the display terminal; the sensing device is used to sense whether there are tourists and no tour guide in the tour site, and to determine whether the tourists need an explanation of the tour site; the display terminal is used to prepare a three-dimensional tour scene video in advance; and is also used to play the three-dimensional tour scene video when the sensing judgment result of the sensing device is yes; the three-dimensional tour scene video includes the three-dimensional tour scene and the tour guide's explanation of the tour site.
[0188] "Coupled" can mean that two or more components are in direct physical or electrical contact; it can also mean that two or more components are not in direct contact but still cooperate or interact with each other. Each sub-venue within a tourist attraction, such as a museum, can be equipped with one or more display terminals and one or more sensing devices.
[0189] Alternatively, as Figure 8 As shown, the display terminal includes: a voice collector, a sound feature encoder, a text vector generator, a speech synthesizer and a vocoder; the voice collector, the sound feature encoder, the speech synthesizer and the vocoder are connected in sequence; the text vector generator is connected to the speech synthesizer; the voice collector is used to collect the voice clips of the tour guide; the sound feature encoder is used to receive the voice clips of the tour guide and extract the voice feature information of the tour guide from the voice clips; the text vector generator is used to receive the input text information that the tour guide needs to explain about the tourist site and extract the text vector from the text information; the speech synthesizer is used to receive the sound feature information and the text vector and synthesize the sound feature information and the text vector into a speech spectrum; the vocoder is used to receive the speech spectrum and convert the speech spectrum into the explanation speech information.
[0190] The 3D tour scene display system can provide tourists with a better immersive tour experience at the tour site by producing a 3D tour scene video on a display terminal and displaying it through different 3D display modes that the display terminal can realize. At the same time, by using a sensing device to sense whether there is a tour guide at the tour site and whether the tourists need a tour guide, the display terminal can be intelligently controlled to play the 3D tour scene video in a timely manner, thereby solving the technical problem of how to enable tourists to better tour and view exhibitions when there is no tour guide at the tour site, thereby improving the tourists' tour and exhibition viewing experience.
[0191] An embodiment of the present invention further provides a display terminal, comprising the three-dimensional tour scene display system in the above embodiment.
[0192] The display terminal in the three-dimensional tour scene display system may be a display of a display terminal, and the sensing device in the three-dimensional tour scene display system may be a camera configured on the display.
[0193] By adopting the three-dimensional tour scene display system in the above embodiment, the display terminal solves the technical problem of how to enable tourists to better tour and view exhibitions when there is no tour guide at the tour site, thereby improving the tourists' tour and exhibition viewing experience.
[0194] The display terminal provided by the present invention can be any product or component with a display function, such as an LCD panel, an LCD TV, an OLED panel, an OLED TV, an LED panel, an LED TV, a monitor, a mobile phone, a navigator, or the like.
[0195] It will be understood that the above embodiments are merely exemplary embodiments for illustrating the principles of the present invention, and the present invention is not limited thereto. Those skilled in the art will appreciate that various modifications and improvements can be made without departing from the spirit and substance of the present invention, and such modifications and improvements are also considered to be within the scope of protection of the present invention.
Claims
1. A method for displaying a three-dimensional tour scene, wherein the three-dimensional tour scene is a three-dimensional tour scene created based on a tour site, characterized in that: include: Prepare a 3D tour scene video on the display terminal in advance; The three-dimensional tour scene video includes the three-dimensional tour scene and the tour guide's explanation of the tour site; the tour site is equipped with the display terminal and the sensing device; The sensing device senses whether there are tourists at the tourist site and no guide, and determines whether the tourists need an explanation at the tourist site; If yes, controlling the display terminal to play the three-dimensional tour scene video; The display terminal is configured with a recording mode; The step of producing a 3D tour scene video on a display terminal in advance includes: The display terminal records the video of the tour site in real time; The sensing device senses whether the tour guide is present at the tourist site; If yes, then determine whether the lecturer is in the lecture state; If yes, the display terminal intercepts the video of the tour guide's explanation from the tour site video, and converts the video of the tour guide's explanation into the three-dimensional tour scene video; The determining whether the lecturer is in the lecture state includes: Determining whether the number of people in the tourist site other than the tour guide is greater than or equal to 1, and determining whether there is any audio related to the tour site in the audio of the tourist site; If yes, it is determined that the lecturer is in the lecture state; Alternatively, determining whether the stay time of the person at the tourist site is greater than or equal to a first set time; If yes, it is determined that the lecturer is in the lecture state; When the display terminal intercepts the video of the guide's explanation from the video of the tourist site, the start time of the intercepted video is customized to be when the number of people in the tourist site other than the guide is greater than or equal to 1, and the guide's explanation voice begins to appear in the video of the tourist site; The end time of the intercepted video is customized to the time when people disappear from the tour site and there is no sound.
2. The method for displaying a three-dimensional tour scene according to claim 1, characterized in that: The sensing device senses whether the tour guide is present at the tour site, including: The sensing device captures the portraits of people in the tourist site; Comparing the photographed head portrait with the head portrait of the tour guide stored in the display terminal; If there is a head portrait among the photographed heads that is consistent with the head portrait of the tour guide, it is determined that the tour guide is at the tourist site.
3. The method for displaying a three-dimensional tour scene according to claim 1, characterized in that: The sensing device senses whether there are tourists in the tourist site and there is no guide, including: The sensing device collects facial information of people in the tourist site; Comparing the facial information with the facial information of the tour guide stored in the display terminal; If the comparison results are inconsistent, it is determined that there are tourists at the tourist site but no tour guide.
4. The method for displaying a three-dimensional tour scene according to claim 1, characterized in that: The determining whether the tourist needs an explanation of the tourist site includes: The sensing device identifies whether the tourist is looking at a target object in the tourist site; and / or, the sensing device identifies whether the visitor stays in front of the target object for more than a second set time; If at least one recognition result is yes, it is determined that the tourist needs an explanation of the tourist site.
5. A three-dimensional tour scene display system, wherein the three-dimensional tour scene is a three-dimensional tour scene created based on a tour site, and the three-dimensional tour scene display system runs the three-dimensional tour scene display method according to any one of claims 1 to 4, characterized in that: The three-dimensional tour scene display system includes: A display terminal is placed in the tourist site; A sensing device is configured in the tourist site; the sensing device is coupled to the display terminal; The sensing device is used to sense whether there are tourists and no tour guide at the tourist site, and to determine whether the tourists need a tour guide at the tourist site; The display terminal is used to prepare a three-dimensional tour scene video in advance; and is also used to play the three-dimensional tour scene video when the sensing judgment result of the sensing device is yes; The three-dimensional tour scene video includes the three-dimensional tour scene and the tour guide's explanation of the tour site.
6. A display terminal, characterized in that: Including the three-dimensional tour scene display system described in claim 5.
Citation Information
Patent Citations
Scenario video recording system and method based on 3D virtual synthesis technology and scenario training learning method
CN105306862A
Museum explanation system based on electronic induction and tourist behavior analysis
CN108597417A
Voice-controlled automatic intelligent demonstration exhibiting hall
CN109377921A
Three -dimensional holographic virtual guide display system
CN205050534U