Mixed reality image method, device and equipment based on Mesh networking and storage medium
By acquiring and processing multi-angle stage camera data, identifying and reconstructing dynamic objects, mixing reality synthesis, and using Mesh network transmission, the problem of insufficient personalized experience and interactivity in the existing technology is solved, and high-quality mixed reality image generation and user interaction experience are achieved.
Patent Information
- Application Number
- CN202510197092.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-05-30
AI Technical Summary
The existing technology has shortcomings in personalized experience and interactivity, and cannot dynamically adjust audio-visual content, lacks in-depth interaction, and the picture integration between the cloud and on-site equipment has not yet achieved efficient collaboration, which limits the real-time and flexibility of stage effects.
By obtaining multi-angle stage camera data, converting it into streaming media and uploading it to the cloud, identifying dynamic objects for film removal and stripping and three-dimensional reconstruction, mixed reality synthesis, and using Mesh network to send mixed reality images.
It realizes high-quality mixed reality image generation, enhances the user's audio-visual experience and interactivity, and provides an immersive, interactive and highly customized experience.
Smart Images

Figure CN120070815A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of mixed reality, and particularly to a mixed reality imaging method, apparatus, device, and computer-readable storage medium based on Mesh networking. Background Art
[0002] In recent years, the rapid development of mobile communication and wireless services has provided new technical support for live performances and audio-visual experiences. Through intelligent devices and network connections, users can not only obtain high-quality audio-visual content in real time, but can even obtain immersive experiences through technologies such as mixed reality (MR) and augmented reality (AR). At the same time, the combination of cloud technology and projection devices has made the presentation of complex stage scenes more feasible, thereby reducing the high cost of physical set construction. However, the rich scene applications have also driven the demand for diverse audio-visual experiences, and there are still many deficiencies in the existing technical solutions in terms of personalized experience and interactivity. First, existing audio-visual devices cannot dynamically adjust audio-visual content according to the preferences and needs of different users, resulting in a low level of personalization of the user experience. Second, the current stage effects and live broadcast technologies lack in-depth interaction with the audience and are difficult to meet the user's needs for a sense of participation and immersion. In addition, the integration of the cloud and on-site device images has not achieved efficient coordination, resulting in limitations in the real-time performance and flexibility of the stage effects.
[0003] Therefore, how to achieve more intelligent mixed reality imaging generation while ensuring efficient transmission, and enhance the user's audio-visual experience and interactivity is an issue to be solved currently. Summary of the Invention
[0004] According to the embodiments of the present application, a mixed reality imaging solution based on Mesh networking is provided, which can achieve more intelligent mixed reality imaging generation and greatly enhance the user's audio-visual experience and interactivity.
[0005] In the first aspect of the present application, a mixed reality imaging method based on Mesh networking is provided. The method includes: Obtaining multi-angle stage camera data, converting the multi-angle stage camera data into a streaming media, and uploading the streaming media to the cloud; Identifying dynamic objects in the streaming media according to an arrangement instruction, performing film stripping and three-dimensional reconstruction on the dynamic objects to obtain dynamic virtual images; Performing mixed reality synthesis based on the dynamic virtual images and the streaming media to construct a mixed reality image; Sending the mixed reality image to a first display space by the cloud through Mesh networking.
[0006] In a possible implementation, dynamic objects in the streaming media are identified according to the choreography instructions, and the dynamic objects are subjected to film stripping and three-dimensional reconstruction to obtain dynamic virtual images, including: By using a pre-trained YOLO model, object detection and extraction are performed on the images in the streaming media according to the choreography instructions to obtain dynamic objects; By using a pre-trained Gaussian-SLAM model, film stripping and three-dimensional reconstruction are performed on the dynamic objects to obtain dynamic virtual images.
[0007] In a possible implementation, the method further includes: Obtain the stage projection instruction and the audience emotion data, and generate a dynamic digital human interaction image according to the stage projection instruction and the audience emotion data; The cloud sends the dynamic digital human interaction image to the second display space through Mesh networking.
[0008] In a possible implementation, the method further includes: Obtain the user somatosensory data of the smart terminal and the coordinate data of the smart terminal located by Mesh networking, and generate user personalized display materials according to the user somatosensory data and the coordinate data; The cloud sends the user personalized display materials to the third display space through Mesh networking.
[0009] Optionally, the method for obtaining the coordinate data of the smart terminal located by Mesh networking includes: Calculate the azimuth angle and the horizontal distance from the smart terminal to the Mesh networking according to the signal strength; Calculate the vertical distance of the smart terminal according to the gravity sensor in the smart terminal; Determine the three-dimensional coordinate data of the smart terminal according to the vertical distance, the azimuth angle from the smart terminal to the Mesh networking, and the horizontal distance.
[0010] In a possible implementation, the method further includes: Obtain the smart terminal projection instruction, segment the projection screen in the smart terminal projection instruction to obtain smart terminal pixel blocks; According to the smart terminal projection instruction and the coordinate position of the smart terminal, send the smart terminal pixel blocks to the smart terminal and display the smart terminal pixel blocks on the smart terminal.
[0011] In a possible implementation, the method further includes: Archive and set the playback of the video data of the stage performance, and the playback settings include time point selection and stage position selection; The cloud regenerates the mixed reality image according to the selected playback screen and the playback instruction, and sends it to the smart terminal.
[0012] In the second aspect of the present application, a mixed reality imaging device based on Mesh networking is provided. The device includes: A first acquisition module, configured to acquire multi-angle stage camera data, convert the multi-angle stage camera data into a streaming media, and upload the streaming media to the cloud; A second acquisition module, configured to identify dynamic objects in the streaming media according to an orchestration instruction, perform demembrane peeling and three-dimensional reconstruction on the dynamic objects, and obtain dynamic virtual images; A construction module, configured to perform mixed reality synthesis based on the dynamic virtual images and the streaming media, and construct mixed reality images; A sending module, which sends the mixed reality images to a first display space through Mesh networking by the cloud.
[0013] In the third aspect of the present application, an electronic device is provided. The electronic device includes: a memory and a processor. A computer program is stored on the memory, and when the processor executes the program, the method as described above is implemented.
[0014] In the fourth aspect of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the method according to the first aspect of the present application is implemented.
[0015] The mixed reality imaging method based on Mesh networking provided by the embodiments of the present application acquires multi-angle stage camera data, converts the multi-angle stage camera data into a streaming media, and uploads the streaming media to the cloud. Then, according to the orchestration instruction, dynamic objects in the streaming media are identified, demembrane peeling and three-dimensional reconstruction are performed on the dynamic objects to obtain dynamic virtual images, and then mixed reality synthesis is performed based on the dynamic virtual images and the streaming media to construct mixed reality images. Finally, the cloud sends the mixed reality images to the first display space through Mesh networking. The present application not only realizes the generation of high-quality mixed reality images, but also provides an immersive, highly interactive and highly customized experience for the audience through efficient dynamic object processing, Mesh networking transmission, dynamic digital human interaction and personalized content generation.
[0016] It should be understood that the content described in the summary of the invention section is not intended to limit the key or important features of the embodiments of the present application, nor to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Combined with the drawings and referring to the following detailed description, the above and other features, advantages and aspects of the embodiments of the present application will become more obvious. In the drawings, the same or similar reference numerals represent the same or similar elements, where: Figure 1Flowchart of a mixed reality imaging method based on Mesh networking according to an embodiment of the present application; Figure 2 Structural schematic diagram of Mesh connection according to an embodiment of the present application; Figure 3 Structural schematic diagram of image acquisition and display according to an embodiment of the present application; Figure 4 Block diagram of a mixed reality imaging device based on Mesh networking according to an embodiment of the present application; Figure 5 Structural schematic diagram of a terminal device or server suitable for implementing the embodiments of the present application. Detailed implementation manners
[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are some, but not all, of the embodiments of the present disclosure. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present disclosure without creative efforts shall fall within the scope of protection of the present disclosure.
[0019] In addition, the term "and / or" in this article is only a relationship describing associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.
[0020] Figure 1 Shows a flowchart of a mixed reality imaging method based on Mesh networking according to an embodiment of the present disclosure. Refer to Figure 1 This method includes: S101, obtain multi-angle stage camera data, convert the multi-angle stage camera data into a streaming media, and upload the streaming media to the cloud.
[0021] In this embodiment, by obtaining multi-angle stage camera data, rich stage images are provided. Then, the multi-angle stage camera data is converted into a streaming media, which provides a basis for subsequent processing of the stage images. Moreover, processing the multi-angle stage camera data in the cloud saves a large amount of storage space and computing pressure.
[0022] S102, identify dynamic objects in the streaming media according to the choreography instruction, perform film stripping and three-dimensional reconstruction on the dynamic objects, and obtain dynamic virtual images.
[0023] For example, in a stage performance, visual enhancement of the actor's performance is required, and the choreography instruction is to add a flame special effect around the actor. Then, through video analysis technology, the actor in the streaming media is recognized, and the image information of the actor is stripped from the streaming media, including but not limited to action details such as the actor's movements and expressions. Then, the separated image information of the actor is reconstructed in three dimensions, and a flame special effect is added around the actor according to the choreography instruction. First, on the basis of the three-dimensional contour of the actor, a dynamic flame special effect layer is constructed to ensure that the flame special effect is perfectly synchronized with the actor's movements. Then, the generated flame special effect is fused with the three-dimensional model of the actor through real-time rendering to generate a dynamic virtual image, ensuring the smoothness and naturalness of the special effect.
[0024] In this embodiment, the construction of the dynamic virtual image is realized through stripping and three-dimensional reconstruction, improving the overall visual effect and enhancing the audience's stage perception.
[0025] Optionally, the dynamic objects in the streaming media are recognized according to the choreography instruction, and the dynamic objects are stripped and reconstructed in three dimensions to obtain a dynamic virtual image, including: Through the pre-trained YOLO model, object detection and extraction are performed on the images in the streaming media according to the choreography instruction to obtain dynamic objects; Through the pre-trained Gaussian-SLAM model, the dynamic objects are stripped and reconstructed in three dimensions to obtain a dynamic virtual image.
[0026] Among them, the YOLO (You Only Look Once) model is a real-time object detection algorithm that can quickly and accurately detect objects in images. It segments the image into multiple grid cells, extracts the feature information in the grid cells and detects the objects in the grid cells, and then predicts the boundaries and categories of the objects in the grid cells, so as to realize the detection and extraction of dynamic objects. The Gaussian-SLAM (Gaussian-Simultaneous Localization and Mapping) model is an algorithm for simultaneous localization and mapping that can accurately reconstruct the three-dimensional structure of an object. It receives the key pins of the dynamic object, subsamples the key pins of the dynamic object and identifies the color gradient of the dynamic object, projects the sampled points into three-dimensional space, and adds a new three-dimensional Gaussian to the initial Gaussian distribution according to the initial Gaussian distribution created in three-dimensional space, so as to realize the construction of the dynamic virtual image. In addition, Gaussian-SLAM also calculates the depth and color loss of the initial Gaussian distribution and the new three-dimensional Gaussian distribution to facilitate the optimization of the rendered three-dimensional effect.
[0027] In the present invention, the pre-training data of the YOLO model includes but is not limited to the human body recognition image dataset and the object detection image dataset, and the pre-training data of the Gaussian-SLAM model includes but is not limited to the stage video dataset.
[0028] In this embodiment, by using the pre-trained YOLO model and the pre-trained Gaussian-SLAM model, more intelligent demoulding and three-dimensional reconstruction of dynamic objects are realized. Compared with the traditional virtual image construction method, the efficiency is improved and the manpower is greatly saved.
[0029] S103, perform mixed reality synthesis according to the dynamic virtual image and the streaming media to construct a mixed reality image.
[0030] Among them, the real-time rendering software for performing mixed reality synthesis of the dynamic virtual image and the streaming media includes but is not limited to 3D Studio Max and SketchUp. 3D Max is a professional 3D modeling, animation and rendering software developed by Autodesk, and is widely used in fields such as game development, film production and architectural visualization. SketchUp is an easy-to-use 3D modeling software developed by Trimble, and is famous for its intuitive interface and fast modeling ability.
[0031] In this embodiment, the dynamic virtual image is combined with the original streaming media content to construct the final mixed reality image, which enhances the visual expression and artistic effect of the stage.
[0032] S104, the cloud sends the mixed reality image to the first display space through Mesh networking.
[0033] In addition, the Mesh networking can also be shunted through the on-site WLAN network, giving full play to the multicast characteristics of the network communication equipment and the reuse of the internal network traffic, avoiding large-scale data uploading and downloading, and shunting the image data with the help of the local network terminal.
[0034] The nodes in the Mesh networking include multi-angle stage camera devices, projection devices, the cloud and intelligent terminals. The intelligent terminals include but are not limited to mobile phones, tablet computers and AR / VR glasses. Through the Mesh networking, the cloud can also send the stage program list to the intelligent terminals, which is convenient for on-site audiences and online audiences to view. And, intelligent devices such as drones can also be connected in the Mesh networking to support complex hardware interaction instructions.
[0035] Figure 2 For the structural schematic diagram of the Mesh connection according to the embodiment of the present application, as Figure 2 shown: The cloud server, as the core processing unit, can obtain real-time camera data streams from multi-angle stage camera devices through any one or more nodes in the Mesh network and receive user data from intelligent terminals. At the same time, the cloud can also send precise stage projection instructions to the corresponding projection devices to achieve the dynamic presentation of stage lighting, special effects, and virtual scenes. In addition, the Mesh network can dynamically select and configure network connections according to the actual physical locations and application requirements of multi-angle stage camera devices, projection devices, the cloud, and intelligent terminals. For example, in areas where high bandwidth and low latency are required, more camera devices and projection devices can be preferentially connected to form a denser network topology; while in the audience area, the connection between intelligent terminals and the network can be focused on optimizing to ensure a smooth mixed reality experience.
[0036] In this embodiment, the flexibility and connectivity of the Mesh network enable the stage images to adapt to different display devices and environments, improving the efficiency of stage image transmission.
[0037] In this embodiment, the distributed architecture of the Mesh network ensures the real-time transmission and synchronization of images, and can greatly reduce the latency for the audience to view the mixed reality image content.
[0038] Optionally, the method further includes: Obtaining an intelligent terminal projection instruction, segmenting the projection image in the intelligent terminal projection instruction to obtain intelligent terminal pixel blocks; According to the intelligent terminal projection instruction and the coordinate position of the intelligent terminal, sending the intelligent terminal pixel blocks to the intelligent terminal and displaying the intelligent terminal pixel blocks on the intelligent terminal.
[0039] For example, in a large-scale concert with 50,000 seats, each seat has an intelligent terminal that can establish a connection with the cloud, including but not limited to the audience's mobile phones, smart tablets, and intelligent display terminals provided by the stage contractor. Each intelligent terminal has a coordinate position matching the seat, and the intelligent terminals on adjacent seats correspond to the coordinate positions of adjacent seats. Then, the cloud segments the projection image in the intelligent terminal projection instruction, and each segmented intelligent terminal pixel block corresponds to an intelligent terminal on-site. Finally, the cloud sends each intelligent terminal pixel block to its corresponding intelligent terminal through the Mesh network according to the coordinate position of the intelligent terminal, and displays the received intelligent terminal pixel block on the intelligent terminal. Thus, the display images of the intelligent terminals in the 50,000 seats can be pieced together into a complete projection image, and each intelligent terminal pixel block corresponds to a part of the projection image.
[0040] In addition, the intelligent terminal needs to authenticate with the cloud before receiving the intelligent terminal pixel blocks. After the first successful authentication, the intelligent terminal does not need to perform secondary authentication during the transmission of intelligent terminal pixel blocks, thus simplifying the data interaction process. Moreover, each intelligent terminal only receives the intelligent terminal pixel blocks and pixel point control instructions corresponding to its own coordinate position, avoiding conflicts of pixel instructions. Furthermore, each intelligent terminal caches the intelligent terminal pixel blocks of adjacent intelligent terminals to avoid the problem that adjacent intelligent terminals lose pixel images due to failures, and provides backup pixel point control instructions or pixel images for adjacent intelligent terminals, eliminating the need to obtain them from the cloud and saving data transmission time.
[0041] In this embodiment, through an efficient data processing and reliable data transmission mechanism, large-scale stage intelligent terminal display control is achieved, enhancing the visual experience of the audience and improving the interactive attraction to the audience.
[0042] Optionally, the method further includes: Obtaining a stage projection instruction and audience emotion data, and generating a dynamic digital human interactive image based on the stage projection instruction and audience emotion data; Sending the dynamic digital human interactive image to the second display space by the cloud through Mesh networking.
[0043] Among them, the shape of the projection screen for dynamic digital human projection includes but is not limited to a four-sided spire shape, a cylindrical shape, and a spherical shape. The basic image source of the dynamic digital human can be stage elements directly related to the live stage or the image of the dynamic digital human generated based on stage actors. In addition, the projection device of the dynamic digital human is provided with a separate microphone and speaker, enabling separate interaction between the dynamic digital human and the audience after the stage performance.
[0044] For example, after the start of the stage performance, a water wave lake surface and other images can be projected according to the stage projection instruction, and the water wave image pressed by a finger is displayed in the projection display area. Then, combined with the audience emotion data, dynamic digital human actions are generated to make the image in the second display space more three-dimensional and full.
[0045] In some embodiments, when the current stage performance is a slow song, the stage projection instruction is soft, and the audience emotion data is "calm", the stage lights turn blue, and a dynamic digital human image sitting on a cloud is generated through virtual rendering, gently waving its wings and emitting a soft blue light, echoing the stage atmosphere. Finally, the generated dynamic digital human interactive image is sent to the second display space by the cloud through Mesh networking. Among them, the sources of the audience emotion data include but are not limited to audience interactive voting and the recognition and analysis of the audience's emotions through intelligent devices.
[0046] In this embodiment, a dynamic digital human interactive image is generated based on the stage projection instruction and the audience emotion data. The dynamic digital human can make corresponding responses according to the audience emotion data, realizing a higher-level audience interactive stage.
[0047] Optionally, the method further includes: Obtaining the user somatosensory data of the intelligent terminal and the coordinate data of the intelligent terminal located by Mesh networking, and generating user personalized display materials according to the user somatosensory data and the coordinate data; Sending the user personalized display materials to the third display space by the cloud through Mesh networking.
[0048] Among them, the user somatosensory data includes but is not limited to body temperature and heart rate.
[0049] In a possible implementation manner, the coordinate data of the intelligent terminal located by Mesh networking is row 3, seat 4, and the heart rate in the user somatosensory data of the intelligent terminal at this position obtained is the highest. The preset rendering picture in the stage audience area is a virtual giant dragon flying across the sky. Then, fixed-point actions can be performed over the position of the user with a higher heart rate according to the user somatosensory data.
[0050] In a possible implementation manner, an expression pattern matching the user somatosensory data can also be displayed above the user position according to the user somatosensory data of the intelligent terminal and the coordinate data of the intelligent terminal located by Mesh networking, so as to achieve the display effect of one image for a thousand people.
[0051] In this embodiment, user personalized display materials are generated according to the user somatosensory data and the coordinate data, enabling the user to obtain a unique mixed reality experience related to their own state and position, and enhancing the personalization degree of the stage mixed reality image.
[0052] Optionally, the method for obtaining the coordinate data of the intelligent terminal located by Mesh networking includes: Calculating the azimuth angle and horizontal distance of the intelligent terminal to the Mesh networking according to the signal strength; Calculating the vertical distance of the intelligent terminal according to the gravity sensor in the intelligent terminal; Determining the three-dimensional coordinate data of the intelligent terminal according to the vertical distance, the azimuth angle of the intelligent terminal to the Mesh networking, and the horizontal distance.
[0053] In a possible implementation manner, the signal strengths of the intelligent terminal to Mesh node A and Mesh node B are respectively and , and the actual distances of the intelligent terminal to Mesh node A and Mesh node B are respectively and , then the azimuth angle and the horizontal distance The calculation formula is as follows: , , where and are the angles of Mesh node A and Mesh node B relative to the reference direction. Then, the gravity sensor of the intelligent terminal measures the tilt angle relative to the horizontal plane as , the vertical distance from the intelligent terminal to the Mesh network is , then the vertical distance of the intelligent terminal has the following calculation formula: , Then, finally, the three-dimensional coordinate data of the intelligent terminal are determined based on the vertical distance, the azimuth angle from the intelligent terminal to the Mesh network, and the horizontal distance, which are respectively: .
[0054] In this embodiment, by simultaneously using the signal strength, the gravity sensor, and the triangulation principle, the position of the intelligent terminal in the three-dimensional space is accurately calculated, providing a basis for the generation of personalized display materials in the subsequent third display space.
[0055] Figure 3 is a schematic structural diagram of image acquisition and display according to an embodiment of the present application, as shown in Figure 3 : Among them, the first display space is located in the main display area of the stage, presenting the overall performance of the stage. The second display space is located on both sides around the first display space, so that the audience can watch the dynamic digital human interaction images from different angles. The third display space is located above the stage audience to display user personalized display materials, enhancing the in-depth interaction between the audience and the stage performance.
[0056] Optionally, the method further includes: Archiving and playback setting for the image data of the stage performance, and the playback setting includes time point selection and stage position selection; The cloud regenerates the mixed reality image according to the selected playback screen and playback instruction, and sends it to the intelligent terminal.
[0057] In a possible implementation manner, if user A selects to playback to 3 minutes and 15 seconds and switches to the camera view at the left side of the stage, the cloud will load the stage data at 3 minutes and 15 seconds and send it to user A's intelligent device. After user A watches the stage image at the left side of the stage at 3 minutes and 15 seconds, user A can select to regenerate the mixed reality image. For example, user A can select to add an expression effect to the performer on the stage.
[0058] Among them, the playback settings of different users are stored separately to ensure user privacy. In addition, for the mixed reality images regenerated by the user, new and old markings are made so that the user can distinguish the original stage performance.
[0059] In this embodiment, through the playback settings, the audience is no longer limited to a single perspective and a fixed timeline. When the stage performance ends, the audience can actively explore all aspects of the stage performance in the smart device and choose to generate more personalized stage mixed reality images according to their preferences, obtaining a stronger sense of presence and immersion.
[0060] According to the embodiments of the present disclosure, the following technical effects are achieved: 1) Through the first display screen, the second display screen, and the third display screen, the stage screen, the virtual digital human screen, and the personalized interaction screen are presented to the audience respectively, enriching the stage content and enhancing the audio-visual experience of the audience.
[0061] 2) Combining the pre-trained YOLO model and the pre-trained Gaussian-SLAM model to achieve more intelligent mixed reality image generation in the cloud, saving a large amount of storage space and providing real-time generation feedback.
[0062] 3) Through Mesh networking, the interconnection and communication of multi-angle stage camera devices, projection devices, the cloud, and smart terminals are realized, improving the transmission efficiency of image data.
[0063] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to the present application.
[0064] The above is the introduction of the method embodiments. The following further illustrates the solution of the present application through device embodiments.
[0065] Figure 4 The block diagram of the mixed reality image device based on Mesh networking according to the embodiment of the present application is shown, as Figure 4 shown including: The first acquisition module 401 is used to acquire multi-angle stage camera data, convert the multi-angle stage camera data into a streaming media, and upload the streaming media to the cloud; The second acquisition module 402 is configured to identify dynamic objects in the streaming media according to the choreography instruction, perform demembrane peeling and three-dimensional reconstruction on the dynamic objects, and acquire dynamic virtual images; The construction module 403 is configured to perform mixed reality synthesis based on the dynamic virtual image and the streaming media to construct a mixed reality image; The sending module 404 sends the mixed reality image to the first display space through Mesh networking by the cloud.
[0066] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the described modules can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0067] Figure 5 The structure diagram of the terminal device or server suitable for implementing the embodiments of the present application is shown.
[0068] As Figure 5 shown, the terminal device or server includes a central processing unit (CPU) 501, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 502 or the program loaded from the storage section 508 into the random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the terminal device or server are also stored. The CPU 501, ROM 502, and RAM 503 are connected to each other through a bus 504. The input / output (I / O) interface 505 is also connected to the bus 504.
[0069] The following components are connected to the I / O interface 505: an input section 506 including a keyboard, a mouse, etc.; an output section 507 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 508 including a hard disk, etc.; and a communication section 509 including a network interface card such as a LAN card, a modem, etc. The communication section 509 performs communication processing via a network such as the Internet. The drive 510 is also connected to the I / O interface 505 as needed. A removable medium 511, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 510 as needed so that the computer program read from it can be installed into the storage section 508 as needed.
[0070] In particular, according to the embodiments of the present application, the above method flow steps can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a machine-readable medium, and the computer program contains program codes for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 509, and / or installed from the removable medium 511. When the computer program is executed by the central processing unit (CPU) 501, the above functions defined in the system of the present application are executed.
[0071] It should be noted that the computer-readable medium shown in the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or combined with an instruction execution system, apparatus, or device. In the present application, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable medium can send, propagate, or transmit a program for use by or combined with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0072] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0073] The units or modules involved in the embodiments described in the present application can be implemented in software or in hardware. The described units or modules can also be provided in a processor. Among them, the names of these units or modules do not, in some cases, constitute a limitation on the units or modules themselves.
[0074] On the other hand, the present application also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiments; or may exist separately and not be assembled into the electronic device. The above computer-readable storage medium stores one or more programs, and when the above-mentioned programs are executed by one or more processors, they implement the methods described in the present application.
[0075] The above description is only a preferred embodiment of the present application and an explanation of the technical principles applied. Those skilled in the art should understand that the scope of the application involved in the present application is not limited to the technical solutions formed by the specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above application concept. For example, the technical solutions formed by mutually replacing the above features with other technical features having similar functions (but not limited to) described in the present application.
Claims
1. A mixed reality imaging method based on Mesh networking, characterized in that: include: Acquire multi-angle stage camera data, convert the multi-angle stage camera data into streaming media, and upload the streaming media to the cloud; According to the arrangement instruction, the dynamic object in the streaming media is identified, and the dynamic object is stripped and three-dimensionally reconstructed to obtain a dynamic virtual image; Performing mixed reality synthesis based on the dynamic virtual image and the streaming media to construct a mixed reality image; The cloud sends the mixed reality image to the first display space through Mesh networking.
2. The mixed reality imaging method based on Mesh networking according to claim 1, characterized in that: The step of identifying the dynamic object in the streaming media according to the arrangement instruction, performing film stripping and three-dimensional reconstruction on the dynamic object, and obtaining a dynamic virtual image includes: By using a pre-trained YOLO model, object detection and extraction are performed on the image in the streaming media according to the arrangement instructions to obtain the dynamic object; By pre-training the Gaussian-SLAM model, the dynamic object is stripped and three-dimensionally reconstructed to obtain a dynamic virtual image.
3. The mixed reality imaging method based on Mesh networking according to claim 1, characterized in that: The method further comprises: Acquiring stage projection instructions and audience emotion data, and generating dynamic digital human interactive images according to the stage projection instructions and audience emotion data; The cloud sends the dynamic digital human interactive image to the second display space through Mesh networking.
4. The mixed reality imaging method based on Mesh networking according to claim 1, characterized in that: The method further comprises: Acquire user somatosensory data of a smart terminal and coordinate data of the smart terminal located by the Mesh network, and generate user personalized display materials according to the user somatosensory data and the coordinate data; The cloud sends the user personalized display material to the third display space through Mesh networking.
5. The mixed reality imaging method based on Mesh networking according to claim 4 is characterized in that: The method for acquiring the coordinate data of the smart terminal positioned by the Mesh network includes: Calculate the azimuth and horizontal distance from the smart terminal to the Mesh network according to the signal strength; Calculating a vertical distance of the smart terminal according to a gravity sensor in the smart terminal; The three-dimensional coordinate data of the smart terminal is determined according to the vertical distance, the azimuth angle and the horizontal distance from the smart terminal to the Mesh network.
6. The mixed reality imaging method based on Mesh networking according to claim 1, characterized in that: The method further comprises: Obtaining a smart terminal projection instruction, segmenting the projection picture in the smart terminal projection instruction, and obtaining smart terminal pixel blocks; The smart terminal pixel block is sent to the smart terminal according to the smart terminal projection instruction and the coordinate position of the smart terminal, and the smart terminal pixel block is displayed on the smart terminal.
7. The mixed reality imaging method based on Mesh networking according to claim 4, characterized in that: The method further comprises: Archiving and replaying the video data of the stage performance, wherein the replaying settings include time point selection and stage position selection; The cloud regenerates the mixed reality image according to the selected playback screen and playback instruction, and sends it to the smart terminal.
8. A mixed reality imaging device based on Mesh networking, characterized in that: include: A first acquisition module is used to acquire multi-angle stage camera data, convert the multi-angle stage camera data into streaming media, and upload the streaming media to the cloud; A second acquisition module is used to identify dynamic objects in the streaming media according to the arrangement instructions, perform film stripping and three-dimensional reconstruction on the dynamic objects, and obtain dynamic virtual images; A construction module, used for performing mixed reality synthesis based on the dynamic virtual image and the streaming media to construct a mixed reality image; The sending module sends the mixed reality image to the first display space through the Mesh network by the cloud.
9. An electronic device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: When the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.