Omnidirectional video display method

By synchronously controlling the directional sub-eye to acquire image frames and perform physical stitching and rendering, the problem of space-time consistency and real-time nature of omnidirectional video display in the prior art is solved, real-time dynamic display and direction perception of omnidirectional video are realized, and user experience is improved.

CN120529048APending Publication Date: 2025-08-22WUHAN UNIV OF TECH CHONGQING RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510610290.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-08-22

AI Technical Summary

Technical Problem

The prior art cannot achieve real-time, accurate and direction-aware omnidirectional video display, resulting in poor user experience. Especially in the field of video surveillance, the difference in image coordinates and time differences caused by fixed installation positions of multiple cameras lack time and space consistency.

Method used

By generating control instructions, the directional sub-eye synchronously collects image frames, and physically splicing and rendering of image frames based on posture information and spatial positioning labels, ensuring spatial and temporal consistency and real-time dynamics in the same time and space. Multi-threading and OpenGL modules are used to realize image rendering.

Benefits of technology

Real-time dynamic display of omnidirectional video is realized, and users can perceive direction on multiple display screens, improving the user experience and time-space consistency of omnidirectional video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120529048A_ABST
    Figure CN120529048A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an omni-directional video display method, and relates to real-time dynamic three-dimensional panorama construction, and the method comprises the steps: responding to a first operation, generating a control instruction, and transmitting the control instruction to an image collection end, so as to enable the image collection end to call a directional sub-eye according to the control instruction, acquiring an image frame acquired by each directional sub-eye, a spatial positioning label and a moment label according to the same moment and the same frame rate; obtaining attitude information of the display screen, and respectively determining the attitude offset of the spatial positioning label of each directional sub-eye at the moment corresponding to the moment label relative to the attitude information; and rendering all-directional video frames obtained by physically splicing the image frames of the labels at the same moment to the corresponding display screens based on the attitude offset. According to the method and the device, the same-moment and same-frame-rate acquisition of the omnidirectional video can be realized, the space-time consistency and the real-time dynamism of the omnidirectional video can be realized, a real-time dynamic scene with accurate space-time measurement can be provided for a user, and the quantitative perception of the azimuth and the direction of the scene is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to an omnidirectional video display method. Background Art

[0002] Currently, when users capture images through related devices, they are limited by the shooting range of a single lens and cannot capture more targets in the image. In general daily applications, the shooting range can be expanded by using a wide-angle lens, but the expanded range is also relatively limited. Alternatively, panoramic images can be captured by panning the lens, but the quality of the image captured in this way depends on the user's stability when panning the lens. If the lens has large up and down displacement fluctuations, the captured image will have defects such as misalignment and distortion, and can only capture images on one plane. The shooting range is still limited, and it is impossible to continuously output omnidirectional videos consisting of multiple video frames.

[0003] In the field of video surveillance, although it is possible to set up lenses at multiple points in a large scene, that is, use multiple cameras to capture images at different azimuths, and then complete panoramic monitoring by viewing the multiple images captured by the multiple cameras at the same time, the installation positions of the multiple cameras are fixed, and the common monitoring image areas between adjacent cameras are often not effectively (incomplete or unclear) displayed, and the real-time performance of the captured images is poor. In addition, the position parameters of multiple lenses in different positions are different, which will lead to differences in the image coordinates of each video channel, and the shooting time may also be different, lacking "time and space consistency", which in turn makes users perceive directions differently when watching omnidirectional videos, affecting the viewing effect, and the rendered scene lacks "sense of direction".

[0004] Therefore, whether in daily user use or in the field of video surveillance, there is a lack of a real-time, accurate, and directionally aware omnidirectional video display method. Users have a poor experience when obtaining and viewing omnidirectional videos, and cannot dynamically render omnidirectional videos that are consistent with the actual scene in time and space. Summary of the Invention

[0005] The embodiments of the present application provide an omnidirectional video display method to address the deficiencies in the above-mentioned related technologies. The omnidirectional video rendering obtained by the present application can meet the requirements of spatiotemporal consistency for the same scene in the same time and space, and can be rendered and changed synchronously with each video frame according to the scene changes, thereby meeting the real-time dynamics from shooting to rendering and display. The technical solution is as follows: In a first aspect, an embodiment of the present application provides an omnidirectional video display method, which is applied to a user terminal and includes: In response to a first operation, generating a control instruction; the first operation is an interactive operation for acquiring omnidirectional video; Sending the control instruction to the image acquisition end, so that the image acquisition end determines the directional sub-eye to be called according to the control instruction, so that all the directional sub-eyes synchronously capture image frames at the same time based on the same acquisition frame rate; Obtaining the image frames collected by each directional sub-eye uploaded by the image acquisition end and the spatial positioning tags and time tags corresponding to the image frames; wherein the image frames collected by each directional sub-eye uploaded at the same time have the same time tag, temporal resolution, and spatial resolution; Acquire the posture information of the display screen at the time corresponding to the moment label, and respectively determine the posture offset of the spatial positioning tag of each directional sub-eye at the time corresponding to the moment label relative to the posture information; Determine the rendering area of ​​each image frame on the display screen according to the posture offset, and physically splice each image frame with the same moment label to obtain an omnidirectional video frame at the moment corresponding to the moment label; The omnidirectional video frame is rendered onto the display screen based on the correspondence between each image frame and the rendering area, so that the display screen displays the omnidirectional video dynamically and in real time according to the time sequence corresponding to the omnidirectional video frame.

[0006] In an optional solution of the first aspect, the method further includes: In response to a second operation, updating the posture information of the display screen in real time; the second operation is an interactive operation that causes the posture information of the display screen to change; Based on the updated posture information, obtain a time tag after the posture information is updated, obtain a spatial positioning tag of each directional sub-eye at the time corresponding to the updated time tag, and update the posture offset of each directional sub-eye relative to the updated posture information based on the spatial positioning tag; determining a rendering area of ​​each image frame on the display screen based on the updated posture offset; Physically splicing each image frame with the same time tag to obtain an omnidirectional video frame at the time corresponding to the time tag, and rendering the omnidirectional video frame onto the display screen based on the correspondence between each image frame and the rendering area, so that the display screen dynamically and in real time displays the omnidirectional video according to the time sequence corresponding to the omnidirectional video frames. In an optional solution of the first aspect, physically splicing each image frame with the same time tag to obtain the omnidirectional video frame at the time corresponding to the time tag specifically includes: Acquire the range of the field of view angle of each two adjacent directional sub-eyes in real time, and determine the field of view intersection point of each two adjacent directional sub-eyes according to the range of the field of view angle of each two adjacent directional sub-eyes; Processing each image frame with the same label at the same time based on the field of view intersection, deleting the overlapping portion of the image frame of one directional sub-eye with the image frame of another directional sub-eye; Physically splicing and deleting the image frames after the overlapping parts to obtain an omnidirectional video frame at the time corresponding to the time label. In an optional solution of the first aspect, determining the rendering area of ​​each image frame on the display screen based on the posture offset further includes: Acquiring a display resolution of the display screen, and determining a display range of the corresponding display screen based on the display resolution; The rendering position of each image frame in the display range is determined in combination with the posture offset; and the rendering area of ​​each image frame in the display range is determined according to the resolution and rendering position of each image frame.

[0007] In an optional solution of the first aspect, the control instruction is further used to obtain omnidirectional video of a historical period; The method further comprises: Sending the control instruction to the image acquisition end so that the image acquisition end traverses all image frames corresponding to the historical period and the corresponding spatial positioning tags and time tags according to the stored identifier of each image frame; Arrange all video frames in chronological order, obtain the posture offset of the spatial positioning label of each image frame with the same moment label relative to the posture information for each historical moment, and physically stitch each image frame based on the posture offset to obtain the omnidirectional video frame corresponding to the historical moment; The omnidirectional video frame of each of the historical moments is sequentially rendered onto the corresponding display screen according to the time sequence.

[0008] In an optional solution of the first aspect, when the number of display screens is the same as the number of called directional sub-eyes, the method further includes: A corresponding relationship between a directional sub-eye and a display screen is determined respectively, and an image frame captured by each of the called directional sub-eyes is rendered respectively onto the display screen having the corresponding relationship.

[0009] On the other hand, the embodiment of the present application also provides an omnidirectional video display method The method is applied to an image acquisition end, wherein the image acquisition end includes a plurality of directional sub-eyes; The method comprises: Obtain control instructions issued by the user end in real time; Determining the directional sub-eye called by the user terminal according to the control instruction; Allocate a sub-thread to the handle of the directional sub-eye called by the user end, read the image frames captured by the directional sub-eye corresponding to each sub-thread through the main thread, and obtain the spatial positioning label and time label of each called directional sub-eye in real time; The image frame and the corresponding spatial positioning tag and time tag are sent to the user terminal in real time, so that the user terminal obtains an omnidirectional video based on the image frame and the corresponding spatial positioning tag and time tag fusion and dynamically renders the omnidirectional video on the display screen in real time.

[0010] In an optional solution of the second aspect, before sending the image frame and the corresponding spatial positioning tag and time tag to the user terminal, the method further includes: The image frames and the corresponding spatial positioning tags and time tags are stored in a cache area, and an identifier is assigned to the directional sub-eye corresponding to each image frame. A mapping relationship between the directional sub-eye and each image frame and the corresponding spatial positioning tag and time tag is determined through each identifier.

[0011] In a third aspect, an embodiment of the present application further provides an electronic device comprising a memory, a processor, and a computer program stored on the memory and runnable on the processor, wherein when the processor executes the program, the method provided by the first aspect or any one of the implementations of the first aspect or the second aspect or any one of the implementations of the second aspect of the present application is implemented.

[0012] In a fourth aspect, the present application also provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, it implements the method provided by the first aspect or any one of the implementations of the first aspect or the second aspect or any one of the implementations of the second aspect of the embodiment of the present application.

[0013] The beneficial effects of the technical solutions provided by some embodiments of the present application include at least: The embodiment of the present application provides an omnidirectional video display method, which allows users to capture omnidirectional videos. By using the corresponding azimuth image rendering method, it can save the time consumption caused by omnidirectional image stitching and image transformation, and provide users with real-time dynamic direction perception. The image rendered on the display screen is related to the direction, so that users can perceive the direction in real time in application scenarios such as mobile phones, VR devices, and multi-displays configured on PCs, better watch omnidirectional videos, improve the user experience, and provide a real-time, accurate, and direction-aware omnidirectional video display method. The omnidirectional video rendering obtained by the present application corresponds to the same scene in the same time and space, which can meet the requirements of time and space consistency, and is rendered and changed synchronously according to each video frame of the scene change, so as to meet the real-time dynamics from shooting to rendering display, thereby constructing an omnidirectional video with "real-time dynamics" and the user can get "direction perception". BRIEF DESCRIPTION OF THE DRAWINGS

[0014] In order to more clearly illustrate the technical solutions in this application or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0015] Figure 1 This is a flow chart of an omnidirectional video display method provided in an embodiment of the present application; Figure 2 This is a flow chart of an omnidirectional video display method provided in an embodiment of the present application; Figure 3 This is a structural diagram of an omnidirectional video display device provided in an embodiment of the present application; Figure 4 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0016] To make the objectives, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making any creative efforts are within the scope of protection of this application.

[0017] The terms "including" and "having," and any variations thereof, in the specification and claims of this application and the accompanying drawings are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or modules is not limited to the listed steps or modules, but may optionally include steps or modules not listed, or may optionally include other steps or modules inherent to the process, method, product, or apparatus.

[0018] It should be noted that the terms "first" and "second" used in this application are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It is understood that the terms "first" and "second" may interchangeably represent a specific order or precedence, where permitted. It should be understood that the objects distinguished by "first" and "second" may interchangeably represent a specific order or precedence, where appropriate, such that the embodiments of the present application described herein can be implemented in an order other than that described or illustrated herein.

[0019] In some embodiments, an omnidirectional video display method provided by the present application can be implemented on a user end and an image acquisition end, wherein the user end can be a terminal that can interact with a user and output control instructions for controlling the above-mentioned image acquisition end, such as a mobile terminal, a computer and other devices. These terminal devices can be connected to one or more display screens. The display screen can be a built-in screen of a device such as a mobile phone screen and / or an external display screen of the device. The display screen can be a curved screen, a folding screen, etc., and the embodiment of the present application is not limited to this; the image acquisition end can be a device for shooting panoramic images. The image acquisition end in the embodiment of the present application is composed of multiple directional sub-eyes, and each directional sub-eye has a different direction, and shoots different areas respectively.

[0020] In some preferred embodiments, six directional sub-eyes may be selected to form an image acquisition end, which is roughly a cube structure, and a directional sub-eye with an optical axis perpendicular to the normal and facing outward is provided at the center point of each surface.

[0021] It can be understood that the user can input operations on the user end, such as controlling the image acquisition end to turn on and take images, controlling the image acquisition end to turn off, selecting several lenses on the image acquisition end to take images, obtaining images taken by the image acquisition end for a certain period of time in the past, etc.; the image acquisition end can then receive the control instructions from the user end, thereby executing the shooting behavior determined by the control instructions.

[0022] Among the related technologies, most of them use the "refractive imaging" method to achieve omnidirectional video acquisition, which increases the complexity of the acquisition system; a few related technologies use a camera installed on a pan-tilt head to achieve 360° rotation in the horizontal direction. However, this method cannot acquire images in all directions. Although it can display image data in a certain direction at a certain moment, it lacks "real-time dynamics". The orientation information corresponding to the actual acquired image cannot be reflected in real time on the display terminal; some existing technologies display the panorama through multi-screen display, but in this way, the user cannot perceive the spatial positioning label corresponding to each video frame, and can only display the image separately through each display screen, and cannot achieve temporal and spatial consistency. The user cannot view the panoramic image from the perspective of the user, resulting in poor display effect.

[0023] The present application is described in detail below with reference to specific embodiments.

[0024] Next, combine Figure 1 , taking the user end executing an omnidirectional video display method as an example, an omnidirectional video display method provided by the embodiment of the present application is introduced. Figure 1 , Figure 1 This is a flow chart of an omnidirectional video display method provided in an embodiment of the present application. Figure 1 As shown, the method includes the following steps: S101: Generate a control instruction in response to a first operation; the first operation is an interactive operation for acquiring omnidirectional video.

[0025] S102: Send the control instruction to the image acquisition end, so that the image acquisition end determines the called directional sub-eye according to the control instruction, so that all the directional sub-eyes synchronously capture image frames at the same time based on the same capture frame rate.

[0026] S103, obtaining the image frames captured by each directional sub-eye and the spatial positioning tags and time tags corresponding to the image frames uploaded by the image acquisition terminal; wherein the image frames captured by each directional sub-eye uploaded at the same time have the same time tag, temporal resolution, and spatial resolution. S104: Acquire the posture information of the display screen at the time corresponding to the time label, and respectively determine the posture offset of the spatial positioning tag of each directional sub-eye at the time corresponding to the time label relative to the posture information.

[0027] S105 , determining a rendering area of ​​each image frame on the display screen according to the posture offset, and physically splicing each image frame with the same moment label to obtain an omnidirectional video frame at the moment corresponding to the moment label.

[0028] S106 , rendering the omnidirectional video frame onto the display screen based on the correspondence between each image frame and the rendering area, so that the display screen dynamically and in real time displays the omnidirectional video according to the time sequence corresponding to the omnidirectional video frame.

[0029] It should be noted that the omnidirectional video rendering obtained by shooting this application corresponds to the same scene in the same space and time and can meet the requirements of spatiotemporal consistency and real-time dynamics, and is rendered and changed synchronously with each video frame that changes the scene, so as to meet the real-time dynamics from shooting to rendering display. Finally, an omnidirectional video that is consistent with the time and space of the actual shooting scene in 720° space (horizontal 360° + vertical 360°) can be obtained, and can be rendered in real time according to each video frame shot, ensuring that the spatial information and time information corresponding to each video frame displayed on the display screen are highly consistent with the actual scene, and can dynamically render each video frame according to the actual acquisition frame rate, thereby realizing omnidirectional video with spatiotemporal consistency and real-time dynamics.

[0030] It can be understood that real-time dynamics can be understood as the acquisition frame rate of multiple directional sub-eyes used to capture image frames being the same and having the same spatiotemporal resolution, so that within the acquisition sequence, the multiple directional sub-eyes can capture any moving entity in the physical space where each directional sub-eye works, and the image frames captured according to the acquisition frame rate and presentation frame rate can characterize the spatial position, motion posture and motion trajectory of the moving entity in a timeable, measurable, calibrable, alignable, correlated and traceable manner.

[0031] The minimum acquisition frame rate is the number of image frames that can be captured per second by the directional sub-eye. It can be obtained by: the acquisition operation period of each directional sub-eye within the acquisition sequence (from the time when the acquisition instruction is issued to the time when the acquisition data is output) is less than the length of the acquisition frame peak period (the period between two consecutive frame peak moments).

[0032] Among them, the minimum presentation frame rate is the number of bitmap images that appear continuously on the display per second. It can be selected according to the hardware level and actual situation to avoid the influence of ghost error and ghost interaction delay error. The ghost error is the drift error between the position, posture and motion trajectory of any displayed physical object and the actual position, posture and motion trajectory caused by the imbalance between the acquisition frame rate and the presentation frame rate; the ghost interaction delay error refers to when the user interacts with the virtual object, due to the influence of factors such as lighting, projection and tracking, there is a certain visual difference and delay between the virtual object and the real scene.

[0033] Specifically, in S101, the first operation is an interactive operation input by the user on the user end, which is specifically used to obtain omnidirectional video. For example, if the user end is a mobile phone, the first operation can be: the user directly selects a lens on the touch screen, enters the shooting time, etc. For example, if the user end is a computer, the first operation can be: the user inputs information through an interactive device such as a mouse and keyboard. The embodiment of the present application does not limit this.

[0034] Specifically, in S101, the control instruction can be understood as an interactive message generated by the user end based on the interactive operation for controlling the image acquisition end, so that the image acquisition end can realize video shooting according to the content of the control instruction. For example, the user selects all the lenses on the image acquisition end on the interactive interface of the user end to shoot a panoramic image between time A and time B, then the corresponding control instruction is generated and S102 is executed to send the control instruction to the image acquisition end, so that the image acquisition end determines the directional sub-eye to be called according to the control instruction.

[0035] Specifically, in S102 , the user may select at least one directional sub-eye, so that all the directional sub-eyes synchronously capture image frames at the same time based on the same capture frame rate.

[0036] Specifically, in S103, after the image acquisition end determines the directional sub-eye selected by the user according to the control instruction, it obtains image frames through the directional sub-eye and records the image frames acquired by each directional sub-eye and the spatial positioning tag and time tag corresponding to the image frame; wherein, the image frames acquired by each directional sub-eye uploaded at the same time have the same time tag, temporal resolution, and spatial resolution.

[0037] It can be understood that the spatial positioning tag corresponding to the image frame is the direction information of the directional sub-eye that collects the image frame. The image acquisition end can determine the optical axis of each directional sub-eye through the built-in gyroscope, thereby determining the spatial positioning tag of the corresponding image frame, which can specifically be the pitch angle, heading angle and roll angle of the optical axis of the directional sub-eye relative to the reference coordinate system.

[0038] Specifically, in S104, the attitude information of the display screen can also be determined. The central axis of the display screen can be selected as the normal direction, and the pitch angle, heading angle and roll angle of the normal direction relative to the reference coordinate system can be calculated to obtain the attitude information (various attitude angles) of the display screen at the corresponding moment.

[0039] For example, taking a mobile phone as a display screen, the gyroscope on board the mobile phone can be used to determine the pitch angle, heading angle and roll angle of the mobile phone screen relative to the reference coordinate system.

[0040] In some embodiments, the posture information of each display screen can also be set through the terminal, for example, determining that the normal of a certain display screen is the same as the X-axis in the reference coordinate system, so that real-time dynamics can be achieved by synchronously changing the rendering of the video frame according to the scene change. This embodiment of the present application is not limited to this.

[0041] It is understandable that the user can select a display screen for displaying the image frame through the above first operation, which may include the display screen of the terminal or an external display screen connected to the terminal device. This embodiment of the present application does not limit this.

[0042] Specifically, the attitude offset is obtained by comparing the attitude angle corresponding to the attitude information of the display screen with the spatial positioning tag (including the attitude angle) of the directional sub-eye, so that the direction difference between the display screen and the image frame can be determined based on the attitude offset.

[0043] Then, S105 is executed to determine the rendering area of ​​each image frame on the display screen according to the posture offset, and each image frame with the same moment label is physically spliced ​​to obtain an omnidirectional video frame at the moment corresponding to the moment label.

[0044] Specifically, the display range of the corresponding display screen can be determined according to the display resolution of the display screen; The rendering position of each image frame in the display range is determined in combination with the posture offset; and the rendering area of ​​each image frame in the display range of the display screen can be determined according to the resolution and rendering position of each image frame.

[0045] Thereby, the relative position relationship of each image frame on the display screen is determined. There may be overlapping areas between different image frames, and the image frames can be spliced ​​into omnidirectional image frames by physical splicing.

[0046] It can be understood that the image frames used for physical splicing in S105 are image frames with the same time tag.

[0047] The video frame size displayed on the display screen is determined according to the value of the posture offset, that is, the rendering area is determined, and the omnidirectional video frame obtained by physical splicing is rendered on the corresponding display screen. Further, S106 is executed to render the omnidirectional video frame onto the display screen based on the correspondence between each image frame and the rendering area, so that the display screen dynamically and in real time displays the omnidirectional video according to the time sequence corresponding to the omnidirectional video frame.

[0048] In some embodiments, as each directional sub-eye at the image acquisition end continuously acquires image frames, steps S101 to S106 are sequentially executed according to the transmitted multi-frame information, thereby enabling real-time omnidirectional video display on the user end, thereby achieving spatiotemporal consistency from video frame acquisition to video rendering, and achieving spatiotemporal consistency in 720° space (horizontal 360° + vertical 360°) based on the actual acquired lens, thereby improving the effect of the acquired omnidirectional video.

[0049] In some embodiments, the rendering of image frames can be implemented through modules such as OpenGL, which is not limited in this embodiment of the present application.

[0050] In some application scenarios, taking mobile phones as an example, users can click on the mobile phone's interactive interface to determine the image acquisition device used to shoot omnidirectional video and select directional sub-eyes. For example, directional sub-eyes on the six end faces of a hexahedral image acquisition device are selected to capture a 360° panoramic video. Users can also determine the display screen used to display the panoramic video. For example, if a mobile phone screen is selected, the posture information of the mobile phone screen can be determined, and the posture offset of each directional sub-eye relative to the mobile phone screen can be calculated, so that the displayed image can be determined according to the real-time posture of the mobile phone screen.

[0051] In some embodiments, after each directional sub-eye captures the image frames and the spatial positioning tags and time tags corresponding to the image frames, the user terminal physically splices each image frame with the same time tag to obtain an omnidirectional video frame at the time corresponding to the time tag, including: Acquire the range of the field of view angle of each two adjacent directional sub-eyes in real time, and determine the field of view intersection point of each two adjacent directional sub-eyes according to the range of the field of view angle of each two adjacent directional sub-eyes; Processing each image frame with the same label at the same time based on the field of view intersection, deleting the overlapping portion of the image frame of one directional sub-eye with the image frame of another directional sub-eye; The image frames after the overlapping parts are physically spliced ​​and deleted to obtain the omnidirectional video frame at the moment corresponding to the moment label.

[0052] It is understandable that if there is no overlap between the selected directional sub-eyes, there is no need to stitch them together to obtain the omnidirectional video.

[0053] It should be noted that the method provided in the embodiment of the present application only needs to calculate the field of view intersection of the directional sub-eyes, crop the image according to the field of view intersection and physically splice it, without the need for complex image change operations, and can efficiently and quickly generate omnidirectional video.

[0054] In some embodiments, since the user needs to view images in different directions when using the user terminal, the posture information of the display screen of the user terminal is changed to view omnidirectional videos. The method may further include the following steps: In response to a second operation, updating the posture information of the display screen in real time; the second operation is an interactive operation that causes the posture information of the display screen to change; Based on the updated posture information, obtain a time tag after the posture information is updated, obtain a spatial positioning tag of each directional sub-eye at the time corresponding to the updated time tag, and update the posture offset of each directional sub-eye relative to the updated posture information based on the spatial positioning tag; Determine the rendering area of ​​each image frame on the display screen based on the updated posture offset, and physically splice each image frame with the same time label to obtain the omnidirectional video frame at the time corresponding to the time label; The omnidirectional video frame is rendered onto the display screen based on the correspondence between each image frame and the rendering area, so that the display screen displays the omnidirectional video dynamically and in real time according to the time sequence corresponding to the omnidirectional video frame.

[0055] Specifically, the second operation is an interactive operation input by the user on the user side, which will cause the posture information of the display screen to change. The second operation may include: the user moves and adjusts the display screen to change the posture of the display screen, and may also include: the user changes the direction reference of the display screen through gestures and buttons on the mobile phone screen, and drags the direction reference on the mobile phone screen to change the posture of the display screen relative to the direction reference.

[0056] In some embodiments, the user terminal may also be a wearable VR head display device, and the second operation may also include: the user controls the posture of the display screen by turning the head direction or using a handheld operating device. This embodiment of the present application is not limited to this.

[0057] In some embodiments, the display resolution of the display screen may be obtained, and the display range of the corresponding display screen may be determined based on the display resolution; The rendering position of each image frame in the display range is determined in combination with the posture offset; and the rendering area of ​​each image frame in the display range is determined according to the resolution and rendering position of each image frame.

[0058] It can be understood that for an application scenario with only one display screen, a direction reference can be selected, that is, the reference coordinate system in the above embodiment, and the posture information of the display screen in the reference coordinate system can be determined, that is, the posture angle difference relative to the direction reference. The image frames of multiple directional sub-eyes can respectively determine the posture angle difference relative to the direction reference according to the spatial positioning label, and then determine the posture offset relative to the display screen. The relative position relationship between the image frame and the display screen can be determined based on the posture offset, that is, the rendering position of each image frame in the display range is determined in combination with the posture offset, and then the rendering area of ​​each image frame in the display range is determined according to the resolution and rendering position of each image frame.

[0059] It should be noted that omnidirectional video rendering corresponding to the same scene in the same space and time satisfies spatiotemporal consistency, and video frame rendering changes synchronously with scene changes to meet real-time dynamics.

[0060] It can be understood that for each display screen, based on the spatial positioning tag and the display range, at least one image frame with an intersection in the display range is determined, and the image frame can be cropped, scaled, and other operations can be performed according to the display resolution to generate an omnidirectional video frame corresponding to the display screen resolution, which is then displayed on the display screen after being rendered by the user end.

[0061] For example, for mobile phones, the screen display range is relatively small, so a portion of the image frame can be captured for rendering. Users can then watch omnidirectional video by moving the phone or dragging the screen. For VR headsets, the VR device display can display a larger range, and the rendered image can be determined as the user's visual direction changes. This is equivalent to projecting the image frame captured by each directional sub-eye onto the image plane. As the user's visual direction changes, the VR device display window moves on the image plane, thereby determining the image that needs to be rendered on the VR device display.

[0062] In some embodiments, for application scenarios with multiple display screens, the attitude angle of each display screen relative to a direction reference or a reference coordinate system can be determined separately, and the image frame captured by each directional sub-eye can be projected onto the reference coordinate system. The display range of each display screen can be easily determined based on the display resolution and image zoom ratio of each display screen. Based on the attitude offset of the image frame captured by each directional sub-eye relative to each display screen, the image frame falling into the corresponding display screen can be determined and then rendered to the corresponding display screen respectively.

[0063] In some embodiments, the control instructions in the above embodiments are further used to obtain omnidirectional videos of historical periods. That is, the user can obtain omnidirectional videos of historical periods cached by the image acquisition end through interaction with the user end, specifically including: Sending the control instruction to the image acquisition end so that the image acquisition end traverses all image frames corresponding to the historical period and the corresponding spatial positioning tags and time tags according to the stored identifier of each image frame; Arrange all video frames in chronological order, obtain the posture offset of the spatial positioning label of each image frame with the same moment label relative to the posture information for each historical moment, and physically stitch each image frame based on the posture offset to obtain the omnidirectional video frame corresponding to the historical moment; The omnidirectional video frame of each of the historical moments is sequentially rendered onto the corresponding display screen according to the time sequence.

[0064] Specifically, the user can also select image frames in one or several directions in the past period to display. That is, the image acquisition end stores the image frames and corresponding spatial positioning tags collected by each directional sub-eye respectively, and the user end can select the required image frames and render them separately.

[0065] In some embodiments, when the number of display screens is the same as the number of called directional sub-eyes, the corresponding relationship between the directional sub-eyes and the display screens can be determined respectively, and the image frames captured by each directional sub-eye are rendered respectively on the display screens with the corresponding relationship.

[0066] Next, combine Figure 2 , taking the image acquisition end executing an omnidirectional video display method as an example, an omnidirectional video display method provided by the embodiment of the present application is introduced. Figure 2 , Figure 2 This is a flow chart of an omnidirectional video display method provided in an embodiment of the present application. Figure 2 As shown, the method includes the following steps: S201, obtaining control instructions sent by the user terminal in real time.

[0067] S202: Determine the directional sub-eye called by the user terminal according to the control instruction.

[0068] S203: Allocate a sub-thread to the handle of the directional sub-eye called by the user end, read the image frames captured by the directional sub-eye corresponding to each sub-thread through the main thread, and obtain the spatial positioning tag and time tag of each called directional sub-eye.

[0069] S204, sending the image frame and the corresponding spatial positioning tag and time tag to the user terminal in real time, so that the user terminal obtains an omnidirectional video based on the fusion of the image frame and the corresponding spatial positioning tag and time tag and dynamically renders the omnidirectional video to the corresponding display screen in real time.

[0070] Specifically, in S201-S202, the image acquisition end may obtain the control instruction sent by the user end, parse the control instruction, determine the directional sub-eye to be called by the user end, and capture image frames respectively through each called directional sub-eye.

[0071] In some embodiments, the control instruction can also be used to determine the start time of capturing the image frame, that is, the image acquisition end responds to the control instruction, starts at the corresponding start time, and stops capturing at a given end time; the control instruction can also be used to determine the storage location of the captured image frames and other parameters, such as storage locally and / or in the cloud; the image acquisition end can also respond to the control instruction to adjust the lens parameters of each directional sub-eye, which is not limited in this embodiment of the present application.

[0072] In some embodiments, due to the large number of directional sub-eyes, it is difficult to read the captured image frames in a short time using a single thread, and the image frames captured by multiple directional sub-eyes are easily confused. Therefore, in order to achieve synchronous display of multiple videos, a multi-threading method can be used to reduce the operating pressure caused by the operation of multiple lenses in the same thread. Specifically, in S203, a sub-thread can be allocated to the handle of each directional sub-eye, that is, each directional sub-eye is bound to a separate sub-thread, thereby ensuring that data confusion will not occur during the operation of the image acquisition end. The image frames corresponding to the directional sub-eye corresponding to each sub-thread are read separately by the main thread.

[0073] Specifically, when transmitting image frame data, the main thread identifies the directional sub-eye handle and then passes it to multiple sub-threads and binds the camera handles respectively, and transmits the image data to the main thread for display and transmission operations.

[0074] Specifically, in S203, when the directional sub-eye captures each image frame, it is also necessary to record the spatial positioning tag and time tag of each image frame, that is, the attitude angle of each directional sub-eye relative to the reference coordinate system at the time of shooting is determined by the built-in gyroscope, and S104 is executed to send the image frame, spatial positioning tag and time tag to the user end.

[0075] Specifically, when transmitting the image frame, spatial positioning tag, and time tag to the user end, each frame of the image is continuously sent in the form of a video stream. For example, taking the Daheng Image Mercury II Pro ME2P-U3 series camera as an example, the camera resolution is 5120x5120, and the data rate per second of one camera can be calculated according to the following formula: ; Where Q is the amount of image data collected by one camera per second, W and H are the size of the collected image, i.e., the resolution, f is the number of frames, and d is the pixel depth. The calculation yields: ; According to the above calculations, it can be concluded that the bus bandwidth required to transmit one camera needs to be at least greater than 566.25MB / s. The hardware platform used to process and transmit data can be determined based on this transmission bandwidth. The embodiment of the present application does not impose any restrictions on this. Specifically, the bandwidth of each bus can be guaranteed by directly connecting the CPU with different PCI-E interfaces using a PCI-E to USB3.0 capture card.

[0076] In some embodiments, the image frames, spatial positioning tags, and time tags collected by each directional sub-eye can be compressed. Specifically, a socket can be created in the main thread, and an image data transmission function can be added to the function of reading image data in the main thread. This can realize sending the image after reading the image. Relying on the reliability of the image data transmission function TCP, the image data will not be lost.

[0077] In some embodiments, a cache area can be created in the image data transmission function TCP, and the compressed image frame, spatial positioning tag and time tag can be stored in the cache area. The compression method can be qCompress method, which can control the compression ratio. After the image compression is completed, Base64 encoding can be performed. Finally, the output stream version is controlled and the compressed image data is written to the socket, and S204 is executed to send it to the user end.

[0078] In some embodiments, when the image frames are stored in a cache area or directly sent to the user end, an identifier can be assigned to each image frame, and the identifier can be used to distinguish different directional sub-eyes for capturing the image frames. For example, the identifier is determined to be "camera number-image number" to distinguish different directional sub-eyes and different image numbers, thereby avoiding image transmission data disorder during data transmission and allowing the image data to be sorted according to given image data, so that the image frames are transmitted to the user end in the correct sequence.

[0079] Specifically, the current image frame and other information can also be stored in a database, and related data can be stored in the database in the form of camera handle + event stamp, thereby realizing the function of image information backtracking.

[0080] In some embodiments, after the image frame, spatial positioning tag and time tag are sent to the user terminal in S204, the user terminal can implement steps S104-S105 to render the omnidirectional video frame to the corresponding display screen. Please refer to S104-S105 for details and will not be repeated here.

[0081] In some embodiments, a 6-axis attitude sensor can be combined to quickly obtain Euler angles (P, R, Y angles), and the Euler angle data can be passed to the Camera class in the OpenGL graphics library. The Camera class member functions can be directly used to implement the user's orientation perception when watching omnidirectional videos. First, the Euler angles are initialized and all values ​​are set to 0 to determine the Euler angle baseline. Then, the Euler angles of the corresponding display screen on the user end are compared with the Euler angles of each lens. The difference in viewing angle between the display screen viewed by the user and each lens can be determined, so that the user can dynamically perceive the direction when watching omnidirectional videos, and the image rendered on the display screen changes with the actual viewing direction.

[0082] In some embodiments, for the image acquisition end, taking a 720° omnidirectional acquisition device as an example, that is, an image acquisition end in the shape of a six-sided cube, a directional sub-eye is set on each end face. The acquired image can be seamlessly restored to a spatial six-sided omnidirectional video after processing. However, due to the errors caused by the geometric structure limitations of the directional sub-eyes, it is difficult to directly obtain a spatial six-sided omnidirectional video from the image frames captured by all the directional sub-eyes, resulting in a certain range of acquisition blind spots in the image acquisition end. For details, see Figure 3 Schematic diagram of the lens acquisition range and blind area of ​​view.

[0083] For example, Figure 3 As shown, the blue lines represent the boundaries of the directional sub-eye's field of view, which can be used to determine the maximum field of view angle of each directional sub-eye. In this arrangement, the geometric optical properties and overall field of view distribution are symmetrical about the principal optical axis and the cutting reference line, indicating a high degree of symmetry in the model. This model has blind spots within a small range, but two adjacent lenses intersect at a certain distance. This means that image rendering based on this intersection prevents image overlap and enables single-point, blind-spot capture. The field of view boundary line is determined based on the maximum field of view angle of each directional sub-eye, thereby determining the field of view intersection point of the field of view boundary lines of two adjacent directional sub-eyes. This field of view intersection point is used as the image cutting point to cut the overlapping image portions, which are then physically spliced ​​together to produce omnidirectional video, ensuring spatiotemporal consistency between the captured image frames, the real-time rendered image frames, and the actual scene.

[0084] Specifically, the embodiments of the present application provide users with an immersive, directional, and omnidirectional video display method through interaction between a user terminal and an image acquisition terminal, thereby achieving spatiotemporal consistency and dynamic real-time omnidirectional video display, specifically including: The user performs a first operation on the user terminal to determine the directional sub-eye to be called, the start time and end time of shooting, the parameters of the directional sub-eye, and the display screen, etc.; According to the rules of spatiotemporal consistency and real-time dynamics, the user terminal generates a control instruction containing information related to the above operation in response to the first operation, and sends it to the image acquisition terminal.

[0085] After receiving the corresponding control instruction, the image acquisition end parses the content of the control instruction, opens the corresponding called directional sub-eye at the given shooting start time, reads the handle of the directional sub-eye, and configures the sub-threads respectively. It reads the image frames captured by the directional sub-eye corresponding to each sub-thread at the same time and frame rate, and records the image frames captured by each directional sub-eye, as well as the spatial positioning tag and time tag corresponding to the image frame.

[0086] The image acquisition end can read the image frame data through the main thread, store the data in a given cache area, and configure an identifier for each image frame to distinguish each lens and the corresponding image frame, and transmit the image frame and the corresponding spatial positioning tag to the user end in sequence according to the time sequence of the moment tag; the image frame and the spatial positioning tag can also be directly sent to the user end in the time sequence of the moment tag; the image frame and the corresponding spatial positioning tag and moment tag can be compressed before transmission and storage.

[0087] The user end can receive image frame data frame by frame, and identify the corresponding directional sub-eye information through the identifier. For each image frame with the same time tag, the attitude angle difference with the display screen is determined based on the spatial positioning tag, and then the attitude offset is determined.

[0088] The user end can process the image frame according to the posture offset to obtain the omnidirectional video frame at the corresponding moment, and then determine the display range according to the display resolution, determine the area that needs to be rendered in the omnidirectional video frame, and then render the omnidirectional video frame on the display screen. In this way, the rendering can be realized in real time and dynamically according to the orientation information corresponding to the video frame, thereby rendering the actual scene into a 720° space (horizontal 360° + vertical 360°), thereby achieving the spatiotemporal consistency of the omnidirectional video and the dynamic real-time performance of the captured image frames and rendered video frames.

[0089] During use, the user terminal responds to a second operation such as user movement, shaking, gesture operation, etc. that causes the display screen posture information to change, that is, the original posture angle of the display screen changes relative to the direction reference or reference coordinate system. At this time, it is necessary to update the real-time posture information of the display screen, and then update the omnidirectional video frame that needs to be rendered.

[0090] In this way, users can capture omnidirectional videos and use the corresponding azimuth image rendering method, which can not only save the time consumption caused by omnidirectional image stitching and image transformation, but also provide users with real-time dynamic direction perception. The rendered image and direction on the display are related, allowing users to perceive the direction in real time in application scenarios such as mobile phones, VR devices, and PC configurations with multiple monitors, better watch omnidirectional videos, and improve user experience.

[0091] In a specific embodiment, the present application also provides an omnidirectional video display system, including an acquisition system control module, a data acquisition module, a data processing module, a data network transmission module, and an OpenGL-based omnidirectional video stream display module, wherein the acquisition system control module controls the data acquisition module and the data processing module to perform image acquisition and grayscale image conversion, and the data network transmission module and the OpenGL-based omnidirectional video stream display module are connected to realize remote image display.

[0092] The omnidirectional video rendering obtained through this setting can meet the spatiotemporal consistency of the same scene in the same time and space, and realize real-time dynamics by synchronously changing the video frame rendering according to the scene changes.

[0093] The acquisition system control module primarily focuses on front-end camera control. The software required for development uses QT, which handles user interface design, thread creation, data transmission, six-way video display, image conversion, image compression, and camera parameter control. The module was developed using the SDK (Software Development Kit) provided by the camera developer. This enabled real-time display of six channels of video and built the system control software, providing a software platform for adjusting camera parameters. To enable synchronous display of six channels of video, a multithreaded approach was adopted to reduce the operational burden of running the cameras on the same thread. The camera handles were bound to the threads to ensure data integrity. A thread lock (for image data) was added to each thread to prevent blocking or null pointer errors when the main thread reads data. After the main thread identifies the camera handle, it passes it to the multithreaded thread, binds the camera handle, performs image conversion, and transfers the image data to the main thread for display and transmission. This illustrates the overall operational flow of the server software.

[0094] The data acquisition module is designed primarily for camera structure. Addressing the limitations of existing 360° panoramic images and the characteristics of multi-camera systems, and addressing the need for omnidirectional video surveillance, six cameras working simultaneously can capture and present scenes in six directions in real time. This model achieves a physical model for single-point, real-time, and seamless acquisition through adaptive optical parameter adjustment and image segmentation based on camera geometry. The six-camera video is based on real-time dynamic images of the six-camera panoramic view, replacing static images with video streams. This video is continuously streamed and presented, enabling omnidirectional video surveillance. The video acquisition device (image acquisition device) uses Daheng Image's Mercury II Pro ME2P-U3 series camera, which has a high-resolution 5120x5120 resolution. The bus bandwidth required to transmit one camera can be calculated, allowing a traditional PC to be used as the platform for image acquisition. For example, a core processor such as the Intel i9-10980XE can be paired with an ASUS Pro WS X299 SAGE II motherboard, allowing for multiple PCI-E Gen3×16 ports directly connected to the CPU to handle simultaneous image processing of multiple signals. To handle multiple channels of simultaneous video capture and the large amount of video data being transmitted, a split bus, a PCI-E to USB 3.0 capture card, and direct CPU connections between different PCI-E ports can be used to ensure maximum bus bandwidth for each channel.

[0095] The data network transmission module should consider the following aspects. First, it needs to address the lack of wired networks in some operating locations. Second, given the need to transmit large amounts of image data, network bandwidth is limited. Furthermore, the module requires long-term operation, server compatibility, and cost. Therefore, a 4G development module equipped with the CAT1 module EC600NCNLADongle can be selected, with an uplink speed of 40MB / s and a downlink speed of 140MB / s. Given that the bandwidth calculated above may not be sufficient for direct transmission of the original complete image, the data processing module can be used to compress the image. By creating a socket in the main thread of the acquisition system control module and adding an image data transmission function to the main thread's image data reading function, the image can be sent as soon as it is read. Thanks to the reliability of TCP, image data will not be lost.

[0096] The TCP image data transmission function pre-creates a buffer area, compresses and stores the image data in this buffer, and then compresses it using qCompress (with a controllable compression ratio). Once the image is compressed, it is then Base64-encoded. Finally, the output stream version is controlled, and the compressed image data is written to the socket and sent to the client. Six-way image transmission can cause data disorganization, but the software requires the image data to be sorted according to a predetermined sequence for subsequent use. Therefore, an identifier (a numeric symbol) can be added when writing the image data to identify the source of the image data.

[0097] Specifically, the OpenGL-based omnidirectional video stream display module creates an OpenGL class in QT and promotes the OpenGL widget control to the created class, implementing the required functionality within the class. TCP image reception and OpenGL image rendering are connected via signals and slots. After OpenGL initialization, the received image data is sequentially rendered to achieve a two-dimensional presentation of omnidirectional video, providing an artificial sense of direction and achieving omnidirectional spatiotemporal consistency and real-time dynamics.

[0098] This application solves the problem that the time loss caused by panoramic image stitching cannot meet the "real-time dynamic" problem; the traditional method of displaying panoramic views through multiple screens causes users to lack orientation perception. By updating the Euler angle data of the 6-axis attitude sensor, the image information of the corresponding orientation on the user's display device can be redrawn, greatly satisfying the orientation perception.

[0099] The following are device embodiments of the present application, which can be used to implement the method embodiments of the present application. For details not disclosed in the device embodiments of the present application, please refer to the method embodiments of the present application.

[0100] See next Figure 3 , a schematic diagram of the structure of an omnidirectional video display device provided in an exemplary embodiment of the present application. This device can be implemented as all or part of a terminal through software, hardware, or a combination of both, or can be integrated into a server as a standalone module. The omnidirectional video display device 30 in this embodiment of the present application includes a user terminal 301 and an image acquisition terminal 302.

[0101] According to the rules of spatiotemporal consistency and real-time dynamics, the user terminal 301 is used to generate a control instruction in response to a first operation; the first operation is an interactive operation input by the user to the user terminal for obtaining an omnidirectional video; Sending the control instruction to the image acquisition end, so that the image acquisition end determines the directional sub-eye to be called according to the control instruction, so that all the directional sub-eyes synchronously capture image frames at the same time based on the same acquisition frame rate; Obtaining the image frames collected by each directional sub-eye uploaded by the image acquisition end and the spatial positioning tags and time tags corresponding to the image frames; wherein the image frames collected by each directional sub-eye uploaded at the same time have the same time tag, temporal resolution, and spatial resolution; Acquiring posture information of a display screen for displaying an image at the same moment and frame rate corresponding to the moment label, and determining a posture offset of the directional information space positioning tag of each directional sub-eye at the moment corresponding to the moment label relative to the posture information; determining a rendering area of ​​each image frame on the display screen based on the posture offset, and physically splicing each image frame with the same moment label to obtain an omnidirectional video frame at the moment corresponding to the moment label; Based on the correspondence between each image frame and the rendering area, the omnidirectional video frame is rendered onto the display screen, so that the display screen dynamically and in real time displays the omnidirectional video according to the time sequence corresponding to the omnidirectional video frame. Among them, the image acquisition terminal 302 is used to obtain the control instruction issued by the user terminal; Determining the directional sub-eye called by the user terminal according to the control instruction; Allocate a sub-thread to the handle of the directional sub-eye called by the user end, read the image frames captured by the directional sub-eye corresponding to each sub-thread through the main thread, and obtain the spatial positioning label and time label of each called directional sub-eye; The image frame and the corresponding spatial positioning tag and time tag are sent to the user terminal, so that the user terminal obtains an omnidirectional video based on the image frame and the corresponding spatial positioning tag and renders the omnidirectional video on the corresponding display screen.

[0102] It should be noted that the device 30 provided in the above embodiment, when performing the omnidirectional video display method, is merely illustrated by the division of the aforementioned functional modules. In actual applications, the aforementioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to perform all or part of the functions described above. Furthermore, the device provided in the above embodiment and the omnidirectional video display method embodiment are based on the same concept. The implementation process is detailed in the method embodiment and will not be repeated here.

[0103] An embodiment of the present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the steps of the method of any of the above embodiments are implemented.

[0104] See Figure 4 , is a structural block diagram of an electronic device provided in an embodiment of the present application.

[0105] like Figure 4 As shown, the electronic device 400 includes a processor 401 and a memory 402 .

[0106] In the embodiment of the present application, the processor 401 is the control center of the computer system and can be the processor of a physical machine or the processor of a virtual machine. The processor 401 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 401 can be implemented in the form of at least one hardware selected from the group consisting of a DSP (Digital Signal Processing), an FPGA (Field-Programmable Gate Array), and a PLA (Programmable Logic Array).

[0107] The processor 401 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state.

[0108] The memory 402 may include one or more computer-readable storage media, which may be non-transitory. The memory 402 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash memory storage devices. In some embodiments of the present application, the non-transitory computer-readable storage medium in the memory 402 is used to store at least one instruction, which is used to be executed by the processor 401 to implement the method in the embodiment of the present application.

[0109] In some embodiments, the electronic device 400 further includes a peripheral device interface 403 and at least one peripheral device 404. The processor 401, memory 402, and peripheral device interface 403 may be connected via a bus or signal lines. Each peripheral device 404 may be connected to the peripheral device interface 403 via a bus, signal lines, or circuit boards. Specifically, the peripheral devices 404 include a display screen, a camera, and an audio circuit. The peripheral device interface 403 may be used to connect at least one I / O (Input / Output)-related peripheral device to the processor 401 and memory 402.

[0110] In some embodiments of the present application, the processor 401, the memory 402, and the peripheral device interface 403 are integrated on the same chip or circuit board; in some other embodiments of the present application, any one or two of the processor 401, the memory 402, and the peripheral device interface 403 may be implemented on separate chips or circuit boards. This embodiment of the present application is not specifically limited to this.

[0111] The electronic device structure block diagram shown in the embodiment of the present application does not constitute a limitation on the electronic device 400. The electronic device 400 may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component arrangement.

[0112] The present application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method of any of the aforementioned embodiments. The computer-readable storage medium may include, but is not limited to, any type of disk, including a floppy disk, an optical disk, a DVD, a CD-ROM, a microdrive and a magneto-optical disk, a ROM, a RAM, an EPROM, an EEPROM, a DRAM, a VRAM, a flash memory device, a magnetic or optical card, a nanosystem (including a molecular memory IC), or any other type of medium or device suitable for storing instructions and / or data.

[0113] Through the description of the above embodiments, those skilled in the art will clearly understand that each embodiment can be implemented using software plus a necessary general-purpose hardware platform, or of course, hardware. Based on this understanding, the essence of the above technical solution, or the portion that contributes to the relevant technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, or an optical disk, and includes a number of instructions for causing a computer device (such as a personal computer, server, or network device) to execute the methods described in each embodiment or certain portions of the embodiments.

[0114] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. An omnidirectional video display method, characterized in that: The method is applied to a user terminal, and includes: In response to a first operation, generating a control instruction; the first operation is an interactive operation for acquiring omnidirectional video; Sending the control instruction to the image acquisition end, so that the image acquisition end determines the directional sub-eye to be called according to the control instruction, so that all the directional sub-eyes synchronously capture image frames at the same time based on the same acquisition frame rate; Obtaining the image frames collected by each directional sub-eye uploaded by the image acquisition end and the spatial positioning tags and time tags corresponding to the image frames; wherein the image frames collected by each directional sub-eye uploaded at the same time have the same time tag, temporal resolution, and spatial resolution; Acquire the posture information of the display screen at the time corresponding to the moment label, and respectively determine the posture offset of the spatial positioning tag of each directional sub-eye at the time corresponding to the moment label relative to the posture information; Determine the rendering area of ​​each image frame on the display screen according to the posture offset, and physically splice each image frame with the same moment label to obtain an omnidirectional video frame at the moment corresponding to the moment label; The omnidirectional video frame is rendered onto the display screen based on the correspondence between each image frame and the rendering area, so that the display screen displays the omnidirectional video dynamically and in real time according to the time sequence corresponding to the omnidirectional video frame.

2. The omnidirectional video display method according to claim 1, characterized in that: The method further comprises: In response to a second operation, updating the posture information of the display screen in real time; the second operation is an interactive operation that causes the posture information of the display screen to change; Based on the updated posture information, obtain a time tag after the posture information is updated, obtain a spatial positioning tag of each directional sub-eye at the time corresponding to the updated time tag, and update the posture offset of each directional sub-eye relative to the updated posture information based on the spatial positioning tag; Determine the rendering area of ​​each image frame on the display screen based on the updated posture offset, and physically splice each image frame with the same time label to obtain the omnidirectional video frame at the time corresponding to the time label; The omnidirectional video frame is rendered onto the display screen based on the correspondence between each image frame and the rendering area, so that the display screen displays the omnidirectional video dynamically and in real time according to the time sequence corresponding to the omnidirectional video frame.

3. The omnidirectional video display method according to claim 1, wherein: Physically splicing each image frame with the same time label to obtain an omnidirectional video frame at the time corresponding to the time label specifically includes: Acquire the range of the field of view angle of each two adjacent directional sub-eyes in real time, and determine the field of view intersection point of each two adjacent directional sub-eyes according to the range of the field of view angle of each two adjacent directional sub-eyes; Processing each image frame with the same label at the same time based on the field of view intersection, deleting the overlapping portion of the image frame of one directional sub-eye with the image frame of another directional sub-eye; The image frames after the overlapping parts are physically spliced ​​and deleted to obtain an omnidirectional video frame at the time corresponding to the time label.

4. The omnidirectional video display method according to claim 3, wherein: The step of determining a rendering area of ​​each image frame on a display screen according to the posture offset further includes: Acquiring a display resolution of the display screen, and determining a display range of the corresponding display screen based on the display resolution; The rendering position of each image frame in the display range is determined in combination with the posture offset; and the rendering area of ​​each image frame in the display range is determined according to the resolution and rendering position of each image frame.

5. The omnidirectional video display method according to any one of claims 1 to 4, characterized in that: The control instructions are also used to obtain omnidirectional video of historical periods; The method further comprises: Sending the control instruction to the image acquisition end so that the image acquisition end traverses all image frames corresponding to the historical period and the corresponding spatial positioning tags and time tags according to the stored identifier of each image frame; Arrange all video frames in chronological order, obtain the posture offset of the spatial positioning label of each image frame with the same moment label relative to the posture information for each historical moment, and physically stitch each image frame based on the posture offset to obtain the omnidirectional video frame corresponding to the historical moment; The omnidirectional video frame of each of the historical moments is sequentially rendered onto the corresponding display screen according to the time sequence.

6. The omnidirectional video display method according to any one of claims 1 to 4, characterized in that: When the number of display screens is the same as the number of called directional sub-eyes; The method further comprises: A corresponding relationship between a directional sub-eye and a display screen is determined respectively, and an image frame captured by each of the called directional sub-eyes is rendered respectively onto the display screen having the corresponding relationship.

7. An omnidirectional video display method, characterized in that: The method is applied to an image acquisition end, wherein the image acquisition end includes a plurality of directional sub-eyes; The method comprises: Obtain control instructions issued by the user end in real time; Determining the directional sub-eye called by the user terminal according to the control instruction; Allocate a sub-thread to the handle of the directional sub-eye called by the user end, read the image frames captured by the directional sub-eye corresponding to each sub-thread through the main thread, and obtain the spatial positioning label and time label of each called directional sub-eye in real time; The image frame and the corresponding spatial positioning tag and time tag are sent to the user terminal in real time, so that the user terminal obtains an omnidirectional video based on the image frame and the corresponding spatial positioning tag and time tag fusion and dynamically renders the omnidirectional video on the display screen in real time.

8. The omnidirectional video display method according to claim 7, characterized in that: Before sending the image frame and the corresponding spatial positioning tag and time tag to the user terminal, the method further includes: The image frames and the corresponding spatial positioning tags and time tags are stored in a cache area, and an identifier is assigned to the directional sub-eye corresponding to each image frame. A mapping relationship between the directional sub-eye and each image frame and the corresponding spatial positioning tag and time tag is determined through each identifier.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the method according to any one of claims 1 to 6 or the steps of the method according to any one of claims 7 to 8 are implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 or the steps of the method according to any one of claims 7 to 8 are implemented.