Video fusion method, computer program product, client and storage medium
By directly integrating video in the client's three-dimensional engine, the problems of high hardware performance and large transmission delay in the existing technology are solved, and more efficient video fusion and better user experience are achieved.
Patent Information
- Application Number
- CN202111376391.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-19
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-11-19
AI Technical Summary
The existing video fusion technology has defects in high hardware performance requirements, large video streaming delay, and poor user experience.
By directly fusion of video in the client's three-dimensional engine, a virtual camera is used to synchronize the pictures on the playback panel and the pictures of the virtual scene, generate the fused video frames, and render and output on the client.
It reduces the requirements for hardware performance, reduces the transmission delay of video streams, improves the user experience, and enables the video fusion method to be effectively implemented on low computing devices.
Smart Images

Figure CN114358112B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of virtual reality technology, and more specifically to a video fusion method, a computer program product, a client, and a storage medium. Background Art
[0002] Video fusion technology is a branch of virtual reality technology, or a development stage of virtual reality. Video fusion technology refers to the fusion of one or more real videos of a scene or model captured by real cameras with a related virtual scene to generate a new video that is a mixture of real video and virtual scene.
[0003] In the existing video fusion technology, the conventional practice is to transmit the video stream captured by the real camera (i.e., the real video stream or the real video) to the server, and the server performs some processing on the real video stream (such as decoding, splicing multiple video streams into one video stream, etc.) and outputs a processed real video stream. In addition, the 3D engine of the server also outputs a virtual video stream containing the virtual scene after rendering the virtual scene. Subsequently, the server cuts and splices the real video stream and the virtual video stream together to form the final fused video stream. Finally, the server encodes the fused video stream and transmits it to the client, and the client decodes the fused video stream and outputs it to the display device for display.
[0004] Existing video fusion methods can achieve the fusion of real video and virtual scenes, but they have the following disadvantages: video cropping and splicing have high requirements on hardware performance, and may even require additional hardware devices to improve performance; video streams are transmitted between the server and the client, which requires redundant encoding and decoding operations, resulting in high latency in the client receiving the fused video stream and poor user experience. Summary of the invention
[0005] The present invention is proposed in view of the above problems. The present invention provides a video fusion method, a computer program product, a client and a storage medium.
[0006] According to one aspect of the present invention, a video fusion method is provided, comprising: obtaining real configuration information of a real camera and a video stream captured by the real camera, wherein the real configuration information includes initial pose information and viewing angle information, and the initial pose information includes initial position information and initial posture information; transforming the initial pose information based on a specific transformation relationship to obtain transformed pose information; assigning virtual configuration information to a virtual camera of a three-dimensional engine, the virtual configuration information including viewing angle information and transformed pose information; generating a playback panel, wherein a plane where the playback panel is located is parallel to an imaging plane of the virtual camera; for any current video frame in the video stream, matching the current video frame to the playback panel; and synchronously capturing a picture on the playback panel and a picture of a specific virtual scene through a virtual camera to obtain a fused video frame corresponding to the current video frame.
[0007] Exemplarily, after synchronously capturing the image on the playback panel and the image of a specific virtual scene through a virtual camera to obtain a fused video frame corresponding to the current video frame, the method further includes: rendering and outputting the fused video frame to a display device of the client for display on the display device.
[0008] Exemplarily, after the fused video frame is rendered and output to a display device of the client for display on the display device, the method further includes: receiving a modification instruction for virtual configuration information input by a user; modifying the virtual configuration information based on the modification instruction; and assigning the modified virtual configuration information to the virtual camera and returning to the step of synchronously capturing the picture on the playback panel and the picture of a specific virtual scene through the virtual camera to obtain a fused video frame corresponding to the current video frame.
[0009] Exemplarily, before synchronously capturing the image on the playback panel and the image of a specific virtual scene through a virtual camera to obtain a fused video frame corresponding to the current video frame, the method also includes: performing edge blurring on the current video frame to obtain a blurred video frame corresponding to the current video frame.
[0010] Exemplarily, performing edge blur processing on the current video frame to obtain a blurred video frame corresponding to the current video frame includes: performing color extraction on each pixel in the current video frame to obtain first color information; obtaining a specific blurred template image, the blur style of the specific blurred template image is a default style or is set based on style setting information input by a user; performing color extraction on each pixel in the blurred template image to obtain second color information, wherein the first color information and the second color information are both represented by a three-dimensional tensor, and the three dimensions in the three-dimensional tensor respectively represent the height of the image, the width of the image, and the color channel, wherein the color channel includes a transparency channel and a specific number of color channels; performing a tensor product calculation on the first color information and the second color information to obtain the blurred video frame.
[0011] Exemplarily, for any current video frame in the video stream, matching the current video frame to the playback panel includes: scaling the current video frame based on the size of the playback panel so that the size of the current video frame is consistent with the playback panel; and loading the video frame information of the scaled current video frame into the playback panel.
[0012] Exemplarily, before matching any current video frame in the video stream to the playback panel, the method further includes: decoding the video stream to obtain texture information of each video frame in the video stream, wherein the video frame information includes texture information of the current video frame.
[0013] According to another aspect of the present invention, a computer program product is provided, comprising computer program instructions, wherein the computer program instructions are used to execute the above-mentioned video fusion method when running.
[0014] According to another aspect of the present invention, a client is provided, including a processor and a memory, wherein the memory stores computer program instructions, and the computer program instructions are used to execute the above-mentioned video fusion method when the processor runs.
[0015] Exemplarily, the client also includes a display device, which is connected to the processor, and the display device is used to display the fused video frame. After the computer program instructions are executed by the processor to synchronously capture the images on the playback panel and the images of the specific virtual scene through a virtual camera to obtain a fused video frame corresponding to the current video frame, the computer program instructions are also used by the processor to execute: rendering and outputting the fused video frame to the display device for display on the display device.
[0016] Exemplarily, the client further includes a communication device, which is connected to the processor, and is used to receive real configuration information and video streams from the server and transmit the real configuration information and video streams to the processor.
[0017] Exemplarily, the client also includes a real camera, which is connected to the processor. The real camera is used to capture video streams and transmit real configuration information and video streams to the processor.
[0018] According to another aspect of the present invention, a storage medium is provided, on which program instructions are stored, and the program instructions are used to execute the above-mentioned video fusion method when running.
[0019] According to the video fusion method, computer program product, client and storage medium of the embodiment of the present invention, there is no need to perform cropping and splicing of video streams as in the prior art, which can reduce the performance requirements for hardware, which in turn enables the video fusion method of the present invention to be better implemented on the client, and even makes it possible to implement the method on client devices with relatively low computing power (such as mobile terminals). In addition, since the video fusion solution provided by the present invention directly performs fusion at the back end, there is no need to transmit the fused video stream between the server and the client, thus saving unnecessary encoding and decoding operations, reducing the reception delay of the fused video stream, and improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The above and other purposes, features and advantages of the present invention will become more apparent by describing the embodiments of the present invention in more detail in conjunction with the accompanying drawings. The accompanying drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings, the same reference numerals generally represent the same components or steps.
[0021] Figure 1 A schematic block diagram showing an example electronic device for implementing the video fusion method and apparatus according to an embodiment of the present invention;
[0022] Figure 2 A schematic flow chart of a video fusion method according to an embodiment of the present invention is shown;
[0023] Figure 3 A schematic diagram showing camera posture transformation according to an embodiment of the present invention is shown;
[0024] Figure 4 A schematic diagram showing the positional relationship among a virtual camera, a playback panel and a virtual scene according to an embodiment of the present invention;
[0025] Figure 5 A schematic diagram of a process of performing edge blurring processing on a video frame according to an embodiment of the present invention is shown;
[0026] Figure 6 A schematic diagram of a process of video fusion according to an embodiment of the present invention is shown;
[0027] Figure 7 A schematic block diagram showing a video fusion device according to an embodiment of the present invention; and
[0028] Figure 8 A schematic block diagram of a client according to an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0029] In recent years, research on computer vision, deep learning, machine learning, image processing, image recognition and other technologies based on artificial intelligence has made important progress. Artificial Intelligence (AI) is an emerging science and technology that studies and develops theories, methods, technologies and application systems for simulating and extending human intelligence. Artificial intelligence is a comprehensive discipline involving many types of technologies such as chips, big data, cloud computing, the Internet of Things, distributed storage, deep learning, machine learning, neural networks, etc. Computer vision, as an important branch of artificial intelligence, specifically allows machines to recognize the world. Computer vision technology usually includes face recognition, video fusion, fingerprint recognition and anti-counterfeiting verification, biometric recognition, face detection, pedestrian detection, target detection, pedestrian recognition, image processing, image recognition, image semantic understanding, image retrieval, text recognition, video processing, video content recognition, behavior recognition, 3D reconstruction, virtual reality, augmented reality, simultaneous localization and mapping (SLAM), computational photography, robot navigation and positioning and other technologies. With the research and advancement of artificial intelligence technology, this technology has been applied in many fields, such as security, urban management, traffic management, building management, park management, facial access, facial attendance, logistics management, warehouse management, robots, intelligent marketing, computational photography, mobile phone imaging, cloud services, smart homes, wearable devices, unmanned driving, automatic driving, smart medical care, facial payment, facial unlocking, fingerprint unlocking, identity verification, smart screens, smart TVs, cameras, mobile Internet, live streaming, beauty, makeup, medical beauty, and smart temperature measurement.
[0030] In order to make the purpose, technical scheme and advantages of the present invention more obvious, the exemplary embodiments according to the present invention will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments of the present invention, and it should be understood that the present invention is not limited to the exemplary embodiments described herein. Based on the embodiments of the present invention described in the present invention, all other embodiments obtained by those skilled in the art without creative work should fall within the protection scope of the present invention.
[0031] As mentioned above, the existing video fusion method adopts a working mode of back-end rendering and fusion with front-end display, wherein the back-end is the server and the front-end is the client. In the existing working mode of back-end rendering and fusion with front-end display, the server provides video fusion services to users, and the client is mainly responsible for displaying the fusion results. In this working mode, since the server must transmit the fused video stream to the client after completing the fusion, it is inevitable to encode the video stream. Based on this, in this working mode, rendering the virtual video stream and then cutting and splicing the video together with the real video stream is a relatively low-development fusion method. Therefore, limited by the working mode of back-end rendering and fusion with front-end display, the prior art generally adopts the above-mentioned fusion method of cutting and splicing video streams to achieve video fusion.
[0032] The embodiment of the present invention provides a video fusion method, a computer program product, a client, and a storage medium. According to the video fusion method of the embodiment of the present invention, video fusion is achieved by directly fusing on the client (specifically in the three-dimensional engine of the client). This method abandons the existing working mode of back-end rendering and fusing front-end display, so the drawbacks of the existing working mode can be solved. The video fusion technology according to the embodiment of the present invention can be applied to any field that requires virtual and real video fusion.
[0033] First, refer to Figure 1 An exemplary electronic device 100 for implementing the video fusion method and apparatus according to an embodiment of the present invention is described.
[0034] like Figure 1 As shown, the electronic device 100 includes one or more processors 102 and one or more storage devices 104. Optionally, the electronic device 100 may also include an input device 106, an output device 108, and an image acquisition device 110, and these components are interconnected via a bus system 112 and / or other forms of connection mechanisms (not shown). It should be noted that Figure 1 The components and structures of the electronic device 100 shown are merely exemplary and non-limiting. The electronic device may also have other components and structures as required.
[0035] The processor 102 can be implemented in at least one hardware form of a digital signal processor (DSP), a field programmable gate array (FPGA), a programmable logic array (PLA), and a microprocessor. The processor 102 can be a central processing unit (CPU), a graphics processing unit (GPU), an application specific integrated circuit (ASIC), or one or a combination of other forms of processing units with data processing capabilities and / or instruction execution capabilities, and can control other components in the electronic device 100 to perform desired functions.
[0036] The storage device 104 may include one or more computer program products, and the computer program product may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory (cache), etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 102 may run the program instructions to implement the client functions (implemented by the processor) and / or other desired functions in the embodiments of the present invention described below. Various applications and various data, such as various data used and / or generated by the application, may also be stored in the computer-readable storage medium.
[0037] The input device 106 may be a device used by a user to input instructions, and may include one or more of a keyboard, a mouse, a microphone, a touch screen, and the like.
[0038] The output device 108 can output various information (such as images and / or sounds) to the outside (such as a user), and can include one or more of a display, a speaker, etc. Optionally, the input device 106 and the output device 108 can be integrated together and implemented using the same interactive device (such as a touch screen).
[0039] The image acquisition device 110 can acquire images and store the acquired images in the storage device 104 for use by other components. The image acquisition device 110 can be a separate camera or a camera in a mobile terminal, etc. It should be understood that the image acquisition device 110 is only an example, and the electronic device 100 may not include the image acquisition device 110. In this case, other devices with image acquisition capabilities can be used to acquire images and send the acquired images to the electronic device 100.
[0040] Exemplarily, the exemplary electronic device for implementing the video fusion method and apparatus according to the embodiments of the present invention may be implemented on a device such as a personal computer or a remote server.
[0041] Next, we will refer to Figure 2 A video fusion method according to an embodiment of the present invention is described. Figure 2 A schematic flow chart of a video fusion method 200 according to an embodiment of the present invention is shown. The video fusion method 200 is used to run in a 3D engine of a client, that is, implemented by the 3D engine of the client.
[0042] Although this article divides the devices involved in video fusion into a server and a client when describing the technical problems to be solved by the present invention and the technical effects of the present invention, it should be noted that this does not mean a restriction on the implementation form of these devices themselves. The client described in this article can be understood as a client device used by a user who requests video fusion, and the user can interact with the client to control the process of video fusion, view the results of video fusion, etc. The client itself can be implemented using any suitable device, including but not limited to a personal computer, a mobile terminal, or a server device with server functions, etc.
[0043] The 3D engine described in this article can be any existing or future virtual modeling engine, including but not limited to one or more of the following: Unreal4 engine, Unity3D engine, etc. The 3D engine is a tool for virtual scene modeling and rendering. Implementing the entire video fusion method directly in the 3D engine is conducive to the rapid fusion and rendering of video streams.
[0044] Optionally, the video fusion method 200 described herein can be developed using a client / server (C / S) architecture, where the server is mainly used to transmit the real configuration information of the real camera and the video stream captured by the real camera to the client, and the client is responsible for video fusion (or fusion and rendering). The development language used to implement the video fusion method 200 described herein can be any suitable programming language, including but not limited to programming languages such as C++, UEC++, and Blueprint.
[0045] like Figure 2 As shown, the video fusion method 200 includes steps S210, S220, S230, S240, S250 and S260.
[0046] In step S210, real configuration information of the real camera and a video stream captured by the real camera are obtained, wherein the real configuration information includes initial posture information and viewing angle information, and the initial posture information includes initial position information and initial posture information.
[0047] A real camera is a camera arranged in the real physical world (hereinafter referred to as the real world). The number of real cameras can be arbitrary, and it can be one or more. In the case where the number of real cameras is multiple, the postures of the multiple real cameras need to be consistent. Exemplarily, in the case where the number of real cameras is multiple, the video streams captured by these multiple real cameras can be converged and spliced together to form a video stream as the video stream obtained in step S210 of the present application and the video stream that is subsequently involved in the fusion with the virtual scene.
[0048] In addition, what is obtained in step S210 is the real configuration information of a single real camera. In the case where there are multiple real cameras, the pose information and viewing angle information of the multiple real cameras can be integrated, and the integrated pose information and viewing angle information can be regarded as the pose information and viewing angle information of a single real camera. In this case, what is obtained in step S210 can be the integrated pose information and viewing angle information of multiple real cameras.
[0049] Optionally, the initial position information of the real camera may be geographic coordinate information or three-dimensional coordinate information in a Cartesian coordinate system. The geographic coordinate information may be information such as longitude and latitude.
[0050] The camera's posture information (including initial posture information and transformed posture information) may include the camera's pitch angle, yaw angle, and roll angle. Those skilled in the art may understand the meaning of the above-mentioned camera posture information, which will not be described in detail herein.
[0051] Exemplarily, one or more real cameras can be connected to a camera management system, which is responsible for the overall scheduling and management of the cameras. The camera management system can be implemented using a server device. The camera management system can be connected to a client (such as the above-mentioned electronic device 100) for implementing the video fusion method described herein. The camera management system can send the real configuration information of the real camera to the client described herein, and the client performs subsequent video fusion based on the real configuration information. Of course, the video stream captured by the real camera can also be sent to the client by the camera management system. In this case, the camera management system can be regarded as the server described herein.
[0052] In addition, optionally, one or more real cameras may also be directly connected to the client. In this case, each real camera may directly transmit its real configuration information and the collected video stream to the client. In addition, optionally, the client itself may include one or more real cameras. In this case, each real camera may transmit its real configuration information and the collected video stream to the client's processor (such as the above-mentioned processor 102) for video fusion. In addition, optionally, the real configuration information of the real camera may also be directly input to the client by the user, or may be pre-stored by the client in the client's storage device (such as the above-mentioned storage device 104).
[0053] In step S220, the initial posture information is transformed based on a specific transformation relationship to obtain transformed posture information.
[0054] The specific transformation relationship is the relationship information preset for realizing the transformation between the position and posture of the real camera and the position and posture of the virtual camera, which can be represented by any suitable relationship function. Optionally, the specific transformation relationship can also be represented by other suitable algorithm models (such as neural network models).
[0055] There is a certain transformation relationship between the camera in the real world and the virtual camera in the virtual world constructed by the 3D engine. Through the transformation, the initial posture information of the real camera can be transformed into the virtual world to obtain the transformed posture information. The transformed posture information can be used as the posture information of the virtual camera. In this way, the virtual camera in the 3D engine can maintain the same or basically the same observation effect as the real camera.
[0056] The specific transformation relationship may be preset by the user in advance. The user described herein may be any suitable person, including but not limited to a technical developer of the video fusion method and / or a user of the video fusion method. For example, the user may find out the transformation relationship between the position and posture of the camera in the real world and the position and posture of the camera in the virtual world in advance through theory or experiment, and use the relationship as the specific transformation relationship. When the video fusion is actually performed later, the specific transformation relationship may be used to transform the position and posture.
[0057] Figure 3 FIG. 2 is a schematic diagram showing a camera posture transformation according to an embodiment of the present invention. Figure 3 As shown, if the real configuration information of the real camera includes geographic coordinate information, it can be transformed into three-dimensional coordinate information, and then transformed according to the coordinate offset coefficient (i.e., the specific coordinate offset coefficient described in this article) to obtain the transformed position information. If the real configuration information of the real camera originally includes three-dimensional coordinate information, it can be directly transformed according to the coordinate offset coefficient to obtain the transformed position information. The initial posture information of the real camera can be transformed according to the rotation angle coefficient (i.e., the specific rotation angle coefficient described in this article) to obtain the transformed posture information. Among them, the specific transformation relationship can include the above-mentioned coordinate offset coefficient and rotation angle coefficient, and can optionally include the projection transformation matrix used when transforming the geographic coordinate information into three-dimensional coordinate information.
[0058] The viewing angle information refers to the size information of the angle formed by the center point of the camera lens to the two ends of the diagonal of the imaging plane. For a lens, the viewing angle mainly refers to the viewing angle range that it can achieve. The viewing angle of a camera may include the horizontal viewing angle and the vertical viewing angle within its viewing range. Accordingly, the viewing angle information may include the horizontal viewing angle information and / or the vertical viewing angle information.
[0059] The viewing angle information in the real configuration information can be directly assigned to the virtual camera without transformation. For example, if the viewing angle of the real camera is 60 degrees, the viewing angle of the virtual camera can also be set to 60 degrees.
[0060] In this article, the real configuration information obtained in step S210 and subsequently transformed is the real configuration information of a single real camera. As mentioned above, when there are multiple real cameras, the posture information and viewing angle information of multiple real cameras can be integrated. For example, the position information of multiple real cameras can be averaged, and the position information obtained after the average value is regarded as the position information of a single real camera. In this way, the position information obtained after the average value can be used as the initial position information obtained in step S210. As for the viewing angle information, the viewing angle information of multiple real cameras can be spliced, and the spliced viewing angle information can be regarded as the viewing angle information of a single real camera. In this way, the spliced viewing angle information can be used as the viewing angle information obtained in step S210.
[0061] Furthermore, regarding the pose information, as described above, a plurality of real cameras maintain the same pose, and therefore the pose information of any one of these cameras may be used as the initial pose information.
[0062] In step S230, virtual configuration information is assigned to a virtual camera of the three-dimensional engine, where the virtual configuration information includes viewing angle information and transformed posture information.
[0063] The transformed position information and posture information as well as the original viewing angle information may be assigned to a virtual camera, so that a camera having the same or substantially the same viewing effect as a real camera can be constructed in the virtual world.
[0064] In step S240, a playback panel is generated, wherein a plane where the playback panel is located is parallel to an imaging plane of the virtual camera.
[0065] The plane where the playback panel is located is parallel to the imaging plane of the virtual camera, that is, the plane where the playback panel is located is perpendicular to the main optical axis of the virtual camera. The playback panel is placed right in front of the virtual camera.
[0066] The size of the play panel can be set by the user or set as a default value by the 3D engine. The distance between the play panel and the virtual camera can also be set by the user or set as a default value by the 3D engine. For example, the size of the play panel can be determined according to the viewing angle of the virtual camera and the distance between the play panel and the virtual camera.
[0067] The viewing angle of the virtual camera and the distance between the playback panel and the virtual camera can determine the area of the imaging region of the virtual camera on the plane where the playback panel is located. Optionally, the playback panel can be set to be consistent or substantially consistent with the size of the imaging region of the virtual camera on the plane where the playback panel is located, so as to avoid blank images or loss of video stream information as much as possible. For example, the playback panel can be set to be compared with the imaging region of the virtual camera on the plane where the playback panel is located, and the height of the playback panel differs from the height of the imaging region by a first threshold value and the width of the playback panel differs from the width of the imaging region by a second threshold value. The first threshold value and the second threshold value can be set to appropriate values as needed, and these two threshold values can be set as small as possible.
[0068] Figure 4 A schematic diagram showing the positional relationship among a virtual camera, a playback panel, and a virtual scene according to an embodiment of the present invention is shown. Figure 4 , showing the playback panel. Figure 4 The positional relationship between the virtual camera and the playback panel can be understood. Although the virtual world is three-dimensional, the picture ultimately displayed by the display device is two-dimensional. In addition, the images captured by the camera in the real world are usually two-dimensional, so the images captured by the virtual camera also need to be mapped to a two-dimensional plane. The playback panel can be understood as a simulated imaging plane corresponding to the virtual camera.
[0069] In step S250, for any current video frame in the video stream, the current video frame is matched to the play panel.
[0070] The current video frame is matched to the playback panel, so that the video frame can be "presented" through the playback panel. The "presentation" means presenting to the virtual camera, that is, visible to the virtual camera. For the user, the video frame on the playback panel is preferably visible. However, this is not a limitation of the present invention, and the video frame on the playback panel may also be invisible to the user.
[0071] In one example, the current video frame may be each video frame in the video stream captured by the real camera, that is, each video frame in the video stream may be subjected to a fusion operation with a specific virtual scene (steps S250 and S260). In another example, the current video frame may be each video frame in a portion of the video frames in the video stream captured by the real camera, that is, a portion of the video frames may be selected from the video stream to perform a fusion operation with a specific virtual scene. For example, video frames whose quality meets preset requirements may be selected from the video stream to perform a fusion operation with a specific virtual scene. For another example, video frames within a preset period may be selected from the video stream to perform a fusion operation with a specific virtual scene.
[0072] Matching the current video frame to the playback panel may include scaling the current video frame to be consistent with the size of the playback panel to align with the playback panel, and loading the video frame information of the scaled current video frame into the playback panel. Loading the video frame information into the playback panel may be understood as setting the pixel information (such as the texture information described below) of each pixel of the current video frame as the pixel information at the corresponding position of the playback panel. Through the above process, the image content contained in the video frame can be presented on the playback panel.
[0073] In step S260, the images on the playback panel and the images of the specific virtual scene are synchronously captured by a virtual camera to obtain a fused video frame corresponding to the current video frame.
[0074] The specific virtual scene may be any virtual scene, including but not limited to a station scene, a building scene, a street scene, etc. Exemplarily, for any two video frames in a video stream, the specific virtual scenes respectively fused with the two video frames may be the same or different.
[0075] The virtual scene is pre-created by the user in the three-dimensional engine or imported into the three-dimensional engine from outside. The scene model is preferably three-dimensional, but the present application does not exclude the case of a two-dimensional model.
[0076] See also Figure 4 , through the acquisition of the virtual camera, the picture on the playback panel (which comes from the video frame captured by the real camera) and the picture of the specific virtual scene can be acquired synchronously. Therefore, the picture captured by the virtual camera itself is the fused picture (i.e., the fused video frame). Multiple fused video frames can form a fused video stream. Exemplarily, the picture captured by the virtual camera can be directly rendered and output to the display device, and the user can directly view the fusion effect of the fused video frame on the display device. Exemplarily, the picture captured by the virtual camera can also be encoded by any preset encoding method to obtain an encoded video frame (a plurality of encoded video frames is an encoded video stream). Optionally, the encoded video frame can be stored or further transmitted, etc.
[0077] According to the above description, the video frames captured by the real camera in this article are directly imaged in the virtual world, and the virtual camera can directly capture the fused images. This method belongs to a fusion method at the back end (ie, the client).
[0078] Since there is no need to perform cropping and splicing of video streams as in the prior art, compared with the existing working mode of front-end rendering and fusion with back-end display, the working mode of directly fusing at the back end described in this article can avoid the disadvantages caused by cropping and splicing operations. In other words, the video fusion scheme described in this article can reduce the performance requirements for hardware, which in turn enables the video fusion method of the present invention to be better implemented on the client, and even makes it possible to implement the method on client devices (such as mobile terminals) with relatively low computing power. In addition, since the fusion is performed directly at the back end, this working mode does not need to transmit the fused video stream between the server and the client, so it can save unnecessary encoding and decoding operations, reduce the reception delay of the fused video stream, and improve the user experience.
[0079] Exemplarily, the video fusion method according to the embodiment of the present invention may be implemented in a device, apparatus or system having a memory and a processor.
[0080] The video fusion method according to the embodiment of the present invention may be deployed at an image acquisition end, for example, may be deployed at a personal terminal or a server having an image acquisition function.
[0081] Alternatively, the video fusion method according to the embodiment of the present invention can also be deployed in a distributed manner on the server (or cloud) and the personal terminal. For example, the video stream captured by the camera can be received on the server (or cloud) or the video stream can be directly captured, and the server (or cloud) transmits the received or captured video stream to the client, and the client performs video fusion.
[0082] According to an embodiment of the present invention, after synchronously capturing the image on the playback panel and the image of a specific virtual scene through a virtual camera to obtain a fused video frame corresponding to the current video frame (step S260), method 200 may further include: rendering and outputting the fused video frame to a display device of the client for display on the display device.
[0083] The display device may be any suitable device having a display function that is currently available or may appear in the future. For example, the display device may include but is not limited to one or more of the following: a cathode ray tube display (CRT), a plasma display (PDP), a liquid crystal display (LCD), etc.
[0084] Since the present invention performs video fusion directly on the client, the fusion solution of the present invention allows the fused video frame (or fused video stream) to be directly rendered and output for display, without having to go through multiple operations such as intermediate encoding, transmission, and decoding on the client before it can be displayed on the client as in the prior art.
[0085] Optionally, the fused video frame (or fused video stream) can be directly rendered and output while being fused, so that users can view the fusion effect. Compared with the existing working mode of front-end rendering and fusion and back-end display, this solution of rendering and displaying while fusion can greatly reduce the display delay of the fused video frame and improve the performance of the video fusion algorithm, thus further improving the user experience.
[0086] According to an embodiment of the present invention, after rendering and outputting the fused video frame to a display device of the client for display on the display device, method 200 may further include: receiving a modification instruction for virtual configuration information input by a user; modifying the virtual configuration information based on the modification instruction; and assigning the modified virtual configuration information to the virtual camera and returning to step S260.
[0087] Due to the limitations of information collection, the real configuration information of the real camera may have certain errors (collection end errors). In addition, there may also be modeling errors on the virtual scene side. Therefore, if the user finds that there is a certain degree of deviation between the picture collected by the virtual camera and the picture collected by the real camera, you can choose to manually fine-tune the parameters of the virtual camera (that is, the virtual configuration information assigned to the virtual camera) to eliminate the above deviation. The deviation is generally small, and the user can choose whether to adjust it according to needs. The above adjustment process can be repeated until the picture collected by the virtual camera meets the user's needs.
[0088] By allowing the user to independently adjust the virtual configuration information of the virtual camera, it helps to adjust the virtual camera to a more accurate working state, which in turn helps to obtain a better video fusion effect.
[0089] According to an embodiment of the present invention, before synchronously capturing the image on the playback panel and the image of a specific virtual scene through a virtual camera to obtain a fused video frame corresponding to the current video frame (step S260), method 200 may further include: performing edge blur processing on the current video frame to obtain a blurred video frame corresponding to the current video frame.
[0090] Before fusion, the current video frame can be subjected to edge blur processing, so that the edge of the real video frame and the virtual scene can be smoothly transitioned, so that the fusion effect is more real and natural, and the viewing experience is better. The edge blur processing operation can be implemented by any existing or future edge blur technology. For example, the edge blur processing can be performed by using a shader in a 3D engine.
[0091] In the case of performing edge blur processing on the current video frame, the video frame information of the blurred video frame can be loaded into the playback panel. The two operations of scaling the current video frame to the same size as the playback panel to align with the playback panel and performing edge blur processing on the current video frame can be performed in any order. For example, the former can be performed before, after, or simultaneously with the latter. As long as the final result is that the video frame with blurred edges is presented on the playback panel.
[0092] According to an embodiment of the present invention, edge blurring is performed on a current video frame to obtain a blurred video frame corresponding to the current video frame, including: performing color extraction on each pixel in the current video frame to obtain first color information; obtaining a specific blurred template image, wherein the blur style of the specific blurred template image is a default style or is set based on style setting information input by a user; performing color extraction on each pixel in the blurred template image to obtain second color information, wherein the first color information and the second color information are both represented by a three-dimensional tensor, and the three dimensions in the three-dimensional tensor respectively represent the height of the image, the width of the image, and the color channel, wherein the color channel includes a transparency channel and a specific number of color channels; performing a tensor product calculation on the first color information and the second color information to obtain a blurred video frame.
[0093] The specific number may be any suitable number, which may be determined as required, and the present invention is not limited thereto. For example, in the case of color extraction based on the RGB color space, the specific number may be three, that is, there are three color channels: red, green and blue channels.
[0094] Figure 5 A schematic diagram of a process of performing edge blurring processing on a video frame according to an embodiment of the present invention is shown. Figure 5 In the example, the source image refers to any video frame in the video stream captured by the real camera. The blurred image refers to a specific blurred template image. In one example, a default blurred template image can be set in advance in the 3D engine as the specific blurred template image. In another example, the blurred template image can be set by the user. The following describes some exemplary implementations of the user setting the blurred template image by himself.
[0095] The user can input style setting information to the client through an input device (such as the above-mentioned input device 106). The style setting information may include but is not limited to text information, voice information, etc. Optionally, the user can directly input instructions (i.e., style setting information) for indicating which virtual template image to use as a specific virtual template image to the client in the form of text or voice, so as to instruct the client (specifically, the three-dimensional engine on the client) to select or generate a corresponding virtual template image. Optionally, the client can provide an interactive control to the user, and the user determines the specific virtual template image by interacting with the interactive control. For example, the interactive control can be a selection control related to multiple virtual template images, and the user can select one (or more of them) from multiple virtual template images as a specific virtual template image through the selection control. For another example, the interactive control can be a text control respectively related to one or more parameters of the virtual template image, and the user can set the corresponding parameters of the virtual template image through controls such as text controls or selection controls to obtain the desired specific virtual template image. The one or more parameters may include but are not limited to the color value and / or transparency value of each pixel and / or each pixel area (each pixel area may include multiple pixels).
[0096] Optionally, the transparency of any area of a specific blurred template image can be set. Exemplarily, the specific blurred template image can be set so that the transparency of pixels in one or more edge areas is less than a preset threshold value, and the preset threshold value can be any suitable value, and the present invention is not limited to this. For example, the preset threshold value can be 10%, 30%, 50%, 60%, etc. The size of each edge area can also be set arbitrarily, and it can include any number of pixels. The transparency value and color value of pixels between different edge areas can be the same or different. The above-mentioned edge area refers to an area relatively close to the edge of the image. However, this is not a limitation of the present invention, and the transparency of the specific blurred template image can be set in any area (including the central area).
[0097] Optionally, the color value of any area of the specific virtual template image can also be set. The color value described herein can be a value of any suitable color space, including but not limited to RGB value, YUV value, etc.
[0098] like Figure 5As shown, the source image and the specific blurred template image can be RGBA split respectively. RGBA splitting refers to extracting the R (red) value, G (green) value, B (blue) value and A (transparency) value of each pixel in the image respectively. In this way, for each pixel of the image, four channels of color data can be obtained. The above four channels are color channels, which include a transparency channel and three color channels. Assume that the number of pixels of the image is x×y, where x can represent the height of the image (that is, the number of pixels in the vertical direction of the image), and y can represent the width of the image (that is, the number of pixels in the horizontal direction of the image). Performing RGBA splitting on an image of size x×y can obtain a three-dimensional tensor of size x×y×4. The above RGBA splitting can be performed on the current video frame and the specific blurred template image respectively, and the results of the splitting of the two can be calculated by tensor product, and the obtained result can be used as the blurred video frame.
[0099] Figure 5 The operation of RGBA splitting shown is only an example and is not intended to limit the present invention. For example, the three RGB color channels involved in RGBA splitting can also be replaced by other color channels, such as YUV channels. Figure 5 In the example shown, the result of the current video frame splitting (i.e., the first color information) shows R, G, B, and A=1, which means that the split result is the original RGB value of each pixel and the transparency A of each pixel is 1 (i.e., 100%). Figure 5 In the example shown, the result of splitting the specific blurred template image (i.e., the second color information) shows R=1, G=1, B=1, and A, which means that the specific blurred template image does not adjust the color value of each pixel of the current video frame (the color value of each pixel maintains the original RGB value), but multiplies the transparency of each pixel of the current video frame by a corresponding coefficient A (the coefficient A is the transparency of each pixel of the specific blurred template image). The coefficient A corresponding to any two pixels in the current video frame can be the same or different. For example, it can be set so that the closer to the image edge of the current video frame, the smaller the corresponding coefficient A is, so as to form a gradual transparency adjustment mode.
[0100] Figure 5 The blurring method shown is only an example and not a limitation of the present invention, and the present invention may be implemented by other suitable blurring methods. For example, the specific blurring template image may be set to have a color value of at least some pixels that is not 1, so that the color values of the pixels corresponding to the pixels whose color values are not 1 in the current video frame may be adjusted.
[0101] According to an embodiment of the present invention, for any current video frame in the video stream, matching the current video frame to the playback panel (step S250) may include: scaling the current video frame based on the size of the playback panel so that the size of the current video frame is consistent with the playback panel; and loading the video frame information of the scaled current video frame into the playback panel.
[0102] The implementation method of matching the current video frame to the operation on the playback panel has been described above and will not be repeated here.
[0103] According to an embodiment of the present invention, before matching the current video frame to the playback panel (step S250) for any current video frame in the video stream, method 200 may further include: decoding the video stream to obtain texture information of each video frame in the video stream, wherein the video frame information includes texture information of the current video frame.
[0104] The video stream may be decoded using any suitable existing or future decoding method. For example, the OpenCV codec, ffmpeg codec, etc. may be used for decoding. When the 3D engine is implemented using the Unreal4 engine, the OpenCV decoding method may be used to decode the video stream. The OpenCV decoding is relatively compatible with the Unreal4 engine, and the use of this decoding method may enable the video fusion method provided in the embodiment of the present invention to work better in the Unreal4 engine.
[0105] After the video stream is decoded, the texture information of each video frame in the video stream can be obtained. The texture information of any video frame is transmitted to the playback panel, and the picture of the video frame can be presented in the playback panel.
[0106] The above decoding operation is only an example and not a limitation of the present invention, and it can be performed when necessary. For example, in the case where the video stream obtained by the 3D engine is a video stream that has been encoded and cannot be recognized by the engine, subsequent operations such as matching the playback panel can be performed after decoding the video stream. For another example, in the case where the video stream obtained by the 3D engine is a video stream that can be recognized by the engine, subsequent operations such as matching the playback panel can be performed directly on the video stream.
[0107] According to an embodiment of the present invention, transforming the initial posture information based on a specific transformation relationship to obtain transformed posture information (step S220) may include: determining the corresponding three-dimensional coordinate information based on the initial position information; transforming the three-dimensional coordinate information based on a specific coordinate offset coefficient to obtain transformed position information, wherein the transformed posture information includes the transformed position information, and the specific transformation relationship includes the specific coordinate offset coefficient.
[0108] The implementation method for transforming the initial position information has been described above and will not be repeated here.
[0109] According to an embodiment of the present invention, the initial position information is geographic coordinate information, and determining the corresponding three-dimensional coordinate information based on the initial position information may include: transforming the initial position information into a Cartesian coordinate system through a projection transformation matrix to obtain three-dimensional coordinate information, wherein the specific transformation relationship includes the projection transformation matrix.
[0110] When the initial position information is geographic coordinate information, it can be transformed into three-dimensional coordinate information through a projection transformation matrix before subsequent transformation. When the initial position information is three-dimensional coordinate information, determining the corresponding three-dimensional coordinate information based on the initial position information may include: directly determining the initial position information as three-dimensional coordinate information.
[0111] According to an embodiment of the present invention, transforming the initial posture information based on a specific transformation relationship to obtain transformed posture information (step S220) may include: transforming the initial posture information based on a specific rotation angle coefficient to obtain transformed posture information, wherein the transformed posture information includes the transformed posture information, and the specific transformation relationship includes the specific rotation angle coefficient.
[0112] The implementation method of transforming the initial posture information has been described above and will not be repeated here.
[0113] According to an embodiment of the present invention, steps S210 to S260 are implemented through a packaged software development kit, and method 200 may also include: obtaining the software development kit; loading the software development kit into the three-dimensional engine; and in response to a video fusion instruction input by a user, calling the software development kit in the three-dimensional engine to trigger the execution of steps S210 to S260.
[0114] In the existing working mode of front-end rendering and fusion and back-end display, since decoding, rendering, cropping, splicing, encoding and other operations need to be performed on the server first, and then transmitted to the client for decoding, display and other operations, the entire video fusion algorithm is difficult to integrate into the 3D engine as a plug-in or module, resulting in low reuse and flexibility of the video fusion algorithm.
[0115] In the video fusion method according to the embodiment of the present invention, since posture transformation, picture fusion (and possible encoding and decoding) and other operations are implemented in the same client, these functions can be highly encapsulated and used as a separate module. This can effectively reduce the coupling degree of the project and improve the reusability and flexibility of the module.
[0116] The embodiment of the present invention encapsulates the algorithm of the entire video fusion method 200 into a software development kit (SDK), so that authorized users can download and install it on their own clients at any time for use. After the user downloads the SDK to his own client, the SDK can be loaded into the client's three-dimensional engine and used as a functional module (video fusion module) of the three-dimensional engine. When the user operates in the three-dimensional engine, the user can call the video fusion module by inputting a video fusion instruction. When the video fusion module is called, the SDK can start to automatically execute the video fusion method 200.
[0117] Figure 6 A schematic diagram of a video fusion process according to an embodiment of the present invention is shown. Figure 6 In the example shown, the initial position information of the real camera is geographic coordinate information, but as mentioned above, this is only an example and not a limitation of the present invention. Figure 6 , the latitude and longitude coordinates and posture of the real camera can be first transformed by matrix (the specific transformation relationship can be expressed by matrix), and the transformed posture information and the viewing angle information of the real camera can be assigned to the virtual camera. In addition, the video stream collected by the real camera can be decoded by OpenCV, blurred by Shader edges, matched with the playback panel, and the presentation result of the video stream collected by the real camera on the playback panel can be obtained. Subsequently, the virtual camera can be used to collect the playback panel screen and the screen of the virtual scene modeled in the three-dimensional engine to obtain a fused screen (i.e., a fused video frame or a fused video stream). Subsequently, the fused video frame or the fused video stream can be optionally directly rendered and output to the display for display. Of course, optionally, the fused video frame or the fused video stream can also be encoded, and then the encoded video frame or video stream is subjected to one or more of the following operations: storage, transmission to other devices in the client, transmission to other devices outside the client, etc. Of course, the above-mentioned scheme of directly rendering and outputting the fused video frame or fused video stream for display and the scheme of encoding the fused video frame or fused video stream for further processing can be implemented in the same embodiment.
[0118] According to another aspect of the present invention, a video fusion device is provided. Figure 7 A schematic block diagram of a video fusion device 700 according to an embodiment of the present invention is shown.
[0119] like Figure 7 As shown, the video fusion device 700 according to the embodiment of the present invention includes an acquisition module 710, a transformation module 720, an assignment module 730, a generation module 740, a matching module 750 and a collection module 760. The modules of the video fusion device 700 can respectively perform the above combined Figure 2The following describes the various steps / functions of the video fusion method 200. Only the main functions of the various components of the video fusion device 700 are described below, and the details described above are omitted.
[0120] The acquisition module 710 is used to acquire the real configuration information of the real camera and the video stream captured by the real camera, wherein the real configuration information includes initial posture information and viewing angle information, and the initial posture information includes initial position information and initial posture information. The acquisition module 710 can be composed of Figure 1 The processor 102 in the electronic device shown executes program instructions stored in the storage device 104 to implement.
[0121] The transformation module 720 is used to transform the initial posture information based on a specific transformation relationship to obtain transformed posture information. The transformation module 720 can be composed of Figure 1 The processor 102 in the electronic device shown executes program instructions stored in the storage device 104 to implement.
[0122] The assignment module 730 is used to assign virtual configuration information to the virtual camera of the 3D engine, where the virtual configuration information includes viewing angle information and transformed posture information. Figure 1 The processor 102 in the electronic device shown executes program instructions stored in the storage device 104 to implement.
[0123] The generation module 740 is used to generate a playback panel, wherein the plane where the playback panel is located is parallel to the imaging plane of the virtual camera. Figure 1 The processor 102 in the electronic device shown executes program instructions stored in the storage device 104 to implement.
[0124] The matching module 750 is used to match any current video frame in the video stream to the playback panel. The matching module 750 can be composed of Figure 1 The processor 102 in the electronic device shown executes program instructions stored in the storage device 104 to implement.
[0125] The acquisition module 760 is used to synchronously acquire the images on the playback panel and the images of the specific virtual scene through the virtual camera to obtain a fused video frame corresponding to the current video frame. Figure 1 The processor 102 in the electronic device shown executes program instructions stored in the storage device 104 to implement.
[0126] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0127] Figure 8 A schematic block diagram of a client 800 according to an embodiment of the present invention is shown. The client 800 includes a storage device (ie, memory) 810 and a processor 820 . Figure 8 The display device shown is only an example and is not intended to limit the present invention.
[0128] The storage device 810 stores computer program instructions for implementing corresponding steps in the video fusion method 200 according to the embodiment of the present invention.
[0129] The processor 820 is used to run the computer program instructions stored in the storage device 810 to perform corresponding steps of the video fusion method 200 according to the embodiment of the present invention.
[0130] In one embodiment, the computer program instructions are used by the processor 820 to execute the following steps when they are executed: Step S210: Acquire real configuration information of a real camera and a video stream captured by the real camera, wherein the real configuration information includes initial pose information and viewing angle information, and the initial pose information includes initial position information and initial posture information; Step S220: Transform the initial pose information based on a specific transformation relationship to obtain transformed pose information; Step S230: Assign virtual configuration information to a virtual camera of a three-dimensional engine, wherein the virtual configuration information includes viewing angle information and transformed pose information; Step S240: Generate a playback panel, wherein a plane where the playback panel is located is parallel to an imaging plane of the virtual camera; Step S250: For any current video frame in the video stream, match the current video frame to the playback panel; and Step S260: Synchronously capture the picture on the playback panel and the picture of a specific virtual scene through the virtual camera to obtain a fused video frame corresponding to the current video frame.
[0131] Exemplarily, the client 800 may further include a display device 830. The display device 830 is connected to the processor 820, and the display device 830 is used to display the fused video frame. After the computer program instructions are executed by the processor 820 to synchronously capture the picture on the playback panel and the picture of the specific virtual scene through the virtual camera to obtain the fused video frame corresponding to the current video frame, the computer program instructions are also used to execute when the processor 820 is running: rendering and outputting the fused video frame to the display device 830 for display on the display device 830.
[0132] The display device 830 is optional, and the client 800 may not include the display device 830. In this case, the fused video frame may be sent to other devices and displayed by the display devices of the other devices.
[0133] Exemplarily, the client 800 may further include a communication device (not shown), which is connected to the processor, and is used to receive real configuration information and video streams from the server and transmit the real configuration information and video streams to the processor.
[0134] The communication device may be any suitable device with communication function, including but not limited to various existing or future wired or wireless communication devices, such as Bluetooth communication devices, Wi-Fi communication devices, etc.
[0135] Exemplarily, the client may further include a real camera, which is connected to the processor 820 , and is used to capture a video stream and transmit the real configuration information and the video stream to the processor 820 .
[0136] In addition, according to an embodiment of the present invention, a computer program product is also provided, including computer program instructions, which are used to execute the corresponding steps of the video fusion method 200 according to an embodiment of the present invention when the computer program instructions are run by a computer or a processor, and are used to implement the corresponding modules in the video fusion device 700 according to an embodiment of the present invention.
[0137] In addition, according to an embodiment of the present invention, a storage medium is also provided, on which program instructions are stored, and when the program instructions are run by a computer or a processor, the corresponding steps of the video fusion method 200 according to the embodiment of the present invention are executed, and the corresponding modules in the video fusion device 700 according to the embodiment of the present invention are implemented. The storage medium may include, for example, a memory card of a smart phone, a storage component of a tablet computer, a hard disk of a personal computer, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disk read-only memory (CD-ROM), a USB memory, or any combination of the above storage media.
[0138] In one embodiment, when the program instructions are executed by a computer or a processor, the computer or the processor may implement the various functional modules of the video fusion device according to the embodiment of the present invention, and / or may execute the video fusion method according to the embodiment of the present invention.
[0139] In one embodiment, the program instructions are used to perform the following steps at runtime: Step S210: Obtain real configuration information of a real camera and a video stream captured by the real camera, wherein the real configuration information includes initial pose information and viewing angle information, and the initial pose information includes initial position information and initial posture information; Step S220: Transform the initial pose information based on a specific transformation relationship to obtain transformed pose information; Step S230: Assign virtual configuration information to a virtual camera of a three-dimensional engine, wherein the virtual configuration information includes viewing angle information and transformed pose information; Step S240: Generate a playback panel, wherein a plane where the playback panel is located is parallel to an imaging plane of the virtual camera; Step S250: For any current video frame in the video stream, match the current video frame to the playback panel; and Step S260: Synchronously capture the picture on the playback panel and the picture of a specific virtual scene through the virtual camera to obtain a fused video frame corresponding to the current video frame.
[0140] Each module in the client according to an embodiment of the present invention can be implemented by running computer program instructions stored in a memory by a processor of an electronic device implementing video fusion according to an embodiment of the present invention, or can be implemented when computer instructions stored in a computer-readable storage medium of a computer program product according to an embodiment of the present invention are executed by a computer.
[0141] Although example embodiments have been described herein with reference to the accompanying drawings, it should be understood that the above example embodiments are merely exemplary and are not intended to limit the scope of the present invention thereto. Various changes and modifications may be made therein by one of ordinary skill in the art without departing from the scope and spirit of the present invention. All such changes and modifications are intended to be included within the scope of the present invention as required by the appended claims.
[0142] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0143] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed.
[0144] In the description provided herein, a large number of specific details are described. However, it is understood that embodiments of the present invention can be practiced without these specific details. In some instances, well-known methods, structures and techniques are not shown in detail so as not to obscure the understanding of this description.
[0145] Similarly, it should be understood that in order to streamline the present invention and help understand one or more of the various inventive aspects, in the description of the exemplary embodiments of the present invention, the various features of the present invention are sometimes grouped together into a single embodiment, figure, or description thereof. However, the method of the present invention should not be interpreted as reflecting the following intention: the claimed invention requires more features than the features explicitly stated in each claim. More specifically, as reflected in the corresponding claims, the inventive point is that the corresponding technical problem can be solved with less than all the features of a single disclosed embodiment. Therefore, the claims following the specific embodiment are hereby expressly incorporated into the specific embodiment, wherein each claim itself serves as a separate embodiment of the present invention.
[0146] It will be understood by those skilled in the art that, except for mutually exclusive features, all features disclosed in this specification (including the accompanying claims, abstracts and drawings) and all processes or units of any method or device disclosed in this specification may be combined in any combination. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstracts and drawings) may be replaced by an alternative feature that provides the same, equivalent or similar purpose.
[0147] In addition, those skilled in the art will appreciate that, although some embodiments herein include certain features included in other embodiments but not other features, the combination of features of different embodiments is meant to be within the scope of the present invention and form different embodiments. For example, in the claims, any one of the claimed embodiments may be used in any combination.
[0148] The various component embodiments of the present invention may be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. It should be understood by those skilled in the art that a microprocessor or a digital signal processor (DSP) may be used in practice to implement some or all of the functions of some modules in a video fusion device or client according to an embodiment of the present invention. The present invention may also be implemented as a device program (e.g., a computer program and a computer program product) for executing part or all of the methods described herein. Such a program for implementing the present invention may be stored on a computer-readable medium, or may be in the form of one or more signals. Such a signal may be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.
[0149] It should be noted that the above embodiments illustrate the present invention rather than limit it, and that those skilled in the art may devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference symbol between brackets shall not be construed as a limitation on the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "one" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present invention may be implemented by means of hardware comprising a number of different elements and by means of a suitably programmed computer. In a unit claim enumerating a number of devices, several of these devices may be embodied by the same hardware item. The use of the words first, second, and third, etc., does not indicate any order. These words may be interpreted as names.
[0150] The above is only a specific embodiment or description of a specific embodiment of the present invention, and the protection scope of the present invention is not limited thereto. Any person skilled in the art who is familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. The protection scope of the present invention shall be based on the protection scope of the claims.
Claims
1. A video fusion method, used for running in a three-dimensional engine on a client, the method comprising: Acquire real configuration information of a real camera and a video stream captured by the real camera, wherein the real configuration information includes initial posture information and viewing angle information, and the initial posture information includes initial position information and initial posture information; Transforming the initial posture information based on a specific transformation relationship to obtain transformed posture information; Assigning virtual configuration information to a virtual camera of the three-dimensional engine, the virtual configuration information including the viewing angle information and the transformed posture information; Generate a playback panel, wherein a plane where the playback panel is located is parallel to an imaging plane of the virtual camera; For any current video frame in the video stream, matching the current video frame to the playback panel; and The virtual camera synchronously captures the picture on the playback panel and the picture of the specific virtual scene to obtain a fused video frame corresponding to the current video frame.
2. The method of claim 1, wherein: After synchronously capturing the picture on the playback panel and the picture of the specific virtual scene by the virtual camera to obtain a fused video frame corresponding to the current video frame, the method further includes: The fused video frame is rendered and output to a display device of the client to be displayed on the display device.
3. The method of claim 2, wherein: After rendering and outputting the fused video frame to a display device of the client for display on the display device, the method further includes: receiving a modification instruction input by a user for the virtual configuration information; Modifying the virtual configuration information based on the modification instruction; and The modified virtual configuration information is assigned to the virtual camera and the process returns to the step of synchronously capturing the image on the playback panel and the image of the specific virtual scene through the virtual camera to obtain a fused video frame corresponding to the current video frame.
4. The method of claim 1, wherein: Before synchronously capturing the picture on the playback panel and the picture of the specific virtual scene by the virtual camera to obtain a fused video frame corresponding to the current video frame, the method further includes: The current video frame is subjected to edge blurring processing to obtain a blurred video frame corresponding to the current video frame.
5. The method of claim 4, wherein: The performing edge blurring processing on the current video frame to obtain a blurred video frame corresponding to the current video frame comprises: Performing color extraction on each pixel in the current video frame to obtain first color information; Acquire a specific blur template image, wherein the blur style of the specific blur template image is a default style or is set based on style setting information input by a user; Performing color extraction on each pixel in the blurred template image to obtain second color information, wherein the first color information and the second color information are both represented by a three-dimensional tensor, and three dimensions in the three-dimensional tensor respectively represent an image height, an image width, and a color channel, wherein the color channel includes a transparency channel and a specific number of color channels; A tensor product calculation is performed on the first color information and the second color information to obtain the blurred video frame.
6. The method according to any one of claims 1 to 5, wherein: For any current video frame in the video stream, matching the current video frame to the playback panel includes: Scaling the current video frame based on the size of the playback panel so that the size of the current video frame is consistent with the playback panel; The scaled video frame information of the current video frame is loaded into the playback panel.
7. The method of claim 6, wherein: Before matching any current video frame in the video stream to the playback panel, the method further includes: The video stream is decoded to obtain texture information of each video frame in the video stream, wherein the video frame information includes texture information of the current video frame.
8. A computer program product, comprising computer program instructions, wherein the computer program instructions are used to execute the video fusion method according to any one of claims 1 to 7 when running.
9. A client comprising a processor and a memory, wherein: The memory stores computer program instructions, which are used by the processor to execute the video fusion method according to any one of claims 1 to 7 when the processor is running the computer program instructions.
10. The client according to claim 9, wherein: The client further includes a display device, which is connected to the processor and is used to display the fused video frame. After the step of synchronously capturing the picture on the playback panel and the picture of the specific virtual scene by the virtual camera to obtain a fused video frame corresponding to the current video frame is executed when the computer program instructions are executed by the processor, the computer program instructions are further used to execute: The fused video frame is rendered and output to the display device for display on the display device.
11. A computer-readable storage medium having program instructions stored thereon, wherein the program instructions are used to execute the video fusion method according to any one of claims 1 to 7 when running.
Citation Information
Patent Citations
Dual-camera video fusion distortion correction and viewpoint micro-adjustment method and system thereof
CN107392853A
Multi-dimensional adjustment rack, image recognition test system and testing method thereof
CN109854906A