Camera Tracking Method, Device, Equipment and Storage Medium

By generating mask information and identification image information, and calculating camera tracking information in combination with Gaussian model and camera internal reference matrix, the problem of limited inertial jitter and mechanical transformation accuracy in extended real-life shooting is solved, and efficient stable synthesis of virtual background and display carrier is achieved.

CN114913308BActive Publication Date: 2025-08-01SHENZHEN UNI-LEADER TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210549117.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-20
Publication Date
2025-08-01
Estimated Expiration
2042-05-20

AI Technical Summary

Technical Problem

The existing cameras cannot effectively solve the relative drift and jitter between virtual background and display carrier in extended real-life shooting.

Method used

By obtaining the initial video information, generating mask information and identification image information, using a mixed Gaussian model segmentation algorithm and expansion and feathering operation to improve character and background segmentation efficiency, combining the camera internal reference and distortion correction matrix to calculate the camera tracking information, and directly render virtual scene information without mechanical transformation.

Benefits of technology

It improves the degree of matching between character video information and virtual scene information, reduces jitter error, improves playback effect and rendering efficiency, and avoids mechanical transformation of the gimbal and rocker arms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114913308B_ABST
    Figure CN114913308B_ABST
Patent Text Reader

Abstract

The present application relates to a camera tracking method, apparatus, device and storage medium. The method includes: obtaining initial video information, generating mask information according to the initial video information; generating identification image information according to the mask information; obtaining camera tracking information according to the identification image information; rendering virtual scene information according to the camera tracking information, where the virtual scene information includes virtual camera information from the perspective of the camera and virtual scene information displayed on the identification image information; obtaining replacement video information, where the replacement video information includes human video information and virtual scene video information containing the identification image information; and synthesizing the virtual scene information and the replacement video information to obtain synthesized video information. The technical effect of the present application is: to provide a camera tracking method in extended reality shooting.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of extended reality, and in particular, to a camera tracking method, device, equipment, and storage medium. Background Art

[0002] Extended Reality (XR) refers to combining the real and the virtual through a computer to create a virtual environment for human-computer interaction, which is also a collective term for various technologies such as AR, VR, and MR; Extended Reality is widely used in virtual production, program shooting, and live streaming; the main form is that a computer renders a virtual three-dimensional scene and outputs it to a relatively large-size display carrier, such as an LED screen, a projector, a television, etc., as the virtual background of the host. The camera shoots the host and the display carrier, and the captured image of the camera is input into the computer. The virtual background other than the display of the display carrier is also rendered by the computer.

[0003] When an existing camera may move during shooting, such as panning or tilting, in order to avoid relative drift and jitter between the virtual background and the virtual background displayed on the display carrier, it is necessary to input the motion state parameters of the camera for each frame into the computer. Based on these parameters, the relative motion of the virtual camera in the virtual scene is calculated. Gears are installed at the positions of the camera pan-tilt head, boom, and the rotating shaft on the camera to obtain the rotation angle parameters of the pan-tilt head and boom, and thus calculate the position and shooting angle of the camera to complete virtual camera tracking.

[0004] Based on the existing use of extended reality, the applicant believes that there are at least the following problems: mechanical modification of the pan-tilt head and boom is required, and the accuracy is limited by the accuracy of mechanical processing and sensors, and errors caused by problems such as jitter due to inertia cannot be solved. Summary of the Invention

[0005] In order to improve the problems of mechanical modification of the pan-tilt head and boom, limited accuracy due to mechanical processing and sensor accuracy, and jitter caused by inertia, the present application provides a camera tracking method and system.

[0006] In a first aspect, the present application provides a camera tracking method, adopting the following technical solution: The method includes:

[0007] Obtain initial video information, where the initial video information includes human video information and background information;

[0008] Generate mask information according to the initial video information;

[0009] Generate identification image information according to the mask information;

[0010] The camera captures video information containing the identification image;

[0011] Obtain camera tracking information in real time according to the video information, where the camera tracking information includes camera position information and camera angle information;

[0012] Render virtual scene information according to the camera tracking information, where the virtual scene information includes virtual camera information from the camera's perspective and virtual scene information for displaying identification image information;

[0013] Obtain replacement video information, where the replacement video information includes human video information and virtual scene video information containing identification image information;

[0014] Synthesize the virtual scene information and the replacement video information to obtain synthesized video information.

[0015] Through the above technical solution, first obtain the initial video information, generate mask information based on the human video information in the initial video information. The setting of the mask information improves the matching degree between the human video information and the virtual scene information, thereby improving the playback effect; generate identification image information in the corresponding initial video information according to the mask information, calculate the camera tracking information based on the generated identification image information, render the virtual scene information according to the camera tracking information, synthesize the replacement video information, where the replacement video information includes human video information and virtual scene video information containing identification image information, and synthesize the virtual scene information and the replacement video information to generate synthesized video information. There is no need to mechanically modify the pan-tilt head and jib, directly calculate the camera tracking information through the generated identification information, and then directly obtain the corresponding virtual scene information according to the camera tracking information, thus providing a camera tracking method in extended reality shooting.

[0016] In a specific feasible implementation, the generating mask information according to the initial video information includes:

[0017] Obtain the corresponding algorithm time information, Gaussian mixture length information, background ratio information, and noise intensity information according to the initial video information;

[0018] Process the initial video information according to the background-foreground segmentation algorithm of the Gaussian mixture model to generate video segmentation information; apply the video segmentation information to the initial video information to obtain mask information that matches the human video information.

[0019] Through the above technical solution, calculate and generate segmentation information through the background-foreground segmentation algorithm of the Gaussian mixture model, and quickly and conveniently segment the human and background in the video information through the preset algorithm time, Gaussian mixture intensity, background ratio, and noise intensity, thereby improving the segmentation efficiency of the human and background.

[0020] In a specific feasible implementation, after applying the video segmentation information to the initial video information to obtain mask information that matches the human video information, the following steps are further included:

[0021] Obtain a preset expansion iteration coefficient and an expansion detection matrix;

[0022] Perform an expansion operation on the mask information according to the expansion iteration coefficient and the expansion detection matrix;

[0023] Set the mask information after the expansion operation as the mask information.

[0024] Through the above technical solution, first obtain the human video information, and obtain a preset expansion iteration coefficient and an expansion detection matrix, and generate corresponding mask information according to the human video information, reducing the possibility that the mask area is relatively close to the edge area of the actual person, resulting in the person or the virtual background being easily mis-occluded when the person waves the arm or moves quickly, thereby directly improving the degree of fit between the human video information and the virtual scene information.

[0025] In a specific feasible implementation, after setting the mask information after the expansion operation as the mask information, the following steps are further included:

[0026] Obtain a preset average value blurring information and a feathering detection matrix;

[0027] Perform a feathering operation on the mask information according to the average value blurring information and the feathering detection matrix;

[0028] Set the mask information after the feathering operation as the mask information.

[0029] Through the above technical solution, by obtaining the average value blurring information and the feathering detection matrix, the staff can adjust the feathering degree of the mask information by adjusting the feathering detection matrix, thereby improving the feathering effect and reducing the possibility that the mask area is relatively close to the edge area of the actual person, resulting in the person or the virtual background being easily mis-occluded when the person waves the arm or moves quickly, further improving the degree of fit between the human video information and the virtual scene information.

[0030] In a specific feasible implementation, the generating of the identification image information according to the mask information includes:

[0031] Respectively obtain the pixel values corresponding to the mask information and the human video information;

[0032] Perform an inversion operation on the pixel values corresponding to the mask information and the human video information;

[0033] Overlay the image information with non-zero corresponding pixel values with the preset identification information and set the overlaid image information as the identification image information.

[0034] Through the above technical solution, through the negation operation of the pixel values, the background image information is fully replaced with the identification information, enabling the system to quickly and accurately replace the background image information.

[0035] In a specific feasible implementation, the real-time acquisition of camera tracking information according to the identification image information includes: The real-time acquisition of camera tracking information according to the identification image information includes:

[0036] Obtain the model origin coordinate information of the preset display screen model;

[0037] Establish a world coordinate system with the model origin coordinate information as the origin;

[0038] Statistically analyze the video two-dimensional coordinate matrix containing at least three marker points in the replacement video information;

[0039] Obtain the coordinate matrix of three identification points in the world coordinate system and set the coordinate matrix as the identification coordinate information;

[0040] Obtain the internal parameter matrix preset by the camera, and the internal parameter matrix can be expressed as, where u0, v0 are half of the pixel width and height of the replacement video information, β is the tilt parameter, f is the focal length value of the camera;

[0041] Obtain the preset distortion correction matrix;

[0042] Obtain the camera tracking information according to the video two-dimensional coordinate matrix, identification coordinate information, internal parameter matrix and distortion correction matrix.

[0043] Through the above technical solution, by establishing a world coordinate system, the system can generate camera tracking information quickly and conveniently according to the preset identification points on the identification image information and the marker points on the replacement video information, and according to the internal parameter matrix and distortion correction matrix of the camera, enabling the system to quickly and conveniently generate the corresponding virtual scene information according to the camera tracking information.

[0044] In a specific feasible implementation, the rendering of virtual scene information according to the camera tracking information includes:

[0045] Bind the coordinate information corresponding to the virtual camera to the camera tracking information;

[0046] Obtain the camera tracking information, the internal parameter matrix of the camera and the vector coordinate matrix of the virtual scene relative to the world coordinate system;

[0047] Calculate the vector coordinate matrix of the virtual scene in the camera view based on the camera internal parameter matrix, the camera tracking information, and the vector coordinate matrix of the virtual scene relative to the world coordinate system;

[0048] Obtain the vector coordinate matrix of the background image information in the world coordinate system;

[0049] Calculate the display screen of the background image information in the virtual camera view based on the camera internal parameter matrix, the camera tracking information, and the vector coordinate matrix of the background image information in the world coordinate system.

[0050] Through the above technical solution, after binding the virtual camera with the camera tracking information, the vector coordinate matrix of the virtual scene in the camera view can be calculated according to the camera tracking information, the camera internal parameter matrix, and the vector coordinate matrix of the virtual scene relative to the world coordinate system. After obtaining the vector coordinate matrix of the background image information in the world coordinate system, the display screen of the background image information in the virtual camera view is calculated, thereby improving the matching degree between the virtual scene and the camera, and thus enhancing the authenticity of the virtual scene.

[0051] In a second aspect, the present application provides a camera tracking device, adopting the following technical solution: The device includes: a video information acquisition module for acquiring initial video information, where the initial video information includes human video information and background information;

[0052] A mask information generation module for generating mask information according to the initial video information;

[0053] An identification image generation module for generating identification image information according to the mask information and capturing video information containing the identification image through the camera;

[0054] A camera coordinate acquisition module for acquiring camera tracking information in real time according to the video information, where the camera tracking information includes camera position information and camera angle information;

[0055] A virtual scene rendering module for rendering virtual scene information according to the camera tracking information, where the virtual scene information includes virtual camera information in the camera view and virtual scene information for displaying the identification image information;

[0056] A replacement video acquisition module for acquiring replacement video information, where the replacement video information includes human video information and virtual scene information containing the identification image information;

[0057] A composite video generation module for synthesizing the virtual scene information and the replacement video information and acquiring composite video information.

[0058] Through the above technical solution, initial video information is first obtained, and a mask information is generated based on the person video information in the initial video information. The setting of the mask information improves the matching degree between the person video information and the virtual scene information, thereby improving the playback effect; an identification image information is generated in the corresponding initial video information according to the mask information, the camera tracking information is calculated based on the generated identification image information, the virtual scene information is rendered according to the camera tracking information, and a replacement video information is synthesized. The replacement video information includes the person video information and the identification image information. The virtual scene information and the replacement video information are synthesized to generate a synthesized video information. There is no need to mechanically transform the pan-tilt head and jib, and the camera tracking information is directly calculated through the generated identification information, and then the corresponding virtual scene information is directly obtained according to the camera tracking information, thereby providing a camera tracking method in extended reality shooting.

[0059] In a third aspect, the present application provides a computer device, adopting the following technical solution: including a memory and a processor, and a computer program capable of being loaded and executed by the processor, such as any one of the above camera tracking methods, is stored on the memory.

[0060] Through the above technical solution, initial video information is first obtained, and a mask information is generated based on the person video information in the initial video information. The setting of the mask information improves the matching degree between the person video information and the virtual scene information, thereby improving the playback effect; an identification image information is generated in the corresponding initial video information according to the mask information, the camera tracking information is calculated based on the generated identification image information, the virtual scene information is rendered according to the camera tracking information, and a replacement video information is synthesized. The replacement video information includes the person video information and the identification image information. The virtual scene information and the replacement video information are synthesized to generate a synthesized video information. There is no need to mechanically transform the pan-tilt head and jib, and the camera tracking information is directly calculated through the generated identification information, and then the corresponding virtual scene information is directly obtained according to the camera tracking information, thereby providing a camera tracking method in extended reality shooting.

[0061] In a fourth aspect, the present application provides a computer-readable storage medium, adopting the following technical solution: storing a computer program capable of being loaded and executed by the processor, such as any one of the above camera tracking methods.

[0062] Through the above technical solution, initial video information is first obtained, and a mask information is generated based on the person video information in the initial video information. The setting of the mask information improves the matching degree between the person video information and the virtual scene information, thereby improving the playback effect; an identification image information is generated in the corresponding initial video information according to the mask information, the camera tracking information is calculated based on the generated identification image information, the virtual scene information is rendered according to the camera tracking information, and a replacement video information is synthesized. The replacement video information includes the person video information and the identification image information. The virtual scene information and the replacement video information are synthesized to generate a synthesized video information. There is no need to mechanically transform the pan-tilt head and the jib. The camera tracking information is directly calculated through the generated identification information, and then the corresponding virtual scene information is directly obtained according to the camera tracking information, thereby providing a camera tracking method in extended reality shooting.

[0063] In summary, the present application includes at least one of the following beneficial technical effects:

[0064] 1. First, obtain the initial video information, generate the mask information according to the person video information in the initial video information. The setting of the mask information improves the matching degree between the person video information and the virtual scene information, thereby improving the playback effect; generate the identification image information in the corresponding initial video information according to the mask information, calculate the camera tracking information based on the generated identification image information, render the virtual scene information according to the camera tracking information, synthesize the replacement video information, and the replacement video information includes the person video information and the identification image information. The virtual scene information and the replacement video information are synthesized to generate the synthesized video information. There is no need to mechanically transform the pan-tilt head and the jib. The camera tracking information is directly calculated through the generated identification information, and then the corresponding virtual scene information is directly obtained according to the camera tracking information, thereby providing a camera tracking method in extended reality shooting;

[0065] 2. By establishing a world coordinate system, the system can generate the camera tracking information quickly and conveniently according to the preset identification points on the identification image information and the marker points on the replacement video information, and according to the internal parameter matrix and distortion correction matrix of the camera, so that the system can quickly and conveniently generate the corresponding virtual information according to the camera tracking information. Description of the Drawings

[0066] Figure 1 is a flowchart of the camera tracking method in the embodiment of the present application.

[0067] Figure 2 is a schematic diagram of the enlarged mask operation in the embodiment of the present application.

[0068] Figure 3 is a schematic diagram of the feathered mask operation in the embodiment of the present application.

[0069] Figure 4 It is a structural block diagram of the camera tracking device in the embodiment of the present application.

[0070] Reference numerals: 401, video information acquisition module; 402, mask information generation module; 403, identification image generation module; 404, camera coordinate acquisition module; 405, virtual scene rendering module; 406, replacement video acquisition module; 407, composite video generation module. Specific embodiments

[0071] The following will Figures 1-4 further describe the present application in detail with reference to the accompanying

[0072] The embodiment of the present application discloses a camera tracking method. The method is based on a virtual production system, which includes a display screen arranged in an "L" shape, a camera for acquiring video information, and the camera is moved through a camera movement system to acquire video information from different angles. The camera transmits the captured video information to a processor, and the processor renders a virtual scene based on the video information and finally outputs composite video information including the virtual scene.

[0073] As Figure 1 shown, the method includes the following steps:

[0074] S10, acquire initial video information.

[0075] Among them, the initial video information includes human video information and background image information. The camera first directly shoots the display screen and the people on the display screen and sends the captured initial video information to the processor for waiting for the next step of processing.

[0076] S11, generate mask information.

[0077] Among them, first perform an expansion operation on the human video information to obtain mask information, and obtain feathered mask information according to the mask information.

[0078] S12, generate identification image information according to the mask information.

[0079] Among them, obtain the mask information, and synthesize the remaining part of the mask information with the preset identification image information.

[0080] S13, obtain camera tracking information in real time according to the identification image information.

[0081] Among them, the camera tracking information includes camera position information and camera angle information. Obtain the model origin coordinate information of the display screen model, establish a world coordinate system according to the model coordinate origin information, and inversely calculate the camera tracking information of the camera during the movement according to the identification point information on the identification image information.

[0082] S14, Render the virtual scene information according to the camera tracking information.

[0083] Among them, the virtual scene information includes the virtual camera information from the perspective of the camera and the virtual scene information displayed on the identification image information. By obtaining the camera tracking information, the perspective information of the camera can be obtained, and then the virtual scene information under the current perspective can be rendered.

[0084] S15, Obtain the synthesized video information.

[0085] Among them, the synthesized video information includes the person video information and the identification image information. By synthesizing the virtual scene information on the identification image information, and then synthesizing the person video information with the virtual scene information, the obtained synthesized video information includes the person video information captured by the camera, the background image information displayed on the display screen, and the virtual scene information rendered by the processor.

[0086] In one embodiment, considering that the background image information displayed on the display screen needs to be replaced, an operation of segmenting the initial video information is required. The segmentation operation can be specifically performed as follows:

[0087] Obtain the corresponding algorithm time information, Gaussian mixture length information, background ratio information, and noise intensity information according to the initial video information. Generate video segmentation information according to the background foreground segmentation algorithm of the Gaussian mixture model. Logically call the corresponding initial video information with the practical video segmentation information to obtain the mask information matching the person video information; enabling the system to automatically screen out the corresponding segmentation information according to the person video information, and apply the segmentation information to the corresponding initial video information, so that the system can quickly and conveniently generate the relevant mask information.

[0088] In one embodiment, as Figure 2 shown, considering that the background image information needs to be replaced, in order to reduce the possibility of the person video information or the virtual background information being wrongly blocked during the person's movement, an operation of expanding the mask based on the person video information is required. The specific operation of expanding the mask can be specifically performed as follows:

[0089] Assign corresponding pixel points to the acquired initial video information, where the person video information is assigned pixel point "1" and the background image information is assigned pixel point "0". First, obtain the dilation iteration coefficient, which is generally set to 1; the dilation detection matrix, which is generally set to a fifth-order matrix; the mask information in the initial video information. Detect the pixel points of the mask information in the initial video information and the pixel points of the surrounding dilation detection matrix. If the pixel points within the range of the dilation detection matrix in the initial video information contain "1", then set the corresponding pixel points to "1"; if all the pixel points within the range of the dilation detection matrix in the initial video information are "0", then set the corresponding pixel points to "0" to obtain the corresponding mask information. This reduces the possibility that the person video information or virtual scene information is blocked during the movement of the person video information, thereby improving the matching degree between the person video information and the virtual scene information.

[0090] In one embodiment, as Figure 3 shown, considering that the background information needs to be replaced, in order to further reduce the possibility that the person video information or virtual background information is wrongly blocked during the movement of the person, it is necessary to perform a feathering operation based on the mask information. The specific feathering operation can be specifically executed as follows:

[0091] Based on the pixel points corresponding to the mask information, obtain the preset average blur information, generally using average blur, and the feathering detection matrix, which is generally set to an eleventh-order matrix. Detect the pixel points corresponding to the mask information. If the pixel points within the range of the feathering detection matrix in the initial video information contain "1", then set the corresponding pixel points to "1"; otherwise, set the corresponding pixel points to "0" to obtain the corresponding mask information. This reduces the possibility that the person virtual scene information is wrongly blocked during the movement due to too many sharp corners in the mask information, and further improves the matching degree between the person video information and the virtual scene information.

[0092] In one embodiment, considering that it is necessary to perform a synthesis operation on the remaining background image information and the identification image information in the mask information of the initial video information, the specific synthesis operation can be specifically executed as follows:

[0093] Obtain the mask information and the background image information, and respectively obtain the pixel values corresponding to the mask information and the background image information. The pixel value corresponding to the mask information is "0", and the pixel value corresponding to the background image information is "1". The superposition formula for the mask information and the background image information is: M = (1 - N) * P, where M is the identification pattern with an alpha channel and superimposed with the identification image information, N is the mask information, and P is the image containing several identification point patterns. The form and quantity of the identification points can be set by the staff; this improves the identification efficiency of the system for the identification points in different environments.

[0094] In one embodiment, considering the need to obtain camera tracking information based on identification image information, the acquisition operation may be specifically performed as follows:

[0095] Obtain the origin information of the display screen model, establish a world coordinate system with the origin information of the display screen model as the origin, obtain the identification coordinate matrix of the preset identification point on the identification image information in the world coordinate system, count the video coordinate matrix of at least three marking points in the replacement video information in the world coordinate system, and obtain the camera preset intrinsic parameter matrix. The intrinsic parameter matrix can be expressed as, where u0 and v0 are half the pixel width and height of the replacement video information, β is the tilt parameter, f is the focal length of the camera, and the preset distortion correction matrix is obtained. The distortion correction matrix can be obtained through camera calibration, and the formula [M position , M rotation ]=OpenCV.solvePnP(M world ,M v ,M camera ,M distortion ) can calculate the camera tracking information, where [M position , M rotation ] is the position and rotation matrix of the camera relative to the origin in the world coordinate system, that is, the camera tracking information in the world coordinate system; M world is the coordinate matrix of several markers displayed on the display screen in the world coordinate system, M v To replace the coordinate matrix of at least three marker points in the video information, M camera is the intrinsic parameter matrix of the camera, M distortion It is the distortion correction matrix of the camera; it enables the system to automatically calculate the position tracking information of the camera corresponding to the shooting according to the existing camera position, identification coordinate matrix, video coordinate matrix and the camera's internal parameter matrix, so that the system can quickly and conveniently obtain the shooting angle of the virtual camera, and then render the virtual scene of the shooting angle through the processor, thereby improving the rendering efficiency of the system.

[0096] In one embodiment, considering the need to render virtual scene information based on camera tracking information, the specific rendering operation may be performed as follows:

[0097] After binding the virtual camera to the camera's rotation matrix and position matrix, the virtual scene of the camera's perspective can be obtained according to the camera perspective algorithm, which can be specifically expressed as V Clip =M camera ×[M position ,M rotation ]×V world , where V Clip is the vector coordinate matrix of the virtual scene in the camera's perspective, [Mposition , M rotation is the position and rotation matrix of the camera relative to the origin in the world coordinate system, and V world is the vector coordinate matrix of the virtual scene relative to the world coordinate system. After obtaining the vector coordinate matrix V clip of the camera, the calculation can be performed according to the formula V clip_screen = M camera × [M position , M rotation × V screen , where V clip-screen is the vector coordinate matrix of the display screen model in the perspective of the virtual camera, and V screen is the vector coordinate matrix of the display screen in the world coordinate system. Based on the vector coordinate matrix of the display screen in the perspective of the virtual camera, the display screen's display image in the perspective of the virtual camera can be obtained; enabling the system to automatically calculate and obtain the vector coordinate matrix of the display screen in the perspective of the virtual camera and the vector coordinate matrix of the display screen in the world coordinate system according to the camera position information, thereby improving the matching degree between the rendered scene information and the character video information in the perspective of the virtual camera.

[0098] The implementation principle of the embodiments of this application is as follows:

[0099] The implementation principle of the embodiments of this application is as follows: First, obtain the initial video information, which includes character video information and background image information. Generate the background based on the Gaussian mixture model background-foreground segmentation algorithm, generate the corresponding mask information according to the character video information, overlay the mask with the image containing a specific identification pattern, obtain the position and angle information of the camera in the world coordinate system, render the virtual scene according to the camera's position information, and synthesize the virtual scene and the video information to output the corresponding synthesized video information.

[0100] Based on the above method, the embodiments of this application also disclose a camera tracking device.

[0101] As Figure 4 shown, the device includes the following modules:

[0102] A video information acquisition module 401, configured to acquire initial video information, where the initial video information includes character video information and background image information;

[0103] A mask information generation module 402, configured to generate mask information according to the initial video information;

[0104] An identification image generation module 403, configured to generate identification image information according to the mask information;

[0105] The camera coordinate acquisition module 404 is used to obtain the camera tracking information in real time according to the identification image information, and the camera tracking information includes the camera position information and the camera angle information;

[0106] The virtual scene rendering module 405 is used to render the virtual scene information according to the camera tracking information, and the virtual scene information includes the virtual camera information from the camera's perspective and the virtual scene information displayed on the identification image information;

[0107] The replacement video acquisition module 406 is used to obtain the replacement video information, and the replacement video information includes the person video information and the virtual scene video information containing the identification image information;

[0108] The composite video generation module 407 is used to synthesize the virtual scene information and the replacement video information and obtain the composite video information.

[0109] In one embodiment, the mask information generation module 402 is further used to generate the mask information according to the initial video information, including: obtaining the preset algorithm time information, Gaussian mixture length information, background ratio information, and noise intensity information according to the initial video information; generating video segmentation information according to the background and foreground segmentation algorithm of the Gaussian mixture model; applying the video segmentation information to the initial video information to obtain the mask information matching the person video information.

[0110] In one embodiment, after applying the video segmentation information to the initial video information to obtain the mask information matching the person video information, the mask information generation module 402 further includes: obtaining the preset dilation iteration coefficient and dilation detection matrix; performing a dilation operation on the mask information according to the dilation iteration coefficient and the dilation detection matrix; setting the mask information after the dilation operation as the mask information.

[0111] In one embodiment, after setting the mask information after the dilation operation as the mask information, the mask information generation module 402 further includes: obtaining the preset average blur information and feathering detection matrix; performing a feathering operation on the mask information according to the average blur information and the feathering detection matrix; setting the mask information after the feathering operation as the mask information.

[0112] In one embodiment, the identification image generation module 403 is further used to generate the identification image information according to the mask information, including: respectively obtaining the pixel values corresponding to the mask information and the person video information; performing an inversion operation on the pixel values corresponding to the final mask information and the person video information; superimposing the image information with non-zero corresponding pixel values and the preset identification information and setting the superimposed image information as the identification image information.

[0113] In one embodiment, the camera coordinate acquisition module 404 is further configured to obtain camera tracking information in real time according to the identification image information, including: setting the origin of the display screen model as the origin to establish a world coordinate system; obtaining the identification coordinate matrix of the preset identification points on the identification image information in the world coordinate system; counting the video coordinate matrix of at least three marker points included in the replacement video information in the world coordinate system; obtaining the internal parameter matrix preset by the camera, and the internal parameter matrix can be expressed as, where u0 and v0 are half of the pixel width and height of the replacement video information, and β is the tilt parameter, f is the focal length value of the camera; obtaining a preset distortion correction matrix; and obtaining camera tracking information according to the identification coordinate information, video coordinate information, internal parameter matrix and distortion correction matrix.

[0114] In one embodiment, the camera coordinate acquisition module 404 is further configured to render virtual scene information according to the camera tracking information, including: binding the virtual camera to the camera tracking information; obtaining the camera tracking information, the internal parameter matrix of the camera, and the vector coordinate matrix of the virtual scene relative to the world coordinate system; calculating the vector coordinate matrix of the virtual scene in the camera view according to the internal parameter matrix of the camera, the camera tracking information, and the vector coordinate matrix of the virtual scene relative to the world coordinate system; obtaining the vector coordinate matrix of the background image information in the world coordinate system; and calculating the display screen of the background image information in the virtual camera view according to the internal parameter matrix of the camera, the camera tracking information, and the vector coordinate matrix of the background image information in the world coordinate system.

[0115] The embodiment of the present application also discloses a computer device.

[0116] Specifically, the computer device includes a memory and a processor, and a computer program capable of being loaded and executed by the processor for the above camera tracking method is stored on the memory.

[0117] The embodiment of the present application also discloses a computer-readable storage medium.

[0118] Specifically, the computer-readable storage medium stores a computer program capable of being loaded and executed by the processor for the camera tracking method as described above. The computer-readable storage medium includes, for example, various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.

[0119] This specific embodiment is only an explanation of the present invention, and it is not a limitation of the present invention. Those skilled in the art can make modifications to this embodiment without creative contributions according to needs after reading this specification, but as long as they are within the scope of the claims of the present invention, they are protected by the patent law.

Claims

1. A camera tracking method, characterized in that, The method includes: Obtaining initial video information, where the initial video information includes human video information and background information; Generating mask information according to the initial video information; Generating identification image information according to the mask information; A camera captures video information containing the identification image; Obtaining camera tracking information in real time according to the identification image information, where the camera tracking information includes camera position information and camera angle information; Rendering virtual scene information according to the camera tracking information, where the virtual scene information includes virtual camera information from the camera's perspective and virtual scene information for displaying the identification image information; Obtaining replacement video information, where the replacement video information includes human video information and virtual scene video information containing the identification image information; Combining the virtual scene information with the replacement video information and obtaining combined video information; Obtaining camera tracking information in real time according to the identification image information includes: obtaining the origin coordinate information of a preset display screen model, establishing a world coordinate system with the origin coordinate information of the display screen model as the origin, obtaining the coordinate matrix of three preset identification points on the identification image information in the world coordinate system and setting the coordinate matrix as the identification coordinate matrix, counting the video two-dimensional coordinate matrix of at least three marker points in the world coordinate system in the replacement video information, obtaining the internal parameter matrix preset for the camera, obtaining a preset distortion correction matrix, and the distortion correction matrix is obtained through camera calibration; Obtaining camera tracking information according to the video two-dimensional coordinate matrix, the identification coordinate matrix, the internal parameter matrix, and the distortion correction matrix; by establishing a world coordinate system, the system generates camera tracking information based on the preset identification points on the identification image information and the marker points on the replacement video information, and according to the internal parameter matrix and the distortion correction matrix of the camera, so that the system generates corresponding virtual scene information according to the camera tracking information.

2. The method according to claim 1, characterized in that, The generating mask information according to the initial video information includes: Obtaining corresponding algorithm time information, Gaussian mixture length information, background ratio information, and noise intensity information according to the initial video information; Processing the initial video information by the background-foreground segmentation algorithm of the Gaussian mixture model to generate video segmentation information; Applying the video segmentation information to the initial video information to obtain mask information that matches the human video information.

3. The method according to claim 2, characterized in that After the applying the video segmentation information to the initial video information to obtain mask information that matches the human video information, it further includes: Obtaining a preset dilation iteration coefficient and a dilation detection matrix; Performing a dilation operation on the mask information according to the dilation iteration coefficient and the dilation detection matrix; Setting the mask information after the dilation operation as the mask information.

4. The method according to claim 3, wherein After the setting the mask information after the dilation operation as the mask information, it further includes: Obtaining a preset average blur information and a feathering detection matrix; Performing a feathering operation on the mask information according to the average blur information and the feathering detection matrix; Setting the mask information after the feathering operation as the mask information.

5. The method according to claim 1, characterized in that, The generating identification image information according to the mask information includes: Respectively obtaining the pixel values corresponding to the mask information and the human video information; Invert the pixel values corresponding to the mask information and the human video information; Overlay the image information with non-zero corresponding pixel values and the preset identification information, and set the overlaid image information as the identification image information.

6. The method according to claim 1, wherein The rendering of the virtual scene information according to the camera tracking information includes: Bind the coordinate information corresponding to the virtual camera to the camera tracking information; Obtain the camera tracking information, the camera intrinsic matrix, and the vector coordinate matrix of the virtual scene relative to the world coordinate system; Calculate the vector coordinate matrix of the virtual scene in the camera view according to the camera intrinsic matrix, the camera tracking information, and the vector coordinate matrix of the virtual scene relative to the world coordinate system; Obtain the vector coordinate matrix of the background image information in the world coordinate system; Calculate the display picture of the background image information in the virtual camera view according to the camera intrinsic matrix, the camera tracking information, and the vector coordinate matrix of the background image information in the world coordinate system.

7. A camera tracking device, which applies the method according to any one of claims 1-6, characterized in that The device includes: A video information acquisition module (401) for acquiring initial video information, where the initial video information includes human video information and background information; A mask information generation module (402) for generating mask information according to the initial video information; An identification image generation module (403) for generating identification image information according to the mask information and capturing video information containing the identification image through a camera; A camera coordinate acquisition module (404) for acquiring camera tracking information in real time according to the identification image information, where the camera tracking information includes camera position information and camera angle information; A virtual scene rendering module (405) for rendering virtual scene information according to the camera tracking information, where the virtual scene information includes virtual camera information in the camera view and virtual scene information for displaying the identification image information; A replacement video acquisition module (406) for acquiring replacement video information, where the replacement video information includes human video information and virtual scene video information containing identification image information; A composite video generation module (407) for synthesizing the virtual scene information and the replacement video information and acquiring composite video information.

8. A computer device, characterized in that, Includes a memory and a processor, and a computer program capable of being loaded and executed by the processor, which is the method according to any one of claims 1 to 6, is stored on the memory.

9. A computer-readable storage medium, characterized in that, Stores a computer program capable of being loaded and executed by the processor, which is the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Augmented reality studio implementation method and device

    CN109688343A

  • An augmented reality image processing method and device based on optical positioning

    CN109840949A