Scene multi-focal-length target video coding method
By working in conjunction with multiple fixed-focus recording systems, the zoom recording system identifies and merges the highest-resolution fixed-focus video frame image data into the zoom video frame, solving the problem of multiple focal length targets being difficult to display clearly in conventional recording systems, and achieving high-definition display of multiple targets in the zoom video stream.
Patent Information
- Application Number
- CN202511018605.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2025-10-17
AI Technical Summary
Conventional video recording systems struggle to clearly display multiple targets at different focal lengths simultaneously in the same video stream, such as key elements like athletes, referees, and scoreboards in sports competitions.
By employing a zoom video recording system and multiple fixed-focus video recording systems working in concert, the system identifies the feature vectors of key targets, matches and fuses the image data of the highest-resolution fixed-focus camera video frames into the zoom camera video frames, thereby achieving multi-focal-length target video encoding.
It enhances the clarity of multiple targets in the recording system, ensuring that multiple targets in the zoom video stream maintain high clarity, and solves the problem that multiple focal length targets are difficult to display clearly at the same time in existing technologies.
Smart Images

Figure CN120812412A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of video coding, in particular to a scene multi-focus target video coding method. BACKGROUND
[0002] The conventional video recording system is difficult to clearly display various targets at multiple focus positions in a video stream when recording a video. For example, in a sports match, key elements such as athletes, referees and scoreboards are difficult to be clearly displayed in one picture simultaneously because they are scattered in different positions. Therefore, the conventional video recording system has certain use limitations. SUMMARY
[0003] In view of the defects in the prior art, the present application provides a scene multi-focus target video coding method, which can effectively solve the above problems.
[0004] The technical scheme adopted by the present application is as follows:
[0005] The present application provides a scene multi-focus target video coding method, comprising the following steps:
[0006] Step S1, determining key targets that need to be identified, obtaining a key target feature representation vector of each key target, and forming a key target identification feature matrix F={F1, F2,..., Fm} by the key target feature representation vectors. m};wherein, Fj represents the feature representation vector of the key target j; j=1, 2,..., m, and m is the number of key targets; j
[0007] Step S2, setting a zoom recording system and n fixed-focus recording systems facing a shooting scene according to the shooting scene requirements; wherein, each fixed-focus recording system is configured to have different shooting angles and focus positions; and the shooting angles of each fixed-focus recording system cover the shooting scene.
[0008] Step S3, simultaneously starting the zoom recording system and the n fixed-focus recording systems to perform multi-recording system cooperative shooting, and simultaneously obtaining a zoom camera video frame S and a fixed-focus camera video frame set {C1, C2,..., Cn} at the same collection time t; wherein, Ci represents the fixed-focus camera video frame collected by the fixed-focus recording system i at the collection time t, i=1, 2,..., n. t 1,t 2,t n,t i,t
[0009] Step S4, based on the key target identification feature matrix F={F1, F2,..., Fm}, identifying the key targets in the zoom camera video frame S and the fixed-focus camera video frame set {C1, C2,..., Cn}. m t m key targets are identified, and image data and position information of each identified key target j in the zoom video frame S t are obtained;
[0010] Based on the key target identification feature matrix F = {F1, F2,..., F m}, key target identification is performed on each fixed-focus video frame C i,t , k i,t key targets are identified, and image data and position information of each identified key target l in the fixed-focus video frame C i,t are obtained; where l = 1, 2,..., k i,t ;
[0011] Step S5, according to the image data and the position information, the same key target in the zoom video frame S t and each fixed-focus video frame C i,t is matched;
[0012] For each key target j, the image sharpness in each fixed-focus video frame C i,t and in the zoom video frame S t is compared, if the image sharpness in the zoom video frame S t is the highest, no image fusion processing is performed; otherwise, the image data and the position information of the highest sharpness of the key target j in each fixed-focus video frame C i,t are obtained;
[0013] Step S6, the obtained image data of the highest sharpness of the key target j is fused into the position of the same key target in the zoom video frame S t , to obtain the fused zoom video frame S t ;
[0014] Step S7, for the m key targets in the zoom video frame S t , steps S5 to S6 are performed, and then the processed zoom video frame S t is video compression encoded to obtain the multi-focal fusion encoded zoom video frame S t ;
[0015] Step S8, the encoded zoom video frame S t is output; t = t + 1; return to step S3.
[0016] Preferably, in step S1, the feature representation vector F j of the key target j = {F j,visual , Fj,spatial F j,temporal};
[0017] wherein:
[0018] F j,visual is the appearance feature vector of the key target j, which is obtained by extracting the color histogram, texture feature and shape description feature of the key target j;
[0019] F j,spatial is the position feature vector of the key target j, which is used to describe the position distribution and moving track pattern of the key target j in the picture;
[0020] F j,temporal is the time feature vector of the key target j, which is used to describe the time pattern and duration statistics of the key target j.
[0021] Preferably, the zoom video recording system is used to record the overall scene and dynamically adjust the focal length according to the scene requirement to obtain a panoramic video stream containing all the key targets when recording.
[0022] Preferably, each fixed-focus video recording system i is configured with a shooting angle θ i and a specific fixed focal length f i , and when recording the scene, the shooting angle θ i and the specific fixed focal length f i remain unchanged; wherein the shooting angle θ i is the shooting angle relative to the zoom video recording system.
[0023] Preferably, the image data and the position information of the key target j in the zoom video frame S t are represented as: E s,j,t ={T s,j,t , P s,j,t , B s,j,t}; wherein T s,j,t represents the image content matrix of the key target j in the zoom video frame S t at the collection time t, which is obtained by cropping the image of the key target j from the zoom video frame S t ; P s,j,t represents the position coordinate vector of the center of the key target j in the zoom video frame S t at the collection time t; and B s,j,t represents the boundary box coordinate vector of the minimum circumscribed rectangle region of the key target j in the zoom video frame S t at the collection time t.
[0024] The image data and the position information of each recognized key target l in the fixed-focus video frame C i,t are represented as: Ei,l,t ={T i,l,t ,P i,l,t ,B i,l,t}; where T i,l,t Represents the key target l in the fixed-focus camera video frame C at the acquisition time t i,t The image content matrix is obtained by taking the fixed-focus camera video frame C i,t The image of the key target l is obtained by cropping; i,l,t Represents the center of the key target l in the fixed-focus camera video frame C at the acquisition time t i,t Position coordinate vector of B i,l,t Represents the key target l in the fixed-focus camera video frame C at the acquisition time t i,t The bounding box coordinate vector of the minimum bounding rectangular area.
[0025] Preferably, step S4 further includes:
[0026] For the zoom camera video frame S t Each key target j identified in the image, and in each of the fixed-focus camera video frames C i,t For each key target l identified in step S1, the confidence evaluation function is used to calculate the cosine similarity Confidence with the feature representation vector of the corresponding key target in step S1. The recognition is considered valid only when the cosine similarity Confidence is greater than the threshold τ_threshold; among them, τ_threshold is set to 0.7-0.9.
[0027] Preferably, in step S5, the zoom camera video frame S is matched according to the image data and its position information. t and each of the fixed-focus camera video frames C i,t The same key objectives as in the 2015 / 2016 / Revised Implementation Plan, specifically:
[0028] Step S5.1: focus target j on zoom camera video frame S according to acquisition time t t The image content matrix T s,j,t , and the focus target l in the fixed-focus camera video frame C at the acquisition time t i,t The image content matrix T i,l,t , calculate the appearance similarity Simvisual(j,l) between the key target j and the key target l;
[0029] Step S5.2: focus target j in zoom camera video frame S based on acquisition time t t The position coordinate vector P s,j,t , with the center of the focus target l in the fixed-focus camera video frame C at acquisition time t i,t The position coordinate vector P i,l,tsimspatial(j, l) = exp(-||P
[0030] simspatial(j, l) = exp(-||P s,j,t -transform(P i,l,t ,M_i)|| 2 / 2σ 2 )
[0031] where transform(P i,l,t ,M_i) represents a transform function of P i,l,t from the fixed-focus video system i coordinate system to the zoom video system coordinate system; M_i is a 3x3 unit transform matrix, which is obtained by camera calibration of the fixed-focus video system i; and σ is a position tolerance parameter, which is set to 10%-20% of the image size of the key target l cropped in the fixed-focus video frame C i,t .
[0032] Step S5.3, the comprehensive similarity between the key target j and the key target l is calculated by using the following formula:
[0033] Sim(j, l) = a - Simvisual(j, l) + β - Simspatial(j, l)
[0034] where a and β are appearance similarity weight and position similarity weight respectively; a + β = 1, a, β > 0; for a scene with large target appearance change, the value of β is increased; for a scene with relatively stable target appearance, the value of a is increased.
[0035] Step S5.4, if the comprehensive similarity Sim(j, l) between the key target j and the key target l is greater than the matching threshold θ_match, it means that the key target j and the key target l match successfully, which are the same key target.
[0036] Preferably, for each key target j, the image sharpness in each of the fixed-focus video frames C i,t and in the zoom video frame S t is compared, specifically:
[0037] For each video frame, the sharpness quantitative score is obtained by using the following method, and the comparison of image sharpness is based on the sharpness quantitative score:
[0038] Step A: the video frame to be quantitatively scored for sharpness is called a target image, denoted as target image T, which has a width of M pixels and a height of N pixels, and the target image T is converted from the spatial domain to the frequency domain by using the following formula:
[0039]
[0040] wherein:
[0041] T(x, y) represents the pixel value at (x, y) position in spatial domain; j is imaginary unit in the formula; F(u, v) represents the complex coefficient in frequency domain; (x, y) position in spatial domain corresponds to (u, v) position in frequency domain;
[0042] Step B: for the complex coefficient F(u, v) in frequency domain, the following filtering operation is performed to retain high frequency components and filter out low frequency components to obtain the filtered image H(u, v):
[0043] if then let H(u, v) = 1; otherwise, let H(u, v) = 0;
[0044] wherein: ω_cutoff is the cutoff frequency threshold;
[0045] Step C: multiply the filtered image H(u, v) and the complex coefficient F(u, v) in frequency domain by element to realize frequency domain filtering and extract high frequency information in the image to obtain high frequency information representation F_high(u, v) using the following formula:
[0046] F_high(u, v) = F(u, v) ⊙ H(u, v)
[0047] wherein: ⊙ represents element-wise multiplication;
[0048] Step D: obtain the sharpness quantization score Sharpness_Score(T) using the following formula:
[0049]
[0050] wherein: U and V are the lengths of rows and columns in frequency domain, respectively.
[0051] Preferably, in step S6, the highest sharpness image data of the key target j obtained is fused into the position of the same key target in the zoom video frame S t to obtain the fused zoom video frame S t , specifically:
[0052] Step S6.1: the highest sharpness image data of the key target j obtained is represented as so that the center coordinates of the image data and the center coordinates of the same key target in the zoom video frame S t coincide, and if the sizes of the two do not match, the sizes are adjusted using bilinear interpolation, so that the image data Replaced into the zoom video frame S t the same key target in the zoom video frame S
[0053] Step S6.2, color and brightness correction is performed on the image data embedded in the zoom video frame S t the same key target in the zoom video frame S to make it the same as the color and brightness distribution in the zoom video frame S t , to obtain the first image data
[0054] Step S6.3, for each pixel point p in the first image data , its weight w(p) is calculated using the following formula:
[0055]
[0056] wherein: σ is a parameter for controlling the transition smoothness, set to 1 / 6 to 1 / 4 of the size of the first image data ;
[0057] Step S6.4, the first image data is fused to obtain the fused image I_fused using the following formula:
[0058]
[0059] wherein: represents the pixel value of the pixel point p in the first image data ; S_region(p) represents the pixel value of the pixel point p in the same key target region in the zoom video frame S t ; I_fused(p) represents the pixel value of the pixel point p in the fused image;
[0060] Step S6.5, the fused image I_fused is subjected to boundary smoothing processing to obtain the boundary smoothing processed image I_fused" using the following formula:
[0061] I_fused" = I_fused * G_smooth
[0062] wherein: G_smooth is a smoothing filter; G_smooth = (1 / 16) [1, 2, 1; 2, 4, 2; 1, 2, 1];
[0063] Step S6.6, the boundary smoothing processed image I_fused" is re-embedded into the target region in the zoom video frame S t to obtain the final fused zoom video frame S_enhanced(t);
[0064] S_enhanced(t) = S t + Mask 0 (I_fused" S_region)
[0065] Wherein: Mask is the target region mask matrix in which the same key target in the zoom video frame S t is located; S_region is the target region in which the same key target in the zoom video frame S t is located; S t represents the zoom video frame;
[0066] Step S6.7, video compression encoding is performed on the final fused zoom video frame S_enhanced(t), and output is performed.
[0067] The scene multi-focal length target video encoding method provided by the application has the following advantages:
[0068] The application provides a scene multi-focal length target video encoding method, which is a method capable of enhancing the definition of specific multiple targets in a video recording system. Specifically, the zoom video recording system and the fixed-focus video recording system are used to cooperatively work, the specific targets in the fixed-focus system and the zoom system are intercepted, the target with the highest definition is calculated through an algorithm, and the target data is fused into the frame of the zoom video stream, so that multiple targets in the zoom video stream can maintain high definition. BRIEF DESCRIPTION OF DRAWINGS
[0069] Figure 1 The flowchart of the scene multi-focal length target video encoding method provided by the application is shown.
[0070] Figure 2 The embodiment flowchart of the scene multi-focal length target video encoding method provided by the application is shown. DETAILED DESCRIPTION
[0071] In order to make the technical problems, technical solutions and beneficial effects of the application clearer, the application will be further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the application and do not limit the application.
[0072] The application provides a scene multi-focal length target video encoding method, which is a method capable of enhancing the definition of specific multiple targets in a video recording system. Specifically, the zoom video recording system and the fixed-focus video recording system are used to cooperatively work, the specific targets in the fixed-focus system and the zoom system are intercepted, the target with the highest definition is calculated through an algorithm, and the target data is fused into the frame of the zoom video stream, so that multiple targets in the zoom video stream can maintain high definition.
[0073] Reference is made toFigure 1 and Figure 2 The present invention provides a scene multi-focal length target video encoding method, comprising the following steps:
[0074] Step S1: determine the key targets that need to be identified, obtain the key target feature representation vector of each key target, and each of the key target feature representation vectors forms a key target identification feature matrix F = {F1, F2, ..., F m}; Among them, F j The feature representation vector representing the key target j; j = 1, 2, ..., m, where m is the number of key targets;
[0075] Specifically, the feature representation vector F of the key target j j ={F j,visual ,F j,spatial ,F j,temporal}; where: F j,visual is the appearance feature vector of the key target j, which is obtained by extracting the color histogram, texture features, and shape description features of the key target j; F j,spatial is the position feature vector of the key target j, which is used to describe the position distribution and movement trajectory pattern of the key target j in the picture; F j,temporal is the temporal feature vector of key target j, which is used to describe the temporal pattern and duration statistics of key target j.
[0076] Step S2: according to the shooting scene requirements, setting a zoom recording system and n fixed-focus recording systems facing the shooting scene; wherein each of the fixed-focus recording systems is configured to have a different shooting angle and focal length position; and the shooting angle of each of the fixed-focus recording systems covers the shooting scene;
[0077] When shooting, the zoom recording system is used to capture the entire scene, dynamically adjust the focal length according to the needs of the scene, and obtain a panoramic video stream containing all key targets.
[0078] Each fixed-focus video system i is configured with a shooting angle θ i With a specific fixed focal length f i , when shooting a scene, the shooting angle θ i With a specific fixed focal length f i remains unchanged; where the shooting angle θ i is the shooting angle relative to the zoom recording system.
[0079] Step S3, simultaneously start the zoom video system and n fixed focus video systems to perform multi-video system collaborative shooting, and simultaneously obtain zoom video frames S at the same acquisition time t t And fixed focus camera video frame set {C 1,t ,C2,t ,...,C n,t}; Among them, C i,t represents the fixed-focus video frame captured by the fixed-focus video system i at the capture time t, i = 1, 2, ..., n;
[0080] Step S4, based on the key target recognition feature matrix F={F1,F2,...,F m}, in the zoom camera video frame S t Identify m key targets in the zoom camera video frame S and obtain each identified key target j in the zoom camera video frame S t Image data and location information in ;
[0081] Based on the key target recognition feature matrix F={F1,F2,...,F m}, for each fixed-focus camera video frame C i,t Perform key target recognition and identify k i,t key targets, and obtain each identified key target l in the fixed-focus camera video frame C i,t Image data and its location information in; where l=1,2,...,k i,t ;
[0082] Specifically, the focus target j is in the zoom camera video frame S t The image data and its position information in E are expressed as: s,j,t ={T s,j,t ,P s,j,t ,B s,j,t}; where T s,j,t Represents the focus target j in the zoom camera video frame S at the acquisition time t t The image content matrix is obtained by taking the zoom camera video frame S t The image of the key target j is obtained by cropping; s,j,t represents the center of the focused target j in the zoom camera video frame S at the acquisition time t t Position coordinate vector of B s,j,t Represents the focus target j in the zoom camera video frame S at the acquisition time t t The bounding box coordinate vector of the minimum bounding rectangle area;
[0083] Each identified key target l is located in the fixed-focus camera video frame C i,t The image data and its position information in E are expressed as: i,l,t ={T i,l,t ,P i,l,t ,B i,l,t}; where T i,l,t Represents the key target l in the fixed-focus camera video frame C at the acquisition time t i,tThe image content matrix is obtained by taking the fixed-focus camera video frame C i,t The image of the key target l is obtained by cropping; i,l,t Represents the center of the key target l in the fixed-focus camera video frame C at the acquisition time t i,t Position coordinate vector of B i,l,t Represents the key target l in the fixed-focus camera video frame C at the acquisition time t i,t The bounding box coordinate vector of the minimum bounding rectangular area.
[0084] For the zoom camera video frame S t Each key target j identified in the image, and in each of the fixed-focus camera video frames C i,t For each key target l identified in step S1, the confidence evaluation function is used to calculate the cosine similarity Confidence with the feature representation vector of the corresponding key target in step S1. The recognition is considered valid only when the cosine similarity Confidence is greater than the threshold τ_threshold; among them, τ_threshold is set to 0.7-0.9.
[0085] Step S5, matching the zoom camera video frame S according to the image data and its position information t and each of the fixed-focus camera video frames C i,t The same key objectives in
[0086] The specific matching method is:
[0087] Step S5.1: focus target j in zoom camera video frame S according to acquisition time t t The image content matrix T s,j,t , and the focus target l in the fixed-focus camera video frame C at the acquisition time t i,t The image content matrix T i,l,t , calculate the appearance similarity Simvisual(j,l) between the key target j and the key target l;
[0088] Step S5.2: focus target j in zoom camera video frame S based on acquisition time t t The position coordinate vector P s,j,t , and the center of the focus target l at the acquisition time t is in the fixed-focus camera video frame C i,t The position coordinate vector P i,l,t , calculate the position similarity Simspatial(j,l) between the key target j and the key target l;
[0089] simspatial(j,l)=exp(-||P s,j,t -transform(P i,l,t ,M_i)||2 / 2σ 2 )
[0090] Where: transform(P i,l,t ,M_i) represents the P in the i coordinate system of the fixed focus video system i,l,t Transformation function to the zoom video system coordinate system; Mi is the 3×3 unit transformation matrix, obtained by calibrating the camera of the fixed-focus video system i; σ is the position tolerance parameter, set to the fixed-focus video frame C i,t 10%-20% of the image size of the focused object l cropped to the image;
[0091] In step S5.3, the comprehensive similarity between key target j and key target l is calculated using the following formula:
[0092] Sim(j,l)=α·Simvisual(j,l)+β·Simspatial(j,l)
[0093] Where: α and β are the appearance similarity weight and position similarity weight respectively; α+β=1,α,β>0; for scenes with large changes in target appearance, increase the value of β; for scenes with relatively stable target appearance, increase the value of α;
[0094] In step S5.4, if the comprehensive similarity Sim(j,l) between the key target j and the key target l is greater than the matching threshold θ_match, it means that the key target j and the key target l are successfully matched and are the same key target.
[0095] For each focused target j, compare its i,t and in the zoom camera video frame S t The image clarity in the zoom camera video frame S t If the image clarity in the image is the highest, no image fusion processing is performed; otherwise, the key target j is obtained in each of the fixed-focus camera video frames C i,t The highest definition image data and its location information;
[0096] Specifically, for each video frame, the following method is used to obtain its clarity quantification score; then, based on the clarity quantification score, the image clarity is compared:
[0097] Step A: The video frame that needs to be quantified for clarity is called the target image, which is expressed as: target image T, with a width of M pixels and a height of N pixels. The target image T is converted from the spatial domain to the frequency domain using the following formula:
[0098]
[0099] wherein:
[0100] T(x, y) represents the pixel value at (x, y) position in spatial domain; j is imaginary unit in the formula; F(u, v) represents the complex coefficient in frequency domain; (x, y) position in spatial domain corresponds to (u, v) position in frequency domain;
[0101] Step B: for the complex coefficient F(u, v) in frequency domain, the following filtering operation is performed to retain high frequency components and filter out low frequency components to obtain the filtered image H(u, v):
[0102] If then let H(u, v) = 1; otherwise, let H(u, v) = 0;
[0103] wherein: ω_cutoff is the cutoff frequency threshold;
[0104] Step C: multiply the filtered image H(u, v) and the complex coefficient F(u, v) in frequency domain by element to realize frequency domain filtering and extract high frequency information in the image to obtain high frequency information representation F_high(u, v) using the following formula:
[0105] F_high(u, v) = F(u, v) ⊙ H(u, v)
[0106] wherein: ⊙ represents element-wise multiplication;
[0107] Step D: obtain the sharpness quantization score Sharpness_Score(T) using the following formula:
[0108]
[0109] wherein: U and V are the lengths of rows and columns in frequency domain, respectively.
[0110] Step S6, fuse the highest sharpness image data of the key target j obtained to the same key target position in the zoom video frame S t to obtain the fused zoom video frame S t .
[0111] This step specifically includes:
[0112] Step S6.1, represent the highest sharpness image data of the key target j obtained as so that the center coordinates of the image data and the center coordinates of the same key target in the zoom video frame S t coincide, and if the sizes of the two do not match, perform size adjustment using bilinear interpolation, so that the image data is replaced to the zoom video frame St within the region with the same key objectives;
[0113] Step S6.2, embedding the zoom camera video frame S t Image data of the same key target Perform color and brightness correction to match the zoom camera video frame S t The color and brightness distribution in the image are the same, and the first image data is obtained.
[0114] Step S6.3, the first image data For each pixel point p in , the following formula is used to calculate its weight w(p):
[0115]
[0116] Where: σ is the parameter that controls the transition smoothness, set to the first image data 1 / 6 to 1 / 4 of the size;
[0117] Step S6.4, using the following formula, the first image data The pixels p in the image are fused to obtain the fused image I_fused:
[0118]
[0119] in: Represents the first image data The pixel value of the pixel point p in the image; S_region(p) represents the zoom camera video frame S t I_fused(p) represents the pixel value of the pixel point p in the same key target area in the fused image.
[0120] In step S6.5, the fused image I_fused is subjected to boundary smoothing using the following formula to obtain a boundary smoothed image I_fused":
[0121] I_fused"=I_fused*G_smooth
[0122] Where: G_smooth is the smoothing filter; G_smooth = (1 / 16) [1, 2, 1; 2, 4, 2; 1, 2, 1];
[0123] Step S6.6, re-embedding the image I_fused" after boundary smoothing into the zoom camera video frame S t The target area in the image is obtained to obtain the final fused zoom camera video frame S_enhanced(t);
[0124] S_enhanced(t)=S t +Mask⊙(I_fused"-S_region)
[0125] Wherein: Mask is a target region mask matrix in which the same key target in the zoom video frame S t is located; S_region is a target region in which the same key target in the zoom video frame S t is located; S t represents the zoom video frame;
[0126] Step S6.7, video compression encoding is performed on the final fused zoom video frame S_enhanced(t), and the zoom video frame S_enhanced(t) is outputted.
[0127] Step S7, for m key targets in the zoom video frame S t , steps S5 to S6 are performed, and video compression encoding is performed on the processed zoom video frame S t , to obtain a multi-focal fused encoded zoom video frame S t .
[0128] Step S8, the encoded zoom video frame S t is outputted; t is set to t+1; and the step S3 is returned.
[0129] The scene multi-focal target video encoding method provided by the application combines the cooperative work of the zoom video system and the multiple fixed-focus video systems. Firstly, according to the preset scene requirements, the shooting angle and the focal length position of each fixed-focus video system are accurately set to capture the target details in different depth of field ranges; through the zoom video system and the multiple fixed-focus video systems, one zoom video stream and multiple fixed-focus video streams can be obtained. Using the artificial intelligence image recognition technology, the target regions corresponding to the multiple target pictures or names predefined by the user are automatically recognized and accurately intercepted from these video streams. Through the position information and the feature matching algorithm, the corresponding relationship between the specific targets in the zoom video stream and the fixed-focus video stream is established. For each target, the algorithm traverses all the target instances in the related video streams, and selects the highest definition target image data in the frequency domain after the band-pass filtering. Then the highest definition target image data is replaced into the corresponding target position in the zoom video stream, and the replaced region is smoothed, so that the definition of all key targets in the finally generated zoom video stream reaches the optimal state. Finally, the video compression encoding technology is used to compress the optimized zoom video stream.
[0130] An embodiment will be introduced below:
[0131] (1) Predefine the target recognition database that needs to be enhanced, and the target that needs to be enhanced can be a photo, a name or a feature description;
[0132] Specifically, the system constructs a database containing all important target feature information, including optical properties, spatial position and timing relationship elements, forming a complete target feature library.
[0133] Specifically, the user needs to predefine the important target set O={O1, O2,..., O_m} that he wants to keep high definition in the video, where m is the total number of targets. This predefinition solves the technical problem that the traditional recording system cannot distinguish important targets from the background. The target can be specified in the following three flexible ways:
[0134] Provide a reference photo of the target: directly upload the standard image of the target, suitable for fixed targets with known appearance;
[0135] Enter the name of the target: such as "athlete", "referee", "scoreboard", the system will call the pre-trained semantic recognition model;
[0136] Describe the features of the target: such as "person wearing a red jersey", supporting flexible description based on attributes.
[0137] In order to realize accurate recognition, the system will establish a multi-dimensional feature vector for each target, which is one of the key technical innovations of the invention. The feature vector contains three types of complementary information:
[0138] Appearance feature vector: extract the color histogram, texture feature, shape description and other visual characteristics of the target, usually 512-1024 dimensions;
[0139] Position feature vector: records the common position distribution and movement trajectory pattern of the target in the picture, usually 64-128 dimensions;
[0140] Time feature vector: describes the time pattern and duration statistics of the target, usually 32-64 dimensions.
[0141] The complete feature representation of each target uses vector splicing, that is: the appearance feature vector, position feature vector and time feature vector are spliced, and the total dimension of the spliced feature vector is usually between 600-1200 dimensions.
[0142] The feature files of all targets are combined into a target recognition database matrix, which serves as a lookup table for subsequent recognition algorithms, supporting fast similarity calculation and target matching.
[0143] (2) The zoom video system and multiple fixed-focus video systems are started simultaneously for recording. At the same time, the frame S1 in the zoom video stream and the frame set {C1, C2, ..., Cn} in the fixed-focus video stream are obtained respectively.
[0144] Through multi-camera collaborative recording, the technical problem of a single camera being unable to simultaneously obtain high-definition images of multiple targets at different distances is solved. Through the collaborative work of multiple cameras, a combination of panoramic coverage and local high-definition is achieved.
[0145] The system adopts a "one main and multiple auxiliary" camera configuration strategy, and simultaneously activates two types of recording devices to work together:
[0146] A zoom camera, Z, serves as the main camera, capturing the entire scene. It dynamically adjusts its focus based on the scene's needs, producing a panoramic video stream encompassing all objects. Its advantage is a wide field of view, capturing the entire scene, but its clarity for distant objects is limited.
[0147] A set of n fixed-focus cameras, F = {F1, F2, ..., Fn}, serves as auxiliary cameras. Each camera is pre-set at a specific focal length and angle, specifically responsible for capturing objects within a specific distance range. This design ensures that a dedicated camera provides high-definition images of the target at each distance level.
[0148] At time t, the system achieves strict time synchronization and synchronously obtains the main video stream frame from the zoom camera and the auxiliary video stream frame set from each fixed-focus camera. Different cameras can have different resolution configurations.
[0149] Time synchronization: All cameras use hardware synchronization signals or Network Time Protocol (NTP) to ensure that the time difference between frames is less than 1 / 60 second, avoiding target position deviation caused by time asynchrony.
[0150] Camera parameter configuration matrix: P_camera = [f1, θ1; f2, θ2; ...; f_n, θ_n]; where f_i is the focal length of the i-th fixed-focus camera (unit: mm), and θ_i is its shooting angle relative to the main camera (unit: degrees). This matrix records the geometric configuration information of the entire camera array and is used for subsequent coordinate transformation and object matching.
[0151] (3) Using advanced artificial intelligence recognition technology, identify the preset key targets in frames S1 and {C1, C2, ..., Cn}. Record the image data and precise location information of each identified key target.
[0152] Specifically, the system uses an artificial intelligence image recognition algorithm AI_Detect based on convolutional neural networks, which combines target detection and feature matching techniques to automatically find and identify preset important targets in the video streams of all cameras. The technical advantage of the algorithm is that it can handle complex situations such as target size changes, light changes, and partial occlusions.
[0153] The specific recognition process uses a parallel processing architecture:
[0154] The system uses the target recognition database matrix Φ established in step one as a reference standard and loads it into the GPU memory to improve query speed.
[0155] Each frame of the main video stream S(t) is scanned in real time, and all preset targets are identified using sliding windows and multi-scale detection techniques.
[0156] The same recognition process is performed on the frames of each auxiliary video stream, and all camera recognition tasks are executed in parallel to improve processing efficiency.
[0157] For each identified target, the system accurately records the triple (T, P, B):
[0158] T: The specific image content matrix of the target, which is obtained by cropping the target area from the original image and maintaining the original pixel values; P: The center position coordinate vector of the target in the frame, measured in pixels with the origin at the top left corner of the image; B: The bounding box coordinate vector of the target, which defines the smallest rectangular area containing the target.
[0159] To ensure recognition accuracy, the system introduces a confidence evaluation mechanism: a confidence evaluation function is used, which calculates the cosine similarity between the detected target features and the target features in the database. The closer the value is to 1, the higher the matching degree.
[0160] Only when Confidence > τ_threshold, where τ_threshold is usually set to 0.7-0.9, can the recognition be considered valid. The value can be adjusted according to the accuracy requirements of the application scenario.
[0161] Quality control of recognition results: The system also records the quality indicators of each recognition result, including target integrity (whether it is occluded), image clarity, lighting conditions, etc., to provide a basis for subsequent target selection.
[0162] (4) According to the image data and position information, accurately match the correspondence relationship of the same key targets between the zoom video frame S1 and the fixed focus video frame set {C1, C2,..., Cn}.
[0163] This step solves the key technical problem in the multi-camera system: how to accurately identify whether the targets photographed in different cameras are the same. This is a prerequisite for achieving target clarity enhancement.
[0164] Since the same target may appear in the main video stream and multiple auxiliary video streams at the same time, but due to the differences in shooting angle, distance, and lighting conditions, the appearance of the same target in different cameras will be different. The system needs to establish the correspondence between these targets and determine which are different shooting angles of the same target. The technical innovation of the matching algorithm lies in considering multiple dimensions of information comprehensively to avoid the limitations of single feature matching. The matching judgment is based on the weighted combination of two main factors:
[0165] First, calculate the appearance similarity: compare the visual feature vectors of the two target images, including color distribution, texture pattern, shape contour, etc. The advantage of cosine similarity is that it is not sensitive to image brightness changes, making it suitable for target matching under different lighting conditions.
[0166] Then calculate the position similarity: this function considers the spatial position relationship of the target in different cameras.
[0167] Finally, the appearance similarity and position similarity are weighted and summed to get the comprehensive similarity.
[0168] (5) For each matched key target, convert the image from the spatial domain to the frequency domain. Apply a specially designed high-pass filter algorithm that allows frequency information above a preset threshold to pass through to extract high-frequency components. By comparing the high-frequency component strength and quantity of different key target images, the highest clarity target F1 is evaluated and selected.
[0169] Specifically, for each matched target, the system needs to select the highest clarity version from the main video stream and each auxiliary video stream. The principle of clarity evaluation is that a clear image contains more high-frequency components, while a blurred image has fewer high-frequency components. Clarity evaluation uses frequency domain analysis technology, which has the advantages of objectivity, accuracy, and is not affected by human subjective perception.
[0170] Further, the clarity is quantitatively scored by calculating the proportion of high-frequency energy to total energy, with a higher value indicating a clearer image. The normalization of the denominator ensures the comparability between images of different sizes.
[0171] To avoid noise interference, the difference between the highest score and the second highest score can be checked. When the difference is less than a threshold, a comprehensive judgment will be made based on other quality indicators (such as contrast, saturation) to select the highest clarity image version for each important target, providing the best material for subsequent image fusion.
[0172] (6) The image data of the target F1 with the highest definition is fused to the corresponding position in the zoom video frame S1. During the fusion process, a smoothing algorithm is used to process the edge area of the target to ensure a natural transition and consistency between F1 and the background of S1.
[0173] Specifically, this invention seamlessly blends high-definition target images from different sources into the main video stream, maintaining both the target's high definition and ensuring the naturalness and coherence of the overall image. To ensure a natural and smooth fusion of the resulting image, an intelligent weighted fusion technique based on Gaussian weights is employed. The core concept of this technique is to fully utilize the high-definition image in the target's center area and gradually transition to the original image at the edges, avoiding harsh boundary effects.
[0174] (7) The fused zoom video frame S1 is processed using a streaming compression encoding method to reduce the size of the video file and optimize transmission efficiency. The compressed video stream is output, and the clarity of multiple key targets in the video stream is significantly improved while maintaining the overall visual perception quality and smooth playback experience.
[0175] This step is the final link of the system. It is necessary to use appropriate H264 / H265 video for efficient compression encoding while maintaining the enhancement effect.
[0176] The scene multi-focal length target video encoding method provided by the present invention is a video encoding method that can effectively improve the clarity of multiple targets in a multi-focal length and fixed-focal length collaborative video recording system.
[0177] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A scene multi-focal length target video encoding method, characterized in that: The following steps are involved: Step S1: determine the key targets that need to be identified, obtain the key target feature representation vector of each key target, and each key target feature representation vector forms a key target identification feature matrix F = {F1, F2, ..., F m }; Among them, F j The feature representation vector representing the key target j; j = 1, 2, ..., m, where m is the number of key targets; Step S2: according to the requirements of the shooting scene, setting a zoom recording system and n fixed-focus recording systems facing the shooting scene; wherein each of the fixed-focus recording systems is configured to have a different shooting angle and focal length position; and the shooting angle of each of the fixed-focus recording systems covers the shooting scene; Step S3, simultaneously start the zoom video system and n fixed focus video systems to perform multi-video system collaborative shooting, and simultaneously obtain zoom video frames S at the same acquisition time t t And fixed focus camera video frame set {C 1,t ,C 2,t ,...,C n,t }; Among them, C i,t represents the fixed-focus video frame captured by the fixed-focus video system i at the capture time t, i = 1, 2, ..., n; Step S4, based on the key target recognition feature matrix F={F1,F2,...,F m }, in the zoom camera video frame S t Identify m key targets in the zoom camera video frame S and obtain each identified key target j in the zoom camera video frame S t Image data and location information in ; Based on the key target recognition feature matrix F={F1,F2,...,F m }, for each fixed-focus camera video frame C i,t Perform key target recognition and identify k i,t key targets, and obtain each identified key target l in the fixed-focus camera video frame C i,t Image data and its location information in; where l=1,2,...,k i,t ; Step S5, matching the zoom camera video frame S according to the image data and its position information t and each of the fixed-focus camera video frames C i,t The same key objectives in For each focused target j, compare its i,t and in the zoom camera video frame S t The image clarity in the zoom camera video frame S t If the image clarity in the image is the highest, no image fusion processing is performed; otherwise, the key target j is obtained in each of the fixed-focus camera video frames C i,t The highest definition image data and its location information; Step S6: The acquired image data of the key target j with the highest definition is integrated into the zoom camera video frame S t The position of the same focus target in the image is obtained by merging the zoom camera video frame S t ; Step S7: for the zoom camera video frame S t For each of the m key targets in the image, steps S5 to S6 are executed, and then the zoom camera video frame S after processing is processed. t Perform video compression encoding to obtain the zoom camera video frame S after encoding with multiple focal length fusion t ; Step S8: Output the encoded zoom camera video frame S t ; Let t = t + 1; return to step S3.
2. The scene multi-focal length target video encoding method according to claim 1, characterized in that: In step S1, the feature representation vector F of the key target j j ={F j,visual ,F j,spatial ,F j,temporal }; in: F j,visual is the appearance feature vector of the key target j, which is obtained by extracting the color histogram, texture features, and shape description features of the key target j; F j,spatial is the position feature vector of the key target j, which is used to describe the position distribution and movement trajectory pattern of the key target j in the picture; F j,temporal is the temporal feature vector of key target j, which is used to describe the temporal pattern and duration statistics of key target j.
3. The scene multi-focal length target video encoding method according to claim 1, characterized in that: The zoom video recording system is used to capture the entire scene during filming, dynamically adjust the focal length according to the scene requirements, and obtain a panoramic video stream containing all key targets.
4. The scene multi-focal length target video encoding method according to claim 1, characterized in that: Each fixed-focus video system i is configured with a shooting angle θ i With a specific fixed focal length f i , when shooting a scene, the shooting angle θ i With a specific fixed focal length f i remains unchanged; where the shooting angle θ i is the shooting angle relative to the zoom recording system.
5. The scene multi-focal length target video encoding method according to claim 1, characterized in that: Focus target j in the zoom camera video frame S t The image data and its position information in E are expressed as: s,j,t ={T s,j,t ,P s,j,t ,B s,j,t }; where T s,j,t Represents the focus target j in the zoom camera video frame S at the acquisition time t t The image content matrix is obtained by taking the zoom camera video frame S t The image of the key target j is obtained by cropping; s,j,t represents the center of the focused target j in the zoom camera video frame S at the acquisition time t t Position coordinate vector of B s,j,t Represents the focus target j in the zoom camera video frame S at the acquisition time t t The bounding box coordinate vector of the minimum bounding rectangle area; Each identified key target l is located in the fixed-focus camera video frame C i,t The image data and its position information in E are expressed as: i,l,t ={T i,l,t ,P i,l,t ,B i,l,t }; where T i,l,t Represents the key target l in the fixed-focus camera video frame C at the acquisition time t i,t The image content matrix is obtained by taking the fixed-focus camera video frame C i,t The image of the key target l is obtained by cropping; i,l,t Represents the center of the key target l in the fixed-focus camera video frame C at the acquisition time t i,t Position coordinate vector of B i,l,t Represents the key target l in the fixed-focus camera video frame C at the acquisition time t i,t The bounding box coordinate vector of the minimum bounding rectangular area.
6. The scene multi-focal length target video encoding method according to claim 5, characterized in that: Step S4 further includes: For the zoom camera video frame S t Each key target j identified in the image, and in each of the fixed-focus camera video frames C i,t For each key target l identified in step S1, the confidence evaluation function is used to calculate the cosine similarity Confidence with the feature representation vector of the corresponding key target in step S1. The recognition is considered valid only when the cosine similarity Confidence is greater than the threshold τ_threshold; among them, τ_threshold is set to 0.7-0.
9.
7. The scene multi-focal length target video encoding method according to claim 6, characterized in that: Step S5, matching the zoom camera video frame S according to the image data and its position information t and each of the fixed-focus camera video frames C i,t The same key objectives as in the 2015 / 2016 / Revised Implementation Plan, specifically: Step S5.1: focus target j on zoom camera video frame S according to acquisition time t t The image content matrix T s,j,t , and the focus target l in the fixed-focus camera video frame C at the acquisition time t i,t The image content matrix T i,l,t , calculate the appearance similarity Simvisual(j,l) between the key target j and the key target l; Step S5.2: focus target j in zoom camera video frame S based on acquisition time t t The position coordinate vector P s,j,t , with the center of the focus target l in the fixed-focus camera video frame C at acquisition time t i,t The position coordinate vector P i,l,t , calculate the position similarity Simspatial(j,l) between the key target j and the key target l; simspatial(j,l)=exp(-||P s,j,t -transform(P i,l,t ,M_i)|| 2 / 2σ 2 ) Where: transform(P i,l,t ,M_i) represents the P in the i coordinate system of the fixed focus video system i,l,t Transformation function to the zoom video system coordinate system; Mi is the 3×3 unit transformation matrix, obtained by calibrating the camera of the fixed-focus video system i; σ is the position tolerance parameter, set to the fixed-focus video frame C i,t 10%-20% of the image size of the cropped focus object l; In step S5.3, the comprehensive similarity between key target j and key target l is calculated using the following formula: Sim(j,l)=α·Simvisual(j,l)+β·Simspatial(j,l) Where: α and β are the appearance similarity weight and position similarity weight respectively; α+β=1,α,β>0; for scenes with large changes in target appearance, increase the value of β; for scenes with relatively stable target appearance, increase the value of α; In step S5.4, if the comprehensive similarity Sim(j,l) between the key target j and the key target l is greater than the matching threshold θ_match, it means that the key target j and the key target l are successfully matched and are the same key target.
8. The scene multi-focal length target video encoding method according to claim 7, characterized in that: For each focused target j, compare its i,t and in the zoom camera video frame S t The image clarity in , specifically: For each video frame, the following method is used to obtain its clarity quantification score. Based on the clarity quantification score, the image clarity is compared: Step A: The video frame that needs to be quantified for clarity is called the target image, which is expressed as: target image T, with a width of M pixels and a height of N pixels. The target image T is converted from the spatial domain to the frequency domain using the following formula: in: T(x,y) represents the pixel value at position (x,y) in the spatial domain; j in the formula is the imaginary unit; F(u,v) represents the complex coefficient in the frequency domain; the position (x,y) in the spatial domain corresponds to the position (u,v) in the frequency domain; Step B: For the complex coefficients F(u,v) in the frequency domain, perform the following filtering operation to retain the high-frequency components and filter out the low-frequency components to obtain the filtered image H(u,v): if Then let H(u,v)=1; otherwise, let H(u,v)=0; Where: ω_cutoff is the cutoff frequency threshold; Step C: Use the following formula to perform element-by-element multiplication of the filtered image H(u,v) and the complex coefficient F(u,v) in the frequency domain to implement frequency domain filtering, extract the high-frequency information in the image, and obtain the high-frequency information representation F_high(u,v): F_high(u,v)=F(u,v)⊙H(u,v) Among them: ⊙ represents element-by-element product; Step D: Use the following formula to obtain the sharpness quantitative score Sharpness_Score(T): Where: U and V are the lengths of rows and columns in the frequency domain respectively.
9. The scene multi-focal length target video encoding method according to claim 6, characterized in that: Step S6: The acquired image data of the key target j with the highest definition is integrated into the zoom camera video frame S t The position of the same focus target in the image is obtained by merging the zoom camera video frame S t , specifically: Step S6.1: The acquired image data of the key target j with the highest definition is expressed as Make image data The center coordinates and zoom camera video frame S t The center coordinates of the same focus objects in the image are coincident. If the sizes of the two do not match, bilinear interpolation is used to adjust the size so that the image data Replace with zoom camera video frame S t within the region with the same key objectives; Step S6.2, embedding the zoom camera video frame S t Image data of the same key target Perform color and brightness correction to match the zoom camera video frame S t The color and brightness distribution in the image are the same, and the first image data is obtained. Step S6.3, the first image data For each pixel point p in , the following formula is used to calculate its weight w(p): Where: σ is the parameter that controls the transition smoothness, set to the first image data 1 / 6 to 1 / 4 of the size; Step S6.4, using the following formula, the first image data The pixels p in the image are fused to obtain the fused image I_fused: in: Represents the first image data The pixel value of the pixel point p in the middle; S_region(p) represents the zoom camera video frame S t I_fused(p) represents the pixel value of the pixel point p in the same key target area in the fused image. In step S6.5, the fused image I_fused is subjected to boundary smoothing using the following formula to obtain a boundary smoothed image I_fused": I_fused"=I_fused*G_smooth Where: G_smooth is the smoothing filter; G_smooth = (1 / 16) [1, 2, 1; 2, 4, 2; 1, 2, 1]; Step S6.6, re-embedding the image I_fused" after boundary smoothing into the zoom camera video frame S t The target area in the image is obtained to obtain the final fused zoom camera video frame S_enhanced(t); S_enhanced(t)=S t +Mask⊙(I_fused"-S_region) Among them: Mask is the zoom camera video frame S t The target area mask matrix where the same key target is located; S_region is the zoom camera video frame S t The target area where the same key target is located; S t Represents a zoom camera video frame; Step S6.7: perform video compression encoding on the final fused zoom camera video frame S_enhanced(t) and output it.
Citation Information
Patent Citations
Method for cooperative acquisition of multi-target videos under different resolutions by variable-focus array camera
CN101720027A
Multi-focal length photo fusion method and system and photographing device
CN102542545A
Multi-layer focusing fusion method for microscope image
CN107481213A
Monitoring method and device
CN110830756A
Image processing method and device, electronic equipment and storage medium
CN113936154A
Cited By
Automatic focusing method and three-proofing mobile phone
CN120980197A