Method and System for Generating an Overhead View of a Game Situation Based on Multi-Camera Information Fusion

Through multi-camera information fusion technology, we can identify and match the rectangular vertex coordinates of the stadium and generate the optimal dimensional transformation matrix, which solves the problem that a single camera is difficult to generate a global top view, and achieves real-time global top view generation and improvement of target positioning accuracy for sports events.

CN119963966BActive Publication Date: 2025-07-01NANJING UNIV OF POSTS & TELECOMM
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510429066.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-07-01
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

It is difficult for the prior art to generate a global top view of sports events in real time through a single camera, especially when tracking high-speed motion targets and analyzing occlusion scenes, data completeness and reliability are insufficient.

Method used

Using the method of multi-camera information fusion, the HSV color space separation and linear intersection recognition of the current frame image of each camera is used to determine the coordinates of the four vertices of the rectangle, and the optimal dimensional transformation matrix is ​​generated through the rectangle matching and error analysis algorithm to realize the global target positioning of the stadium and the real-time top view generation.

Benefits of technology

Real-time global top view generation of sports events is achieved, the accuracy of target positioning and data reliability are improved, and the dynamic evolution of the situation in the field can be better understood.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119963966B_ABST
    Figure CN119963966B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for generating an overhead view of a game situation based on multi-camera information fusion. Cameras that can cover the entire field and form overlapping areas are deployed at the four corners of the stadium. The HSV color space is used to separate and extract straight lines and intersection points in the images, calculate and construct a set of characteristic points of the field rectangle, map the local perspective to a unified overhead coordinate system based on geometric features and coordinate transformation matrices, screen out the coordinates of the optimally matched rectangle through rectangle matching scoring, calculate the coordinate differences of athletes in the overlapping area when transformed to the overhead view from different perspectives, assign weights to the rectangle matching error and the coordinate error of the athletes through the CRITIC weight method, solve the objective function, and determine the optimal dimension transformation matrix. The present invention can realize the real-time generation of an overhead view in sports events.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision and image processing, and specifically relates to a method and system for generating a bird's-eye view of a game situation based on multi-camera information fusion. Background Art

[0002] When analyzing sports events tactically, it is usually more intuitive to look down at the game. However, most current cameras are set up around the venue, making it difficult to shoot from a bird's-eye view. Therefore, image processing technology is needed to convert images from other perspectives into a bird's-eye view.

[0003] At present, in the field of computer vision and intelligent perception, the detection and recognition of scene targets and the generation of bird's-eye views are mainly achieved by using the target detection algorithm under monocular vision with BEV projection transformation technology, or by using the perspective transformation method based on OpenCV. The technologies applied include multi-object recognition and positioning target detection based on deep learning, feature extraction and classification methods based on traditional machine learning, target detection based on instance segmentation, and multi-scale target detection for objects of different sizes. Based on the above technologies, it is currently possible to identify multiple objects in the video and give the category identification and screen position data of each object. For example, a method for real-time generation of a bird's-eye view of a sports event with a dynamic perspective is disclosed in the publication number CN118154728A. This method uses a single camera to perform real-time image capture and coordinate initialization, image preprocessing and key point extraction, initial bird's-eye view generation and region division, rectangle screening and error optimization, rematching recognition and real-time bird's-eye view generation.

[0004] However, due to factors such as camera viewing angle, shooting distance and clarity, it is difficult to conduct in-depth data analysis based solely on the position data of a single-frame image from a single camera. When conducting a time series analysis of an event, a large amount of manpower is often required to manually acquire data. Taking sports events as an example, it is difficult to fully grasp the overall situation of the stadium by processing the single-frame image recognition of a single camera with a fixed position. For example, a real-time generation method for a bird's-eye view of a sports event with a dynamic perspective is disclosed in publication number CN118154728A. It only uses one camera, has a limited viewing angle, and lacks accuracy when covering the entire standard stadium. It is also limited by optical distortion and resolution, and it is difficult to meet the requirements of real-time tactical analysis for target positioning accuracy and panoramic situation awareness. Especially under complex working conditions such as high-speed moving target tracking and occlusion scene analysis, the data completeness and reliability of the single-perspective system face severe challenges. Summary of the invention

[0005] The present invention aims to solve one of the technical problems existing in the related art at least to a certain extent.

[0006] An object of the present invention is to provide a method for generating a top view of a game situation based on multi-camera information fusion. This method realizes the global target positioning in a complete stadium by fusing the local information of multiple cameras, and converts the image dimension frame by frame to generate a complete game situation in real time, so as to assist coaches and audiences in better grasping the state of sports events.

[0007] Another object of the present invention is to provide a system for generating a top view of a game situation based on multi-camera information fusion for implementing the above method.

[0008] To achieve the above object, on the one hand, the present invention provides a method for generating a top view of a game situation based on multi-camera information fusion, including:

[0009] S100. Cameras are respectively placed at the four corners of the sports field to simultaneously record the real-time shooting of the stadium, and the current frame image in the video is intercepted; taking one of the corners of the stadium as the coordinate origin, a plane rectangular coordinate system is established, and the relative positions of the four cameras, as well as the height and width of the stadium, are recorded.

[0010] S200. Separate the current frame image of each camera in the HSV color space, set the value range to identify all the straight lines and intersections on the image, arbitrarily take a corner point on the image as a reference point, calculate and determine the intersection point closest to the reference point, and then find other intersections related to it by analyzing the straight line where the intersection point is located. Through these intersections, the coordinates of the four vertices of the rectangle are gradually determined and stored in the point set in sequence. ; At moment, create a legend of the full-field top view, and store all the intersections in the top view into the point set after transformation by the transformation matrix. In, divide the field into rectangles by connecting the intersections, record the coordinates of all rectangles in the order from left to right and from small to large, and store them as the rectangle point set obtained based on the top view. ; S300. According to the four vertices of the rectangle on a certain frame image obtained in step S200, find and screen the corresponding matrix coordinate points in the top view, restore all the coordinates in this frame image to the top view by the method of coordinate restoration, and then match the target rectangle and the selected rectangle through the rectangle matching and error analysis algorithms to accurately screen the rectangles; S400. Introduce a method for calculating the athlete coordinate transformation error based on multi-view camera fusion, obtain the athlete coordinates in the overlapping area between cameras, and use matrix transformation to the top view perspective to calculate the Euclidean distance of the same athlete in the top view to obtain the athlete coordinate transformation error; S500. Superimpose the errors of the four-dimensional transformation matrices obtained from the non-overlapping areas of the four cameras to form an error point set. ; Use the athlete coordinate transformation error in the overlapping area of the four cameras to construct an error point set. ; Set coefficients for two sets of error points respectively, and take the minimum weighted sum as the objective function to obtain the optimal dimension transformation matrix; S600. After the current time sequence ends, switch to the next time sequence, and perform dimension transformation on the entire sports field according to the optimal dimension transformation matrix obtained in step 500. Save all the points of the required field condition information after transformation in the point set In it, generate a top view through the information stored in this point set, clear other point sets, intercept the frame image of the next time sequence in the video, and repeat steps S100~S500 to continuously obtain the top view. A further preferred technical solution of the present invention is that in step S100, a plane rectangular coordinate system is established with the lower left corner of the stadium as the coordinate origin, and the relative positions of the four cameras are recorded as , , , respectively, the height of the stadium is and the width is .

[0011] Preferably, in step S200, the current frame image of each camera is separated in the HSV color space, and a value range is set to identify all the straight lines and intersections on the image. Arbitrarily take a corner point on the image as a reference point, calculate and determine the intersection point closest to the reference point, and then find other intersections related to it by analyzing the straight line where the intersection point is located. Through these intersection points, the four vertex coordinates of the rectangle are gradually determined and stored in the point set in sequence; The specific method is as follows: S210. Convert the current frame image from the RGB three-color channel to the HSV color space, and establish a plane rectangular coordinate system with the lower left corner of the current frame image as the origin, set a suitable value range, and identify all the straight line equations , calculate all the intersection points of the straight lines and record them;

[0012] S220. Arbitrarily take a boundary point of one of the four vertices of the current frame image as the reference point. Assume the length of the image is , and the height is . The intersection point with the minimum distance from the point is denoted as . The point is generated by the intersection of two straight lines. Substitute the point coordinates to obtain the analytical formulas of the two straight lines passing through the point, denoted as and ;

[0013] S230. Combine with the identified straight line equations except to obtain the intersection point as , , … , similarly, With The identified straight lines outside are combined to obtain the intersection point , , … ;calculate and and Distance of points and , respectively select and The corresponding points are denoted by and , the other straight line where the two points lie is and ;

[0014] S240, straight line With straight line The intersection point is obtained by combining the analytical expressions of , and then the four points obtained As vertex coordinates are stored in order in the point set middle;

[0015] S250, repeat steps S210 to S240, traverse according to the difference in the distance between the two points, find the point set of the rectangle on the original image, until no new distance difference appears, then all the rectangular point sets on the original image are found, and stored in the point set in sequence As a preferred embodiment, in step S200, Always create a top view legend for the entire scene, and store all intersection points in the top view into a point set after transforming them through the transformation matrix In the example, the site is divided into Rectangles, record the coordinates of all rectangles in order from left to right and from small to large, and store them as a rectangular point set based on the top view ; The specific method is:

[0016] S260, when the timing is When the image of the entire playing field or the dimensions of the field is obtained, a top-view coordinate system with the vertices of the playing field as the coordinate origin is created. , get the coordinates of all intersection points on the site , according to the coordinates of the four vertices of the site, the coordinates on the matching coordinate system Calculate a three-row and three-column transformation matrix, and multiply all points on the field by this transformation matrix to get the new intersection coordinates , so as to obtain and record the top view F of the site, and the time sequence After that, directly obtain the top view F in the initial state;

[0017] S270. Among the obtained in step S260, take any two points with different horizontal and vertical coordinates, and then the four points of a rectangle can be determined, and they are stored in the array in clockwise order . Connect all the points to generate n horizontal lines and m vertical lines, dividing the entire top view site into several small rectangles and large rectangles composed of small rectangles. Thus, the site is divided into rectangles . Record the coordinates of all rectangles in order from left to right and from small to large, and store them in an eight-row column array to obtain the rectangle point set .

[0018] Preferably, in step S300, according to the four vertices of the rectangle on a certain frame of image obtained in step S200, find and screen the corresponding matrix coordinate points in the top view, restore all the coordinates in this frame of image to the top view by means of coordinate restoration, and then, through the rectangle matching and error analysis algorithms, match the target rectangle and the selected rectangle to accurately screen the rectangle, including:

[0019] S310. Based on the four vertices of the rectangle on a certain frame of image obtained in step S200, introduce a rectangle matching scoring algorithm model to initially identify and screen out suitable rectangles;

[0020] S320. Adopt a loop method, that is, when selecting the target rectangle, perform a shift process on the point set of the obtained coordinates, and then repeat step S310 to obtain all the optimal solutions;

[0021] S330. After obtaining the most suitable rectangle pair, calculate a dimension transformation matrix.

[0022] Preferably, in step S310, based on the four vertices of the rectangle on a certain frame of image obtained in step S200, introduce a rectangle matching scoring algorithm model to initially identify and screen out suitable rectangles; the specific method is:

[0023] S311. Based on the point set obtained in step S200 and Obtain a new top view F1. First, compare the boundary lengths of the top views F and F1. If the boundary length of a certain side of F1 is greater than the corresponding boundary length in F, it indicates that the currently selected rectangular point is deviated or inappropriate. Then, compare the rectangles in the figure. If the length of a certain side of the rectangle in F1 is significantly greater than the length of the corresponding side of the rectangle in F, it can be determined that the current target rectangle does not match the size and shape of the selected rectangle. Remove this part of the rectangle and store the remaining rectangle coordinate points in a new point set. in;

[0024] S312. Based on the point set , calculate the length Length, width Width of each group of rectangles under two perspectives, and the included angle between two sides of the rectangle. At the same time, obtain the centroid coordinates of each group of rectangles. Define the width-to-height difference ratio as

[0025] , the angle difference as , and the centroid distance as . Calculate the Hausdorff distance of the point sets of the rectangle coordinates under two perspectives, expressed as:

[0026] According to the top view creation method in step S200, obtain the restored top view F2 under the perspective of this camera, and thus obtain two straight lines representing the perspective in the top view F2. Extend these two straight lines in the reverse direction to intersect to obtain the coordinates of a virtual camera . By comparing with the camera coordinates in step S100, obtain a Euclidean distance difference ;

[0027] Then, normalize the data and calculate the normalized similarity score :

[0028]

[0029] where = 0.12, = 0.15, = 0.13, = 0.10, = 0.5. Traverse and all the rectangle coordinate groups in, calculate the similarity scores of all matching combinations , and store them in the point sets , respectively. Compare the scores to obtain the matching rectangles and update them in one-to-one correspondence and 。

[0030] S320. It should be noted that when comparing the coordinates of the target rectangle and the selected rectangle, it cannot be guaranteed that the order of the coordinate points in their corresponding point sets is the same. Therefore, a loop method is adopted, that is, the point set of the coordinates obtained when selecting the target rectangle is subjected to a shift process, and then step S310 is repeated. Since a rectangle can select four coordinate points, only four loops are required to obtain all the optimal solutions; S330. After obtaining the most suitable rectangle pair, by comparing the point sets and in the coordinates, a dimensional transformation matrix can be calculated.

[0031] S400. Transform the athletes captured by the camera into the stadium under the top-down view, and calculate the coordinate differences of the athletes after different perspective transformations as the error basis of the dimensional transformation matrix. The specific method is as follows:

[0032] S410. Select two perspectives of cameras A and B, use YOLO to identify all athletes in the camera perspectives and take the center point coordinates (X, Y) as the position where the athlete is located;

[0033] S420. Calculate the coordinates of the athletes in the overlapping area after dimensional transformation in each perspective and store them in the point set. The specific implementation steps are as follows:

[0034] S421. According to the dimensional change matrix calculated in step S300, perform dimensional transformation on the positions of the athletes in each camera perspective to obtain the coordinates after dimensional transformation (X', Y'), and store the coordinates of all athletes captured by the i-th camera after dimensional change in the point set Si;

[0035] S422. Calculate the coordinate range of the overlapping area between the cameras, and filter out the coordinates that meet the range in the point set and store them in the new point set S'i;

[0036] S430. Take the point set S'1 and the point set S'2, calculate the Euclidean distance between all point pairs in the two perspectives as the matching cost, and construct a cost matrix to apply the Hungarian algorithm to match the point pairs; the specific implementation steps are as follows:

[0037] S431. Define the detection point set of camera A as , and the detection point set of camera B as . Calculate the Euclidean distance between all point pairs in the two perspectives as the matching cost, and construct a cost matrix as , and its elements are S432. Solve the row index and column index sets , such that , and obtain an ordered matching index list 。

[0038] S440. For each matched row-column index (i, j), the point corresponding to camera A , and the point of camera B , reorder the point sets according to the matching results to ensure that the coordinates at the corresponding positions belong to the same athlete. For each pair of matching points ( , ), calculate the coordinate error ;

[0039] Step 450. Repeat the process of steps S410 to S440 for views B and C, views C and D, and views D and A, and superimpose the error values of the overlapping area parts to obtain the final error 。

[0040] S500. Introduce a method for determining the optimal dimensional transformation matrix based on the minimum total error: In the sports event venue, there are overlapping parts in the shooting ranges of the cameras at the four corners of the venue, forming a globally covered monitoring area; there are 4 dimensional transformation matrices obtained from the non-overlapping areas of 4 cameras respectively. Define the error superposition of the rectangles of these 4 dimensional transformation matrices to form an error point set and the set of coordinate transformation errors of athletes in the overlapping areas of each camera . Assign appropriate coefficients to these two errors respectively, and use the minimum weighted sum as the objective function to obtain the optimal dimensional transformation matrix. The specific method of step S500 is as follows:

[0041] S510. Obtain four dimensional transformation matrices from the rectangles preliminarily screened by a single camera, and superimpose the errors of the rectangles of the four dimensional transformation matrices to form an error point set , , represents the error point set of the dimensional transformation matrix from the th camera; Denote the sum of the error values of the coordinate transformation of the athletes in the common area photographed by the four cameras as ; After unifying the scale and standardizing the two error sets and , assign appropriate weights and to and respectively through the CRITIC weight method, and solve the total minimum error through the weighted sum to determine the optimal dimensional transformation matrix; The specific method is as follows: S511. Since the sources and scales of the two error sets are different, denote , , set the target length as , where , The original element numbers of and are used to adjust the error point set and the error sum to the same length . For subsequent alignment of quantiles, the length of the shorter set is extended using linear interpolation. Let the shorter set be extended to obtain ; The Min - Max normalization method is applied to normalize and respectively, and the calculation is as follows:

[0042]

[0043]

[0044] to obtain and ;

[0045] S512. Appropriate weights are assigned to the error point set and the error sum respectively. Using the CRITIC weight method, the normalized and are defined as two indicators, and a matrix is constructed:

[0046]

[0047] Then the standard deviation is calculated, The standard deviation of

[0048]

[0049] The standard deviation of

[0050]

[0051] Calculate the Pearson correlation coefficient between the two columns:

[0052]

[0053] According to the obtained Pearson correlation coefficient, calculate the information content of and respectively. The information content of is: , The information content of is: , ; Thus, the weights

[0054] S520. Calculate the total minimum error and construct an objective function, expressed as:

[0055]

[0056] Satisfy the constraint conditions:

[0057]

[0058] Then repeat step S511 - step S512, traverse all rectangles for which the dimensional transformation matrix is obtained and the coordinates of the athletes after dimensional transformation, and the dimensional transformation matrix with the minimum total error obtained is the optimal dimensional transformation matrix.

[0059] On the other hand, the present invention provides a system for generating a top - view of the game situation based on multi - camera information fusion for implementing the above - mentioned method, including a module for acquiring current - frame information of the video, a module for rectangle matching and dimensional transformation, a module for calculating coordinate transformation error, a module for weighted evaluation of heterogeneous errors, and a module for fusing local information to restore the top - view perspective. Among them:

[0060] The module for acquiring current - frame information of the video is used to perform HSV color - space separation processing on the image of the current frame in the video, obtain all information points on the current field captured by a certain camera, and store them in a point set;

[0061] The module for rectangle matching and dimensional transformation is used to extend and intersect the straight lines of the top - view, record the rectangle coordinates, obtain the matching rectangle pairs according to the similarity, and use the rectangle coordinates of the original perspective and the top - view perspective of the same camera to obtain the error of these rectangle coordinates and generate a dimensional transformation matrix;

[0062] The module for calculating coordinate transformation error is used to transform the coordinates of the athletes in the overlapping area to the top - view perspective, calculate the Euclidean distance of the same athlete in the top - view, and obtain the athlete coordinate transformation error;

[0063] The module for weighted evaluation of heterogeneous errors is used to assign corresponding weight coefficients to different types of errors according to their respective importance levels and influence ranges, and comprehensively evaluate the overall error situation through weighted error summation;

[0064] The module for fusing local information to restore the top - view perspective is used to fuse the local information of the cameras at the four corners of the field of the current frame according to the overlapping area captured by the camera, obtain the global target of the stadium, and thus perform dimensional transformation on the full - field information of the stadium in the current frame to obtain the corresponding full - field top - view.

[0065] On the other hand, the present invention provides a non - transient computer - readable storage medium, on which computer instructions are stored, and these computer instructions cause the computer to execute the above - mentioned method for generating a top - view of the game situation based on multi - camera information fusion.

[0066] In another aspect, the present invention provides an electronic device, including: a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus. The processor calls the logical instructions in the memory to execute the above method for generating a top view of the game situation based on multi-camera information fusion.

[0067] In yet another aspect, the present invention provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer executes the above method for generating a top view of the game situation based on multi-camera information fusion.

[0068] Beneficial effects: The method of the present invention for generating a top view of the game situation based on multi-camera information fusion performs dimensional transformation after fusing each frame of the images captured by multiple cameras, thereby generating a top view of the full-field information in real time. Ultimately, it can realize the generation of a real-time top view of a sports event, helping users better grasp the information on the field and helping users deeply understand the dynamic evolution of the game situation on the field.

[0069] The method of the present invention for generating a top view of the game situation based on multi-camera information fusion relies on the original video scene features obtained through comprehensive shooting and conducts various analysis tasks by means of multiple computational analysis methods, effectively eliminating the influence of external interference factors. By converting the original video covering the entire stadium into a top-view image, the system overcomes the problems of field condition observation caused by occlusion, bad viewing angles, incomplete pictures, etc., and provides more convenient and comprehensive support for users during the real-time viewing and field condition analysis of sports events.

[0070] The present invention performs dimensional transformation by identifying and screening rectangles, fuses the field condition information from different perspectives according to the overlapping parts, and constructs a top view of the game venue, enabling a more reasonable and accurate grasp of the game situation of sports events in real time. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] Figure 1 It is the overall flowchart of the method for generating a top view of the game situation based on multi-camera information fusion. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0072] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of them, and they should not be construed as limiting the present invention. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention. In the description of the present invention, it should be understood that the terms used are only for the purpose of description and cannot be construed as indicating or implying relative importance.

[0073] The following will describe Figure 1 a method and system for generating a top view of a game situation based on multi-camera information fusion provided by the present invention.

[0074] Embodiment 1: This embodiment provides a method for generating a top view of a game situation based on multi-camera information fusion, as Figure 1 shown, including the following steps:

[0075] S100. Place a camera at each of the four corners of the sports field. The cameras at the four corners of the field simultaneously perform real-time shooting and recording of the stadium, intercept the current frame image in the video, record the current time sequence as , and then establish a plane rectangular coordinate system with the lower left corner of the stadium as the coordinate origin, and simultaneously record the relative positions of the four cameras. Denote the height and width of the stadium as and .

[0076] S200. Perform HSV color separation on the current frame image of each camera, set the value range of the required information points or lines, extract the required information on the current frame image and record it. Arbitrarily select a corner point on the image as the reference point, and by continuously determining straight lines and calculating the closest points, four points forming a rectangle are obtained and placed in the point set in the order; rectangles at different distances are obtained respectively according to the above operations, and the four points are stored in the point set in the form of a two-dimensional array; create a legend for the full-field top view at the time sequence of . According to all the obtained boundary lines on the field, extend the boundary lines to divide the field into N rectangles, record the coordinates of all rectangles in order, and store the point set of the rectangles; after the four cameras respectively perform step S200, corresponding rectangle point sets and will be obtained:

[0077] S210, convert the current frame image from the RGB three-color channel to the HSV color space, take the lower left corner of the current frame image as the origin, and establish Plane rectangular coordinate system, set the appropriate value range, and identify all straight line equations , calculate the intersection points of all straight lines and record them;

[0078] S220, randomly select one of the four vertices of the current frame image as the reference point, and assume that the length of the image is , the height is , will be with the point The distance is the minimum The intersection of ,point It is produced by the intersection of two straight lines. Point coordinates, get the The analytical expression of the two straight lines of the point is recorded as and ;

[0079] S230, With The intersection point is obtained by combining the equations of the identified lines outside , , … , similarly, With The identified straight lines outside are combined to obtain the intersection point , , … ;calculate and and Distance of points and , respectively select and The corresponding points are denoted by and , the other straight line where the two points lie is and ;

[0080] S240, straight line With straight line The intersection point is obtained by combining the analytical expressions of , and then the four points obtained As vertex coordinates are stored in order in the point set middle;

[0081] S250. Repeat steps S210 - S240, traverse according to the different calculated distances between two points, find the point set of the rectangles on the original image, until no new distance differences appear, then all the rectangle point sets on the original image are found and stored in the point set in sequence. in different rows;

[0082] S260. At the time sequence of , obtain a picture containing the entire competition venue or the size of the venue, and thus create a top - view coordinate system with the vertex of the competition venue as the coordinate origin , obtain the coordinates of all intersection points on the venue , match the coordinates on the coordinate system corresponding to the coordinates of the four vertices of the venue Calculate a transformation matrix of three rows and three columns, multiply all points on the field by this transformation matrix to obtain new intersection coordinates , thus obtain and record the top - view F of the venue. After the time sequence , directly obtain the top - view F in the initial state;

[0083] S270. Arbitrarily select two points with different horizontal and vertical coordinates in the obtained in step S260, then the four points of a rectangle can be determined, and stored in the array in clockwise order. Connect all the points to generate n horizontal lines and m vertical lines, divide the entire top - view venue into several small rectangles and large rectangles composed of small rectangles. Thus, the venue is divided into rectangles , record the coordinates of all rectangles in order from left to right and from small to large, and store them using an eight - row eight - column array to obtain the rectangle point set .

[0084] S300. Propose a rectangle screening and scoring algorithm: According to the four vertices of the rectangle on a certain frame of image obtained in step S200, find and screen the corresponding matrix coordinate points in the top - view, so as to restore all the coordinates in this frame of image to the top - view through the method of restoring coordinates. Then, through the rectangle matching and error analysis algorithms, match the target rectangle and the selected rectangle to further accurately screen the rectangles. The specific method is as follows:

[0085] S310. Based on the four vertices of the rectangle on a certain frame of image obtained in step S200, introduce a rectangle matching and scoring algorithm model to initially identify and screen suitable rectangles, including:

[0086] S311. Based on the point sets and Obtain a new top view F1. First, compare the boundary lengths of the top view F and F1. If the boundary length of a certain side of F1 is greater than the corresponding boundary length in F, it indicates that the currently selected rectangular point is deviated or inappropriate. Then, compare the rectangles in the figure. If the length of a certain side of the rectangle in F1 is significantly greater than the length of the corresponding side of the rectangle in F, it can be determined that the current target rectangle does not match the size and shape of the selected rectangle. Remove this part of the rectangles and store the remaining rectangle coordinate points in a new point set. in;

[0087] S312. Based on the point set , calculate the length Length, width Width of each group of rectangles and the included angle of the two sides of the rectangle in two perspectives. At the same time, obtain the centroid coordinates of each group of rectangles. Define the width-to-height difference ratio as

[0088] , the angle difference as , and the centroid distance as . Calculate the Hausdorff distance of the point sets of the rectangle coordinates in two perspectives, expressed as:

[0089]

[0090] According to the top view creation method in step S200, obtain the restored top view F2 from the perspective of this camera, and thus obtain two straight lines representing the perspectives in the top view F2. Extend these two straight lines in the reverse direction to intersect to obtain the coordinates of a virtual camera . By comparing with the camera coordinates in step S100, obtain a Euclidean distance difference ;

[0091] Then, normalize the data and calculate the normalized similarity score :

[0092]

[0093] where = 0.12, = 0.15, = 0.13, = 0.10, = 0.5. Traverse and all the rectangle coordinate groups in, calculate the similarity scores of all matching combinations , and store them in the point sets , Among them, compare the scores to obtain the matching rectangles, and update them in a one-to-one correspondence order. and .

[0094] S320. It should be noted that when comparing the coordinates of the target rectangle and the selected rectangle, it cannot be guaranteed that the order of the coordinate points appearing in their corresponding point sets is the same. Therefore, a loop method is adopted, that is, the point set of the coordinates obtained when selecting the target rectangle is subjected to a shift process, and then step S310 is repeated. Since a rectangle can select four coordinate points, only four loops are required to obtain all the optimal solutions;

[0095] S330. After obtaining the most suitable rectangle pair, by comparing the point sets and in the coordinates, a dimensional transformation matrix can be calculated.

[0096] S400. Transform the athletes captured by the camera into the stadium under the top-down view, and calculate the coordinate differences of the athletes after different perspective transformations as the error basis of the dimensional transformation matrix. The specific method is as follows:

[0097] S410. Select two perspectives of cameras A and B, use YOLO to identify all athletes in the camera perspectives and take the center point coordinates (X, Y) as the position where the athlete is located;

[0098] S420. Calculate the coordinates of the athletes in the overlapping area after dimensional transformation in each perspective and store them in the point set. The specific implementation steps are as follows:

[0099] S421. According to the dimensional change matrix calculated in step S300, perform dimensional transformation on the positions of the athletes in each camera perspective to obtain the coordinates after dimensional transformation (X’, Y’), and store the coordinates of all athletes captured by the i-th camera after dimensional change in the point set Si;

[0100] S422. Calculate the coordinate range of the overlapping area between the cameras, and screen out the coordinates that meet the range in the point set and store them in the new point set S’i;

[0101] S430. Take the point set S’1 and the point set S’2, calculate the Euclidean distance between all point pairs in the two perspectives as the matching cost, and construct a cost matrix to apply the Hungarian algorithm to match the point pairs; the specific implementation steps are as follows:

[0102] S431. Define the detection point set of camera A as , and the detection point set of camera B as , calculate the Euclidean distance between all point pairs in the two perspectives as the matching cost, and construct the cost matrix as , whose elements are ;

[0103] S432. Solve the set of row indices and column indices such that , and obtain an ordered list of matching indices .

[0104] S440. For each matched row-column index (i, j), the point corresponding to camera A , and the point of camera B , reorder the point set according to the matching result to ensure that the coordinates at the corresponding positions belong to the same athlete. For each pair of matched points ( , ), calculate the coordinate error ;

[0105] Step 450. Repeat the process of steps S410 to S440 for perspectives B and C, perspectives C and D, and perspectives D and A, and superimpose the error values of the overlapping area parts to obtain the final error .

[0106] S500. Introduce a method for determining the optimal dimensional transformation matrix based on the minimum total error: In a sports event venue, there are overlapping parts in the shooting ranges of the cameras at the four corners of the venue, forming a globally covered monitoring area; there are 4 dimensional transformation matrices obtained from the non-overlapping areas of the 4 cameras respectively, and the error superposition of the rectangles defining these 4 dimensional transformation matrices forms an error point set and the set of coordinate transformation errors of the athletes in the overlapping areas of each camera . Appropriate coefficients are assigned to these two errors respectively, and the weighted sum is minimized as the objective function to obtain the optimal dimensional transformation matrix. The specific method of step S500 is as follows:

[0107] S510. Obtain four dimensional transformation matrices from the rectangles preliminarily screened by a single camera, and superimpose the errors of the four dimensional transformation matrix rectangles to form an error point set , , represents the error point set of the dimensional transformation matrix from the th camera; Denote the sum of the error values of the coordinate transformation of the athletes in the common area photographed by the four cameras as ; After unifying the scale and standardizing the two error sets and , through the CRITIC weight method, appropriate weights and are assigned to and , the total minimum error is solved by weighted sum to determine the optimal dimensional transformation matrix; the specific method is as follows:

[0108] S511. Since the sources and scales of the two error sets are different, denote , , and set the target length as , where , are respectively and the original number of elements. Adjust the error point set and the error sum to the same length for subsequent alignment of quantiles. Use linear interpolation to expand the length of the shorter set. Suppose the shorter set is expanded to obtain ; Apply the Min-Max normalization method to standardize and respectively, and calculate:

[0109]

[0110]

[0111] to obtain and ;

[0112] S512. Assign appropriate weights to the error point set and the error sum respectively. Apply the CRITIC weight method and define the normalized and as two indicators to construct a matrix:

[0113]

[0114] Then calculate the standard deviation. The standard deviation of

[0115]

[0116] The standard deviation of

[0117]

[0118] Calculate the Pearson correlation coefficient between the two columns:

[0119]

[0120] Calculate and The amount of information of is: , The amount of information of ; From this, the weight , ;

[0121] S520. Calculate the total minimum error and construct an objective function, expressed as:

[0122]

[0123] Satisfy the constraint conditions:

[0124]

[0125] Then repeat steps S511 - S512, traverse all rectangles for obtaining the dimension transformation matrix and the coordinates of the athletes after dimension transformation, and the dimension transformation matrix with the minimum total error obtained is the optimal dimension transformation matrix.

[0126] S600. After the current time sequence ends, switch to the next time sequence, perform dimension transformation on the entire sports field according to the optimal dimension transformation matrix obtained in step 500, save all points of the required field condition information after transformation in the point set , generate a top view through the information stored in this point set, clear other point sets, intercept the frame image of the next time sequence in the video, and repeat steps S100 - S500 to continuously obtain the top view.

[0127] The present invention will perform dimension change after fusing each frame of images captured by multiple cameras, thereby generating a top view of the full - field information in real time. Finally, it can realize the generation of a real - time top view of a sports event, helping users better master the information on the field and helping users deeply understand the dynamic evolution of the game situation.

[0128] The present invention designs a method for generating a global top view of a sports event based on multi - camera information fusion. By performing dimension transformation on the same - frame video images captured by multiple cameras, it generates a top view in real time, which is applicable to the video broadcast and field condition analysis scenarios of sports events. Relying on the original video scene features obtained by comprehensive shooting, it conducts various analysis works with the help of multiple computing and analysis means, effectively eliminating the influence of external interference factors. By converting the original video covering the entire sports field into a top - view perspective image, the system overcomes the problems of field condition observation caused by occlusion, bad viewing angles, incomplete pictures, etc., and provides more convenient and comprehensive support for users during the real - time viewing and field condition analysis of sports events.

[0129] The present invention performs dimensional transformation by identifying and screening rectangles, fuses the field conditions information from different perspectives according to the overlapping parts, constructs a top view of the competition venue, and can have a more reasonable and accurate grasp of the field conditions of sports events in real time.

[0130] Embodiment 2: This embodiment provides a system for generating a top view of the game situation based on multi-camera information fusion to implement the method of Embodiment 1. The system includes a module for obtaining the current frame information of the video, a rectangle matching and dimensional transformation module, a coordinate transformation error calculation module, a heterogeneous error weighted evaluation module, and a module for fusing local information to restore the top view perspective.

[0131] Among them: The module for obtaining the current frame information of the video is used to perform HSV color space separation processing on the image of the current frame in the video, obtain all the information points on the current field captured by a certain camera, and store them in a point set;

[0132] The rectangle matching and dimensional transformation module is used to extend and intersect the straight lines of the top view, record the rectangle coordinates, obtain the matching rectangle pairs according to the similarity, and use the rectangle coordinates of the original perspective and the top view perspective of the same camera to obtain the errors of these rectangle coordinates and generate a dimensional transformation matrix;

[0133] The coordinate transformation error calculation module is used to transform the coordinates of the athletes in the overlapping area to the top view perspective, calculate the Euclidean distance of the same athlete in the top view, and obtain the athlete coordinate transformation error;

[0134] The heterogeneous error weighted evaluation module is used to assign corresponding weight coefficients to different types of errors according to their respective importance levels and influence ranges, and comprehensively evaluate the overall error situation through error weighted summation;

[0135] The module for fusing local information to restore the top view perspective is used to fuse the local information of the cameras at the four corners of the current frame venue according to the overlapping area captured by the camera, obtain the global target of the stadium, and thus perform dimensional transformation on the full-field information of the stadium in the current frame to obtain the corresponding full-field top view.

[0136] Embodiment 3: This embodiment provides a non-transitory computer-readable storage medium, on which computer instructions are stored. The computer instructions cause the computer to execute the method for generating a top view of the game situation based on multi-camera information fusion. The method includes the following steps:

[0137] S100. Place cameras at the four corners of the sports field respectively, perform real-time shooting and recording of the stadium simultaneously, and intercept the current frame image in the video; take one of the corners of the stadium as the coordinate origin, establish a plane rectangular coordinate system, and record the relative positions of the four cameras as well as the height and width of the stadium;

[0138] S200. Separate the current frame image of each camera in the HSV color space, set the value range to identify all the straight lines and intersection points on the image. Arbitrarily select a corner point on the image as a reference point, calculate and determine the intersection point closest to the reference point, and then find other intersection points related to it by analyzing the straight line where the intersection point is located. Through these intersection points, gradually determine the four vertex coordinates of the rectangle and store them in the point set in order. ;

[0139] At moment, create a legend of the full - field top - view, and store all the intersection points in the top - view after transformation through the transformation matrix in the point set . Divide the field into rectangles by connecting the intersection points, record the coordinates of all rectangles in the order from left to right and from small to large, and store them as the rectangle point set obtained based on the top - view. ;

[0140] S300. According to the four vertices of the rectangle on a certain frame image obtained in step S200, find and screen the corresponding matrix coordinate points in the top - view, restore all the coordinates in this frame image to the top - view through the coordinate restoration method, and then match the target rectangle and the selected rectangle through the rectangle matching and error analysis algorithm to accurately screen the rectangle.

[0141] S400. Introduce a method for calculating the athlete coordinate transformation error based on multi - perspective camera fusion, obtain the athlete coordinates in the overlapping area between cameras, and use matrix transformation to the top - view perspective to calculate the Euclidean distance of the same athlete in the top - view to obtain the athlete coordinate transformation error.

[0142] S500. Superimpose the errors of the four - dimensional transformation matrices obtained from the non - overlapping areas of the four cameras to form an error point set ; Construct an error point set with the athlete coordinate transformation errors in the overlapping areas of the four cameras ; Set coefficients for the two error point sets respectively, and use the minimum weighted sum as the objective function to obtain the optimal dimensional transformation matrix.

[0143] S600. After the end of the current time series, switch to the next time series. Perform dimensional transformation on the entire sports field according to the optimal dimensional transformation matrix obtained in step 500, save all the points of the required field conditions after transformation in the point set . Generate a top - view through the information stored in this point set, clear other point sets, intercept the frame image of the next time series in the video, and repeat steps S100 - S500 to continuously obtain the top - view.

[0144] Embodiment 4: This embodiment provides an electronic device, which may include: a processor, a communications interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus. The processor can call the logical instructions in the memory to execute a method for generating a top view of the game situation based on multi-camera information fusion. The method includes the following steps:

[0145] S100. Place cameras at the four corners of the sports field respectively, simultaneously perform real-time shooting and recording of the stadium, and intercept the current frame image in the video; take one of the corners of the stadium as the coordinate origin, establish a plane rectangular coordinate system, and record the relative positions of the four cameras as well as the height and width of the stadium.

[0146] S200. Perform HSV color space separation on the current frame image of each camera, set the value range to identify all the straight lines and intersection points on the image, randomly select a corner point on the image as a reference point, calculate and determine the intersection point closest to the reference point, and then find other intersection points related to it by analyzing the straight line where the intersection point is located. Through these intersection points, gradually determine the four vertex coordinates of the rectangle and store them in the point set in order. ;

[0147] At moment, create a legend of the top view of the whole field, and store all the intersection points in the top view into the point set after transformation by the transformation matrix. In, divide the field into rectangles by connecting the intersection points, record the coordinates of all the rectangles in the order from left to right and from small to large, and store them as the rectangle point set obtained based on the top view. ;

[0148] S300. According to the four vertexes of the rectangle on a certain frame image obtained in step S200, find and screen the corresponding matrix coordinate points in the top view, restore all the coordinates in this frame image to the top view by the method of coordinate restoration, and then match the target rectangle and the selected rectangle through the rectangle matching and error analysis algorithms to accurately screen the rectangles.

[0149] S400. Introduce a method for calculating the athlete coordinate transformation error based on multi-view camera fusion, obtain the athlete coordinates in the overlapping area between cameras, and use matrix transformation to the top view perspective to calculate the Euclidean distance of the same athlete in the top view to obtain the athlete coordinate transformation error.

[0150] S500. Superimpose the errors of the four-dimensional transformation matrices obtained from the non-overlapping areas of the four cameras to form an error point set. ; Construct an error point set based on the coordinate transformation error of the athletes in the overlapping area of the four cameras. ; Set coefficients for the two error point sets respectively, and use the minimum weighted sum as the objective function to obtain the optimal dimension transformation matrix.

[0151] S600. After the current time sequence ends, switch to the next time sequence, perform dimension transformation on the entire sports field according to the optimal dimension transformation matrix obtained in step 500, and save all the points of the required field condition information after transformation in the point set. Generate a top view through the information stored in this point set, clear other point sets, intercept the frame image of the next time sequence in the video, and repeat steps S100~S500 to continuously obtain the top view.

[0152] In addition, when the logical instructions in the above-mentioned memory can be implemented in the form of software function units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), magnetic disks, or optical discs that can store program codes.

[0153] Embodiment 5: The computer program product provided in this embodiment includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the method for generating a top view of the game situation based on multi-camera information fusion. The method includes the following steps:

[0154] S100. Place cameras at the four corners of the sports field respectively, simultaneously record the sports field in real time, and intercept the current frame image in the video; use one of the corners of the sports field as the coordinate origin to establish a plane rectangular coordinate system, and record the relative positions of the four cameras, as well as the height and width of the sports field.

[0155] S200. Separate the current frame image of each camera in the HSV color space, set the value range to identify all the straight lines and intersection points on the image, arbitrarily select a corner point on the image as a reference point, calculate and determine the intersection point closest to the reference point, and then find other intersection points related to it by analyzing the straight line where the intersection point is located. Through these intersection points, gradually determine the four vertex coordinates of the rectangle and store them in the point set in sequence. ;

[0156] At moment, create a legend of the full - field top - view, and store all the intersection points in the top - view after transformation through the transformation matrix in the point set . Divide the field into rectangles by connecting the intersection points, record the coordinates of all rectangles in the order from left to right and from small to large, and store them as the rectangle point set obtained based on the top - view. ;

[0157] S300. According to the four vertexes of the rectangle on a certain frame image obtained in step S200, find and screen the corresponding matrix coordinate points in the top - view, restore all the coordinates in this frame image to the top - view through the coordinate restoration method, and then match the target rectangle and the selected rectangle through the rectangle matching and error analysis algorithms to accurately screen the rectangle.

[0158] S400. Introduce a method for calculating the athlete coordinate transformation error based on multi - perspective camera fusion, obtain the athlete coordinates in the overlapping area between cameras, and use matrix transformation to the top - view perspective to calculate the Euclidean distance of the same athlete in the top - view to obtain the athlete coordinate transformation error.

[0159] S500. Superimpose the errors of the four - dimensional transformation matrices obtained from the non - overlapping areas of the four cameras to form an error point set ; Construct an error point set with the athlete coordinate transformation errors in the overlapping areas of the four cameras ; Set coefficients for the two error point sets respectively, and take the minimum weighted sum as the objective function to obtain the optimal dimensional transformation matrix.

[0160] S600. After the current time sequence ends, switch to the next time sequence, perform dimensional transformation on the entire sports field according to the optimal dimensional transformation matrix obtained in step 500, save all the points of the required field conditions after transformation in the point set . Generate a top - view through the information stored in this point set, clear other point sets, intercept the frame image of the next time sequence in the video, and repeat steps S100 - S500 to continuously obtain the top - view.

[0161] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0162] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0163] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for generating a bird's-eye view of a game situation based on multi-camera information fusion, characterized in that: include: S100, placing cameras at the four corners of the stadium, simultaneously shooting and recording the stadium in real time, and capturing the current frame image in the video; Using one of the corners of the stadium as the origin, a rectangular coordinate system is established to record the relative positions of the four cameras and the height and width of the stadium; S200, perform HSV color space separation on the current frame image of each camera, set a value range to identify all straight lines and intersections on the image, randomly select a corner point on the image as a reference point, calculate and determine the intersection closest to the reference point, and then find other intersections related to the straight line where the intersection is located by analyzing the straight line, and gradually determine the coordinates of the four vertices of the rectangle through these intersections and store them in the point set in order. ; exist Always create a top view legend for the entire scene, and store all intersection points in the top view into a point set after transforming them through the transformation matrix In the example, the site is divided into Rectangles, record the coordinates of all rectangles in order from left to right and from small to large, and store them as a rectangular point set based on the top view ; S300, according to the four vertices of the rectangle on a certain frame image obtained in step S200, find and filter the corresponding matrix coordinate points in the top view, restore all the coordinates in the frame image to the top view by a coordinate restoration method, and then match the target rectangle with the selected rectangle by a rectangle matching and error analysis algorithm to accurately filter the rectangle; S400, introducing a method for calculating the coordinate transformation error of an athlete based on multi-view camera fusion, obtaining the coordinates of the athlete in the overlapping area between cameras, and using a matrix to transform to a bird's-eye view, calculating the Euclidean distance of the same athlete in the bird's-eye view, and obtaining the coordinate transformation error of the athlete; S500, superimposing the errors of the four-dimensional transformation matrices obtained in the non-overlapping areas of the four cameras to form an error point set ; Construct the error point set based on the athlete coordinate transformation errors in the overlapping area of ​​the four cameras ; Set coefficients for the two error point sets respectively, take the weighted sum minimum as the objective function, and obtain the optimal dimensional transformation matrix; S600: After the current sequence ends, switch to the next sequence, perform dimension transformation on the entire sports field according to the optimal dimension transformation matrix obtained in step 500, and save all points of the required field information after transformation in the point set In the process, a bird's-eye view is generated by the information stored in the point set, and other point sets are cleared, and the frame image of the next time sequence in the video is captured, and steps S100 to S500 are repeated to continuously obtain a bird's-eye view.

2. The method for generating a bird's-eye view of a game situation based on multi-camera information fusion according to claim 1, characterized in that: In step S100, the lower left corner of the stadium is used as the coordinate origin to establish Plane rectangular coordinate system, recording the relative positions of the four cameras , , , The height of the stadium is and width is .

3. The method for generating a bird's-eye view of a game situation based on multi-camera information fusion according to claim 2, characterized in that: In step S200, the current frame image of each camera is separated in HSV color space, and a value range is set to identify all straight lines and intersections on the image. A corner point on the image is randomly selected as a reference point, and the intersection closest to the reference point is calculated and determined. Then, by analyzing the straight line where the intersection point is located, other intersections related to it are found. Through these intersections, the coordinates of the four vertices of the rectangle are gradually determined and stored in the point set in order. ; The specific method is: S210, convert the current frame image from the RGB three-color channel to the HSV color space, take the lower left corner of the current frame image as the origin, and establish Plane rectangular coordinate system, set the appropriate value range, and identify all straight line equations , calculate the intersection points of all straight lines and record them; S220, randomly select one of the four vertices of the current frame image as the reference point, and assume that the length of the image is , the height is , will be with the point The distance is the minimum The intersection of ,point It is produced by the intersection of two straight lines. Point coordinates, get the The analytical expression of the two straight lines of the point is recorded as and ; S230, With The intersection point is obtained by combining the equations of the identified lines outside , , … , similarly, With The identified straight lines outside are combined to obtain the intersection point , , … ;calculate and and Distance of points and , respectively select and The corresponding points are denoted as and , the other straight line where the two points lie is and ; S240, straight line With straight line The intersection point is obtained by combining the analytical expressions of , and then the four points obtained As vertex coordinates are stored in order in the point set middle; S250, repeat steps S210 to S240, traverse according to the difference in the distance between the two points, find the point set of the rectangle on the original image, until no new distance difference appears, then all the rectangular point sets on the original image are found, and stored in the point set in sequence different rows.

4. The method for generating a bird's-eye view of a game situation based on multi-camera information fusion according to claim 3, characterized in that: In step S200, Always create a top view legend for the entire scene, and store all intersection points in the top view into a point set after transforming them through the transformation matrix In the example, the site is divided into Rectangles, record the coordinates of all rectangles in order from left to right and from small to large, and store them as a rectangular point set based on the top view ; The specific method is: S260, when the timing is When the image of the entire playing field or the dimensions of the field is obtained, a top-view coordinate system with the vertices of the playing field as the coordinate origin is created. , get the coordinates of all intersection points on the site , according to the coordinates of the four vertices of the site, the coordinates on the matching coordinate system Calculate a three-row and three-column transformation matrix, and multiply all points on the field by this transformation matrix to get the new intersection coordinates , so as to obtain the top view F of the site and record the time series Afterwards, directly obtain the top view F of the initial state; S270, obtained in step S260 By taking any two points whose horizontal and vertical coordinates are different, we can determine the four points of a rectangle and store them in an array in clockwise order. In the example, all the points are connected to generate n horizontal lines and m vertical lines, which divide the entire top view site into several small rectangles and large rectangles composed of small rectangles. Rectangle , record the coordinates of all rectangles in order from left to right and from small to large, using eight lines The obtained rectangular point set is stored in an array of columns. .

5. The method for generating a bird's-eye view of a game situation based on multi-camera information fusion according to claim 4, characterized in that: In step S300, according to the four vertices of the rectangle on a frame image obtained in step S200, the corresponding matrix coordinate points in the top view are found and screened, all the coordinates in the frame image are restored to the top view by the coordinate restoration method, and then the target rectangle and the selected rectangle are matched by the rectangle matching and error analysis algorithm to accurately screen the rectangle, including: S310, based on the four vertices of the rectangle on a certain frame image obtained in step S200, a rectangle matching scoring algorithm model is introduced to preliminarily identify and select a suitable rectangle; S320, adopting a loop method, that is, selecting the point set of the coordinates obtained when the target rectangle is selected Perform shift processing, and then repeat step S310 to obtain all optimal solutions; S330: After obtaining the most suitable rectangle pair, a dimensional transformation matrix is ​​calculated.

6. The method for generating a bird's-eye view of a game situation based on multi-camera information fusion according to claim 5, characterized in that: In step S310, based on the four vertices of the rectangle on a certain frame image obtained in step S200, a rectangle matching scoring algorithm model is introduced to preliminarily identify and select suitable rectangles; the specific method is: S311, based on the point set obtained in step S200 and Get a new top view F1, first compare the boundary lengths of the top view F and F1. If a boundary length of F1 is greater than the corresponding boundary length in F, it means that the currently selected rectangle point is deviated or inappropriate. Then compare the rectangles in the figure. If the length of a side of the rectangle in F1 is obviously greater than the length of the corresponding side of the rectangle in F, it can be determined that the current target rectangle is not compatible with the selected rectangle in size and shape. Remove this part of the rectangle and store the remaining rectangle coordinate points in a new point set. Medium; S312, based on point set , Calculate the length Length, width Width and the width of each rectangle under two viewing angles. , and obtain the centroid coordinates of each set of rectangles at the same time, and define the aspect ratio as , the angle difference is , the center of gravity distance is , calculate the Hausdorff distance of the point set of the rectangular coordinates under two perspectives, expressed as: According to the top view creation method in step S200, the top view F2 restored under the camera's viewing angle is obtained, thereby obtaining two straight lines representing the viewing angle in the top view F2, and the two straight lines are extended in the opposite direction and intersected to obtain the coordinates of an imaginary camera. , by comparing with the camera coordinates in step S100 Compare and get a Euclidean distance difference ; Then normalize the data and calculate the normalized similarity score : in =0.12, =0.15, =0.13, =0.10, =0.5, traverse and All rectangular coordinate groups in , calculate the similarity scores of all matching combinations , stored in point sets , In the example, compare the scores, get the matching rectangles, and update them in a one-to-one order. and .

7. The method for generating a bird's-eye view of a game situation based on multi-camera information fusion according to claim 5, characterized in that: In step S400, the coordinates of the athletes in the overlapping area between the cameras are obtained, and the coordinates are transformed to the top view using a matrix, and the Euclidean distance of the same athlete in the top view is calculated to obtain the coordinate transformation error of the athlete; the specific method is: S410, selecting two viewing angles of camera A and camera B, using YOLO to identify all athletes in the viewing angle of the camera and taking the coordinates (X, Y) of their center points as the positions of the athletes; S420, calculating the coordinates of the athletes in the overlapping area after dimension transformation at each viewing angle, and storing them in a point set. The specific implementation steps are as follows: S421, according to the dimension change matrix calculated in step S300, the position of the athlete under each camera's view is dimensionally transformed to obtain the coordinates (X', Y') after the dimension change, and the coordinates of all athletes photographed by the i-th camera after the dimension change are stored in the point set Si; S422, calculating the coordinate range of the overlapping area between cameras, selecting the coordinates that meet the range in the point set and storing them in a new point set S'i; S430, taking point set S'1 and point set S'2, calculating the Euclidean distance between all point pairs in the two perspectives as the matching cost, constructing a cost matrix, and applying the Hungarian algorithm to match the point pairs; S440, for each matching row and column index (i, j), the point corresponding to camera A , and the point of camera B , reorder the point set according to the matching results to ensure that the coordinates of the corresponding positions belong to the same player, and for each matching point pair ( , ) to obtain the coordinate error ; Step 450: Repeat the process from step S410 to step S440 for viewing angles B and C, viewing angles C and D, and viewing angles D and A, and superimpose the error values ​​of the overlapping areas to obtain a final error value. .

8. The method for generating a bird's-eye view of a game situation based on multi-camera information fusion according to claim 7, characterized in that: The specific implementation steps of step S430 are: S431, define the detection point set of camera A as , the detection point set of camera B is , calculate the Euclidean distance between all point pairs in the two perspectives as the matching cost, and construct the cost matrix as , whose elements are ; S432. Solve the row index and column index set , so that , get the ordered matching index list .

9. The method for generating a bird's-eye view of a game situation based on multi-camera information fusion according to claim 7, characterized in that: The specific method of step S500 is: S510, obtain a four-dimensional transformation matrix from a rectangle obtained by preliminary screening of a single camera, and superimpose the errors of the four-dimensional transformation matrix rectangles to form an error point set , , Indicates that from The error point set of the dimensional transformation matrix of the four cameras; the error value of the coordinate transformation of the athlete in the common area captured by the four cameras is recorded as ; After unifying the scale and standardizing the two error sets and , through the CRITIC weight method and Assign appropriate weights to each and , by weighted summing and solving the total minimum error to determine the optimal dimension transformation matrix; the specific method is as follows: S511. Since the sources and scales of the two error sets are different, , , let the target length be ,in , They are and The number of original elements of the error point set and error and Adjust to the same length , used for subsequent alignment of quantiles, using linear interpolation to extend the length of the shorter set, assuming that the shorter set is extended get ; Min-Max normalization method is used to and Perform standardization and calculate: get and ; S512, error point set and error and Assign appropriate weights to each of them and apply the CRITIC weight method to normalize the and Defined as two indicators, construct the matrix: Then calculate the standard deviation, The standard deviation of is: The standard deviation of is: Compute the Pearson correlation coefficient between two columns: According to the obtained Pearson correlation coefficient, and The amount of information, The amount of information is: , The amount of information is: ; The weight is obtained from this , ; S520, calculate the total minimum error and construct the objective function, which is expressed as: Satisfy the constraints: Then, step S511 to step S512 are repeated to traverse all rectangles of the dimensional transformation matrix and the coordinates of the athletes after the dimensional transformation. The dimensional transformation matrix with the smallest total error is the optimal dimensional transformation matrix.

10. A system for generating a bird's-eye view of a game situation based on multi-camera information fusion for implementing any one of the methods of claims 1 to 9, characterized in that: It includes a module for obtaining the information of the current video frame, a module for rectangle matching and dimension transformation, a module for calculating the coordinate transformation error, a module for weighted evaluation of heterogeneous errors, and a module for restoring the bird's-eye view by fusing local information, wherein: the module for obtaining the information of the current video frame is used to perform HSV color space separation processing on the image of the current frame in the video, and obtain all the information points on the current scene captured by a certain camera and store them in a point concentration; The rectangle matching and dimension transformation module is used to extend and intersect the straight lines of the top view, record the rectangle coordinates, obtain the matching rectangle pair based on similarity matching, use the rectangle coordinates of the original view and the top view of the same camera, obtain the error of these rectangle coordinates and generate the dimension transformation matrix; A coordinate transformation error calculation module is used to transform the coordinates of the athletes in the overlapping area to a bird's-eye view, calculate the Euclidean distance of the same athlete in the bird's-eye view, and obtain the coordinate transformation error of the athlete; The heterogeneous error weighted assessment module is used to assign corresponding weight coefficients to the error types from different sources according to their respective importance and impact range, and to evaluate the overall error situation through error weighting and comprehensive assessment; The module for restoring the bird's-eye view by fusing local information is used to fuse the local information of the cameras at the four corners of the venue in the current frame according to the overlapping areas captured by the cameras, and obtain the global target of the stadium, thereby transforming the dimension of the full-field information of the stadium in the current frame to obtain the corresponding full-field bird's-eye view.

Citation Information

Patent Citations

  • Video positioning and speed measuring method, device and equipment for swimmer and storage medium

    CN116309686A

  • Dynamic-view-angle aerial view real-time generation method for sports event site condition

    CN118154728A

  • Method for determining calibration precision of binocular deflection system, medium and equipment

    CN119394165A