A method, device, electronic device and medium for switching close-up of an object in a specific scene

By setting up multiple free-view cameras and telephoto zoom cameras in the ball game venue, identifying and tracking the target object, selecting the best viewing angle and switching to the close-up screen, the problem of being difficult to achieve clear close-up of star players' ball-holding attacks in the prior art is solved, and the audience's needs for close-up viewing of the game are met.

CN115529468BActive Publication Date: 2025-05-06HISENSE GRP HLDG CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110709047.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-06-25
Publication Date
2025-05-06
Estimated Expiration
2041-06-25

AI Technical Summary

Technical Problem

It is difficult for the existing technology to achieve a clear close-up picture of star players playing the ball during live ball game, resulting in the audience being unable to meet the requirements for star players to watch the game in close-up.

Method used

By setting up multiple free viewing cameras and telephoto zoom cameras on the site, identify and track the target object, select the best free viewing angle, and switch to the close-up picture taken by the corresponding telephoto zoom camera.

Benefits of technology

It has achieved a clear close-up display of star players' ball attacks during the live broadcast, meeting the audience's need for star players to watch the game in close-up.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115529468B_ABST
    Figure CN115529468B_ABST
Patent Text Reader

Abstract

This application provides a method, apparatus, and electronic device for switching close-up shots of objects in a specific scenario. The method includes: processing video data streams from multiple free-view cameras to identify target objects from multiple free-view perspectives; tracking the target objects in the video data streams from multiple free-view perspectives based on the identification information of the target objects from multiple free-view perspectives; if a target object is detected holding a target object at a target time, selecting a target free-view perspective to display the target object from the multiple free-view perspectives based on the video data streams from multiple free-view perspectives acquired at the target time; controlling a playback device to switch the playback screen to the shooting screen of the target free-view perspective, and then switching to a close-up shot of the target object taken by a telephoto zoom camera corresponding to the target free-view perspective. This application can achieve close-up shots of a designated star player's ball-handling attack, satisfying the audience's requirement for full-court close-up viewing of a designated star player.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a method, device and electronic device for switching close-up of an object in a specific scene. Background Art

[0002] When watching a live broadcast of a ball game (such as a basketball game, a football game, etc.), the audience hopes to watch a close-up of the basketball star at the best free viewing angle every time their favorite star player attacks with the ball, so that they can see every detail of the star player's attack with the ball.

[0003] However, currently, a short-focus fixed-focus camera is used to shoot images with a free perspective. When the images shot by the camera are enlarged, the images will be distorted, and the details of the star players' offensive movements with the ball cannot be clearly seen. As a result, it is difficult to achieve a close-up of the star players' offensive movements with the ball during the live broadcast, and the audience's demand for a full-game close-up of their favorite star players cannot be met. Summary of the invention

[0004] The present application provides a method, device and electronic device for switching close-up of an object in a specific scenario, which are used to achieve a close-up of a designated star player's attack with the ball during a live broadcast, thereby satisfying the audience's requirement for watching a full-court close-up of the designated star player.

[0005] The specific technical solutions provided by the embodiments of this application are as follows:

[0006] In a first aspect, an embodiment of the present application provides a method for switching a close-up of an object in a specific scene, wherein a plurality of free-viewing angle cameras and a plurality of telephoto zoom cameras are arranged in a venue corresponding to the specific scene, and the specific scene includes a plurality of objects, and the method includes:

[0007] Processing the video data streams of the multiple free-viewing angles captured by the multiple free-viewing angle cameras to identify the target objects in the multiple free-viewing angles;

[0008] Tracking the target object in the video data streams of the multiple free perspectives based on identification information of the target object under the multiple free perspectives;

[0009] If it is detected at a target moment that the target object holds a target object, selecting a target free perspective for displaying the target object from a plurality of free perspectives according to a plurality of free perspective video data streams acquired at the target moment;

[0010] After controlling the playback device to switch the playback screen to the shooting screen of the target free viewing angle, it switches to the close-up screen of the target object shot by the telephoto zoom camera corresponding to the target free viewing angle.

[0011] In the embodiment of the present application, by identifying the target object under multiple free perspectives, the target object is tracked. When it is determined that the target object holds the target object, the target free perspective of the target object can be selected from multiple free perspectives to display; then the playback device is controlled to switch the playback screen to the shooting screen of the target free perspective, and then switch to the close-up picture of the target object taken by the telephoto zoom camera corresponding to the target free perspective, so as to achieve a close-up of the target object when holding the target object. When the specific scene is a live game scene and the target object is a star player, a close-up of the star player's ball-holding offense can be achieved during the live broadcast, satisfying the audience's requirement for a full-court close-up of their favorite star players.

[0012] In some exemplary embodiments, the processing of the plurality of free-viewing video data streams respectively to identify target objects in the plurality of free-viewing angles includes:

[0013] For the multiple free-viewing angle video data streams, the following operations are performed respectively:

[0014] Performing object detection on each free image frame of a free-viewing angle video data stream to obtain a plurality of object detection frames corresponding to each free image frame;

[0015] Identify each object detection frame in the obtained multiple object detection frames to obtain identification information of each object detection frame;

[0016] The identification information of each object detection frame is matched with the identification information of the target object in a preset object recognition information library to determine a target object detection frame where the identification information of the target object is located, and the target object detection frame represents the target object.

[0017] In the embodiment of the present application, for each free image frame of a free-viewing video data stream, multiple object detection frames in each free image frame are first detected, and then each object detection frame is identified, and the identification information of each object detection frame is matched with the identification information of the target object to determine the target object detection frame where the identification information of the target object is located, thereby identifying the target object in the free viewing angle. In this way, the target objects in multiple free viewing angles can be accurately identified.

[0018] In some exemplary embodiments, after performing target detection on each free image frame in a free-view video data stream to obtain a plurality of object detection frames corresponding to each free image frame, the method further includes:

[0019] Performing object recognition on each magnified image frame of a plurality of magnified video data streams captured by the plurality of telephoto zoom cameras, respectively, to obtain a plurality of object recognition information of each magnified image frame in each magnified image frame; wherein the telephoto zoom camera is pre-zoomed to a set focal length for capturing the magnified video data stream;

[0020] The step of respectively identifying a plurality of object detection frames corresponding to each free image frame in each free image frame to obtain identification information of the plurality of object detection frames of each free image frame includes:

[0021] Respectively identifying a plurality of object detection frames corresponding to each free image frame in each free image frame to obtain candidate recognition information of the plurality of object detection frames of each free image frame;

[0022] The candidate recognition information of the multiple object detection frames of each free image frame is combined with the multiple object recognition information of a corresponding frame of enlarged image taken by each of the multiple telephoto zoom cameras to obtain the recognition information of the multiple object detection frames of each free image frame.

[0023] In the embodiment of the present application, since the object detection frame of each frame of the free image may not be clearly identified, and the object corresponding to the object detection frame in each frame of the enlarged image is more easily identified, the candidate identification information of multiple object detection frames of each frame of the free image is combined with the corresponding multiple object identification information of each frame of the enlarged image, the identification information of the multiple object detection frames of each frame of the free image can be obtained more accurately.

[0024] In some exemplary embodiments, if it is detected at a target moment that the target object holds a target object, selecting a target free perspective for photographing the target object from a plurality of free perspectives according to a plurality of free perspective video data streams acquired at the target moment includes:

[0025] If it is detected at the target moment that the target object holds the target object, then the perspective matching degrees of the multiple free perspectives are determined respectively according to the perspective information of the target object in the video data streams of the multiple free perspectives acquired at the target moment; wherein the perspective matching degrees are used to characterize the matching degree between the corresponding free perspective and the target free perspective;

[0026] A target free viewing angle for photographing the target object is selected from the multiple free viewing angles according to the viewing angle matching degrees of the multiple free viewing angles.

[0027] Through the above implementation, the perspective matching degree of multiple free perspectives is determined according to the perspective information of the target object in the video data streams of multiple free perspectives obtained at the target moment, and then the target free perspective for shooting the target object is selected from the multiple free perspectives according to the perspective matching degree of the multiple free perspectives, so that the best free perspective for shooting the target object can be selected.

[0028] In some exemplary embodiments, determining the viewing angle matching degrees of the plurality of free viewing angles respectively according to the viewing angle information of the target object in the plurality of free viewing angle video data streams acquired at the target time includes:

[0029] For the multiple free-viewing angle video data streams acquired at the target moment, the following operations are performed respectively:

[0030] Performing face recognition on the target object in a specified frame of free image of a video data stream to obtain the face confidence of the target object in the frame of free image; wherein the target object in the frame of free image holds a target object;

[0031] Determining, according to the depth map of the one frame of free image, a target distance between the target object and a free-viewing angle camera that shoots the one frame of free image;

[0032] The viewing angle matching degree of the free viewing angle corresponding to the one frame of free image is determined according to the face confidence and the target distance.

[0033] In the above implementation, the face confidence of the target object can be identified to represent the degree to which the face of the target object is facing the corresponding free-view camera. The face confidence is combined with the target distance between the target object and the corresponding free-view camera to determine the perspective matching degree of the corresponding free view.

[0034] In some exemplary embodiments, after performing face recognition on the target object in a free image frame of a video data stream and obtaining the face confidence of the target object in the free image frame, the method further includes:

[0035] If the face confidence is less than a preset threshold, the free viewing angle corresponding to the one frame of free image is used as a non-target free viewing angle.

[0036] In the above embodiment, if the face confidence of the target object under a certain free viewing angle is less than a preset threshold, it means that the free viewing angle is not suitable as the target free viewing angle for shooting the target object (which can be understood as the optimal viewing angle). The free viewing angle can be directly excluded without calculating the target distance between the target object and the corresponding free viewing angle camera.

[0037] In some exemplary embodiments, the telephoto zoom camera corresponding to the target free viewing angle is:

[0038] The telephoto zoom camera that is closest to the free view angle camera corresponding to the target free view angle; or

[0039] The telephoto zoom camera is bound to the free view camera corresponding to the target free view.

[0040] In the above implementation, the telephoto zoom camera that is closest to the free view camera corresponding to the target free view can capture a close-up image of the target object more clearly. The telephoto zoom camera bound to the free view camera corresponding to the target free view can quickly determine the close-up image captured by the telephoto zoom camera that needs to be switched.

[0041] In a second aspect, an embodiment of the present application provides a device for switching close-up of an object in a specific scene, wherein a plurality of free-viewing angle cameras and a plurality of telephoto zoom cameras are arranged in a venue corresponding to the specific scene, and the specific scene includes a plurality of objects, and the device includes:

[0042] An object recognition module, used for processing the video data streams of the multiple free-viewing angles captured by the multiple free-viewing angle cameras to recognize target objects under the multiple free-viewing angles;

[0043] An object tracking module, configured to track the target object in the video data streams of the multiple free perspectives based on identification information of the target object in the multiple free perspectives;

[0044] A perspective selection module, configured to select a target free perspective for displaying the target object from a plurality of free perspectives according to a plurality of free perspective video data streams acquired at the target moment if the target object is detected to be holding a target object at the target moment;

[0045] The control switching module is used to control the playback device to switch the playback screen to the shooting screen of the target free viewing angle, and then switch to the close-up picture of the target object shot by the telephoto zoom camera corresponding to the target free viewing angle.

[0046] In some exemplary embodiments, the object recognition module further includes:

[0047] The detection submodule is used to perform object detection on each free image frame of a free-viewing angle video data stream, and obtain a plurality of object detection frames corresponding to each free image frame;

[0048] A first recognition submodule is used to recognize each object detection frame in the obtained multiple object detection frames to obtain recognition information of each object detection frame;

[0049] The matching submodule is used to match the identification information of each object detection frame with the identification information of the target object in a preset object recognition information library to determine the target object detection frame where the identification information of the target object is located, and the target object detection frame represents the target object.

[0050] In some exemplary embodiments, the object recognition module further includes a second recognition submodule, configured to:

[0051] Performing object recognition on each magnified image frame of a plurality of magnified video data streams captured by the plurality of telephoto zoom cameras, respectively, to obtain a plurality of object recognition information of each magnified image frame in each magnified image frame; wherein the telephoto zoom camera is pre-zoomed to a set focal length for capturing the magnified video data stream;

[0052] The first identification submodule is further used for:

[0053] Respectively identifying a plurality of object detection frames corresponding to each free image frame in each free image frame to obtain candidate recognition information of the plurality of object detection frames of each free image frame;

[0054] The candidate recognition information of the multiple object detection frames of each free image frame is combined with the multiple object recognition information of a corresponding frame of enlarged image taken by each of the multiple telephoto zoom cameras to obtain the recognition information of the multiple object detection frames of each free image frame.

[0055] In some exemplary embodiments, the perspective selection module includes:

[0056] A determination submodule, for determining the viewing angle matching degree of a plurality of free viewing angles respectively according to the viewing angle information of the target object in the video data streams of a plurality of free viewing angles acquired at the target moment if it is detected that the target object holds the target object at the target moment; wherein the viewing angle matching degree is used to characterize the matching degree between the corresponding free viewing angle and the target free viewing angle;

[0057] A selection submodule is used to select a target free viewing angle for photographing the target object from the multiple free viewing angles according to the viewing angle matching degrees of the multiple free viewing angles.

[0058] In some exemplary embodiments, the determining submodule is further configured to:

[0059] For the multiple free-viewing angle video data streams acquired at the target moment, the following operations are performed respectively:

[0060] Performing face recognition on the target object in a free image frame of a video data stream to obtain a face confidence of the target object in the free image frame; wherein the target object in the free image frame holds a target object;

[0061] Determining, according to the depth map of the one frame of free image, a target distance between the target object and a free-viewing angle camera that shoots the one frame of free image;

[0062] The viewing angle matching degree of the free viewing angle corresponding to the one frame of free image is determined according to the face confidence and the target distance.

[0063] In some exemplary embodiments, the determining submodule is further configured to:

[0064] If the face confidence is less than a preset threshold, the free viewing angle corresponding to the one frame of free image is used as a non-target free viewing angle.

[0065] In some exemplary embodiments, the telephoto zoom camera corresponding to the target free viewing angle is:

[0066] The telephoto zoom camera that is closest to the free view angle camera corresponding to the target free view angle; or

[0067] The telephoto zoom camera is bound to the free view camera corresponding to the target free view.

[0068] In a third aspect, an embodiment of the present application provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program that can be executed on the processor, and when the computer program is executed by the processor, the processor implements any method described in the first aspect.

[0069] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described in any one of the first aspects is implemented.

[0070] The technical effects brought about by any implementation method of the second to fourth aspects can refer to the technical effects brought about by the corresponding implementation method in the first aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0072] Figure 1 A flowchart of a method for switching a close-up of an object in a specific scenario provided in an embodiment of the present application;

[0073] Figure 2 A schematic diagram of the arrangement of a telephoto zoom camera in a specific scenario provided in an embodiment of the present application;

[0074] Figure 3 A schematic diagram of the structure of an object close-up switching device in a specific scenario provided in an embodiment of the present application;

[0075] Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0076] In order to enable those skilled in the art to better understand the technical solutions in this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this application.

[0077] In the description of this application, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the feature. In the description of this application, unless otherwise specified, "plurality" means two or more.

[0078] At present, a short-focus fixed-focus camera is used to shoot images with a free perspective. When the images shot by the camera are enlarged, the images will be distorted, and the details of the star players' offensive movements with the ball cannot be clearly seen. As a result, it is impossible to achieve a close-up of the star players' offensive movements with the ball during the live broadcast, and the audience's demand for a full-game close-up of their favorite star players cannot be met.

[0079] In view of this, the embodiments of the present application provide a method, device and electronic device for switching close-ups of objects in a specific scenario, which are used to achieve a close-up of a target object when the target object holds the target object in a specific scenario. When the specific scene is a live game scene and the target object is a star player, a close-up of the star player's ball-holding attack can be achieved during the live broadcast, thereby satisfying the audience's requirements for watching their favorite star players in close-up throughout the game.

[0080] The object close-up switching method in a specific scenario of the present application is described in detail below with reference to the accompanying drawings and specific embodiments.

[0081] In a method for switching a close-up of an object in a specific scene provided in an embodiment of the present application, a plurality of free-viewing angle cameras and a plurality of telephoto zoom cameras are arranged in a venue corresponding to the specific scene, and the specific scene includes a plurality of objects. For example, the specific scene may be a live broadcast scene of a sports game, such as a basketball game, a football game, a volleyball game, etc. In this case, the object may be an athlete. In addition, the specific scene may also be other suitable scenes, which are not limited here.

[0082] In the venue corresponding to a specific scene, the number of free-viewpoint cameras and the number of telephoto zoom cameras can be set as needed; multiple free-viewpoint cameras can cover the entire viewing angle of the venue, and the position of each free-viewpoint camera can be determined based on the free viewing angle range captured by the free-viewpoint camera; multiple telephoto zoom cameras can be distributed at different positions in the venue, and can be set as needed, without limitation here.

[0083] Figure 1 The present invention provides a method for switching object close-up in a specific scenario, which can be executed by a server. Figure 1 As shown, the method may include the following steps:

[0084] Step S101 : Processing a plurality of free-viewing angle video data streams shot by a plurality of free-viewing angle cameras to identify target objects under a plurality of free-viewing angles.

[0085] In an embodiment of the present application, each free-viewpoint camera can upload a free-viewpoint video data stream to a server. When the server processes the received multiple free-viewpoint video data streams, it processes each frame of the video data stream (hereinafter referred to as a free image) for a free-viewpoint video data stream shot by a free-viewpoint camera to identify the target object in each free image frame. Specifically, each free image frame may include multiple objects, and each object can be detected first, and then the target object can be identified from each object. The target object can be one or more specified objects. For example, in a basketball game scene, including multiple basketball players, one or more specified basketball stars can be used as target objects.

[0086] Step S102 : tracking the target object in the video data streams of the multiple free viewpoints based on the identification information of the target object in the multiple free viewpoints.

[0087] In this step, after the target object is identified, identification information of the target object can be obtained, such as facial features, clothing features, etc. For a designated basketball star, clothing features may be clothing color, clothing numbers, etc. Based on the identification information of the target object, a target tracking algorithm can be used to track the target object in multiple free-view video data streams.

[0088] Optionally, because the target being tracked may appear anywhere on the free image, the feature representation learned by the target tracking algorithm needs to be spatially invariant. For example, a visual tracking algorithm based on a twin deep network (SiamRPN++) can be used. This algorithm uses a spatially aware sampling strategy to keep the learned features spatially invariant, which can better track the target object.

[0089] Specifically, after a target object in a free perspective is identified at the current moment in step S102, the position of the target object in a certain frame of free image in the video data stream of the free perspective at the current moment can be used as the starting position, and the target object can be continuously tracked using a target tracking algorithm.

[0090] Step S103, if it is detected at the target moment that the target object holds the target object, a target free perspective for displaying the target object is selected from the multiple free perspectives according to the video data streams of the multiple free perspectives acquired at the target moment.

[0091] In the process of tracking the target object, by detecting the position of the target object, combined with the position of the target object, it can be determined whether the target object is holding the target object, for example, in a basketball game scenario, the target object is a basketball. If the target object is detected to be holding the target object at the target moment, the video data streams of multiple free perspectives obtained at the target moment can be processed to determine the target free perspective that is most suitable for viewing the target object among the various free perspectives, that is, the best perspective.

[0092] Step S104, controlling the playback device to switch the playback screen to the shooting screen of the target free viewing angle, and then switching to the close-up screen of the target object shot by the telephoto zoom camera corresponding to the target free viewing angle.

[0093] In this step, the playback device may be a terminal device with a video playback function, such as a mobile phone, a television, a notebook, a tablet computer, a personal computer, a vehicle-mounted terminal, etc., which is not limited here.

[0094] After determining the target free viewpoint, the server can send the video data stream of the target free viewpoint to the playback device so that the playback device plays the shot image of the target free viewpoint; and send the enlarged video data stream shot by the telephoto zoom camera corresponding to the target free viewpoint to the playback device so that the playback device plays the shot image of the target free viewpoint and then plays the close-up image of the target object.

[0095] In some possible implementations, the telephoto zoom camera corresponding to the target free perspective may be: a telephoto zoom camera that is closest to the free perspective camera corresponding to the target free perspective; or, a telephoto zoom camera bound to the free perspective camera corresponding to the target free perspective. For example, one telephoto zoom camera may be bound to multiple free perspective cameras, which may be determined specifically according to the setting location.

[0096] In this embodiment, the telephoto zoom camera that is closest to the free view camera corresponding to the target free view can capture a close-up image of the target object more clearly. The telephoto zoom camera bound to the free view camera corresponding to the target free view can quickly determine the close-up image captured by the telephoto zoom camera that needs to be switched.

[0097] In the embodiment of the present application, by identifying the target object under multiple free perspectives, the target object is tracked. When it is determined that the target object holds the target object, the target free perspective of the target object can be selected from multiple free perspectives to display; then the playback device is controlled to switch the playback screen to the shooting screen of the target free perspective, and then switch to the close-up picture of the target object taken by the telephoto zoom camera corresponding to the target free perspective, so as to achieve a close-up of the target object holding the target object in a specific scene. When the specific scene is a live game scene and the target object is a star player, a close-up of the star player's ball-holding offense can be achieved during the live broadcast, satisfying the audience's requirements for a full-court close-up of their favorite star players.

[0098] Further, after step S104, when it is detected that the target subject does not hold the target object, the close-up image is switched to the best full-field viewing angle under the free viewing angle.

[0099] In some embodiments, when identifying the target object under multiple free perspectives in the above step S101, the following steps may be performed respectively for the video data streams under multiple free perspectives:

[0100] A1. Perform object detection on each free image frame of a free-viewing angle video data stream to obtain a plurality of object detection frames corresponding to each free image frame.

[0101] In this step, a target detection model can be used to perform target detection on each frame of free image. For example, the target detection model can use an object recognition and positioning algorithm based on a deep neural network, such as the YOLO (You Only LookOnce) algorithm. The target detection model can be a yolov5 model with a faster detection speed and a smaller model, or other models, which are not limited here. When using the target detection model for object detection, each object can correspond to a unique ID to generate an object detection frame for each object.

[0102] A2. Identify each object detection frame among the obtained multiple object detection frames to obtain identification information of each object detection frame.

[0103] The identification information of each object detection frame may be features that can uniquely characterize the corresponding object, such as facial features, clothing features, and other features, wherein clothing features and other features may be identified by the above-mentioned target detection model, and facial features may be obtained by a face recognition algorithm.

[0104] A3. Match the identification information of each object detection frame with the identification information of the target object in a preset object recognition information library to determine the target object detection frame where the identification information of the target object is located, and the target object detection frame represents the target object.

[0105] In this embodiment, the identification information of the target object can be obtained in advance, and then the identification information of the target object and the corresponding identity information can be stored in the object recognition information library. The identification information of the target object can be features that can uniquely characterize the target object, such as facial features, clothing features, and other features, and the identity information can include name, role, etc.

[0106] For example, when the target object is a basketball star, the identification information of the basketball star may include facial features, jersey color, jersey number, etc.; the identity information may include name, team, position in the team (center, power forward, small forward, shooting guard, point guard), etc.

[0107] In the embodiment of the present application, for each free image frame of a free-viewing video data stream, multiple object detection frames in each free image frame are first detected, and then each object detection frame is identified, and the identification information of each object detection frame is matched with the identification information of the target object to determine the target object detection frame where the identification information of the target object is located, thereby identifying the target object in the free viewing angle. In this way, the target objects in multiple free viewing angles can be accurately identified.

[0108] Considering that when multiple object detection frames in each free image frame are identified, there may be unclear identification, such as failure to recognize facial features. In order to more clearly identify each object detection frame, each free image frame can be combined with an enlarged image taken by a corresponding telephoto zoom camera to identify multiple object detection frames in each free image frame.

[0109] In some embodiments, after performing target detection on each free image frame in a free-view video data stream to obtain a plurality of object detection frames corresponding to each free image frame, the following steps may be further performed:

[0110] Object recognition is performed on each magnified image frame of multiple magnified video data streams shot by multiple telephoto zoom cameras to obtain multiple object recognition information of each magnified image frame; wherein the telephoto zoom camera is pre-zoomed to a set focal length for shooting the magnified video data stream.

[0111] The object identification information of each frame of the enlarged image may be a feature that can uniquely characterize the corresponding object, such as a face feature (which can be obtained through a face recognition algorithm), a clothing feature, or other features.

[0112] Furthermore, in the above step A2, the multiple object detection frames corresponding to each free image frame in each free image frame are respectively identified to obtain the identification information of the multiple object detection frames of each free image frame, which may include the following steps:

[0113] B1. Respectively identify multiple object detection frames corresponding to each free image frame in each free image frame to obtain candidate recognition information of multiple object detection frames of each free image frame.

[0114] B2. Combine the candidate recognition information of multiple object detection frames of each free image frame with the recognition information of multiple objects of a corresponding frame of enlarged image taken by each of the multiple telephoto zoom cameras to obtain the recognition information of multiple object detection frames of each free image frame.

[0115] For a certain frame of free image, a corresponding frame of enlarged image shot by each of the multiple telephoto zoom cameras is shot at the same time as the frame of free image.

[0116] For example, in a frame of free image, the position of each object detection frame in the image can be determined, and then the actual position of each object detection frame in the venue can be obtained. Then, in the corresponding enlarged image, each object can be identified, the position of each object can be obtained, and then the actual position of each object in the venue can be obtained. In this way, based on the actual position in the venue, the object in the enlarged image corresponding to the object detection frame in the free image can be determined. In addition, the candidate recognition information of multiple object detection frames in each frame of free image can be directly matched with the corresponding multiple object recognition information of each frame of enlarged image to determine which object in the enlarged image a certain object detection frame corresponds to.

[0117] For example, taking a basketball game scene as an example, based on a frame of free image, it is identified that the jersey color of a certain object detection frame is red and the jersey number is 4, but the facial features cannot be identified; based on a frame of enlarged image corresponding to the frame of free image, it can be identified that the jersey color of a certain object is red and the jersey number is 4, and the facial features are also identified. In this way, the recognition information of the above-mentioned object detection frame can be obtained, including the jersey color, jersey number, and facial features, and the recognition information of the object detection frame is matched with the recognition information of the target object in the object recognition information library, so as to quickly determine whether the object detection frame is the detection frame where the target object is located.

[0118] In the embodiment of the present application, since the object detection frame of each frame of the free image may not be clearly identified, and the object corresponding to the object detection frame in each frame of the enlarged image is more easily identified, the candidate identification information of multiple object detection frames of each frame of the free image is combined with the corresponding multiple object identification information of each frame of the enlarged image, the identification information of the multiple object detection frames of each frame of the free image can be obtained more accurately.

[0119] In some embodiments, if the target object is detected to be holding a target object at the target moment in step S103, then selecting a target free perspective for photographing the target object from the multiple free perspectives according to the video data streams of the multiple free perspectives acquired at the target moment may include the following steps:

[0120] a1. If it is detected at the target moment that the target object is holding the target object, the perspective matching degrees of the multiple free perspectives are determined respectively according to the perspective information of the target object in the video data streams of the multiple free perspectives obtained at the target moment; wherein the perspective matching degree is used to characterize the matching degree between the corresponding free perspective and the target free perspective.

[0121] a2. Selecting a target free viewing angle for photographing a target object from the multiple free viewing angles according to the viewing angle matching degrees of the multiple free viewing angles.

[0122] In some exemplary implementations, in the above step a1, for the multiple free-view video data streams acquired at the target moment, the following steps may be performed respectively:

[0123] b1. Performing face recognition on a target object in a specified free image frame of a video data stream to obtain a face confidence of the target object in the free image frame; wherein the target object in the free image frame holds a target object;

[0124] b2. determining the target distance between the target object and the free view camera that shoots the free image according to the depth map of the free image;

[0125] b3. Determine the perspective matching degree of the free perspective corresponding to a frame of free image according to the face confidence and the target distance.

[0126] In this embodiment, the target free view angle of the target object weighs the face confidence c of the target object at each free view angle and the target distance d of the target object from the free view angle camera at each free view angle. At each free view angle, a color image and a depth image of a frame of free image are acquired, and the face of the target object at the free view angle is identified using a face recognition algorithm based on the color image to obtain the face confidence c. At the same time, based on the depth image of the free view angle, the target distance d of the target object from the free view angle camera of the free view angle is calculated. The calculation formula of the matching degree v of each free view angle of the target object at the free view angle is as follows:

[0127] v=m×c+n×d

[0128] Among them, the larger the face confidence c, the better, and the smaller the target distance d, the better; m and n are setting coefficients, and m>n, and the specific values ​​of m and n can be set according to the scene.

[0129] By calculating the matching degree v of each free perspective, the free perspective with the largest matching degree v is obtained as the target free perspective.

[0130] Through the above implementation, the face confidence of the target object can be identified to represent the degree to which the face of the target object is facing the corresponding free-view camera. The face confidence is combined with the target distance between the target object and the corresponding free-view camera to determine the perspective matching degree of the corresponding free view.

[0131] Based on the above embodiment, after performing face recognition on a target object in a free image frame of a video data stream in step b1 and obtaining the face confidence of the target object in the free image frame, the following steps may be further performed:

[0132] If the face confidence is less than a preset threshold, the free view corresponding to a frame of free image is taken as a non-target free view.

[0133] It is understandable that when the face confidence of the target object detected in a certain free viewing angle is less than the preset threshold, the depth image in the free viewing angle will not be used to calculate the target distance of the target object from the free viewing angle camera of the free viewing angle, that is, the free viewing angle will not be selected.

[0134] In the above embodiment, if the face confidence of the target object under a certain free viewing angle is less than a preset threshold, it means that the free viewing angle is not suitable as the target free viewing angle for shooting the target object (which can be understood as the optimal viewing angle). The free viewing angle can be directly excluded without calculating the target distance between the target object and the corresponding free viewing angle camera.

[0135] Figure 2 A schematic diagram of the arrangement of telephoto zoom cameras in a specific scenario provided by an embodiment of the present application is shown.

[0136] Reference Figure 2 As shown, the specific scene takes a basketball game scene as an example, and the target object is a designated basketball star. In order to allow basketball fans to clearly see every front close-up picture of the designated basketball star holding the ball and attacking throughout the game, six telephoto zoom cameras are set at six positions on the basketball court (the number and positions of the telephoto zoom cameras can be set according to the actual scene). Every moment of the designated basketball star holding the ball and attacking can be zoomed in and out, satisfying the viewing experience of basketball fans for their favorite basketball stars.

[0137] When the game starts, after the server identifies the designated basketball star in multiple free perspectives, it tracks the designated basketball star in each free perspective. When the designated basketball star is not holding the ball, the server will not obtain the enlarged video stream shot by the telephoto zoom camera, that is, the audience will not see the front close-up picture of the designated basketball star. The server can recommend the best viewing angle to play the entire game.

[0138] When the designated basketball star holds the ball, the best free perspective (i.e., the target free perspective) for viewing the designated basketball star is selected from multiple free perspectives. Figure 2At the middle left half court position, the server takes the free view angle corresponding to the free view angle camera located at the bottom corner of the left half court as the best free view angle, and then controls the playback device to switch the screen. Since multiple free view angle cameras are distributed around the basketball court, in order to ensure the smoothness of the screen switching, the playback device can be controlled to switch from the current free view angle shooting screen to the best free view angle shooting screen clockwise or counterclockwise. At the same time, the telephoto zoom camera 1 located at the bottom line of the left half court will shoot the offensive close-up screen of the designated basketball star. When the playback device switches to the shooting screen of the best free view angle, it will switch to the close-up screen shot by the telephoto zoom camera 3 that is closest to the best free view angle.

[0139] When the designated basketball star holds the ball in Figure 2 At position two in the middle left half of the court, the designated basketball star chooses to play a back-to-the-basket singles after holding the ball (with his face facing the telephoto zoom camera 2). Similar to the above process, the server uses the free viewpoint corresponding to the free viewpoint camera located at the top corner of the left half of the court as the best free viewpoint, controls the playback device to switch the playback screen to the shooting screen of the best free viewpoint, and then switches to the front close-up picture of the back-to-the-basket singles shot by the telephoto zoom camera 2; after the designated basketball star plays a back-to-the-basket singles, he suddenly chooses to turn around and break through the baseline for a slam dunk. During this process, the best free viewpoint is switched to a free viewpoint at the baseline of the left half of the court, and finally the playback screen is switched to the front close-up picture of the designated basketball star turning around and breaking through the baseline for a slam dunk shot by the telephoto zoom camera 1 at the baseline of the left half of the court.

[0140] Based on the above process, the audience can have a close-up front view of every ball-holding attack by a designated basketball star on the court, which greatly enhances the viewing experience.

[0141] Based on the same inventive concept, an embodiment of the present application provides an object close-up switching device in a specific scene, wherein a plurality of free-viewpoint cameras and a plurality of telephoto zoom cameras are arranged in a venue corresponding to the specific scene, and the specific scene includes a plurality of objects; the principle of solving the problem by the device is similar to the method of the above-mentioned embodiment, and therefore the implementation of the device can refer to the implementation of the method, and the repeated parts will not be repeated.

[0142] Reference Figure 3 As shown, an object close-up switching device in a specific scenario provided by an embodiment of the present application includes:

[0143] The object recognition module 31 is used to process the video data streams of multiple free viewpoints shot by multiple free viewpoint cameras to recognize the target objects under the multiple free viewpoints;

[0144] An object tracking module 32, configured to track a target object in a plurality of free-viewing angles of view based on identification information of the target object in the plurality of free-viewing angles of view;

[0145] A perspective selection module 33 is used to select a target free perspective for displaying the target object from a plurality of free perspectives according to a plurality of free perspective video data streams acquired at the target moment if it is detected that the target object holds the target object at the target moment;

[0146] The control switching module 34 is used to control the playback device to switch the playback screen to the shooting screen of the target free viewing angle, and then switch to the close-up screen of the target object shot by the telephoto zoom camera corresponding to the target free viewing angle.

[0147] In some exemplary embodiments, the object recognition module 31 further includes:

[0148] The detection submodule is used to perform object detection on each free image frame of a free-viewing angle video data stream, and obtain a plurality of object detection frames corresponding to each free image frame;

[0149] A first recognition submodule is used to recognize each object detection frame in the obtained multiple object detection frames to obtain recognition information of each object detection frame;

[0150] The matching submodule is used to match the identification information of each object detection frame with the identification information of the target object in a preset object recognition information library to determine the target object detection frame where the identification information of the target object is located, and the target object detection frame represents the target object.

[0151] In some exemplary embodiments, the object recognition module 31 further includes a second recognition submodule, which is used to:

[0152] According to a plurality of object detection frames corresponding to each frame of free image under a free viewing angle, object recognition is performed on each frame of enlarged image in an enlarged video data stream shot by a telephoto zoom camera corresponding to a free viewing angle, so as to obtain a plurality of object recognition information of each enlarged image in each frame of enlarged image; wherein the telephoto zoom camera is pre-zoomed to a set focal length for shooting the enlarged video data stream;

[0153] The first identification submodule is further used for:

[0154] Respectively identifying multiple object detection frames corresponding to each free image frame in each free image frame to obtain candidate recognition information of multiple object detection frames of each free image frame;

[0155] The candidate recognition information of the multiple object detection frames of each frame of the free image is combined with the corresponding multiple object recognition information of each frame of the enlarged image to obtain the recognition information of the multiple object detection frames of each frame of the free image.

[0156] In some exemplary embodiments, the perspective selection module 33 includes:

[0157] A determination submodule, for determining the view matching degrees of the multiple free view angles respectively according to the view information of the target object in the video data streams of the multiple free view angles acquired at the target time if the target object is detected to hold the target object at the target time; wherein the view matching degree is used to characterize the matching degree between the corresponding free view angle and the target free view angle;

[0158] The selection submodule is used to select a target free viewing angle for photographing a target object from a plurality of free viewing angles according to the viewing angle matching degrees of the plurality of free viewing angles.

[0159] In some exemplary embodiments, the determining submodule is further configured to:

[0160] For multiple free-view video data streams acquired at the target time, perform the following operations respectively:

[0161] Performing face recognition on a target object in a free image frame of a video data stream to obtain a face confidence of the target object in the free image frame; wherein the target object in the free image frame holds a target object;

[0162] Determine, based on a depth map of a frame of free image, a target distance between a target object and a free view camera that shoots the frame of free image;

[0163] According to the face confidence and the target distance, the perspective matching degree of the free perspective corresponding to a frame of free image is determined.

[0164] In some exemplary embodiments, the determining submodule is further configured to:

[0165] If the face confidence is less than a preset threshold, the free view corresponding to a frame of free image is taken as a non-target free view.

[0166] In some exemplary embodiments, the telephoto zoom camera corresponding to the target free viewing angle is:

[0167] The telephoto zoom camera that is closest to the free view camera corresponding to the target free view; or

[0168] The telephoto zoom camera bound to the free view camera corresponding to the target free view.

[0169] Based on the same inventive concept, the embodiment of the present application provides an electronic device, which can be a server or a terminal device, which is not limited here. The principle of solving the problem by the electronic device is similar to the method of the above embodiment, so the implementation of the electronic device can refer to the implementation of the method, and the repeated parts will not be repeated.

[0170] like Figure 4As shown, the electronic device includes: a processor 41, a communication interface 43, a memory 42 and a communication bus 44, wherein the processor 41, the communication interface 43, and the memory 42 communicate with each other through the communication bus 44;

[0171] The memory 42 stores a computer program. When the program is executed by the processor 41, the processor 41 executes the object close-up switching method in any specific scene of the above-mentioned embodiments.

[0172] The above-mentioned processor can be a general-purpose processor, including a central processing unit, a network processor (Network Processor, NP), etc.; it can also be a digital signal processing processor (Digital Signal Processing, DSP), an application-specific integrated circuit, a field programmable gate array or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, etc.

[0173] The communication bus mentioned in the above electronic device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.

[0174] The communication interface is used for communication between the above electronic device and other devices.

[0175] The memory may include a random access memory (RAM) or a non-volatile memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.

[0176] An embodiment of the present application provides a computer-readable storage medium, which stores computer-executable instructions, and the computer-executable instructions are used to enable a computer to execute the object close-up switching method in any specific scenario of the above embodiments.

[0177] The computer-readable storage medium in the above embodiments can be any available medium or data storage device that can be accessed by the processor in the electronic device, including but not limited to magnetic storage such as floppy disks, hard disks, magnetic tapes, magneto-optical disks (MO), etc., optical storage such as CDs, DVDs, BDs, HVDs, etc., and semiconductor storage such as ROM, EPROM, EEPROM, non-volatile memory (NAND FLASH), solid-state drives (SSD), etc.

[0178] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.

[0179] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0180] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0181] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0182] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.

Claims

1. A method for switching close-up of an object in a specific scene, characterized in that: A plurality of free-viewing angle cameras and a plurality of telephoto zoom cameras are arranged in a venue corresponding to the specific scene, and the specific scene includes a plurality of objects. The method includes: Processing the video data streams of the multiple free-viewing angles captured by the multiple free-viewing angle cameras to identify the target objects in the multiple free-viewing angles; Tracking the target object in the video data streams of the multiple free perspectives based on identification information of the target object under the multiple free perspectives; If it is detected at the target moment that the target object holds the target object, then the perspective matching degrees of the multiple free perspectives are respectively determined according to the perspective information of the target object in the video data streams of the multiple free perspectives acquired at the target moment; wherein the perspective matching degrees are used to characterize the matching degree between the corresponding free perspective and the target free perspective; According to the viewpoint matching degrees of the multiple free viewpoints, selecting a target free viewpoint for displaying the target object from the multiple free viewpoints; After controlling the playback device to switch the playback screen to the shooting screen of the target free viewing angle, it switches to the close-up screen of the target object shot by the telephoto zoom camera corresponding to the target free viewing angle.

2. The method according to claim 1, characterized in that The processing of the video data streams of the multiple free-viewing angles captured by the multiple free-viewing angle cameras to identify the target objects in the multiple free-viewing angles includes: For the multiple free-viewing angle video data streams, the following operations are performed respectively: Performing object detection on each free image frame of a free-viewing angle video data stream to obtain a plurality of object detection frames corresponding to each free image frame; Identify each object detection frame in the obtained multiple object detection frames to obtain identification information of each object detection frame; The identification information of each object detection frame is matched with the identification information of the target object in a preset object recognition information library to determine a target object detection frame where the identification information of the target object is located, and the target object detection frame represents the target object.

3. The method according to claim 2, characterized in that After performing object detection on each free image frame of a free-viewing angle video data stream to obtain a plurality of object detection frames corresponding to each free image frame, the method further includes: Performing object recognition on each magnified image frame of a plurality of magnified video data streams captured by the plurality of telephoto zoom cameras, respectively, to obtain a plurality of object recognition information of each magnified image frame in each magnified image frame; wherein the telephoto zoom camera is pre-zoomed to a set focal length for capturing the magnified video data stream; Respectively identifying multiple object detection frames corresponding to each free image frame in each free image frame to obtain identification information of multiple object detection frames of each free image frame, including: Respectively identifying a plurality of object detection frames corresponding to each free image frame in each free image frame to obtain candidate recognition information of the plurality of object detection frames of each free image frame; The candidate recognition information of the multiple object detection frames of each free image frame is combined with the multiple object recognition information of a corresponding frame of enlarged image taken by each of the multiple telephoto zoom cameras to obtain the recognition information of the multiple object detection frames of each free image frame.

4. The method according to any one of claims 1 to 3, characterized in that: Determining the viewing angle matching degrees of the plurality of free viewing angles respectively according to the viewing angle information of the target object in the video data streams of the plurality of free viewing angles acquired at the target moment includes: For the multiple free-viewing angle video data streams acquired at the target moment, the following operations are performed respectively: Performing face recognition on the target object in a free image frame of a video data stream to obtain a face confidence of the target object in the free image frame; wherein the target object in the free image frame holds a target object; Determining, according to the depth map of the one frame of free image, a target distance between the target object and a free-viewing angle camera that shoots the one frame of free image; The viewing angle matching degree of the free viewing angle corresponding to the one frame of free image is determined according to the face confidence and the target distance.

5. The method according to claim 4, characterized in that After performing face recognition on the target object in a free image frame of a video data stream to obtain the face confidence of the target object in the free image frame, the method further includes: If the face confidence is less than a preset threshold, the free viewing angle corresponding to the one frame of free image is used as a non-target free viewing angle.

6. The method according to any one of claims 1 to 3, characterized in that: The telephoto zoom camera corresponding to the target free viewing angle is: The telephoto zoom camera that is closest to the free view angle camera corresponding to the target free view angle; or The telephoto zoom camera is bound to the free view camera corresponding to the target free view.

7. A device for switching close-up of an object in a specific scene, characterized in that: A plurality of free-viewing angle cameras and a plurality of telephoto zoom cameras are arranged in a venue corresponding to the specific scene, and the specific scene includes a plurality of objects. The device includes: An object recognition module, used for processing the video data streams of the multiple free-viewing angles captured by the multiple free-viewing angle cameras to recognize target objects under the multiple free-viewing angles; An object tracking module, configured to track the target object in the video data streams of the multiple free perspectives based on identification information of the target object in the multiple free perspectives; A perspective selection module, for determining the perspective matching degrees of the multiple free perspectives respectively according to the perspective information of the target object in the video data streams of the multiple free perspectives acquired at the target moment if it is detected that the target object holds the target object at the target moment; wherein the perspective matching degree is used to characterize the matching degree between the corresponding free perspective and the target free perspective; and selecting a target free perspective for displaying the target object from the multiple free perspectives according to the perspective matching degrees of the multiple free perspectives; The control switching module is used to control the playback device to switch the playback screen to the shooting screen of the target free viewing angle, and then switch to the close-up picture of the target object shot by the telephoto zoom camera corresponding to the target free viewing angle.

8. An electronic device, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the computer program is executed by the processor, the processor implements the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored therein, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Multi-camera intelligent control method and device

    CN101572804A

  • Intelligent multi-target active tracking monitoring method and system

    CN105338248A