A vision-based positioning and shooting method and system

By performing character detection and lip feature information analysis on the video image frame sequence, real-time positioning and positioning shooting of speakers without increasing costs is achieved, and the problem of high cost using microphone arrays is solved.

CN114241570BActive Publication Date: 2025-06-10REMO TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111576107.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-21
Publication Date
2025-06-10
Estimated Expiration
2041-12-21

AI Technical Summary

Technical Problem

In the case of space limitations, how to achieve real-time positioning and positioning shooting of speakers without increasing costs, avoiding the use of expensive microphone arrays.

Method used

By detecting the video image frame sequence, obtaining face feature information, lip feature information and coordinate information, maintaining the character position sequence list, and using the lip feature information to determine whether the character is talking, thereby cropping or controlling the camera for close-up shooting.

Benefits of technology

It realizes real-time positioning and positioning shooting of speakers without increasing costs, improving the video shooting experience, and is especially suitable for space-constrained environments such as small conference rooms.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114241570B_ABST
    Figure CN114241570B_ABST
Patent Text Reader

Abstract

The present invention discloses a vision-based positioning and shooting method and system. The method includes: performing human detection on a video image frame sequence to obtain human detection information of each human in the current frame video image, where the human detection information includes: face feature information, lip feature information, and coordinate information; maintaining a human position sequence list according to the human detection information; determining whether a human is speaking based on the lip feature information in the human position sequence list; cropping a close-up shot of the speaking human from the current frame video image according to the coordinate information corresponding to the current frame video image, or controlling a camera to perform a close-up shot of the speaking human according to the coordinate information corresponding to the current frame video image. Without increasing costs, the present invention realizes real-time positioning of the speaker and positioning shooting under the condition of limited space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present invention relate to the field of shooting technology, and in particular, to a vision-based positioning shooting method and system. Background Art

[0002] In video conferencing or in a recording and broadcasting system, it is very helpful to locate the speaker in real time, then zoom in on the speaker for close-up shooting, or crop the close-up shot of the speaker from the video panoramic image. Currently, microphone array voice positioning is used for video positioning, but using a microphone array requires additional equipment, which is costly and increases the system complexity. For situations with limited space such as small meeting rooms or classrooms, the cost of using a microphone array is very high. Without increasing the cost, how to achieve real-time positioning of the speaker and perform positioning shooting in a space-limited situation is a technical problem that we urgently need to solve. Summary of the Invention

[0003] The present invention provides a vision-based positioning shooting method and system to achieve real-time positioning of the speaker and perform positioning shooting in a space-limited situation without increasing the cost.

[0004] In a first aspect, embodiments of the present invention provide a vision-based positioning shooting method, and the vision-based positioning shooting method includes:

[0005] Perform human detection on a video image frame sequence to obtain human detection information of each person in the current frame video image, where the human detection information includes: face feature information, lip feature information, and coordinate information;

[0006] Maintain a human position sequence list according to the human detection information;

[0007] Judge whether there is a person speaking according to the lip feature information in the human position sequence list;

[0008] Crop the close-up shot of the person speaking from the current frame video image according to the coordinate information corresponding to the current frame video image, or control the camera to perform close-up shooting on the person speaking according to the coordinate information corresponding to the current frame video image.

[0009] Wherein, the judging whether there is a person speaking according to the lip feature information in the human position sequence list includes:

[0010] Successively extract K frames backward from the current frame in the lip feature information of the human position sequence information corresponding to each person at an interval of N frames, and send the lip feature information of the K-frame video images extracted into a speech classifier to obtain the real-time score of each person;

[0011] Calculate the average value of the real-time scores calculated for each person in the first M times to obtain the speaking score of the current frame of video image. If the speaking score is greater than or equal to a preset threshold, it is determined that the person is speaking.

[0012] Among them, the person detection of the video image frame sequence to obtain the person detection information of each person in the current frame of video image is specifically:

[0013] Use a face detection algorithm to perform face detection on the people in the image frame sequence; obtain the face feature information of each person according to the face detection result; and extract the lip feature information from the face detection result;

[0014] Obtain the coordinate information of each person in the current frame of video image.

[0015] Among them, the extraction of the lip feature information from the face detection result includes:

[0016] Perform key point detection on the face in the face detection result to detect the lip position;

[0017] Taking the lip position as the center, crop an image block and scale it to a fixed size to obtain a lip image;

[0018] Input the lip image into a convolutional neural network to obtain the lip feature information.

[0019] Among them, the maintenance of the person position sequence table according to the person detection information is specifically:

[0020] Match the face feature information of the person detected in the current frame image with the face feature information of the person in the person position sequence table. If there is a matching person, update the corresponding position sequence information of the person; if there is no matching person, construct a new person position sequence information for the person; or

[0021] Match the coordinate information detected in the current frame image with the coordinate information of the person in the person position sequence table. If there is a matching person, update the corresponding position sequence information of the person; if there is no matching person, construct a new person position sequence information for the person.

[0022] In a second aspect, an embodiment of the present invention further provides a vision-based positioning shooting system, and the vision-based positioning shooting system includes:

[0023] A person detection module for performing person detection on a video image frame sequence to obtain the person detection information of each person in the current frame of video image, where the person detection information includes: face feature information, lip feature information, and coordinate information;

[0024] A sequence maintenance module, configured to maintain a person location sequence list according to the person detection information;

[0025] A judgment module, configured to judge whether a person is speaking according to the lip feature information in the person location sequence list;

[0026] A shooting module, configured to crop a close-up of the person who is speaking from the current frame video image according to the coordinate information corresponding to the current frame video image, or control a camera to perform a close-up shooting on the person who is speaking according to the coordinate information corresponding to the current frame video image.

[0027] Among them, the judgment module is specifically configured to:

[0028] Sequentially extract K frames backward from the current frame at an interval of N frames for the lip feature information in the person location sequence information corresponding to each person, and send the lip feature information of the K-frame video images extracted into a speech classifier to obtain the real-time score of each person;

[0029] Calculate the average value of the real-time scores calculated for the first M times of each person to obtain the speech score of the current frame video image. If the speech score is greater than or equal to a preset threshold, it is judged that the person is speaking.

[0030] Among them, the person detection module includes:

[0031] A face detection unit, configured to perform face detection on the persons in the image frame sequence by using a face detection algorithm; obtain the face feature information of each person according to the face detection result; and extract the lip feature information from the face detection result;

[0032] A coordinate acquisition unit, configured to acquire the coordinate information of each person in the current frame video image.

[0033] Among them, the extraction of the lip feature information from the face detection result includes:

[0034] Perform key point detection on the face in the face detection result to detect the lip position;

[0035] Taking the lip position as the center, crop an image block and scale it to a fixed size to obtain a lip image;

[0036] Input the lip image into a convolutional neural network to obtain the lip feature information.

[0037] Among them, the sequence maintenance module includes:

[0038] A matching unit, configured to match the facial feature information detected in the current frame image with the facial feature information of a person in the person position sequence list; or match the coordinate information detected in the current frame image with the coordinate information of a person in the person position sequence list;

[0039] An updating unit, configured to update the position sequence information corresponding to the person if the matching unit determines that there is a matching person;

[0040] A constructing unit, configured to construct a new person position sequence information for the person if the matching unit determines that there is no matching person.

[0041] The present invention performs face detection on the persons in the image frame sequence by using a face detection algorithm; maintains a person position sequence list according to the face detection results; wherein, a position sequence information is constructed for each person, and the position sequence information includes facial feature information, lip feature information, coordinate information and time stamps, and all the position sequence information constitutes the person position sequence list; determines whether a person is speaking according to the position sequence information; crops a close-up of the speaking person from the current frame video image according to the coordinate information, or controls a camera to perform a close-up shot of the speaking person according to the coordinate information. Without increasing costs, the present invention realizes real-time positioning of a speaker and performs positioning shooting under the condition of limited space. Description of the Drawings

[0042] Figure 1 is a method flow chart of a vision-based positioning shooting method provided in Embodiment 1 of the present invention;

[0043] Figure 2 is a method flow chart of another vision-based positioning shooting method provided in Embodiment 2 of the present invention;

[0044] Figure 3 is a sub-method flow chart of a vision-based positioning shooting method provided in Embodiment 2 of the present invention;

[0045] Figure 4 is a schematic flow chart of extracting lip feature information from face detection results provided in Embodiment 2 of the present invention;

[0046] Figure 5 is another sub-method flow chart of a vision-based positioning shooting method provided in Embodiment 2 of the present invention;

[0047] Figure 6 is a structural block diagram of a vision-based positioning shooting system provided in Embodiment 3 of the present invention;

[0048] Figure 7It is a structural block diagram of another vision-based positioning and shooting system provided by the fourth embodiment of the present invention. Detailed implementation manners

[0049] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. Additionally, it should be noted that for the sake of description, only the parts related to the present invention rather than all the structures are shown in the drawings.

[0050] Embodiment 1

[0051] Figure 1 It is a method flow chart of a vision-based positioning and shooting method provided by Embodiment 1 of the present invention. This embodiment is applicable to the situation of target tracking and close-up shooting. This method can be executed by a computer or a camera, and specifically includes the following steps:

[0052] Step 110: Perform person detection on the video image frame sequence to obtain the person detection information of each person in the current frame of the video image.

[0053] Among them, the person detection information includes: face feature information, lip feature information, and coordinate information. In this embodiment, a face detection algorithm is used to perform face detection on the persons in the video image frame sequence to obtain the face feature information and lip feature information of each person in the current frame, and the coordinate information of each person is obtained from the current frame of the video image.

[0054] Step S120: Maintain a person position sequence list according to the person detection information.

[0055] Among them, a person position sequence information is constructed for each person. The person position sequence information includes face feature information, lip feature information, coordinate information, and the time stamp of the current frame image. All the person position sequence information constitutes the person position sequence list. Preferably, maintaining the person position sequence list according to the person detection information detected in each frame of the video image specifically includes: matching the face feature information with the face feature information of the persons in the person position sequence list. If there is a matching person, the person position sequence information corresponding to this person is updated; if there is no matching person, a new person position sequence information is constructed for this person; or matching the coordinate information detected in the current frame image with the coordinate information of the persons in the person position sequence list. If there is a matching person, the position sequence information corresponding to this person is updated; if there is no matching person, a new person position sequence information is constructed for this person. Matching according to the face feature information has a better matching accuracy. However, for places with fixed seats, especially for places with fixed and neatly arranged seats, updating the person position sequence list according to the coordinate information has a smaller system load and a faster response speed.

[0056] Step S130: Determine whether a person is speaking according to the lip feature information in the person position sequence list.

[0057] Each person corresponds to a sequence of lip feature information in the person position sequence list, that is, a lip trajectory recording the lip change process. According to the sequence of lip feature information, it can be accurately determined whether the corresponding person is speaking.

[0058] Step S140: Crop a close-up shot of the speaking person from the current frame video image according to the coordinate information corresponding to the current frame video image, or control the camera to take a close-up shot of the speaking person according to the coordinate information corresponding to the current frame video image.

[0059] When it is determined that a person is speaking, a close-up shot of the corresponding person is cropped from the current frame video image according to the coordinate information corresponding to the speaking person in the current frame video image, or the camera is controlled to take a close-up shot of the corresponding person according to the coordinate information corresponding to the speaking person in the current frame video image, completing the automatic positioning and shooting of the speaker.

[0060] In this embodiment, by performing person detection on the video image frame sequence, the face feature information, lip feature information, and coordinate information of each person in the current frame video image are obtained. According to a sequence of lip features (i.e., lip trajectory) of the person, it can be accurately determined whether the person is speaking. If there is a person speaking, the close-up shot of the speaking person is cropped from the video image or the camera is controlled to take a close-up shot of the speaking person according to the person coordinate information obtained from the current frame video image. Without increasing costs, this embodiment realizes real-time positioning of the speaker and positioning shooting under limited space conditions.

[0061] Embodiment 2

[0062] Figure 2 The flowchart of another vision-based positioning and shooting method provided by Embodiment 2 of the present invention. This embodiment is applicable to the case of target tracking and close-up shooting. This method can be executed by a computer or a camera, and specifically includes the following steps:

[0063] Step 210: Use a face detection algorithm to perform face detection on the persons in the image frame sequence; obtain the face feature information of each person according to the face detection result; and extract the lip feature information from the face detection result.

[0064] Use a face detection algorithm to perform face detection on the persons in the image frame sequence to obtain the face feature information and lip feature information of each person in the current frame. In some embodiments, such asFigure 3 As shown, extracting lip feature information from the face detection result includes steps S211 to S213, and the specific content is as follows:

[0065] Step S211: Detect key points of the face in the face detection result to detect the lip position.

[0066] Step S212: Taking the lip position as the center, crop an image patch and scale it to a fixed size to obtain a lip image.

[0067] Step S213: Input the lip image into a convolutional neural network to obtain lip feature information.

[0068] Specifically, the process of extracting lip feature information from the face detection result is as Figure 4 shown. Face detection and the processing of lip feature information are close to real-time processing. A lightweight network architecture and fast image processing algorithms can be used, which is beneficial to improving the processing efficiency of the device and the response speed.

[0069] Step S220: Obtain the coordinate information of each person in the current frame of video image.

[0070] Among them, the person detection information includes: face feature information, lip feature information, and coordinate information. Use a face detection algorithm to detect the faces of the people in the video image frame sequence to obtain the face feature information and lip feature information of each person in the current frame, and obtain the coordinate information of each person from the current frame of video image.

[0071] Step S230: Maintain a person position sequence list according to the person detection information.

[0072] Among them, a person position sequence information is constructed for each person. The person position sequence information includes face feature information, lip feature information, coordinate information, and the timestamp of the current frame image. All the person position sequence information constitutes the person position sequence list. Preferably, the person position sequence list is maintained according to the person detection information detected in each frame of video image, which specifically includes: matching the face feature information with the face feature information of the person in the person position sequence list. If there is a matching person, the person position sequence information corresponding to this person is updated; if there is no matching person, a new person position sequence information is constructed for this person; or matching the coordinate information detected in the current frame image with the coordinate information of the person in the person position sequence list. If there is a matching person, the position sequence information corresponding to this person is updated; if there is no matching person, a new person position sequence information is constructed for this person. In this embodiment, matching according to the face feature information can improve the accuracy of detection, and further improve the accuracy of positioning the close-up shooting; for places with fixed seats, especially for places with fixed and neatly arranged seats, updating the person position sequence list according to the coordinate information has a small system load and a faster response speed on the premise of ensuring accuracy.

[0073] Step S240: Determine whether a person is speaking according to the lip feature information in the person position sequence list.

[0074] Each person corresponds to a sequence of lip feature information in the person position sequence list, that is, a lip trajectory recording the lip change process. According to the sequence of lip feature information, it can be accurately determined whether the corresponding person is speaking.

[0075] In some embodiments, as Figure 5 shown, step S240 specifically includes step S241 to step S242, and the specific content is as follows:

[0076] Step S241: Sequentially extract K frames backward from the current frame for the lip feature information in the person position sequence information corresponding to each person at an interval of N frames. Send the lip feature information corresponding to the extracted K-frame video images into the speech classifier to obtain the real-time score of each person.

[0077] The real-time score of the lip trajectory of each person is calculated every once in a while. Extract K frames backward from the current frame for the lip feature information at an interval of N frames. For the lip feature information corresponding to the extracted K-frame video images, splice these lip feature information in sequence and then send them into the speech classifier to obtain the real-time score of the lip trajectory of this person.

[0078] Step S242: Calculate the average value of the real-time scores calculated for each person in the first M times to obtain the speaking score of the current frame of video image. If the speaking score is greater than or equal to the preset threshold, it is determined that the person is speaking.

[0079] The lip trajectory of each person consists of a sequence of a series of real-time scores. The higher the value, the higher the possibility that the corresponding person is speaking. The average of the real-time scores calculated for each person in the first M times is used as the speaking score of the current frame of video image. If the speaking score of a person in the current frame of video image is greater than or equal to the preset threshold, it is determined that the person is speaking. The value of M can be the number of frames captured by the camera within 1 s. If the value of M is too small, the calculation load will be too large, affecting the system function; if the value of M is too large, the real-time performance will decline.

[0080] In some embodiments, the extraction of lip feature information can be processed by a lightweight processor for real-time people, while the calculation of the speaking score of people can be processed by a heavyweight processor for non-real-time tasks, so as to balance the calculation load and reduce the overall calculation latency. At the same time, this solution can effectively reduce the calculation load in actual operation.

[0081] Step S250: Crop the close-up shot of the person who is speaking from the current frame of video image according to the coordinate information corresponding to the current frame of video image, or control the camera to take a close-up shot of the person who is speaking according to the coordinate information corresponding to the current frame of video image.

[0082] When it is determined that someone is speaking, the close-up shot of the corresponding person is cropped from the current frame of video image according to the coordinate information corresponding to the person who is speaking in the current frame of video image, or the camera is controlled to take a close-up shot of the corresponding person according to the coordinate information corresponding to the person who is speaking in the current frame of video image, completing the automatic positioning shooting of the speaker. If the number of people who are speaking is greater than or equal to 2, the close-up shots of the corresponding people can be cropped respectively, or the close-up shot corresponding to the smallest bounding box containing all the people who are speaking can be cropped; or the cameras can be controlled to take close-up shots of the corresponding people respectively according to the coordinate information corresponding to the people in the current frame of video image, that is, one person corresponds to one camera, or the minimum abscissa, maximum abscissa, minimum ordinate, and maximum ordinate in the coordinate information of all the people who are speaking can be obtained, and the camera is controlled to take a close-up shot of all the people who are speaking according to the minimum abscissa, maximum abscissa, minimum ordinate, and maximum ordinate.

[0083] The technical solution of this embodiment can accurately determine whether a person is speaking by judging based on a series of lip features (i.e., lip trajectories) of the person. If it is determined that a person is speaking, a close-up of the speaking person is cropped from the current frame of the video image according to the person coordinate information obtained from the image, or the camera is controlled to take a close-up of the speaking person. This embodiment locates the speaker in the video through visual processing technology, eliminating the need for a microphone array in the video conferencing system, which is of great significance for small meeting rooms and low-cost meeting room solutions. It can have a positioning shooting function without increasing costs, improving the user experience.

[0084] Embodiment III

[0085] The vision-based positioning and shooting system provided by the embodiments of the present invention can execute the vision-based positioning and shooting method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.

[0086] Figure 6 FIG. 10 is a structural block diagram of a vision-based positioning and shooting system provided by Embodiment III of the present invention. This embodiment is applicable to the case of target tracking and close-up shooting. The system can be executed by a computer or a camera, and includes a person detection module 10, a sequence maintenance module 20, a judgment module 30, and a shooting module 40. The specific content is as follows:

[0087] The person detection module 10 is configured to detect persons in a video image frame sequence, and obtain person detection information of each person in the current frame of the video image, where the person detection information includes: face feature information, lip feature information, and coordinate information.

[0088] The sequence maintenance module 20 is configured to maintain a person position sequence list according to the person detection information.

[0089] The judgment module 30 is configured to judge whether a person is speaking according to the lip feature information in the person position sequence list.

[0090] The shooting module 40 is configured to crop a close-up of the speaking person from the current frame of the video image according to the coordinate information corresponding to the current frame of the video image, or control the camera to take a close-up of the speaking person according to the coordinate information corresponding to the current frame of the video image.

[0091] In this embodiment, by performing human detection on the video image frame sequence, the face feature information, lip feature information, and coordinate information of each person in the current frame of the video image are obtained. According to a series of lip features (i.e., lip trajectories) of the person, it is determined whether the person is speaking. It can accurately determine whether the person is speaking. If there is a person speaking, according to the person coordinate information obtained from the current frame of the video image, a close-up shot of the speaking person is cropped from the video image or the camera is controlled to take a close-up shot of the speaking person. Without increasing costs, this embodiment realizes real-time positioning of the speaker and performs positioning shooting under limited space conditions.

[0092] Embodiment 4

[0093] The vision-based positioning and shooting system provided by the embodiments of the present invention can execute the vision-based positioning and shooting method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.

[0094] Such as Figure 7 As shown in the structure block diagram of another vision-based positioning and shooting system provided by Embodiment 4 of the present invention, this embodiment is applicable to the case of target tracking and close-up shooting. This system can be executed by a computer or a camera, and includes a human detection module 10, a sequence maintenance module 20, a judgment module 30, and a shooting module 40. The specific content is as follows:

[0095] The human detection module 10 is used to perform human detection on the video image frame sequence and obtain the human detection information of each person in the current frame of the video image, where the human detection information includes: face feature information, lip feature information, and coordinate information.

[0096] In some embodiments, the human detection module 10 includes a face detection unit 11 and a coordinate acquisition unit 12. The specific content is as follows:

[0097] The face detection unit 11 is used to perform face detection on the people in the image frame sequence by using a face detection algorithm; obtain the face feature information of each person according to the face detection result; and extract the lip feature information from the face detection result.

[0098] Preferably, the extracting the lip feature information from the face detection result includes:

[0099] Performing key point detection on the face in the face detection result to detect the lip position;

[0100] Taking an image block with the lip position as the center, cropping it, and scaling it to a fixed size to obtain a lip image;

[0101] Inputting the lip image into a convolutional neural network to obtain the lip feature information.

[0102] A coordinate acquisition unit 12 is configured to acquire the coordinate information of each person in the current frame of video image.

[0103] A sequence maintenance module 20 is configured to maintain a person position sequence list according to the person detection information.

[0104] In some embodiments, the sequence maintenance module 20 includes a matching unit 21, an updating unit 22, and a constructing unit 23. The specific details are as follows:

[0105] The matching unit 21 is configured to match the face feature information detected in the current frame image with the face feature information of the person in the person position sequence list; or match the coordinate information detected in the current frame image with the coordinate information of the person in the person position sequence list.

[0106] The updating unit 22 is configured to update the position sequence information corresponding to the person if the matching unit determines that there is a matching person.

[0107] The constructing unit 23 is configured to construct a new person position sequence information for the person if the matching unit determines that there is no matching person.

[0108] A judgment module 30 is configured to judge whether a person is speaking according to the lip feature information in the person position sequence list.

[0109] In some embodiments, the judgment module 30 is specifically configured to:

[0110] Successively extract K frames backward from the current frame at intervals of N frames for the lip feature information in the person position sequence information corresponding to each person, and send the lip feature information of the K-frame video images extracted into a speech classifier to obtain the real-time score of each person;

[0111] Calculate the average value of the real-time scores calculated for the previous M times of each person to obtain the speech score of the current frame of video image. If the speech score is greater than or equal to a preset threshold, it is determined that the person is speaking.

[0112] A shooting module 40 is configured to crop a close-up of the person speaking from the current frame of video image according to the coordinate information corresponding to the current frame of video image, or control a camera to take a close-up of the person speaking according to the coordinate information corresponding to the current frame of video image.

[0113] In this embodiment, by performing human detection on the video image frame sequence, the face feature information, lip feature information, and coordinate information of each person in the current frame of the video image are obtained. Whether a person is speaking is determined based on a sequence of lip features (i.e., lip trajectories) of the person, and it can accurately determine whether a person is speaking. If there is a person speaking, the close-up of the speaking person is cropped from the video image according to the person coordinate information obtained from the current frame of the video image, or the camera is controlled to perform a close-up shot of the speaking person. This embodiment locates the speaker in the video through visual processing technology to eliminate the microphone array of the video conferencing system, which is of great significance for small meeting rooms and low-cost meeting room solutions. It can have the function of positioning and shooting without increasing costs, improving the user experience.

[0114] Note that the above is only the preferred embodiment of the present invention and the applied technical principles. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein. Various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, more other equivalent embodiments can be included, and the scope of the present invention is determined by the scope of the appended claims.

Claims

1. A vision-based positioning and shooting method, characterized in that, the vision-based positioning and shooting method includes: Performing human detection on a video image frame sequence to obtain human detection information of each human in the current frame video image, where the human detection information includes: face feature information, lip feature information, and coordinate information; Maintaining a human position sequence list according to the human detection information; Judging whether a human is speaking according to the lip feature information in the human position sequence list; Cropping a close-up shot of the speaking human from the current frame video image according to the coordinate information corresponding to the current frame video image, or controlling a camera to perform a close-up shot of the speaking human according to the coordinate information corresponding to the current frame video image; Wherein, the maintaining of the human position sequence list according to the human detection information is specifically: Matching the face feature information detected in the current frame image with the face feature information of the humans in the human position sequence list. If there is a matching human, updating the position sequence information corresponding to the human; if there is no matching human, constructing a new human position sequence information for the human; or Matching the coordinate information detected in the current frame image with the coordinate information of the humans in the human position sequence list. If there is a matching human, updating the position sequence information corresponding to the human; if there is no matching human, constructing a new human position sequence information for the human.

2. The vision-based positioning and shooting method according to claim 1, characterized in that, the judging whether a human is speaking according to the lip feature information in the human position sequence list includes: Successively extracting K frames backward from the current frame at intervals of N frames for the lip feature information in the position sequence information corresponding to each human, and sending the lip feature information of the K-frame video images extracted into a speech classifier to obtain the real-time score of each human; Calculating the average value of the real-time scores calculated for the first M times of each human to obtain the speech score of the current frame video image. If the speech score is greater than or equal to a preset threshold, it is judged that the human is speaking.

3. The vision-based positioning and shooting method according to claim 1, characterized in that, the performing human detection on a video image frame sequence to obtain human detection information of each human in the current frame video image is specifically: Using a face detection algorithm to perform face detection on the humans in the image frame sequence; obtaining the face feature information of each human according to the face detection result; and extracting lip feature information from the face detection result; Obtaining the coordinate information of each human in the current frame video image.

4. The vision-based positioning and shooting method according to claim 3, characterized in that, the extracting lip feature information from the face detection result includes: Performing key point detection on the face in the face detection result to detect the lip position; Taking the lip position as the center, cropping an image block and scaling it to a fixed size to obtain a lip image; Inputting the lip image into a convolutional neural network to obtain lip feature information.

5. A vision-based positioning and shooting system, characterized in that, The vision-based positioning and shooting system includes: A person detection module, which is used to detect people in the video image frame sequence, and obtain the person detection information of each person in the current frame video image, where the person detection information includes: face feature information, lip feature information, and coordinate information; A sequence maintenance module, which is used to maintain a person position sequence list according to the person detection information; A judgment module, which is used to judge whether a person is speaking according to the lip feature information in the person position sequence list; A shooting module, which is used to crop the close-up of the person who is speaking from the current frame video image according to the coordinate information corresponding to the current frame video image, or control the camera to take a close-up of the person who is speaking according to the coordinate information corresponding to the current frame video image; Among them, the sequence maintenance module includes: A matching unit, which is used to match the face feature information detected in the current frame image with the face feature information of the person in the person position sequence list; or match the coordinate information detected in the current frame image with the coordinate information of the person in the person position sequence list; An update unit, which is used to update the position sequence information corresponding to the person if the matching unit determines that there is a matching person; A construction unit, which is used to construct a new person position sequence information for the person if the matching unit determines that there is no matching person.

6. The vision-based positioning and shooting system according to claim 5, wherein, The judgment module is specifically used for: Successively extract K frames backward from the current frame in the lip feature information of the person position sequence information corresponding to each person at an interval of N frames, and send the lip feature information of the K-frame video images extracted into a speech classifier to obtain the real-time score of each person; Calculate the average value of the real-time scores calculated for each person in the previous M times to obtain the speech score of the current frame video image. If the speech score is greater than or equal to a preset threshold, it is determined that the person is speaking.

7. The vision-based positioning and shooting system according to claim 5, wherein, The person detection module includes: A face detection unit, which is used to detect the faces of people in the image frame sequence by using a face detection algorithm; obtain the face feature information of each person according to the face detection result; and extract the lip feature information from the face detection result; A coordinate acquisition unit, which is used to acquire the coordinate information of each person in the current frame video image.

8. The vision-based positioning and shooting system according to claim 7, wherein, The extraction of the lip feature information from the face detection result includes: Perform key point detection on the face in the face detection result to detect the lip position; Take the lip position as the center, crop an image block, and scale it to a fixed size to obtain a lip image; Input the lip image into a convolutional neural network to obtain the lip feature information.

Citation Information

Patent Citations

  • Image processing method, apparatus and system

    CN109492506A