Screen positioning method, device, electronic device and storage medium
By setting multi-level confidence thresholds and region matching, the problem of inaccurate keyword recognition in video frames is solved, and the accuracy of picture positioning is improved, especially when keywords appear word by word or gradually.
Patent Information
- Application Number
- CN202310666895.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-07
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2043-06-07
AI Technical Summary
In the video frame, the keyword recognition accuracy is too low because the keywords appear from dark to bright or appear word by word, resulting in a large error in picture positioning.
By setting a larger first confidence threshold, the candidate start and end frames are determined, and secondary screening is performed before and after the candidate frames. The target forward and backward frames are screened using the second confidence threshold to ensure keyword confidence and regional consistency, thereby improving the accuracy of picture positioning.
It effectively reduces the possibility of misjudgment and missed judgment and improves the accuracy of picture positioning, especially when keywords appear one by one or from dark to bright.
Smart Images

Figure CN116758455B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image detection technology, and in particular to a picture positioning method, device, electronic equipment and storage medium. Background Art
[0002] In some scenarios, it is necessary to locate the screen where the keyword appears.
[0003] When locating scenes where keywords appear, the system typically determines whether the video frame contains the preset keywords and whether the keyword recognition confidence level is above a preset threshold to determine whether the video frame displays the preset keywords. However, because keywords in some scenes may appear from dark to bright or appear word by word, in these cases, the keywords in the video frame may be missing or appear lighter in color, resulting in low keyword recognition accuracy. This may result in scenes displaying the preset keywords being missed, leading to large errors in scene positioning. Summary of the Invention
[0004] The purpose of the embodiment of the present invention is to provide a screen positioning method to improve the accuracy of positioning the screen displaying preset keywords. The specific technical solution is as follows:
[0005] In a first aspect of the present invention, a method for positioning a picture is provided, the method comprising:
[0006] Determining the confidence level of the presence of a preset keyword in each video frame of the target video, wherein the confidence level is used to represent the possibility that the preset keyword is contained in the video frame;
[0007] Determine the start frame of each video frame whose confidence is greater than a preset first confidence threshold as a candidate start frame; determine the end frame of each video frame whose confidence is greater than the preset first confidence threshold as a candidate end frame;
[0008] Determine the first forward frame among the forward frames of the candidate start frame as the target forward frame; determine the first backward frame among the backward frames of the candidate end frame as the target backward frame; wherein the first forward frame is a forward frame whose target confidence is greater than the preset second confidence threshold, and the first backward frame is a backward frame whose target confidence is greater than the preset second confidence threshold; wherein the target confidence is the confidence of a preset keyword present in the candidate start frame and the candidate end frame, and the second confidence threshold is less than the first confidence threshold;
[0009] The target forward frame is determined as the target start frame for starting to display the preset keyword, and is used as the target start frame of the picture; the target backward frame is determined as the target end frame for ending to display the preset keyword, and is used as the target end frame of the picture.
[0010] In a possible embodiment, determining the first forward frame among the forward frames of the candidate start frame as the target forward frame includes:
[0011] Determining, in each of the forward frames, a first forward frame in which a difference between an area where the target keyword is located and an area where the target keyword is located in the candidate start frame is less than a preset threshold, as a target forward frame;
[0012] The determining the first backward frame among the backward frames of the candidate termination frame as the target backward frame includes:
[0013] The first backward frame in which the difference between the region where the target keyword is located in each backward frame and the region where the target keyword is located in the candidate termination frame is less than a preset threshold is determined as the target backward frame.
[0014] In a possible embodiment, determining the target forward frame as the target start frame for starting to display the preset keyword and using it as the target start frame of the picture includes:
[0015] If there are multiple target forward frames, determine the starting frame among the target forward frames as the target starting frame, and use it as the target starting frame of the picture;
[0016] The step of determining the target backward frame as the target termination frame for ending the display of the preset keyword and using it as the target termination frame of the picture includes:
[0017] If there are multiple target backward frames, the end frame in the target backward frames is determined as the target end frame, and is used as the target end frame of the picture.
[0018] In a possible embodiment, determining a start frame in the target forward frame as the target start frame and using it as the target start frame of the picture includes:
[0019] Selecting a second forward frame from a plurality of target forward frames of the candidate start frame, if the second backward frame of the candidate end frame satisfies a target confidence greater than the preset second confidence threshold, and the difference between the area where the target keyword is located and the area where the target keyword is located in the candidate start frame is less than the preset threshold, taking the second forward frame as a new candidate start frame, and returning to the step of determining the first forward frame among the forward frames of the candidate start frame as the target forward frame, until the forward frame of the candidate start frame does not satisfy the target confidence greater than the preset second confidence threshold, and the difference between the area where the target keyword is located and the area where the target keyword is located in the candidate start frame is less than the preset threshold;
[0020] Taking the candidate start frame as the target start frame, and taking it as the target start frame of the picture;
[0021] The step of determining the end frame in the target backward frame as the target end frame and using the end frame as the target end frame of the picture includes:
[0022] Selecting a second backward frame from a plurality of target backward frames of the candidate termination frame, and if the second backward frame of the candidate termination frame satisfies a target confidence greater than the preset second confidence threshold, and a difference between an area where the target keyword is located and an area where the target keyword is located in the candidate termination frame is less than a preset threshold, taking the second backward frame as a new candidate termination frame, and returning to the step of determining the first backward frame among the backward frames of the candidate termination frame as the target backward frame, until the target backward frame of the candidate termination frame does not satisfy a target confidence greater than the preset second confidence threshold, and a difference between an area where the target keyword is located and an area where the target keyword is located in the candidate termination frame is less than a preset threshold;
[0023] The candidate end frame is used as the target end frame, and the target end frame of the picture is used as the target end frame.
[0024] In a possible embodiment, determining the first forward frame among the forward frames of the candidate start frame as the target forward frame; and determining the first backward frame among the backward frames of the candidate end frame as the target backward frame include:
[0025] Determine a first forward frame among the forward frames of the candidate start frame as a target forward frame, and perform the first forward frame determination on the forward frames of the candidate start frame starting from the candidate start frame;
[0026] A first backward frame among the backward frames of the candidate termination frame is determined as a target backward frame, and the first backward frame determination is performed on the backward frames of the candidate termination frame starting from the candidate termination frame.
[0027] In a possible embodiment, the method further includes:
[0028] The obtained target start frame and the obtained target end frame are converted into a time format and output, a target time period is obtained based on the time of the target start frame and the time of the target end frame, and a preset process is performed within the time period including the target time period.
[0029] In a second aspect of the present invention, a picture positioning device is provided, comprising:
[0030] A first matching module is used to determine the confidence level of the presence of a preset keyword in each video frame of the target video, wherein the confidence level is used to represent the possibility that the preset keyword is contained in the video frame;
[0031] A first determining module is configured to determine a start frame in each video frame whose confidence is greater than a preset first confidence threshold as a candidate start frame; and determine an end frame in each video frame whose confidence of a preset keyword is greater than the preset first confidence threshold as a candidate end frame;
[0032] a second matching module, configured to determine a first forward frame among the forward frames of the candidate start frame as a target forward frame; and determine a first backward frame among the backward frames of the candidate end frame as a target backward frame; wherein the first forward frame is a forward frame whose target confidence is greater than the preset second confidence threshold, and the first backward frame is a backward frame whose target confidence is greater than the preset second confidence threshold; wherein the target confidence is a confidence of a preset keyword present in the candidate start frame and the candidate end frame, and the second confidence threshold is less than the first confidence threshold;
[0033] The second determining module is used to determine the target forward frame as the target start frame for starting to display the preset keyword, and use it as the target start frame of the picture; determine the target backward frame as the target end frame for ending to display the preset keyword, and use it as the target end frame of the picture.
[0034] In another aspect of the present invention, an electronic device is provided, comprising a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other via the communication bus;
[0035] Memory for storing computer programs;
[0036] The processor is configured to implement any of the above-mentioned steps of the screen positioning method when executing the program stored in the memory.
[0037] In another aspect of the present invention, a computer-readable storage medium is provided, wherein a computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, any of the above-mentioned steps of the picture positioning method is implemented.
[0038] In another aspect of the present invention, a computer-readable storage medium is provided, wherein a computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, any of the above-mentioned picture positioning methods is implemented.
[0039] In another aspect of the present invention, a computer program product comprising instructions is provided, which, when executed on a computer, enables the computer to execute any of the above-mentioned picture positioning methods.
[0040] The picture positioning method provided by the embodiment of the present invention determines the confidence level of the presence of preset keywords in each video frame of the target video, wherein the confidence level is used to characterize the possibility of the preset keywords being contained in the video frame; determines the starting frame of each video frame whose confidence level is greater than a preset first confidence threshold as a candidate starting frame; determines the ending frame of each video frame whose confidence level of the preset keywords is greater than the preset first confidence threshold as a candidate ending frame; determines the first forward frame of each forward frame of the candidate starting frame as the target forward frame; determines the first backward frame of each backward frame of the candidate ending frame as the target backward frame; wherein the first backward frame of each backward frame of the candidate ending frame is determined as the target backward frame. A forward frame is a forward frame whose target confidence is greater than the preset second confidence threshold, and a backward frame is a backward frame whose target confidence is greater than the preset second confidence threshold; wherein the target confidence is the confidence of the preset keyword present in the candidate start frame and the candidate end frame, and the second confidence threshold is less than the first confidence threshold; the target forward frame that meets the preset matching conditions is determined to be the target start frame for starting to display the preset keyword, and is used as the target start frame of the picture; the target backward frame that meets the preset matching conditions is determined to be the target end frame for ending to display the preset keyword, and is used as the target end frame of the picture. Applying the solution of the present application, by setting a larger first confidence threshold, some candidate start frames and candidate end frames with higher confidence are found. However, because displaying the "episode card" is a continuous process, the text will appear one by one or from dark to bright. Even if some video frames have lower confidence, they are still video frames that are currently displaying the "episode card". Therefore, a second confidence threshold is set to perform a secondary screening of the video frames before the candidate start frame and after the candidate end frame, reducing the possibility of misjudgment. At the same time, since the second confidence threshold is lower than the first confidence threshold, the possibility of missed judgment is also reduced, and the situation where words appear one by one or words appear from dark to bright can be better handled, thereby improving the accuracy of screen positioning of preset keywords. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for describing the embodiments or the prior art.
[0042] Figure 1 A schematic diagram of a flow chart of a picture positioning method provided by an embodiment of the present invention;
[0043] Figure 2 A schematic diagram of another flow chart of the image positioning method provided by an embodiment of the present invention;
[0044] Figure 3 A schematic structural diagram of a picture positioning device provided by an embodiment of the present invention;
[0045] Figure 4A schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0046] The technical solutions in the embodiments of the present invention will be described below with reference to the accompanying drawings in the embodiments of the present invention.
[0047] When inserting ads into a video for playback, they're typically placed before or after the episode card, at the end of the main film, or during the video. The episode card is the frame that displays the episode number, such as "Episode 1." When inserting ads after the episode card for playback, it's usually necessary to locate the episode card first.
[0048] When locating video episode cards, the system typically determines whether a video frame contains preset keywords and whether the keyword recognition confidence level is above a preset threshold to determine whether the video frame is an episode card. However, because some episode card keywords appear from bright to dark or word for word, in these cases, the episode card keywords may be missing or appear lighter in color, resulting in low keyword recognition accuracy. This can lead to episode cards being identified as non-episode cards, resulting in large errors in episode card location.
[0049] In order to improve the accuracy of positioning such a screen displaying keywords, the present invention provides a screen positioning method, device, electronic device and storage medium.
[0050] The image positioning method provided by the present invention can be applied to electronic devices, and the electronic devices can be servers, computers, mobile terminals, etc.
[0051] like Figure 1 As shown, Figure 1 A schematic flow chart of a method for positioning an image provided by an embodiment of the present invention may include the following steps:
[0052] Step S101: determining the confidence level of the presence of a preset keyword in each video frame of the target video, wherein the confidence level is used to represent the possibility of the preset keyword being contained in the video frame;
[0053] Step S102: determining a start frame in each video frame whose confidence is greater than a preset first confidence threshold as a candidate start frame; determining an end frame in each video frame whose confidence is greater than the preset first confidence threshold as a candidate end frame;
[0054] Step S103: Determine the first forward frame among the forward frames of the candidate start frame as the target forward frame; determine the first backward frame among the backward frames of the candidate end frame as the target backward frame; wherein the first forward frame is a forward frame whose target confidence is greater than the preset second confidence threshold, and the first backward frame is a backward frame whose target confidence is greater than the preset second confidence threshold; wherein the target confidence is the confidence of a preset keyword present in the candidate start frame and the candidate end frame, and the second confidence threshold is less than the first confidence threshold;
[0055] Step S104: determining the target forward frame as the target start frame for starting to display the preset keyword, and using it as the target start frame of the picture; determining the target backward frame as the target end frame for ending to display the preset keyword, and using it as the target end frame of the picture.
[0056] According to an embodiment of the present invention, the confidence level of the preset keywords in each video frame of the target video is determined, and the confidence level is used to characterize the possibility that the preset keywords are contained in the video frame. The starting frame in each video frame whose confidence level is greater than the preset first confidence threshold is determined as a candidate starting frame, and the ending frame in each video frame whose confidence level of the preset keywords is greater than the preset first confidence threshold is determined as a candidate ending frame. The first forward frame and the first backward frame in each forward frame of the candidate starting frame and the first backward frame in each backward frame of the candidate ending frame whose target confidence level is greater than the preset second confidence threshold are determined as target forward frames and target backward frames. The target confidence level is the confidence level of the preset keywords in the candidate starting frame and the candidate ending frame, and the second confidence threshold is less than the first confidence threshold. The target forward frame is determined to be the target starting frame for starting to display the preset keywords, and is used as the target starting frame of the picture. The target backward frame is determined to be the target ending frame for ending the display of the preset keywords, and is used as the target ending frame of the picture. By applying the solution of the present application, a larger first confidence threshold is set to find some candidate start frames and candidate end frames with higher confidence. However, because displaying the "episode card" is a continuous process, the text will appear one by one or from dark to bright. Even if some video frames have lower confidence, they are still video frames that are displaying the "episode card". Therefore, a second confidence threshold is set to perform a secondary screening of the video frames before the candidate start frame and after the candidate end frame, thereby reducing the possibility of misjudgment. At the same time, since the second confidence threshold is lower than the first confidence threshold, the possibility of missed judgment is also reduced, and the situation where text appears one by one or text appears from dark to bright can be better handled, thereby improving the accuracy of the screen positioning of the preset keywords.
[0057] The above steps S101-S104 are exemplarily described below:
[0058] In step S101, the target video may be any video that needs to be positioned. The target video may be a video showing the title sequence number, such as "Episode 1" or "Episode 2," or a video showing the ending credits.
[0059] In film and television drama videos, unified keywords are usually used when displaying information such as episode numbers and cast lists. For example, when displaying episode number information, the display method of "first episode" or "episode 1" is usually adopted, that is, the display method of "episode x" is adopted, and "x" represents text or numbers, etc. When displaying cast information, the word "cast" will also be displayed on the screen. Therefore, before locating these pictures, keywords can be set in advance. When locating the picture, the picture can be detected according to the preset keywords, and the positioning result can be obtained based on the detection result. The above preset keywords can be set according to actual needs, such as setting specific "first episode", "episode 2", etc., or setting keywords such as "episode", "episode" or "episode x".
[0060] After obtaining the target video, the target video can be subjected to frame extraction to obtain the video frames of the target video. In one embodiment, the target video as a whole can be subjected to frame extraction at the second level to obtain the respective video frames. Frame extraction at the second level means extracting a frame once per second of the video. Of course, since the appearance time of the opening and ending credits is usually relatively fixed, in one embodiment, frame extraction at the second level can also be performed on the video frames in a preset time period of the target video. The preset time period can be set according to actual conditions, such as the first five minutes of the video, the last five minutes of the video, and so on.
[0061] In step S101, when determining the confidence that a preset keyword exists in each video frame of the target video, text detection can be performed on each of the above video frames to obtain the content, location, and confidence information of the text contained in each of the above video frames. The detected text content is matched with the preset keyword, and then the confidence that the preset keyword exists in each video frame is obtained. The confidence represents the possibility that the preset keyword is contained in the video frame. When matching each video frame according to the preset keyword to determine the confidence that a preset keyword exists in each video frame of the target video, any feasible keyword matching method can be used, and the present invention does not make specific limitations on this. Exemplarily, the text content detected from the video frame can be converted into a vector form, and the similarity between it and the text vector corresponding to the preset keyword is calculated, thereby obtaining a keyword matching result between the text in the video and the preset keyword, and obtaining the confidence that the keyword exists.
[0062] Of course, you can also use regular matching to obtain the regular expression of the text detected from the video, and detect whether there are preset keywords in the regular expression of the video frame text, such as the regular expression of "Episode x", to obtain the keyword matching results of the text in the video frame and the preset keywords, and obtain the confidence level of the existence of the keywords.
[0063] In step S102, the starting frame of each video frame whose confidence is greater than the preset first confidence threshold is determined as a candidate starting frame, and the ending frame of each video frame whose preset keyword confidence is greater than the preset first confidence threshold is determined as a candidate ending frame. If the above-mentioned keyword confidence is greater than the preset first confidence threshold, it can be considered that the preset keyword exists in the video frame. The above-mentioned preset first confidence threshold can be set according to actual needs, such as being set to 0.9, 0.85, etc. Each video frame usually includes a timestamp to mark the appearance time of the video frame in the target video. Therefore, in step S102, the video frame with the earliest appearance time among the successfully matched video frames can be used as a candidate starting frame, and the video frame with the latest appearance time can be used as a candidate ending frame.
[0064] There are many ways for keywords to appear in film and television dramas, including simultaneous appearance, overall gradual appearance, and word-by-word appearance, etc. Except for simultaneous appearance, other keyword appearance methods will result in the inability to display complete text in the video frame, resulting in a low confidence level in text detection in the video frame, and thus causing such video frames containing keywords to be missed. Therefore, in an embodiment of the present invention, a broader condition can be set, and based on the above-mentioned preset keywords, the video frames near the above-mentioned candidate start frame and the candidate end frame are matched again. Missed detection usually means missing the video frame that appears before the candidate start frame and / or missing the video frame that appears after the candidate end frame. Therefore, based on the above-mentioned preset keywords, the forward frame of the above-mentioned candidate start frame and the backward frame of the candidate end frame can be matched twice.
[0065] As a specific implementation, in step S103, the forward frame of the candidate starting frame and the backward frame of the candidate ending frame can be matched twice based on the preset keywords. Since the keywords may appear gradually as a whole, in this case, the overall gradual appearance of the keywords will cause the keyword confidence to be greater than the above-mentioned first preset confidence threshold in the video frame, and the keyword confidence will first increase and then decrease as the video frame timestamp increases. That is to say, if the target forward frame is the forward frame of the candidate starting frame, and the target backward frame is the backward frame of the candidate ending frame, then the target confidence in the target forward frame and the target backward frame must be less than the above-mentioned first confidence threshold. Therefore, in order to avoid false detection, a second confidence threshold can be set. The target confidence must be greater than the second confidence threshold to be determined as the target forward frame and the target backward frame. The second confidence threshold is less than the above-mentioned first confidence threshold, such as 0.4, 0.5, etc.
[0066] In one possible embodiment, the keyword may appear word for word. In this case, determining the first forward frame among the forward frames of the candidate start frame as the target forward frame includes: determining the first forward frame in each forward frame where the difference between the area where the target keyword is located and the area where the target keyword is located in the candidate start frame is less than a preset threshold, as the target forward frame. Determining the first backward frame among the backward frames of the candidate end frame as the target backward frame includes: determining the first backward frame in each backward frame where the difference between the area where the target keyword is located and the area where the target keyword is located in the candidate end frame is less than a preset threshold, as the target backward frame.
[0067] Determine that the first forward frame in each forward frame is a forward frame whose target confidence is greater than a preset second confidence threshold, and the difference between the area where the target keyword is located and the area where the target keyword is located in the candidate start frame is less than a preset threshold. This means that there is a target keyword consistent with the preset keyword in the first forward frame, and the difference between the area where the target keyword is located and the area where the keyword in the candidate start frame is located is less than the preset threshold, which means that the forward frame may have begun to display part of the content of the displayed keyword in the above-mentioned candidate start frame. Therefore, the first forward frame can be determined as the target forward frame. Determine that the first backward frame in each backward frame is a backward frame whose target confidence is greater than a preset second confidence threshold, and the difference between the area where the target keyword is located and the area where the target keyword is located in the candidate end frame is less than a preset threshold. This indicates that there is a target keyword that is consistent with the preset keyword in the first backward frame, and the difference between the area where the target keyword is located and the area where the keyword in the candidate termination frame is located is less than a preset threshold, which indicates that the backward frame may still be displaying part of the content of the keyword displayed in the above-mentioned candidate termination frame. Therefore, the first backward frame can be determined as the target backward frame. Applying the solution of the embodiment of the present application, the target forward frame and the target backward frame are determined based on the keyword confidence and the area where the keyword is located, rather than only based on the keyword confidence, which can better handle the situation where words appear one by one and improve the accuracy of keyword screen positioning.
[0068] In a possible embodiment, the first forward frame among the forward frames of the candidate start frame is determined as the target forward frame, and the first forward frame determination is performed on the forward frames of the candidate start frame starting from the candidate start frame. The first backward frame among the backward frames of the candidate end frame is determined as the target backward frame, and the first backward frame determination is performed on the backward frames of the candidate end frame starting from the candidate end frame.
[0069] The above-mentioned first forward frame determination refers to determining whether a video frame is the first forward frame. As described above, the first forward frame refers to a video frame whose target confidence is greater than a preset second confidence threshold. That is, here it is to determine whether the forward frame of the candidate starting frame satisfies the target confidence greater than the preset second confidence threshold. In a common method, each video frame of the target video is matched to obtain the confidence of the preset keyword. When obtaining the first forward frame, it is possible to start from the first frame of the target video and match all the forward frames of the above-mentioned candidate starting frame, or to match some of the forward frames of the candidate starting frame. Applying the solution of this embodiment, it is possible to start from the candidate starting frame and perform a first forward frame determination on the forward frame of the candidate starting frame to obtain the target forward frame more quickly. Correspondingly, when matching the backward frame of the above-mentioned candidate terminating frame, it is possible to start from the candidate terminating frame and perform a first backward frame determination on the backward frame of the candidate terminating frame to obtain the target backward frame more quickly, thereby improving the rate of keyword screen positioning.
[0070] In step S104, the target forward frame is determined to be the target start frame for starting to display the preset keyword, and is used as the target start frame of the picture; the target backward frame is determined to be the target end frame for ending to display the preset keyword, and is used as the target end frame of the picture.
[0071] The target starting frame of the above picture is the first frame of the picture to be located, and the target ending frame of the picture is the last frame of the picture to be located, that is, the first frame and the last frame of the picture showing the "episode card" to be located. Based on this, the position of the desired video frame in the picture can be obtained, thereby realizing picture positioning.
[0072] After determining the target start frame and target end frame, advertisement video insertion, related video recommendation, etc. can be performed based on the target start frame and target end frame. Taking advertisement insertion as an example, the advertisement video can be inserted before the target start frame or after the target end frame.
[0073] There may be multiple target forward frames and / or target backward frames that meet the target confidence greater than the preset second confidence threshold, and the difference between the area where the target keyword is located and the area where the target keyword is located in the candidate starting frame is less than the preset threshold.
[0074] In a possible embodiment, if there are multiple target forward frames, the starting frame in the target forward frames is determined as the target starting frame, and is used as the starting frame of the picture; if there are multiple target backward frames, the ending frame in the target backward frames is determined as the target ending frame, and is used as the ending frame of the picture.
[0075] By applying the solution of the embodiment of the present application, the frontmost frame among multiple target forward frames that meet the requirements of keyword confidence and keyword area is determined as the target start frame, and the backmost frame among the target backward frames is determined as the target end frame, and they are used as the first frame and the last frame of the picture to be located, so that the located picture containing the keywords is more complete, thereby improving the accuracy of keyword picture positioning.
[0076] In one possible embodiment, determining the starting frame in the target forward frame as the target starting frame and using it as the target starting frame of the picture includes: selecting a second forward frame from multiple target forward frames of the candidate starting frame, if the second backward frame of the candidate ending frame satisfies the target confidence greater than the preset second confidence threshold, and the difference between the area where the target keyword is located and the area where the target keyword is located in the candidate starting frame is less than the preset threshold, using the second forward frame as the new candidate starting frame, and returning to the step of determining the first forward frame in each forward frame of the candidate starting frame as the target forward frame, until the forward frame of the candidate starting frame does not satisfy the target confidence greater than the preset second confidence threshold, and the difference between the area where the target keyword is located and the area where the target keyword is located in the candidate starting frame is less than the preset threshold. Using the candidate starting frame as the target starting frame, and using it as the target starting frame of the picture.
[0077] Determine the terminating frame in the target backward frame as the target terminating frame and use it as the target terminating frame of the picture, including: selecting a second backward frame from multiple target backward frames of the candidate terminating frame, if the second backward frame of the candidate terminating frame satisfies the target confidence greater than the preset second confidence threshold, and the difference between the area where the target keyword is located and the area where the target keyword is located in the candidate terminating frame is less than the preset threshold, then use the second backward frame as the new candidate terminating frame, and return to the step of determining the first backward frame in each backward frame of the candidate terminating frame as the target backward frame, until the target backward frame of the candidate terminating frame does not satisfy the target confidence greater than the preset second confidence threshold, and the difference between the area where the target keyword is located and the area where the target keyword is located in the candidate terminating frame is less than the preset threshold. Use the candidate terminating frame as the target terminating frame and use it as the target terminating frame of the picture.
[0078] That is to say, if the continuous forward frames of candidate start frame all meet the target confidence and are greater than the preset second confidence threshold, and the difference between the region where the target keyword resides and the region where the target keyword resides in the candidate start frame is less than the preset threshold, then the frontmost frame in the continuous forward frames can be determined as the target start frame. If the continuous backward frames of candidate termination frame all meet the target confidence and are greater than the preset second confidence threshold, and the difference between the region where the target keyword resides and the region where the target keyword resides in the candidate termination frame is less than the preset threshold, then the last frame in the continuous backward frames can be determined as the target termination frame. That is, search forward from selected candidate start frame always, search backward from selected candidate termination frame always, until video frame no longer meets the target confidence and is greater than the preset second confidence threshold, and the difference between the region where the target keyword resides and the region where the target keyword resides in the candidate termination frame is less than the condition of the preset threshold, stop circulation. Apply the scheme of the embodiment of the application, further improve the accuracy of keyword picture positioning.
[0079] In one possible embodiment, the obtained target start frame and target end frame are converted into a time format and output, a target time period is obtained based on the time of the target start frame and the time of the target end frame, and a preset process is performed within the time period including the target time period. Exemplarily, the preset process can be performed before, after, or within the target time period based on the obtained target time period. Exemplarily, the preset process can be inserting advertisements or making related video recommendations.
[0080] like Figure 2 As shown, Figure 2 Another flowchart of the image positioning method provided by the present invention may include the following steps:
[0081] Step 1: Extract frames from the TV series video to obtain second-level results.
[0082] Step ②: Perform text detection and recognition on each video frame obtained by extraction based on pre-set keywords to obtain episode card key frames.
[0083] The key frames of the episode card are the start and end frames in the video frames where there are preset keywords and the keyword confidence in the video frame is greater than the preset first confidence threshold, that is, the candidate start frame and candidate end frame in this article.
[0084] Step 3: According to the actual situation, according to the above-mentioned keyword word-by-word appearance logic or keyword overall gradual change logic, the gradual episode card points are obtained based on the above-mentioned episode card key frames.
[0085] The above-mentioned keyword word-by-word appearance logic or keyword overall gradual change logic is to determine the first forward frame in which the difference between the area where the target keyword is located in each forward frame and the area where the target keyword is located in the candidate start frame is less than a preset threshold, and use it as the target forward frame. The first backward frame in which the difference between the area where the target keyword is located in each backward frame and the area where the target keyword is located in the candidate end frame is less than a preset threshold is determined as the target backward frame. Among them, the first forward frame is a forward frame whose target confidence is greater than the preset second confidence threshold. The first backward frame is a backward frame whose target confidence is greater than the preset second confidence threshold.
[0086] By applying the embodiments of the present invention, the text content and the position of the text are used to effectively detect the episode cards that appear word by word, and the text confidence is used to effectively detect the episode cards with overall brightness changes, thereby improving the accuracy of episode card positioning, and thus improving the accuracy of advertising delivery.
[0087] In another aspect of the present invention, a picture positioning device is provided. Figure 3 As shown, the device may include:
[0088] A first matching module 301 is configured to determine a confidence level of the presence of a preset keyword in each video frame of a target video, wherein the confidence level is used to represent the likelihood that the video frame contains the preset keyword;
[0089] A first determining module 302 is configured to determine a start frame in each video frame whose confidence is greater than a preset first confidence threshold as a candidate start frame; and determine an end frame in each video frame whose confidence of a preset keyword is greater than the preset first confidence threshold as a candidate end frame;
[0090] The second matching module 303 is configured to determine the first forward frame among the forward frames of the candidate start frame as the target forward frame; and determine the first backward frame among the backward frames of the candidate end frame as the target backward frame; wherein the first forward frame is a forward frame whose target confidence is greater than the preset second confidence threshold, and the first backward frame is a backward frame whose target confidence is greater than the preset second confidence threshold; wherein the target confidence is the confidence of a preset keyword present in the candidate start frame and the candidate end frame, and the second confidence threshold is less than the first confidence threshold;
[0091] The second determining module 304 is configured to determine the target forward frame as the target start frame for starting to display the preset keyword, and use it as the target start frame of the picture; and determine the target backward frame as the target end frame for ending to display the preset keyword, and use it as the target end frame of the picture.
[0092] The embodiment of the present invention further provides an electronic device, such as Figure 4 As shown, it includes a processor 401, a communication interface 402, a memory 403 and a communication bus 404, wherein the processor 401, the communication interface 402, and the memory 403 communicate with each other through the communication bus 404.
[0093] Memory 403, used for storing computer programs;
[0094] The processor 401 is configured to execute the program stored in the memory 403 by performing the following steps:
[0095] Determining the confidence level of the presence of a preset keyword in each video frame of the target video, wherein the confidence level is used to represent the possibility that the preset keyword is contained in the video frame;
[0096] Determine the start frame of each video frame whose confidence is greater than a preset first confidence threshold as a candidate start frame; determine the end frame of each video frame whose confidence is greater than the preset first confidence threshold as a candidate end frame;
[0097] Determine the first forward frame among the forward frames of the candidate start frame as the target forward frame; determine the first backward frame among the backward frames of the candidate end frame as the target backward frame; wherein the first forward frame is a forward frame whose target confidence is greater than the preset second confidence threshold, and the first backward frame is a backward frame whose target confidence is greater than the preset second confidence threshold; wherein the target confidence is the confidence of a preset keyword present in the candidate start frame and the candidate end frame, and the second confidence threshold is less than the first confidence threshold;
[0098] The target forward frame is determined as the target start frame for starting to display the preset keyword, and is used as the target start frame of the picture; the target backward frame is determined as the target end frame for ending to display the preset keyword, and is used as the target end frame of the picture.
[0099] The communication bus mentioned in the terminal can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, only one thick line is used in the figure, but this does not mean that there is only one bus or only one type of bus.
[0100] The communication interface is used for communication between the above terminal and other devices.
[0101] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage. Alternatively, the memory may be at least one storage device located away from the processor.
[0102] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, and discrete hardware components.
[0103] In another embodiment of the present invention, a computer-readable storage medium is provided, wherein a computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, the screen positioning method described in any one of the above embodiments is implemented.
[0104] In another embodiment of the present invention, a computer program product including instructions is provided. When the computer program product is run on a computer, the computer executes the picture positioning method described in any one of the above embodiments.
[0105] In the above embodiments, all or part of the embodiments can be implemented by software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of the present invention are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).
[0106] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.
[0107] Each embodiment in this specification is described in a related manner. Similar portions between embodiments can be referenced to each other. Each embodiment focuses on the differences from other embodiments. In particular, the device, electronic device, storage medium, and program product embodiments are generally similar to the method embodiments, so their descriptions are relatively simplified. For related portions, reference can be made to the descriptions of the method embodiments.
[0108] The above description is only a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention are included in the scope of protection of the present invention.
Claims
1. A picture positioning method, characterized in that: The method comprises: Determining the confidence level of the presence of a preset keyword in each video frame of the target video, wherein the confidence level is used to represent the possibility that the preset keyword is contained in the video frame; Determining the confidence level of the presence of the preset keyword in each video frame of the target video includes: Performing text detection on each video frame of the target video to obtain text content, matching the text content with preset keywords, and obtaining a confidence level that the preset keywords exist in each video frame; Determine the start frame of each video frame whose confidence is greater than a preset first confidence threshold as a candidate start frame; determine the end frame of each video frame whose confidence is greater than the preset first confidence threshold as a candidate end frame; Determine a first forward frame among the forward frames of the candidate start frame as a target forward frame; determine a first backward frame among the backward frames of the candidate end frame as a target backward frame; wherein the first forward frame is a forward frame whose target confidence is greater than a preset second confidence threshold, and the first backward frame is a backward frame whose target confidence is greater than the preset second confidence threshold; wherein the target confidence is a confidence of a preset keyword present in the candidate start frame and the candidate end frame, and the second confidence threshold is less than the first confidence threshold; The target forward frame is determined as the target start frame for starting to display the preset keyword, and is used as the target start frame of the picture; the target backward frame is determined as the target end frame for ending to display the preset keyword, and is used as the target end frame of the picture.
2. The method according to claim 1, characterized in that The determining the first forward frame among the forward frames of the candidate start frame as the target forward frame includes: Determining, in each of the forward frames, a first forward frame in which a difference between an area where the target keyword is located and an area where the target keyword is located in the candidate start frame is less than a preset threshold, as a target forward frame; The determining the first backward frame among the backward frames of the candidate termination frame as the target backward frame includes: The first backward frame in which the difference between the region where the target keyword is located in each backward frame and the region where the target keyword is located in the candidate termination frame is less than a preset threshold is determined as the target backward frame.
3. The method according to claim 2, characterized in that The step of determining the target forward frame as the target start frame for starting to display the preset keyword and using it as the target start frame of the picture includes: If there are multiple target forward frames, determine the starting frame among the target forward frames as the target starting frame, and use it as the target starting frame of the picture; The step of determining the target backward frame as the target termination frame for ending the display of the preset keyword and using it as the target termination frame of the picture includes: If there are multiple target backward frames, the end frame in the target backward frames is determined as the target end frame, and is used as the target end frame of the picture.
4. The method according to claim 3, characterized in that The step of determining a starting frame in the target forward frame as the target starting frame and using the starting frame as the target starting frame of the picture includes: Selecting a second forward frame from a plurality of target forward frames of the candidate start frame, if the second backward frame of the candidate end frame satisfies a target confidence greater than the preset second confidence threshold, and the difference between the area where the target keyword is located and the area where the target keyword is located in the candidate start frame is less than the preset threshold, taking the second forward frame as a new candidate start frame, and returning to the step of determining the first forward frame among the forward frames of the candidate start frame as the target forward frame, until the forward frame of the candidate start frame does not satisfy the target confidence greater than the preset second confidence threshold, and the difference between the area where the target keyword is located and the area where the target keyword is located in the candidate start frame is less than the preset threshold; Taking the candidate start frame as the target start frame, and taking it as the target start frame of the picture; The step of determining the end frame in the target backward frame as the target end frame and using the end frame as the target end frame of the picture includes: Selecting a second backward frame from a plurality of target backward frames of the candidate termination frame, and if the second backward frame of the candidate termination frame satisfies a target confidence greater than the preset second confidence threshold, and a difference between an area where the target keyword is located and an area where the target keyword is located in the candidate termination frame is less than a preset threshold, taking the second backward frame as a new candidate termination frame, and returning to the step of determining the first backward frame among the backward frames of the candidate termination frame as the target backward frame, until the target backward frame of the candidate termination frame does not satisfy a target confidence greater than the preset second confidence threshold, and a difference between an area where the target keyword is located and an area where the target keyword is located in the candidate termination frame is less than a preset threshold; The candidate end frame is used as the target end frame, and the target end frame of the picture is used as the target end frame.
5. The method according to claim 1, wherein determining the first forward frame among the forward frames of the candidate start frame as the target forward frame; Determining the first backward frame among the backward frames of the candidate termination frame as the target backward frame includes: Determine a first forward frame among the forward frames of the candidate start frame as a target forward frame, and perform the first forward frame determination on the forward frames of the candidate start frame starting from the candidate start frame; A first backward frame among the backward frames of the candidate termination frame is determined as a target backward frame, and the first backward frame determination is performed on the backward frames of the candidate termination frame starting from the candidate termination frame.
6. The method according to claim 1, characterized in that The method further comprises: The obtained target start frame and the obtained target end frame are converted into a time format and output, a target time period is obtained based on the time of the target start frame and the time of the target end frame, and a preset process is performed within the time period including the target time period.
7. A picture positioning device, characterized in that: The device comprises: A first matching module is used to determine the confidence level of the presence of a preset keyword in each video frame of the target video, wherein the confidence level is used to represent the possibility that the preset keyword is contained in the video frame; The first matching module is specifically configured to perform text detection on each video frame of the target video to obtain text content, match the text content with preset keywords, and obtain a confidence level that the preset keywords exist in each video frame; A first determining module is configured to determine a start frame in each video frame whose confidence is greater than a preset first confidence threshold as a candidate start frame; and determine an end frame in each video frame whose confidence of a preset keyword is greater than the preset first confidence threshold as a candidate end frame; a second matching module, configured to determine a first forward frame among the forward frames of the candidate start frame as a target forward frame; and determine a first backward frame among the backward frames of the candidate end frame as a target backward frame; wherein the first forward frame is a forward frame whose target confidence is greater than a preset second confidence threshold, and the first backward frame is a backward frame whose target confidence is greater than the preset second confidence threshold; wherein the target confidence is a confidence of a preset keyword present in the candidate start frame and the candidate end frame, and the second confidence threshold is less than the first confidence threshold; The second determining module is used to determine the target forward frame as the target start frame for starting to display the preset keyword, and use it as the target start frame of the picture; determine the target backward frame as the target end frame for ending to display the preset keyword, and use it as the target end frame of the picture.
8. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus; Memory for storing computer programs; A processor, configured to implement the method according to any one of claims 1 to 6 when executing a program stored in a memory.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Method and device for determining key frame, storage medium and electronic equipment
CN114429606A
Video clip extraction method and device, electronic equipment and storage medium
CN115103225A