Video text recognition method and device, storage medium and electronic device
By integrating the destop words and editing distance of frame text subsets for video text, the problem of inaccurate edge text recognition in video frames is solved, and the accuracy of recognition of video text content is improved.
Patent Information
- Application Number
- CN202110462515.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-27
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2041-04-27
AI Technical Summary
In the prior art, the text content recognition accuracy in video frames is low, especially the text recognition of edge positions is not accurate enough, resulting in the problem of low recognition accuracy when recognizing video text content.
By obtaining a subset of frame text in the video text, removing text carrying stop words, calculating the editing distance between text segments in the frame text subset, and integrating candidate text based on the editing distance to improve the recognition accuracy of video text.
The fine processing of video text is realized, the recognition accuracy of video text content is improved, and the problem of low recognition accuracy caused by text extraction tools to ignore some text information in video frames is solved.
Smart Images

Figure CN113762038B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computers, and in particular to a video text recognition method and device, a storage medium and an electronic device. Background Art
[0002] Nowadays, many video playback platforms allow users to upload videos to be released. As a video content creator, in order to attract more users to watch, you often edit and create rich content in the video. In addition to providing intuitive video frames, it is more important to provide relevant text descriptions of the video content.
[0003] In order to facilitate video management, the background often needs to identify and analyze the text content in the video frame. At present, the common way to identify the text content in the above video frames is to rely on extraction tools such as optical character recognition (OCR). However, the extraction tools provided in the relevant technology have low recognition accuracy and often ignore some texts at relatively edge positions in the video frame, resulting in low recognition accuracy when identifying the content of the video text.
[0004] To address the above-mentioned problems, no effective solution has been proposed yet. Summary of the invention
[0005] Embodiments of the present invention provide a method and device for recognizing video text, a storage medium, and an electronic device, so as to at least solve the technical problem of low accuracy in content recognition of video text caused by text extraction tools ignoring part of text information in video frames.
[0006] According to one aspect of an embodiment of the present invention, a method for recognizing video text is provided, comprising: obtaining video text extracted from a target video to be recognized, wherein the video text includes frame text subsets corresponding to respective video frames of the target video; determining a target frame text subset carrying stop words from the video text; removing the stop words carried in the target frame text subset to update the video text to a candidate text; determining an edit distance between text segments in the frame text subsets corresponding to any two video frames in the candidate text; and integrating the candidate texts according to the edit distance to obtain a target text recognized for the target video.
[0007] According to another aspect of an embodiment of the present invention, a video text recognition device is also provided, including: a first acquisition unit, used to acquire video text extracted from a target video to be recognized, wherein the video text includes frame text subsets corresponding to each video frame of the target video; a first determination unit, used to determine a target frame text subset carrying stop words from the video text; a filtering unit, used to remove the stop words carried in the target frame text subset to update the video text to a candidate text; a second determination unit, used to determine the edit distance between text segments in the frame text subsets corresponding to any two video frames in the candidate text; and an integrated recognition unit, used to integrate the candidate texts according to the edit distance to obtain a target text recognized for the target video.
[0008] According to another aspect of the embodiments of the present invention, a computer-readable storage medium is provided, in which a computer program is stored, wherein the computer program is configured to execute the above-mentioned video text recognition method when running.
[0009] According to another aspect of an embodiment of the present invention, there is provided an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the video text recognition method through the computer program.
[0010] In an embodiment of the present invention, after obtaining the frame text subsets corresponding to each video frame in the video text extracted from the target video, the target frame text subset carrying stop words is determined from the above video text, and the stop words carried in the above target frame text subset are removed to update the above video text to the candidate text. Then the edit distance between the text fragments in the frame text subsets corresponding to any two video frames in the candidate text is determined, and the candidate texts are integrated according to the edit distance to obtain the target text finally recognized for the target video. That is to say, the text fragments in the frame text subset extracted by the extraction tool are filtered and integrated, the stop words are removed, and the text fragments in the frame text subset in the candidate text filtered out of the stop words are integrated based on the edit distance, so as to achieve the fine processing of the extracted video text, and achieve the purpose of more accurately identifying the frame text information in the target video, instead of relying solely on the text extraction tool, thereby solving the technical problem of low content recognition accuracy of the video text caused by the text extraction tool ignoring part of the text information in the video frame. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0012] Figure 1 is a schematic diagram of a hardware environment of an optional video text recognition method according to an embodiment of the present invention;
[0013] Figure 2 is a flow chart of an optional video text recognition method according to an embodiment of the present invention;
[0014] Figure 3 is a schematic diagram of an optional video text recognition method according to an embodiment of the present invention;
[0015] Figure 4 is a flowchart of another optional video text recognition method according to an embodiment of the present invention;
[0016] Figure 5 is a flowchart of another optional video text recognition method according to an embodiment of the present invention;
[0017] Figure 6 is a schematic diagram of a calculation matrix in an optional video text recognition method according to an embodiment of the present invention;
[0018] Figure 7 is a schematic diagram of a calculation matrix in another optional video text recognition method according to an embodiment of the present invention;
[0019] Figure 8 is a schematic structural diagram of an optional video text recognition device according to an embodiment of the present invention;
[0020] Fig. 9 It is a schematic diagram of the structure of an optional electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0021] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0022] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0023] According to one aspect of an embodiment of the present invention, a method for recognizing video text is provided. Optionally, as an optional implementation, the method for recognizing video text can be applied to, but is not limited to, Figure 1 In the video text recognition system in the hardware environment shown, the video text recognition system may include but is not limited to a terminal device 102, a network 104, and a server 106. The terminal device 102 runs a client (such as a client computer) logged in using a target user account. Figure 1 The terminal device 102 includes a human-computer interaction screen 1022, a processor 1024 and a memory 1026. The human-computer interaction screen 1022 is used to present the target video being played in the above-mentioned playback client (such as Figure 1 The target video about the game published by publisher A is shown in the figure, wherein the target video carries text information; it is also used to provide a human-computer interaction interface to receive human-computer interaction operations performed on the human-computer interaction interface of the playback client (such as interactive operations on the target video: collection operation, sharing operation, evaluation operation, etc.). The processor 1024 is used to generate an interaction instruction in response to the above human-computer interaction operation, and send the interaction instruction to the server corresponding to the playback client to complete the interaction process of the target video. The memory 1026 is used to store the above target video.
[0024] In addition, the server 106 includes a database 1062 and a processing engine 1064. The database 1062 is used to store the target video sent by the terminal device 102 and the target text in the recognized target video. The processing engine 1064 is used to perform denoising and recognition processing on the video text extracted from the target video to obtain the target text.
[0025] The specific process is as follows: As in step S102, the server 106 obtains the target video to be identified published by publisher A in the terminal device 102. Steps S104 to S112 are executed in the server 106: the video text extracted from the target video is obtained, and the video text here includes the frame text subsets corresponding to each video frame in the target video. Then, the target frame text subset carrying stop words is determined from the above video text. The stop words carried in the above target frame text subset are removed to update the above video text to the candidate text. The edit distance between the text fragments in the frame text subset corresponding to any two video frames in the candidate text is determined, and the candidate texts are integrated according to the edit distance to obtain the target text finally identified for the target video. In order to facilitate the classification of the target video based on the above accurately identified target text, the target video can be accurately recommended to more users according to the type.
[0026] It should be noted that, in this embodiment, after obtaining the frame text subsets corresponding to each video frame in the video text extracted from the target video, the target frame text subset carrying stop words is determined from the above video text, and the stop words carried in the above target frame text subset are removed to update the above video text to the candidate text. Then the edit distance between the text fragments in the frame text subsets corresponding to any two video frames in the candidate text is determined, and the candidate texts are integrated according to the edit distance to obtain the target text finally recognized for the target video. In other words, the text fragments in the frame text subset extracted by the extraction tool are filtered and integrated, the stop words are removed, and the text fragments in the frame text subset in the candidate text filtered out of the stop words are integrated based on the edit distance, so as to achieve the fine processing of the extracted video text, and achieve the purpose of more accurately identifying the frame text information in the target video, instead of relying solely on the text extraction tool, thereby solving the technical problem of low content recognition accuracy of the video text caused by the text extraction tool ignoring part of the text information in the video frame.
[0027] Optionally, in this embodiment, the terminal device may be a terminal device configured with a target client, which may include but is not limited to at least one of the following: a mobile phone (such as an Android phone, an iOS phone, etc.), a laptop, a tablet computer, a PDA, a MID (Mobile Internet Devices), a PAD, a desktop computer, a smart TV, etc. The target client may be an application client with a video playback function plug-in, such as a video client, an instant messaging client, a browser client, an education client, etc. The network may include but is not limited to: a wired network, a wireless network, wherein the wired network includes: a local area network, a metropolitan area network and a wide area network, and the wireless network includes: Bluetooth, WIFI and other networks that realize wireless communication. The server may be a single server, or a server cluster consisting of multiple servers, or a cloud server. The above is only an example, and no limitation is made to this in this embodiment.
[0028] Optionally, as an optional implementation, as Figure 2 As shown, the above-mentioned video text recognition method includes:
[0029] S202, obtaining video text extracted from the target video to be identified, wherein the video text includes frame text subsets corresponding to respective video frames of the target video;
[0030] S204, determining a target frame text subset carrying stop words from the video text;
[0031] S206, removing stop words carried in the target frame text subset to update the video text into candidate text;
[0032] S208, determining the edit distance between text segments in the frame text subsets corresponding to any two video frames in the candidate text;
[0033] S210, integrating the candidate texts according to the edit distance to obtain the target text recognized for the target video.
[0034] Optionally, in this embodiment, the above video text recognition method can be applied to, but is not limited to, a client with a video playback function plug-in. For example, a playback client for playing long videos, a video sharing client for playing short videos, a playback platform client for playing videos released and shared by individuals or enterprises (such as a web page with a playback function plug-in, a community sharing space, etc.). The above is an example, and this embodiment does not make any limitation to this.
[0035] Optionally, in this embodiment, the video text extracted from the target video to be identified may be, but is not limited to, text content identified by a text extraction tool (such as an OCR tool). That is to say, before obtaining the video text of the target video, the text content of the video text in each video frame (such as a frame text subset) is first extracted from the target video using a text extraction tool or method of related technology. The text content here may include, but is not limited to, at least one of the following: the title text of the target video, the content text in the target video, and the hash tag text in the target video. Among them, the hash tag text (hashtag text) here may be, but is not limited to, topic text in the video, such as text information introduced by the symbol "#". For example, in a video such as Figure 3 The text content in the target video shown includes: Title text 302: "One shot and it's a hit!" (eg Figure 3 ), content text 304: "Publisher A", "Game Special" and "Oli Gei!" (such as Figure 3 The text in the dotted circle shown), topic text 306: "#GameTime..." (such as Figure 3 (text in the solid circle shown). Figure 3 The figure shows an example of text content in a video frame. In this embodiment, the type and content of the video text contained in a specific scene are not limited.
[0036] It should be noted that a stop word set including each stop word may be constructed in advance, but is not limited to, wherein the stop word set may be constructed in the following manner, but is not limited to:
[0037] 1) Add the publishing information of the playback platform where the video is located to the stop word set, wherein the publishing information includes the watermark of the login account registered by the publisher and the platform identifier of the playback platform.
[0038] 2) Words that appear in each video frame of the video with a frequency higher than a threshold are added to the stop word set.
[0039] That is to say, invalid text fragments (such as words or phrases) that appear frequently in each video frame of the video and have no substantive meaning are used as stop words, so as to ensure the recognition accuracy of valid text fragments when the invalid text fragments are displayed obscured by other identified valid text fragments.
[0040] Optionally, in this embodiment, determining the edit distance between text segments in the frame text subsets corresponding to any two video frames in the filtered candidate text may include but is not limited to: sequentially comparing and calculating the frame text subsets corresponding to each video frame to obtain the edit distance between text segments in different frame text subsets. It should be noted that the method for calculating the edit distance here may be but is not limited to the Leveinshtein distance method: calculating the cost required to convert the source string into the target string, where the lower the cost, the higher the similarity, and the higher the cost, the lower the similarity. In this embodiment, it may be but is not limited to calculating the cost required to convert a string contained in a text segment in a frame text subset into a string contained in a text segment in another frame text subset.
[0041] Further, in this embodiment, in the process of calculating the edit distance between the above text segments, a calculation matrix for calculating the string distance between strings of different lengths in the text segments can be constructed, but is not limited to it. In the calculation matrix, the characters in the above two text segments to be calculated are used as matrix reference elements, and the matrix reference elements here can be, but are not limited to, the first row element or the first column element.
[0042] In addition, in this embodiment, before calculating the edit distance between different text segments, it is also necessary to obtain the string length corresponding to each text segment, and determine the integration processing method for each text segment based on the comparison result of the string length combined with the edit distance. In this way, similar texts in the text segments can be deduplicated, further ensuring the accuracy of the recognition results.
[0043] Optionally, in this embodiment, after the candidate texts are integrated according to the edit distance to obtain the target text identified for the target video, the method further includes: determining the type label corresponding to the target video according to the target text; and marking the target video with the type label. In addition, when the type label of the target video is accurately determined, the accuracy of pushing the target video will be further improved, so that the target video will be pushed to the user account that is more interested in the type.
[0044] Specific combination Figure 4 The process steps shown are explained:
[0045] Assume that the target video is a short video about a game published by publisher A. As in step S402, the text in each video frame of the short video is extracted using an OCR tool to obtain a video text. The video text may include at least one of the following: a title text, a content text, and a topic text. Then, step S404 is performed to search for stop words matching the reference stop words recorded in the stop word set in the video text. When matching stop words are found, step S406 is performed to filter and remove the stop words found to obtain a candidate text. Then, as in step S408, the edit distance is calculated for the text fragments in the frame text subset corresponding to any two video frames in the candidate text, and the text fragments are integrated according to the edit distance to obtain the object text fragments in the target text. Finally, as in step S410, the object text fragments in the target text are spliced in sequence according to the order of appearance of the video frames to obtain the target text finally identified for the target video.
[0046] Above Figure 4 The process shown is only an example, and this embodiment does not limit the execution order, the exemplary screens involved, and the names of the tools involved.
[0047] Through the embodiments provided by the present application, the text fragments in the frame text subset that have been extracted by the extraction tool are filtered and integrated to remove stop words, and the text fragments in the frame text subset in the candidate text from which the stop words have been filtered out are integrated based on the edit distance, thereby achieving fine processing of the extracted video text and achieving the purpose of more accurately identifying the frame text information in the target video without relying solely on the text extraction tool, thereby solving the technical problem of low content recognition accuracy of the video text caused by the text extraction tool ignoring part of the text information in the video frame.
[0048] As an optional solution, determining a target frame text subset carrying stop words from the video text includes:
[0049] S1, searching the video text for stop words that match the reference stop words recorded in the stop word set;
[0050] S2, when a stop word is found, determining the frame text subset where the stop word is located as the target frame text subset.
[0051] It should be noted that, in this embodiment, the process is performed on the premise that the video text of the target video extracted by the extraction tool is obtained, and a refined re-recognition process is performed on the extracted video text.
[0052] Optionally, in this embodiment, the stop word set may include, but is not limited to, traditional text stop words, such as "you", "I", "he", "of", "already", etc., and sensitive words such as "awesome".
[0053] In addition, in this embodiment, the stop word set may further include, but is not limited to, the release information of the target video when it is published on the playback platform. Here, the release information may include, but is not limited to, publicly disclosed logo watermark words (such as "XX Video", "YY TV Station", etc.).
[0054] Furthermore, in this embodiment, the stop word set may further include, but is not limited to: words with a high word frequency. It should be noted that some logos are not common words in the public dictionary, so the word frequency can be further used to distinguish and identify stop words.
[0055] Through the embodiment provided by this application, in each frame text subset of the video text, when a stop word matching the reference stop word in the stop word set is found, the frame text subset where the stop word is located is determined as the target frame text subset, so as to directly perform efficient filtering processing on the target frame text subset, shorten the denoising process in text recognition, and improve the recognition efficiency.
[0056] As an optional solution, before obtaining the video text extracted from the target video to be recognized, it further includes:
[0057] S1, determining the release information of the reference video on the playback platform, where the release information includes at least one of the following: the watermark of the login account registered by the reference video on the playback platform, the platform identifier of the playback platform;
[0058] S2, adding the release information to the stop word set.
[0059] It should be noted that the platform identifier in the release information of the reference video here can be, but is not limited to, a logo certified in the public dictionary, so as to facilitate directly adding it to the stop word set as a stop word.
[0060] Specifically combined with Figure 3 the following example is used for illustration: Assume that the reference video is a short video about games published by publisher A. Before using the recognition method provided in this embodiment for the target video, if the platform identifier watermark "publisher A" of the playback platform is found in the content text 302 of the reference video, such release information as this type of logo watermark can be added to the stop word set to construct a stop word set for subsequent stop word recognition of the target video.
[0061] Through the embodiments provided by the present application, the release information of the reference video in the playback platform is determined, and the release information is added to the stop word set. Thus, in the video recognition process, invalid text segments are filtered out in time to avoid display interference of invalid text segments on valid text segments, thereby improving the recognition accuracy of valid text segments.
[0062] As an optional solution, after obtaining the video text extracted from the target video to be identified, the following is further included:
[0063] S1, counting the word frequency of each text segment in each frame text subset of the video text, wherein the word frequency of the text segment is used to indicate the number of times the text segment appears in the video text;
[0064] S2, obtaining the ratio between the word frequency of each text segment and the total number of text segments contained in the target video;
[0065] S3, adding the target text segment whose ratio is greater than the first threshold to the stop word set.
[0066] Optionally, in this embodiment, after the video text is acquired, the stop word set can be updated in time. That is, the station logos that are not certified in the public dictionary can also be effectively identified through the statistical analysis of word frequency, and the release information such as the watermarks of these station logos that are not certified in the public dictionary is added to the stop word set, so as to achieve the purpose of real-time updating of the stop word set.
[0067] For example, when the word frequencies of different text segments in the video OCR results are obtained, the ratio between the word frequencies and the total number of frames of the video is calculated, and the target text segments whose ratio is greater than a specific threshold K (the K value is a value close to 1) are added to the stop word set.
[0068] Through the embodiments provided by the present application, after obtaining the video text, the stop word set can be updated in real time based on the word frequency of each text segment counted in the video text to further ensure the accuracy of filtering and identifying the video text of the video.
[0069] As an optional solution, determining the edit distance between text segments in the frame text subsets corresponding to any two video frames in the candidate text includes:
[0070] S1, determining a current frame text subset to be processed from the candidate texts;
[0071] S2, traverse the current frame text subset to obtain the first text segment to be compared;
[0072] S3, obtaining a second text segment to be compared from the reference frame text subset other than the current frame text subset in the candidate text;
[0073] S4, calculating the edit distance between the first text segment and the second text segment.
[0074] Specific combination Figure 5 The example shown is for illustration:
[0075] A current frame text subset and a reference frame text subset other than the current frame text subset are determined from the candidate texts. Then, as in step S502-1, a first text segment is obtained from the current frame text subset, and as in step S502, a second text segment is obtained from the reference frame text subset.
[0076] It should be noted that the current frame text subset and the reference frame text subset to be processed can be determined at the same time, or the current frame text subset can be determined first, and then the reference frame text subset can be determined from the non-current frame text subset. Figure 5 As shown and described above, the order and method of obtaining the two frame text subsets are not limited. In addition, the current frame text subset to be processed and the reference frame text subset can be determined at the same time. The first text segment in the current frame text subset can be determined first, and then the second text segment can be determined from the reference frame text subset. The first text segment and the second text segment can also be obtained asynchronously at the same time. This is to explain that after obtaining the first text segment and the second text segment, the two are compared and calculated. Figure 5 The illustration and the above text description do not impose any limitation on the order and method of obtaining the two.
[0077] After obtaining the first text segment and the second text segment, the edit distance between the two is calculated as in step S504. Assume that the video text segment extracted by the OCR result in the first frame is "This parabola is amazing", and the video text segment extracted by the OCR result in the third frame is "This parabola stretches", and the video text segment extracted by the OCR result in the fifth frame may have another one identified as "This parabola is amazing". Then, the edit distance can be calculated by comparing the above text segments to determine the similar text segments appearing in different video frames.
[0078] After the integration process of the first text segment and the second text segment as shown in step S506 (such as performing deduplication and denoising process on similar text segments or retaining process on non-similar texts), the object text segment in the target text is obtained. Then, as shown in step S508, the object text segments in the target text are spliced according to the order of appearance of the video frames in which they are located, so as to obtain the target text finally recognized by the target video.
[0079] Through the embodiment provided by the present application, when a first text segment is obtained from the current frame text subset and a second text segment is obtained from the reference frame text subset, the two text segments are compared to calculate the edit distance between them. Thus, similar text segments are deduplicated and de-noised according to the edit distance, so as to determine the correct text segment from the similar text segments, thereby achieving the purpose of improving text recognition efficiency.
[0080] As an optional solution, calculating the edit distance between the first text segment and the second text segment includes:
[0081] S1, determining a first character string length corresponding to a first text segment, and a second character string length corresponding to a second text segment;
[0082] S2, constructing a calculation matrix based on the first character string length and the second character string length, wherein each first character contained in the first text segment and each second character contained in the second text segment are used as matrix reference elements of the calculation matrix, and the matrix reference elements are the first row elements or the first column elements in the calculation matrix;
[0083] S3, sequentially traversing each character included in the calculation matrix, determining a first character string to be calculated based on the first character, and determining a second character string to be calculated based on the second character;
[0084] S4, calculating a string distance between the first string and the second string, wherein the edit distance between the first text segment and the second text segment includes a plurality of string distances.
[0085] Optionally, in this embodiment, before obtaining the second text segment to be compared from the reference frame text subset other than the current frame text subset in the candidate text, it also includes: determining a candidate frame text subset other than the current frame text subset from the candidate text; and when the candidate frame text subset has not been used to calculate the edit distance, determining the candidate frame text subset as the reference frame text subset.
[0086] The following example is used to illustrate: the length of a first character string corresponding to a first text segment and the length of a second character string corresponding to a second text segment are counted. When the length of the first character string is the same as the length of the second character string, the following operations are performed:
[0087] The string distance of the strings contained in the above two text fields is calculated by the following formula:
[0088]
[0089] Among them LD a,b(i, j) represents the cost of converting between a string a of length i and a string b of length j, that is, the string distance between string a and string b. It should be noted that, in this embodiment, the characters included in the above text fragments can be combined into different strings in sequence, so the edit distance between the text fragments can be determined by calculating the string distances of different strings.
[0090] For example, here we take the two strings Meng and man as examples to calculate the edit distance (also the string distance). Assume that the lengths of the two strings are constructed as follows Figure 6 In the calculation matrix shown, the characters in the string Meng are used as the first column elements and the identifier "0" is added, and the characters in the string man are used as the first row elements and the identifier "0" is added. It should be noted that the identifier "0" here is used to distinguish the strings, which is only an example and not a limitation.
[0091] Assume that the string "0" constructed with the characters in the string "0Meng" is used as the first string (eg Figure 6 ), and the string "0ma" constructed with the characters in "0man" as the second string (as shown in the vertical square box in the first column of Figure 6 ), then the edit distance is calculated to be 2 (as shown in the first horizontal square box in the figure). Figure 6 (shown in the circular box).
[0092] The above is just an example. By iterating the calculation in sequence, the string distance corresponding to the string constructed by each character will be obtained as follows: Figure 7 The contents of the matrix are shown.
[0093] Through the embodiments provided by the present application, after determining the length of the first string corresponding to the first text fragment and the length of the second string corresponding to the second text fragment, a calculation matrix is constructed based on the above string lengths, and the string distance between the first string constructed in sequence by different first characters in the first text fragment and the second string constructed in sequence by different second characters in the second text fragment is calculated in the calculation matrix, so as to finally obtain the accurate editing distance between the text fragments.
[0094] As an optional implementation scheme, the candidate texts are integrated according to the edit distance to obtain the target text recognized for the target video, including:
[0095] 1) When the length of a first character string corresponding to the first text segment is the same as the length of a second character string corresponding to the second text segment, and the edit distance indication between the first text segment and the second text segment is greater than a first threshold, both the first text segment and the second text segment are used as object text segments in the target text;
[0096] 2) When the length of a first character string corresponding to the first text segment is the same as the length of a second character string corresponding to the second text segment, and the edit distance indication between the first text segment and the second text segment is less than a second threshold, determine that the first text segment and the second text segment are similar text segments, and select a text segment from the first text segment and the second text segment as the object text segment in the target text, wherein the second threshold is less than the first threshold.
[0097] As another optional implementation, the candidate texts are integrated according to the edit distance to obtain the target text recognized for the target video, including:
[0098] 1) When the length of a first character string corresponding to a first text segment is different from the length of a second character string corresponding to a second text segment, and the edit distance indication between the first text segment and the second text segment is less than a third threshold, the first text segment and the second text segment are determined to be similar text segments, and a text segment is selected from the first text segment and the second text segment as an object text segment in a target text.
[0099] Optionally, in this embodiment, selecting a text segment from the first text segment and the second text segment as the object text segment in the target text includes: obtaining a first number of occurrences of the first text segment in the target video, and a second number of occurrences of the second text segment in the target video; when the first number of occurrences is greater than the second number of occurrences, replacing the second text segment with the first text segment, and determining the first text segment as the object text segment; when the second number of occurrences is greater than the first number of occurrences, replacing the first text segment with the second text segment, and determining the second text segment as the object text segment.
[0100] Assuming that a first text segment with a string length of 5 is obtained, and a second text segment with a string length of 5 is obtained, and the edit distance between the two is calculated to be 1, the two can be regarded as similar text segments. Then, the frequency of occurrence of the two in the entire video is calculated respectively, and the text segment with a low frequency of occurrence is replaced with the text segment with a high frequency of occurrence, and the text segment with a high frequency of occurrence is determined as the correct object text segment.
[0101] Through the embodiments provided in the present application, based on the character string length corresponding to the text fragments and the edit distance between the text fragments, the accurate object text fragment in the target text is determined from the first text fragment and the second text fragment, thereby ensuring the accuracy of text content recognition of the target video.
[0102] As an optional solution, the candidate texts are integrated according to the edit distance to obtain the target text recognized for the target video, including:
[0103] S1, splicing the acquired object text segments in the target text according to the order of the frames appearing in the target video to obtain the target text corresponding to the target video.
[0104] When all object text segments of the target text are obtained, each object text segment can be spliced from front to back, from left to right, and from top to bottom in the order in which the video frames in which they are located appear in the target video to obtain the correct target text content of the target video.
[0105] Through the embodiments provided in the present application, the accuracy of the target text identified for the target video is further guaranteed by integrating the refined object text segments.
[0106] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present invention.
[0107] According to another aspect of the embodiment of the present invention, a video text recognition device for implementing the above-mentioned video text recognition method is also provided. Figure 8 As shown, the device comprises:
[0108] 1) A first acquisition unit 802 is used to acquire video text extracted from a target video to be identified, wherein the video text includes a frame text subset corresponding to each video frame of the target video;
[0109] 2) A first determining unit 804, configured to determine a target frame text subset carrying stop words from the video text;
[0110] 3) A filtering unit 806, used to remove stop words carried in the target frame text subset to update the video text into a candidate text;
[0111] 4) A second determining unit 808, configured to determine the edit distance between text segments in the frame text subsets corresponding to any two video frames in the candidate text;
[0112] 5) An integration recognition unit 810 is used to integrate the candidate texts according to the edit distance to obtain the target text recognized for the target video.
[0113] As an optional solution, the first determining unit includes:
[0114] A search module, for searching the video text for stop words that match the reference stop words recorded in the stop word set;
[0115] The first determination module is used to determine the frame text subset where the stop word is located as the target frame text subset when a stop word is found.
[0116] Optionally, in this embodiment, the embodiments here can refer to the above method embodiments, which will not be repeated here.
[0117] As an optional solution, it also includes:
[0118] A third determining unit is used to determine the release information of the reference video in the playback platform before obtaining the video text extracted from the target video to be identified, wherein the release information includes at least one of the following: a watermark of a login account registered by the reference video in the playback platform, and a platform identifier of the playback platform;
[0119] The first adding unit is used to add the published information to the stop word set.
[0120] Optionally, in this embodiment, the embodiments here can refer to the above method embodiments, which will not be repeated here.
[0121] As an optional solution, it also includes:
[0122] A counting unit is used to count the word frequency of each text segment in each frame text subset of the video text after obtaining the video text extracted from the target video to be identified, wherein the word frequency of the text segment is used to indicate the number of times the text segment appears in the video text;
[0123] A second acquisition unit is used to acquire the ratio between the word frequency of each text segment and the total number of text segments included in the target video;
[0124] The second adding unit is used to add the target text segment whose ratio is greater than the first threshold to the stop word set.
[0125] Optionally, in this embodiment, the embodiments here can refer to the above method embodiments, which will not be repeated here.
[0126] As an optional solution, the second determining unit includes:
[0127] A second determination module is used to determine a current frame text subset to be processed from the candidate texts;
[0128] A first acquisition module, used for traversing the current frame text subset to obtain the first text segment to be compared;
[0129] A second acquisition module is used to acquire a second text segment to be compared from a reference frame text subset other than a current frame text subset in the candidate text;
[0130] The calculation module is used to calculate the edit distance between the first text segment and the second text segment.
[0131] Optionally, in this embodiment, the embodiments here can refer to the above method embodiments, which will not be repeated here.
[0132] As an optional solution, the calculation module includes:
[0133] A first determination submodule, used to determine a first character string length corresponding to the first text segment, and a second character string length corresponding to the second text segment;
[0134] A construction submodule, used to construct a calculation matrix based on the first character string length and the second character string length, wherein each first character contained in the first text segment and each second character contained in the second text segment are used as matrix reference elements of the calculation matrix, and the matrix reference elements are the first row elements or the first column elements in the calculation matrix;
[0135] A third determination submodule, used for sequentially traversing each character included in the calculation matrix, determining a first character string to be calculated based on the first character, and determining a second character string to be calculated based on the second character;
[0136] The calculation submodule is used to calculate the string distance between the first string and the second string, wherein the edit distance between the first text segment and the second text segment includes a plurality of string distances.
[0137] Optionally, in this embodiment, the embodiments here can refer to the above method embodiments, which will not be repeated here.
[0138] As an optional solution, it also includes:
[0139] A fourth determining unit is used to determine a candidate frame text subset other than the current frame text subset from the candidate text before obtaining the second text segment to be compared from the reference frame text subset other than the current frame text subset in the candidate text;
[0140] The fifth determining unit is configured to determine the candidate frame text subset as the reference frame text subset when the candidate frame text subset has not been used to calculate the edit distance.
[0141] Optionally, in this embodiment, the embodiments here can refer to the above method embodiments, which will not be repeated here.
[0142] As an optional solution, the integrated identification unit includes:
[0143] a first recognition module, configured to, when a first character string length corresponding to the first text segment is the same as a second character string length corresponding to the second text segment and an edit distance indication between the first text segment and the second text segment is greater than a first threshold, regard both the first text segment and the second text segment as object text segments in the target text;
[0144] A second recognition module is used to determine that the first text segment and the second text segment are similar text segments when the length of a first character string corresponding to the first text segment is the same as the length of a second character string corresponding to the second text segment, and the edit distance indication between the first text segment and the second text segment is less than a second threshold, and select a text segment from the first text segment and the second text segment as an object text segment in the target text, wherein the second threshold is less than the first threshold.
[0145] Optionally, in this embodiment, the embodiments here can refer to the above method embodiments, which will not be repeated here.
[0146] As an optional solution, the integrated identification unit includes:
[0147] A third recognition module is used to determine that the first text segment and the second text segment are similar text segments when the length of a first character string corresponding to the first text segment is different from the length of a second character string corresponding to the second text segment, and the edit distance indication between the first text segment and the second text segment is less than a third threshold, and select a text segment from the first text segment and the second text segment as the object text segment in the target text.
[0148] Optionally, in this embodiment, the embodiments here can refer to the above method embodiments, which will not be repeated here.
[0149] As an optional solution, the second identification module and the third identification module each include:
[0150] A first acquisition submodule, used to acquire a first occurrence number of the first text segment in the target video, and a second occurrence number of the second text segment in the target video;
[0151] A first replacement submodule, configured to replace the second text segment with the first text segment when the first number of occurrences is greater than the second number of occurrences, and determine the first text segment as the object text segment;
[0152] The second replacement submodule is used to replace the first text segment with the second text segment when the second number of occurrences is greater than the first number of occurrences, and determine the second text segment as the object text segment.
[0153] Optionally, in this embodiment, the embodiments here can refer to the above method embodiments, which will not be repeated here.
[0154] As an optional solution, the integrated identification unit includes:
[0155] The splicing module is used to splice the object text segments in the acquired target text according to the order of frames appearing in the target video to obtain the target text corresponding to the target video.
[0156] Optionally, in this embodiment, the embodiments here can refer to the above method embodiments, which will not be repeated here.
[0157] As an optional solution, it also includes:
[0158] a sixth determination unit, configured to determine a type label corresponding to the target video according to the target text after integrating the candidate texts according to the edit distance to obtain the target text identified for the target video;
[0159] The tagging unit is used to tag the target video with a type tag.
[0160] Optionally, in this embodiment, the embodiments here can refer to the above method embodiments, which will not be repeated here.
[0161] According to another aspect of the embodiment of the present invention, an electronic device for implementing the above-mentioned video text recognition method is also provided. The electronic device may be Figure 1 The terminal device or server shown in the figure. This embodiment is described by taking the electronic device as a server as an example. Fig. 9 As shown, the electronic device includes a memory 902 and a processor 904. The memory 902 stores a computer program, and the processor 904 is configured to execute the steps in any of the above method embodiments through the computer program.
[0162] Optionally, in this embodiment, the electronic device may be located in at least one network device among a plurality of network devices of a computer network.
[0163] Optionally, in this embodiment, the processor may be configured to perform the following steps through a computer program:
[0164] S1, obtaining video text extracted from a target video to be identified, wherein the video text includes a frame text subset corresponding to each video frame of the target video;
[0165] S2, determining a target frame text subset carrying stop words from the video text;
[0166] S3, removing the stop words carried in the target frame text subset to update the video text into the candidate text;
[0167] S4, determining the edit distance between text segments in the frame text subsets corresponding to any two video frames in the candidate text;
[0168] S5, integrating the candidate texts according to the edit distance to obtain the target text recognized for the target video.
[0169] Alternatively, a person skilled in the art may understand that: Fig. 9 The structure shown is for illustration only, and the electronic device may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (MID), a PAD, and other terminal devices. Fig. 9 The electronic device and the electronic equipment described above are not limited in structure. Fig. 9 More or fewer components (such as network interfaces, etc.) as shown in, or with Fig. 9 Different configurations are shown.
[0170] Among them, the memory 902 can be used to store software programs and modules, such as the program instructions / modules corresponding to the video text recognition method and device in the embodiment of the present invention. The processor 904 executes various functional applications and data processing by running the software programs and modules stored in the memory 902, that is, realizing the above-mentioned video text recognition method. The memory 902 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 902 may further include a memory remotely arranged relative to the processor 904, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. Among them, the memory 902 can be specifically, but not limited to, used to store information such as target videos and video texts. As an example, such as Fig. 9As shown, the memory 902 may include, but is not limited to, the first acquisition unit 802, the first determination unit 804, the filtering unit 806, the second determination unit 808, and the integrated recognition unit 810 in the video text recognition device. In addition, other module units in the video text recognition device may also be included but are not limited to, which will not be repeated in this example.
[0171] Optionally, the transmission device 906 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wired network and a wireless network. In one example, the transmission device 906 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices and routers via a network cable so as to communicate with the Internet or a local area network. In one example, the transmission device 906 is a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0172] In addition, the electronic device further includes: a display 908 for displaying the target video; and a connection bus 910 for connecting various module components in the electronic device.
[0173] In other embodiments, the terminal device or server may be a node in a distributed system, wherein the distributed system may be a blockchain system, and the blockchain system may be a distributed system formed by connecting the multiple nodes through network communication. Among them, the nodes may form a peer-to-peer (P2P, Peer To Peer) network, and any form of computing device, such as a server, terminal and other electronic devices, may become a node in the blockchain system by joining the peer-to-peer network.
[0174] According to one aspect of the present application, a computer program product or computer program is provided, the computer program product or computer program comprising computer instructions, the computer instructions being stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device performs the above method. The computer program is configured to execute the steps of any one of the above method embodiments when it is run.
[0175] Optionally, in this embodiment, the computer-readable storage medium may be configured to store a computer program for performing the following steps:
[0176] S1, obtaining video text extracted from a target video to be identified, wherein the video text includes a frame text subset corresponding to each video frame of the target video;
[0177] S2, determining a target frame text subset carrying stop words from the video text;
[0178] S3, removing the stop words carried in the target frame text subset to update the video text into the candidate text;
[0179] S4, determining the edit distance between text segments in the frame text subsets corresponding to any two video frames in the candidate text;
[0180] S5, integrating the candidate texts according to the edit distance to obtain the target text recognized for the target video.
[0181] Optionally, in this embodiment, a person of ordinary skill in the art may understand that all or part of the steps in the various methods of the above embodiments may be completed by instructing hardware related to the terminal device through a program, and the program may be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a disk or an optical disk, etc.
[0182] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0183] If the integrated units in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in the above computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling one or more computer devices (which can be personal computers, servers or network devices, etc.) to perform all or part of the steps of the methods described in various embodiments of the present invention.
[0184] In the above embodiments of the present invention, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0185] In the several embodiments provided in the present application, it should be understood that the disclosed client can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0186] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0187] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0188] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A video text recognition method, characterized in that: include: Acquire video text extracted from a target video to be identified, wherein the video text includes frame text subsets corresponding to respective video frames of the target video; Determining a target frame text subset carrying stop words from the video text; Removing the stop words carried in the target frame text subset to update the video text into candidate text; Determine a current frame text subset to be processed from the candidate texts; Traversing the current frame text subset to obtain the first text segment to be compared; Acquire a second text segment to be compared from a reference frame text subset other than the current frame text subset in the candidate text; Determining a first character string length corresponding to the first text segment, and a second character string length corresponding to the second text segment; Calculating an edit distance between the first text segment and the second text segment based on the first character string length and the second character string length; Integrate the candidate texts according to the edit distance to determine the object text segment in the target text; The object text segments in the acquired target text are spliced in the order of frames appearing in the target video to obtain the target text identified for the target video.
2. The method according to claim 1, characterized in that Determining a target frame text subset carrying stop words from the video text comprises: Searching the video text for stop words that match reference stop words recorded in a stop word set; When the stop words are found, the frame text subset where the stop words are located is determined as the target frame text subset.
3. The method according to claim 2, characterized in that Before obtaining the video text extracted from the target video to be identified, the method further includes: Determine the release information of the reference video in the playback platform, wherein the release information includes at least one of the following: a watermark of a login account registered by the reference video in the playback platform, and a platform identifier of the playback platform; The published information is added to the stop word set.
4. The method according to claim 2, characterized in that: After obtaining the video text extracted from the target video to be identified, the method further includes: Counting the word frequency of each text segment in each frame text subset of the video text, wherein the word frequency of the text segment is used to indicate the number of times the text segment appears in the video text; Obtaining the ratio between the word frequency of each of the text segments and the total number of text segments contained in the target video; The target text segment having a ratio greater than a first threshold is added to the stop word set.
5. The method according to claim 1, characterized in that The calculating the edit distance between the first text segment and the second text segment based on the first character string length and the second character string length includes: A calculation matrix is constructed based on the first character string length and the second character string length, wherein each first character included in the first text segment and each second character included in the second text segment are used as matrix reference elements of the calculation matrix, and the matrix reference elements are first row elements or first column elements in the calculation matrix; Sequentially traverse each character included in the calculation matrix, determine a first character string to be calculated based on the first character, and determine a second character string to be calculated based on the second character; A string distance between the first string and the second string is calculated, wherein the edit distance between the first text segment and the second text segment includes a plurality of the string distances.
6. The method according to claim 1, characterized in that Before obtaining the second text segment to be compared from the reference frame text subset other than the current frame text subset in the candidate text, the method further includes: Determine a candidate frame text subset other than the current frame text subset from the candidate texts; In a case where the candidate frame text subset has not been used to calculate the edit distance, the candidate frame text subset is determined as the reference frame text subset.
7. The method according to claim 1, characterized in that The step of integrating the candidate texts according to the edit distance to determine the object text segment in the target text comprises: When the length of a first character string corresponding to the first text segment is the same as the length of a second character string corresponding to the second text segment, and the edit distance indication between the first text segment and the second text segment is greater than a first threshold, both the first text segment and the second text segment are used as object text segments in the target text; When the length of the first character string corresponding to the first text segment is the same as the length of the second character string corresponding to the second text segment, and the edit distance indication between the first text segment and the second text segment is less than a second threshold, the first text segment and the second text segment are determined to be similar text segments, and a text segment is selected from the first text segment and the second text segment as the object text segment in the target text, wherein the second threshold is less than the first threshold.
8. The method according to claim 1, characterized in that The step of integrating the candidate texts according to the edit distance to determine the object text segment in the target text comprises: When the length of a first character string corresponding to the first text segment is different from the length of a second character string corresponding to the second text segment, and the edit distance indication between the first text segment and the second text segment is less than a third threshold, the first text segment and the second text segment are determined to be similar text segments, and a text segment is selected from the first text segment and the second text segment as the object text segment in the target text.
9. The method according to claim 7 or 8, characterized in that: Selecting a text segment from the first text segment and the second text segment as the object text segment in the target text includes: Obtaining a first occurrence number of the first text segment in the target video and a second occurrence number of the second text segment in the target video; When the first number of occurrences is greater than the second number of occurrences, the second text segment is replaced by the first text segment, and the first text segment is determined as the object text segment; When the second number of occurrences is greater than the first number of occurrences, the first text segment is replaced with the second text segment, and the second text segment is determined as the object text segment.
10. The method according to any one of claims 1 to 8, characterized in that After the object text segments in the acquired target text are spliced in the order of frames appearing in the target video to obtain the target text identified for the target video, the method further includes: Determine a type label corresponding to the target video according to the target text; The target video is marked with the type tag.
11. A video text recognition device, characterized in that: include: A first acquisition unit is used to acquire video text extracted from a target video to be identified, wherein the video text includes a frame text subset corresponding to each video frame of the target video; A first determining unit, configured to determine a target frame text subset carrying stop words from the video text; A filtering unit, configured to remove the stop words carried in the target frame text subset, so as to update the video text into a candidate text; A second determination unit is used to determine a current frame text subset to be processed from the candidate texts; traverse the current frame text subset to obtain a first text segment to be compared; obtain a second text segment to be compared from a reference frame text subset other than the current frame text subset in the candidate texts; determine a first character string length corresponding to the first text segment, and a second character string length corresponding to the second text segment; and calculate an edit distance between the first text segment and the second text segment based on the first character string length and the second character string length; An integration and recognition unit, used for integrating the candidate texts according to the edit distance to determine the object text segment in the target text; The device is also used to, after integrating the candidate texts according to the edit distance and determining the object text segments in the target text, splice the acquired object text segments in the target text in the order of the frames appearing in the target video to obtain the target text identified for the target video.
12. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein the program executes the method described in any one of claims 1 to 10 when executed.
13. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 10 through the computer program.
Citation Information
Patent Citations
Text detection method and device and computer readable storage medium
CN110728167A
Video screening method and device, electronic equipment and storage medium
CN110990631A