Video recognition method, device, apparatus and computer storage medium

By integrating the audio, text, and visual information features of videos, this technology solves the problem of low accuracy in video originality recognition in existing technologies, achieving more efficient video deduplication and improving video quality and user experience on video platforms.

CN115631447BActive Publication Date: 2026-04-14MIGU CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-11
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

The accuracy of existing technologies for identifying the originality of videos is low, resulting in a high rate of video duplication on video platforms, which affects the user's viewing and creation experience.

Method used

By combining the audio and text information and the visual information of the video to perform feature fusion, the fused video features of the video to be identified are determined and compared with a preset video database to identify the originality of the video.

Benefits of technology

It improved the accuracy of video originality identification, reduced the video duplication rate on video platforms, and enhanced video quality and user viewing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115631447B_ABST
    Figure CN115631447B_ABST
Patent Text Reader

Abstract

The embodiment of the present application relates to the technical field of computer data processing, and discloses a video identification method, which comprises the following steps: determining voice text information and picture information of a to-be-identified video; performing feature fusion on the voice text information and the picture information to obtain fusion video features of the to-be-identified video; and determining an originality determination result of the to-be-identified video according to the fusion video features and a preset video database. In the above manner, the accuracy of original video determination is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer data processing technology, specifically to a video recognition method, apparatus, device, and computer storage medium. Background Technology

[0002] On video playback platforms, such as those where users can upload their own videos, the platform typically identifies the originality of videos and filters out duplicates to improve the user viewing experience.

[0003] The inventors of this application discovered during the implementation of the embodiments of the present invention that the industry's video originality recognition only identifies and filters based on the video's visual content, which has the problem of low accuracy. Summary of the Invention

[0004] In view of the above problems, embodiments of the present invention provide a video recognition method to solve the problem of low accuracy in video originality recognition in the prior art.

[0005] According to one aspect of the present invention, a video recognition method is provided, the method comprising:

[0006] Determine the audio-visual information and video information of the video to be recognized;

[0007] The voice text information and the video information are fused to obtain the fused video features of the video to be identified;

[0008] Based on the fused video features and the preset video database, the originality determination result of the video to be identified is determined.

[0009] In one alternative approach, the screen information includes screen text information;

[0010] For each video frame included in the video to be identified, several selectable recognition regions in the video frame are marked respectively to obtain the marked video frame;

[0011] Using the center of the marked video frame as the scaling center point, the marked video frame is scaled multiple times to obtain several overlapping video frames of different sizes corresponding to the marked video frame.

[0012] The optional recognition regions in several overlapping video frames of different sizes are merged to obtain at least one target recognition region corresponding to the video frame;

[0013] The on-screen text information is determined based on the text recognition information within the target recognition area corresponding to the video frame.

[0014] In an alternative approach, the method further includes:

[0015] Based on the overlapping area between the several optional recognition regions and the text recognition results corresponding to each optional recognition region, the optional recognition regions are merged to obtain the target recognition region.

[0016] In an alternative approach, the method further includes:

[0017] The optional recognition regions whose overlapping area is greater than a preset area threshold and whose similarity of the text recognition results is greater than a preset similarity threshold are determined as associated recognition regions;

[0018] The associated identification regions are merged to obtain the target identification region.

[0019] In an alternative approach, the method further includes:

[0020] Speech recognition is performed on each video frame in the video to be identified to obtain the foreground speech information and background speech information corresponding to each video frame;

[0021] The foreground speech information and background speech information are respectively converted into text to obtain foreground speech text and background speech text;

[0022] The foreground speech text is deduplicated based on the background speech text to obtain the speech text information.

[0023] In an alternative approach, the method further includes:

[0024] The voice text information is matched with the screen text information to obtain the first matching text and the non-matching text information;

[0025] The approximate text corresponding to the mismatched text information is matched to obtain the second matching text; the approximate text is obtained by performing at least one of the following on the mismatched text: phonetic similarity processing, shape similarity processing, and semantic similarity processing.

[0026] The fused video features are determined based on the first and second matching texts.

[0027] In one optional embodiment, the video database includes at least one of the fused video features, speech and text information, and image information of several pre-stored original videos; the method further includes:

[0028] The fused video features are matched with at least one of the fused video features, speech and text information, and image information of the pre-stored original video. When the match is successful, the originality determination result is determined to be non-original.

[0029] According to another aspect of the present invention, a video recognition device is provided, comprising:

[0030] The determination module is used to determine the audio-visual information and video information of the video to be recognized;

[0031] The fusion module is used to perform feature fusion on the voice text information and the image information to obtain the fused video features of the video to be identified;

[0032] The determination module is used to determine the originality determination result of the video to be identified based on the fused video features and a preset video database.

[0033] According to another aspect of the present invention, a video recognition device is provided, comprising:

[0034] The processor, memory, communication interface, and communication bus are provided, wherein the processor, memory, and communication interface communicate with each other via the communication bus.

[0035] The memory is used to store at least one executable instruction that causes the processor to perform the operation of the video recognition method as described in any of the preceding claims.

[0036] According to another aspect of the present invention, a computer-readable storage medium is provided, the storage medium storing at least one executable instruction that causes a video recognition device to perform the operation of the video recognition method as described in any of the preceding claims.

[0037] This invention, through its embodiments, determines the audio-text information and video information of a video to be identified. The video information in this embodiment may include both image information and text information. Feature fusion is performed on the audio-text information and the video information to obtain fused video features of the video to be identified. Based on the fused video features and a pre-set video database, the originality determination result of the video to be identified is determined. By fusing audio-text information with video information including text information, and in addition to repeat video recognition based on the image content of the video, image text and audio-text are added as dimensions for video recognition. The fused video features of the video to be identified are comprehensively determined through multi-dimensional feature information such as the image and text corresponding to the video. By comparing these fused video features with the features of original videos in the pre-stored video database under the corresponding dimensions, the accuracy of video originality recognition can be improved, the video duplication rate on video platforms can be reduced, and the video quality and user viewing experience on video playback platforms can be enhanced.

[0038] The above description is merely an overview of the technical solutions of the embodiments of the present invention. In order to better understand the technical means of the embodiments of the present invention and to implement them in accordance with the contents of the specification, and to make the above and other objects, features and advantages of the embodiments of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description

[0039] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:

[0040] Figure 1 A flowchart illustrating the video recognition method provided in an embodiment of the present invention is shown;

[0041] Figure 2 This diagram illustrates the marked optional recognition regions in the video recognition method provided by an embodiment of the present invention;

[0042] Figure 3 This diagram illustrates overlapping video frames of several sizes in the video recognition method provided by an embodiment of the present invention.

[0043] Figure 4 This diagram illustrates the segmentation of the target recognition region in the video recognition method provided by an embodiment of the present invention;

[0044] Figure 5 A schematic diagram of the structure of the video recognition device provided in an embodiment of the present invention is shown;

[0045] Figure 6 A schematic diagram of the structure of the video recognition device provided in an embodiment of the present invention is shown. Detailed Implementation

[0046] Exemplary embodiments of the invention will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the invention are shown in the drawings, it should be understood that the invention can be implemented in various forms and should not be limited to the embodiments set forth herein.

[0047] Before describing the embodiments of the present invention, the prior art and its existing problems will be further explained:

[0048] The contemporary internet is an era where content is king. With technological advancements and network development, the creation of short videos for self-media has become a trend. However, some self-media creators, in their initial pursuit of rapid follower growth, often fail to maintain originality and instead plagiarize trending (or popular) videos. To ensure rights such as recommendation stability and originality, video originality detection and deduplication have become essential steps on short video platforms. Existing video deduplication solutions primarily involve analyzing keyframes and using deep learning and other technologies to compare image similarities. If similar, the video is considered a duplicate; otherwise, it is not. However, these methods are essentially limited to analyzing and comparing video frames.

[0049] In summary, existing technical solutions are based solely on video keyframes. They use deep learning algorithms and other techniques to compare a series of keyframes. If a keyframe is similar to an existing video stored in a database, it's considered a duplicate; otherwise, it's considered acceptable. The problem is that when users alter the video footage, duplication cannot be detected, creating loopholes for plagiarists. For example, existing videos can be "packaged" with watermarks, filters, stickers, and background music, then uploaded as original works. These pseudo-creative short videos lead to excessive video duplication on video platforms, lowering the overall video quality and impacting the user experience. Furthermore, they significantly discourage original videos, affecting the user's video creation experience.

[0050] Therefore, a method for accurately identifying non-original videos is needed to solve the problem that existing technologies cannot effectively identify duplicate videos, thus affecting the user's video viewing experience.

[0051] Figure 1 A flowchart of a video recognition method provided in an embodiment of the present invention is shown. This method is executed by a computer processing device. The computer processing device may include a mobile phone, a laptop computer, etc. Figure 1 As shown, the method includes the following steps:

[0052] Step 10: Determine the audio-visual information and video information of the video to be recognized.

[0053] In one embodiment of the present invention, the video to be identified can be a video uploaded by a user in a target application, such as a short video clip created by the user. The target application can be a short video playback application, etc. The voice and text information includes text information converted from the audio track information contained in the video to be identified. The image information includes the image frame content information of the video to be identified, wherein the image frame content information includes image content information and text content information of the image. The image content information can be obtained by image recognition of the image frame, and the text content information can be obtained by text recognition technology such as OCR (Optical Character Recognition) of the image frame.

[0054] Specifically, when performing OCR recognition on image frames, to improve recognition efficiency and accuracy, several optional recognition regions can be pre-marked for each image frame. This includes marking the center position and four coordinate quadrants of the image frame, ensuring the optional recognition regions cover the center and surrounding areas. Then, using the center of the image frame as the scaling center, the image frame is scaled to obtain several overlapping image frames of different sizes with the same center. The optional recognition regions in the overlapping image frames are then filtered and merged based on their similarity, ultimately resulting in a smaller number of target recognition regions that can cover the text content of the image frames. The similarity of the optional recognition regions can be determined based on their overlap and the similarity of the text content within each region.

[0055] Therefore, in one embodiment of the present invention, the screen information includes screen text information; step 10 further includes:

[0056] Step 101: For each video frame included in the video to be identified, mark several selectable recognition regions in the video frame to obtain the marked video frame.

[0057] In one embodiment of the present invention, in order to ensure the coverage of the recognition results, the selection of the optional recognition area can be as follows: Figure 2 As shown, the video frame is divided into sections with the center of the video frame as the origin and lines parallel to the length and width of the video frame as the coordinate axes. Figure 2 The four coordinate quadrants are denoted as quadrants I, II, III, and IV. Selectable recognition regions are marked at the center, within each quadrant, and on the boundaries between quadrants, as shown below. Figure 2 Within each coordinate quadrant, two optional identification regions with overlapping centers are marked in a cross shape (e.g., P1 and P2 in quadrant I). At least one optional identification region is marked on the boundary line between the coordinate systems of each quadrant (e.g., ...). Figure 2P9 and P10 in the video frame are marked with two overlapping optional recognition regions in the form of a cross at the center of the video frame. Figure 2 (P12 and P11 in the text).

[0058] Specifically, let the width of each video frame in the FMs set be W and the height be H. Two sets of positioning regions with aspect ratios of 1:4 and 4:1 are selected. Positioning sampling is performed within the W*H video frames, resulting in a total of 2*(5+1) = 12 regions, as follows: Figure 2 As shown. "5" represents the four quadrants plus the center position in the video frame. "1" indicates a position where subtitles (or video text) frequently appear, such as... Figure 2 The positions P9 and P10 in the text are used to add specific analysis to video frames to avoid missing text content.

[0059] It should be noted that, in order to improve data processing efficiency while ensuring recognition accuracy, unlike existing technologies that only select key frames of the video for analysis, before marking the selectable recognition region, video frames in the video to be recognized can be filtered according to the frame rate to compress the number of video frames. Specifically, assuming the video length to be recognized is TL seconds and the frame rate is SF, then the video has a total of TL*SF video frames. The SF / 2-1th frame per second is taken as the set of compressed video frames, thus obtaining TL video frames, denoted as the FMs video frame set. Considering that the text in the video frames analyzed in this embodiment is video content that the human eye can clearly see, the video frames used for text recognition need to last for 1 second or more.

[0060] Step 102: Using the center of the marked video frame as the scaling center point, perform multiple scaling processes on the marked video frame to obtain several overlapping video frames of different sizes corresponding to the marked video frame.

[0061] In one embodiment of the present invention, it is readily understood that when scaling is performed with the center of the video frame as the scaling center point, the optional marker regions in the video frame will move closer to or further away from the common scaling center as the video frame is scaled. This results in overlap between optional marker regions. Therefore, based on the overlap, several optional marker regions included in the overlapping video frames can be simplified and aggregated to obtain the target recognition region. The target recognition region obtained through multiple scaling and aggregation can cover the entire video frame, thereby improving both detection efficiency and detection accuracy.

[0062] Specifically, the optional identification area of ​​the marker can be as follows: Figure 2 The diagram shows 12 regions P1-P12. Figure 2This represents 12 selectable recognition regions at the initial scale without scaling. Then... Figure 2 Based on this, the video frames are scaled using an n-layer pyramid scheme (the scaling ratio can be adjusted according to the actual size of the video frames, such as 0.8), resulting in 12 detection regions at the remaining scales. When n is 3, the three-layer pyramid scaling is as follows: Figure 3 As shown, after scaling, we get Figure 3 The three overlapping video frames of different sizes shown contain a total of 3*12=36 text calculation regions. Among them, Figure 3 The outermost layer of the Middle Pyramid, S3, is... Figure 2 The initial scale of the selectable detection area is shown. The rectangle in the middle layer S2 represents the result of the video frame after scaling by 0.8, and correspondingly, the rectangle in the innermost layer S1 represents the result of the original video frame after scaling by 0.8*0.8. Through pyramid scaling, different areas of the video frame can be covered using two regions with aspect ratios of 1:4 and 4:1.

[0063] Step 103: Merge the selectable recognition regions in several overlapping video frames of different sizes to obtain at least one target recognition region corresponding to the video frame.

[0064] In one embodiment of the invention, the selectable recognition region in the overlapping video frames is also scaled accordingly, such as... Figure 3 As shown, optional recognition regions may overlap between video frames of different sizes. Therefore, optional recognition regions can be merged based on the overlap area and the text content included in each optional recognition region to obtain at least one target recognition region. Specifically, when merging optional recognition regions, a group of optional recognition regions with an overlap area higher than a preset area threshold can be identified first. Then, the text similarity of the optional recognition regions within the group can be determined. When the text similarity rate is greater than a preset similarity threshold, it can be determined that the optional recognition regions with high similarity are different detection regions detecting the same text in similar regions. Therefore, the optional recognition regions can be merged, that is, the optional recognition regions with a text similarity higher than the preset similarity threshold can be merged into one target recognition region.

[0065] Therefore, in one embodiment of the present invention, step 103 further includes:

[0066] Step 1031: Based on the overlapping area between the plurality of optional recognition regions and the text recognition result corresponding to each optional recognition region, the optional recognition regions are merged to obtain the target recognition region.

[0067] In one embodiment of the present invention, selectable recognition regions with overlapping areas greater than a preset area threshold are determined as associated recognition regions. The similarity of text content within the associated recognition regions is further determined. When the similarity of text content is greater than a preset similarity threshold, the associated recognition regions are merged to obtain the target recognition region.

[0068] Therefore, in one embodiment of the present invention, step 1031 further includes:

[0069] Step 10311: The optional recognition region whose overlapping area is greater than a preset area threshold and whose similarity of the text recognition result is greater than a preset similarity threshold is determined as the associated recognition region.

[0070] In one embodiment of the present invention, text recognition is performed on each optional recognition region to obtain its corresponding text recognition result. The text recognition results of optional recognition regions with overlapping area values ​​greater than an area threshold are compared pairwise to obtain associated recognition regions.

[0071] Specifically, in combination Figure 3 The process of merging associated identification regions is explained below. First, regarding... Figure 3 The 36 selectable recognition regions included in the three overlapping video frames of different sizes are each processed by OCR. When no result is output, it means there is no text in that region, and the region is discarded. When text is recognized in that region, regions Fa, Fb, and Fc are identified and output. A loop is then entered, and the number of checks is [number missing]. In the above assumption, n=3, so three judgments are needed (Fa and Fb, Fa and Fc, Fb and Fc). Taking the judgment of Fa and Fb as an example, the overlapping area of ​​the two regions is calculated. When the overlapping area is greater than a threshold, the similarity of the recognized text between Fa and Fb is judged (text matching). When the text similarity rate is greater than a threshold, it is determined that Fa and Fb are different recognition regions that detect the same text in similar regions, and regions Fa and Fb are merged. In summary, region merging is only performed when the regions overlap and the recognized content is similar.

[0072] Step 10312: Merge the associated identification regions to obtain the target identification region.

[0073] In one embodiment of the present invention, for every two associated identification regions, a merged identification region is determined based on the coordinate coverage range of the associated identification regions in the video frame. Specifically, the coordinate coverage range is determined based on the extreme values ​​of the coordinates of the associated identification regions, and the region within the coordinate coverage range is determined as the merged identification region. For example, let the coordinates of the region be (x, y, w, h), where x and y are the upper left corner of the region, and w and h are the width and height of the region. It can be seen that the quadruple (x, y, w, h) can uniformly determine a region in the video frame. The coordinates of the upper left corner (x, y) and the lower right corner (x+w, y+h) of the region are calculated, and the two point values ​​of the two regions are compared respectively. The two minimum values ​​of the upper left corner and the two maximum values ​​of the lower right corner are taken to form the merged identification region, where Wf*Hf is the width and height of the merged identification region.

[0074] Subsequently, the merged recognition region is segmented to obtain several sub-recognition regions. Specifically, as shown in... Figure 4 The merged region is divided horizontally and vertically, with the width of each division set according to the aspect ratio of the selectable recognition regions marked in step 102. Based on the text recognition results corresponding to each sub-recognition region, the sub-recognition regions are merged to obtain the target recognition region.

[0075] In one embodiment of the present invention, when the text recognition result of a sub-recognition region is empty, the sub-recognition regions can be merged and compressed to reduce the range. When the OCR recognition results of two adjacent segmented regions are similar, the two sub-recognition regions are merged, and the width of the merged region is 1.3 times the width of the detection region.

[0076] Step 104: Determine the text information of the screen based on the text recognition information within the target recognition area corresponding to the video frame.

[0077] In one embodiment of the present invention, considering the existence of several video frames and the inclusion of several target recognition regions within each video frame, the following steps are taken: First, for each video frame, the text recognition information within all target recognition regions is concatenated. By iteratively concatenating all target recognition regions, the longest text corresponding to that video frame is obtained and determined as the image text information corresponding to that video frame. Then, duplicates are removed from all video frames to finally obtain the image text information corresponding to the video to be recognized. In each iteration, it is determined whether the text obtained after concatenating the text recognition results of two related recognition regions is semantically complete. If it is complete, the two related recognition regions are merged until the longest text is obtained. Finally, the region corresponding to the longest text is taken as the target recognition region.

[0078] Specifically, assuming there are N target video regions detected in the video frame FM, then N corresponding texts will be output, denoted as Txt1, Txt2, ..., Txtn. Find the longest text among the N texts (if multiple texts have the same length, randomly select one), such as Txtj. Match the remaining texts (Txti, 1 <= i <= n and i != j) with Txtj. If Txti is contained within Txtj, discard Txti; if the end of Txti matches the beginning of Txtj, concatenate Txti and Txtj and assign the result to Txtj; similarly, if the beginning of Txti matches the end of Txtj, concatenate Txtj and Txti and assign the result to Txtj; if Txti and Txtj have no match or few matching values, save Txti separately and continue processing after completing one matching step. After comparing all Txti with Txtj, the resulting Txtj is the longest, and a list of mismatched Txti is also obtained. This list of mismatched Txti is then iterated through again, matching each match with the latest and longest Txtj (as described above). If a match still doesn't exist, the length of each Txti is calculated; those less than one-third the length of Txtj are discarded, while those exceeding one-third are retained. The final Txtj and / or other text lists are returned.

[0079] Next, the text recognition information corresponding to all video frames in the video to be recognized is merged. Specifically, duplicate text recognition information is removed, and the text recognition information obtained after removal is arranged in the order of the video frames to finally obtain the on-screen text information corresponding to the video to be recognized.

[0080] In one embodiment of the present invention, step 10 further includes the extraction of speech and text information from the video to be recognized:

[0081] Step 105: Perform speech recognition on each video frame in the video to be identified to obtain the foreground speech information and background speech information corresponding to each video frame.

[0082] In one embodiment of the present invention, the foreground audio information includes the voice of the main speaker in the video to be identified, such as the voice of an actor or singer in the scene, the voice of a news broadcast or narration, etc., and the background audio information includes the background music and ambient sounds of the video to be identified. The audio track data of the video to be identified is extracted, audio feature recognition is performed on the audio track data, and the audio track data is divided into foreground audio information and background audio information based on the audio feature recognition results.

[0083] Optionally, the system can also match the preset background sound database with the audio track data to extract the matched background voice information.

[0084] Step 106: Perform text conversion on the foreground speech information and background speech information respectively to obtain foreground speech text and background speech text.

[0085] In one embodiment of the present invention, the foreground speech information is processed by speech-to-text to obtain foreground speech text, such as the dialogue text of an actor in the video to be identified or the text of narration. The background speech information is processed by speech-to-text to obtain background speech text, such as the lyrics text of background music contained in the video to be identified.

[0086] Step 1043: Perform deduplication processing on the foreground speech text based on the background speech text to obtain the speech text information.

[0087] In one embodiment of the present invention, background audio text is compared with foreground audio text, and text information identical to background audio text is removed from the foreground audio text to obtain audio text information. This reduces the influence of background noise on the image text information extracted from the video to be processed, and avoids the situation where the video content is substantially the same but different background music is used, resulting in inaccurate deduplication.

[0088] Step 20: Perform feature fusion on the voice text information and the video information to obtain the fused video features of the video to be identified.

[0089] In one embodiment of the present invention, feature fusion is performed on the speech-text information and the image information from two dimensions: image and text. Specifically, when performing feature fusion from the text dimension, the speech-text information and the image-text information are compared. The text information that exists in both is determined as the first matching text information. The remaining unmatched information in the image-text information and speech-text information is approximated, and the approximated information is matched again. The common text information obtained from the two matches is added to the fused video features. Simultaneously, image features are extracted from the image to obtain image feature information, which is added to the fused video features. Therefore, by comparing video features from multiple dimensions, including image feature information, speech, and common text information from the image, video deduplication can be performed, improving the accuracy of video deduplication recognition.

[0090] Therefore, in one embodiment of the present invention, step 20 further includes:

[0091] Step 201: Match the voice text information with the screen text information to obtain the first matching text and the non-matching text information.

[0092] In an embodiment of the present invention, considering that the text and the voice commentary in the video are both for expressing the video content and serving the main theme, they should be matched, such as being basically the same. The first matching text includes the voice text and the matching text part in the picture text, and the unmatched text information is the remaining text part after removing the first matching text from the voice text information and the picture text information. Among them, the matching of the voice text and the picture text can be that the texts of the two are exactly the same or the similarity is greater than a preset threshold.

[0093] Step 202: Match the approximate text corresponding to the unmatched text information to obtain a second matching text; the approximate text is obtained by performing at least one of phonetic approximation processing, shape approximation processing, and semantic approximation processing on the unmatched text.

[0094] In an embodiment of the present invention, in order to improve the accuracy of the matching of the voice text and the picture text information, so as to ensure the integrity of the subsequent extraction of the fusion feature information for the video to be recognized, considering that the videos made or uploaded by users may have problems such as inaccurate pronunciation, typos, or the subtitles not matching the voiceover, these situations will cause the picture text information and the voice text information to be unmatched. Therefore, the approximate text of the unmatched text can be matched again to avoid the unmatched caused by the above situations, so as to avoid the omission of video features. Among them, phonetic approximation processing includes determining words with similar pronunciations in the text, such as "食用 (shí yòng)" and "使用 (shǐ yòng)". Shape approximation processing includes determining words with similar shapes in the text, such as "人 (rén), 入 (rù), 八 (bā)" and "己 (jǐ), 已 (yǐ), 巳 (sì)". Semantic approximation processing includes determining words with similar meanings in the text, such as "秀美 (xiù měi)" and "优美 (yōu měi)". The same text in the approximate texts corresponding to the voice text information and the picture text information in the unmatched text information is determined as the second matching text. For example, when the unmatched text information includes the voice text information "食用 (shí yòng)", words such as "使用 (shǐ yòng)", "试用 (shì yòng)", and "适用 (shì yòng)" are added to the second matching text.

[0095] In yet another embodiment of the present invention, let the text content obtained in the first stage (video frame text extraction and analysis) be TCI(Image), and let the text content obtained in the second stage (video voice extraction and analysis) be TCA(Audio). Perform a free text pattern matching between TCI and TCA, similar to the process of paper plagiarism checking. Let TCI be the master text, split TCA into sentences, judge whether the sentence appears in TCI, and if it appears, calculate the position where it appears and the number of text characters (counted by the number of Chinese characters). It should be emphasized that: when judging whether a sentence appears in TCI, there is a core "phonetic similarity processing" (based on the principle of text reading), which effectively avoids the transcription errors in speech recognition during text transcription, such as "practical" being transcribed as "edible" after recognition. Let the TCI text be [asbserRTGSUJLLIfvnfurLRUTLS], with one letter representing one Chinese character. Among them, a sentence in TCA is [JLLIfVMfur]. Through text matching (A = A, a = a, A!= a), it is easy to obtain the similarity result between [JLLIfVMfur] and [asbserRTGSUJLLIfvnfurLRUTLS]. When the vast majority of Chinese characters in the sentence are matched and the number of unmatched characters is small, perform "phonetic similarity processing" on the unmatched characters, that is, perform pronunciation processing on [VM] ("reading", but without making a sound), and find the corresponding text position in TCI according to the similarity result, and take out the text [vn] for pronunciation processing ("reading", but without making a sound). When the results of the pronunciation processing of [VM] and [vn] are the same, it is determined that [VM] and [vn] are the same, and then correct the matching of the sentence in the TCI text to [asbserRTGSUJLLIfvnfurLRUTLS].

[0096] Step 203: Determine the fused video feature according to the first matching text and the second matching text.

[0097] In one embodiment of the present invention, the set of the first matching text and the second matching text is determined as the fused video feature.

[0098] Step 30: Determine the originality determination result of the video to be recognized according to the fused video feature and a preset video database.

[0099] In one embodiment of the present invention, the preset video database includes a number of pre-stored original videos, where the original videos include videos uploaded by other users in history and certified as original. The originality determination result is used to determine whether the video to be recognized is an original video. Among them, when its feature overlap with the videos in the video database is higher than a certain threshold, it is determined that the video to be recognized is a non-original video.

[0100] In another embodiment of the present invention, in order to improve the accuracy of video deduplication recognition, the video features of several dimensions such as fused video features, voice and text information, and image information can be compared. When the overlap of video features in at least one dimension is higher than the threshold, the video to be recognized is determined to be a non-original video.

[0101] Specifically, the video database includes at least one of the fused video features, voice and text information, and image information of several pre-stored original videos;

[0102] Step 30 further includes: Step 301: Match the fused video features with at least one of the fused video features, voice and text information, and image information of the pre-stored original video. When the match is successful, determine that the originality determination result is not original.

[0103] In one embodiment of the present invention, if the fused video features overlap with one or more of the pre-stored fused video features, speech-text information, and image information, then the match is determined to be successful, i.e., the video to be identified is determined to lack originality. Correspondingly, if the fused video features do not match any of the pre-stored fused video features, speech-text information, and image information of the original video, then the originality identification result of the video to be identified is determined to be original.

[0104] This invention, through its embodiments, determines the audio-text information and video information of a video to be identified. The video information in this embodiment may include both image information and text information. Feature fusion is performed on the audio-text information and the video information to obtain fused video features of the video to be identified. Based on the fused video features and a pre-set video database, the originality determination result of the video to be identified is determined. By fusing audio-text information with video information including text information, and in addition to repeat video recognition based on the image content of the video, image text and audio-text are added as dimensions for video recognition. The fused video features of the video to be identified are comprehensively determined through multi-dimensional feature information such as the image and text corresponding to the video. By comparing these fused video features with the features of original videos in the pre-stored video database under the corresponding dimensions, the accuracy of video originality recognition can be improved, the video duplication rate on video platforms can be reduced, and the video quality and user viewing experience on video playback platforms can be enhanced.

[0105] Figure 5 A schematic diagram of the structure of a video recognition device provided in an embodiment of the present invention is shown. Figure 5 As shown, the device 40 includes: a determination module 401, a fusion module 402, and a judgment module 403.

[0106] The determining module 401 is used to determine the audio-text information and video information of the video to be recognized.

[0107] The fusion module 402 is used to perform feature fusion on the voice text information and the image information to obtain the fused video features of the video to be identified;

[0108] The determination module 403 is used to determine the originality determination result of the video to be identified based on the fused video features and the preset video database.

[0109] The operation process of the video recognition device provided in this embodiment of the invention is largely the same as that of the aforementioned method embodiment, and will not be described again.

[0110] The video recognition device provided in this embodiment of the invention determines the audio-text information and video information of the video to be recognized. The video information in this embodiment may include video image information and video text information. Feature fusion is performed on the audio-text information and the video information to obtain the fused video features of the video to be recognized. Based on the fused video features and a preset video database, the originality determination result of the video to be recognized is determined. By fusing the audio-text information with the video information including video text information, and based on repeated video recognition according to the image content of the video, video text and audio-text are added as dimensions for video recognition. The fused video features of the video to be recognized are comprehensively determined through multi-dimensional feature information such as the image and text corresponding to the video. By comparing the fused video features with the features of original videos in the corresponding dimensions pre-stored in the video database, the accuracy of video originality recognition can be improved, the video duplication rate on the video platform can be reduced, and the video quality and user viewing experience on the video playback platform can be improved.

[0111] Figure 6 The diagram shows a structural schematic of a video recognition device provided in an embodiment of the present invention. The specific embodiments of the present invention do not limit the specific implementation of the video recognition device.

[0112] like Figure 6 As shown, the video recognition device may include: a processor 502, a communications interface 504, a memory 506, and a communications bus 508.

[0113] The processor 502, communication interface 504, and memory 506 communicate with each other via communication bus 508. Communication interface 504 is used to communicate with other network elements such as clients or other servers. The processor 502 executes program 510, specifically performing the relevant steps described above in the video recognition method embodiment.

[0114] Specifically, program 510 may include program code, which includes computer-executable instructions.

[0115] Processor 502 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention. The video recognition device includes one or more processors, which may be processors of the same type, such as one or more CPUs; or they may be processors of different types, such as one or more CPUs and one or more ASICs.

[0116] Memory 506 is used to store program 510. Memory 506 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.

[0117] Specifically, program 510 can be called by processor 502 to cause the video recognition device to perform the following operations:

[0118] Determine the audio-visual information and video information of the video to be recognized;

[0119] The voice text information and the video information are fused to obtain the fused video features of the video to be identified;

[0120] Based on the fused video features and the preset video database, the originality determination result of the video to be identified is determined.

[0121] The operation process of the video recognition device provided in this embodiment of the invention is largely the same as that of the aforementioned method embodiment, and will not be repeated here.

[0122] The video recognition device provided in this embodiment of the invention determines the audio-text information and video information of the video to be recognized. The video information in this embodiment may include image information and text information. Feature fusion is performed on the audio-text information and the video information to obtain the fused video features of the video to be recognized. Based on the fused video features and a preset video database, the originality determination result of the video to be recognized is determined. By fusing the audio-text information with the video information including the text information, and based on repeated video recognition according to the image content of the video, image text and audio-text are added as dimensions for video recognition. The fused video features of the video to be recognized are comprehensively determined through multi-dimensional feature information such as the image and text corresponding to the video. By comparing the fused video features with the features of original videos in the corresponding dimensions pre-stored in the video database, the accuracy of video originality recognition can be improved, the video duplication rate on the video platform can be reduced, and the video quality and user viewing experience on the video playback platform can be improved.

[0123] This invention provides a computer-readable storage medium storing at least one executable instruction that, when executed on a video recognition device, causes the video recognition device to perform the video recognition method described in any of the above method embodiments.

[0124] Specifically, the executable instructions can be used to cause the video recognition device to perform the following operations:

[0125] Determine the audio-visual information and video information of the video to be recognized;

[0126] The voice text information and the video information are fused to obtain the fused video features of the video to be identified;

[0127] Based on the fused video features and the preset video database, the originality determination result of the video to be identified is determined.

[0128] The operation process of the executable instructions stored in the computer storage medium provided in this embodiment of the invention is largely the same as that in the aforementioned method embodiments, and will not be described again.

[0129] The executable instructions stored in the computer storage medium provided in this embodiment of the invention determine the audio-text information and video information of the video to be identified. The video information in this embodiment may include video image information and video text information. Feature fusion is performed on the audio-text information and the video information to obtain the fused video features of the video to be identified. Based on the fused video features and a preset video database, the originality determination result of the video to be identified is determined. By fusing the audio-text information with the video information including video text information, and based on repeated video recognition according to the image content of the video, video text and audio-text are added as dimensions for video recognition. The fused video features of the video to be identified are comprehensively determined through multi-dimensional feature information such as the image and text corresponding to the video. By comparing the fused video features with the features of original videos in the corresponding dimensions pre-stored in the video database, the accuracy of video originality recognition can be improved, the video duplication rate on the video platform can be reduced, and the video quality and user viewing experience on the video playback platform can be improved.

[0130] This invention provides a video recognition device for performing the video recognition method described above.

[0131] This invention provides a computer program that can be called by a processor to cause a video recognition device to execute the video recognition method in any of the above method embodiments.

[0132] This invention provides a computer program product, which includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions, which, when executed on a computer, cause the computer to perform the video recognition method described in any of the above method embodiments.

[0133] The algorithms or displays provided herein are not inherently related to any particular computer, virtual system, or other device. Various general-purpose systems can also be used in conjunction with the teachings herein. The required structure for constructing such systems is apparent from the above description. Furthermore, the embodiments of the present invention are not directed to any particular programming language. It should be understood that the content of the invention described herein can be implemented using various programming languages, and the above description of specific languages ​​is for the purpose of disclosing the best mode of implementation of the invention.

[0134] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of the invention may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.

[0135] Similarly, it should be understood that, in order to simplify the invention and aid in understanding one or more aspects of the invention, features of the embodiments of the invention are sometimes grouped together in a single embodiment, figure, or description thereof in the above description of exemplary embodiments of the invention. However, this disclosure should not be construed as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim.

[0136] Those skilled in the art will understand that modules in the device of the embodiments can be adaptively changed and placed in one or more devices different from those of the embodiments. Modules, units, or components in the embodiments can be combined into a single module, unit, or component, and can be divided into several sub-modules, sub-units, or sub-components. Except where at least some of such features and / or processes or units are mutually exclusive, any combination can be used to combine all features disclosed in this specification (including the accompanying claims, abstract, and drawings) and all processes or units of any method or device so disclosed. Unless expressly stated otherwise, each feature disclosed in this specification (including the accompanying claims, abstract, and drawings) may be replaced by an alternative feature that serves the same, equivalent, or similar purpose.

[0137] It should be noted that the above embodiments are illustrative of the invention and not restrictive, and that those skilled in the art can devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The invention can be implemented by means of hardware comprising a plurality of different elements and by means of a suitably programmed computer. In the unit claims enumerating a plurality of means, a plurality of such means may be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words may be interpreted as names. The steps in the above embodiments, unless otherwise specified, should not be construed as limiting the order of execution.

Claims

1. A video recognition method, characterized in that, The method includes: The audio-visual information and video information of the video to be recognized are determined. The video information includes video-text information. The video-text information is determined through the following steps: For each video frame in the video to be recognized, several optional recognition regions in the video frame are marked to obtain marked video frames; the marked video frames are scaled several times with the center of the marked video frames as the scaling center point to obtain several overlapping video frames of different sizes corresponding to the marked video frames; based on the overlap area between each optional recognition region and the text recognition result corresponding to each optional recognition region, the optional recognition regions are merged to obtain the target recognition region; the video-text information is determined based on the text recognition information in the target recognition region corresponding to the video frame. The voice text information and the video information are fused to obtain the fused video features of the video to be identified; Based on the fused video features and the preset video database, the originality determination result of the video to be identified is determined.

2. The method according to claim 1, characterized in that, The step of merging the optional recognition regions based on the overlap area between each optional recognition region and the text recognition result corresponding to each optional recognition region to obtain the target recognition region includes: The optional recognition regions whose overlapping area is greater than a preset area threshold and whose similarity of the text recognition results is greater than a preset similarity threshold are determined as associated recognition regions; The associated identification regions are merged to obtain the target identification region.

3. The method according to claim 1, characterized in that, The voice text information is determined through the following steps: Speech recognition is performed on each video frame in the video to be identified to obtain the foreground speech information and background speech information corresponding to each video frame in the video to be identified. The foreground speech information and background speech information are respectively converted into text to obtain foreground speech text and background speech text; The foreground speech text is deduplicated based on the background speech text to obtain the speech text information.

4. The method according to claim 1, characterized in that, The video information includes video text information; the fused video features are determined through the following steps: The voice text information is matched with the screen text information to obtain the first matching text and the non-matching text information; The approximate text corresponding to the mismatched text information is matched to obtain the second matching text; the approximate text is obtained by performing at least one of the following on the mismatched text: phonetic similarity processing, shape similarity processing, and semantic similarity processing. The fused video features are determined based on the first and second matching texts.

5. The method according to claim 1, characterized in that, The video database includes at least one of the fused video features, audio-text information, and video information of several pre-stored original videos; the originality determination result is determined through the following steps: The fused video features are matched with at least one of the fused video features, speech and text information, and image information of the pre-stored original video. When the match is successful, the originality determination result is determined to be non-original.

6. A video recognition device, characterized in that, The device includes: A determining module is used to determine the audio-text information and video information of the video to be recognized; the video information includes video-text information; the video-text information is determined through the following steps: for each video frame in the video to be recognized, several optional recognition regions in the video frame are marked to obtain marked video frames; the marked video frames are scaled several times with the center of the marked video frames as the scaling center point to obtain several overlapping video frames of different sizes corresponding to the marked video frames; based on the overlap area between each optional recognition region and the text recognition result corresponding to each optional recognition region, the optional recognition regions are merged to obtain a target recognition region; the video-text information is determined based on the text recognition information in the target recognition region corresponding to the video frame. The fusion module is used to perform feature fusion on the voice text information and the image information to obtain the fused video features of the video to be identified; The determination module is used to determine the originality determination result of the video to be identified based on the fused video features and a preset video database.

7. A video recognition device, characterized in that, include: The processor, memory, communication interface, and communication bus are provided, wherein the processor, memory, and communication interface communicate with each other via the communication bus. The memory is used to store at least one executable instruction that causes the processor to perform the operation of the video recognition method as described in any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The storage medium stores at least one executable instruction, which, when executed on the video recognition device, causes the video recognition device to perform the operation of the video recognition method as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Video recognition method and device and computer readable storage medium

    CN112580599A

  • Original short video AI intelligent detection method, system and device

    CN114928764A