Video identification method, device, equipment, medium and product

By comprehensively considering similarity scores from visual, textual, multimodal, and facial feature dimensions in video recognition, the problem of low recognition accuracy in existing technologies based on a single feature dimension is solved, achieving higher accuracy in video originality recognition.

CN121330587APending Publication Date: 2026-01-13ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511632121.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-07
Publication Date
2026-01-13

AI Technical Summary

Technical Problem

Existing technologies rely on a single feature dimension for video originality recognition, resulting in low recognition accuracy and a high risk of misjudgment due to bias or accidental similarity.

Method used

By obtaining the comprehensive similarity scores of the video to be identified and each candidate video in the candidate video set in terms of visual feature dimensions, text feature dimensions, multimodal feature dimensions, and facial feature dimensions, the originality of the video to be identified is determined by analyzing the comprehensive similarity scores of multiple feature dimensions using a video recognition model.

Benefits of technology

It improves the accuracy of video originality recognition by combining similarity analysis of multiple feature dimensions, reducing false positives and increasing recognition accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121330587A_ABST
    Figure CN121330587A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a video recognition method and device, equipment, a medium and a product. The scheme comprises the steps of obtaining a to-be-recognized video; obtaining a candidate video set for identifying the originality of the video to be identified; for each candidate video in the candidate video set, comprehensive similarity scores between the candidate video and the video to be recognized are determined in at least two target feature dimensions in specific feature dimensions, and a plurality of comprehensive similarity scores are obtained, the specific feature dimensions comprise at least two feature dimensions of a visual feature dimension, a text feature dimension, a multi-modal feature dimension and a face feature dimension; inputting at least part of the generated comprehensive similar scores into a video recognition model to obtain an original video probability value of the to-be-recognized video output by the video recognition model based on the at least part of the comprehensive similar scores; and if the original video probability value is greater than or equal to a preset threshold value, determining the to-be-identified video as the original video.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] One or more embodiments of the present specification relate to the technical field of video recognition, and in particular to a video recognition method. One or more embodiments of the present specification also relate to a video recognition apparatus, a computing device, a computer-readable storage medium, and a computer program product. BACKGROUND

[0002] With the rapid iteration and popularity of video technology, various video playback software emerge in an endless stream, and the market competition of video playback software is increasingly fierce. For a video playback software, original video content has high revenue in various consumption dimensions, and non-original video content, such as videos of carrying or videos edited on the basis of original videos, often has problems such as insufficient clarity, poor timeliness of content, and lack of interactivity. Therefore, it is crucial to identify the originality of the uploaded videos of users.

[0003] Therefore, how to achieve high-precision identification of original videos has become a technical problem to be solved. SUMMARY

[0004] Therefore, one or more embodiments of the present specification provide a video recognition method, apparatus, device, medium, and product to achieve high-precision identification of original videos.

[0005] According to a first aspect of one or more embodiments of the present specification, a video recognition method is provided, including:

[0006] obtaining a to-be-identified video;

[0007] obtaining a candidate video set for identifying originality of the to-be-identified video;

[0008] determining, for each candidate video in the candidate video set, a comprehensive similarity score of the to-be-identified video and the each candidate video in at least two target feature dimensions in a specific feature dimension, to obtain a plurality of comprehensive similarity scores; the specific feature dimension includes at least two feature dimensions in a visual feature dimension, a text feature dimension, a multi-modal feature dimension, and a face feature dimension, and the comprehensive similarity score of one candidate video is obtained by concatenating similarity scores of the one candidate video and the to-be-identified video in the at least two target feature dimensions;

[0009] inputting at least part of the plurality of comprehensive similarity scores into a video recognition model to obtain an original video probability value of the to-be-identified video output by the video recognition model based on the at least part of the comprehensive similarity scores;

[0010] if the original video probability value is greater than or equal to a preset threshold, determining the to-be-identified video as an original video.

[0011] According to a second aspect of one or more embodiments of the present specification, a video recognition device is provided, comprising:

[0012] a first obtaining module configured to obtain a to-be-recognized video;

[0013] a second obtaining module configured to obtain a candidate video set for recognizing originality of the to-be-recognized video;

[0014] a first determining module configured to determine, for each candidate video in the candidate video set, a comprehensive similarity score of the to-be-recognized video and the each candidate video in at least two target feature dimensions in specific feature dimensions, to obtain a plurality of comprehensive similarity scores; the specific feature dimensions comprise at least two feature dimensions in visual feature dimensions, text feature dimensions, multi-modal feature dimensions, and face feature dimensions, and a comprehensive similarity score of a candidate video is obtained by concatenating similarity scores of the candidate video and the to-be-recognized video in the at least two target feature dimensions;

[0015] an input module configured to input at least part of the plurality of comprehensive similarity scores into a video recognition model, to obtain an original video probability value of the to-be-recognized video output by the video recognition model based on the at least part of the plurality of comprehensive similarity scores;

[0016] a second determining module configured to determine the to-be-recognized video as an original video if the original video probability value is greater than or equal to a preset threshold.

[0017] According to a third aspect of one or more embodiments of the present specification, a computing device is provided, comprising a memory, a processor, and computer instructions stored on the memory and executable on the processor, wherein the processor executes the computer instructions to implement steps of the video recognition method.

[0018] According to a fourth aspect of one or more embodiments of the present specification, a computer readable storage medium is provided, which stores computer instructions, wherein the computer instructions are executed by a processor to implement steps of the video recognition method.

[0019] According to a fifth aspect of an embodiment of the present specification, a computer program product is provided, comprising computer programs or instructions, wherein the computer programs / instructions are executed by a processor to implement steps of the above-mentioned video recognition method.

[0020] The one embodiment of the specification can achieve at least the following beneficial effects: by determining the comprehensive similarity scores of each candidate video in the candidate video set and the to-be-identified video in at least two target feature dimensions in specific feature dimensions, obtaining a plurality of comprehensive similarity scores, wherein the specific feature dimensions include at least two feature dimensions in a visual feature dimension, a text feature dimension, a multi-modal feature dimension, and a face feature dimension, inputting at least part of the comprehensive similarity scores into a video recognition model, and enabling the video recognition model to analyze the originality of the to-be-identified video based on the at least part of the comprehensive similarity scores. Since the comprehensive similarity score is a similarity score generated by splicing each similarity score in at least two feature dimensions, the comprehensive similarity score can reflect the similarity between the to-be-identified video and the candidate video in multiple feature dimensions, so that the originality of the to-be-identified video can be identified in multiple feature dimensions, thereby improving the recognition accuracy of the originality identification of the to-be-identified video. In addition, at least part of the comprehensive similarity scores between the to-be-identified video and at least part of the candidate videos are provided to the video recognition model, so that the video recognition model can comprehensively analyze the originality of the to-be-identified video based on the cross comparison of the plurality of candidate videos, thereby further improving the recognition accuracy of the originality identification of the to-be-identified video. BRIEF DESCRIPTION OF DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the specification or the prior art, the following will briefly introduce the drawings needed to be used in the embodiment or prior art description. Obviously, the drawings in the following description are only some embodiments of the specification, and for those skilled in the art, other drawings can also be obtained without creative labor.

[0022] Figure 1 is an application scenario schematic diagram of a video recognition method provided by an embodiment of the specification;

[0023] Figure 2 is a flow schematic diagram of a video recognition method provided by an embodiment of the specification;

[0024] Figure 3 is a process schematic diagram of originality identification of a product explanation video provided by an embodiment of the specification;

[0025] Figure 4 is a structural schematic diagram of a video recognition device provided by an embodiment of the specification;

[0026] Figure 5 is a structural block diagram of a computing device provided by an embodiment of the specification. DETAILED DESCRIPTION

[0027] In order to better understand the technical solutions in the specification, the technical solutions in the specification will be clearly and completely described below in combination with the drawings in the specification. Obviously, the described embodiments are only some of the embodiments of the specification, not all the embodiments. Based on the embodiments in the specification, all other embodiments obtained by those of ordinary skill in the art without creative labor should be within the scope of protection of the specification.

[0028] The specification uses specific terms to describe the embodiments of the specification. As "one embodiment", "an embodiment", and / or "some embodiments" means a certain feature, structure or characteristic related to at least one embodiment of the specification. Therefore, it should be emphasized and noted that the "an embodiment" or "one embodiment" or "one alternative embodiment" mentioned in the specification twice or more in different positions does not necessarily mean the same embodiment. In addition, those skilled in the art can combine and combine different embodiments or examples described in the specification and the features of different embodiments or examples without contradiction.

[0029] The terms used in one or more embodiments of the specification are only for the purpose of describing specific embodiments, and are not intended to limit one or more embodiments of the specification. The singular forms "a", "an", "one", "said" and "the" used in one or more embodiments of the specification and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of the specification includes any or all possible combinations of one or more associated listed items.

[0030] The term "include", "contain" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, product or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or includes elements inherent to such process, method, product or device. Without more limitations, it does not exclude the presence of other same or equivalent elements in the process, method, product or device including the elements.

[0031] Although the terms first, second, etc. can be used herein to describe various information, these terms are not intended to denote a temporal sequence, but to distinguish between different types of information. For example, a first entity discussed below could be termed a second entity without departing from the scope of one or more embodiments of the present disclosure. The first, second, etc. designations can be used herein for purposes of

[0032] Depending on the context, the word "if" as used herein can be interpreted to mean "when" or "upon" or "in response to determining."

[0033] In the present specification, unless specifically stated otherwise, "receiving and sending data" can be direct receiving and sending, or indirect receiving and sending. For example, A device receiving data sent by B device can be understood as A device directly receiving data sent by B device, or A device indirectly receiving data sent by B device through C device or other subject. Similarly, B device sending data to A device can be understood as B device directly sending data to A device, or B device indirectly sending data to A device through C device or other subject. Here, C device can be one subject, or two or more subjects.

[0034] In the present specification, unless specifically stated otherwise, the association relationship between structures can be direct or indirect. For example, when describing "A device is connected with B device", unless it is specifically stated that A device is directly connected with B device, it should be understood that A device can be directly connected with B device, or indirectly connected with B device. For another example, when describing "A device is on B device", unless it is specifically stated that A device is directly on B device (A device is adjacent to B device and A device is on B device), it should be understood that A device can be directly on B device, or A device can be indirectly on B device (A device is separated from B device by other elements, and A device is on B device). Similarly.

[0035] ​The user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of the present specification are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant laws, regulations and standards in the relevant region, and provide corresponding operation portal for user to choose authorization or refusal.

[0036] In the prior art, when identifying the originality of a video, it is often limited to a single feature dimension, such as a multi-modal feature dimension or a visual feature dimension, and the to-be-identified video is compared with all videos in the video library one by one in the feature dimension, and the target video with the highest similarity is filtered out. If the similarity of the target video and the to-be-identified video in this dimension meets the preset standard, it can be determined that the to-be-identified video is an original video.

[0037] However, this judgment method which depends on a single feature dimension and takes the most similar target video as a reference ignores the comparison information of other feature dimensions and does not consider other videos that may exist in association in the video library, and is prone to misjudgment due to the one-sidedness of a single dimension or the accidental similarity of an individual video. The recognition accuracy of this recognition method is relatively low.

[0038] The technical solutions provided by the embodiments of the present specification will be described in detail below with reference to the accompanying drawings.

[0039] Figure 1 is an application scenario diagram of a video recognition method provided by an embodiment of the present specification.

[0040] As shown in Figure 1 The application scenario diagram includes a terminal device 101 and a server 102.

[0041] In an embodiment of the present specification, the terminal device 101 can include but is not limited to at least one of a smart phone, a tablet computer, a notebook computer, a smart interactive device, a wearable device, a vehicle-mounted intelligent terminal, etc. The wearable device can include but is not limited to at least one of a smart bracelet, a smart watch, smart glasses, etc.

[0042] The server 102 can include but is not limited to at least one of any device, apparatus, platform, device cluster or cloud computing service center having computing and processing capabilities.

[0043] A communication connection is established between terminal device 101 and server 102. This communication connection can be, but is not limited to, a local area network (LAN) connection, a wide area network (WAN) connection, an internet connection, a short-range communication (SMR) connection, or other types of data network connections. Short-range communication connections include, but are not limited to, Near Field Communication (NFC), LAN, Bluetooth, and infrared connections. Users can send a recognition request for a video to be recognized to server 102 via terminal device 101. Server 102 responds to the recognition request by performing originality recognition on the user-provided video and displays the result to the user via terminal device 101. It is understood that if terminal device 101 meets the requirements for deploying the recognition program corresponding to the aforementioned video recognition method, it can also recognize user-provided videos.

[0044] This application provides a video recognition method, and also relates to a video recognition device, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.

[0045] Figure 2 This is a flowchart illustrating a video recognition method provided in an embodiment of this specification.

[0046] From a programming perspective, the entity executing the process can be a program mounted on a device used for video recognition. From a hardware perspective, this method can be executed by any device, equipment, platform, or cluster of devices with computing and processing capabilities; for example, it can be a terminal device, a server, or both.

[0047] like Figure 2 As shown, the process may include the following steps:

[0048] Step 202: Obtain the video to be identified.

[0049] In the embodiments of this specification, the video to be identified can be a video pre-uploaded by the user or already uploaded to the video playback software. The types of videos can include product or service explanation videos, home life skills videos, pet raising tutorial videos, workplace skills training videos, traditional intangible cultural heritage display videos, and other videos with specific content attributes that are available for user viewing.

[0050] In actual applications, the recognition system can obtain a to-be-recognized video submitted by a user, or the recognition system can also obtain a to-be-recognized video sent by a video playing software, where the recognition system can be a system for video recognition based on the foregoing video recognition method. The user can submit the to-be-recognized video to the recognition system, the recognition system can generate a corresponding recognition result for the to-be-recognized video, and the recognition system can feed back the generated recognition result to the user. If the recognition result reflects that the to-be-recognized video does not have a security risk and meets the standard of an original video, the user can upload the to-be-recognized video and the corresponding recognition result to the video playing software. The recognition system can also determine whether to upload the to-be-recognized video to the video playing software based on the recognition result for the to-be-recognized video. If the recognition result indicates that the to-be-recognized video does not have a security risk and meets the standard of an original video, the recognition system can upload the to-be-recognized video to the video playing software. If the recognition result indicates that the to-be-recognized video has a security risk and / or does not meet the standard of an original video, the recognition system can not upload the to-be-recognized video to the video playing software.

[0051] The user can also directly upload the to-be-recognized video to the video playing software. If the to-be-recognized video is not associated with a corresponding recognition result, the video playing software can send the to-be-recognized video to the recognition system for recognition, determine whether to play the to-be-recognized video based on the recognition result returned by the recognition system, and if the recognition result indicates that the to-be-recognized video does not have a security risk and meets the standard of an original video, the video playing software can play the to-be-recognized video. If the recognition result indicates that the to-be-recognized video has a security risk and / or does not meet the standard of an original video, the video playing software can not play the to-be-recognized video.

[0052] Step 204: Obtain a candidate video set for identifying the originality of the to-be-recognized video.

[0053] In the embodiments of the present specification, after the recognition system obtains a recognition request for a to-be-recognized video, the recognition system can obtain a candidate video set for identifying the originality of the to-be-recognized video. Originality can be understood as whether the core topic, material, and presentation form of the to-be-recognized video are original to the user, rather than directly copying or simply modifying the works of others. The candidate videos in the candidate video set can be videos determined based on the to-be-recognized video, and the video similarity between each candidate video and the to-be-recognized video can be greater than or equal to a preset threshold. In other words, the candidate video can be a video that has a certain similarity with the to-be-recognized video after preliminary screening. By using such a candidate video with similarity to the to-be-recognized video for subsequent originality identification, interference from meaningless low-similarity videos can be avoided, so that the recognition process can focus on effective comparison objects, thereby improving the accuracy of originality identification for the to-be-recognized video.

[0054] Step 206: determining, for each candidate video in the candidate video set, a comprehensive similarity score of the to-be-identified video and the each candidate video in at least two target feature dimensions in a specific feature dimension, to obtain a plurality of comprehensive similarity scores.

[0055] The specific feature dimension includes at least two of a visual feature dimension, a text feature dimension, a multi-modal feature dimension, and a face feature dimension, and the comprehensive similarity score of a candidate video is obtained by concatenating similarity scores of the candidate video and the to-be-identified video in the at least two target feature dimensions.

[0056] In the embodiments of the present specification, for each candidate video in the candidate video set, a video pair can be constructed with the to-be-identified video, and a comprehensive similarity score can be generated for each video pair. The comprehensive similarity score is a similarity score generated by concatenating similarity scores in at least two target feature dimensions in a specific feature dimension, wherein a similarity score is calculated for each of the at least two target feature dimensions.

[0057] In the embodiments of the present specification, the visual feature dimension can be understood as a dimension for extracting features of visual elements in a video. The visual elements can be all visual elements that can be directly perceived by the human eye and constitute the basic constituent part of the picture content, and are the elements that carry picture information and embody picture differences. The text feature dimension can be understood as a dimension for extracting features of text information in a video. The text information can include explicit text in the video and text extracted based on audio. The multi-modal feature dimension can be understood as a dimension for extracting features of fusion information obtained by fusing the above-mentioned visual elements and text information. The face feature dimension can be understood as a dimension for extracting features of face information in a video, wherein the face information can include facial contours, eye corner positions, mouth corner positions, nose wing positions, and ear spacing.

[0058] Step 208: inputting at least part of the plurality of comprehensive similarity scores into a video recognition model to obtain an original video probability value of the to-be-identified video output by the video recognition model based on the at least part of the comprehensive similarity scores.

[0059] In the embodiments of the present specification, the at least part of the comprehensive similarity scores can be two comprehensive similarity scores, three comprehensive similarity scores, or comprehensive similarity scores of other numbers. It can be understood that the at least part of the comprehensive similarity scores can also be one comprehensive similarity score under the condition of meeting the user identification accuracy requirement, and no limitation is made thereto.

[0060] In the embodiments of the present specification, the video recognition model can be a model deployed in the identification system or a model deployed outside the identification system, and the identification system can call the model through instructions. The video recognition model can be a binary classification task model or a multi-classification task model.

[0061] In actual application, in the scenario where the video recognition model is a binary classification task model, the output result of the video recognition model can include a probability value of the to-be-identified video being an original video and a probability value of the to-be-identified video being a non-original video.

[0062] In actual application, in the scenario where the video recognition model is a multi-classification task model, the output result of the video recognition model can be more detailed. In addition to the probability value of the original video, the types of non-original videos can be further divided, such as probability values of original two-cut videos, low-creation two-cut videos, and carrying videos. The original video can be understood as a to-be-identified video that has no repeated content with all videos in the video library, and the core idea, presentation form, and video content of the video are independently created by the current video publisher. The original two-cut video can be understood as a to-be-identified video that has similar content with a certain video in the video library, but the current video publisher adds significant gain content to the video based on the similar content, such as adding original commentary content or supplementing in-depth analysis content. The low-creation two-cut video can be understood as a to-be-identified video that has similar content with a certain video in the video library, but the modification made by the current video publisher to the video is a surface modification and does not produce significant gain, such as adding transition effects. The transition effect is used to connect the transition effect of two different shots or video clips, and the core function is to switch the picture from one segment to another segment naturally, avoiding the harshness of direct jumping. The carrying video can be understood as a to-be-identified video that has a repetition degree greater than or equal to a preset threshold with a certain video content in the video library, and there is no substantial difference between the two.

[0063] Step 210: If the original video probability value is greater than or equal to a preset threshold, the to-be-identified video is determined as an original video.

[0064] In the embodiments of the present specification, according to the output result of the video recognition model, the originality of the to-be-identified video can be determined. If the original video probability value is greater than or equal to a preset threshold, the to-be-identified video can be determined as an original video, and the identification system can add an original video label to the to-be-identified video.

[0065] Although the method steps are provided in the embodiments or flowcharts of the specification, it can be understood that the order of the steps listed in the embodiments or flowcharts is only one of the many execution orders, and does not represent the only execution order. The order of some steps can be adjusted according to actual needs, or some steps can be omitted. When the claims refer to the method steps, the adjustment of the order of the steps or the parallelism between the steps is also within the protection scope of the claims.

[0066] Figure 2 In the method, the comprehensive similarity score between the to-be-identified video and the candidate video is generated based on the splicing of the respective similarity scores in at least two feature dimensions, so that the comprehensive similarity score can reflect the similarity of the to-be-identified video and the candidate video in multiple feature dimensions, thereby enabling the originality of the to-be-identified video to be identified in multiple feature dimensions, so as to improve the identification accuracy of the originality identification of the to-be-identified video. In addition, at least part of the comprehensive similarity scores between the to-be-identified video and at least part of the candidate videos are provided to the video identification model, so that the video identification model can comprehensively analyze the originality of the to-be-identified video based on the cross comparison of the multiple candidate videos, thereby further improving the identification accuracy of the originality identification of the to-be-identified video.

[0067] Based on the method of Figure 2 The embodiments of the specification also provide some improved implementations of the method, which are described below.

[0068] In one or more embodiments of the specification, the comprehensive similarity score refers to both the global features of the to-be-identified video, such as the multi-modal features, and the local features of the to-be-identified video, such as the visual features, the text features, and the like.

[0069] Optionally, the determining, for each candidate video in the set of candidate videos, a comprehensive similarity score between the to-be-identified video and the each candidate video in at least two target feature dimensions of the specific feature dimensions can specifically include: determining, for any candidate video in the each candidate video, a first similarity score between the to-be-identified video and the any candidate video in the visual feature dimension; determining a second similarity score between the to-be-identified video and the any candidate video in the text feature dimension; determining a third similarity score between the to-be-identified video and the any candidate video in the multi-modal feature dimension; and splicing the first similarity score, the second similarity score, and the third similarity score to obtain the comprehensive similarity score between the to-be-identified video and the any candidate video.

[0070] In the embodiments of the present specification, in the process of identifying the to-be-identified video, a comprehensive similarity score is generated for each candidate video in the candidate video set. The method for generating a comprehensive similarity score for other candidate videos in the candidate video set can refer to the method for generating a comprehensive similarity score for any candidate video.

[0071] In one or more embodiments of the present specification, a specific embodiment of determining the first similarity score in the visual feature dimension is also proposed.

[0072] Optionally, the determining the first similarity score of the to-be-identified video and the any candidate video in the visual feature dimension can specifically include: extracting a first quantity of first video frames from the to-be-identified video; extracting a second quantity of second video frames from the any candidate video; determining, for any first video frame and any second video frame, a visual feature similarity score between the any first video frame and the any second video frame; if the visual feature similarity score is greater than or equal to a first preset threshold, determining the any first video frame and the any second video frame as a video frame similar pair; and determining the first similarity score of the to-be-identified video and the any candidate video in the visual feature dimension based on the number of video frame similar pairs determined based on the first quantity of first video frames and the second quantity of second video frames and the visual feature similarity scores corresponding to each video frame similar pair.

[0073] In the embodiments of the present specification, extracting a first quantity of first video frames from the to-be-identified video can be understood as completely extracting all video frames included in the to-be-identified video, that is, the first quantity is equal to the total number of video frames of the to-be-identified video. In this way, the first similarity score can be calculated based on all video frames to improve the accuracy of the determined first similarity score. Alternatively, it can also be understood as extracting part of the total video frames, that is, the first quantity is less than the total number of video frames. In this way, the system computing resources can be saved, and the efficiency of video identification can be improved. In the process of extracting part of the total video frames, part of the total video frames can be extracted according to a preset time interval in a uniform extraction manner. For example, one video frame is extracted every 1 second. The video frames obtained by the uniform video frame interval extraction manner can cover multiple video scenes included in the to-be-identified video in terms of time proportion, so that the visual similarity score calculated subsequently is more consistent with the true situation of the to-be-identified video. It can be understood that the first video frames can also be extracted from the total video frames in a random extraction manner.

[0074] In the embodiments of the present specification, the extraction of the second quantity of second video frames from any candidate video can refer to the above-mentioned extraction of the first quantity of first video frames from the video to be recognized, and will not be repeated here. It should be noted that the first quantity and the second quantity can be the same or different, and no limitation is made thereto.

[0075] In the embodiments of the present specification, determining the visual feature similarity score between the any first video frame and the any second video frame can include extracting a first visual feature of the any first video frame, extracting a second visual feature of the any second video frame, and calculating a visual feature similarity score between the first visual feature and the second visual feature. The visual feature can be a frame-level visual feature, which can be a visual feature extracted from a frame of video.

[0076] In actual applications, the first visual feature can be extracted from the any first video frame by using an existing visual feature extraction model, and the second visual feature can be extracted from the any second video frame by using the visual feature extraction model. The visual feature extraction model can refer to a technical model capable of processing video frame pictures and extracting visual information. The role of the visual frame model is to convert images in videos into feature data that can be understood and calculated by computers.

[0077] In actual applications, calculating the visual feature similarity score between the first visual feature and the second visual feature can include calculating any one of a cosine similarity, an Euclidean distance, and a Hamming distance between the first visual feature and the second visual feature, to obtain the visual feature similarity score.

[0078] In the embodiments of the present specification, a video frame similarity pair can refer to a similarity pair formed according to two similar video frames. Each first video frame and each second video frame can only be used to form one video frame similarity pair.

[0079] In the embodiments of the present specification, the first quantity of first video frames and the second quantity of second video frames are compared with each other according to the method of determining whether the any first video frame and the any second video frame can form a video frame similarity pair, and the total number of the formed video frame similarity pairs is counted.

[0080] For example, it is assumed that the first quantity of first video frames is five first video frames, namely video frame 1, video frame 2, video frame 3, video frame 4, and video frame 5. The second quantity of second video frames is six video frames, namely video frame 6, video frame 7, video frame 8, video frame 9, video frame 10, and video frame 11. The method for determining the video frame similar pair can include selecting a video frame, such as video frame 2, from the first video frames in a certain order or randomly, calculating the visual feature similarity score between video frame 2 and each second video frame, and determining video frame 2 and video frame 10 as a video frame similar pair if the visual feature similarity score between video frame 2 and video frame 10 meets the requirements of the video frame similar pair. Then, another first video frame, such as video frame 5, is selected from the remaining first video frames other than video frame 2, the visual feature similarity score between video frame 5 and each second video frame other than video frame 10 is calculated, and it is determined whether video frame 5 can form a video frame similar pair with other second video frames based on the calculated visual feature similarity scores. If video frame 2 does not meet the requirements of constructing a video frame similar pair with each second video frame, the visual feature similarity score between video frame 5 and each second video frame is calculated to determine whether video frame 5 can form a video frame similar pair with each second video frame. The first quantity of first video frames and the second quantity of second video frames are compared two by two in a manner that the determined video frames in the video frame similar pair are not put back, and the total number of the constructed video frame similar pairs is counted.

[0081] In the embodiments of the present specification, the third quantity of the determined video frame similar pairs for the first quantity of first video frames and the second quantity of second video frames is counted, and each visual feature similarity score corresponding to each video frame similar pair is counted. The first ratio between the third quantity and the first quantity is calculated, that is, the proportion of the first video frame in each video frame similar pair in all first video frames. The second ratio between the third quantity and the second quantity is calculated, that is, the proportion of the second video frame in each video frame similar pair in all second video frames. The visual feature average similarity score of each visual feature similarity score is calculated, and the calculated first ratio, second ratio, and visual feature average similarity score are spliced to obtain the first similarity score of the to-be-identified video and any candidate video in the visual feature dimension.

[0082] In the embodiments of the present specification, the first similarity score in the visual feature dimension is obtained by splicing the first ratio of the first video frame in all first video frames, the second ratio of the second video frame in all second video frames, and the average of the visual feature similarity scores corresponding to each video frame similarity pair. This determination method can comprehensively analyze the visual similarity between the to-be-identified video and the candidate video from three aspects of the first video frame coverage ratio, the second video frame coverage ratio, and the average level of the similarity score, thereby improving the accuracy of judging the originality of the to-be-identified video according to the first similarity score.

[0083] In one or more embodiments of the present specification, a specific embodiment of determining the second similarity score in the text feature dimension is also proposed.

[0084] Optionally, the determining the second similarity score in the text feature dimension between the to-be-identified video and any candidate video can specifically include: obtaining first text from first preset information of the to-be-identified video; obtaining second text from second preset information of any candidate video, the first preset information including at least one of title information, image information, and audio information in the video; the second preset information belongs to the same type of data as the first preset information; and determining the second similarity score in the text feature dimension between the to-be-identified video and any candidate video based on first text features extracted from the first text and second text features extracted from the second text.

[0085] In the embodiments of the present specification, the text information type can be understood as the acquisition channel type of the text information, and specifically can include three channel types. The first acquisition channel can be to directly extract text information from the title of the to-be-identified video to obtain title text information. The second acquisition channel can be to extract text information by recognizing images in the to-be-identified video to obtain image text information, such as extracting text information from subtitles or text identifiers in the images. The third acquisition channel can be to extract text information by recognizing audio in the to-be-identified video to obtain audio text information.

[0086] In actual applications, the optical character recognition (OCR) technology can be used to extract text information from the images of the to-be-identified video, and the automatic speech recognition (ASR) technology can be used to extract text information from the audio of the to-be-identified video.

[0087] In the embodiments of this specification, the correspondence between the text information type of the second text information and the text information type of the first text information means that when the first text information includes the text information of the target acquisition channel, the second text information also includes the text information of the target acquisition channel, wherein the target acquisition channel can be any one acquisition channel, any two acquisition channels, or three acquisition channels.

[0088] In practical applications, when the first text information includes text information from any two or three acquisition channels, the individual text information can be concatenated to obtain the first text information. For example, if the first text information includes title text information and image text information, and the title text information is "Travel Guide" and the image text information is "Day 1 Scenic Spot 1, Day 2 Scenic Spot 2, Day 3 Scenic Spot 3, etc.", then the first text information can be obtained by directly concatenating the image text information after the title text information, i.e., "Travel Guide, Day 1 Scenic Spot 1, Day 2 Scenic Spot 2, Day 3 Scenic Spot 3, etc." The explanation of the second text information can refer to the explanation of the first text information, and will not be repeated here.

[0089] In the embodiments of this specification, a first text feature can be extracted from the first text information using an existing text feature extraction model, and a second text feature can be extracted from the second text information using the same text feature extraction model. A text feature extraction model can refer to an algorithmic model used to process text information and extract text features. A text model can transform unstructured text data into a computer-quantifiable and comparable feature vector. The calculated text feature similarity between the first and second text features can be used to determine the second similarity score between the video to be identified and any candidate video in the text feature dimension.

[0090] In practical applications, methods for calculating the text feature similarity between a first text feature and a second text feature can include calculating the semantic similarity between the first and second text features, and calculating the edit distance between the first and second text features. Edit distance is a metric that measures the difference between two strings or text sequences; it can be understood as the minimum number of editing operations required to transform one string into another. The smaller the edit distance, the more similar the first and second text features are.

[0091] In practical applications, the semantic similarity and edit distance calculated for the first and second text features are concatenated to obtain the second similarity score in the text feature dimension.

[0092] In the embodiments of this specification, since there are multiple possible combinations of text types for the first text information, such as title text information, title text information plus image text information, title text information plus image text information plus audio text information, etc., the corresponding second text information can also match these text type combinations. In practical applications, for each type combination, a set of second similarity scores in the text feature dimension is determined, and the second similarity scores in the text feature dimension of each set are concatenated to obtain the second similarity score between the video to be identified and any candidate video in the text feature dimension.

[0093] For example, in scenarios where both the first and second text information include title text information, a first semantic similarity and a first edit distance are calculated; in scenarios where both the first and second text information include image text information, a second semantic similarity and a second edit distance are calculated; in scenarios where both the first and second text information include audio text information, a third semantic similarity and a third edit distance are calculated; in scenarios where both the first and second text information include title text information and image text information, a fourth semantic similarity and a fourth edit distance are calculated; in scenarios where both the first and second text information include title text information and audio text information, a fifth semantic similarity and a fifth edit distance are calculated; in scenarios where both the first and second text information include image text information and audio text information, a sixth semantic similarity and a sixth edit distance are calculated; and in scenarios where both the first and second text information include title text information, image text information, and audio text information, a seventh semantic similarity and a seventh edit distance are calculated. The first semantic similarity, first edit distance, second semantic similarity, second edit distance, third semantic similarity, third edit distance, fourth semantic similarity, fourth edit distance, fifth semantic similarity, fifth edit distance, sixth semantic similarity, sixth edit distance, seventh semantic similarity, and seventh edit distance are concatenated to obtain the second similarity score between the video to be identified and any candidate video in the text feature dimension. By concatenating various semantic similarities and edit distances under multiple text type combinations, the second similarity score in the text feature dimension can cover multi-dimensional text information, avoiding the one-sidedness of a single perspective. This allows for a more realistic reflection of the text similarity between the video to be identified and the candidate videos, improving the accuracy of originality identification for the video to be identified.

[0094] In one or more embodiments described herein, a specific embodiment for determining a third similarity score on a multimodal feature dimension is also proposed.

[0095] Optionally, determining the third similarity score between the video to be identified and any candidate video in the multimodal feature dimension may specifically include: obtaining first multimodal information from the video to be identified, the first multimodal information including video frame information extracted from the video to be identified and text information extracted from the video to be identified; obtaining second multimodal information from any candidate video, the second multimodal information including video frame information extracted from the candidate video and text information extracted from the candidate video; and determining the third similarity score between the video to be identified and any candidate video in the multimodal feature dimension based on the first multimodal features extracted from the first multimodal information and the second multimodal features extracted from the second multimodal information.

[0096] In the embodiments of this specification, the first multimodal information may include video frame information and text information extracted from the video to be identified. In practical applications, the video frame information may be all video frames extracted from the video to be identified, or it may be a portion of the video frames extracted from the video to be identified. The text information may include at least one of the following: title text information extracted from the video to be identified, image text information identified based on images in the video to be identified, and audio text information identified based on audio in the video to be identified. The explanation of the second multimodal information can be referred to the explanation of the first multimodal information, and will not be repeated here.

[0097] In the embodiments of this specification, a first multimodal feature can be extracted from the first multimodal information using an existing multimodal feature extraction model, and a second multimodal feature can be extracted from the second multimodal information using the same multimodal feature extraction model. A multimodal feature extraction model is a model capable of simultaneously processing, understanding, and fusing multiple different modalities of information. Multimodal feature extraction models can overcome the limitations of single-modality models and, by mining the correlation information between different modalities, can extract more comprehensive features that are closer to the overall semantics of the video, thereby improving the accuracy of video recognition.

[0098] In practical applications, the third similarity score between the video to be identified and any candidate video on the multimodal feature dimension can be obtained by calculating any one of the cosine similarity, Euclidean distance, and Manhattan distance between the first multimodal feature and the second multimodal feature.

[0099] In practical applications, multimodal information can be constructed in various combinations based on the number of video frames or the type of text information included. For each combination, a corresponding multimodal dimension similarity score, i.e., a third similarity score, is calculated. These third similarity scores are then concatenated to obtain the final similarity score in the multimodal feature dimension. For example, a first multimodal information 'a' can be constructed based on some video frame information and title text information in the video to be identified, and a second multimodal information 'a' can be constructed based on some video frame information and title text information in the candidate video. A third similarity score in the multimodal feature dimension is determined based on the first multimodal information 'a' and the second multimodal information 'a'. Alternatively, a first multimodal information 'b' can be constructed based on all video frame information in the video to be identified and text information composed of title text information and image text information, and a second multimodal information 'b' can be constructed based on all video frame information in the candidate video and text information composed of title text information and image text information. Another third similarity score in the multimodal feature dimension is determined based on the first multimodal information 'b' and the second multimodal information 'b'. Furthermore, multimodal information can be constructed based on other combinations to calculate more third similarity scores, which will not be listed here. By concatenating these third similarity scores under different combinations, the comparison results of multi-dimensional multimodal features can be integrated, so that the final third similarity score can more comprehensively reflect the overall similarity between the video to be identified and the candidate video at the multimodal level.

[0100] When the video to be identified contains a face image, the comprehensive similarity score can also include a fourth similarity score based on the face feature dimension.

[0101] Optionally, the video to be identified includes a face image. The step of concatenating the first similarity score, the second similarity score, and the third similarity score to obtain the comprehensive similarity score between the video to be identified and any candidate video may specifically include: determining a fourth similarity score between the video to be identified and any candidate video in the face feature dimension; and concatenating the first similarity score, the second similarity score, the third similarity score, and the fourth similarity score to obtain the comprehensive similarity score between the video to be identified and any candidate video.

[0102] In the embodiments of this specification, the facial images in the video to be identified may include one or more facial images. Since the candidate videos in the candidate video set are videos that meet certain similarity requirements with the video to be identified, the candidate videos may also include facial images. The number of facial images in the candidate videos may be the same as or different from the number of facial images in the video to be identified; this is not limited. It is understood that the candidate videos may also not contain facial images.

[0103] In one or more embodiments described herein, a specific embodiment for determining a fourth similarity score on the facial feature dimension is also proposed.

[0104] Optionally, determining the fourth similarity score between the video to be identified and any candidate video in the facial feature dimension may specifically include: extracting a first facial image from the video to be identified; extracting a second facial image from any candidate video; and determining the fourth similarity score between the video to be identified and any candidate video in the facial feature dimension based on the first facial features extracted from the first facial image and the second facial features extracted from the second facial image.

[0105] In the embodiments of this specification, the method for extracting a first face image from a video to be identified may include: firstly, selecting representative video frames from the video to be identified; specifically, extracting several independent static frames from the video using keyframe sampling as the basic material for face detection. Keyframes can refer to frames with significant changes in content, such as frames containing the appearance or turning of a person. Secondly, for each sampled static frame, a specialized face detection algorithm scans the image to identify and locate the face region. The face detection algorithm analyzes the pixel features of the image to distinguish between faces and background. If a face is present in the image, it outputs the bounding box coordinates of the face, clarifying the range of the face in the image and eliminating interference from non-face regions. The bounding box coordinates can indicate the specific location of the face in the image, such as the coordinates of the upper left and lower right corners. Finally, based on the located bounding box coordinates, the image region containing the face is cropped from the video frame to obtain the extracted face image. In practical applications, the method for extracting a second face image can refer to the method for extracting the first face image, and will not be elaborated further here.

[0106] In the embodiments of this specification, a first facial feature can be extracted from a first facial image using an existing facial feature extraction model, and a second facial feature can be extracted from a second facial image using the same model. A facial feature extraction model can refer to a model used to extract unique and quantifiable facial features from a facial image. This model can transform a visually visible facial image into a computer-computable and comparable feature vector, providing a data foundation for subsequent similarity calculations of facial feature dimensions.

[0107] In practical applications, the fourth similarity score between the video to be identified and any candidate video on the facial feature dimension can be obtained by calculating any one of the cosine similarity, Euclidean distance, and Manhattan distance between the first facial feature and the second facial feature.

[0108] In practical applications, in scenarios where the video to be identified includes multiple first face images and any candidate video includes multiple second face images, each first face image is compared with each second face image. After comparing one first face image with one second face image, a face feature similarity score can be generated, thus obtaining several face feature similarity scores. These several face feature similarity scores are then concatenated to obtain the fourth similarity score between the video to be identified and any candidate video in the face feature dimension.

[0109] For example, suppose there are X first face images and Y second face images. Each of the X first face images is compared with each of the Y second face images. Each comparison generates a face feature similarity score. The total number of comparisons is XY, thereby generating XY face feature similarity scores. The XY face feature similarity scores are then concatenated to obtain a fourth similarity score in the face feature dimension.

[0110] In this embodiment, the comprehensive similarity score between the video to be identified and any of the candidate videos is generated by concatenating a first similarity score in the visual feature dimension, a second similarity score in the text feature dimension, a third similarity score in the multimodal feature dimension, and a fourth similarity score in the face feature dimension. The first similarity score includes the first ratio of the first video frame in a video frame similarity pair to all first video frames, the second ratio of the second video frame in a video frame similarity pair to all second video frames, and the average score of the visual feature similarity scores corresponding to each video frame similarity pair. The second similarity score covers text comparison results in seven scenarios, specifically including text semantic similarity and character-level edit distance in each scenario. The third similarity score consists of multiple similarity scores calculated from various combinations of multimodal information. The fourth similarity score can include multiple similarity scores generated from multiple sets of face images. By combining the similarity scores from the four dimensions mentioned above, the comprehensive similarity score can fully cover comparison information from multiple levels, including visual, text, multimodal, and facial data. It includes both global dimensions, such as multimodal dimensions, and local dimensions, such as visual dimensions, presenting the similarity features between the video to be identified and the candidate video from multiple dimensions. This can effectively improve the accuracy of originality identification for the video to be identified.

[0111] Each candidate video in the candidate video set can generate a corresponding comprehensive similarity score with the video to be identified. Therefore, multiple comprehensive similarity scores can be generated for the entire candidate video set. These comprehensive similarity scores can be sorted based on the similarity score along the target feature dimension to obtain a comprehensive similarity score sequence. A portion of these comprehensive similarity scores can then be selected from this sequence and input into the video recognition model.

[0112] Optionally, the step of inputting at least a portion of the multiple comprehensive similarity scores into a video recognition model to obtain the original video probability value of the video to be recognized output by the video recognition model based on the at least a portion of the comprehensive similarity scores may specifically include: obtaining the weight values ​​corresponding to each feature dimension; sorting the comprehensive similarity scores according to the similarity scores corresponding to the feature dimensions of the target weight values ​​to obtain a comprehensive similarity score sequence; and inputting the top preset number of comprehensive similarity scores in the comprehensive similarity score sequence into the video recognition model to obtain the original video probability value of the video to be recognized output by the video recognition model based on the top preset number of comprehensive similarity scores.

[0113] In the embodiments of this specification, the range of each feature dimension is consistent with the feature dimensions used to generate the comprehensive similarity score. For example, if the comprehensive similarity score is generated by concatenating similarity scores generated under the visual feature dimension, text feature dimension, and multimodal feature dimension, then each feature dimension may include the visual feature dimension, text feature dimension, and multimodal feature dimension.

[0114] In the embodiments of this specification, the weight value corresponding to the feature dimension can be a quantitative value determined according to the importance of the feature dimension in video originality recognition, wherein the importance can be determined based on experience or data accumulated in the historical video recognition process.

[0115] In the embodiments of this specification, the feature dimension of the target weight value can refer to the largest feature dimension of the weight value. The corresponding similarity score can refer to the similarity score between the video to be identified and the candidate video on that feature dimension. Based on the similarity score of the feature dimension of the target weight value, the comprehensive similarity scores are sorted. For example, if the feature dimension of the target weight value is a visual feature dimension, the comprehensive similarity scores are sorted in descending order of the similarity score of the visual feature dimension.

[0116] In the embodiments of this specification, the first preset number of comprehensive similarity scores can be the first two comprehensive similarity scores, the first three comprehensive similarity scores, or other comprehensive similarity scores. It is understood that, provided the user's recognition accuracy requirements are met, the first preset number of comprehensive similarity scores can also be the first comprehensive similarity score, and this is not limited.

[0117] In the embodiments of this specification, the top preset number of comprehensive similarity scores from the comprehensive similarity score sequence are input into the video recognition model. This allows the video recognition model to focus on the comparison data of the candidate videos most relevant to the video to be identified, and also avoids introducing a large number of low-similarity candidate videos, thereby reducing noise interference. Furthermore, the comprehensive similarity scores input to the model correspond to the comparison results between the video to be identified and multiple candidate videos. This allows the model to combine the features of multiple candidate videos for comprehensive analysis, thereby improving the accuracy of identifying the originality of the video to be identified.

[0118] If the video to be identified is generated by editing multiple published videos—for example, by extracting a segment from published video a, a segment from published video b, and a segment from published video c—and then generating the video to be identified based on these segments, then sorting the comprehensive similarity scores according to a certain feature dimension and selecting the first preset number of comprehensive similarity scores from the sequence may have the following drawbacks.

[0119] Because the pre-set number of comprehensive similarity scores all have high similarity scores in this feature dimension, video recognition models are prone to over-focusing on features in this dimension. For example, the model might primarily focus on the high similarity with video 'a' in the visual feature dimension. However, the video to be identified is a fusion of multiple different videos. The similarity between the video to be identified and videos 'b' and 'c' may be scattered across other feature dimensions, and the high similarity scores corresponding to these dimensions may not have been included in the pre-set number of comprehensive similarity scores. In this case, the video recognition model may focus on high similarity information in a single dimension, ignoring the similarity between the video to be identified and other published videos 'b' or 'c'. This would prevent the video recognition model from fully capturing the multi-source editing nature of the video to be identified, thus affecting the accuracy of originality identification for the video.

[0120] Optionally, inputting at least a portion of the multiple comprehensive similarity scores into the video recognition model may include: sorting the comprehensive similarity scores based on the similarity scores corresponding to the visual feature dimension to obtain a first sequence of comprehensive similarity scores; sorting the comprehensive similarity scores based on the similarity scores corresponding to the text feature dimension to obtain a second sequence of comprehensive similarity scores; sorting the comprehensive similarity scores based on the similarity scores corresponding to the multimodal feature dimension to obtain a third sequence of comprehensive similarity scores; sorting the comprehensive similarity scores based on the similarity scores corresponding to the facial feature dimension to obtain a fourth sequence of comprehensive similarity scores; and inputting the first comprehensive similarity score in the first sequence, the first comprehensive similarity score in the second sequence, the first comprehensive similarity score in the third sequence, and the first comprehensive similarity score in the fourth sequence into the video recognition model.

[0121] In this embodiment, based on the similarity scores corresponding to each feature dimension, all comprehensive similarity scores are sorted to obtain multiple comprehensive similarity score sequences. The first comprehensive similarity score in each sequence is then input into the video recognition model, enabling the model to accurately address the originality recognition requirements of multi-source edited videos. When the video to be recognized is generated by fusing multiple published video clips, the similarity between the video to be recognized and each video source is often dispersed across different feature dimensions. By selecting the first comprehensive similarity score from the comprehensive similarity score sequence corresponding to each feature dimension, the video recognition model can capture the similarity between the video to be recognized and each video source. This allows the model to integrate this multi-source similarity information and accurately determine the non-originality of the spliced ​​edits in the video to be recognized, thereby improving the accuracy of originality recognition for the video to be recognized.

[0122] By filtering candidate videos using multi-dimensional features to construct a candidate video set, the candidate video set can cover videos that are related to the video to be identified in more feature dimensions, thereby improving the practicality of the candidate video set.

[0123] Optionally, obtaining a candidate video set for identifying the originality of the video to be identified may specifically include: determining a first candidate video set for identifying the originality of the video to be identified based on the visual feature dimension; the similarity score between each first candidate video in the first candidate video set and the video to be identified on the visual feature dimension is greater than or equal to a first preset threshold; determining a second candidate video set for identifying the originality of the video to be identified based on the text feature dimension; the similarity score between each second candidate video in the second candidate video set and the video to be identified on the text feature dimension is greater than or equal to a second preset threshold; determining a third candidate video set for identifying the originality of the video to be identified based on the multimodal feature dimension; the similarity score between each third candidate video in the third candidate video set and the video to be identified on the multimodal feature dimension is greater than or equal to a third preset threshold; merging the first candidate video set, the second candidate video set, and at least two candidate video sets in the third candidate video set to obtain the candidate video set.

[0124] In the embodiments of this specification, the explanations of visual feature dimension, text feature dimension, and modal feature dimension can refer to the foregoing explanations, and will not be repeated here. The first preset threshold, the second preset threshold, and the third preset threshold can be all the same threshold, some of the same threshold, or all different thresholds; there is no limitation in this regard.

[0125] In the embodiments of this specification, the first candidate video set, the second candidate video set, and the third candidate video set may contain duplicate candidate videos. Therefore, before merging at least two candidate video sets, it is necessary to perform deduplication processing on each candidate video set. Based on the deduplicated candidate videos, the candidate video sets are aggregated to obtain the candidate video sets.

[0126] In one or more embodiments of this specification, a specific embodiment of a method for determining a first candidate video set is proposed.

[0127] Optionally, determining a first candidate video set for identifying the originality of the video to be identified based on the visual feature dimensions may include: extracting multiple video frames from the video to be identified; extracting visual features from each video frame to obtain each visual feature; for any visual feature, filtering multiple candidate visual features similar to the visual feature from a visual feature library; determining the first video to which each candidate video frame corresponding to each candidate visual feature belongs; and aggregating the first videos determined based on each visual feature to obtain the first candidate video set.

[0128] In the embodiments of this specification, all or part of the video frames can be extracted from the video to be identified. For example, if L video frames are extracted from the video to be identified, visual features are extracted for each video frame, resulting in L visual features. For each visual feature, multiple candidate visual features with relatively high similarity to the visual feature can be selected from the visual feature library. Specifically, the visual feature similarity score between the visual feature and each candidate visual feature in the visual feature library is determined. From each visual feature similarity score, multiple candidate visual features with the highest similarity scores are selected sequentially. For example, if M visual features are selected, a total of LM candidate visual features can be selected from the visual feature library. Based on the visual feature similarity scores corresponding to the candidate visual features, visual features with similarity scores less than a preset threshold are deleted from the LM candidate visual features. For example, after deletion, N candidate visual features are obtained. The corresponding candidate video frames can be recalled using the N candidate visual features. The corresponding original candidate videos can be queried based on each recalled video frame. The queried original candidate videos are aggregated to obtain the first candidate video set.

[0129] In practical applications, a visual feature library can refer to a database that stores visual features extracted from video frames of other published videos. Based on the visual features extracted from the video to be identified, when selecting candidate visual features from the visual feature library, different numbers of candidate visual features can be selected for different visual features. For example, M1 candidate visual features can be selected for visual feature L1, and M2 candidate visual features can be selected for visual feature L2. It is understandable that the same number of candidate visual features can also be selected for different visual features.

[0130] In one or more embodiments of this specification, a specific embodiment of a method for determining a second candidate video set is proposed.

[0131] Optionally, determining a second candidate video set for identifying the originality of the video to be identified based on the text feature dimension may include: extracting title text information from the video to be identified to obtain a first text; using an OCR model to recognize image information in the video to be identified to obtain a second text; using an ASR model to recognize audio information in the video to be identified to obtain a third text; extracting target text features from the target text, wherein the target text includes at least one of the first text, the second text, and the third text; filtering multiple candidate text features similar to the target text features from a text feature library; determining the second video to which each candidate text corresponding to each candidate text belongs; and aggregating the second videos to obtain a second candidate video set.

[0132] In the embodiments of this specification, the explanation of the target text can refer to the explanation of the first text information described above, and will not be repeated here. For the target text extracted from the video to be identified, target text features are extracted. For the target text features, multiple candidate text features with relatively high similarity to the target text features can be selected from the text feature library. Specifically, the text feature similarity score between the target text feature and each candidate text feature in the text feature library is determined. From each text feature similarity score, multiple candidate text features with the highest similarity scores are selected sequentially. Multiple candidate texts can be recalled through these multiple candidate text features. Based on each recalled candidate text, the corresponding original candidate video can be queried. The queried original candidate videos are aggregated to obtain the second candidate video set.

[0133] In practical applications, a text feature library can refer to a library that stores text features extracted from the text of other published videos. The target text extracted from the video to be identified can yield various text types depending on the content extracted. For example, extracting the first text mentioned above yields target text of type first; extracting both the first and second texts yields target text of type second; and extracting the first, second, and third texts yields target text of type third. Other types are not listed here. For each type of target text, a second candidate video set can be constructed. Merging these second candidate video sets yields a merged second candidate video set.

[0134] In one or more embodiments of this specification, a specific embodiment of a method for determining a third candidate video set is proposed.

[0135] Optionally, determining a third candidate video set for identifying the originality of the video to be identified based on the multimodal feature dimensions may include: extracting multiple video frames from the video to be identified; extracting title text information from the video to be identified to obtain first text; using an OCR model to recognize image information in the video to be identified to obtain second text; using an ASR model to recognize audio sound in the video to be identified to obtain third text; extracting target multimodal features from multimodal information, wherein the multimodal information includes at least some video frames from the multiple video frames and target text, wherein the target text includes at least one of the first text, the second text, and the third text; filtering multiple candidate multimodal features similar to the target multimodal features from a multimodal feature library; determining the third video to which each candidate multimodal information corresponding to each candidate multimodal feature belongs; and aggregating the third videos to obtain a third candidate video set.

[0136] In the embodiments of this specification, the explanation of constructing multimodal information can be referred to the aforementioned explanation of the first multimodal information, and will not be repeated here. For the multimodal information extracted from the video to be identified, target multimodal features are extracted. For the target multimodal features, multiple candidate multimodal features with relatively high similarity to the target multimodal feature can be selected from the multimodal feature library. Specifically, the multimodal feature similarity score between the target multimodal feature and each candidate multimodal feature in the multimodal feature library is determined. From each multimodal feature similarity score, the multiple candidate multimodal features with the highest similarity scores are selected sequentially. Multiple candidate multimodal information can be recalled through these multiple candidate multimodal features. Based on each recalled candidate multimodal information, the corresponding original candidate video can be queried. The retrieved original candidate videos are aggregated to obtain the third candidate video set.

[0137] In practical applications, a multimodal feature library can refer to a library that stores multimodal features extracted from multimodal information of other published videos. The multimodal information extracted from the video to be identified can be of various types depending on the content extracted. For example, extracting a first number of video frames and the aforementioned first text yields the first type of multimodal information; extracting a second number of video frames and simultaneously extracting the aforementioned first and second texts yields the second type of multimodal information; extracting all video frames and simultaneously extracting the aforementioned first, second, and third texts yields the third type of multimodal information. Other types are not listed here. For each type of multimodal information, a third candidate video set can be constructed. Merging these third candidate video sets yields the merged third candidate video set.

[0138] In scenarios where the video to be identified includes facial images, a candidate video set for identifying the originality of the video to be identified can also be constructed based on the facial images.

[0139] Optionally, the video to be identified includes a face image. The step of merging at least two video sets from the first candidate video set, the second candidate video set, and the third candidate video set to obtain the candidate video set may specifically include: determining a fourth candidate video set for identifying the originality of the video to be identified based on the face feature dimension; ensuring that each fourth candidate video in the fourth candidate video set has a similarity score greater than or equal to a fourth preset threshold with the video to be identified on the face feature dimension; and merging the first candidate video set, the second candidate video set, the third candidate video set, and at least two video sets from the fourth candidate video set to obtain the candidate video set.

[0140] In the embodiments of this specification, the explanation of the facial feature dimensions can be referred to the foregoing explanation, and will not be repeated here. The first preset threshold, the second preset threshold, the third preset threshold, and the fourth preset threshold can all be the same threshold, some can be the same threshold, or they can all be different thresholds, and there is no limitation in this regard.

[0141] In the embodiments of this specification, the first candidate video set, the second candidate video set, the third candidate video set, and the fourth candidate video set may have duplicate candidate videos. Therefore, before merging at least two candidate video sets, it is necessary to perform deduplication processing on each candidate video set. Based on the deduplicated candidate videos, the candidate video sets are aggregated to obtain the candidate video sets.

[0142] In the embodiments of this specification, a candidate video set is constructed from multiple feature dimensions, including visual, textual, multimodal, and facial features, so that the candidate video set can cover the comparison information related to the video to be identified under each feature dimension. This allows for a comprehensive analysis from multiple feature dimensions, such as visual image similarity, text content relevance, and consistency of on-screen faces, when performing originality identification on the video to be identified based on this candidate video set during the subsequent video recognition stage. This effectively improves the accuracy of originality identification for the video to be identified.

[0143] In one or more embodiments of this specification, a specific embodiment of a method for determining a fourth candidate video set is proposed.

[0144] Optionally, determining a fourth candidate video set for identifying the originality of the video to be identified based on the facial feature dimension may include: obtaining a target facial image from the video to be identified; extracting target facial features from the target facial image; filtering multiple candidate facial features similar to the target facial features from a facial feature library; determining the fourth video to which each candidate facial image corresponding to each candidate facial feature belongs; and aggregating the fourth videos to obtain the fourth candidate video set.

[0145] In the embodiments of this specification, for a target facial feature, multiple candidate facial features with relatively high similarity to the target facial feature can be selected from the facial feature database. Specifically, the facial feature similarity score between the target facial feature and each candidate facial feature in the facial feature database is determined, and multiple candidate facial features with the highest similarity scores are selected sequentially from each facial feature similarity score. Multiple candidate facial features can be used to recall corresponding candidate facial information. Based on each recalled candidate facial information, the corresponding original candidate video can be queried. The retrieved original candidate videos are aggregated to obtain a fourth candidate video set.

[0146] In practical applications, a facial feature library can refer to a library that stores facial features extracted from facial information in other published videos. If the video to be identified contains multiple facial images, a fourth candidate video set can be constructed for each facial image. The various fourth candidate video sets are then merged to obtain the merged fourth candidate video set.

[0147] To facilitate understanding of this solution by those skilled in the art, a specific embodiment of the method for originality identification of promotional videos provided by e-commerce platforms for promoting products is also proposed in one or more embodiments of this specification.

[0148] Figure 3 This is a flowchart illustrating an embodiment of the originality identification process for product demonstration videos provided by e-commerce platforms. It includes steps 302 to 342.

[0149] Step 302: Obtain video frames, text information, product information, and facial information from the product demonstration videos provided by the e-commerce platform.

[0150] In the embodiments of this specification, the explanations of video frames, text information, and facial information can refer to the above content, and will not be repeated here. Product information may include product shape, product size, and product appearance, etc., and product appearance may refer to pattern information on the product, etc.

[0151] Step 304: Extract visual features from the video frames using a visual feature extraction model.

[0152] Step 306: Extract text features from the text information using a text feature extraction model.

[0153] Step 308: Use a multimodal feature extraction model to extract multimodal features from the multimodal information composed of at least two types of information, including the video frame, the text information, the product information, and the face information.

[0154] Step 310: Extract facial features from the facial information using a facial feature extraction model.

[0155] Step 312: Extract the item features from the product information using the item feature extraction model.

[0156] In the embodiments of this specification, the item feature extraction model can be a model used to automatically mine and extract the essential attributes or key features of an item from item-related data. The item feature extraction model is used to transform the original information of an item into feature data that can be understood and calculated by a computer, so as to provide a basis for subsequent tasks such as item identification, classification, and comparison.

[0157] Step 314: Recall candidate visual features from the visual feature library based on visual features.

[0158] Step 316: Retrieve candidate text features from the text feature library based on text features.

[0159] Step 318: Recall candidate multimodal features from the multimodal feature library based on multimodal features.

[0160] Step 320: Retrieve candidate facial features from the facial feature database based on facial features.

[0161] Step 322: Retrieve candidate item features from the item feature library based on item features.

[0162] In the embodiments of this specification, the item feature library may refer to a library that stores item features extracted from item information of other published videos.

[0163] Step 324: Query the corresponding video using the visual information corresponding to the candidate visual features to obtain the first candidate video set.

[0164] Step 326: Query the corresponding video using the text information corresponding to the candidate text features to obtain the second candidate video set.

[0165] Step 328: Query the corresponding videos using the multimodal information corresponding to the candidate multimodal features to obtain the third candidate video set.

[0166] Step 330: Query the corresponding video by using the facial information corresponding to the candidate's facial features to obtain the fourth candidate video set.

[0167] Step 332: Query the corresponding video by using the item information corresponding to the candidate item features to obtain the fifth candidate video set.

[0168] Step 334: Aggregate the first candidate video set, the second candidate video set, the third candidate video set, the fourth candidate video set, and the fifth candidate video set to obtain a candidate video set.

[0169] Step 336: For any candidate video in the candidate video set, calculate the visual feature similarity score, text feature similarity score, multimodal feature similarity score, facial feature similarity score, and object feature similarity score between the candidate video and the video to be promoted and explained.

[0170] Step 338: Concatenate the visual feature similarity score, the text feature similarity score, the multimodal feature similarity score, the face feature similarity score, and the object feature similarity score to obtain the comprehensive similarity score of any candidate video.

[0171] Step 340: Determine the comprehensive similarity score between each candidate video in the candidate video set and the promotional video using the method for determining the comprehensive similarity score of any candidate video, and obtain multiple comprehensive similarity scores.

[0172] Step 342: Input at least a portion of the comprehensive similarity scores from the plurality of comprehensive similarity scores into the video recognition model to obtain the recognition result output by the video recognition model based on the at least a portion of the comprehensive similarity scores.

[0173] Understandable Figure 3 When constructing the candidate video set, the fifth candidate video set can be constructed without being based on object features, and the object feature similarity score can be omitted when generating the comprehensive similarity score; there are no restrictions on this. If there are no face images in the video to be identified, the fourth candidate video set can be constructed without face features, and the face feature similarity score can be omitted when generating the comprehensive similarity score.

[0174] The various technical features in the above embodiments can be combined arbitrarily, as long as there is no conflict or contradiction between the combinations of features. However, due to space limitations, they have not been described one by one. Therefore, the arbitrary combination of various technical features in the above embodiments is also within the scope of this specification.

[0175] Based on the same idea, embodiments of this specification also provide apparatus corresponding to the above methods.

[0176] Figure 4 The embodiments provided in this specification correspond to Figure 2 A schematic diagram of the structure of a video recognition device.

[0177] like Figure 4 As shown, the device may include:

[0178] The first acquisition module 402 is used to acquire the video to be recognized.

[0179] The second acquisition module 404 allows the user to acquire a set of candidate videos for identifying the originality of the video to be identified.

[0180] The first determining module 406 is used to determine, for each candidate video in the candidate video set, the comprehensive similarity score between the video to be identified and each candidate video in a specific feature dimension, at least two of the target feature dimensions, to obtain multiple comprehensive similarity scores; the specific feature dimensions include at least two of the following: visual feature dimension, text feature dimension, multimodal feature dimension, and face feature dimension. The comprehensive similarity score of a candidate video is obtained by concatenating the similarity scores of the candidate video and the video to be identified in the at least two target feature dimensions.

[0181] The input module 408 is used to input at least a portion of the comprehensive similarity scores from the plurality of comprehensive similarity scores into the video recognition model to obtain the original video probability value of the video to be recognized output by the video recognition model based on the at least a portion of the comprehensive similarity scores.

[0182] The second determining module 410 is used to determine the video to be identified as an original video if the original video probability value is greater than or equal to a preset threshold.

[0183] based on Figure 4 The embodiments of this specification also provide some specific implementation schemes of the method, which are described below.

[0184] Optionally, the first determining module 406 may specifically include:

[0185] The first determining unit is configured to determine, for any candidate video among the candidate videos, a first similarity score between the video to be identified and any candidate video on the visual feature dimension.

[0186] The second determining unit is used to determine the second similarity score between the video to be identified and any candidate video on the text feature dimension.

[0187] The third determining unit is used to determine the third similarity score between the video to be identified and any candidate video on the multimodal feature dimension.

[0188] The splicing unit is used to splice the first similarity score, the second similarity score, and the third similarity score to obtain the comprehensive similarity score between the video to be identified and any candidate video.

[0189] Optionally, the first determining unit may specifically include:

[0190] The first extraction subunit is used to extract a first number of first video frames from the video to be identified.

[0191] The second extraction subunit is used to extract a second number of second video frames from any of the candidate videos.

[0192] The first determining subunit is used to determine the visual feature similarity score between any first video frame and any second video frame for any first video frame and any second video frame.

[0193] The second determining subunit is used to determine any first video frame and any second video frame as a video frame similarity pair if the visual feature similarity score is greater than or equal to a first preset threshold.

[0194] The third determining subunit is used to determine the first similarity score of the video to be identified and any candidate video in the visual feature dimension based on the number of video frame similarity pairs determined by the first number of first video frames and the second number of second video frames, and the visual feature similarity score corresponding to each video frame similarity pair.

[0195] Optionally, the second determining unit may specifically include:

[0196] The first acquisition subunit is used to acquire first text information from the video to be identified; the text information type of the first text information includes at least one of title text information in the video to be identified, image text information identified based on the image in the video to be identified, and audio text information identified based on the audio in the video to be identified.

[0197] The second acquisition subunit is used to acquire second text information from any of the candidate videos, wherein the text information type of the second text information corresponds to the text information type of the first text information.

[0198] The fourth determining subunit is used to determine the second similarity score of the video to be identified and any candidate video on the text feature dimension based on the first text feature extracted from the first text information and the second text feature extracted from the second text information.

[0199] Optionally, the third determining unit may specifically include:

[0200] The third acquisition subunit is used to acquire first multimodal information from the video to be identified. The first multimodal information includes video frame information extracted from the video to be identified and text information acquired from the video to be identified.

[0201] The fourth acquisition subunit is used to acquire second multimodal information from any of the candidate videos. The second multimodal information includes video frame information extracted from any of the candidate videos and text information acquired from any of the candidate videos.

[0202] The fifth determining subunit is used to determine the third similarity score between the video to be identified and any candidate video in the multimodal feature dimension based on the first multimodal feature extracted from the first multimodal information and the second multimodal feature extracted from the second multimodal information.

[0203] Optionally, the video to be identified may include a facial image.

[0204] Optionally, the splicing unit may specifically include:

[0205] The sixth determining subunit is used to determine the fourth similarity score between the video to be identified and any candidate video on the facial feature dimension;

[0206] The splicing subunit is used to splice the first similarity score, the second similarity score, the third similarity score, and the fourth similarity score to obtain the comprehensive similarity score between the video to be identified and any candidate video.

[0207] Optionally, the sixth determining subunit may be used to: extract a first face image from the video to be identified; extract a second face image from any candidate video; and determine a fourth similarity score between the video to be identified and any candidate video in the face feature dimension based on the first face feature extracted from the first face image and the second face feature extracted from the second face image.

[0208] Optionally, the input module 408 may specifically include:

[0209] The acquisition unit is used to obtain the weight values ​​corresponding to each feature dimension.

[0210] The sorting unit is used to sort the comprehensive similarity scores based on the similarity scores corresponding to the feature dimensions of the target weight value, so as to obtain a comprehensive similarity score sequence.

[0211] The input unit is used to input the first preset number of comprehensive similarity scores located in the comprehensive similarity score sequence into the video recognition model, and to obtain the original video probability value of the video to be recognized output by the video recognition model based on the first preset number of comprehensive similarity scores.

[0212] Optionally, the second acquisition module 404 may specifically include:

[0213] The fourth determining unit is used to determine a first candidate video set for identifying the originality of the video to be identified based on the visual feature dimension; the similarity score of each first candidate video in the first candidate video set with the video to be identified on the visual feature dimension is greater than or equal to a first preset threshold.

[0214] The fifth determining unit is used to determine a second candidate video set for identifying the originality of the video to be identified based on the text feature dimension; the similarity score of each second candidate video in the second candidate video set with the video to be identified on the text feature dimension is greater than or equal to a second preset threshold.

[0215] The sixth determining unit is used to determine a third candidate video set for identifying the originality of the video to be identified based on the multimodal feature dimension; the similarity score of each third candidate video in the third candidate video set and the video to be identified on the multimodal feature dimension is greater than or equal to a third preset threshold.

[0216] The merging unit is used to merge at least two video sets from the first candidate video set, the second candidate video set, and the third candidate video set to obtain the candidate video set.

[0217] Optionally, the merging unit may specifically include:

[0218] The seventh determining subunit is used to determine a fourth candidate video set for identifying the originality of the video to be identified based on the facial feature dimension; the similarity score of each fourth candidate video in the fourth candidate video set with the video to be identified on the facial feature dimension is greater than or equal to a fourth preset threshold.

[0219] The merging subunit is used to merge at least two video sets from the first candidate video set, the second candidate video set, the third candidate video set, and the fourth candidate video set to obtain the candidate video set.

[0220] It is understood that the modules mentioned above refer to computer programs or program segments used to perform one or more specific functions. Furthermore, the distinction between these modules does not imply that the actual program code must also be separate.

[0221] For ease of description, the above devices are described by dividing them into various modules or units based on their functions. Of course, when implementing one or more of these specifications, the functions of each module or unit can be implemented in the same or different software and / or hardware, or a module that performs the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.

[0222] The above is an illustrative scheme of a video recognition device according to this embodiment. It should be noted that the technical solution of this video recognition device and the technical solution of the video recognition method described above belong to the same concept. For details not described in detail in the technical solution of the video recognition device, please refer to the description of the technical solution of the video recognition method described above.

[0223] Based on the same idea, this specification also provides devices corresponding to the above methods in its embodiments.

[0224] Figure 5 A structural block diagram of a computing device 500 provided according to an embodiment of this specification is shown.

[0225] The computing device 500 includes:

[0226] Memory 510 and processor 520;

[0227] The memory 510 is used to store computer programs / instructions, and the processor 520 is used to execute the computer programs / instructions, which, when executed by the processor 520, implement the steps of the video recognition method.

[0228] Specifically, the components of the computing device 500 include, but are not limited to, a memory 510 and a processor 520. The processor 520 is connected to the memory 510 via a bus 530, and the database 550 is used to store data.

[0229] The computing device 500 also includes an access device 540, which enables the computing device 500 to communicate via one or more networks 560. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 540 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.

[0230] In one embodiment of this specification, the aforementioned components of the computing device 500 and Figure 5 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 5 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this application. Those skilled in the art can add or replace other components as needed.

[0231] Computing device 500 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). Computing device 500 can also be a mobile or stationary server.

[0232] The processor 520 implements the video recognition method when executing the computer instructions.

[0233] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solution of the video recognition method described above belong to the same concept. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the video recognition method described above.

[0234] An embodiment of this specification also provides a computer-readable storage medium storing computer instructions that, when executed by a processor, implement the steps of the video recognition method as described above.

[0235] The above is an illustrative scheme of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the video recognition method described above belong to the same concept. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the video recognition method described above.

[0236] An embodiment of this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the video recognition method described above.

[0237] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the video recognition method described above belong to the same concept. For details not described in detail in the technical solution of the computer program product, please refer to the description of the technical solution of the video recognition method described above.

[0238] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for the embodiments of apparatus, devices, and systems, since they are basically similar to the method embodiments, the descriptions are relatively simple, and relevant parts can be referred to the descriptions of the method embodiments. The apparatus, devices, systems, and methods provided in the embodiments of this specification correspond to each other; therefore, the apparatus, devices, and systems also have similar beneficial technical effects to the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the corresponding apparatus, devices, and systems will not be repeated here.

[0239] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0240] In the 1990s, improvements to a technology could be clearly distinguished as either hardware improvements (e.g., improvements to the circuit structure of diodes, transistors, switches, etc.) or software improvements (improvements to the methodology). However, with technological advancements, many methodological improvements today can be considered direct improvements to the hardware circuit structure. Designers almost always obtain the corresponding hardware circuit structure by programming the improved methodology into the hardware circuit. Therefore, it cannot be said that a methodological improvement cannot be implemented using hardware physical modules. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logic function is determined by the user programming the device. Designers can program a digital system themselves to "integrate" it onto a PLD, without needing chip manufacturers to design and manufacture dedicated integrated circuit chips. Furthermore, nowadays, instead of manually manufacturing integrated circuit chips, this programming is mostly implemented using "logic compiler" software. Similar to the software compiler used in program development, the original code before compilation must be written in a specific programming language, called a Hardware Description Language (HDL). There are many HDLs, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, and RHDL (Ruby Hardware Description Language). Currently, the most commonly used are VHDL (Very-High-Speed ​​Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should understand that by simply performing some logic programming on the method flow using one of these hardware description languages ​​and programming it into an integrated circuit, the hardware circuit implementing the logical method flow can be easily obtained.

[0241] The controller can be implemented in any suitable manner. For example, it can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicon Labs C8051 F320. A memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also recognize that, in addition to implementing the controller in purely computer-readable program code form, the same functionality can be achieved by logically programming the method steps to make the controller take the form of logic gates, switches, ASICs, programmable logic controllers, and embedded microcontrollers. Therefore, such a controller can be considered a hardware component, and the means included therein for implementing various functions can also be considered as structures within the hardware component. Alternatively, the means for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.

[0242] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.

[0243] For ease of description, the above devices are described separately by function as various units. Of course, in implementing this application, the functions of each unit can be implemented in one or more software and / or hardware.

[0244] Those skilled in the art will understand that one or more embodiments of this specification can be provided as a method, system, or computer program product. Therefore, the invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0245] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0246] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0247] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0248] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0249] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0250] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital character versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0251] This application can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0252] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A video recognition method, comprising: Obtain the video to be recognized; Obtain a candidate video set for identifying the originality of the video to be identified; For each candidate video in the candidate video set, a comprehensive similarity score is determined between the video to be identified and each candidate video in at least two target feature dimensions in a specific feature dimension, resulting in multiple comprehensive similarity scores. The specific feature dimensions include at least two feature dimensions from visual feature dimension, text feature dimension, multimodal feature dimension, and face feature dimension. The comprehensive similarity score of a candidate video is obtained by concatenating the similarity scores of the candidate video and the video to be identified in the at least two target feature dimensions. At least some of the comprehensive similarity scores are input into the video recognition model to obtain the original video probability value of the video to be identified, which is output by the video recognition model based on the at least some comprehensive similarity scores. If the probability value of the original video is greater than or equal to a preset threshold, then the video to be identified is determined to be an original video.

2. The method according to claim 1, wherein determining the comprehensive similarity score between the video to be identified and each candidate video in the candidate video set in at least two target feature dimensions in a specific feature dimension specifically includes: For any candidate video among the candidate videos, determine the first similarity score between the video to be identified and any candidate video on the visual feature dimension; Determine the second similarity score between the video to be identified and any candidate video on the text feature dimension; Determine the third similarity score between the video to be identified and any candidate video on the multimodal feature dimension; By concatenating the first similarity score, the second similarity score, and the third similarity score, the comprehensive similarity score between the video to be identified and any candidate video is obtained.

3. The method according to claim 2, wherein determining the first similarity score between the video to be identified and any candidate video on the visual feature dimension specifically includes: Extract a first number of first video frames from the video to be identified; Extract a second number of second video frames from any of the candidate videos; For any first video frame and any second video frame, determine the visual feature similarity score between the first video frame and the second video frame; If the visual feature similarity score is greater than or equal to the first preset threshold, then any first video frame and any second video frame are determined as a video frame similarity pair. Based on the number of video frame similarity pairs determined by the first number of first video frames and the second number of second video frames, and the visual feature similarity score corresponding to each video frame similarity pair, the first similarity score of the video to be identified and any candidate video in the visual feature dimension is determined.

4. The method according to claim 2, wherein determining the second similarity score between the video to be identified and any candidate video on the text feature dimension specifically includes: Obtain first text information from the video to be identified; The text information type of the first text information includes at least one of the title text information in the video to be identified, image text information identified based on the image in the video to be identified, and audio text information identified based on the audio in the video to be identified; Obtain second text information from any of the candidate videos, wherein the text information type of the second text information corresponds to the text information type of the first text information; Based on the first text features extracted from the first text information and the second text features extracted from the second text information, a second similarity score is determined between the video to be identified and any candidate video in the text feature dimension.

5. The method according to claim 2, wherein determining the third similarity score between the video to be identified and any candidate video on the multimodal feature dimension specifically includes: First multimodal information is obtained from the video to be identified, the first multimodal information including video frame information extracted from the video to be identified and text information obtained from the video to be identified; Second multimodal information is obtained from any of the candidate videos, the second multimodal information including video frame information extracted from any of the candidate videos and text information obtained from any of the candidate videos; Based on the first multimodal features extracted from the first multimodal information and the second multimodal features extracted from the second multimodal information, a third similarity score is determined between the video to be identified and any candidate video in the multimodal feature dimension.

6. The method according to claim 2, wherein the video to be identified includes a face image, and the step of concatenating the first similarity score, the second similarity score, and the third similarity score to obtain the comprehensive similarity score between the video to be identified and any candidate video specifically includes: Determine the fourth similarity score between the video to be identified and any candidate video on the facial feature dimension; By concatenating the first similarity score, the second similarity score, the third similarity score, and the fourth similarity score, the comprehensive similarity score between the video to be identified and any candidate video is obtained.

7. The method according to claim 6, wherein determining the fourth similarity score between the video to be identified and any candidate video on the facial feature dimension specifically includes: Extract the first face image from the video to be identified; Extract a second face image from any of the candidate videos; Based on the first facial features extracted from the first facial image and the second facial features extracted from the second facial image, a fourth similarity score is determined between the video to be identified and any candidate video in the facial feature dimension.

8. The method according to claim 1, wherein inputting at least a portion of the comprehensive similarity scores from the plurality of comprehensive similarity scores into a video recognition model to obtain an original video probability value of the video to be recognized output by the video recognition model based on the at least a portion of the comprehensive similarity scores, specifically includes: Obtain the weight values ​​corresponding to each feature dimension; Based on the similarity scores corresponding to the feature dimensions of the target weight value, the comprehensive similarity scores are sorted to obtain a comprehensive similarity score sequence; The first preset number of comprehensive similarity scores located in the comprehensive similarity score sequence are input into the video recognition model to obtain the original video probability value of the video to be recognized, which is output by the video recognition model based on the first preset number of comprehensive similarity scores.

9. The method according to claim 1, wherein obtaining a candidate video set for identifying the originality of the video to be identified specifically includes: Based on the aforementioned visual feature dimensions, a first candidate video set is determined for identifying the originality of the video to be identified; The similarity score of each first candidate video in the first candidate video set and the video to be identified in the visual feature dimension is greater than or equal to the first preset threshold. Based on the text feature dimension, a second candidate video set is determined for identifying the originality of the video to be identified; the similarity score of each second candidate video in the second candidate video set with the video to be identified on the text feature dimension is greater than or equal to a second preset threshold. Based on the multimodal feature dimension, a third candidate video set is determined for identifying the originality of the video to be identified; the similarity score of each third candidate video in the third candidate video set with the video to be identified on the multimodal feature dimension is greater than or equal to a third preset threshold. The candidate video set is obtained by merging the first candidate video set, the second candidate video set, and at least two video sets from the third candidate video set.

10. The method according to claim 9, wherein the video to be identified includes a face image, and the step of merging at least two video sets from the first candidate video set, the second candidate video set, and the third candidate video set to obtain the candidate video set specifically includes: Based on the facial feature dimensions, a fourth candidate video set is determined for identifying the originality of the video to be identified; The similarity score between each fourth candidate video in the fourth candidate video set and the video to be identified in the facial feature dimension is greater than or equal to the fourth preset threshold. The candidate video set is obtained by merging at least two video sets from the first candidate video set, the second candidate video set, the third candidate video set, and the fourth candidate video set.

11. A video recognition device, comprising: The first acquisition module is used to acquire the video to be recognized; The second acquisition module allows the user to acquire a set of candidate videos for identifying the originality of the video to be identified; The first determining module is used to determine, for each candidate video in the candidate video set, the comprehensive similarity score between the video to be identified and each candidate video in a specific feature dimension, at least two target feature dimensions, to obtain multiple comprehensive similarity scores; the specific feature dimension includes at least two feature dimensions from visual feature dimension, text feature dimension, multimodal feature dimension and face feature dimension, and the comprehensive similarity score of a candidate video is obtained by concatenating the similarity scores of the candidate video and the video to be identified in the at least two target feature dimensions; The input module is used to input at least a portion of the comprehensive similarity scores from the plurality of comprehensive similarity scores into the video recognition model, and to obtain the original video probability value of the video to be recognized output by the video recognition model based on the at least a portion of the comprehensive similarity scores; The second determining module is used to determine the video to be identified as an original video if the probability value of the original video is greater than or equal to a preset threshold.

12. A computing device, comprising: Memory and processor; The memory is used to store computer programs or instructions, and the processor is used to execute the computer programs or instructions, wherein when the computer programs or instructions are executed by the processor, they implement the steps of the method according to any one of claims 1 to 10.

13. A computer-readable storage medium storing computer instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 10.

14. A computer program product comprising a computer program or instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 10.