Method, device, medium and electronic device for determining evaluation index
By determining the detection box and sample box in the video timestamp coordinate system, and calculating the evaluation index with intersection and side length information, the problem of insufficient robustness of video copy detection in the prior art is solved, and more accurate copy fragment recognition and evaluation is achieved.
Patent Information
- Application Number
- CN202210143724.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-18
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-02-18
AI Technical Summary
In the prior art, when detecting video copy, the evaluation indicators are insufficiently robust, making it difficult to effectively identify and evaluate copy clips between videos.
By constructing a target coordinate system with video timestamps as coordinates, the detection box and sample box are determined, and evaluation indicators are calculated using information such as intersection and side length to improve the robustness of copying video detection results.
It improves the robustness of video copy detection results, can more accurately identify and evaluate copy clips between videos, and enhances the accuracy of infringement positioning.
Smart Images

Figure CN114612818B_ABST
Abstract
Description
Technical Field
[0001] The present specification relates to the field of video processing technology, and in particular to a method and device for determining an evaluation index, a computer-readable storage medium, and an electronic device. Background Art
[0002] The video obtained by transforming and editing the original video material, such as photometric transformation, geometric transformation, editing transformation, encoding format, and time series editing operations, can be called a copied video. Copying a video will infringe the original video material.
[0003] In order to avoid the harm caused by video infringement, it is necessary to detect the video resources to determine the copied videos. The prior art also provides an evaluation scheme for the detection results of copied videos, however, the robustness of the evaluation scheme provided by the related art needs to be improved.
[0004] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of this specification, and therefore may include information that does not constitute prior art known to ordinary technicians in the field. Summary of the invention
[0005] The purpose of this specification is to provide a method, device, computer-readable storage medium and electronic device for determining an evaluation index, which improve the robustness of the copy video detection result at least to a certain extent.
[0006] Other features and advantages of the present specification will become apparent from the following detailed description, or may be learned in part from the practice of the present specification.
[0007] According to one aspect of the present specification, a method for determining an evaluation index is provided, the method comprising: constructing a target coordinate system with a timestamp sequence of a first video as a horizontal axis coordinate and a timestamp sequence of a second video as a vertical axis coordinate, wherein each coordinate point in the target coordinate system corresponds to a video frame pair with a corresponding timestamp; determining a detection frame in the target coordinate system, wherein the detection frame corresponds to a suspected copy segment pair between the first video and the second video, and a predicted value of an inter-frame similarity of the video frame pair corresponding to the coordinate points in the detection frame is greater than a preset value; determining a sample frame in the target coordinate system, wherein an actual value of an inter-frame similarity of the video frame pair corresponding to the coordinate points in the sample frame is greater than the preset value; and determining an evaluation index for a prediction result of the suspected copy segment pair based on an intersection between the sample frame and the detection frame, a side length of the sample frame, and a side length of the detection frame.
[0008] According to another aspect of the present specification, a device for determining an evaluation index is provided, the device comprising: a coordinate system construction module, a detection frame determination module, a sample frame determination module, and an evaluation module.
[0009] Among them, the above-mentioned coordinate system construction module is used to construct a target coordinate system with the timestamp sequence of the first video as the horizontal axis coordinate and the timestamp sequence of the second video as the vertical axis coordinate, and each coordinate point in the target coordinate system corresponds to a video frame pair with a corresponding timestamp; the above-mentioned detection frame determination module is used to determine a detection frame in the target coordinate system, the detection frame corresponds to a suspected copy segment pair between the first video and the second video, and the predicted value of the inter-frame similarity of the video frame pair corresponding to the coordinate point in the detection frame is greater than a preset value; the above-mentioned sample frame determination module is used to determine a sample frame in the target coordinate system, and the actual value of the inter-frame similarity of the video frame pair corresponding to the coordinate point in the sample frame is greater than the preset value; and the above-mentioned evaluation module is used to determine the evaluation index of the prediction result of the suspected copy segment pair according to the intersection between the sample frame and the detection frame, the side length of the sample frame and the side length of the detection frame.
[0010] According to another aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method for determining the evaluation index in the above embodiment is implemented.
[0011] According to one aspect of the present specification, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements a method for determining an evaluation index as in the above-mentioned embodiment when executing the computer program.
[0012] The methods, devices and electronic devices for determining evaluation indicators provided in the embodiments of this specification have the following technical effects:
[0013] In the scheme provided by the exemplary embodiments of this specification, two videos in which copied video clips may exist are input, and a target coordinate system is determined based on the timestamps of the two videos, so that the chronological relationship of the video frames in the video can be effectively represented, which is conducive to marking the suspected copy clip pairs and the actual copy clip pairs that actually have an infringing relationship in the target coordinate system. Specifically, the suspected copy clips are represented as detection boxes in the above target coordinate system, and the actual copy clips are represented as sample boxes in the above target coordinate system. In order to enable the evaluation indicators provided by the embodiments of this specification to take into account the equivalent characteristics of video segmentation to increase the robustness of the evaluation indicators, the embodiments of this specification further determine the evaluation indicators for the prediction results of the above suspected copy clip pairs based on the intersection between the sample box and the detection box, the side length of the sample box, and the side length of the detection box.
[0014] The above-mentioned video segmentation equivalent characteristic means that it is sometimes difficult to determine the boundary of the copy segment, such as Figure 1 As shown, the middle frames of the video part are modified and / or briefly inserted into other video frames. It is considered that in these cases, it is reasonable to mark the copied video segments as a whole segment and multiple continuous segments. In other words, the evaluation index value (such as recall rate, precision rate) of "marked as a whole segment" should be equal to the evaluation index value (such as recall rate, precision rate) of "marked as multiple segments". However, through the evaluation scheme provided in the embodiment of the specification, the calculation results in the two cases are the same. However, in the evaluation scheme provided by the related art, the evaluation index value of "marked as a whole segment" is not necessarily equal to the evaluation index value of "marked as multiple segments". It can be seen that the scheme provided in the embodiment of this specification has high robustness.
[0015] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present specification. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The accompanying drawings herein are incorporated into and constitute a part of the specification, illustrate embodiments consistent with the specification, and together with the specification are used to explain the principles of the specification. Obviously, the accompanying drawings described below are only some embodiments of the specification, and for those of ordinary skill in the art, other accompanying drawings can be obtained based on these accompanying drawings without creative work.
[0017] Figure 1 A schematic diagram of copying a video clip according to an embodiment of the present specification.
[0018] Figure 2 This is a schematic diagram of the architecture of the travel information sharing system provided in the embodiments of this specification.
[0019] Figure 3A flowchart of a method for determining an evaluation index provided in an embodiment of this specification.
[0020] Figure 4 A flowchart of a method for determining a detection frame provided in an embodiment of this specification.
[0021] Figure 5 A schematic diagram of determining a detection frame based on a target coordinate system provided in an embodiment of this specification.
[0022] Figure 6 A schematic flow chart of a method for determining a sample frame provided in an embodiment of this specification.
[0023] Figure 7 A schematic diagram of determining a sample frame based on a target coordinate system according to an embodiment of this specification.
[0024] Figure 8 A flowchart of a method for calculating a recall rate of video clip copy detection provided in an embodiment of the present specification.
[0025] Fig. 9 A schematic diagram of determining a recall rate based on a target coordinate system according to an embodiment of this specification.
[0026] Fig.10 A flowchart of a method for calculating the accuracy of video clip copy detection provided in an embodiment of the present specification.
[0027] Fig.11 A schematic diagram of determining accuracy based on a target coordinate system according to an embodiment of this specification.
[0028] Fig.12 A schematic diagram of determining an evaluation index based on a target coordinate system provided in one embodiment of this specification.
[0029] Fig.13 A schematic diagram of the structure of a device for determining evaluation indicators provided in one embodiment of this specification.
[0030] Fig.14 A schematic diagram of the structure of a device for determining evaluation indicators provided in another embodiment of this specification.
[0031] Fig.15 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this specification. DETAILED DESCRIPTION
[0032] In order to make the purpose, technical solutions and advantages of this specification more clear, the embodiments of this specification will be further described in detail below with reference to the accompanying drawings.
[0033] When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this specification. Instead, they are only examples of devices and methods consistent with some aspects of this specification as detailed in the attached claims.
[0034] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as being limited to the examples set forth herein; on the contrary, these embodiments are provided so that this specification will be more comprehensive and complete, and the concept of the example embodiments will be fully conveyed to those skilled in the art. The described features, structures, or characteristics may be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the embodiments of this specification. However, those skilled in the art will appreciate that the technical solutions of this specification may be practiced while omitting one or more of the specific details, or other methods, components, devices, steps, etc. may be adopted. In other cases, known technical solutions are not shown or described in detail to avoid obscuring various aspects of this specification.
[0035] In addition, the drawings are only schematic illustrations of the present specification and are not necessarily drawn to scale. The same reference numerals in the figures represent the same or similar parts, and their repeated description will be omitted. Some of the block diagrams shown in the drawings are functional entities and do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0036] Video copy detection is divided into copy detection at the video granularity and copy detection at the fragment granularity. Among them, copy detection at the video granularity only needs to output whether there are repeated infringing fragments in the two videos, while the fragment granularity needs to specifically output the time position of the infringing fragments of the two videos. In other words, copy detection at the fragment granularity can realize video infringement positioning. Specifically, video infringement positioning means: since the infringing video fragment only contains part of the original video and is embedded in another video content, the video infringement positioning needs to locate the infringing time fragment of the infringing video and the infringed time fragment of the original video, so as to realize the effective pointing of the video infringing fragment in the actual copyright protection process, thereby improving the efficiency and reliability of the detection system.
[0037] The above-mentioned video-level copy detection has been relatively mature in previous research and applications. The embodiment of this specification is aimed at an evaluation scheme for copy detection at a segment granularity.
[0038] In the fragment-level copy detection scheme provided by the related art, for example, the numerators of the precision rate and the recall rate are all correctly detected fragments, where a correctly detected fragment is defined as a correctly detected fragment as long as there is one frame overlap with the actual infringing fragment. The denominator of the precision rate is the number of all detected fragments, and the denominator of the recall rate is the number of fragments that are actually marked as true copies. The fragment-level precision / recall rate and the frame-level precision / recall rate mentioned in the above related art have their limitations.
[0039] Specifically, the evaluation index provided by the related technology is only suitable for copy detection of clips and videos, that is, it requires marked infringing clips and possibly infringing videos as input, but is not applicable to any two complete videos. Therefore, this evaluation method is unrealistic in actual scenarios, that is, it has poor practicality. At the same time, for the clip precision / recall rate, as long as the detected clip and the actual marked clip overlap by one frame, it is considered to be the correct calculation method, which will lead to a low perception of the accuracy of the evaluation index in locating infringements. In addition, the solution provided by the related technology does not take into account the equivalent characteristics of the segmentation of video copies. In other words, due to the difficulty in determining the boundaries of the copied clips, such as Figure 1 As shown in the figure, in the case where the middle frame of the video part is modified and / or briefly inserted into other video frames, it is reasonable to mark the copied video segment as a whole segment and multiple consecutive segments. However, based on the solution provided by the relevant technology, the evaluation index value (such as recall rate and precision rate) of "marking as a whole segment" is not equal to the evaluation index value (such as recall rate and precision rate) of "marking as multiple segments".
[0040] The embodiments of this specification can solve the above technical problems and provide this technical solution. Specifically, Figure 2 A schematic diagram of the system architecture of the federated learning implementation scheme provided in the embodiments of this specification.
[0041] like Figure 2 As shown, the system architecture may include a network 130 and a computing device 140. The video 10 and the video 20 may be sent to the computing device 140 via the network 130. Of course, the video 10 and the video 20 may also be transmitted to the computing device 140 via a data line or the like.
[0042] Exemplarily, video 10 and video 20 may be any two videos, for example Figure 1 The video 110 and the video 120 shown in FIG. 110 and FIG. 120. It can be seen that the solution provided in the embodiment of this specification can be applied to any two videos and used to evaluate the prediction results of the suspected copy segments between the two videos. Among them, the time axis T can represent the timestamp of the video 120. Specifically, Figure 1 , video frames with time stamps from tm to video frames with time stamps tn in video 120 are shown.
[0043] Exemplarily, the network 130 may be a communication medium of various connection types capable of providing a communication link, such as a wired communication link, a wireless communication link, or a fiber optic cable, etc., and this specification does not limit this.
[0044] Exemplarily, the computing device can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, as well as big data and artificial intelligence platforms. It can also be a notebook with computing capabilities, etc.
[0045] In the solution provided by the embodiment of this specification, after receiving video 10 and video 20, the computing device 140 uses the timestamp sequence of video 10 as the horizontal axis T1 coordinate and the timestamp sequence of video 20 as the vertical axis T2 coordinate to construct a target coordinate system. Each coordinate point in the target coordinate system corresponds to a video frame pair with a corresponding timestamp. Figure 2 The coordinate point s in the left line of the target corresponds to the video frame pair of the video frame with time stamp t1 in video 10 and the video frame with time stamp t2 in video 20. The order of the video frame sequences of the two videos can be reflected by means of the coordinate system.
[0046] Then, the computing device 140 determines the detection frame B1 and the detection frame B3 in the above target coordinate system. Each detection frame corresponds to a suspected copy segment pair between video 10 and video 20, and the predicted value of the inter-frame similarity of the video frame pair corresponding to the coordinate point in each detection frame is greater than the preset value. And, the computing device 140 also determines the sample frame B2 in the above target coordinate system. The actual value of the inter-frame similarity of the video frame pair corresponding to the coordinate point in the sample frame is greater than the preset value.
[0047] Furthermore, the computing device 140 determines an evaluation index for the prediction result of the suspected copy segment pair according to the intersection between the sample frame and the detection frame, the side length of the sample frame, and the side length of the detection frame.
[0048] The following first passes Figures 3 to 12 The following is a detailed description of the method for determining the evaluation index provided in this specification:
[0049] For example, Figure 3 A flowchart of a method for determining an evaluation index provided in an embodiment of this specification. Figure 3 , the method shown in this embodiment includes:
[0050] S310, constructing a target coordinate system with the timestamp sequence of the first video as the horizontal axis coordinate and the timestamp sequence of the second video as the vertical axis coordinate, wherein each coordinate point in the target coordinate system corresponds to a video frame pair with a corresponding timestamp.
[0051] Exemplarily, the first video and the second video may be any two videos. Figure 2 The target coordinate system uses the timestamp sequence of video 10 (first video) as the horizontal axis T1 coordinate, so that the order of the video frame sequence in the first video can be reflected based on the horizontal axis; and uses the timestamp sequence of video 20 (second video) as the vertical axis T2 coordinate, so that the order of the video frame sequence in the second video can be reflected based on the vertical axis.
[0052] Among them, the coordinate system not only reflects the order of the video frame sequences of the two videos through the horizontal and vertical axes respectively, but also, since each coordinate point can simultaneously correspond to a video frame of the first video (recorded as video frame a) and a video frame of the second video (recorded as video frame b), the video frame a and the video frame b corresponding to the same coordinate point are recorded as a video frame pair. Therefore, the coordinate system can also simply and clearly reflect the video frame pair between the two videos; further, the line segment parallel to the horizontal axis in the coordinate system can reflect a video clip of the first video (recorded as video clip a), and the line segment parallel to the vertical axis in the coordinate system can reflect a video clip of the second video (recorded as video clip b), wherein the above two line segments can construct a rectangular area, and the video clip a and the video clip b corresponding to the rectangular area are recorded as a video clip pair. Therefore, the coordinate system can also simply and clearly reflect the video clip pair between the two videos.
[0053] S320, determining a detection frame in the target coordinate system, the detection frame corresponding to the suspected copy segment pair between the first video and the second video, and the predicted value of the inter-frame similarity of the video frame pair corresponding to the coordinate points in the detection frame is greater than a preset value. And, S330, determining a sample frame in the target coordinate system, the actual value of the inter-frame similarity of the video frame pair corresponding to the coordinate points in the sample frame is greater than a preset value.
[0054] In this embodiment, reference Figure 2 , a detection frame B1 and a detection frame B3 are determined in the above target coordinate system. Each detection frame corresponds to a suspected copy segment pair between video 10 (first video) and video 20 (second video), and the predicted value of the inter-frame similarity of the video frame pair corresponding to the coordinate point in each detection frame is greater than a preset value. And, a sample frame B2 is determined in the above target coordinate system. The actual value of the inter-frame similarity of the video frame pair corresponding to the coordinate point in the sample frame B2 is greater than a preset value.
[0055] It should be noted that there is no particular order in which S320 and S330 are executed. S320 may be executed first and then S330, S330 may be executed first and then S320, or S320 and S330 may be executed together.
[0056] S340 , determining an evaluation index for the prediction result of the suspected copy segment pair according to the intersection between the sample frame and the detection frame, the side length of the sample frame, and the side length of the detection frame.
[0057] Exemplarily, the above evaluation indicators include: Recall, Precision, F-Score, etc. The specific implementation method for determining the evaluation indicators will be described in detail in the subsequent embodiments.
[0058] according to Figure 3 It can be seen from the scheme provided by the illustrated embodiment that, compared with the evaluation index provided by the related art which is only suitable for copy detection of clips and videos, the scheme provided by the embodiment of this specification is applicable to copy detection between any two videos. Therefore, the scheme provided by the embodiment of this specification has a wide range of applications and strong practicality. On the other hand, the scheme provided by the embodiment of this specification determines the evaluation index based on the intersection between the detection frame and the sample frame, rather than adopting the calculation method in the related art that as long as the detected clip and the actual marked clip have one frame overlap, it is considered to be correct. It can be seen that the scheme provided by the embodiment of this specification can improve the perception of the accuracy of infringement positioning. On the other hand, the evaluation index provided by the embodiment of this specification takes into account the equivalent characteristics of video segmentation (which will be explained in detail in the following embodiments), which effectively improves the robustness of the evaluation index.
[0059] The following combination Figures 4 to 12 The embodiment shown in the figure is Figure 3 The specific implementation methods of each step of the embodiment shown are described in detail:
[0060] As a specific implementation of S310, refer to Figure 2 , using the timestamp sequence of video 10 as the horizontal axis T1 coordinate, and the timestamp sequence of video 20 as the vertical axis T2 coordinate, to construct the above-mentioned target coordinate system. It should be noted that the positive direction of the horizontal axis can be the direction of increasing timestamps of the first video or the direction of decreasing timestamps. Similarly, the positive direction of the vertical axis can be the direction of increasing timestamps of the second video or the direction of decreasing timestamps. Among them, the coordinate system can not only reflect the chronological relationship of the video frame sequences of the two videos, but also can simply and clearly reflect the video frame pairs and video clip pairs between the two videos. For example, it can simply and clearly reflect the detection box and sample box in the embodiments of this specification.
[0061] As a specific implementation of S320, Figure 4A flow chart of a method for determining a detection frame provided in an embodiment of the present specification. The detection frame corresponds to a suspected copy segment pair between a first video and a second video. Exemplarily, the suspected copy segment pair may come from a sample data set that has been manually marked, wherein the marked portion is the suspected copy segment pair. The suspected copy segment pair may also be predicted based on a relevant algorithm. For example, by using a dynamic matching algorithm or a time-domain neural network algorithm, at least one segment pair between the first video and the second video whose similarity is greater than a preset value is obtained, and the suspected copy segment pair may be obtained.
[0062] In an exemplary embodiment, in a case where the suspected copy segment pair includes a first segment in a first video and a second segment in a second video, Figure 4 The illustrated embodiment includes:
[0063] S410, determining the horizontal coordinate starting point and the horizontal coordinate end point of the detection frame according to the start frame timestamp and the end frame timestamp of the first segment. S420, determining the vertical coordinate starting point and the vertical coordinate end point of the detection frame according to the start frame timestamp and the end frame timestamp of the second segment. And, S430, determining the detection frame in the target coordinate system according to the horizontal coordinate starting point and the horizontal coordinate end point of the detection frame and the vertical coordinate starting point and the vertical coordinate end point of the detection frame.
[0064] Exemplary, reference Figure 5 According to the start frame timestamp and the end frame timestamp of the first segment in the first video, the relevant coordinate position is determined in the horizontal axis T1, and the line segment X1 corresponding to the first segment is obtained, and the horizontal axis coordinate starting point and the horizontal axis coordinate end point of the detection frame B1 are also determined. According to the start frame timestamp and the end frame timestamp of the second segment in the second video, the relevant coordinate position is determined in the vertical axis T2, and the line segment X2 corresponding to the second segment is obtained, and the vertical axis coordinate starting point and the horizontal axis coordinate end point of the detection frame B1 are also determined. Further, according to the horizontal axis coordinate starting point and the horizontal axis coordinate end point of the detection frame B1 and the vertical axis coordinate starting point and the vertical axis coordinate end point of the detection frame B1, the detection frame B1 is determined in the target coordinate system.
[0065] Similarly, according to the start frame timestamp and end frame timestamp of another segment X1' in the first video, the relevant coordinate position is determined in the horizontal axis T1, and the line segment X1' corresponding to the first segment is obtained, and the horizontal axis coordinate starting point and the horizontal axis coordinate end point of the detection frame B2 are also determined. According to the start frame timestamp and end frame timestamp of another segment X2' in the second video, the corresponding coordinate position is determined in the vertical axis T2, and the line segment X2' corresponding to the second segment is obtained, and the vertical axis coordinate starting point and the horizontal axis coordinate end point of the detection frame B2 are also determined. Further, according to the horizontal axis coordinate starting point and the horizontal axis coordinate end point of the detection frame B2 and the vertical axis coordinate starting point and the vertical axis coordinate end point of the detection frame B2, the detection frame B2 is determined in the target coordinate system.
[0066] pass Figure 4 It can be seen from the corresponding embodiment that all suspected copy segment pairs between the two videos can be clearly and concisely marked in the above target coordinate system, and the inter-frame sequence relationship in the suspected copy segment can also be reflected through the block diagram in the above target coordinate system, which has an accurate, practical and simple effect. Among them, the predicted value of the inter-frame similarity of the video frame pair corresponding to the coordinate point in each detection frame is greater than the preset value.
[0067] As a specific implementation of S330, Figure 6 A flow chart of a method for determining a sample frame provided in an embodiment of the present specification. The sample frame corresponds to an actual copy segment pair between a first video and a second video. Exemplarily, the actual copy segment pair may be known data, for example, it is known that video segment a and video segment b in the first video are copy segments. The actual copy segment pair may also be determined by similarity calculation. For details, refer to Figure 6 The embodiment shown.
[0068] In S510, the similarity pxy between the xth video frame with timestamp x in the first video and the yth video frame in the second video is calculated, where x is taken in ascending order of the timestamp of the first video, and y is the timestamp of the second video.
[0069] Exemplarily, the similarity between two video frames can be calculated by extracting image features from the video frames. Considering the amount of calculation required to determine the similarity, in this embodiment, a video frame pair for calculating the similarity is determined at intervals of a preset number of frames. For example, a video frame pair for calculating the similarity is determined at intervals of ten frames in both the first video and the second video, the similarity between the first frame of the first video and the first frame of the second video is calculated, the similarity between the 11th frame of the first video and the 11th frame of the second video is calculated, and so on.
[0070] In S520, a mapping relationship is established between the similarity pxy and the coordinate point (x, y) in the target coordinate system. And, in S530, a coordinate region with a similarity greater than a preset value is determined as a region within a sample frame to obtain a sample frame.
[0071] In this embodiment, a mapping relationship is established between the similarity of the video frame pairs between the two videos and the coordinate points in the target coordinate system. In other words, an association relationship is established between the coordinate points reflecting the video frame pairs and the similarity values of the corresponding video frame pairs. Thus, the target coordinate system can reflect the similarity of each video frame pair between the two videos. Figure 7 As shown, the similarity px1y1 between the x1th video frame in the first video and the y1th video frame in the second video is associated with the coordinate point corresponding to the video frame pair (the x1th video frame and the y1th video frame); the similarity px2y2 between the x2th video frame in the first video and the y2th video frame in the second video is associated with the coordinate point corresponding to the video frame pair (the x2th video frame and the y2th video frame).
[0072] Since the above similarity is the actual similarity of the video frame pair, and the actual value of the inter-frame similarity of the video frame pair corresponding to the coordinate point in the sample frame is greater than the preset value, the actual copy segment pair can be further determined based on these similarities. Therefore, the coordinate area with a similarity greater than the preset value is determined as the internal area of the sample frame, and then all sample frames of the actual copy segment pair between the two videos are determined in the target coordinate system.
[0073] pass Figure 6 It can be seen from the corresponding embodiment that all actual copy segment pairs between the two videos can be clearly and concisely marked in the above-mentioned target coordinate system. At the same time, the frame sequence relationship in the actual copy segment can also be reflected through the block diagram in the above-mentioned target coordinate system, which has an accurate, practical and simple and clear effect.
[0074] The following introduction Figure 3 The specific implementation of S340 in the embodiment shown is as follows:
[0075] In an exemplary embodiment, when the evaluation indicator is the recall rate, Figure 8 This is a flowchart of a method for calculating a recall rate of video clip copy detection provided in an embodiment of the present specification. The embodiment shown in the flowchart includes S810-S830.
[0076] In S810, the intersection of each sample frame and all detection frames in the target coordinate system is obtained to obtain a first set of intersections.
[0077] Exemplary, reference Fig. 9 , for the i-th sample frame G i , get the sample frame G iThe intersection O between all detection boxes in the target coordinate system i If the target coordinate system contains a total of m sample frames, then in this embodiment Recorded as the first set of intersection. Fig. 9 The solid line box represents the detection box, the dotted line box represents the sample box, and the shaded area represents the intersection.
[0078] In S820, the union lengths of all first sides corresponding to the first group of intersections are obtained to obtain a first union length, the union lengths of the first sides corresponding to all sample frames are obtained to obtain a second union length, and the ratio of the first union length to the second union length is calculated to obtain a first ratio, and the first side is parallel to the horizontal axis of the target coordinate system.
[0079] Exemplary, reference Fig. 9 , with sample frame G i The corresponding intersection O i For example, the intersection O i The first edges are projected onto the horizontal axis T1 of the target coordinate system, and the length of the line segment obtained by the projection is determined as the length of the union In the case where there are a total of m sample frames in the target coordinate system, the first set of intersections The first sides of all intersections in the target coordinate system are projected onto the horizontal axis T1, and the length of the line segment obtained by the projection is determined as the length of the first union.
[0080] With sample frame G i For example, take the sample frame G i The first edges are projected onto the horizontal axis T1 of the target coordinate system, and the length of the line segment obtained by the projection is determined as the length of the union When the target coordinate system contains a total of m sample frames, the first sides of all sample frames are projected onto the horizontal axis T1 of the target coordinate system, and the length of the line segment obtained by the projection is determined as the second union length
[0081] Then we can further get the first ratio can be expressed as:
[0082] Continue to refer Figure 8 In S820', the union lengths of all second sides corresponding to the first group of intersections are obtained to obtain a third union length, the union lengths of the second sides corresponding to all sample frames are obtained to obtain a fourth union length, and the ratio of the third union length to the fourth union length is calculated to obtain a second ratio, and the second side is parallel to the longitudinal axis of the target coordinate system.
[0083] For example, see Fig. 9 , with sample frame G i The corresponding intersection O i For example, the intersection Oi The second side is projected onto the longitudinal axis T2 of the target coordinate system, and the length of the line segment obtained by the projection is determined as the length of the union In the case where there are a total of m sample frames in the target coordinate system, the first set of intersections The second sides of all intersections in the target coordinate system are projected onto the longitudinal axis T2, and the length of the projected line segment is determined as the length of the third union.
[0084] With sample frame G i For example, take the sample frame G i The second side is projected onto the longitudinal axis T2 of the target coordinate system, and the length of the line segment obtained by the projection is determined as the length of the union When the target coordinate system contains a total of m sample frames, the second sides of all sample frames are projected onto the longitudinal axis T2 of the target coordinate system, and the length of the line segment obtained by the projection is determined as the fourth union length
[0085] Then we can further get the second ratio can be expressed as:
[0086] In S830, a recall rate of the predicted suspected copy segment pairs is determined according to the first ratio and the second ratio.
[0087] Exemplarily, the product of the first ratio and the second ratio is determined as the recall rate for predicting the suspected copy segment pair:
[0088]
[0089] In an exemplary embodiment, when the evaluation indicator is accuracy, Fig.10 This is a flowchart of a method for calculating the accuracy of video clip copy detection provided in an embodiment of this specification. The embodiment shown in this figure includes S1010-S1030.
[0090] In S1010, the intersection of each detection frame and all sample frames in the target coordinate system is obtained to obtain a second set of intersections.
[0091] Exemplary, reference Fig.11 , for the jth detection box P j , get the detection box P j The intersection Q of all sample frames in the target coordinate system j . If the above target coordinate system contains a total of n detection boxes, then in this embodiment Recorded as the second set of intersection. Fig.11 The solid line box represents the detection box, the dotted line box represents the sample box, and the shaded area represents the intersection.
[0092] In S1020, the union lengths of all first edges corresponding to the second group of intersections are obtained to obtain the fifth union length, the union lengths of the first edges corresponding to all detection boxes are obtained to obtain the sixth union length, and the ratio of the fifth union length to the sixth union length is calculated to obtain a third ratio, and the first edge is parallel to the horizontal axis of the target coordinate system.
[0093] Exemplary, reference Fig.11 , to detect box P j The corresponding intersection Q j For example, the intersection Q j The first edges are projected onto the horizontal axis T1 of the target coordinate system, and the length of the line segment obtained by the projection is determined as the length of the union When there are n detection boxes in the target coordinate system, the second set of intersections The first sides of all intersections in the target coordinate system are projected onto the horizontal axis T1, and the length of the line segment obtained by the projection is determined as the length of the fifth union.
[0094] Take the detection box P j For example, the detection box P j The first edges are projected onto the horizontal axis T1 of the target coordinate system, and the length of the line segment obtained by the projection is determined as the length of the union When there are n detection frames in the target coordinate system, the first sides of all detection frames are projected onto the horizontal axis T1 of the target coordinate system, and the length of the line segment obtained by the projection is determined as the sixth union length.
[0095] Then we can further get the third ratio can be expressed as:
[0096] Continue to refer Figure 8 In S1020', the union lengths of all second sides corresponding to the second group of intersections are obtained to obtain a seventh union length, the union lengths of the second sides corresponding to all detection boxes are obtained to obtain an eighth union length, and the ratio of the seventh union length to the eighth union length is calculated to obtain a fourth ratio, and the second side is parallel to the longitudinal axis of the target coordinate system.
[0097] For example, see Fig.11 , to detect box P j The corresponding intersection Q j For example, the intersection Q j The second side is projected onto the longitudinal axis T2 of the target coordinate system, and the length of the line segment obtained by the projection is determined as the length of the union When there are n detection boxes in the target coordinate system, the second set of intersections The second sides of all intersections in the target coordinate system are projected onto the longitudinal axis T2, and the length of the line segment obtained by the projection is determined as the length of the seventh union.
[0098] Take the detection box P j For example, the detection box P j The second side is projected onto the longitudinal axis T2 of the target coordinate system, and the length of the line segment obtained by the projection is determined as the length of the union When there are n detection frames in the target coordinate system, the second sides of all detection frames are projected onto the longitudinal axis T2 of the target coordinate system, and the length of the line segment obtained by the projection is determined as the eighth union length.
[0099] Then it can be further obtained that the fourth ratio can be expressed as:
[0100] In S1030, the accuracy of predicting the suspected copy segment pair is determined according to the third ratio and the fourth ratio.
[0101] Exemplarily, the product of the third ratio and the fourth ratio is determined as the accuracy rate for predicting the suspected copy segment pair:
[0102]
[0103] Finally, by combining recall and precision, we can get the f-score as an evaluation parameter.
[0104] After determining the recall rate Recall and precision rate Precision through the above embodiment, other evaluation indicators such as F-score can be further calculated. Among them:
[0105] F-score = (2Recall×Precision) / (Recall+Precision) (3)
[0106] For example, for a database, this embodiment can also obtain the F-score for each video and then average the scores of multiple infringing video pairs in the database, or first obtain the Precision and Recall of each infringing video pair, and then calculate the F-score after averaging.
[0107] It should be noted that the evaluation indicators provided in the embodiments of this specification take into account the intersection area of the detection box and the sample (the union of the projection of the side lengths of the intersection area) and the side length of the detection box / side length of the sample box. Compared with the related art, the solution provided in the embodiments of this specification can improve the perception of the accuracy of infringement positioning. Fig.12 And Table 1 to illustrate, among which, Fig.12 The solid line box represents the detection box, the dotted line box represents the sample box, and the shaded area represents the intersection.
[0108] for Fig.12 For the four cases (a) to (d), the recall rate and precision rate calculated according to the two related technologies and the solution provided in the embodiments of this specification are shown in Table 1:
[0109] Table 1
[0110]
[0111] Among them, the solution provided by the first related technology is a video-level evaluation index. Specifically, the correctly detected segments are defined as those that have one frame overlap with the actual infringing segments. The denominator of the precision rate Precision (SP) is the number of all detected segments, and the numerator is the number of correctly detected segments (such as formula (4)); the denominator of the recall rate Recall (SR) is the number of actual marked groundtruth copy segments, and the numerator is the number of correctly detected segments (such as formula (5)).
[0112]
[0113] The second related technology provides a frame-level evaluation index, which is similar to the solution provided by the first related technology, except that the segments are replaced by frames. The denominator of the specific precision Precision (SP) is the number of all detected frames (all detected frames), and the numerator is the number of correctly detected frames (correctly detected frames) (such as formula (6)); the denominator of the recall rate Recall (SR) is the number of groundtruth copy frames that are actually marked, and the numerator is the number of correctly detected frames (correctly detected frames) (such as formula (7)).
[0114]
[0115] On the other hand, the evaluation index provided in the embodiment of this specification can take into account the above-mentioned video segmentation equivalent characteristics, thereby facilitating the improvement of the robustness of the evaluation index. Fig.12It is introduced that the evaluation index provided in the embodiment of this specification takes into account the video segmentation equivalence characteristic, while the evaluation index provided in the related art does not take into account the video segmentation equivalence characteristic.
[0116] Reference again Figure 1 , compared with video 110, video frame C is inserted into video 120. In view of the video segmentation equivalence feature, it can be considered that: the video frame with timestamp tm in video 120 to the video frame with timestamp tn ( Figure 1 The entire paragraph shown) and Figure 1 The video 110 shown in FIG. 1 is a pair of copy segments; it can also be considered that: the video frame with timestamp tm in the video 120 to the video frame A ( Figure 1 part shown) and Figure 1 The video 110 shown in FIG. 1 is a copy segment pair. At the same time, it is considered that the video frame B in the video 120 and the video frame with the timestamp tn ( Figure 1 Another part shown) and Figure 1 The video 110 shown in FIG. 1 is a pair of copied segments (ie, "marked as multiple segments").
[0117] Combined with reference Fig.12 The sample frame b in the video 120 can be considered as: the video frame with time stamp tm to the video frame with time stamp tn ( Figure 1 The entire paragraph shown) and Figure 1 The video 110 shown in FIG. 1 is a pair of copied segments; Fig.12 The detection box in b can be considered as: the video frame with time stamp tm in video 120 to video frame A ( Figure 1 part shown) and Figure 1 The video 110 shown in FIG. 1 is a copy segment pair. At the same time, it is considered that the video frame B in the video 120 and the video frame with the timestamp tn ( Figure 1 Another part shown) and Figure 1 The video 110 shown in FIG. 1 is a pair of copied segments (ie, "marked as multiple segments").
[0118] In this case, the recall rate R1 and the precision rate P1 calculated by the embodiments of this specification are both 1.
[0119] Combined with reference Fig.12 The sample frame c in the video 120 can be considered as: the video frame with time stamp tm to the video frame A ( Figure 1 part shown) and Figure 1 The video 110 shown in FIG. 1 is a copy segment pair. At the same time, it is considered that the video frame B in the video 120 and the video frame with the timestamp tn ( Figure 1 Another part shown) and Figure 1 The video 110 shown in FIG. 1 is a pair of copied segments; Fig.12The detection box in c can be considered as: the video frame with time stamp tm in video 120 to the video frame with time stamp tn ( Figure 1 The entire paragraph shown) and Figure 1 The video 110 shown in FIG. 1 is a pair of copied segments (ie, “marked as a whole segment”).
[0120] In this case, the recall rate R2 and the precision rate P2 calculated by the embodiments of this specification are both 1.
[0121] In summary, for Figure 1 In the case where both "marked as a whole paragraph" and "marked as multiple paragraphs" are reasonable, the solution provided in the embodiment of this specification is used to calculate "marked as a whole paragraph (such as Fig.12 c)”, and the recall rate R2 of “annotated as multiple segments (such as Fig.12 b)” has the same recall rate R1 as “annotated as a whole paragraph (e.g. Fig.12 c)”, and the accuracy P2 of “marked as multiple segments (such as Fig.12 It can be seen that the evaluation index provided in the embodiment of this specification takes into account the equivalent characteristics of video segmentation, and therefore has higher robustness.
[0122] It should be noted that the above figures are only schematic illustrations of the processes included in the method according to the exemplary embodiments of the present specification, and are not intended to be limiting. It is easy to understand that the processes shown in the above figures do not indicate or limit the time sequence of these processes. In addition, it is also easy to understand that these processes can be performed synchronously or asynchronously, for example, in multiple modules.
[0123] The following are device embodiments of this specification, which can be used to implement the method embodiments of this specification. For details not disclosed in the device embodiments of this specification, please refer to the method embodiments of this specification.
[0124] in, Fig.13 The structural diagram of the device for determining the evaluation index provided in one embodiment of the present specification is as follows. The device for determining the evaluation index 1300 in the embodiment of the present specification includes: a coordinate system construction module 1310, a detection frame determination module 1320, a sample frame determination module 1330, and an evaluation module 1340.
[0125] Among them, the above-mentioned coordinate system construction module 1310 is used to construct a target coordinate system with the timestamp sequence of the first video as the horizontal axis coordinate and the timestamp sequence of the second video as the vertical axis coordinate, and each coordinate point in the above-mentioned target coordinate system corresponds to a video frame pair with a corresponding timestamp; the above-mentioned detection frame determination module 1320 is used to determine a detection frame in the above-mentioned target coordinate system, the above-mentioned detection frame corresponds to the suspected copy segment pair between the above-mentioned first video and the above-mentioned second video, and the predicted value of the inter-frame similarity of the video frame pair corresponding to the coordinate point in the above-mentioned detection frame is greater than the preset value; the above-mentioned sample frame determination module 1330 is used to determine a sample frame in the above-mentioned target coordinate system, and the actual value of the inter-frame similarity of the video frame pair corresponding to the coordinate point in the above-mentioned sample frame is greater than the above-mentioned preset value; and the above-mentioned evaluation module 1340 is used to determine the evaluation index of the prediction result of the above-mentioned suspected copy segment pair according to the intersection between the above-mentioned sample frame and the above-mentioned detection frame, the side length of the above-mentioned sample frame and the side length of the above-mentioned detection frame.
[0126] In an exemplary embodiment, Fig.14 The schematic diagram shows the structure of another exemplary device for determining evaluation indicators provided in this specification. Fig.14 :
[0127] In an exemplary embodiment, based on the aforementioned scheme, the evaluation index is the recall rate;
[0128] The evaluation module 1340 includes: a first intersection determination unit 13401 , a first ratio calculation unit 13402 , a second ratio calculation unit 13403 and a recall rate determination unit 13404 .
[0129] Among them, the above-mentioned first intersection determination unit 13401 is used to: obtain the intersection of each of the above-mentioned sample frames and all the detection frames in the above-mentioned target coordinate system to obtain a first group of intersections; the above-mentioned first ratio calculation unit 13402 is used to: obtain the union length of all first edges corresponding to the above-mentioned first group of intersections to obtain a first union length, obtain the union length of the first edges corresponding to all sample frames to obtain a second union length, and calculate the ratio of the first union length to the second union length to obtain a first ratio, and the above-mentioned first edge is parallel to the horizontal axis of the above-mentioned target coordinate system; the above-mentioned second ratio calculation unit 13403 is used to: obtain the union length of all second edges corresponding to the above-mentioned first group of intersections to obtain a third union length, obtain the union length of the second edges corresponding to all sample frames to obtain a fourth union length, and calculate the ratio of the third union length to the fourth union length to obtain a second ratio, and the above-mentioned second edge is parallel to the vertical axis of the above-mentioned target coordinate system; and the above-mentioned recall rate determination unit 13404 is used to: determine the recall rate of the pair of suspected copy segments predicted according to the above-mentioned first ratio and the above-mentioned second ratio.
[0130] In an exemplary embodiment, based on the above scheme, the first ratio calculation unit 13402 is specifically used to: project the first sides of all intersections in the first group of intersections onto the horizontal axis of the target coordinate system, and determine the line segment length obtained by the projection as the first union length; and specifically used to: project the first sides of all sample frames onto the horizontal axis of the target coordinate system, and determine the line segment length obtained by the projection as the second union length;
[0131] In an exemplary embodiment, based on the above-mentioned scheme, the above-mentioned second ratio calculation unit 13403 is specifically used to: project the second sides of all intersections in the above-mentioned first group of intersections onto the longitudinal axis of the above-mentioned target coordinate system, and determine the length of the line segment obtained by the projection as the above-mentioned third union length; and, is also specifically used to: project the second sides of all sample frames onto the longitudinal axis of the above-mentioned target coordinate system, and determine the length of the line segment obtained by the projection as the above-mentioned fourth union length.
[0132] In an exemplary embodiment, based on the above scheme, the recall rate determining unit 13404 is specifically used to: determine the product of the first ratio and the second ratio as the recall rate for predicting the suspected copy segment pair.
[0133] In an exemplary embodiment, based on the aforementioned scheme, the evaluation index is accuracy;
[0134] The evaluation module 1340 further includes: a second intersection determination unit 13405 , a third ratio calculation unit 13406 , a fourth ratio calculation unit 13407 and an accuracy determination unit 13408 .
[0135] The second intersection determination unit 13405 is used to obtain the intersection of each detection frame and all sample frames in the target coordinate system to obtain a second group of intersections; the third ratio calculation unit 13406 is used to obtain the union lengths of all first sides corresponding to the second group of intersections to obtain a fifth union length, obtain the union lengths of the first sides corresponding to all detection frames to obtain a sixth union length, and calculate the ratio of the fifth union length to the sixth union length to obtain a third ratio, wherein the first side is parallel to the horizontal axis of the target coordinate system; the fourth ratio calculation unit 13407 is used to obtain the union lengths of all second sides corresponding to the second group of intersections to obtain a seventh union length, obtain the union lengths of the second sides corresponding to all detection frames to obtain an eighth union length, and calculate the ratio of the seventh union length to the eighth union length to obtain a fourth ratio, wherein the second side is parallel to the vertical axis of the target coordinate system; and the accuracy determination unit 13408 is used to determine the accuracy of the prediction result of the suspected copy segment pair according to the third ratio and the fourth ratio.
[0136] In an exemplary embodiment, based on the above scheme, the third ratio calculation unit 13406 is specifically used to: project the first sides of all intersections in the second group of intersections onto the horizontal axis of the target coordinate system, and determine the length of the line segment obtained by the projection as the fifth union length; and is also specifically used to: project the first sides of all detection boxes onto the horizontal axis of the target coordinate system, and determine the length of the line segment obtained by the projection as the sixth union length;
[0137] In an exemplary embodiment, based on the above-mentioned scheme, the above-mentioned fourth ratio calculation unit 13407 is specifically used to: project the second sides of all intersections in the above-mentioned second group of intersections onto the longitudinal axis of the above-mentioned target coordinate system, and determine the length of the line segment obtained by the projection as the above-mentioned seventh union length; and, is also specifically used to: project the second sides of all detection boxes onto the longitudinal axis of the above-mentioned target coordinate system, and determine the length of the line segment obtained by the projection as the above-mentioned eighth union length.
[0138] In an exemplary embodiment, based on the above scheme, the accuracy determination unit 13408 is specifically used to: determine the product of the third ratio and the fourth ratio as the accuracy of predicting the suspected copy segment pair.
[0139] In an exemplary embodiment, based on the above scheme, the suspected copy segment pair includes a first segment in the first video and a second segment in the second video;
[0140] The above-mentioned detection frame determination module 1320 is specifically used to: determine the horizontal axis coordinate starting point and the horizontal axis coordinate end point of the above-mentioned detection frame according to the start frame timestamp and the end frame timestamp of the above-mentioned first segment; determine the vertical axis coordinate starting point and the vertical axis coordinate end point of the above-mentioned detection frame according to the start frame timestamp and the end frame timestamp of the above-mentioned second segment; and determine the above-mentioned detection frame in the above-mentioned target coordinate system according to the horizontal axis coordinate starting point and the horizontal axis coordinate end point of the above-mentioned detection frame and the vertical axis coordinate starting point and the vertical axis coordinate end point of the above-mentioned detection frame.
[0141] In an exemplary embodiment, based on the above solution, the apparatus further includes: a mapping module 1350 .
[0142] The mapping module 1350 is used to: calculate the similarity pxy between the xth video frame with a timestamp of x in the first video and the yth video frame of the second video, wherein x is taken in ascending order according to the timestamp of the first video, and y is the timestamp of the second video; and establish a mapping relationship between the similarity pxy and the coordinate point (x, y) in the target coordinate system;
[0143] The sample frame determination module 1330 is specifically used to determine the coordinate region whose similarity is greater than the preset value as the region within the sample frame to obtain the sample frame.
[0144] In an exemplary embodiment, based on the above solution, the apparatus further includes: a prediction module 1360 .
[0145] The prediction module 1360 is used to obtain a pair of suspected copy segments between the first video and the second video before determining a detection frame in the target coordinate system.
[0146] In an exemplary embodiment, based on the aforementioned scheme, the prediction module 1360 is specifically used to obtain at least one pair of segments between the first video and the second video whose similarity is greater than a preset value through a dynamic matching algorithm or a time domain neural network algorithm, and obtain the suspected copy segment pair.
[0147] It should be noted that the device for determining the evaluation index provided in the above embodiment is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0148] In addition, the device for determining the evaluation index provided in the above embodiment and the method for determining the evaluation index belong to the same concept. Therefore, for details not disclosed in the device embodiment of this specification, please refer to the embodiment of the method for determining the evaluation index described above in this specification, which will not be repeated here.
[0149] The serial numbers of the embodiments of this specification are for description only and do not represent the advantages or disadvantages of the embodiments.
[0150] The embodiments of this specification also provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of any of the above-mentioned embodiments when executing the program.
[0151] Fig.15 This is a schematic diagram of the structure of an electronic device provided in the embodiments of this specification. Fig.15 As shown, the electronic device 1500 includes: a processor 1501 and a memory 1502 .
[0152] In the embodiments of this specification, the processor 1501 is the control center of the computer system, which can be the processor of a physical machine or the processor of a virtual machine. The processor 1501 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 1501 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), and programmable logic array (PLA). The processor 1501 may also include a main processor and a coprocessor, the main processor is a processor for processing data in the wake-up state; the coprocessor is a low-power processor for processing data in the standby state.
[0153] In the embodiment of this specification, the processor 1501 may be specifically used for:
[0154] A target coordinate system is constructed with the timestamp sequence of the first video as the horizontal axis coordinate and the timestamp sequence of the second video as the vertical axis coordinate, wherein each coordinate point in the target coordinate system corresponds to a video frame pair with a corresponding timestamp; a detection frame is determined in the target coordinate system, wherein the detection frame corresponds to a suspected copy segment pair between the first video and the second video, and the predicted value of the inter-frame similarity of the video frame pair corresponding to the coordinate points in the detection frame is greater than a preset value; a sample frame is determined in the target coordinate system, wherein the actual value of the inter-frame similarity of the video frame pair corresponding to the coordinate points in the sample frame is greater than the preset value; and, based on the intersection between the sample frame and the detection frame, the side length of the sample frame and the side length of the detection frame, an evaluation index for the prediction result of the suspected copy segment pair is determined.
[0155] Furthermore, the above evaluation index is the recall rate;
[0156] The method of determining the evaluation index of the prediction result of the suspected copy segment pair according to the intersection of the sample frame and the detection frame, the side length of the sample frame and the side length of the detection frame includes: obtaining the intersection of each of the sample frame and all the detection frames in the target coordinate system to obtain a first group of intersections; obtaining the union lengths of all first edges corresponding to the first group of intersections to obtain a first union length, obtaining the union lengths of the first edges corresponding to all sample frames to obtain a second union length, and calculating the ratio of the first union length to the second union length to obtain a first ratio, wherein the first edge is parallel to the horizontal axis of the target coordinate system; obtaining the union lengths of all second edges corresponding to the first group of intersections to obtain a third union length, obtaining the union lengths of the second edges corresponding to all sample frames to obtain a fourth union length, and calculating the ratio of the third union length to the fourth union length to obtain a second ratio, wherein the second edge is parallel to the vertical axis of the target coordinate system; and determining the recall rate of predicting the suspected copy segment pair according to the first ratio and the second ratio.
[0157] Furthermore, the obtaining of the union lengths of all first edges corresponding to the first group of intersections to obtain a first union length includes: projecting the first edges of all intersections in the first group of intersections onto the horizontal axis of the target coordinate system, and determining the line segment length obtained by the projection as the first union length; the obtaining of the union lengths of the first edges corresponding to all sample frames to obtain a second union length includes: projecting the first edges of all sample frames onto the horizontal axis of the target coordinate system, and determining the line segment length obtained by the projection as the second union length; the obtaining of the union lengths of all second edges corresponding to the first group of intersections to obtain a third union length includes: projecting the second edges of all intersections in the first group of intersections onto the vertical axis of the target coordinate system, and determining the line segment length obtained by the projection as the third union length; the obtaining of the union lengths of the second edges corresponding to all sample frames to obtain a fourth union length includes: projecting the second edges of all sample frames onto the vertical axis of the target coordinate system, and determining the line segment length obtained by the projection as the fourth union length.
[0158] Further, the above determining the recall rate for predicting the suspected copy segment pair according to the first ratio and the second ratio includes: determining the product of the first ratio and the second ratio as the recall rate for predicting the suspected copy segment pair.
[0159] Furthermore, the above evaluation index is precision;
[0160] The method of determining the evaluation index of the prediction result of the suspected copy segment pair according to the intersection of the sample frame and the detection frame, the side length of the sample frame and the side length of the detection frame includes: obtaining the intersection of each detection frame and all sample frames in the target coordinate system to obtain a second group of intersections; obtaining the union lengths of all first edges corresponding to the second group of intersections to obtain a fifth union length, obtaining the union lengths of the first edges corresponding to all detection frames to obtain a sixth union length, and calculating the ratio of the fifth union length to the sixth union length to obtain a third ratio, wherein the first edge is parallel to the horizontal axis of the target coordinate system; obtaining the union lengths of all second edges corresponding to the second group of intersections to obtain a seventh union length, obtaining the union lengths of the second edges corresponding to all detection frames to obtain an eighth union length, and calculating the ratio of the seventh union length to the eighth union length to obtain a fourth ratio, wherein the second edge is parallel to the vertical axis of the target coordinate system; and determining the accuracy of the prediction result of the suspected copy segment pair according to the third ratio and the fourth ratio.
[0161] Furthermore, the obtaining of the union lengths of all first edges corresponding to the second group of intersections to obtain a fifth union length includes: projecting the first edges of all intersections in the second group of intersections onto the horizontal axis of the target coordinate system, and determining the length of the line segment obtained by the projection as the fifth union length; the obtaining of the union lengths of the first edges corresponding to all detection frames to obtain a sixth union length includes: projecting the first edges of all detection frames onto the horizontal axis of the target coordinate system, and determining the length of the line segment obtained by the projection as the sixth union length; the obtaining of the union lengths of all second edges corresponding to the second group of intersections to obtain a seventh union length includes: projecting the second edges of all intersections in the second group of intersections onto the vertical axis of the target coordinate system, and determining the length of the line segment obtained by the projection as the seventh union length; the obtaining of the union lengths of the second edges corresponding to all detection frames to obtain an eighth union length includes: projecting the second edges of all detection frames onto the vertical axis of the target coordinate system, and determining the length of the line segment obtained by the projection as the eighth union length.
[0162] Further, the determining of the accuracy of predicting the suspected copy segment pair according to the third ratio and the fourth ratio includes: determining the product of the third ratio and the fourth ratio as the accuracy of predicting the suspected copy segment pair.
[0163] Furthermore, the suspected copy segment pair includes a first segment in the first video and a second segment in the second video;
[0164] The above-mentioned determination of the detection frame in the above-mentioned target coordinate system includes: determining the horizontal axis coordinate starting point and the horizontal axis coordinate end point of the above-mentioned detection frame according to the start frame timestamp and the end frame timestamp of the above-mentioned first segment; determining the vertical axis coordinate starting point and the vertical axis coordinate end point of the above-mentioned detection frame according to the start frame timestamp and the end frame timestamp of the above-mentioned second segment; and determining the above-mentioned detection frame in the above-mentioned target coordinate system according to the horizontal axis coordinate starting point and the horizontal axis coordinate end point of the above-mentioned detection frame and the vertical axis coordinate starting point and the vertical axis coordinate end point of the above-mentioned detection frame.
[0165] Furthermore, the processor 1501 is further configured to:
[0166] Calculate the similarity pxy between the xth video frame with timestamp x in the first video and the yth video frame of the second video, where x is taken in ascending order of the timestamp of the first video, and y is the timestamp of the second video; and establish a mapping relationship between the similarity pxy and the coordinate point (x, y) in the target coordinate system;
[0167] Wherein, determining the sample frame in the target coordinate system includes: determining the coordinate region whose similarity is greater than the preset value as the region within the sample frame to obtain the sample frame.
[0168] Furthermore, the processor 1501 is further configured to:
[0169] Before determining the detection frame in the target coordinate system, a pair of suspected copy segments between the first video and the second video is obtained.
[0170] Furthermore, obtaining a suspected copy segment pair between the first video and the second video includes: obtaining at least one segment pair between the first video and the second video whose similarity is greater than a preset value through a dynamic matching algorithm or a time domain neural network algorithm to obtain the suspected copy segment pair.
[0171] The memory 1502 may include one or more computer-readable storage media, which may be non-transitory. The memory 1502 may also include a high-speed random access memory and a non-volatile memory, such as one or more disk storage devices and flash memory storage devices. In some embodiments of the present specification, the non-transitory computer-readable storage medium in the memory 1502 is used to store at least one instruction, which is used to be executed by the processor 1501 to implement the method in the embodiment of the present specification.
[0172] In some embodiments, the electronic device 1500 further includes: a peripheral device interface 1503 and at least one peripheral device. The processor 1501, the memory 1502 and the peripheral device interface 1503 may be connected via a bus or a signal line. Each peripheral device may be connected to the peripheral device interface 1503 via a bus, a signal line or a circuit board. Specifically, the peripheral device includes: at least one of a display screen 1504, a camera 1505 and an audio circuit 1506.
[0173] The peripheral device interface 1503 may be used to connect at least one peripheral device related to input / output (I / O) to the processor 1501 and the memory 1502. In some embodiments of the present specification, the processor 1501, the memory 1502, and the peripheral device interface 1503 are integrated on the same chip or circuit board; in some other embodiments of the present specification, any one or two of the processor 1501, the memory 1502, and the peripheral device interface 1503 may be implemented on a separate chip or circuit board. This embodiment of the present specification does not specifically limit this.
[0174] The display screen 1504 is used to display a user interface (UI). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 1504 is a touch display screen, the display screen 1504 also has the ability to collect touch signals on the surface or above the surface of the display screen 1504. The touch signal can be input to the processor 1501 as a control signal for processing. At this time, the display screen 1504 can also be used to provide virtual buttons and / or virtual keyboards, also known as soft buttons and / or soft keyboards. In some embodiments of the present specification, the display screen 1504 can be one, and the front panel of the electronic device 1500 is set; in other embodiments of the present specification, the display screen 1504 can be at least two, which are respectively set on different surfaces of the electronic device 1500 or are folded; in some other embodiments of the present specification, the display screen 1504 can be a flexible display screen, which is set on the curved surface or folded surface of the electronic device 1500. Even, the display screen 1504 can also be set to a non-rectangular irregular shape, that is, a special-shaped screen. The display screen 1504 can be made of materials such as a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc.
[0175] Camera 1505 is used to capture images or videos. Optionally, camera 1505 includes a front camera and a rear camera. Usually, the front camera is arranged on the front panel of the electronic device, and the rear camera is arranged on the back of the electronic device. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth of field camera, a wide-angle camera, and a telephoto camera, so as to realize the fusion of the main camera and the depth of field camera to realize the background blur function, the fusion of the main camera and the wide-angle camera to realize the panoramic shooting and the virtual reality (VR) shooting function or other fusion shooting functions. In some embodiments of the present specification, camera 1505 may also include a flash. The flash can be a monochrome temperature flash or a dual color temperature flash. Dual color temperature flash refers to a combination of warm light flash and cold light flash, which can be used for light compensation at different color temperatures.
[0176] The audio circuit 1506 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, and convert the sound waves into electrical signals and input them into the processor 1501 for processing. For the purpose of stereo sound collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the electronic device 1500. The microphone may also be an array microphone or an omnidirectional collection microphone.
[0177] The power supply 1507 is used to power various components in the electronic device 1500. The power supply 1507 can be an alternating current, a direct current, a disposable battery, or a rechargeable battery. When the power supply 1507 includes a rechargeable battery, the rechargeable battery can be a wired rechargeable battery or a wireless rechargeable battery. A wired rechargeable battery is a battery that is charged through a wired line, and a wireless rechargeable battery is a battery that is charged through a wireless coil. The rechargeable battery can also be used to support fast charging technology.
[0178] The electronic device structure block diagram shown in the embodiment of this specification does not constitute a limitation on the electronic device 1500. The electronic device 1500 may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component arrangement.
[0179] In the description of this specification, it should be understood that the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. For those of ordinary skill in the art, the specific meanings of the above terms in this specification can be understood according to specific circumstances. In addition, in the description of this specification, unless otherwise specified, "multiple" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. The character " / " generally indicates that the previously associated objects are in an "or" relationship.
[0180] It should be noted that the above description is of a specific embodiment of the present specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0181] The above is only a specific implementation of this specification, but the protection scope of this specification is not limited to this. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in this specification, which should be included in the protection scope of this specification. Therefore, equivalent changes made according to the claims of this specification are still within the scope covered by this specification.
Claims
1. A method for determining an evaluation index, wherein: The method comprises: A target coordinate system is constructed using a timestamp sequence of the first video as a horizontal coordinate and a timestamp sequence of the second video as a vertical coordinate, wherein each coordinate point in the target coordinate system corresponds to a video frame pair with a corresponding timestamp; Determining a detection frame in the target coordinate system, the detection frame corresponding to the suspected copy segment pair between the first video and the second video, and a predicted value of an inter-frame similarity of a pair of video frames corresponding to coordinate points in the detection frame being greater than a preset value; Determining a sample frame in the target coordinate system, wherein an actual value of inter-frame similarity of a pair of video frames corresponding to coordinate points in the sample frame is greater than the preset value; Determining an evaluation index for the prediction result of the suspected copy segment pair according to the intersection between the sample frame and the detection frame, the side length of the sample frame, and the side length of the detection frame; The suspected copy segment pair includes a first segment in the first video and a second segment in the second video; Determining a detection frame in the target coordinate system includes: Determine a starting point and an end point of a horizontal axis coordinate of the detection frame according to a start frame timestamp and an end frame timestamp of the first segment; Determine a starting point and an end point of a vertical axis coordinate of the detection frame according to a start frame timestamp and an end frame timestamp of the second segment; The detection frame is determined in the target coordinate system according to a starting point and an end point of the horizontal axis coordinate of the detection frame and a starting point and an end point of the vertical axis coordinate of the detection frame.
2. The method according to claim 1, wherein: The evaluation index is recall rate; The step of determining an evaluation index for the prediction result of the suspected copy segment according to the intersection of the sample frame and the detection frame, the side length of the sample frame, and the side length of the detection frame includes: Obtaining the intersection of each of the sample frames and all detection frames in the target coordinate system to obtain a first set of intersections; Obtaining the union lengths of all first sides corresponding to the first group of intersections to obtain a first union length, obtaining the union lengths of first sides corresponding to all sample frames to obtain a second union length, and calculating a ratio of the first union length to the second union length to obtain a first ratio, wherein the first side is parallel to the horizontal axis of the target coordinate system; Obtaining the union lengths of all second sides corresponding to the first group of intersections to obtain a third union length, obtaining the union lengths of the second sides corresponding to all sample frames to obtain a fourth union length, and calculating a ratio of the third union length to the fourth union length to obtain a second ratio, wherein the second side is parallel to the longitudinal axis of the target coordinate system; A recall rate for predicting the suspected copy segment pair is determined according to the first ratio and the second ratio.
3. The method according to claim 2, wherein: The obtaining of the union lengths of all first edges corresponding to the first group of intersections to obtain a first union length includes: projecting the first edges of all intersections in the first group of intersections onto the horizontal axis of the target coordinate system, and determining the line segment length obtained by the projection as the first union length; The obtaining of the union lengths of the first sides corresponding to all sample frames to obtain the second union length includes: projecting the first sides of all sample frames onto the horizontal axis of the target coordinate system, and determining the line segment length obtained by the projection as the second union length; The obtaining of the union lengths of all second sides corresponding to the first group of intersections to obtain a third union length includes: projecting the second sides of all intersections in the first group of intersections onto the longitudinal axis of the target coordinate system, and determining the line segment length obtained by the projection as the third union length; The obtaining of the union lengths of the second sides corresponding to all sample frames to obtain a fourth union length includes: projecting the second sides of all sample frames onto the longitudinal axis of the target coordinate system, and determining the line segment length obtained by the projection as the fourth union length.
4. The method according to claim 2, wherein: The step of determining a recall rate for predicting the suspected copy segment pair according to the first ratio and the second ratio includes: The product of the first ratio and the second ratio is determined as a recall rate for predicting the suspected copy segment pair.
5. The method according to claim 1, wherein: The evaluation index is precision; The step of determining an evaluation index for the prediction result of the suspected copy segment according to the intersection of the sample frame and the detection frame, the side length of the sample frame, and the side length of the detection frame includes: Obtaining the intersection of each detection frame and all sample frames in the target coordinate system to obtain a second set of intersections; Obtaining the union lengths of all first sides corresponding to the second group of intersections to obtain a fifth union length, obtaining the union lengths of the first sides corresponding to all detection boxes to obtain a sixth union length, and calculating a ratio of the fifth union length to the sixth union length to obtain a third ratio, wherein the first side is parallel to the horizontal axis of the target coordinate system; Obtaining the union lengths of all second sides corresponding to the second group of intersections to obtain a seventh union length, obtaining the union lengths of the second sides corresponding to all detection boxes to obtain an eighth union length, and calculating a ratio of the seventh union length to the eighth union length to obtain a fourth ratio, wherein the second side is parallel to the longitudinal axis of the target coordinate system; The accuracy rate of the prediction result of the suspected copy segment pair is determined according to the third ratio and the fourth ratio.
6. The method according to claim 5, wherein: The obtaining of the union lengths of all first edges corresponding to the second group of intersections to obtain the fifth union length includes: projecting the first edges of all intersections in the second group of intersections onto the horizontal axis of the target coordinate system, and determining the line segment length obtained by the projection as the fifth union length; The obtaining of the union lengths of the first sides corresponding to all the detection frames to obtain the sixth union length includes: projecting the first sides of all the detection frames onto the horizontal axis of the target coordinate system, and determining the line segment length obtained by the projection as the sixth union length; The obtaining of the union lengths of all second sides corresponding to the second group of intersections to obtain the seventh union length includes: projecting the second sides of all intersections in the second group of intersections onto the longitudinal axis of the target coordinate system, and determining the line segment length obtained by the projection as the seventh union length; The obtaining of the union lengths of the second sides corresponding to all detection frames to obtain the eighth union length includes: projecting the second sides of all detection frames onto the longitudinal axis of the target coordinate system, and determining the line segment length obtained by the projection as the eighth union length.
7. The method according to claim 5, wherein: The step of determining the accuracy of predicting the suspected copy segment pair according to the third ratio and the fourth ratio comprises: The product of the third ratio and the fourth ratio is determined as the accuracy rate for predicting the suspected copy segment pair.
8. The method according to any one of claims 1 to 7, wherein: The method further comprises: Calculate the similarity p between the xth video frame with timestamp x in the first video and the yth video frame in the second video xy , where x is taken in ascending order according to the timestamp of the first video, and y is the timestamp of the second video; The similarity p xy Establishing a mapping relationship with the coordinate point (x, y) in the target coordinate system; Determining a sample frame in the target coordinate system includes: The coordinate region whose similarity is greater than the preset value is determined as the region within the sample frame to obtain the sample frame.
9. The method according to any one of claims 1 to 7, wherein: Before determining the detection frame in the target coordinate system, the method further includes: Obtain a suspected copy segment pair between the first video and the second video.
10. The method according to claim 9, wherein: Acquiring a suspected copy segment pair between the first video and the second video, including: By using a dynamic matching algorithm or a time-domain neural network algorithm, at least one segment pair between the first video and the second video whose similarity is greater than a preset value is obtained to obtain the suspected copy segment pair.
11. A device for determining an evaluation index, wherein: The device comprises: A coordinate system construction module, used to construct a target coordinate system using a timestamp sequence of the first video as a horizontal axis coordinate and a timestamp sequence of the second video as a vertical axis coordinate, wherein each coordinate point in the target coordinate system corresponds to a video frame pair with a corresponding timestamp; a detection frame determination module, configured to determine a detection frame in the target coordinate system, wherein the detection frame corresponds to a suspected copy segment pair between the first video and the second video, and a predicted value of an inter-frame similarity of a pair of video frames corresponding to coordinate points in the detection frame is greater than a preset value; A sample frame determination module, configured to determine a sample frame in the target coordinate system, wherein an actual value of inter-frame similarity of a pair of video frames corresponding to coordinate points in the sample frame is greater than a preset value; An evaluation module, used to determine an evaluation index for the prediction result of the suspected copy segment pair according to the intersection between the sample frame and the detection frame, the side length of the sample frame and the side length of the detection frame; The suspected copy segment pair includes a first segment in the first video and a second segment in the second video; The detection frame determination module is specifically used to: Determine a starting point and an end point of a horizontal axis coordinate of the detection frame according to a start frame timestamp and an end frame timestamp of the first segment; Determine a starting point and an end point of a vertical axis coordinate of the detection frame according to a start frame timestamp and an end frame timestamp of the second segment; The detection frame is determined in the target coordinate system according to a starting point and an end point of the horizontal axis coordinate of the detection frame and a starting point and an end point of the vertical axis coordinate of the detection frame.
12. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the method for determining the evaluation index according to any one of claims 1 to 10 is implemented.
13. A computer-readable storage medium having a computer program stored thereon, wherein: When the computer program is executed by a processor, the method for determining an evaluation index according to any one of claims 1 to 10 is implemented.
Citation Information
Patent Citations
Video plagiarism detection method and device based on sliding window, equipment and medium
CN111914926A