Video fingerprint generation method, device, electronic equipment and program product

CN122783706APending Publication Date: 2026-09-18XUNLEI COMP SHENZHEN CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610940066.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-26
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

虽然该方案对压缩、噪声等具有一定鲁棒性,但仅关注整体统计信息,容易导致视频匹配准确性偏低

Benefits of technology

[0016] The advantages of this application compared to existing technologies are as follows: To overcome the problem that existing video fingerprint generation schemes only focus on overall statistical information and lack comprehensive representation of video spatial structure and temporal change information, which easily leads to insufficient video fingerprint discrimination ability and low video matching accuracy, this application, in the video fingerprint generation process, first extracts at least one target video segment from the video to be processed, and generates a video fingerprint based on the frame fingerprints corresponding to each target frame with temporal relationship in the target video segment. This allows the generated video fingerprint to represent the change pattern of video content in the temporal dimension, thereby smoothing instantaneous noise and inter-frame jitter, and improving the temporal stability of the video fingerprint. Simultaneously, when generating frame fingerprints, the target frame is divided into multiple grid units, and the frame fingerprint is determined based on the quantized feature values ​​of the visual characteristics of each grid unit. This allows the frame fingerprint to retain the regional distribution characteristics and spatial layout structure of the image, improving the ability to distinguish different video content. Since the video fingerprint is jointly constructed by the frame fingerprint with spatial structure characteristics and the target video segments with temporal relationship, a joint representation of temporal and spatial dimension features is achieved. Temporal aggregation can improve the stability of features, while spatial aggregation can improve the discriminativeness of features. This allows the generated video fingerprint to have both good temporal stability and spatial recognition capabilities, thereby more accurately representing video content and improving the accuracy and robustness of video matching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122783706A_ABST
    Figure CN122783706A_ABST
Patent Text Reader

Abstract

The application discloses a video fingerprint generation method, a video fingerprint generation device, an electronic device and a computer program product. The method generates a video fingerprint of a to-be-processed video based on frame fingerprints corresponding to each target frame in at least one target video segment. Wherein, the video fingerprint is generated based on the frame fingerprints of the target frames with time sequence, which can realize feature extraction in the time dimension, smooth instantaneous noise and frame jitter, and improve the time stability of the video fingerprint; meanwhile, when generating the frame fingerprint corresponding to the target frame, the target frame is divided into multiple grid units, and the frame fingerprint is determined based on the quantized feature values of the visual characteristics of each grid unit, which can realize feature extraction in the space dimension, retain the basic layout structure of the picture, and improve the spatial differentiation ability of the video fingerprint. Therefore, the generation method can make the video fingerprint have good stability and recognition ability at the same time, and thus improve the accuracy and robustness of video matching.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of video analysis technology, and in particular relates to a method for generating video fingerprints, a device for generating video fingerprints, an electronic device, and a computer program product. Background Technology

[0002] With the rapid growth of internet video content, scenarios such as video retrieval, copyright protection, and content regulation place higher demands on video matching technology. To achieve rapid identification of video content, it is typically necessary to first generate a video fingerprint that characterizes the features of the video content.

[0003] Among existing video fingerprinting schemes, the most common is the video fingerprinting scheme based on global statistical features. This type of scheme typically generates video fingerprints based on extracted global features. While this scheme has some robustness to compression and noise, focusing only on overall statistical information can easily lead to low video matching accuracy. Summary of the Invention

[0004] This application provides a method, apparatus, electronic device, and computer program product for generating video fingerprints, which can enable video fingerprints to have both good stability and recognition ability, thereby improving the accuracy and robustness of video matching.

[0005] Firstly, this application provides a method for generating video fingerprints, including: Based on the frame fingerprints corresponding to each target frame in at least one target video segment, the video fingerprint of the video to be processed is obtained; the target video segment is a video segment in the video to be processed; the target frame is some or all of the video frames in the target video segment. For each target frame, perform the following steps to obtain the corresponding frame fingerprint: Divide the current target frame into multiple grid cells; The frame fingerprint of the current target frame is obtained by quantizing the visual features of each grid cell.

[0006] Furthermore, the frame fingerprint of the current target frame is obtained based on the quantized feature values ​​of the visual features of each grid cell, including: Number each grid cell; The numbers of each grid cell are rearranged based on the magnitude of each quantization feature value to obtain a numbering sequence; The frame fingerprint of the current target frame is determined based on the number sequence.

[0007] Furthermore, the frame fingerprint of the current target frame is determined based on the number sequence, including: The target number is determined by selecting the number at a specified position in the numbering sequence; or, two or more numbers are selected consecutively, at intervals, or at a specified position from the numbering sequence as the target number. The target number is determined as the frame fingerprint.

[0008] Furthermore, the frame fingerprint of the current target frame is determined based on the number sequence, including: Generate corresponding unique identifier information based on the sorting results; The unique identifier is determined as a frame fingerprint.

[0009] Further, based on the frame fingerprints corresponding to each target frame in at least one target video segment, a video fingerprint of the video to be processed is obtained, including: For each target video segment, the fingerprints of each target frame in the current target video segment are concatenated in time sequence to obtain the segment fingerprint of the current target video segment. When the number of target video segments is 1, the segment fingerprint of the target video segment is determined as the video fingerprint; When the number of target video segments is greater than 1, the video fingerprint is obtained by concatenating the fingerprints of each segment of the target video segments in time sequence.

[0010] Furthermore, the visual features include at least one of brightness features, color features, and texture features.

[0011] Furthermore, after obtaining the video fingerprint, it also includes: The video fingerprint is compressed to obtain the final video fingerprint.

[0012] Secondly, this application provides a video fingerprint generation apparatus, comprising: The first generation module is used to obtain the video fingerprint of the video to be processed based on the frame fingerprint of each target frame in at least one target video segment; the target video segment is a video segment in the video to be processed; the target frame is some or all of the video frames in the target video segment. The video fingerprint generation device also includes a second generation module, which is triggered and executed before the first generation module runs. For each target frame, the second generation module is used for: Divide the current target frame into multiple grid cells; The frame fingerprint of the current target frame is obtained by quantizing the visual features of each grid cell.

[0013] Thirdly, this application provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the method described in the first aspect.

[0014] Fourthly, this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the method described in the first aspect above.

[0015] Fifthly, this application provides a computer program product comprising a computer program that, when executed by one or more processors, implements the steps of the method described in the first aspect.

[0016] The advantages of this application compared to existing technologies are as follows: To overcome the problem that existing video fingerprint generation schemes only focus on overall statistical information and lack comprehensive representation of video spatial structure and temporal change information, which easily leads to insufficient video fingerprint discrimination ability and low video matching accuracy, this application, in the video fingerprint generation process, first extracts at least one target video segment from the video to be processed, and generates a video fingerprint based on the frame fingerprints corresponding to each target frame with temporal relationship in the target video segment. This allows the generated video fingerprint to represent the change pattern of video content in the temporal dimension, thereby smoothing instantaneous noise and inter-frame jitter, and improving the temporal stability of the video fingerprint. Simultaneously, when generating frame fingerprints, the target frame is divided into multiple grid units, and the frame fingerprint is determined based on the quantized feature values ​​of the visual characteristics of each grid unit. This allows the frame fingerprint to retain the regional distribution characteristics and spatial layout structure of the image, improving the ability to distinguish different video content. Since the video fingerprint is jointly constructed by the frame fingerprint with spatial structure characteristics and the target video segments with temporal relationship, a joint representation of temporal and spatial dimension features is achieved. Temporal aggregation can improve the stability of features, while spatial aggregation can improve the discriminativeness of features. This allows the generated video fingerprint to have both good temporal stability and spatial recognition capabilities, thereby more accurately representing video content and improving the accuracy and robustness of video matching.

[0017] It is understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This is a flowchart illustrating a method for generating video fingerprints according to an embodiment of this application; Figure 2 This is a schematic flowchart illustrating the generation of a frame fingerprint corresponding to the current target frame in another video fingerprint generation method provided in this application embodiment; Figure 3 This is a schematic flowchart illustrating the generation of a frame fingerprint corresponding to the current target frame in another video fingerprint generation method provided in this application embodiment; Figure 4 This is a schematic flowchart illustrating the generation of a frame fingerprint corresponding to the current target frame in another video fingerprint generation method provided in this application embodiment; Figure 5 This is a schematic flowchart illustrating the generation of a video fingerprint in another video fingerprint generation method provided in this application embodiment; Figure 6 This is a flowchart illustrating another method for generating video fingerprints provided in an embodiment of this application; Figure 7 This is a schematic diagram of the structure of the video fingerprint generation device provided in the embodiments of this application; Figure 8 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0020] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0021] Among existing video fingerprinting schemes, the most common is the global statistical feature-based scheme. These schemes typically generate video fingerprints based on extracted global features such as color histograms, average brightness, and edge orientation histograms. While this approach is robust to compression and noise, it focuses solely on overall statistical information and lacks descriptions of the video's spatial structure and temporal variations. Therefore, it is prone to mismatches between videos with similar content but different layouts, resulting in low video matching accuracy.

[0022] To address this issue, this application proposes a method for generating video fingerprints, which enables video fingerprints to possess both good stability and recognition capabilities, thereby improving the accuracy and robustness of video matching. The control method proposed in this application will be described below through specific embodiments.

[0023] The video fingerprint generation method provided in this application embodiment can be applied to electronic devices such as mobile phones, tablets, in-vehicle devices, augmented reality (AR) / virtual reality (VR) devices, laptops, ultra-mobile personal computers (UMPCs), netbooks, and personal digital assistants (PDAs). This application embodiment does not impose any restrictions on the specific type of electronic device.

[0024] To illustrate the technical solutions proposed in this application, the following description will use an electronic device as the execution subject to illustrate various embodiments.

[0025] Figure 1 A schematic flowchart illustrating a video fingerprint generation method provided in an embodiment of this application is shown. The video fingerprint generation method includes: Step 110: For each target frame, the electronic device divides the current target frame into multiple grid cells; and obtains the frame fingerprint of the current target frame based on the quantized feature value of the visual features of each grid cell.

[0026] Step 120: The electronic device obtains the video fingerprint of the video to be processed based on the frame fingerprint corresponding to each target frame in at least one target video segment.

[0027] To accurately generate video fingerprints that characterize the temporal and spatial features of a video, at least one target video segment can be obtained from the video to be processed as the basis for generating the video fingerprint.

[0028] The target video segment can be obtained by dividing the video to be processed, either uniformly or non-uniformly. For example, in the case of uniform division, a fixed duration T can be used as the time window, such as 100 milliseconds, dividing the timeline of the video to be processed into multiple consecutive and non-overlapping time windows. For the k-th time window, video frames with presentation time stamps (PTS) located in the interval [k×T, (k+1)×T) can be collected, and these video frames within that time window are taken as the corresponding target video segment.

[0029] For each target video segment, some or all of its video frames can be used as target frames, and a segment fingerprint corresponding to the target video segment can be generated based on the frame fingerprints corresponding to each target frame. Since there is a natural temporal relationship between the target frames, using multiple target frames from the same target video segment to jointly construct the video fingerprint not only allows the generated video fingerprint to reflect not only the content features of a single frame but also the changing patterns of the video content over time, thus endowing the video fingerprint with temporal characteristics. Even under conditions of slight video editing, inter-frame jitter, or transient noise interference, the relative stability of the video fingerprint can still be maintained, which is beneficial for smoothing transient noise and inter-frame fluctuations, improving the stability and robustness of the video fingerprint in the temporal dimension.

[0030] In addition, in order to preserve the spatial structure information of the video frame, when generating the frame fingerprint corresponding to the target frame, the electronic device does not only extract the overall statistical features of the target frame, but also divides the target frame into grids to divide the target frame into multiple grid units.

[0031] Specifically, the electronic device can employ either a uniform or non-uniform grid partitioning method. The grid partitioning method can be the same or different for different target frames. When using a uniform grid partitioning method, the electronic device can divide the target frame into multiple grid units of the same size, such as 2×2, 3×3, 4×4, and 8×8 grid regions. When using a non-uniform grid partitioning method, the size of different grid units can be different, for example, the size and position of each grid unit can be determined based on the distribution of image content, the importance of the region, or a preset partitioning rule.

[0032] For example, electronic devices can also dynamically determine the grid division method and the number of grid units based on the image resolution, image complexity, scene type, or preset strategy of the target frame, thereby improving the adaptability to different video content. By dividing the target frame into multiple grid units, electronic devices can decompose the overall image into multiple local regions, providing a basis for subsequently extracting the visual features corresponding to each region and constructing a frame fingerprint.

[0033] Based on this, after the grid division is completed, the electronic device can obtain the quantized feature values ​​of the visual characteristics corresponding to each grid cell. Specifically, the quantized feature values ​​can be at least one of brightness features, color features, and texture features. Among them, brightness features can include brightness mean, brightness median, and brightness variance, etc.; color features can include dominant hue, color histogram, average color value, and color distribution, etc.; texture features can include gray-level co-occurrence matrix contrast, energy, entropy, local binary mode (LBP) features, and edge density, etc.

[0034] After obtaining the quantized feature values ​​corresponding to each grid cell, the electronic device can determine the frame fingerprint corresponding to the current target frame based on each quantized feature value. For example, the frame fingerprint can be generated by directly combining, encoding, or concatenating the quantized feature values; the corresponding sorting result can be determined according to the size relationship between the quantized feature values, and the frame fingerprint can be generated based on the sorting result; other methods that can characterize the spatial distribution relationship of each grid cell can also be used to generate the frame fingerprint.

[0035] Compared to the traditional method of representing video content using high-dimensional feature vectors, this embodiment generates frame fingerprints by utilizing quantized feature values ​​of visual characteristics corresponding to grid cells. This achieves effective representation of video content without retaining a large amount of continuous feature data, thereby reducing the data dimensionality and storage overhead of the frame fingerprint. Furthermore, the frame fingerprint structure generated based on quantized feature values ​​is simpler, which helps improve the generation efficiency, storage efficiency, and subsequent matching calculation efficiency of the frame fingerprint.

[0036] Since different grid cells correspond to different spatial regions within the target frame, feature extraction from multiple grid cells separately preserves the relative distribution and layout relationships between regions within the target frame. This allows the generated frame fingerprint to contain not only visual content information but also spatial structural information of the image. Compared to schemes that only use global statistical features, this method can more accurately characterize the layout features and regional differences of the video image, improving the spatial discriminative power of the frame fingerprint.

[0037] In this embodiment, the electronic device performs temporal feature aggregation using target video segments and spatial feature aggregation using grid cells, integrating both temporal and spatial features into the video fingerprint construction process. Temporal aggregation smooths out transient noise and inter-frame jitter, improving the temporal stability of the features; spatial aggregation preserves the basic layout structure of the image, enhancing the spatial distinguishability of the features. The combination of these two features results in a generated video fingerprint with both good stability and recognition capability, thus more accurately representing the video content and improving the accuracy and robustness of subsequent video matching.

[0038] In some embodiments, after obtaining the quantized feature values ​​corresponding to each grid cell, a frame fingerprint corresponding to the current target frame can be further generated based on the relative relationship between the quantized feature values. Unlike directly using the quantized feature values ​​corresponding to each grid cell as the frame fingerprint, the relative magnitude regularity is generally less affected by changes in overall brightness, contrast adjustment, and partial illumination changes, thus improving the stability and robustness of the frame fingerprint. Therefore, this embodiment focuses more on the relative feature distribution between each grid cell, that is, using the magnitude regularity between the quantized feature values ​​corresponding to each grid cell to characterize the image content, rather than relying on the absolute values ​​of the quantized feature values ​​themselves.

[0039] Based on this, see Figure 2 , Figure 2 This paper illustrates a schematic flowchart of another video fingerprint generation method provided in this application, which generates a frame fingerprint corresponding to the current target frame, including: Step 210: The electronic device numbers each grid cell.

[0040] After the target frame is divided into grids, each grid cell can be uniquely identified according to preset rules. For example, each grid cell can be assigned a number from 1 to N in order from left to right and from top to bottom, where N is the total number of grid cells. When the target frame is divided into a 3×3 grid, the cells can be assigned numbers from 1 to 9. By assigning unique numbers to each grid cell, a mapping relationship between the spatial location of the grid cell and its corresponding identifier can be established, providing a basis for subsequent sorting based on quantized feature values.

[0041] Step 220: The electronic device rearranges the numbers of each grid cell based on the magnitude of each quantization feature value to obtain a numbering sequence.

[0042] After numbering the grid cells, the electronic device can sort the grid cells according to the magnitude of their corresponding quantized feature values ​​and rearrange their numbers accordingly. For example, when the quantized feature value is the average brightness, the cells can be sorted from smallest to largest or largest to smallest. When the quantized feature value is the texture feature or color feature, the cells can also be sorted according to the magnitude of their corresponding feature values. Assuming the quantized feature values ​​of the grid cells numbered 1 to 9 are sorted as 3, 7, 1, 5, 9, 2, 8, 4, and 6, the resulting number sequence is 371592846.

[0043] This numbering sequence reflects the relative distribution relationship between the quantized feature values ​​corresponding to each grid cell, rather than the quantized feature values ​​themselves. Therefore, it can reduce the impact of global image transformations such as overall brightness adjustment, contrast adjustment, and local illumination changes on frame fingerprints. When a video undergoes only brightness enhancement, brightness reduction, contrast adjustment, or similar image enhancement processing, although the absolute values ​​of the quantized feature values ​​corresponding to each grid cell may change, their relative magnitude relationships usually remain unchanged or essentially unchanged. Therefore, the generated numbering sequence can still maintain high stability. Based on this, the robustness of video fingerprints to the aforementioned video editing operations can be improved, and the recognition accuracy of identical or similar videos can be increased, thus facilitating the detection of video content that has been simply processed and then re-uploaded.

[0044] Step 230: The electronic device determines the frame fingerprint of the current target frame based on the number sequence.

[0045] Specifically, electronic devices can directly use the number sequence as the frame fingerprint corresponding to the current target frame; for example, when the target frame is divided into a 3×3 grid, the number sequence corresponding to a permutation and combination of 9 numbers can be used as the frame fingerprint.

[0046] Compared to traditional high-dimensional feature vector schemes, this embodiment uses quantized feature values ​​to generate frame fingerprints and further utilizes the relative relationships between quantized feature values ​​to construct fingerprint representations. While retaining the main features of the video content, it can significantly reduce feature dimensions and data scale. This not only improves the storage and transmission efficiency of video fingerprints but also reduces the computational complexity in the video matching process, thereby improving the matching efficiency and engineering deployment feasibility in large-scale video library scenarios.

[0047] To further simplify frame fingerprints, encoding, compression, mapping, or hashing can be performed on the numbering sequence, and the processing result can be used as the frame fingerprint corresponding to the current target frame. For example, the numbering sequence can be mapped to the corresponding permutation index value, encoding value, or compression code. Since frame fingerprints originate from the relative ordering relationship of the quantized feature values ​​of each grid cell, they can reduce the dependence on the absolute values ​​of quantized feature values ​​while preserving the spatial structure features of the image, thereby improving the adaptability of frame fingerprints to changes in brightness, contrast, and some image enhancement operations.

[0048] In this embodiment, the electronic device generates a corresponding frame fingerprint by numbering grid cells and rearranging the grid cell numbers using the magnitude relationship between the corresponding quantized feature values ​​of each grid cell. On the one hand, the electronic device constructs a fingerprint representation by using the relative relationship between quantized feature values ​​rather than the quantized feature values ​​themselves. This significantly reduces the feature dimension and data size while preserving the main spatial structure features of the video content, thereby reducing the storage overhead and matching computation of the video fingerprint and improving the generation and matching efficiency of the video fingerprint. On the other hand, the electronic device characterizes the image content by using the magnitude regularity between the corresponding quantized feature values ​​of each grid cell. This reduces the impact of factors such as overall brightness changes, contrast adjustments, and local illumination changes on the frame fingerprint, making the generated frame fingerprint more stable and robust, thereby improving the recognition accuracy of the same or similar videos.

[0049] In some embodiments, to reduce the storage overhead and matching computation of video fingerprints, see [reference]. Figure 3 , Figure 3 This paper illustrates a schematic flowchart of another video fingerprint generation method provided in this application, which generates a frame fingerprint corresponding to the current target frame, including: Step 310: The electronic device determines the number at a specified position in the number sequence as the target number; or, selects two or more numbers consecutively, at intervals, or at a specified position from the number sequence as the target number.

[0050] After obtaining the number sequence, in order to further reduce the data length of the frame fingerprint, the electronic device does not need to use the complete number sequence as the frame fingerprint, but extracts a portion of the number from the number sequence as the target number.

[0051] For example, an electronic device can determine the target number as the number at a specified position in a number sequence. For instance, when the number sequence is 371592846, the first number 3, the last number 6, and the middle number 9 can be selected, or numbers at other positions can be selected as the target number according to a preset rule.

[0052] For example, the electronic device can also select multiple numbers from the numbering sequence as target numbers. This can be done by continuous selection, such as selecting 371 or 592 consecutively from the numbering sequence; by interval selection, such as selecting 31946 at a preset interval; or by selection at a specified position, such as selecting the numbers corresponding to the 1st, 4th, and 8th positions to obtain 354. Of course, those skilled in the art can use other selection rules according to actual needs, and this is not limited.

[0053] Since different positions in the numbering sequence correspond to the relative positions of different grid cells in the sorting result, whether a single number or multiple numbers are used as the target number, the feature distribution information contained in the original numbering sequence can be preserved to a certain extent. That is, since the preserved target number still originates from the relative sorting relationship between the quantized feature values ​​of each grid cell, rather than the quantized feature values ​​themselves, it can inherit, to some extent, the robustness of sorting encoding to image transformations such as overall brightness changes, contrast adjustments, and local illumination changes, thereby improving the stability of frame fingerprints and the accuracy of video recognition.

[0054] At the same time, compared to using the complete number sequence directly, the target number has a shorter data length, which helps to reduce the storage overhead and matching calculation of subsequent video fingerprints.

[0055] Step 320: The electronic device determines the target number as a frame fingerprint.

[0056] After determining the target number, the electronic device can directly use the target number as the frame fingerprint corresponding to the current target frame. For example, when the target number is a single number, it can be used as the frame fingerprint; when the target number consists of multiple numbers, the combination of the multiple numbers in the selection order can be determined as the frame fingerprint.

[0057] For example, the electronic device may further process the target number before determining it as a frame fingerprint. For instance, the target number may be encoded, compressed, mapped, or hashed, and the result may be used as a frame fingerprint, without limitation.

[0058] In this embodiment, the electronic device constructs a frame fingerprint by retaining a portion of the serial numbers in the sequence. This preserves some of the relative distribution relationships between the quantized feature values ​​of each grid cell while further reducing the amount of fingerprint data, storage overhead, and matching computation complexity. Furthermore, since the target serial number represents the relative ordering of the quantized feature values, compared to directly using quantized feature values ​​to construct the fingerprint, it reduces the impact of factors such as overall brightness adjustment, contrast adjustment, and local illumination changes on the frame fingerprint. This ensures that the compressed frame fingerprint still possesses good stability and robustness, thus balancing video matching efficiency and accuracy.

[0059] In some embodiments, to reduce the storage overhead and matching computation of video fingerprints, see [reference]. Figure 4 , Figure 4 This paper illustrates a schematic flowchart of another video fingerprint generation method provided in this application, which generates a frame fingerprint corresponding to the current target frame, including: Step 410: The electronic device generates corresponding unique identification information based on the sorting results.

[0060] After obtaining the number sequence, since the number sequence essentially corresponds to the sorting result of the quantized feature values ​​of each grid cell, different sorting results can reflect different distributions of image features. Based on this, the electronic device can map the sorting result to the corresponding unique identification information.

[0061] For example, the unique identifier can be a permutation index value. For instance, when the target frame is divided into 9 grid cells, the sorting result corresponds to a permutation of 1 to 9, and there are a total of 9! possible permutations. Therefore, a unique index value can be assigned to each permutation according to a preset permutation order, and the current sorting result can be mapped to the corresponding permutation index value as the unique identifier.

[0062] For example, the unique identification information can also be an encoded value, a compressed code, a mapping code, a numeric identifier, a string identifier, or other information that can uniquely represent the sorting result. For instance, the corresponding unique identification information can be generated through number system conversion, lookup table mapping, encoding compression, or hash mapping. This application does not limit this.

[0063] Because there is a one-to-one correspondence between unique identifiers and sorting results, unique identifiers can retain the characteristic distribution information represented by the sorting results. Furthermore, compared to directly storing the complete sequence of numbers, unique identifiers typically have a shorter data length.

[0064] Step 420: The electronic device determines the unique identification information as a frame fingerprint.

[0065] After generating unique identification information, the electronic device can directly use this unique identification information as the frame fingerprint corresponding to the current target frame for storage, transmission, or subsequent matching processing.

[0066] Once the number sequence 371592846 is mapped to the corresponding permutation index value, the permutation index value can be directly used as the frame fingerprint of the current target frame; alternatively, when the number sequence is encoded and compressed to obtain the corresponding encoded value, the encoded value can also be used as the frame fingerprint of the current target frame.

[0067] In this embodiment, the electronic device further maps the sorting results to unique identification information. On the one hand, this preserves the feature distribution relationship corresponding to the sorting results and the stability and robustness of the sorting encoding, so that the frame fingerprint can still resist the influence of factors such as overall brightness changes, contrast adjustments, and local illumination changes to a certain extent. On the other hand, it can further compress the data length of the frame fingerprint, reduce the storage overhead, transmission overhead, and matching calculation complexity of the video fingerprint, thereby improving the video retrieval efficiency and video matching efficiency in large-scale video library scenarios.

[0068] In some embodiments, see Figure 5 , Figure 5 This paper illustrates a schematic flowchart of another video fingerprint generation method provided in this application, which includes: Step 510: For each target video segment, the electronic device concatenates the frame fingerprints corresponding to each target frame in the current target video segment in time sequence to obtain the segment fingerprint corresponding to the current target video segment.

[0069] After obtaining the frame fingerprints corresponding to each target frame in the target video segment, the electronic device can combine the corresponding frame fingerprints according to the temporal order of each target frame in the target video segment. For example, when a target video segment contains target frames F1, F2, F3, and F4, and the corresponding frame fingerprints are FP1, FP2, FP3, and FP4 respectively, FP1, FP2, FP3, and FP4 can be concatenated in the temporal order of F1→F2→F3→F4 to obtain the corresponding segment fingerprint.

[0070] It should be understood that timing concatenation can be achieved not only by simple splicing, but also by methods that preserve timing relationships, such as encoding combination, compression combination, and mapping combination. This application does not limit this.

[0071] By combining multiple frame fingerprints from the same target video segment in chronological order, the video content change process within the target video segment can be preserved. This allows the obtained segment fingerprint to not only reflect the spatial structural features of a single frame but also the video change patterns within a local time range, thereby improving the temporal representation capability of the segment fingerprint.

[0072] Step 520: When the number of target video segments is 1, the electronic device determines the segment fingerprint of the target video segment as the video fingerprint.

[0073] When the video to be processed corresponds to only one target video segment, there is no need to perform further cross-segment fusion processing. The segment fingerprint corresponding to the target video segment can be directly used as the video fingerprint corresponding to the video to be processed.

[0074] For example, when the video to be processed is divided into only one target video segment, the segment fingerprint corresponding to that target video segment can directly characterize the main features of the entire video to be processed, and therefore can be directly used as the final video fingerprint.

[0075] Step 530: When the number of target video segments is greater than 1, the electronic device obtains the video fingerprint based on the segment fingerprints of each target video segment in time sequence.

[0076] When the video to be processed is divided into multiple target video segments, the electronic device can combine the corresponding segment fingerprints according to the chronological order of each target video segment in the video to be processed. For example, when the video to be processed contains target video segments S1, S2, and S3, and the corresponding segment fingerprints are SP1, SP2, and SP3 respectively, SP1, SP2, and SP3 can be concatenated in the chronological order of S1→S2→S3 to generate the final video fingerprint.

[0077] Similarly, in addition to direct concatenation, the combination of fragment fingerprints can also be achieved using compression coding, mapping coding, or other methods that can preserve the temporal relationship of the fragments.

[0078] In this embodiment, the electronic device first concatenates frame fingerprints within the same target video segment in a temporal sequence to form a segment fingerprint, and then concatenates the segment fingerprints corresponding to multiple target video segments in a temporal sequence to form a video fingerprint, thus constructing a hierarchical fingerprint structure of frame fingerprint—segment fingerprint—video fingerprint. On the one hand, segment fingerprints can preserve the video change features within a local time range; on the other hand, video fingerprints can further preserve the overall temporal relationship between different target video segments. Therefore, the generated video fingerprint not only contains the spatial structure information of the video frame, but also contains the change information of the video content at different time scales, thereby improving the video fingerprint's ability to represent video content and the accuracy and robustness of subsequent video matching.

[0079] In some embodiments, after obtaining the video fingerprint, the method further includes: performing data compression processing on the video fingerprint to obtain the final video fingerprint.

[0080] Since video fingerprints are constructed based on segment fingerprints corresponding to multiple target video segments, and segment fingerprints are constructed based on frame fingerprints corresponding to multiple target frames, the data length of the video fingerprint will increase accordingly as the number of target video segments or target frames increases. To further reduce the storage overhead, transmission overhead, and matching computation load of video fingerprints, the generated video fingerprints can be compressed to obtain the final video fingerprint.

[0081] For example, data compression processing may include encoding compression, number system conversion, mapping compression, indexing, feature aggregation, or other processing methods that can reduce data length. For instance, multiple consecutive frame fingerprints or segment fingerprints can be combined and mapped to corresponding encoded values; or, the video fingerprint can be converted into a shorter string of numbers, character sequences, or binary code streams.

[0082] In this embodiment, by performing data compression on the video fingerprint, the electronic device can further reduce the data size of the video fingerprint while preserving as much of the spatiotemporal feature information represented by the video fingerprint as possible. This reduces storage resource consumption and network transmission burden, and also reduces the computational complexity in the subsequent video matching process, thereby improving the video retrieval efficiency and video matching efficiency in large-scale video library scenarios.

[0083] In some embodiments, in conjunction with the above embodiments, see [reference] Figure 6 Another method for generating video fingerprints can be provided, including: Step 610: The electronic device acquires the video to be processed and divides it into at least one target video segment.

[0084] The electronic device acquires the video to be processed and divides the video into time segments according to a preset time window to obtain at least one target video segment.

[0085] Specifically, a fixed duration T can be used as the time window, such as 100 milliseconds, 200 milliseconds, 500 milliseconds, or 1 second, to divide the timeline of the video to be processed into multiple consecutive and non-overlapping time windows. For the k-th time window, video frames with display timestamps (PTS) located in the interval [k×T, (k+1)×T) are collected, and the corresponding set of video frames is used as the k-th target video segment.

[0086] It should be understood that the fixed time window used in this embodiment is only an example. In other embodiments, dynamic time windows, adaptive time windows, or key event triggering time windows can also be used to divide the target video segments.

[0087] Step 620: The electronic device acquires the target frames in each target video segment and performs grid division.

[0088] For each target video segment, some or all of its video frames are obtained as target frames, and each target frame is divided into a grid.

[0089] Specifically, for each target frame, the image can be divided into multiple grid units using a uniform grid division method. For example, it can be divided into grid areas of 2×2, 3×3, 4×4, 8×8, etc.

[0090] Taking a 3×3 grid as an example, the target frame can be divided into 9 grid cells of the same size, and numbered from 1 to 9 in order from left to right and from top to bottom.

[0091] Step 630: For each target frame, the electronic device determines the quantization feature value corresponding to each grid cell in the target frame.

[0092] For each grid cell in each target video segment, the electronic device can calculate the corresponding visual characteristic quantization feature value. In this embodiment, the quantization feature value uses the brightness feature.

[0093] Specifically, for the i-th grid cell in the k-th target video segment, firstly, for each target frame in the target video segment, calculate the average brightness of all pixels in the corresponding grid cell; then, further average the average brightness of all target frames in the target video segment to obtain the brightness statistics value L(k, i) of the i-th grid cell corresponding to the k-th target video segment.

[0094] Thus, multiple brightness statistics values ​​for the corresponding target video segment can be obtained: L(k,1), L(k,2), ..., L(k,9).

[0095] It should be understood that in other embodiments, the quantization feature value may also be the luminance median, luminance variance, dominant hue, color distribution, texture features, or other features that can characterize the visual content of the grid cell, and this embodiment does not limit this.

[0096] Step 640: For each target frame, the electronic device generates a frame fingerprint based on its corresponding quantized feature value.

[0097] After obtaining the quantization feature values ​​corresponding to each grid cell, the electronic device can generate the corresponding frame fingerprint based on the relative relationship between the quantization feature values.

[0098] Specifically, the electronic device can first pre-number each grid cell, and then sort the numbers according to the magnitude relationship of the corresponding quantization feature values ​​of each grid cell. For example, the electronic device sorts nine brightness statistics in ascending order, and then rearranges the corresponding grid cell numbers according to the sorting result.

[0099] Assuming the grid cell numbers corresponding to the sorting results are 5, 2, 1, 8, 4, 9, 3, 6, 7, then the numbering sequence is: P(k) = [5, 2, 1, 8, 4, 9, 3, 6, 7].

[0100] Subsequently, the electronic device determines the number sequence as the frame fingerprint of the corresponding target video segment; or further maps the number sequence to the corresponding permutation index value, encoding value or other unique identification information, and uses it as the frame fingerprint.

[0101] Step 650: For each target video segment, the electronic device generates a segment fingerprint corresponding to the target video segment based on the frame fingerprints corresponding to each frame of the target video segment.

[0102] For each target video segment, the electronic device can combine the frame fingerprints corresponding to each target frame according to the time sequence between the target frames.

[0103] For example, if a target video segment contains multiple target frames, and their corresponding frame fingerprints are FP1, FP2, FP3, and FP4, the electronic device can concatenate them in chronological order: SP = FP1 + FP2 + FP3 + FP4. This yields the segment fingerprint SP of the corresponding target video segment.

[0104] Step 660: The electronic device generates a video fingerprint based on at least one fragment fingerprint.

[0105] When the video to be processed contains multiple target video segments, the electronic device can combine the fingerprints of each segment according to the time sequence of each target video segment in the video to be processed.

[0106] For example: video clip 1 corresponds to clip fingerprint SP1; video clip 2 corresponds to clip fingerprint SP2; video clip 3 corresponds to clip fingerprint SP3; thus, the electronic device can generate a video fingerprint: VF=SP1+SP2+SP3.

[0107] For example, the electronic device may further encode and compress the video fingerprint, index it, or perform a number system conversion to obtain the final video fingerprint.

[0108] In this embodiment, the electronic device divides the video to be processed into target video segments and constructs a video fingerprint based on multiple target frames within each segment. This allows the generated video fingerprint to characterize the temporal variation patterns of the video content. Simultaneously, the electronic device preserves the spatial structure information of the video frame by dividing the target frames into grids and extracting the quantized feature values ​​corresponding to each grid cell. Furthermore, the electronic device utilizes the relative ordering relationships between the quantized feature values ​​of each grid cell to construct a frame fingerprint. This enables the video fingerprint to adapt well to global transformations such as overall brightness adjustment, contrast adjustment, and some color changes, and to resist the effects of transient noise, inter-frame jitter, and local interference to a certain extent, thereby improving the stability and robustness of the video fingerprint. Moreover, compared to traditional high-dimensional feature vector schemes, this embodiment uses quantized feature values ​​and their ordering relationships to generate video fingerprints. While preserving the main spatiotemporal features of the video, it significantly reduces the feature dimension and data size, decreasing storage overhead and computational complexity during generation. The resulting video fingerprint possesses good temporal stability, spatial discriminative power, and data compactness, enabling more accurate and efficient characterization of video content.

[0109] In some embodiments, after generating the final video fingerprint, the electronic device can store the final video fingerprint in a video fingerprint database and establish an association between the video fingerprint and the corresponding video identifier (Video ID) for subsequent video retrieval and video matching.

[0110] When a new video to be detected is received, the electronic device (or other electronic devices that can use the video fingerprint database) can generate a video fingerprint corresponding to the video to be detected using the method described in the above embodiments, and match the generated video fingerprint with the historical video fingerprints in the video fingerprint database.

[0111] It should be understood that the video to be tested can be the original video or a video that has been edited and re-uploaded. For example, the video to be tested may have been processed by cropping, adjusting brightness, adjusting contrast, adding watermarks, compressing encoding, or adding noise.

[0112] Specifically, electronic devices can determine whether a video matches a historical video by calculating the similarity between the video fingerprint corresponding to the video to be detected. For example, Hamming distance, edit distance, cosine similarity, or a fast retrieval method based on an index structure can be used for matching.

[0113] For example, electronic devices can also build inverted indexes, hash indexes, or other index structures to improve video retrieval efficiency in scenarios with large-scale video libraries.

[0114] Since the video fingerprint generated in this application is constructed by an electronic device based on multiple target frames in the target video segment and the quantized feature values ​​corresponding to multiple grid units in each target frame, and further utilizes the relative ordering relationship between the quantized feature values ​​of each grid unit to characterize the video content, it has good adaptability to global image transformations such as overall brightness adjustment and contrast adjustment. At the same time, through temporal window aggregation and spatial grid aggregation, the electronic device can also reduce the impact of factors such as small-scale image shifts, local noise interference, and watermark coverage on the video fingerprint to a certain extent, thereby improving the stability and robustness of the video fingerprint.

[0115] When an electronic device determines that the similarity between the video fingerprint of the video to be detected and a historical video fingerprint in the database reaches a preset threshold, it can be judged that the video to be detected and the corresponding historical video have a high content similarity. At this time, the electronic device can further output the matching result or trigger the corresponding subsequent processing flow. For example, it can perform duplicate video detection, video content association analysis, copyright risk warning, infringing content identification, content review, or other business processing operations.

[0116] The above methods enable rapid retrieval and efficient matching of massive video data using video fingerprints, reducing storage costs and computational overhead while improving processing efficiency and recognition accuracy in scenarios such as duplicate video recognition, similar video discovery, and video copyright protection.

[0117] Corresponding to the video fingerprint generation method in the above embodiments, Figure 7 A structural block diagram of the video fingerprint generation apparatus 7 provided in an embodiment of this application is shown. For ease of explanation, only the parts related to the embodiment of this application are shown.

[0118] Reference Figure 7 The video fingerprint generation device 7 includes: The first generation module 71 is used to obtain the video fingerprint of the video to be processed based on the frame fingerprint of each target frame in at least one target video segment; the target video segment is a video segment in the video to be processed; the target frame is some or all of the video frames in the target video segment. The video fingerprint generation device also includes a second generation module, which is triggered and executed before the first generation module runs. For each target frame, the second generation module 72 is used for: Divide the current target frame into multiple grid cells; The frame fingerprint of the current target frame is obtained by quantizing the visual features of each grid cell.

[0119] Optionally, the second generation module is specifically used for: Number each grid cell; The numbers of each grid cell are rearranged based on the magnitude of each quantization feature value to obtain a numbering sequence; The frame fingerprint of the current target frame is determined based on the number sequence.

[0120] Optionally, the second generation module is specifically used for: The target number is determined by selecting the number at a specified position in the numbering sequence; or, two or more numbers are selected consecutively, at intervals, or at a specified position from the numbering sequence as the target number. The target number is determined as the frame fingerprint.

[0121] Optionally, the second generation module is specifically used for: Generate corresponding unique identifier information based on the sorting results; The unique identifier is determined as a frame fingerprint.

[0122] Optionally, the first generation module is specifically used for: For each target video segment, the fingerprints of each target frame in the current target video segment are concatenated in time sequence to obtain the segment fingerprint of the current target video segment. When the number of target video segments is 1, the segment fingerprint of the target video segment is determined as the video fingerprint; When the number of target video segments is greater than 1, the video fingerprint is obtained by concatenating the fingerprints of each segment of the target video segments in time sequence.

[0123] Optionally, the visual features include at least one of brightness features, color features, and texture features.

[0124] Optionally, the first generation module is specifically used for: After obtaining the video fingerprint, the video fingerprint is compressed to obtain the final video fingerprint.

[0125] It should be noted that the information interaction and execution process between the above-mentioned devices / units are based on the same concept as the method embodiments of this application. For details on their specific functions and technical effects, please refer to the method embodiments section, and they will not be repeated here.

[0126] Figure 8 This is a schematic diagram of the physical layer structure of an electronic device provided in an embodiment of this application. Figure 8 As shown, the electronic device 8 of this embodiment includes: at least one processor 80 ( Figure 8 The diagram shows only one processor, memory 81, and a computer program 82 stored in memory 81 and executable on at least one processor 80. When processor 80 executes computer program 82, it implements the steps in any of the above-described video fingerprint generation method embodiments, for example... Figure 1Steps 110-120 are shown.

[0127] The processor 80 can be a Central Processing Unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor.

[0128] In some embodiments, memory 81 may be an internal storage unit of electronic device 8, such as a hard disk or memory of electronic device 8. In other embodiments, memory 81 may also be an external storage device of electronic device 8, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc. equipped on electronic device 8.

[0129] Furthermore, the memory 81 may include both internal storage units and external storage devices of the electronic device 8. The memory 81 is used to store operating devices, application programs, bootloaders, data, and other programs, such as program code for computer programs. The memory 81 can also be used to temporarily store data that has been output or will be output.

[0130] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the above device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0131] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps described in the various method embodiments above.

[0132] This application provides a computer program product that, when run on an electronic device, enables the electronic device to perform the steps described in the various method embodiments above.

[0133] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include at least: any entity or device capable of carrying the computer program code to a photographic device / electronic device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, such as a USB flash drive, a portable hard drive, a magnetic disk, or an optical disk.

[0134] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0135] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0136] In the embodiments provided in this application, it should be understood that the disclosed apparatus / network devices and methods can be implemented in other ways. For example, the apparatus / network device embodiments described above are merely illustrative. For instance, the division of modules or units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.

[0137] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0138] The above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.

Claims

1. A method for generating video fingerprints, characterized in that, include: Based on the frame fingerprints corresponding to each target frame in at least one target video segment, the video fingerprint of the video to be processed is obtained. The target video segment is a video segment in the video to be processed; the target frame is some or all of the video frames in the target video segment. For each target frame, perform the following steps to obtain the corresponding frame fingerprint: Divide the current target frame into multiple grid cells; The frame fingerprint of the current target frame is obtained by quantizing the visual features of each grid cell.

2. The generation method as described in claim 1, characterized in that, The step of obtaining the frame fingerprint of the current target frame based on the quantized feature values ​​of the visual features of each grid cell includes: Each of the aforementioned grid cells is numbered; The numbers of each grid cell are rearranged based on the magnitude of each quantized feature value to obtain a numbering sequence; The frame fingerprint of the current target frame is determined based on the number sequence.

3. The generation method as described in claim 2, characterized in that, Determining the frame fingerprint of the current target frame based on the number sequence includes: The target number is determined by selecting the number at a specified position in the numbering sequence; or, two or more numbers are selected consecutively, at intervals, or at a specified position from the numbering sequence as the target number. The target number is determined as the frame fingerprint.

4. The generation method as described in claim 2, characterized in that, Determining the frame fingerprint of the current target frame based on the number sequence includes: Generate corresponding unique identifier information based on the sorting results; The unique identifier information is determined as the frame fingerprint.

5. The generation method according to any one of claims 1-4, characterized in that, The step of obtaining the video fingerprint of the video to be processed based on the frame fingerprints corresponding to each target frame in at least one target video segment includes: For each target video segment, the frame fingerprints corresponding to each target frame in the current target video segment are concatenated in time sequence to obtain the segment fingerprint corresponding to the current target video segment. When the number of target video segments is 1, the segment fingerprint of the target video segment is determined as the video fingerprint; When the number of target video segments is greater than 1, the video fingerprint is obtained by concatenating the fingerprints of each segment of the target video segments in time sequence.

6. The generation method according to any one of claims 1-4, characterized in that, The visual features include at least one of brightness features, color features, and texture features.

7. The generation method according to any one of claims 1-4, characterized in that, Following the video fingerprint, the following is also included: The video fingerprint is compressed to obtain the final video fingerprint.

8. A device for generating video fingerprints, characterized in that, include: The first generation module is used to obtain the video fingerprint of the video to be processed based on the frame fingerprint of each target frame in at least one target video segment. The target video segment is a video segment in the video to be processed; the target frame is some or all of the video frames in the target video segment. The video fingerprint generation device further includes a second generation module, which is triggered and executed before the first generation module runs. For each target frame, the second generation module is used to: Divide the current target frame into multiple grid cells; The frame fingerprint of the current target frame is obtained by quantizing the visual features of each grid cell.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the video fingerprint generation method as described in any one of claims 1 to 7.

10. A computer program product, the computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the video fingerprint generation method as described in any one of claims 1 to 7.