A video map frame accurate identification method and system
Patent Information
- Application Number
- CN202511701591.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-19
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2045-11-19
AI Technical Summary
[0004]为了解决现有技术中存在的问题,本公开提出一种视频地图帧精准识别方法、系统,以解决现有技术无法满足大规模数据处理中的视频地图帧精准识别需求的问题,本公开采用的技术方案为:
[0017] The beneficial effects of this disclosure are as follows: This disclosure provides a method and system for accurate identification of video map frames. By integrating multi-scale scene analysis and visual language models, it significantly reduces the time and cost of manual extraction and overcomes the shortcomings of existing technologies in terms of robustness when dealing with challenges such as complex transitions and geometric distortions. It can meet the needs of accurate identification of video map frames in large-scale data processing and helps to efficiently and compliantly review and govern map images on short video platforms.
Smart Images

Figure CN121789107B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of video map technology, and more specifically, to a method and system for accurate identification of video map frames. Background Technology
[0002] Video map frame extraction technology is used to accurately identify and extract core images containing map information from massive amounts of video data. Given the widespread dissemination of map images on short video platforms, ensuring that their content complies with geographic information compliance requirements has become an urgent task.
[0003] Traditional manual review methods are no longer adequate in terms of accuracy, efficiency, and scalability to meet the current demands of large-scale data processing. While some existing intelligent review methods exist to assist manual review, these methods generally suffer from poor accuracy and insufficient robustness in handling challenges such as complex transitions and geometric distortions. Consequently, they still cannot meet current market needs or the requirements for accurate video map frame recognition in large-scale data processing. Summary of the Invention
[0004] To address the problems existing in the prior art, this disclosure proposes a method and system for accurate video map frame recognition, solving the problem that the prior art cannot meet the requirements for accurate video map frame recognition in large-scale data processing. The technical solution adopted in this disclosure is as follows: In a first aspect, this disclosure provides a method for accurate identification of video map frames, the method comprising: Convert the read video stream data into a discrete video frame sequence; Deep feature extraction is performed on the discrete video frame sequence to obtain the feature vector of each video frame in the discrete video frame sequence; duplicate video frames are eliminated by comparing the consistency of the feature vectors between adjacent video frames to obtain the inter-frame difference sequence. By performing high-sensitivity mutation detection and subtle change detection on the inter-frame difference sequence, the mutation scene boundary point sequence and the subtle scene boundary point sequence are obtained respectively. The mutation scene boundary point sequence and the subtle scene boundary point sequence are merged and deduplicated to obtain the final scene boundary point sequence. Based on the final scene boundary point sequence, the inter-frame difference sequence is divided into several scene segments, and representative video frames are extracted from each scene segment to obtain a set of representative video frames. The representative video frame set is input into the visual language model, and the visual language model performs semantic understanding and recognition on each representative video frame. The visual language model outputs representative video frames containing map elements, which are the final video map frames containing map elements.
[0005] Preferably, before converting the read video stream data into a discrete video frame sequence, the method further includes: The video stream data is read using the video capture module of the OpenCV library.
[0006] Preferably, in the process of converting the read video stream data into a discrete video frame sequence, a fixed-interval frame extraction strategy is used to convert the read video stream data into a discrete video frame sequence.
[0007] Preferably, in the process of extracting deep features from the discrete video frame sequence, a ResNet-50 residual network with the end Softmax layer removed can be used to extract deep features from the discrete video frame sequence.
[0008] Preferably, the step of eliminating duplicate video frames by comparing the consistency of feature vectors between adjacent video frames is achieved by calculating cosine similarity to compare the consistency of feature vectors between adjacent video frames, thereby eliminating duplicate video frames. The application of cosine similarity is transformed into ensuring that there are obvious differences between video frames, thereby constructing a discrete video frame sequence that can quantify the continuous changes in video content.
[0009] Preferably, in the step of performing high-sensitivity mutation detection and subtle change detection on the inter-frame difference sequence in parallel, the high-sensitivity mutation detection path and the subtle change detection path are respectively performed through parallel high-sensitivity mutation detection path and subtle change detection path.
[0010] Preferably, the high-sensitivity mutation detection path is configured with complementary statistical thresholding, Otsu's algorithm, and gradient analysis.
[0011] Preferably, the subtle change detection path is configured with a subtle change detection strategy, which includes: The difference values of adjacent video frames are continuously accumulated within each scene in the inter-frame difference sequence to quantify the cumulative effect of weak amplitude changes within each scene. Video frames that meet the preset subtle scene judgment conditions in each scene are taken as subtle scene boundary points, and all the obtained subtle scene boundary points are collected into the subtle scene boundary point sequence.
[0012] Preferably, the representative video frame is the median frame of the corresponding scene segment. Selecting the median frame as the representative video frame of the scene segment can maximize the elimination of redundant frame data in the discrete video frame sequence or inter-frame difference sequence, thereby efficiently summarizing the core content of the scene.
[0013] Preferably, the large visual language model is MiniCPM-V-2.6.
[0014] A second aspect of this disclosure provides a system for accurate identification of video map frames, the system comprising: The discrete sequence module is used to convert the read video stream data into a discrete video frame sequence; The difference sequence module is used to perform deep feature extraction on the discrete video frame sequence to obtain the feature vector of each video frame in the discrete video frame sequence; by comparing the consistency of the feature vectors between adjacent video frames, duplicate video frames are excluded to obtain the inter-frame difference sequence. The boundary point sequence module is used to perform highly sensitive mutation detection and subtle change detection on the inter-frame difference sequence to obtain the mutation scene boundary point sequence and the subtle scene boundary point sequence respectively. The mutation scene boundary point sequence and the subtle scene boundary point sequence are merged and deduplicated to obtain the final scene boundary point sequence. The representative video frame module is used to divide the inter-frame difference sequence into several scene segments based on the final scene boundary point sequence, and extract the representative video frames from each scene segment to obtain a set of representative video frames. The large model recognition module is used to input the representative video frame set into the visual language large model, perform semantic understanding and recognition on each representative video frame through the visual language large model, and output representative video frames containing map elements. The representative video frames containing map elements are the final video map frames containing map elements.
[0015] In a third aspect, this disclosure provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the video map frame accurate identification method described above.
[0016] In a fourth aspect, this disclosure provides an electronic device including a processor and a memory, the processor being configured to execute a computer program stored in the memory to implement the video map frame accurate identification method described above.
[0017] The beneficial effects of this disclosure are as follows: This disclosure provides a method and system for accurate identification of video map frames. By integrating multi-scale scene analysis and visual language models, it significantly reduces the time and cost of manual extraction and overcomes the shortcomings of existing technologies in terms of robustness when dealing with challenges such as complex transitions and geometric distortions. It can meet the needs of accurate identification of video map frames in large-scale data processing and helps to efficiently and compliantly review and govern map images on short video platforms.
[0018] This disclosure utilizes parallel, highly sensitive mutation detection paths and subtle change detection paths for multi-scale detection and analysis, thereby achieving precise decoupling of inter-frame difference sequences. Furthermore, it incorporates a large visual language model (MiniCPM-V). 2.6)'s powerful semantic reasoning capabilities ensure robust localization with no missed detections for all types of scene boundaries, including mutations and gradations. This disclosure fully utilizes MiniCPM-V Version 2.6's high-precision semantic discrimination function completely solves the problem of recognizing complex transition and geometrically distorted maps.
[0019] This invention relies on the powerful visual understanding capabilities of the visual language large model and its strict adherence to prompts, enabling it to quickly analyze representative video frames and make judgments with clear "yes" or "no" answers. It exhibits strong recognition robustness even when faced with complex backgrounds or geometrically distorted and warped map images. Compared to existing technologies, which suffer from poor accuracy and insufficient robustness in handling complex transitions and geometric distortions, this invention demonstrates superior capabilities.
[0020] This disclosure converts video stream data into a discrete video frame sequence, and further extracts deep features from the discrete video frame sequence and compares the consistency of feature vectors between adjacent video frames to obtain a simplified inter-frame difference sequence. This inter-frame difference sequence is used as input data for scene detection and scene analysis, resulting in a representative set of video frames used as input to a large visual language model. This significantly improves recall while maintaining high accuracy, achieving a complete and efficient technical closed loop from raw video stream data to accurate identification and extraction of map elements, ultimately establishing a high-efficiency, fully automated workflow. Attached Figure Description
[0021] The accompanying drawings, which form part of this application, are used to provide a further understanding of this disclosure. The illustrative embodiments of this disclosure and their descriptions are used to explain this disclosure and do not constitute an undue limitation of this disclosure.
[0022] Figure 1 This is a flowchart of the video map frame accurate identification method described in Embodiment 1 of this disclosure.
[0023] Figure 2 This is a schematic diagram of the ResNet-50 residual network with the terminal Softmax layer removed, as described in Embodiment 1 of this disclosure.
[0024] Figure 3 This is an example diagram illustrating the geometric deformation map recognition capability of the video map frame accurate recognition method described in Embodiment 1 of this disclosure.
[0025] Figure 4This is an example diagram of the discrete video frame sequence described in Embodiment 1 of this disclosure; wherein a, b, and c are discrete video frame sequences converted from three different segments of video stream data, respectively.
[0026] Figure 5 This is an example diagram of the final scene boundary point sequence described in Embodiment 1 of this disclosure; where a, b, and c are the final scene boundary point sequences corresponding to three different segments in the video stream data, respectively.
[0027] Figure 6 This describes the process of recognizing video frames using the large visual language model described in Embodiment 1 of this disclosure.
[0028] Figure 7 This is an example diagram of a video map frame containing map elements as described in Embodiment 1 of this disclosure.
[0029] Figure 8 This is an architecture diagram of the video map frame accurate recognition system described in Embodiment 2 of this disclosure. Detailed Implementation
[0030] The present disclosure will now be described in detail with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in the present application can be combined with each other.
[0031] The following detailed descriptions are exemplary and intended to provide further detailed explanation of this disclosure. Unless otherwise specified, all technical terms used in this disclosure have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments according to this disclosure.
[0032] Example 1: like Figure 1 As shown, this disclosure provides a method for accurate identification of video map frames, the method including steps S100 to S500.
[0033] S100: Convert the read video stream data into a discrete video frame sequence.
[0034] Furthermore, before step S100, which converts the read video stream data into a discrete video frame sequence, the method further includes: The video stream data is read using the video capture module of the OpenCV library.
[0035] During implementation, the high performance of the OpenCV library's video capture module ensures stable reading of video stream data. The video capture module can iteratively execute the read operation, performing the `read()` method frame-by-frame on the video stream data, enabling pixel-level, lossless parsing of the raw video stream data.
[0036] Video frame reading is a key starting point in the video structured analysis process. Its core objective is to transform continuous temporal video signals into a discrete video frame sequence that is computationally feasible. By using the fixed-interval frame extraction strategy to convert video stream data into a discrete video frame sequence, a precise and quantifiable data foundation can be laid for subsequent depth feature extraction, scene analysis, and map recognition.
[0037] Further, in step S100, the read video stream data is converted into a discrete video frame sequence by using a fixed-interval frame extraction strategy. For example... Figure 4 As shown, three segments are selected from continuous video stream data and converted into corresponding discrete video frame sequences. Figure 4 (a) Figure 4 (b) Figure 4 (c) represents discrete video frame sequences a, b, and c, respectively.
[0038] Furthermore, the fixed-interval frame-skipping strategy includes: S101. Using a pre-set fixed step size as the sampling interval, perform equidistant skip sampling in the video stream data based on the sampling interval to obtain several static video frames; the distance between two adjacent static video frames is equal to the sampling interval. S102. All the extracted static video frames are combined into the discrete video frame sequence.
[0039] The fixed-interval frame extraction strategy is based on a predetermined fixed step size n (i.e., sampling interval). The value of n depends on actual needs, and n is preferably an integer greater than or equal to 5. Based on the fixed step size n (i.e., sampling interval), static video frames are accurately extracted in the entire video stream at equal intervals (the distance is equal to the fixed step size n). This optimizes the computational load of subsequent analysis while ensuring sufficient representation of video events. As a systematic and uniform sampling mechanism, the fixed-interval frame extraction strategy effectively avoids the consumption of processing resources by redundant frames and ensures that the extracted image sequence is representative in the time dimension.
[0040] In addition, all the extracted static video frames can be assigned a unique time series number and stored as independent, high-fidelity image files, thus constructing a time-ordered and comprehensive discrete video frame sequence, which greatly improves the efficiency and ease of implementation of data preprocessing, and achieves uniform and high-fidelity coverage of video narrative content with minimal computational cost.
[0041] S200. Perform deep feature extraction on the discrete video frame sequence to obtain the feature vector of each video frame in the discrete video frame sequence; eliminate duplicate video frames by comparing the consistency of the feature vectors between adjacent video frames to obtain the inter-frame difference sequence.
[0042] Furthermore, in step S200, during the deep feature extraction of the discrete video frame sequence, a ResNet-50 residual network with the terminal Softmax layer removed can be used to extract deep features from the discrete video frame sequence, such as... Figure 2 As shown. Of course, other network models can also be used for deep feature extraction.
[0043] Furthermore, the process of eliminating duplicate video frames by comparing the consistency of feature vectors between adjacent video frames is achieved by calculating cosine similarity to compare the consistency of feature vectors between adjacent video frames, thereby eliminating duplicate video frames. The application of cosine similarity is transformed into ensuring that there are obvious differences between video frames, thereby constructing a discrete video frame sequence that can quantify the continuous changes in video content.
[0044] Furthermore, a similarity threshold can be set to evaluate the consistency of feature vectors. For example, if the similarity threshold is set to 95%, and the cosine similarity of two feature vectors is >= 95%, then the two feature vectors are considered to be consistent, and the operation of excluding duplicate video frames should be performed.
[0045] S300: By performing high-sensitivity mutation detection and subtle change detection on the inter-frame difference sequence, a mutation scene boundary point sequence and a subtle scene boundary point sequence are obtained respectively. The mutation scene boundary point sequence and the subtle scene boundary point sequence are merged and deduplicated to obtain the final scene boundary point sequence. Figure 5 As shown, the discrete video frame sequences a, b, and c are processed by S200 and S300 to obtain the final scene boundary point sequences a, b, and c. Furthermore, in step S300, during the parallel high-sensitivity mutation detection and subtle change detection of the inter-frame difference sequence, the high-sensitivity mutation detection path and the subtle change detection path are connected in parallel to perform the high-sensitivity mutation detection and the subtle change detection, respectively. Scene boundary points, as the name suggests, are one or more video frames in the scene sequence (inter-frame difference sequence) that can reflect scene switching.
[0046] During implementation, the inter-frame difference sequence is used as input data, and is fed into parallel high-sensitivity mutation detection paths and subtle change detection paths respectively. High-sensitivity mutation detection and subtle change detection are performed in parallel, thereby achieving multi-path (multi-scale) analysis of the inter-frame difference sequence, facilitating comprehensive coverage of various scene switching modes. After detection by the parallel high-sensitivity mutation detection paths and subtle change detection paths, the obtained mutation scene boundary point sequences and subtle scene boundary point sequences are merged and deduplicated to form the final scene boundary point sequence, thus efficiently and robustly completing the logical division of video content.
[0047] The high-sensitivity mutation detection path performs high-sensitivity mutation detection, focusing on identifying hard-cut scene transitions with significant inter-frame differences, and converting these transitions into a sequence of boundary points for the mutation scene. The subtle change detection path performs subtle change detection, focusing on capturing slow but continuous scene gradations—i.e., subtle scene changes—to compensate for the shortcomings of the high-sensitivity mutation detection.
[0048] Furthermore, the highly sensitive mutation detection path is equipped with complementary statistical thresholding, Otsu's algorithm, and gradient analysis.
[0049] Furthermore, the statistical threshold method calculates the mean and standard deviation of the inter-frame difference sequence, dynamically adjusts the sensitivity coefficient k according to the preset detection requirements, sets an adaptive threshold (mean plus k times the standard deviation) in combination with the sensitivity coefficient k, and takes each video frame in the inter-frame difference sequence whose difference exceeds the adaptive threshold as a first candidate mutation scene boundary point, and assigns each first candidate mutation scene boundary point to the mutation scene boundary point sequence.
[0050] Furthermore, the Otsu algorithm uses the distribution of the inter-frame difference sequence as a gray-level histogram, automatically calculates the optimal segmentation threshold by maximizing the inter-class variance, and takes each video frame in the inter-frame difference sequence whose difference exceeds the optimal segmentation threshold as a second candidate mutation scene boundary point. Each second candidate mutation scene boundary point is then assigned to the mutation scene boundary point sequence, thereby achieving adaptive recognition of significant difference boundaries.
[0051] Furthermore, the gradient analysis method focuses on the rate of change of difference in the inter-frame difference sequence. By calculating the first derivative of the inter-frame difference sequence, the average difference value and standard deviation of the gradient are analyzed, and the same sensitivity coefficient k as the statistical thresholding method is introduced to dynamically set the gradient detection threshold. Each video frame in the inter-frame difference sequence whose gradient value exceeds the gradient detection threshold is regarded as a third candidate mutation scene boundary point. Each third candidate mutation scene boundary point is included in the mutation scene boundary point sequence, thereby effectively identifying frames in the inter-frame difference sequence where the difference suddenly accelerates.
[0052] Preferably, the subtle change detection path is configured with a subtle change detection strategy, which includes: The difference values of adjacent video frames are continuously accumulated within each scene in the inter-frame difference sequence to quantify the cumulative effect of weak amplitude changes within each scene. Video frames that meet the preset subtle scene judgment conditions in each scene are taken as subtle scene boundary points, and all the obtained subtle scene boundary points are collected into the subtle scene boundary point sequence.
[0053] In implementation, the preset subtle scene determination conditions can be determined according to actual needs, preferably including: when the accumulated difference value at a certain position in a scene reaches a preset threshold (e.g., the accumulated difference reaches 0.3 for 30 consecutive frames), or when the scene lasts for a long time without significant changes (e.g., the difference tends to stabilize for 100 consecutive frames), the video frame at the corresponding position is taken as the subtle scene boundary point. The subtle change detection strategy ensures accurate identification of gradual scene changes such as push-pull shots and fade-in / fade-out.
[0054] S400. Based on the final scene boundary point sequence, the inter-frame difference sequence is divided into several scene segments, and representative video frames are extracted from each scene segment to obtain a set of representative video frames.
[0055] Furthermore, the representative video frame is preferably the median frame of the corresponding scene segment. Selecting the median frame as the representative video frame of the scene segment can maximize the elimination of redundant frame data in the discrete video frame sequence or inter-frame difference sequence, thereby efficiently summarizing the core content of the scene.
[0056] It should be noted that S200, S300, and S400 are the steps that connect the preceding and following steps in the video frame scene analysis.
[0057] S500: Input the representative video frame set into the visual language model. The visual language model performs semantic understanding and recognition on each representative video frame. The visual language model outputs representative video frames containing map elements. These representative video frames containing map elements are the final video map frames containing map elements. The process of the visual language model recognizing video frames is as follows: Figure 6 As shown. The final video map frame containing map elements is as follows. Figure 7 As shown.
[0058] Furthermore, the visual language large model is preferably MiniCPM-V-2.6.
[0059] Working Principle: To accurately extract map elements from video stream data, the method converts the video stream data into a discrete video frame sequence. Then, after further deep feature extraction, the consistency of feature vectors between adjacent video frames is compared to eliminate duplicate video frames, resulting in an inter-frame difference sequence after removing redundant data. This inter-frame difference sequence is then subjected to multi-scale detection via parallel high-sensitivity mutation detection paths and subtle change detection paths, yielding a multi-scale scene boundary point sequence. After merging and deduplication, a final, concise scene boundary point sequence is obtained that considers both abrupt and subtle scene changes. Based on this final scene boundary point sequence, scene analysis is performed on the inter-frame difference sequence, resulting in several scene segments. The median frame of each scene segment is extracted as its representative video frame, and these segments are combined into a representative video frame set. This maximizes the elimination of frame sequence redundancy and efficiently summarizes the core content of the scene.
[0060] All images in video frames that are identified as "yes" by the visual language model will be immediately named and stored based on their original timestamp information, completing a fully automated map element extraction process, just as... Figure 6 As shown.
[0061] The representative set of video frames is then input into a powerful visual language model (such as MiniCPM-V-2.6) for deep semantic reasoning and map element discrimination. The visual language model can be pre-defined with judgment logic, which can be implemented in the discrimination stage by constructing highly focused prompts such as: "I am a map review expert; to determine whether there is an obvious map in the image, please only answer 'yes' or 'no'," to guide the visual language model to perform a rigorous binary classification task.
[0062] like Figure 3As shown, this disclosure relies on the powerful visual understanding ability of the visual language big data model and its strict adherence to prompts, which enables it to quickly analyze representative video frames and make a judgment with a clear "yes" or "no" answer. It can also show strong recognition robustness even when faced with complex backgrounds or geometrically distorted and warped map images. Compared with the existing technologies, which have poor accuracy and insufficient robustness in handling complex transitions and geometric distortions, this disclosure has a good ability to cope with these problems.
[0063] To further improve video frame processing efficiency and industrial application value, the method can be configured with a multi-threaded parallel mechanism and queued scheduling to achieve synchronous and efficient processing of multiple video tasks, significantly improving the throughput of large-scale video data processing, and ultimately confirming the effectiveness of the automated solution for efficiently locating, identifying and extracting map elements from video stream data.
[0064] Example 2: like Figure 8 As shown, in a second aspect, this disclosure provides a video map frame accurate recognition system, the system comprising: The discrete sequence module 100 is used to convert the read video stream data into a discrete video frame sequence. The difference sequence module 200 is used to perform deep feature extraction on the discrete video frame sequence to obtain the feature vector of each video frame in the discrete video frame sequence; and to exclude duplicate video frames by comparing the consistency of the feature vectors between adjacent video frames to obtain the inter-frame difference sequence. The boundary point sequence module 300 is used to perform high-sensitivity mutation detection and subtle change detection on the inter-frame difference sequence to obtain the mutation scene boundary point sequence and the subtle scene boundary point sequence respectively, and to obtain the final scene boundary point sequence by merging and deduplicating the mutation scene boundary point sequence and the subtle scene boundary point sequence. The representative video frame module 400 is used to divide the inter-frame difference sequence into several scene segments based on the final scene boundary point sequence, and extract the representative video frames from each scene segment to obtain a set of representative video frames. The large model recognition module 500 is used to input the representative video frame set into the visual language large model, perform semantic understanding and recognition on each representative video frame through the visual language large model, and output representative video frames containing map elements. The representative video frames containing map elements are the final video map frames containing map elements.
[0065] In Example 2, the discrete sequence module 100, the difference sequence module 200, the boundary point sequence module 300, the representative video frame module 400, and the large model recognition module 500 correspond to S100, S200, S300, S400, and S500 in Example 1, respectively.
[0066] Example 3: Embodiment 3 of this disclosure provides a computer-readable storage medium storing a computer program. When executed by a processor, the computer program implements the video map frame accurate recognition method as described in Embodiment 1. Alternatively, a video map frame accurate recognition system as described in Example 2 can be implemented.
[0067] The computer-readable storage medium includes volatile or non-volatile, removable or non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, computer program modules, or other data). Computer-readable storage media include, but are not limited to, RAM (Random Access Memory), ROM (Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), flash memory or other memory technologies, CD-ROM (Compact Disc Read-Only Memory), DVD or other optical disc storage, cartridges, magnetic tapes, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer.
[0068] Example 4: Embodiment 4 of this disclosure provides an electronic device including a processor and a memory, wherein the processor is used to execute a computer program stored in the memory to implement the video map frame accurate recognition method described in Embodiment 1.
[0069] Based on this understanding, all or part of the processes in the methods of the above embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, and read-only memory (ROM).
[0070] Those skilled in the art will understand that embodiments of this disclosure can be provided as methods, systems, or computer program products. Therefore, this disclosure can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this disclosure can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0071] This disclosure is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0072] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0073] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0074] In the description of this specification, references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this disclosure. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0075] In summary, the video map frame accurate recognition method and system provided in embodiments 1-4 of this disclosure, by integrating multi-scale scene analysis and visual language models, significantly reduces the time and cost of manual extraction, and overcomes the shortcomings of insufficient robustness of existing technologies in handling challenges such as complex transitions and geometric distortions. It can meet the needs of accurate video map frame recognition in large-scale data processing and contributes to the efficient and compliant review and governance of map images on short video platforms, etc.
[0076] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this disclosure and not to limit them. Although this disclosure has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of this disclosure. Any modifications or equivalent substitutions that do not depart from the spirit and scope of this disclosure should be covered within the protection scope of the claims of this disclosure.
Claims
1. A method for accurate identification of video map frames, characterized in that, The method includes: Convert the read video stream data into a discrete video frame sequence; Deep feature extraction is performed on the discrete video frame sequence to obtain the feature vector of each video frame in the discrete video frame sequence; duplicate video frames are eliminated by comparing the consistency of the feature vectors between adjacent video frames to obtain the inter-frame difference sequence. By performing high-sensitivity mutation detection and subtle change detection on the inter-frame difference sequence, the mutation scene boundary point sequence and the subtle scene boundary point sequence are obtained respectively. The mutation scene boundary point sequence and the subtle scene boundary point sequence are merged and deduplicated to obtain the final scene boundary point sequence. Based on the final scene boundary point sequence, the inter-frame difference sequence is divided into several scene segments, and representative video frames are extracted from each scene segment to obtain a set of representative video frames. The representative video frame set is input into the visual language model. The visual language model performs semantic understanding and recognition on each representative video frame. The visual language model outputs representative video frames containing map elements. The representative video frames containing map elements are the final video map frames containing map elements. In the process of performing high-sensitivity mutation detection and subtle change detection on the inter-frame difference sequence in parallel, the high-sensitivity mutation detection path and the subtle change detection path are respectively performed in parallel. The highly sensitive mutation detection path is equipped with complementary statistical thresholding, Otsu's algorithm, and gradient analysis. The subtle change detection path is configured with a subtle change detection strategy, which includes: The difference values of adjacent video frames are continuously accumulated within each scene in the inter-frame difference sequence to quantify the cumulative effect of weak amplitude changes within each scene. Video frames that meet the preset subtle scene judgment conditions in each scene are taken as subtle scene boundary points, and all the obtained subtle scene boundary points are collected into the subtle scene boundary point sequence.
2. The video map frame accurate recognition method as described in claim 1, characterized in that, Before converting the read video stream data into a discrete video frame sequence, the method further includes: The video stream data is read using the video capture module of the OpenCV library.
3. The video map frame accurate recognition method as described in claim 1, characterized in that, The process of converting the read video stream data into a discrete video frame sequence involves using a fixed-interval frame extraction strategy. The fixed-interval frame extraction strategy includes: A fixed step size is set in advance as the sampling interval. Based on the sampling interval, the video stream data is sampled at equal intervals to obtain a number of static video frames. The distance between two adjacent static video frames is equal to the sampling interval. All the extracted static video frames are combined into the discrete video frame sequence.
4. The video map frame accurate recognition method as described in claim 1, characterized in that, In the process of extracting deep features from the discrete video frame sequence, a ResNet-50 residual network with the end Softmax layer removed is used to extract deep features from the discrete video frame sequence. The method of eliminating duplicate video frames by comparing the consistency of feature vectors between adjacent video frames is achieved by calculating cosine similarity.
5. The video map frame accurate recognition method as described in claim 1, characterized in that, The representative video frame is the median frame of the corresponding scene segment.
6. The video map frame accurate recognition method as described in claim 1, characterized in that, The large-scale visual language model is MiniCPM-V-2.
6.
7. A video map frame accurate recognition system, characterized in that, The system includes: The discrete sequence module is used to convert the read video stream data into a discrete video frame sequence; The difference sequence module is used to perform deep feature extraction on the discrete video frame sequence to obtain the feature vector of each video frame in the discrete video frame sequence; by comparing the consistency of the feature vectors between adjacent video frames, duplicate video frames are excluded to obtain the inter-frame difference sequence. The boundary point sequence module is used to perform highly sensitive mutation detection and subtle change detection on the inter-frame difference sequence to obtain the mutation scene boundary point sequence and the subtle scene boundary point sequence respectively. The mutation scene boundary point sequence and the subtle scene boundary point sequence are merged and deduplicated to obtain the final scene boundary point sequence. The representative video frame module is used to divide the inter-frame difference sequence into several scene segments based on the final scene boundary point sequence, and extract the representative video frames from each scene segment to obtain a set of representative video frames. The large model recognition module is used to input the representative video frame set into the visual language large model, perform semantic understanding and recognition on each representative video frame through the visual language large model, and output the representative video frame containing map elements. The representative video frame containing map elements is the final video map frame containing map elements. In the process of performing high-sensitivity mutation detection and subtle change detection on the inter-frame difference sequence in parallel, the high-sensitivity mutation detection path and the subtle change detection path are respectively performed in parallel. The highly sensitive mutation detection path is equipped with complementary statistical thresholding, Otsu's algorithm, and gradient analysis. The subtle change detection path is configured with a subtle change detection strategy, which includes: The difference values of adjacent video frames are continuously accumulated within each scene in the inter-frame difference sequence to quantify the cumulative effect of weak amplitude changes within each scene. Video frames that meet the preset subtle scene judgment conditions in each scene are taken as subtle scene boundary points, and all the obtained subtle scene boundary points are collected into the subtle scene boundary point sequence.
Citation Information
Patent Citations
Method for identifying and replacing constituent elements in video and method for recommending video
CN115115979A
Scene change detection method and device, electronic equipment and storage medium
CN120676149A