Key frame identification method and device, computing equipment and computer readable storage medium

By recording video in the target mode of displaying positioning marks in the interactive page, and using the positioning coordinates of positioning marks to automatically identify keyframes, the problem of time-consuming and limited accuracy in the prior art is solved, and efficient and accurate keyframe recognition is achieved.

CN119967228APending Publication Date: 2025-05-09ALI HEALTH TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411908882.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

In the prior art, keyframe recognition relies on manual recording of video frame-disassembly and frame-by-frame analysis, which takes a long time and relies on human subjective judgment, resulting in low recognition efficiency and limited accuracy.

Method used

By recording video in the target mode of displaying positioning marks in the interactive page, the keyframes that perform the trigger operation are automatically identified using the positioning marks' positioning coordinates and set keyframe coordinate constraints.

Benefits of technology

It realizes the automation of keyframe recognition, reduces time-consuming, improves recognition efficiency and accuracy, and does not rely on human subjective judgments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119967228A_ABST
    Figure CN119967228A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a key frame recognition method and device, computing equipment and a computer readable storage medium, and the key frame recognition method comprises the steps that a to-be-recognized video is obtained, the to-be-recognized video is a video obtained by recording an interaction page in a target mode, and the target mode is a mode of displaying a positioning mark in the interaction page; identifying a trigger operation of the interactive page through the positioning mark in the target mode; splitting the to-be-identified video into at least one video frame, and determining a positioning coordinate of the positioning mark in each video frame; and according to the positioning coordinate of the positioning mark in each video frame and a set key frame coordinate constraint, identifying a key frame from each video frame, the key frame being a video frame for executing a trigger operation. According to the invention, the key frame executing the trigger operation in each video frame is automatically identified based on the positioning mark in the target mode, so that the key frame identification efficiency and accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present specification relate to the field of video processing technology, and more particularly to a key frame recognition method, apparatus, computing device, and computer-readable storage medium. Background Art

[0002] With the rapid development of computer technology and the increasing complexity of Internet applications, the user experience of interactive pages has become crucial. Front-end performance testing and performance optimization are one of the key factors to improve user experience, which directly affects the loading speed, interactive response time and visual fluency of interactive pages. In order to effectively evaluate and optimize front-end performance, developers and test engineers usually rely on various tools and technologies to monitor and analyze the changes and rendering process of interactive pages. Keyframes represent the timing when users perform trigger operations on interactive pages. Keyframe recognition is of great significance for front-end performance testing such as interactive page loading speed, interactive response time and visual fluency.

[0003] In the existing technology, the key frames are mainly identified by manually decomposing the recorded video of the interactive page and analyzing it frame by frame. This key frame identification method is time-consuming and relies on human subjective judgment conditions, resulting in low key frame identification efficiency and limited accuracy. Therefore, a more efficient and accurate key frame identification solution is urgently needed. Summary of the invention

[0004] In view of this, an embodiment of this specification provides a key frame recognition method. One or more embodiments of this specification also relate to a key frame recognition device, a computing device, a computer-readable storage medium and a computer program product to solve the technical defects existing in the prior art.

[0005] According to a first aspect of an embodiment of this specification, a key frame recognition method is provided, including: Acquire a video to be identified, wherein the video to be identified is a video recorded on an interactive page in a target mode, the target mode is a mode in which a positioning mark is displayed in the interactive page, and a triggering operation of the interactive page is identified by the positioning mark in the target mode; Splitting the to-be-recognized video into at least one video frame, and determining the positioning coordinates of the positioning mark in each video frame; According to the positioning coordinates of the positioning marks in the video frames and the set key frame coordinate constraints, key frames are identified from the video frames, wherein the key frames are video frames for performing trigger operations.

[0006] According to a second aspect of an embodiment of this specification, a key frame identification device is provided, including: an acquisition module, configured to acquire a video to be identified, wherein the video to be identified is a video obtained by recording an interactive page in a target mode, the target mode is a mode in which a positioning mark is displayed in the interactive page, and a triggering operation of the interactive page is identified by the positioning mark in the target mode; A determination module is configured to split the video to be identified into at least one video frame and determine the positioning coordinates of the positioning mark in each video frame; The identification module is configured to identify key frames from the video frames according to the positioning coordinates of the positioning marks in the video frames and the set key frame coordinate constraints, wherein the key frames are video frames for performing trigger operations.

[0007] According to a third aspect of an embodiment of this specification, a computing device is provided, including: Memory and processor; The memory is used to store computer executable instructions, and the processor is used to execute the computer executable instructions. When the computer executable instructions are executed by the processor, the steps of the key frame recognition method are implemented.

[0008] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores computer-executable instructions, and when the instructions are executed by a processor, the steps of the key frame identification method are implemented.

[0009] According to a fifth aspect of the embodiments of this specification, a computer program product is provided, including a computer program / instruction, which implements the steps of the above-mentioned key frame identification method when executed by a processor.

[0010] An embodiment of the present specification provides a key frame recognition method, which obtains a video to be recognized, wherein the video to be recognized is a video obtained by recording an interactive page in a target mode, and the target mode is a mode in which a positioning mark is displayed in the interactive page, and a trigger operation of the interactive page is identified by the positioning mark in the target mode; the video to be recognized is split into at least one video frame, and the positioning coordinates of the positioning mark in each video frame are determined; and key frames are identified from each video frame according to the positioning coordinates of the positioning mark in each video frame and the set key frame coordinate constraints, wherein the key frames are video frames for performing trigger operations.

[0011] An embodiment of the present specification realizes that an interactive page is recorded in a target mode to obtain a video to be identified. A positioning mark can be displayed in the interactive page in the target mode. The trigger operation of the interactive page is identified by the positioning mark in the target mode. Subsequently, the positioning coordinates of the positioning mark in each video frame can be identified to automatically identify the key frames that perform the trigger operation from each video frame. The key frames that perform the trigger operation in each video frame are automatically identified based on the positioning mark in the target mode. No human input is required, which greatly reduces the time consumption of key frame identification and improves the identification efficiency. It does not rely on human subjective judgment conditions, and improves the accuracy of key frame identification. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 is a flow chart of a key frame identification method provided by an embodiment of this specification; Figure 2a is a schematic diagram of a first interactive page provided by an embodiment of this specification; Figure 2b is a schematic diagram of a second interactive page provided by an embodiment of this specification; Figure 2c is a schematic diagram of a third interactive page provided by an embodiment of this specification; Figure 2d It is a schematic diagram of a coordinate change curve of a customized mark provided by an embodiment of this specification; Figure 2e is a schematic diagram of a region to be identified in a video frame provided by an embodiment of the present specification; Figure 3 is a process flow chart of a key frame identification method provided by an embodiment of this specification; Figure 4 is a structural diagram of a key frame identification device provided by an embodiment of this specification; Figure 5 It is a structural block diagram of a computing device provided by an embodiment of this specification. DETAILED DESCRIPTION

[0013] Many specific details are described in the following description to facilitate a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the connotation of this specification, so this specification is not limited to the specific implementation disclosed below.

[0014] The terms used in one or more embodiments of this specification are only for the purpose of describing specific embodiments, and are not intended to limit one or more embodiments of this specification. The singular forms of "a" and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0015] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, this information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0016] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0017] First, the terms involved in one or more embodiments of this specification are explained.

[0018] OCR (Optical Character Recognition): is a technology that converts text content in an image into computer-readable text. Specifically, OCR can analyze and understand the text information contained in an image file or scanned document through a computer program, and convert it into computer-editable and searchable text data.

[0019] SSIM (Structural Similarity Index Measure): is an indicator used to measure the similarity between two images. It is widely used in image processing and video compression to evaluate image quality. It can also be applied to scenarios where it is necessary to quantify the similarity between images. SSIM aims to more accurately reflect the human visual system's perception of image quality by considering three important visual characteristics of an image: luminance, contrast, and structure.

[0020] Hough Transform: It is a feature extraction technology widely used in the fields of image processing and computer vision, mainly used to detect geometric shapes in images, such as straight lines, circles, etc. The basic idea of ​​Hough Transform is to map points in the image space to curves or surfaces in the parameter space, and then determine the shape in the original image space by finding the point with the largest intersection in the parameter space. It can effectively detect shapes in the case of noise and partial occlusion.

[0021] iOS: An operating system for front-end devices, known for its simple and intuitive user interface, highly optimized performance, and strict security and privacy protection. It has a rich application ecosystem and introduces new features and technological improvements through regular updates to ensure the consistency and advancement of the user experience.

[0022] Android: It is an open source operating system for front-end devices. It is built on the Linux kernel (an operating system kernel that manages hardware resources and provides underlying services). It is widely used on a variety of front-end devices such as smartphones and tablets. It attracts many manufacturers and developers with its openness and flexibility. It supports high customization, allowing users to adjust system settings and interfaces according to their personal preferences. It provides a large number of applications and services, and the modular architecture of the Android system also promotes technological innovation and diversified product forms.

[0023] It should be noted that, at present, in the process of extracting key frames for front-end performance detection, manual frame-by-frame analysis is relied upon to identify the key frames that have executed the triggering operation after manually splitting the recorded video of the front-end interactive page. This leads to problems such as time-consuming and low efficiency of key frame recognition, inconsistent human subjective judgment conditions, and limited accuracy.

[0024] The embodiments of this specification provide a key frame recognition solution, which realizes automatic key frame positioning by introducing image processing algorithms, without human input, greatly improving work efficiency, and obtaining more accurate key frame recognition results, significantly improving the speed and accuracy of the execution efficiency of frame rendering performance testing. Specifically, for the iOS operating system, the auxiliary touch circle can be used as a positioning mark, and the characteristics of the touch circle combined with technologies such as Hough transform can improve the cross-application compatibility and algorithm recognition accuracy; for the Android operating system, the pointer coordinates of the developer mode can be used as a positioning mark, and interference can be eliminated through dynamic cropping strategies, etc. At the same time, the integration of OCR and SSIM technology enhances the adaptability and accuracy for different applications.

[0025] In this specification, a key frame recognition method is provided. This specification also relates to a key frame recognition device, a computing device, and a computer-readable storage medium, which are described in detail one by one in the following embodiments.

[0026] See also Figure 1 , Figure 1 A flowchart of a key frame recognition method provided according to an embodiment of the present specification is shown, which specifically includes the following steps.

[0027] Step 102: Obtain the video to be identified, wherein the video to be identified is a video recorded on an interactive page in a target mode, and the target mode is a mode in which a positioning mark is displayed on the interactive page. In the target mode, the triggering operation of the interactive page is identified by the positioning mark.

[0028] The key frame recognition method provided in the embodiments of this specification can be applied to a front-end performance automation test tool, through which the front-end interactive page is analyzed to test the performance of the front-end. Of course, in actual implementation, it can also be applied to other video / image processing tools that need to extract key frames.

[0029] It should be noted that the video to be identified is a video recorded in the target mode on the interactive page displayed on the front end. The interactive page is a page provided by the front end to realize user interaction in the corresponding scenario. For example, the interactive page can be the interface of any application or a web page in a search engine. The target mode is a development or debugging mode provided by the front end operating system. The target mode can display a corresponding positioning mark in the interactive page, and use the positioning mark to identify the trigger operation performed by the user in the interactive page displayed on the front end. The trigger operation can be a click, slide, or other operation performed on the interactive page.

[0030] In actual implementation, the target modes corresponding to different operating systems may be different, such as the target mode in the first operating system may be a graphic-assisted touch mode, and the target mode in the second operating system may be a pointer positioning mode. The graphic-assisted touch mode refers to providing a touch graphic to realize the positioning of the mark that triggers the operation in the interactive page, and the pointer positioning mode refers to providing a coordinate area of ​​the touch pointer in the interactive page, and displaying the positioning coordinates of the touch pointer in the coordinate area.

[0031] In one implementation, taking the first operating system as the iOS system as an example, the target mode is the circle assisted touch mode. For the iOS system, enter the accessibility option through the settings page in the front end and select "Touch", and then select "Assisted Touch" to call out the touch circle and set the touch circle to the maximum to highlight the characteristics of the touch circle and eliminate interference from other circular components. The user needs to perform a trigger operation in the interactive page. For example, when clicking a component, the touch circle can be dragged to the position of the component to click, and then the click can be jumped to the details page corresponding to the component. By recording the interactive page of the front end in the circle assisted touch mode for a period of time, the video to be identified can be obtained.

[0032] Figure 2a is a schematic diagram of a first interactive page provided by an embodiment of this specification, such as Figure 2a As shown, taking the interactive page of a medical electronic platform as an example, in the circle-assisted touch mode of the iOS system, a touch circle is displayed on the front-end interactive page. The user can drag the touch circle on the interactive page to perform the trigger operation at the corresponding position.

[0033] In another implementation method, taking the second operating system as the Android system as an example, the target mode is the pointer positioning mode of the Android system, such as the "developer mode". In this mode, the interactive page includes a coordinate area for the touch pointer, and the coordinate area can display the positioning coordinates of the touch pointer. That is, when the interactive page is clicked, the coordinates of the click operation will appear in the coordinate area. When the interactive page is not clicked, the coordinates are not displayed in the coordinate area, that is, the coordinate value is empty.

[0034] Figure 2b is a schematic diagram of a second interactive page provided by an embodiment of this specification, such as Figure 2bAs shown, taking the interactive page of a medical electronic platform as an example, in the "developer mode" of the Android system, a coordinate area is displayed in the front-end interactive page. When the interactive page is not clicked, the coordinate area displays "dx:--;dy:--;". If a certain position is clicked on the interactive page (such as clicking on the "Children's Area"), the coordinate area displays "X:527.5;Y:405.1", where "527.5, 405.1" is the position of the click operation. After clicking, you can enter the corresponding detail page to achieve subsequent interaction.

[0035] It should be noted that different target modes can be selected based on the different operating systems of the front-end devices, and the interactive page can be recorded in the target mode to obtain the video to be identified, so as to facilitate the subsequent analysis of the video to be identified based on the positioning marks in the target mode, and automatically identify the key frames that performed the triggering operation.

[0036] Step 104: Split the video to be identified into at least one video frame, and determine the positioning coordinates of the positioning mark in each video frame.

[0037] In actual implementation, you can choose a suitable tool or library to process the video to be identified. Commonly used tools include OpenCV, FFmpeg, etc. These tools provide rich functions to read and operate video data. Then, load the video to be identified from the front end, and capture images from the video frame to be identified as video frames according to the set frame rate, so as to decompose the original continuous video to be identified into a series of static video frames, one video frame is a static image, and then it is easy to identify and determine the positioning coordinates of the positioning mark in each video frame, so as to realize the automatic recognition of key frames from each video frame.

[0038] It should be noted that the form of the positioning mark is different in the target mode of different operating systems, and thus the specific implementation process of determining the positioning coordinates of the positioning mark in each video frame is also different. The positioning coordinates of the positioning mark in each video frame can be determined based on the specific form of the positioning mark in the target mode.

[0039] In an optional implementation of this embodiment, the target mode is a graphic-assisted touch mode of the first operating system, and the positioning mark is a touch graphic; determining the positioning coordinates of the positioning mark in each video frame includes: According to the shape of the touch pattern, a set shape detection algorithm is called to identify the touch pattern in each video frame; Determine the location coordinates of the touch graphic in each video frame.

[0040] The graphic-assisted touch mode refers to a mode in which a touch graphic is displayed in the interactive page to assist positioning through the touch graphic. In the iOS system, the touch graphic is a touch circle. The touch graphic is any shape used to identify the triggering operation in the interactive page. The shape detection algorithm is set to any algorithm that can detect a touch graphic of a specific shape in the video frame, such as setting the shape detection algorithm to Hough transform.

[0041] In actual implementation, when the target mode is the graphic-assisted touch mode of the first operating system, a touch graphic is displayed in the interactive page as a positioning mark. The user drags the touch graphic in the interactive page to perform a trigger operation in the interactive page. In order to identify the key frame in which the trigger operation is performed, the positioning coordinates of the touch graphic in each video frame can be located. Specifically, according to the shape of the touch graphic, a set shape detection algorithm can be called to identify the touch graphic in each video frame, and then the positioning coordinates of the touch graphic in each video frame can be determined.

[0042] As an example, taking the iOS system as an example, the touch graphic is a touch circle. At this time, each video frame can be converted from the image space to the parameter space through the Hough transform, and the center of the circle and its corresponding radius that meet the conditions are found and screened through the voting mechanism to identify the touch circle in each video frame, and then determine the position of the touch circle in each video frame. That is, through the Hough Circle Transform, the detection and recognition of the touch circle (circle) in each video frame is realized.

[0043] Specifically, video frame preprocessing can be performed before performing the Hough transform. If the current input video frame is in color, the video frame can be grayscaled and converted into a grayscale image to simplify subsequent processing. Then, an edge detection algorithm (such as the Canny operator) can be applied to highlight the edges in the video frame and enhance the edge features in the video frame. Alternatively, background noise in the video frame can be removed through a Gaussian denoising operation to reduce the impact of background noise and improve the accuracy of the Hough transform.

[0044] After that, the parameter space and accumulator can be initialized. For touch circle (i.e. circle) detection, the search range of the circle center (a, b) coordinates can be initialized. Specifically, a reasonable search range can be set according to the image size, and then the minimum and maximum possible radius values ​​r can be set based on the interactive page; create an accumulator and initialize a three-dimensional array as the accumulator, whose dimensions correspond to the (a, b, r) parameter space. The accumulator is used to record the number of times each possible circle parameter combination is voted.

[0045] The voting process is as follows: traverse the edge points, and for each non-zero pixel (i.e., edge point) after edge detection, vote in the parameter space; calculate the parameter value, for each edge point (x, y), assuming it is a point on a circle, then the point can belong to countless different circles. In order to find these possible circles, it is necessary to traverse all possible radius values ​​r and, according to (xa) 2 +(yb) 2 =r 2 , calculate the corresponding center position (a, b). For each fixed radius value r, a series of possible center position (a, b) combinations can be obtained by solving the equation, and the corresponding (a, b, r) position in the accumulator is counted.

[0046] After that, find the local maximum value and set a threshold. Only when the count of a cell in the accumulator exceeds this threshold, a circle is considered to exist. In order to remove redundant detection results, only the local maximum point is retained as the final detected circle parameter, which is non-maximum suppression. After that, the local maximum point found in the parameter space is converted back to the circle equation in the image space, and the detected circle is drawn on the original video frame, or the detection result is output in other forms.

[0047] In addition, you can also configure optimization rules based on actual application scenarios, such as merging similar circles, filtering out circle detection results that do not meet the conditions, etc. And if you are not sure about the size of the circle, you can run the Hough transform at different scales to ensure that no potential circles are missed.

[0048] In the actual implementation process, a small number of points can be randomly selected for voting to speed up the calculation, or the center position and radius range can be limited based on prior knowledge or context information to reduce unnecessary calculations, or the number of voting points can be gradually increased until sufficient evidence is found to support a hypothesis.

[0049] It should be noted that the above example uses a touch circle as the touch graphic and Hough transform detection as an example. Other shape detection algorithms can also be used to detect circular touch graphics, such as contour-based detection, template matching, morphological operations, machine learning and deep learning algorithms; of course, the graphic-assisted touch modes of different operating systems can also be configured with touch graphics of different shapes, such as rectangles, triangles or other irregular graphics, etc., and the embodiments of this specification do not limit this.

[0050] In the embodiment of the present specification, in the graphic-assisted touch mode of the first operating system, the touch graphic of a specific shape in each video frame can be identified by calling a set shape detection algorithm, thereby determining the positioning coordinates of the touch graphic in each video frame. Identifying the positioning coordinates of the touch graphic is simple, accurate and efficient. The touch graphic is used as a marker to reduce dependence on the content of the video frame, thereby ensuring the efficiency and accuracy of subsequent key frame recognition. The touch graphic is applicable to a variety of front-end devices of the first operating system, thereby improving compatibility.

[0051] In an optional implementation of this embodiment, the target mode is a pointer positioning mode of the second operating system, the positioning mark is a coordinate area of ​​the touch pointer, and the coordinate area displays the positioning coordinates of the touch pointer; determining the positioning coordinates of the positioning mark in each video frame includes: Identify text information in each video frame; The positioning coordinates displayed in the coordinate area are determined based on the text information.

[0052] Among them, the pointer positioning mode refers to the coordinate area of ​​the touch pointer displayed in the interactive page. The coordinate area can display the positioning coordinates of the touch pointer, that is, when the user performs a trigger operation on the interactive page, the coordinates of the operation can be displayed in the coordinate area. Under the Android system, the pointer positioning mode can be a developer mode provided by the Android system. When you click on the interactive page, the coordinate area of ​​the touch pointer displays the coordinates of the click position.

[0053] In actual implementation, since the coordinates displayed in the coordinate area of ​​the interactive page are text, the text information in each video frame can be recognized through OCR technology, and then the positioning coordinates displayed in the coordinate area are extracted from the text information. The positioning coordinates are the location where the trigger operation is performed in the interactive page.

[0054] As an example, taking the Android system as an example, the text information in each video frame is recognized by OCR technology. Specifically, before recognizing the text information, the video frame can be preprocessed. If the current input video frame is in color, the video frame can be grayscaled and converted into a grayscale image to simplify the subsequent processing process. Then, the edge detection algorithm (such as the Canny operator) can be applied to highlight the edges in the video frame and enhance the edge features in the video frame. Alternatively, the background noise in the video frame can be removed by a Gaussian denoising operation to reduce the impact of background noise and improve the accuracy of subsequent text recognition.

[0055] After identifying and obtaining text information, the text information of any video frame is first cleaned, such as removing unnecessary symbols, spaces and other interference factors, and standardizing the text information format for subsequent processing. Then, according to the expected coordinate representation (such as latitude and longitude, UTM coordinates, national grid reference, etc.), the possible coordinate format is determined. Common coordinate formats include "latitude, longitude", "X, Y", etc. Natural language processing (NLP) technology or regular expressions are used to identify potential coordinate strings in text information. Regular expressions are a powerful tool that can accurately match text patterns in a specific format; simple rules can be used to check whether the candidate coordinates found meet basic logic (for example, the latitude range should be between -90 and 90, and the longitude range should be between -180 and 180), and the matched strings are converted into numerical form to ensure that each coordinate component is correct. In the case of failure to successfully parse the positioning coordinates, error logs can be recorded and other methods can be tried for identification.

[0056] In the embodiments of the present specification, in the pointer positioning mode of the second operating system, text information in each video frame can be identified, and the positioning coordinates displayed in the coordinate area can be extracted from the text information. Subsequently, based on whether the positioning coordinates can be extracted, it can be determined whether a trigger operation has been performed in the current video frame. By identifying and extracting the positioning coordinates displayed in the coordinate area, automatic identification of key frames in which the trigger operation has been performed can be achieved. The method for determining the positioning coordinates of the positioning mark is simple, accurate and efficient. The touch visual feedback characteristics in the pointer positioning mode are used to detect key frames, which ensures the efficiency and accuracy of subsequent key frame identification. The system is applicable to a variety of front-end devices of the second operating system, thereby improving compatibility.

[0057] It should be noted that, for different operating systems, the trigger operations in the interactive page can be identified through different modes, so as to realize the coordinate recognition of the positioning mark in different modes, and then determine the key frames in each video frame where the trigger operation is performed. It can be compatible with different operating systems and can adapt to a variety of front-end devices.

[0058] Step 106: identifying key frames from each video frame according to the positioning coordinates of the positioning marks in each video frame and the set key frame coordinate constraints, wherein the key frame is a video frame for executing the trigger operation.

[0059] Specifically, the video frame corresponding to the moment when the trigger operation is performed is determined from each video frame, that is, the key frame.

[0060] In actual implementation, the set key frame coordinate constraints are rules satisfied by the positioning coordinates of the positioning mark in the target mode and when a trigger operation occurs. For example, in the graphics-assisted touch mode of the first operating system, the key frame coordinate constraints are that the coordinate change amplitude of the positioning coordinates satisfies the change amplitude constraint rules; in the pointer positioning mode of the second operating system, the key frame coordinate constraints are to determine the positioning coordinates displayed in the coordinate area from the text information, that is, the positioning coordinates can be successfully extracted.

[0061] It should be noted that if the positioning coordinates of the positioning mark in a certain video frame meet the set key frame coordinate constraints, it means that a trigger operation has occurred in the video frame. Therefore, based on the positioning coordinates of the positioning mark in each video frame and the set key frame coordinate constraints, the video frame corresponding to the positioning coordinates that meet the key frame coordinate constraints can be screened out and the video frame can be used as the key frame.

[0062] In an optional implementation of this embodiment, in the graphic assisted touch mode of the first operating system, the key frame coordinate constraint is that the coordinate change amplitude of the positioning coordinates satisfies the change amplitude constraint rule; according to the positioning coordinates of the positioning mark in each video frame and the set key frame coordinate constraint, the key frame is identified from each video frame, including: Determine the change range of the positioning coordinates of each video frame according to the positioning coordinates of the positioning marks in each video frame; The key frames whose positioning coordinate change amplitude satisfies the change amplitude constraint rule are identified from each video frame.

[0063] In actual implementation, in the graphic assisted touch mode of the first operating system, if no trigger operation is performed in the interactive page, the position of the positioning mark in the video frame remains unchanged or shakes slightly. If the user wants to perform a trigger operation in the interactive page, he needs to drag the touch graphic from the current position to the position to be triggered, so the coordinates of the touch graphic will change significantly.

[0064] It should be noted that the change amplitude constraint rule refers to the rule that the change amplitude of the positioning coordinates needs to meet. For example, the change amplitude constraint rule means that the change amplitude of the positioning coordinates of the current video frame relative to the positioning coordinates of the previous video frame is greater than the set threshold, or the change amplitude constraint rule is the first coordinate that stabilizes after the positioning coordinates change significantly.

[0065] In specific implementation, the positioning coordinate change amplitude of each video frame can be determined based on the positioning coordinates of the positioning marks in each video frame. The positioning coordinate change amplitude refers to the change amplitude of the positioning coordinates of the current video frame relative to the positioning coordinates of the previous video frame, and the key frames whose positioning coordinate change amplitude satisfies the change amplitude constraint rule are identified from each video frame.

[0066] For example, Figure 2cis a schematic diagram of a third interactive page provided by an embodiment of this specification, such as Figure 2c As shown, taking the interactive page of a medical electronic platform as an example, in the circle-assisted touch mode of the iOS system, a touch circle is displayed in the front-end interactive page. Initially, the touch circle is located in the lower right corner (i.e., the initial position of the touch circle). Assuming that the user drags the touch circle to the "Children's Zone" (the position after the touch circle is dragged), the "Children's Zone" in the interactive page can be clicked, and then the details page corresponding to the "Children's Zone" can be entered to realize subsequent interaction. In the embodiment of this specification, the video frame at the moment of clicking the "Children's Zone" can be identified as a key frame.

[0067] In the embodiments of the present specification, the change amplitude of the positioning coordinates of each video frame is determined based on the positioning coordinates of the positioning mark in each video frame, and then the key frames whose change amplitudes satisfy the change amplitude constraint rules are screened out. The change amplitude of the positioning coordinates in each video frame is used to indicate whether the user has dragged the touch graphic in each video frame relative to the previous video frame, thereby identifying the key frames that have performed the triggering operation, realizing automatic recognition of key frames, ensuring the recognition efficiency and accuracy of key frames, and can be applied to front-end devices of various first operating systems with high compatibility.

[0068] In an optional implementation of this embodiment, determining the change range of the positioning coordinates of each video frame according to the positioning coordinates of the positioning mark in each video frame includes: Generate a coordinate change curve of the positioning mark according to the positioning coordinates of the positioning mark in each video frame, wherein the horizontal axis of the coordinate change curve is the frame index of each video frame, and the vertical axis is the positioning coordinates of the positioning mark in the corresponding video frame; According to the coordinate change curve of the positioning mark, the change range of the positioning coordinates of each video frame is determined.

[0069] In actual implementation, the frame index of each video frame can be used as the horizontal axis, and the positioning coordinates of the positioning mark in each video frame can be used as the vertical axis to generate a coordinate change curve of the positioning mark. The coordinate change curve of the positioning mark can indicate the transformation of the positioning coordinates, thereby determining the change range of the positioning coordinates of each video frame.

[0070] Specifically, the positioning coordinates are two-dimensional coordinates, so the changes of the horizontal coordinates and vertical coordinates of the positioning coordinates can be represented by two curves, that is, the frame index of each video frame is used as the horizontal coordinate, and the horizontal coordinate of the positioning mark in each video frame is used as the vertical axis to generate the horizontal coordinate change curve of the positioning mark; and the frame index of each video frame is used as the horizontal coordinate, and the vertical coordinate of the positioning mark in each video frame is used as the vertical axis to generate the vertical coordinate change curve of the positioning mark. Alternatively, the change mean of the horizontal coordinate and the vertical coordinate of the positioning coordinate can be determined, and the change mean is used as the vertical axis, and the frame index of each video frame is used as the horizontal coordinate to generate the corresponding coordinate change curve.

[0071] For example, Figure 2d is a schematic diagram of a coordinate change curve of a customized mark provided in one embodiment of this specification, such as Figure 2d As shown, the coordinate change curve includes a horizontal coordinate change curve and a vertical coordinate change curve. The horizontal coordinate change curve represents the change range of the horizontal coordinate of the positioning mark in each video frame, and the vertical coordinate change curve represents the change range of the vertical coordinate of the positioning mark in each video frame.

[0072] In the embodiments of the present specification, the video frames whose change amplitudes of the horizontal and vertical coordinates of the positioning mark are greater than the set threshold (the change amplitudes are large) can be determined as key frames according to the generated coordinate change curve, and the average change amplitudes of the horizontal and vertical coordinates can be determined, and the video frames whose average change amplitudes are greater than the set threshold (the change amplitudes are large) can be used as key frames. In this way, the coordinate change curve can intuitively show the change amplitude of the positioning coordinates of each video frame, so that it is convenient to select video frames with large change amplitudes as key frames based on the coordinate change curve, and quickly and accurately identify key frames.

[0073] In an optional implementation of this embodiment, the change amplitude constraint rule is that the change amplitude of the positioning coordinates changes from an increasing trend to a stable trend; identifying key frames whose change amplitude of the positioning coordinates satisfies the change amplitude constraint rule from each video frame includes: Determine a target coordinate point in the coordinate change curve where the coordinate change amplitude changes from an increasing trend to a stable trend, wherein the increasing trend means that the coordinate change amplitude of the current coordinate point relative to the previous coordinate point is greater than a first amplitude threshold, and the stable trend means that the coordinate change amplitude of the current coordinate point relative to the previous coordinate point is less than a second amplitude threshold; The frame index corresponding to the target coordinate point is determined, and the video frame corresponding to the frame index is determined as a key frame.

[0074] In one implementation, the change amplitude constraint rule is that the change amplitude of the positioning coordinate changes from an increasing trend to a stable trend. The increasing trend means that the coordinate change amplitude of the current coordinate point relative to the previous coordinate point is greater than the first amplitude threshold, and the stable trend means that the coordinate change amplitude of the current coordinate point relative to the previous coordinate point is less than the second amplitude threshold. In other words, the change amplitude constraint rule is the first stable coordinate after the positioning coordinate changes significantly.

[0075] In actual implementation, the target coordinate point where the coordinate change amplitude changes from an increasing trend to a stable trend can be located from the coordinate change curve, and the horizontal axis value corresponding to the target coordinate point can be determined. The horizontal axis value is the frame index, and the video frame corresponding to the frame index is determined as the key frame.

[0076] If the coordinate change curve is a coordinate change curve corresponding to the change mean value of the horizontal and vertical coordinates of the positioning coordinates, then the target coordinate point where the vertical axis value in the coordinate change curve changes from an increasing trend to a stable trend can be determined, and then the corresponding frame index can be determined; if the coordinate change curve includes a horizontal coordinate change curve and a vertical coordinate change curve, then the target coordinate point where the vertical axis value changes from an increasing trend to a stable trend can be determined for the horizontal coordinate change curve and the vertical coordinate change curve respectively, and then the corresponding frame index can be determined. If the frame indexes corresponding to the two are consistent, the video frame corresponding to the frame index is used as the key frame. If the frame indexes corresponding to the two are inconsistent, then the video frames corresponding to the two can both be used as the key frames, and other image processing technologies can be further combined to further identify and analyze the two video frames to determine the key frame that performs the trigger operation.

[0077] Using the above example, Figure 2d As shown, when the frame index is less than 69, the horizontal and vertical coordinates of the center of the touch circle (i.e., the positioning mark) are both stable at one position. Starting from the 69th frame, the change amplitude of the horizontal and vertical coordinates of the center of the touch circle begins to increase, that is, the change amplitude of the positioning coordinates is increasing. After the 72nd frame, the horizontal and vertical coordinates of the center of the circle return to stability, that is, the change amplitude of the positioning coordinates changes from an increasing trend to a stable trend. In other words, the coordinate point corresponding to the 72nd frame is the first coordinate point that stabilizes after the horizontal and vertical coordinates of the center of the touch circle change significantly. The coordinate point corresponding to the 72nd frame is the target coordinate point, and the corresponding frame index is "72". At this time, the "72nd frame" can be determined as a key frame. Through the above method, each key frame in the video to be identified that performs the trigger operation can be determined.

[0078] In the embodiments of the present specification, edge features can be enhanced through image preprocessing, and positioning accuracy and robustness can be improved. For each video frame after image preprocessing, the touch circle in each video frame can be detected by Hough transform circle detection to obtain the positioning coordinates of the touch circle, and generate a corresponding coordinate change curve. The target coordinate point can be determined by the change trend in the coordinate change curve, and then the corresponding frame index can be determined to identify the key frame. Based on the coordinate change curve, the coordinate trend analysis can be quickly performed to quickly and accurately identify the key frame.

[0079] In an optional implementation of this embodiment, in the pointer positioning mode of the second operating system, the key frame coordinate constraint is to determine the positioning coordinates displayed in the coordinate area from the text information; according to the positioning coordinates of the positioning mark in each video frame and the set key frame coordinate constraint, the key frame is identified from each video frame, including: If the positioning coordinates displayed in the coordinate area are successfully determined from the first text information, the first video frame corresponding to the first text information is used as a key frame; If the positioning coordinates displayed in the coordinate area are not determined from the text information of each video frame, the key frame is identified based on the similarity between each video frame and the adjacent video frame.

[0080] In actual implementation, in the pointer positioning mode of the second operating system, if no trigger operation is performed in the interactive page, the coordinate area does not display the positioning coordinates, and the coordinate value of the positioning coordinates is empty, that is, the positioning coordinates displayed in the coordinate area cannot be determined from the text information; if the user performs an execution operation in the interactive page, the coordinate area can display the position coordinates of the execution of the trigger operation, that is, the coordinate value of the positioning coordinates is not empty, and the positioning coordinates displayed in the coordinate area can be successfully determined from the text information. Therefore, in the pointer positioning mode of the second operating system, the key frame coordinate constraints can be configured to successfully determine the positioning coordinates displayed in the coordinate area from the text information.

[0081] It should be noted that after performing text recognition on each video frame to obtain the text information corresponding to each video frame, the positioning coordinates of the positioning mark can be extracted from the text information. If the positioning coordinates displayed in the coordinate area can be successfully determined from the first text information, the user performs a trigger operation in the first video frame corresponding to the first text information, and the coordinate area displays the corresponding position coordinates. At this time, the first video frame can be used as a key frame. If the positioning coordinates displayed in the coordinate area are not determined from the text information of each video frame, it means that the coordinate area may be blocked or subject to other interference, and the accurate coordinate value cannot be extracted. Therefore, at this time, the similarity between each video frame and the adjacent video frame can be further determined to identify the key frame.

[0082] In the embodiments of the present specification, if the positioning coordinates displayed in the coordinate area can be successfully determined, the corresponding video frame is directly used as the key frame. Taking into account that in some cases the coordinate area overlaps with the background image due to differences in device models or specific layouts, affecting the accuracy of text recognition, and the positioning coordinates displayed in the coordinate area cannot be determined from the text information of each video frame, the similarity between each video frame and the adjacent video frame can be further introduced as a supplement to ensure that the key frame in each video frame can be successfully identified, avoiding the key frame recognition error caused by the text recognition error, and further ensuring the accuracy of the key frame recognition.

[0083] In an optional implementation of this embodiment, identifying key frames according to the similarity between each video frame and adjacent video frames includes: Calculate the similarity between each video frame and adjacent video frames; When the similarity between the second video frame and the adjacent video frame is lower than the similarity threshold, the second video frame and the latter video frame of the adjacent video frame are determined as a key frame.

[0084] In actual implementation, the similarity between each video frame and the adjacent video frame can be SSIM, and the SSIM algorithm can evaluate the similarity between two adjacent video frames from three dimensions: brightness, contrast, and structure. Of course, other similarity algorithms can also be used, such as mean square error, peak signal-to-noise ratio, etc., which are not limited in this embodiment of the specification.

[0085] Among them, the adjacent video frame can be the previous video frame of the current video frame, or it can be the next video frame of the current video frame. It is sufficient to calculate the similarity between the current video frame and the adjacent video frame, that is, to analyze whether the picture content of the current video frame has changed relative to the previous video frame or the next video frame.

[0086] Specifically, taking SSIM as an example, image preprocessing can be performed first to confirm whether the two video frames to be compared have the same size, resolution and color mode. If they are inconsistent, one or both video frames need to be scaled, cropped, etc. so that they can directly correspond in subsequent calculations. Although SSIM can be applied to color images, in order to simplify the calculation, the two video frames can be converted to grayscale images to reduce the data dimension of each pixel while retaining the brightness information.

[0087] SSIM evaluates the overall quality based on the similarity of local areas. Therefore, a sliding window of a fixed size (such as 11x11 pixels) can be defined. The window will be moved across the entire video frame according to the set step size to calculate the similarity region by region. The set step size refers to the distance the window moves each time (i.e., the step size). Normally, the step size is set to 1, which means that the window will cover every possible position to ensure that no details are missed.

[0088] For each pair of overlapping window positions, the average brightness, variance, and covariance of the corresponding areas in the two video frames are calculated respectively. These statistics reflect the local brightness, contrast, and structural characteristics of the image. The brightness similarity is calculated using the average brightness to measure the closeness of the average brightness of the two regions; the contrast similarity is calculated using the variance to evaluate the consistency of the contrast of the two regions; the structural similarity is calculated by the covariance to capture the correlation of the structural information between the two regions. Finally, the above three similarities are combined and the final SSIM value is calculated in the form of a weighted product, as shown in the following formula (1): (1) Where, refers to the SSIM value of two regions in the current video frame and the adjacent video frame; It refers to the brightness similarity; It refers to the contrast similarity; It refers to the structural similarity; , and is the weight parameter and can be set to 1 by default.

[0089] For the entire video frame, a global SSIM representing the similarity of the entire video frame can be obtained by averaging the SSIM values ​​of all local window positions.

[0090] It should be noted that if no trigger operation is performed on the interactive page, the coordinate area does not display the positioning coordinates, and the coordinate value of the positioning coordinates is empty, that is, the positioning coordinates displayed in the coordinate area cannot be determined from the text information; if the user performs an execution operation on the interactive page, the coordinate area can display the position coordinates of the execution trigger operation, that is, the coordinate value of the positioning coordinates is not empty, and the positioning coordinates displayed in the coordinate area can be successfully determined from the text information. In actual implementation, if the trigger operation is not performed in the interactive page, the content displayed in the coordinate area of ​​the two adjacent video frames is consistent, and the similarity of the two adjacent video frames is close to 100%; if the trigger operation is performed in the interactive page, the content displayed in the coordinate area of ​​the two adjacent video frames will change, and the similarity of the two adjacent video frames will decrease. Therefore, when the similarity between the second video frame and the adjacent video frame is lower than the similarity threshold, the second video frame and the latter video frame of the adjacent video frame can be determined as a key frame, and the adjacent video frame whose picture content changes between the two adjacent video frames can be identified, and the latter video frame is the key frame that performs the trigger operation.

[0091] In the embodiments of the present specification, the similarity between each video frame and the adjacent video frame can be calculated. If the similarity between the two is lower than the similarity threshold, it means that the content of the two video frames has changed, and the second video frame and the latter video frame of the adjacent video frames are determined to be key frames. Similarity recognition key frames are further introduced, and text recognition technology and similarity calculation technology are integrated. The coordinate values ​​and image differences are read to determine whether a trigger action is executed in the video frame, thereby improving the recognition accuracy of key frames.

[0092] In an optional implementation of this embodiment, identifying text information in each video frame includes: Determine the interference area in the interactive page based on the interactive scenario to which the interactive page belongs; Crop the interference area in each video frame to obtain the area to be identified in each video frame; Perform text recognition on the to-be-recognized area of ​​each video frame to obtain corresponding text information.

[0093] It should be noted that in different interactive scenarios, the interactive page may include interference information, which will interfere with the recognition of text information and the calculation of similarity. Therefore, based on the interactive scenario to which the interactive page belongs, the interference area in the interactive page can be determined, the interference area in each video frame can be cropped, the area to be recognized in each video frame can be obtained, and then text recognition can be performed on the area to be recognized in each video frame to obtain the corresponding text information.

[0094] In actual implementation, the interference area can be other areas containing scrolling information in the interactive page except the coordinate area. Take the search page in a certain medical electronic platform as an example. The coordinate area is above the search box. The recommended search information is scrolled in the search box, and the recommended drugs can be scrolled below the search box. The search box and the area below may interfere with the recognition of text information and similarity calculation results due to the presence of scrolling information. Therefore, the search box and the area below can be used as an interference area. The area to be identified can be obtained by cutting out the interference area. The area to be identified includes the coordinate area of ​​the positioning coordinates. Then, text recognition and similarity calculation are performed on the area to be identified.

[0095] Specifically, for special scenarios with interference factors such as a scrolling search box, the unique pixel features of the search box can be identified to locate the position of the search box, and the specific vertical coordinate of the upper boundary of the search box can be determined. Based on this, the search box and the area below can be cropped out to obtain the area to be identified, thereby effectively eliminating invalid interference caused by the scrolling of recommended search content in the search box and recommended products in the area below the search box.

[0096] Using the above example, Figure 2bAs shown, the interactive page is a search page in a pharmaceutical electronic platform. The interactive page includes a search box. The recommended search content displayed in the search box can scroll and change. The drugs corresponding to the "limited time low price" activity are displayed below the search box. The drugs included in the "limited time low price" will scroll and change. The upper boundary ordinate of the search box on the interactive page is identified, and the area below the upper boundary ordinate is used as the interference area. The interference area in the interactive page is cropped to obtain the following Figure 2e The area to be identified is shown in Figure 2e It is a schematic diagram of a region to be identified in a video frame provided by an embodiment of the present specification.

[0097] In the embodiments of the present specification, based on the interactive scene to which the interactive page belongs, an interference area that may interfere with text information recognition and similarity calculation can be determined. After cropping the interference area, the area to be recognized is obtained, and then text recognition, similarity calculation and other operations are performed on the area to be recognized of each video frame. In the corresponding interactive scene, the interference caused by the interference area is effectively eliminated, and the customization of the interactive scene is realized. In addition, a dynamic cropping strategy is adopted to intelligently adjust the video frame, eliminate interference factors in complex page layouts, and ensure the recognition accuracy of key frames.

[0098] In an optional implementation of this embodiment, after identifying the key frames from each video frame according to the positioning coordinates of the positioning marks in each video frame and the set key frame coordinate constraints, the method further includes: Determine the rendering duration of the content display frame according to the key frame and the content display frame after the key frame; Determine the front-end performance test results based on rendering time.

[0099] Among them, the key frame is the video frame on which the trigger operation is executed, that is, the key frame can indicate the moment when the trigger operation is executed; the content display frame after the key frame refers to the video frame after which the rendering of the page to be displayed by the trigger operation is completed, that is, the content display frame can indicate the moment when the rendering of the page corresponding to the trigger operation is completed. For example, the key frame is the video frame in which the "Children's Zone" is clicked, and the content display frame is the video frame that displays the "Children's Zone" page.

[0100] In actual implementation, the time interval between two video frames can be determined based on the key frame and the content display frame after the key frame, thereby determining the rendering duration of the content display frame and obtaining the front-end performance test result.

[0101] It should be noted that on the front-end performance detection tool, after uploading the video to be identified, the key frames can be automatically identified, and then combined with the content display frame after the key frame, the time interval between the two video frames can be determined. The rendering time of the content display frame can be automatically calculated to realize automatic detection of front-end performance.

[0102] The embodiments of the present specification provide a key frame recognition method, in which an interactive page is recorded in a target mode to obtain a video to be recognized. In the target mode, a positioning mark can be displayed in the interactive page. In the target mode, the trigger operation of the interactive page is identified by the positioning mark. Subsequently, the positioning coordinates of the positioning mark in each video frame can be identified to automatically identify the key frames that have executed the trigger operation from each video frame. The key frames that have executed the trigger operation in each video frame are automatically identified based on the positioning mark in the target mode. No human input is required, which greatly reduces the time consumption of key frame recognition and improves the recognition efficiency. It does not rely on human subjective judgment conditions and improves the accuracy of key frame recognition.

[0103] The following combination Figure 3 , taking the application of the key frame recognition method provided in this specification in the medical electronic platform scenario as an example, the key frame recognition method is further explained. Among them, Figure 3 A process flow chart of a key frame recognition method provided by an embodiment of the present specification is shown, which specifically includes the following steps.

[0104] Step 302: For the front end of the first operating system, control entering the graphic-assisted touch mode; in the graphic-assisted touch mode, perform a trigger operation by dragging the touch graphic displayed in the interactive page of the medical electronic platform to the corresponding position. In the graphic-assisted touch mode, record the trigger operation by dragging the touch graphic to the corresponding position until the entire process of rendering the page corresponding to the trigger operation to be displayed is completed, and obtain the corresponding video to be identified.

[0105] Step 304: For the front end of the second operating system, control entering the pointer positioning mode; in the pointer positioning mode, a coordinate area of ​​the touch pointer is displayed in the interactive page of the medical electronic platform. If the trigger operation is not performed in the interactive page, the positioning coordinates displayed in the coordinate area of ​​the touch pointer are empty. If the trigger operation is performed in the interactive page, the positioning coordinates displayed in the coordinate area of ​​the touch pointer are the position coordinates of the trigger operation; in the pointer positioning mode, record the entire process of clicking on a trigger operation position in the interactive page until the trigger operation corresponds to the page that needs to be displayed. The page rendering is completed to obtain the corresponding video to be identified.

[0106] Step 306: The front-end performance testing tool reads the video to be identified recorded by the front-end.

[0107] Step 308: If the video to be identified is a video recorded by the front end of the first operating system in the graphic-assisted touch mode, call a set shape detection algorithm according to the shape of the touch graphic, identify the touch graphic in each video frame, and determine the positioning coordinates of the touch graphic in each video frame.

[0108] Step 310: Generate a coordinate change curve of the touch pattern according to the positioning coordinates of the touch pattern in each video frame, determine the first stable coordinate point after a large change based on the coordinate change curve, and use the video frame corresponding to the frame index of its horizontal axis as the key frame.

[0109] Step 312: If the video to be identified is a video recorded by the front end of the second operating system in the pointer positioning mode, based on the scene characteristics of the medical electronic platform, the interference area of ​​the interactive page in the medical electronic platform is determined, and the interference area in each video frame is cropped to obtain the area to be identified in each video frame.

[0110] Step 314: Perform text recognition on the to-be-recognized area of ​​each video frame to obtain corresponding text information; determine the positioning coordinates displayed in the coordinate area based on the text information.

[0111] Step 316: Determine whether the positioning coordinates displayed in the coordinate area are successfully determined. If so, execute the following step 318; if not, execute the following step 320.

[0112] Step 318: The first video frame whose positioning coordinates are successfully determined is used as a key frame.

[0113] Step 320: Calculate the similarity between each video frame and the to-be-identified region of the adjacent video frame. When the similarity between the second video frame and the to-be-identified region of the adjacent video frame is lower than a similarity threshold, determine that the second video frame and the latter video frame of the adjacent video frame are key frames.

[0114] The embodiments of this specification provide a key frame recognition method. The front ends of different operating systems adopt different modes and different forms of positioning marks to realize coordinate recognition of positioning marks in different modes, and then determine the key frames that perform trigger operations in each video frame recorded for the interactive page in the medical electronic platform. It can be compatible with different front-end operating systems and can adapt to a variety of front-end devices. It can automatically identify the key frames that perform trigger operations without the need for human input, greatly reducing the time consumption of key frame recognition and improving recognition efficiency. It does not rely on human subjective judgment conditions and improves the accuracy of key frame recognition.

[0115] Corresponding to the above method embodiment, this specification also provides a key frame identification device embodiment, Figure 4 FIG. 2 shows a schematic diagram of a key frame recognition device provided by an embodiment of the present specification. Figure 4 As shown, the device comprises: The acquisition module 402 is configured to acquire a video to be identified, wherein the video to be identified is a video recorded on an interactive page in a target mode, the target mode is a mode in which a positioning mark is displayed on the interactive page, and a triggering operation of the interactive page is identified by the positioning mark in the target mode; The determination module 404 is configured to split the video to be identified into at least one video frame and determine the positioning coordinates of the positioning mark in each video frame; The identification module 406 is configured to identify key frames from each video frame according to the positioning coordinates of the positioning marks in each video frame and the set key frame coordinate constraints, wherein the key frame is a video frame for performing a trigger operation.

[0116] Optionally, the target mode is a graphic-assisted touch mode of the first operating system, and the positioning mark is a touch graphic; the determination module 404 is further configured to: According to the shape of the touch pattern, a set shape detection algorithm is called to identify the touch pattern in each video frame; Determine the location coordinates of the touch graphic in each video frame.

[0117] Optionally, the key frame coordinate constraint is that the coordinate change range of the positioning coordinates satisfies the change range constraint rule; the identification module 406 is further configured to: Determine the change range of the positioning coordinates of each video frame according to the positioning coordinates of the positioning marks in each video frame; The key frames whose positioning coordinate change amplitude satisfies the change amplitude constraint rule are identified from each video frame.

[0118] Optionally, the identification module 406 is further configured to: Generate a coordinate change curve of the positioning mark according to the positioning coordinates of the positioning mark in each video frame, wherein the horizontal axis of the coordinate change curve is the frame index of each video frame, and the vertical axis is the positioning coordinates of the positioning mark in the corresponding video frame; According to the coordinate change curve of the positioning mark, the change range of the positioning coordinates of each video frame is determined.

[0119] Optionally, the change amplitude constraint rule is that the change amplitude of the positioning coordinate changes from an increasing trend to a stable trend; the identification module 406 is further configured to: Determine a target coordinate point in the coordinate change curve where the coordinate change amplitude changes from an increasing trend to a stable trend, wherein the increasing trend means that the coordinate change amplitude of the current coordinate point relative to the previous coordinate point is greater than a first amplitude threshold, and the stable trend means that the coordinate change amplitude of the current coordinate point relative to the previous coordinate point is less than a second amplitude threshold; The frame index corresponding to the target coordinate point is determined, and the video frame corresponding to the frame index is determined as a key frame.

[0120] Optionally, the target mode is a pointer positioning mode of the second operating system, the positioning mark is a coordinate area of ​​the touch pointer, and the coordinate area displays the positioning coordinates of the touch pointer; the determination module 404 is further configured to: Identify text information in each video frame; The positioning coordinates displayed in the coordinate area are determined based on the text information.

[0121] Optionally, the key frame coordinate constraint is to determine the positioning coordinates displayed in the coordinate area from the text information; the identification module 406 is further configured to: If the positioning coordinates displayed in the coordinate area are successfully determined from the first text information, the first video frame corresponding to the first text information is used as a key frame; If the positioning coordinates displayed in the coordinate area are not determined from the text information of each video frame, the key frame is identified based on the similarity between each video frame and the adjacent video frame.

[0122] Optionally, the identification module 406 is further configured to: Calculate the similarity between each video frame and adjacent video frames; When the similarity between the second video frame and the adjacent video frame is lower than the similarity threshold, the second video frame and the latter video frame of the adjacent video frame are determined as the key frame.

[0123] Optionally, the determination module 404 is further configured to: Determine the interference area in the interactive page based on the interactive scenario to which the interactive page belongs; Crop the interference area in each video frame to obtain the area to be identified in each video frame; Perform text recognition on the to-be-recognized area of ​​each video frame to obtain corresponding text information.

[0124] Optionally, the device also includes a performance testing module: Determine the rendering duration of the content display frame according to the key frame and the content display frame after the key frame; Determine the front-end performance test results based on rendering time.

[0125] The embodiments of the present specification provide a key frame recognition device, which records an interactive page in a target mode to obtain a video to be recognized. In the target mode, a positioning mark can be displayed in the interactive page. In the target mode, the trigger operation of the interactive page is identified by the positioning mark. Subsequently, the positioning coordinates of the positioning mark in each video frame can be identified to automatically identify the key frames that have executed the trigger operation from each video frame. The key frames that have executed the trigger operation in each video frame are automatically identified based on the positioning mark in the target mode. No human input is required, which greatly reduces the time consumption of key frame recognition and improves the recognition efficiency. It does not rely on human subjective judgment conditions and improves the accuracy of key frame recognition.

[0126] The above is a schematic scheme of a key frame recognition device of this embodiment. It should be noted that the technical scheme of the key frame recognition device and the technical scheme of the key frame recognition method described above are of the same concept, and the details not described in detail in the technical scheme of the key frame recognition device can be found in the description of the technical scheme of the key frame recognition method described above.

[0127] Figure 5 The block diagram of a computing device according to an embodiment of the present specification is shown. The components of the computing device 500 include but are not limited to a memory 510 and a processor 520. The processor 520 is connected to the memory 510 via a bus 530, and the database 550 is used to store data.

[0128] The computing device 500 also includes an access device 540 that enables the computing device 500 to communicate via one or more networks 560. Examples of these networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. The access device 540 may include one or more of any type of network interface (e.g., a network interface card (NIC)) that is wired or wireless, such as an IEEE 802.11 wireless local area network (WLAN) wireless interface, a world-wide interoperability for microwave access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, and a near field communication (NFC).

[0129] In one embodiment of the present specification, the above components of the computing device 500 and Figure 5 Other components not shown in the figure may also be connected to each other, for example, via a bus. It should be understood that Figure 5 The computing device structure block diagram shown is only for the purpose of illustration, and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.

[0130] The computing device 500 may be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smart phone), a wearable computing device (e.g., a smart watch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 500 may also be a mobile or stationary server.

[0131] The processor 520 is used to execute the following computer executable instructions, which implement the steps of the above-mentioned key frame recognition method when executed by the processor.

[0132] The above is a schematic scheme of a computing device of this embodiment. It should be noted that the technical scheme of the computing device and the technical scheme of the above key frame recognition method belong to the same concept, and the details not described in detail in the technical scheme of the computing device can be referred to the description of the technical scheme of the above key frame recognition method.

[0133] An embodiment of the present specification further provides a computer-readable storage medium storing computer-executable instructions, which can implement the steps of the above-mentioned key frame identification method when executed by a processor.

[0134] The above is a schematic scheme of a computer-readable storage medium of this embodiment. It should be noted that the technical scheme of the storage medium and the technical scheme of the key frame recognition method described above are of the same concept, and the details not described in detail in the technical scheme of the storage medium can be found in the description of the technical scheme of the key frame recognition method described above.

[0135] An embodiment of the present specification further provides a computer program product, including a computer program / instruction, which implements the steps of the above-mentioned key frame identification method when executed by a processor.

[0136] The above is a schematic solution of a computer program product of this embodiment. It should be noted that the technical solution of the computer program product and the technical solution of the key frame recognition method described above are of the same concept, and the details not described in detail in the technical solution of the computer program product can be found in the description of the technical solution of the key frame recognition method described above.

[0137] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0138] Computer instructions include computer program codes, which may be in source code form, object code form, executable files or some intermediate forms, etc. Computer readable media may include: any entity or device capable of carrying computer program codes, recording media, USB flash drives, mobile hard disks, magnetic disks, optical disks, computer memories, read-only memories (ROM), random access memories (RAM), electric carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the contents of computer readable media may be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, computer readable media do not include electric carrier signals and telecommunication signals.

[0139] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.

[0140] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0141] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The optional embodiments do not describe all the details in detail, nor do they limit the invention to only the specific implementation methods described. Obviously, many modifications and changes can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that technicians in the relevant technical field can well understand and use this specification. This specification is only limited by the claims and their full scope and equivalents.

Claims

1. A key frame recognition method, comprising: Acquire a video to be identified, wherein the video to be identified is a video recorded on an interactive page in a target mode, the target mode is a mode in which a positioning mark is displayed in the interactive page, and a triggering operation of the interactive page is identified by the positioning mark in the target mode; Splitting the to-be-identified video into at least one video frame, and determining the positioning coordinates of the positioning mark in each video frame; According to the positioning coordinates of the positioning marks in the video frames and the set key frame coordinate constraints, key frames are identified from the video frames, wherein the key frames are video frames for performing trigger operations.

2. According to the key frame recognition method of claim 1, the target mode is a graphic-assisted touch mode of the first operating system, and the positioning mark is a touch graphic; The determining of the positioning coordinates of the positioning mark in each video frame includes: According to the shape of the touch pattern, calling a set shape detection algorithm to identify the touch pattern in each video frame; Determine the location coordinates of the touch graphics in each video frame.

3. The key frame recognition method according to claim 2, wherein the key frame coordinate constraint is that the coordinate change amplitude of the positioning coordinates satisfies the change amplitude constraint rule; and the key frame is recognized from each video frame according to the positioning coordinates of the positioning marks in each video frame and the set key frame coordinate constraint, comprising: Determining a change range of the positioning coordinates of each video frame according to the positioning coordinates of the positioning marks in each video frame; A key frame whose positioning coordinate change range satisfies the change range constraint rule is identified from each video frame.

4. The key frame recognition method according to claim 3, wherein determining the change range of the positioning coordinates of each video frame according to the positioning coordinates of the positioning marks in each video frame comprises: According to the positioning coordinates of the positioning mark in each video frame, a coordinate change curve of the positioning mark is generated, wherein the horizontal axis of the coordinate change curve is the frame index of each video frame, and the vertical axis is the positioning coordinates of the positioning mark in the corresponding video frame; The change range of the positioning coordinates of each video frame is determined according to the coordinate change curve of the positioning mark.

5. The key frame identification method according to claim 4, wherein the variation range constraint rule is that the variation range of the positioning coordinates changes from an increasing trend to a stable trend; and the step of identifying the key frames whose variation range of the positioning coordinates satisfies the variation range constraint rule from the video frames comprises: Determine a target coordinate point in the coordinate change curve where the coordinate change amplitude changes from an increasing trend to a stable trend, wherein the increasing trend is that the coordinate change amplitude of the current coordinate point relative to the previous coordinate point is greater than a first amplitude threshold, and the stable trend is that the coordinate change amplitude of the current coordinate point relative to the previous coordinate point is less than a second amplitude threshold; A frame index corresponding to the target coordinate point is determined, and a video frame corresponding to the frame index is determined as the key frame.

6. The key frame recognition method according to claim 1, wherein the target mode is a pointer positioning mode of the second operating system, the positioning mark is a coordinate area of ​​the touch pointer, and the coordinate area displays the positioning coordinates of the touch pointer; The determining of the positioning coordinates of the positioning mark in each video frame includes: Identifying text information in each video frame; The positioning coordinates displayed in the coordinate area are determined based on the text information.

7. The key frame recognition method according to claim 6, wherein the key frame coordinate constraint is to determine the positioning coordinates displayed in the coordinate area from the text information; and the key frame is recognized from each video frame according to the positioning coordinates of the positioning mark in each video frame and the set key frame coordinate constraint, comprising: If the positioning coordinates displayed in the coordinate area are successfully determined from the first text information, the first video frame corresponding to the first text information is used as the key frame; If the positioning coordinates displayed in the coordinate area are not determined from the text information of each video frame, the key frame is identified according to the similarity between each video frame and an adjacent video frame.

8. The key frame recognition method according to claim 7, wherein the key frame is recognized based on the similarity between each video frame and an adjacent video frame, comprising: Calculating the similarity between each video frame and the adjacent video frames; When the similarity between the second video frame and the adjacent video frame is lower than a similarity threshold, the second video frame and the subsequent video frame of the adjacent video frame are determined to be the key frame.

9. The key frame recognition method according to claim 6, wherein the step of recognizing text information in each video frame comprises: Determining an interference area in the interactive page based on the interactive scenario to which the interactive page belongs; Cropping the interference area in each video frame to obtain the area to be identified in each video frame; Perform text recognition on the to-be-recognized area of ​​each video frame to obtain corresponding text information.

10. The key frame recognition method according to claim 1, after identifying the key frames from the video frames according to the positioning coordinates of the positioning marks in the video frames and the set key frame coordinate constraints, further comprising: Determining a rendering duration of the content display frame according to the key frame and a content display frame after the key frame; A front-end performance test result is determined based on the rendering duration.

11. A key frame recognition device, comprising: an acquisition module, configured to acquire a video to be identified, wherein the video to be identified is a video obtained by recording an interactive page in a target mode, the target mode is a mode in which a positioning mark is displayed in the interactive page, and a triggering operation of the interactive page is identified by the positioning mark in the target mode; A determination module is configured to split the video to be identified into at least one video frame and determine the positioning coordinates of the positioning mark in each video frame; The identification module is configured to identify key frames from the video frames according to the positioning coordinates of the positioning marks in the video frames and the set key frame coordinate constraints, wherein the key frames are video frames for performing trigger operations.

12. A computing device comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the key frame recognition method according to any one of claims 1 to 10 are implemented.

13. A computer-readable storage medium storing a computer program / instruction, wherein the computer program / instruction, when executed by a processor, implements the steps of the key frame recognition method according to any one of claims 1 to 10.

14. A computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the key frame recognition method according to any one of claims 1 to 10.