Reaction speed determination method, device, equipment and medium of point reading device
The video frames of the point-reading device are automatically extracted through image processing technology, which solves the problem of reaction speed evaluation in the existing technology that is time-consuming, labor-intensive and has large errors, and realizes efficient and accurate reaction speed evaluation.
Patent Information
- Application Number
- CN202110406305.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-15
- Publication Date
- 2025-10-17
- Estimated Expiration
- 2041-04-15
AI Technical Summary
In the prior art, the response speed evaluation of point-and-read devices is time-consuming and labor-intensive, and the evaluation results of different testers vary greatly, which easily introduces random errors.
By acquiring the video recording data of the point-reading device, the first trigger frame and the response frame are extracted using image processing technology, and the reaction speed of the point-reading device is automatically determined based on the similarity of adjacent video frames.
It realizes the fully automated evaluation of the response speed of the point-reading equipment, saves manpower input, improves the evaluation efficiency, and eliminates human errors.
Smart Images

Figure CN115220632B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of point reading, and in particular relates to a reaction speed determination method and device of a point reading device, equipment and medium. BACKGROUND
[0002] To realize the point reading function, the point reading device captures the process of the user specifying the target word through the camera carried by the point reading device, and displays the reaction interface matched with the target word on the screen of the point reading device.
[0003] In the related art, the reaction speed of the point reading device is evaluated in an artificial manner. First, a tester simulates the interaction process between the user and the point reading device, and records the interaction process. Then, the tester analyzes the recording. Finally, the tester determines the reaction speed of the point reading device according to experience.
[0004] Determining the reaction speed of the point reading device in the related art requires artificial evaluation multiple times, which not only consumes a lot of time and labor, but also has large differences in evaluation results obtained by different testers, which easily introduces random errors. SUMMARY
[0005] The present application provides a reaction speed determination method and device of a point reading device, equipment and medium, which can improve the efficiency of determining the reaction speed of the point reading device. The technical solution is as follows:
[0006] According to one aspect of the present application, a reaction speed determination method of a point reading device is provided, which comprises:
[0007] Obtaining a first video, the first video being obtained by recording the process in which the point reading device responds to a point reading operation;
[0008] Based on the similarity between adjacent video frames, a first trigger frame and a response frame of the first video are extracted from the first video, wherein the first trigger frame is a video frame in which the point reading device recognizes the point reading operation, and the response frame is a video frame in which the point reading device starts to respond to the point reading operation;
[0009] Based on the first trigger frame and the response frame, the reaction speed of the point reading device is determined.
[0010] According to one aspect of the present application, a reaction speed determination device of a point reading device is provided, which comprises:
[0011] An acquisition module for acquiring a first video, the first video carrying information of the reaction speed of the point reading device;
[0012] The processing module is configured to extract a first trigger frame and a response frame of the first video by performing image processing on the first video, wherein the first trigger frame is a video frame when the point reading device receives the point reading operation, and the response frame is a video frame when the point reading device starts to respond to the point reading operation.
[0013] The determining module is configured to determine the reaction speed of the point reading device based on the first trigger frame and the response frame.
[0014] According to an aspect of the present application, a computer device is provided, which comprises a processor and a memory, the memory storing a computer program, the computer program being loaded and executed by the processor to implement the reaction speed determination method of the point reading device as described above.
[0015] According to another aspect of the present application, a computer readable storage medium is provided, which stores a computer program, the computer program being loaded and executed by a processor to implement the reaction speed determination method of the point reading device as described above.
[0016] According to another aspect of the present application, a computer program product or computer program is provided, which comprises computer instructions stored in a computer readable storage medium. The processor of the computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions to enable the computer device to perform the reaction speed determination method of the point reading device as described above.
[0017] The technical scheme provided by the embodiments of the present application has at least the following beneficial effects:
[0018] By performing image processing on the first video recording the interaction process between the user and the point reading device, the terminal obtains a video frame when the point reading device receives the point reading operation and a video frame when the point reading device starts to respond to the point reading operation, and determines the reaction speed of the point reading device based on the two video frames. The reaction speed determination method of the point reading device adopts image processing technology, so that no manual intervention is required in determining the reaction speed of the point reading device, and full automation is achieved in data acquisition and analysis results, which not only saves manpower, but also greatly improves the efficiency of determining the reaction speed of the point reading device, and the finally determined reaction speed of the point reading device eliminates human-induced errors. BRIEF DESCRIPTION OF DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0020] Figure 1 is a schematic diagram of a reaction speed determination system of a point reading device provided by an example embodiment of the present application;
[0021] Figure 2 is a schematic diagram of a human-computer interaction process of a point reading device according to an example embodiment of the present application;
[0022] Figure 3 is a flowchart of a reaction speed determination method of a point reading device provided by an example embodiment of the present application;
[0023] Figure 4 is a schematic diagram of a first video frame provided by an example embodiment of the present application;
[0024] Figure 5 is a schematic diagram of a second video frame provided by an example embodiment of the present application;
[0025] Figure 6 is a schematic diagram of a second video frame provided by another example embodiment of the present application;
[0026] Figure 7 is a schematic diagram of a second video frame provided by another example embodiment of the present application;
[0027] Figure 8 is a flowchart of a process of obtaining a second video provided by an example embodiment of the present application;
[0028] Figure 9 is a flowchart of a process of recording a first video provided by an example embodiment of the present application;
[0029] Figure 10 is a flowchart of a reaction speed determination method of a point reading device provided by another example embodiment of the present application;
[0030] Figure 11 is a flowchart of a method of obtaining a first trigger frame of a first video provided by an example embodiment of the present application;
[0031] Figure 12 is a flowchart of a method of obtaining a response frame of a first video provided by an example embodiment of the present application;
[0032] Figure 13 is a structural block diagram of a reaction speed determination apparatus of a point reading device provided by an example embodiment of the present application;
[0033] Figure 14 is a structural block diagram of an electronic device provided by an example embodiment of the present application. DETAILED DESCRIPTION
[0034] In order to make the objectives, technical solutions and advantages of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.
[0035] First, a brief introduction to the terms involved in the embodiments of this application is given:
[0036] The first video: refers to a video used to determine the reaction speed of the point reading device. The first video is a recording of the process of the point reading device responding to the point reading operation, and the first video carries information on the reaction speed of the point reading device. In one embodiment, the first video frame includes a point reading operation area and a point reading response area. In one embodiment, the first video frame includes a point reading operation area, and optionally, the point reading operation area includes a point reading operation sub-area and a type calibration area. The type calibration area is an area that uses manually calibrated visual features in advance to indicate the video frame type of the current video frame. In one embodiment, the first video frame includes a point reading response area. The point reading response area is an area used to respond to the point reading operation. The point reading response area displays a response interface that matches the target word displayed on the screen of the point reading device.
[0037] The second video shows the user specifying the target word. In one embodiment, the second video is a recording of the user-computer interaction process of the point reading device, and a type identification area is pre-marked on the second video. In one embodiment, the second video frame includes the type identification area, and the type identification area is used to mark the second video frame.
[0038] First trigger frame: the video frame when the point reading device recognizes the point reading operation, that is, the first video frame in which the user points to the target word.
[0039] Second trigger frame: the second video frame in which the user clicks on the target word.
[0040] Response frame: This is the video frame when the point reading device begins responding to the point reading operation, that is, the first video frame when the response interface matching the target word is displayed on the point reading device screen. In one embodiment, when the user clicks the target word, the response interface matching the target word is displayed on the point reading device screen.
[0041] Type Marking Area: This area uses pre-defined visual features to indicate the video frame type of the current video frame. Video frame types include at least one of the preceding frame of the point-reading operation, the triggering frame of the point-reading operation, and the following frame of the point-reading operation. It's worth noting that in this application, both the first and second video frames have type marking areas. This is because the first video is a recording of the second video and the point-reading device's response to the point-reading operation played in the second video. In the following discussion, type marking areas are distinguished only by whether they exist in the first or second video frame.
[0042] Labeling operation: refers to adding a label to the video to be processed, reducing or increasing the similarity between the video frame to be extracted and the partial video frame. Here, the partial video frame can be selected as the previous frame of the video frame to be extracted or the next frame of the video frame to be extracted.
[0043] Image processing: a technology for analyzing images with a computer to achieve the desired results. Also known as image processing. Image processing generally refers to digital image processing. Digital image refers to a large two-dimensional array obtained by shooting with an industrial camera, a video camera, a scanner, etc. The elements of the array are called pixels, and the values are called gray values. Image processing techniques generally include image compression, enhancement and restoration, matching, description and identification of three parts.
[0044] Template matching: refers to matching an existing template image with a target image. By searching on the target image, the coordinates of the existing template image on the target image are determined.
[0045] Feature point matching: by comparing the feature points on the two images, the similarity of the two images is determined.
[0046] Artificial intelligence (AI) is to use digital computers or digital computer controlled machines to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology of computer science, which tries to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines, so that machines have the functions of perception, reasoning and decision-making.
[0047] Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, both hardware and software technologies. Artificial intelligence basic technology generally includes technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction system, mechatronics, etc. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology and machine learning / deep learning, etc.
[0048] Computer Vision (CV) is a science that studies how to make machines "see". More specifically, it refers to using cameras and computers to replace human eyes to identify, track and measure targets, and further process graphics so that computers can process images that are more suitable for human observation or transmitted to instruments for detection. As a scientific discipline, computer vision researches related theories and technologies, and attempts to establish artificial intelligence systems that can obtain information from images or multidimensional data. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, OCR (Optical Character Recognition), video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, and other technologies. It also includes common face recognition, fingerprint recognition and other biometric identification technologies.
[0049] Figure 1 is a reaction speed determination system of a point reading device according to an example embodiment of the present application, as shown in Figure 1 The reaction speed determination system 100 of the point reading device includes a second video generation system 101, a first video generation system 102 and an image processing system 103.
[0050] In response to inputting the to-be-processed video into the second video generation system 101, the second video generation system 101 outputs a second video.
[0051] In an embodiment, the to-be-processed video shows a process in which a user points at a target word. First, the second video generation system 101 frames the to-be-processed video to obtain a list of to-be-processed video frames. Then, in response to a calibration operation, the second video generation system 101 sets a second trigger frame in the list of to-be-processed video frames to display a first performance feature, a frame before the second trigger frame to display a second performance feature, and a frame after the second trigger frame to display the first performance feature. Finally, the second video generation system 101 replaces the second trigger frame, the frame before the second trigger frame and the frame after the second trigger frame in the list of to-be-processed video frames with the set second trigger frame, the set frame before the second trigger frame and the set frame after the second trigger frame in sequence, and generates the second video.
[0052] In response to the second video generation system 101 inputting the second video into the first video generation system 102, the first video generation system 102 outputs a first video.
[0053] In an embodiment, the first video generation system 102 controls a camera to record a human-computer interaction process of the point reading device, and takes the recorded video as the first video.
[0054] In one embodiment, first, the first video generation system 102 controls the display device to play the second video in the point reading operation area, wherein the second video is a video obtained by recording the human-computer interaction process of the point reading device, and the second video is pre-labeled with a type labeling area; then, the first video generation system 102 controls the camera to record the second video and the process that the point reading device responds to the point reading operation played in the second video, and takes the recorded video as the first video.
[0055] In response to the first video generation system 102 inputting the first video to the image processing system 103, the image processing system 103 outputs the reaction speed of the point reading device.
[0056] In one embodiment, the image processing system 103 first sets a timestamp for the first video to obtain the first video with the timestamp; then, the image processing system 103 performs a frame operation on the first video with the timestamp to obtain a first video frame list; then, based on the similarity of the point reading operation area in adjacent video frames, the image processing system 103 extracts the first trigger frame from the first video frame list, wherein the point reading operation area is an area for identifying the point reading operation; then, based on the similarity of the point reading response area in adjacent video frames, the image processing system 103 extracts the response frame from the first video frame list, wherein the point reading response area is an area for responding to the point reading operation; then, the image processing system 103 obtains the first timestamp of the first trigger frame and the second timestamp of the response frame; finally, based on the difference between the second timestamp and the first timestamp, the image processing system 103 determines the reaction speed of the point reading device.
[0057] Based on the above-mentioned second video generation system 101, the first video generation system 102 and the image processing system 103, the reaction speed determination system 100 of the point reading device outputs the reaction speed of the point reading device.
[0058] In one embodiment, the above-mentioned reaction speed determination system 100 of the point reading device can at least run on a terminal, or run on a server, or run on a terminal and a server.
[0059] Those skilled in the art can know that the number of the above-mentioned terminal and server can be more or less. For example, the above-mentioned terminal can be only one, or the above-mentioned terminal can be dozens or hundreds, or more. The above-mentioned server can be only one, or the above-mentioned server can be dozens or hundreds, or more. The number of terminals and the type of devices and the number of servers are not limited in the embodiments of the present application.
[0060] The following embodiments take the reaction speed determination system 100 of the point reading device applied to a terminal as an example for explanation and description.
[0061] In one embodiment, Figure 2 A schematic diagram of a human-computer interaction process of a point-reading device according to one example embodiment of the present application is shown.
[0062] In one embodiment, the point-reading device 220 is a smart task lamp, which optionally includes a camera, a screen and a base, wherein the camera is used to acquire a point-reading operation process of a user, the screen is used to display a response interface matched with the point-reading operation process, and the base is used to support the smart task lamp. It should be noted that the structure of the smart task lamp described above is only part of the structure of the smart task lamp related to the present application, and the smart task lamp in practice can have other structures to support other functions, such as a light bulb (supporting the lighting function), a pen box (supporting the storage function).
[0063] As shown in Figure 2 The point-reading device 220 includes a screen 221 of the point-reading device, a camera 222 of the point-reading device and a base 223 of the point-reading device, Figure 2 The word "test" 240 on the book and a user 260 using the point-reading device are also shown.
[0064] In response to the user 260 point-reading the word 240 in the book, the camera 222 of the point-reading device records a process video of the user 260 point-reading the word "test" 240 in the book. The point-reading device 220 receives the video recorded by the camera 222 of the point-reading device, processes and analyzes the video, extracts the word information in the video frame containing the word "test" 240, and then displays a response interface matched with the word "test" 240 on the screen 221 of the point-reading device. Optionally, the content of the response interface includes but is not limited to the target word, the pinyin of the target word, the definition, the example sentence.
[0065] Illustratively, the screen 221 of the point-reading device displays the pinyin "Ce Shi" of "test", the definition of "test" "Test is a measurement with experimental nature, that is, a combination of measurement and experiment. The test means is an instrument. Since test and measurement are closely related, they are often not strictly distinguished in actual use."
[0066] To improve the efficiency of determining the reaction speed of the point-reading device, Figure 3 A method for determining the reaction speed of a point-reading device according to one example embodiment of the present application is shown. Figure 3 The method shown is applied to a reaction speed determination system of a point-reading device, and the method includes:
[0067] Step 320, acquiring a first video;
[0068] The first video is obtained by recording a process in which the point reading device responds to the point reading operation, and carries information of a reaction speed of the point reading device.
[0069] In an embodiment, in response to the terminal controlling the camera to record the process of the human-computer interaction of the point reading device, the terminal takes the recorded video as the first video.
[0070] In an embodiment, first, the terminal controls the display device to play the second video in the point reading operation area, wherein the second video is a video obtained by recording the process of the human-computer interaction of the point reading device, and the second video is pre-marked with a type marking area; then, the terminal controls the camera to record the second video and a process in which the point reading device responds to the point reading operation played in the second video, and takes the recorded video as the first video.
[0071] In an embodiment, the second video refers to a video showing a process in which the user points to a target word.
[0072] Optionally, the terminal directly controls the display device to play the second video in the point reading operation area. Specifically, a code for controlling the display device is set in the terminal, and the control parameters of the display device are set by the code. Optionally, the control parameters of the display device include but are not limited to: starting to play the second video, stopping playing the second video, playing time, and playing frequency of the second video.
[0073] Optionally, the camera is connected to the terminal wirelessly or by wire, and the terminal directly controls the camera. Specifically, a code for controlling the camera is set in the terminal, and the control parameters of the camera are set by the code. Optionally, the control parameters of the camera include but are not limited to: opening, closing, aperture size, shutter speed, and sensitivity. It is worth noting that the camera mentioned herein includes a camera that can directly record and a peripheral camera of the terminal, and the present application does not limit the same.
[0074] Optionally, the camera is not connected to the terminal. In response to the camera recording the process of the human-computer interaction of the point reading device, the terminal takes the recording video of the camera as the first video.
[0075] In step 340, a first trigger frame and a response frame of the first video are extracted from the first video based on the similarity between adjacent video frames.
[0076] The first trigger frame is a video frame in which the point reading device recognizes the point reading operation, and the response frame is a video frame in which the point reading device starts to respond to the point reading operation.
[0077] In one embodiment, the pointing operation is an operation of pointing at a target word by a user. Optionally, the pointing operation includes, but is not limited to, finger pointing (a user specifies a target word by a finger), virtual pointing (e.g., a user controls a mouse to stay on a target word), and real object pointing (e.g., a user uses a pencil to specify a target word). In this application, the pointing operation is exemplified by finger pointing.
[0078] In one embodiment, step 340 can include the following steps:
[0079] First, a timestamp is set for the first video to obtain a first video with a timestamp;
[0080] The timestamp is used to mark the exact time point of the first video frame. Optionally, ffmpeg (a video preprocessing software) is used to add a timestamp to each first video frame.
[0081] In one embodiment, the terminal sets a timestamp for the first video using ffmpeg, and the terminal obtains the first video with a timestamp. Illustratively, the terminal sets the timestamp "00:00:06:769" at the upper right of the first video frame.
[0082] Second, the first video with a timestamp is subjected to a frame splitting operation to obtain a first video frame list;
[0083] In one embodiment, a video is obtained by refreshing a plurality of pictures at a certain frequency and in a certain order, and frame splitting is to extract the original constituent pictures from the video. Optionally, opencv (a frame splitting software) is used to split the first video. Optionally, the terminal splits the first video using ffmpeg.
[0084] Third, a first trigger frame is extracted from the first video frame list based on the similarity of the pointing operation area in adjacent video frames;
[0085] The pointing operation area is an area used to identify the pointing operation.
[0086] In one embodiment, the terminal extracts the first trigger frame from the first video frame list based on the similarity of the pointing operation area in adjacent video frames.
[0087] Optionally, the terminal obtains the first trigger frame based on the similarity between adjacent frames of the first video frame list. Specifically, the terminal performs feature point matching on the pointing operation area between adjacent frames of the first video frame list, and in response to the terminal determining that the similarity of the pointing operation area between adjacent frames reaches a threshold value, the terminal takes the current video frame as the first trigger frame.
[0088] Fourth, a response frame is extracted from the first video frame list based on the similarity of the pointing response area in adjacent video frames.
[0089] The point reading response area is an area used to respond to the point reading operation.
[0090] In one embodiment, based on the similarity of the click response areas in adjacent video frames, the terminal extracts the response frame from the first video frame list. The method of extracting the response frame is similar to that of extracting the first trigger frame and will not be repeated here.
[0091] Step 360: Determine the response speed of the point reading device based on the first trigger frame and the response frame.
[0092] In one embodiment, based on the first trigger frame and the response frame, the terminal determines the reaction speed of the point reading device.
[0093] In one embodiment, the terminal obtains a first timestamp of the first trigger frame and a second timestamp of the response frame; based on the difference between the second timestamp and the first timestamp, the terminal determines the reaction speed of the point reading device.
[0094] Illustratively, the second timestamp of the response frame is “00:00:06:769”, and the first timestamp of the first trigger frame is “00:00:04:566”, and the difference is “00:00:02:203”.
[0095] In summary, by performing image processing on a first video recording the interaction between a user and a point-reading device, the terminal obtains a video frame showing the point-reading device receiving a point-reading operation and a video frame showing the point-reading device beginning to respond to the point-reading operation. Based on these two video frames, the terminal determines the response speed of the point-reading device. The above method for determining the response speed of a point-reading device utilizes image processing technology, eliminating the need for human intervention in determining the response speed of the point-reading device. This fully automates data acquisition and analysis, saving manpower and significantly improving the efficiency of determining the response speed of the point-reading device. Furthermore, the ultimately determined response speed of the point-reading device eliminates errors introduced by humans.
[0096] In order to obtain the first trigger frame of the first video, based on Figure 3 In the optional embodiment shown, step 340 further includes the following steps:
[0097] The first video frame list includes N video frames.
[0098] Step 341 , calculating a first similarity between the touch-reading operation areas in the i-th frame and the i-1-th frame in the first video frame list, and a second similarity between the touch-reading operation areas in the i-th frame and the i+1-th frame in the first video frame list;
[0099] Wherein, N is a positive integer greater than 3, i is a positive integer not greater than N-2, and i is greater than or equal to 2.
[0100] In one embodiment, the pointing operation is an operation of a user pointing at a target word. Optionally, the pointing operation includes but is not limited to: finger pointing (the user specifies the target word by a finger), virtual pointing (such as the user controls a mouse to stay on the target word), real object pointing (such as the user uses a pencil to specify the target word). In this application, the pointing operation is exemplified by finger pointing.
[0101] In one embodiment, the terminal sequentially obtains the video frames in the first video frame list in ascending order, i.e., the initial value i = 2, and the value of i gradually increases.
[0102] In one embodiment, the terminal calculates the first similarity between the pointing operation region in the i-th frame and the i-1-th frame of the first video frame list, and the second similarity between the pointing operation region in the i-th frame and the i+1-th frame of the first video.
[0103] In one embodiment, the first video is obtained by recording the process of the pointing device responding to the pointing operation, and the first video carries information of the reaction speed of the pointing device; the video frames of the first video have a pointing operation region and a pointing response region, wherein the pointing operation region includes a pointing operation sub-region and a type marking region. The type marking region is a region that uses pre-labeled visual features to represent the type of the current video frame, i.e., the type marking region is used to mark the first video frame.
[0104] Illustratively, a video frame of the first video is as shown in Figure 4 Figure 4 The pointing operation sub-region 401, the type marking region 402, and the pointing response region 403 are shown. In response to playing the second video on the pointing operation sub-region 401, the pointing response region 403 displays the corresponding interface. In one embodiment, when the pointing operation sub-region 401 displays the target word pointed by the user, the pointing response region 403 displays the response interface matched with the target word. Optionally, the content of the response interface includes but is not limited to: the target word, the pinyin of the target word, the definition, and the example sentence.
[0105] In one embodiment, the terminal calculates the first similarity between the type marking region of the i-th frame and the type marking region of the i-1-th frame of the first video frame list, and the second similarity between the type marking region of the i-th frame and the type marking region of the i+1-th frame of the first video frame list.
[0106] Wherein, based on the known coordinates of the first video frame, the terminal obtains the type marking region by image processing, cropping, and calculation on the first video frame.
[0107] Step 342, in response to the first similarity being less than the first threshold value and the second similarity being greater than or equal to the first threshold value, determining the ith frame as the first trigger frame;
[0108] In one embodiment, in response to the first similarity being less than the first threshold value and the second similarity being greater than or equal to the first threshold value, the terminal determines the ith frame as the first trigger frame.
[0109] The first threshold value is a similarity threshold value of a type labeling region in the first trigger frame of the first person and a type labeling region of an adjacent frame.
[0110] In summary, by calculating the similarity of adjacent frames of the first video, when the similarity of the current frame and the previous frame is less than the first threshold value and the similarity of the current frame and the next frame is greater than or equal to the first threshold value, the current frame is determined as the first trigger frame of the first video. The above method realizes the full-automatic determination of the first trigger frame of the first video, without the need for human intervention, which not only saves manpower investment, but also greatly improves the efficiency of determining the first trigger frame.
[0111] To realize obtaining the response frame of the first video, based on Figure 3 In the optional embodiment shown in FIG. 34, step 340 further includes the following steps:
[0112] The first video frame list contains N video frames.
[0113] Step 343, calculating a third similarity between the point reading response region in the mth frame of the first video frame list and the (m-1)th frame, and a fourth similarity between the point reading response region in the mth frame of the first video frame list and the (m+1)th frame;
[0114] N is a positive integer greater than 3, m is a positive integer not greater than N-1, and m is greater than or equal to 3.
[0115] In one embodiment, the point reading operation is an operation of a user reading a target word. Optionally, the point reading operation includes but is not limited to: fingertip point reading (the user specifies a target word through a finger), virtual point reading (such as the user controlling a mouse to stay on a target word), and real object point reading (such as the user using a pencil to specify a target word). In this application, the point reading operation is exemplified by fingertip point reading.
[0116] In one embodiment, the terminal obtains the video frames in the first video frame list in reverse order, i.e., the initial value m=N-1, and the value of m gradually decreases.
[0117] In one embodiment, the terminal calculates a third similarity between the point reading response region in the mth frame of the first video frame list and the (m-1)th frame, and a fourth similarity between the mth frame of the first video frame list and the (m+1)th frame.
[0118] Wherein, based on the known coordinates of the first video frame, the terminal obtains the point reading response area through image processing, cropping, and calculation of the first video frame.
[0119] Step 344 , in response to the third similarity being less than the second threshold and the fourth similarity being greater than or equal to the second threshold, determining the mth frame as a response frame;
[0120] In one embodiment, in response to the third similarity being less than the second threshold and the fourth similarity being greater than or equal to the second threshold, the terminal determines that the mth frame is a response frame.
[0121] The second threshold is a similarity threshold between the response frame and the adjacent frames preset by the first person.
[0122] In summary, by calculating the similarity between adjacent frames of the first video, if the similarity between the previous frame and the current frame is less than a second threshold, and the similarity between the current frame and the next frame is greater than or equal to the second threshold, the current frame is determined to be the first trigger frame of the first video. This method achieves fully automatic determination of the response frame of the first video without manual intervention, not only saving manpower but also greatly improving the efficiency of determining the response frame.
[0123] In order to obtain the first video, based on Figure 3 In the embodiment shown, the second video is obtained in the following manner, that is, the above step 320 further includes the following steps:
[0124] S1: Divide the video to be processed into frames to obtain a list of video frames to be processed;
[0125] The video to be processed shows the process of the user clicking and reading the target word;
[0126] In one embodiment, the terminal divides the video to be processed into frames to obtain a list of video frames to be processed.
[0127] Optionally, the video to be processed can be obtained by obtaining existing videos from a video library; optionally, the video to be processed can be obtained by obtaining videos uploaded by users; optionally, the video to be processed can be obtained by the first person recording the user reading the target word. It is worth noting that during recording, the first person should control variables such as the shooting angle and scene lighting intensity of the video to be processed to be consistent with actual conditions, so as to ensure that the display effect of the video to be processed closely resembles the interaction process between the user and the reading device in real life.
[0128] S2: In response to the calibration operation, set the second trigger frame in the list of video frames to be processed to display the first expression feature, the previous frame of the second trigger frame to display the second expression feature, and the next frame of the second trigger frame to display the first expression feature.
[0129] wherein the marking operation refers to the first person adding a mark to the to-be-processed video, aiming to reduce or increase the similarity between the video frame to be extracted and the partial video frame. Here, the partial video frame can be selected as the frame before the video frame to be extracted or the frame after the video frame to be extracted.
[0130] In one embodiment, in response to the marking operation, the terminal sets the second trigger frame in the to-be-processed video frame list to display the first performance feature, the frame before the second trigger frame to display the second performance feature, and the frame after the second trigger frame to display the first performance feature.
[0131] In one embodiment, based on the experience of the same first person, the terminal identifies the second trigger frame on the to-be-processed video. The second trigger frame is the to-be-processed video frame in which the user points to read the target word.
[0132] In one embodiment, the second trigger frame has a type marking area, and the terminal sets the type marking area of the second trigger frame to display the first performance feature, the type marking area of the frame before the second trigger frame to display the second performance feature, and the type marking area of the frame after the second trigger frame to display the first performance feature.
[0133] Optionally, the performance feature includes at least one of a pattern performance feature, a color performance feature, and a text performance feature. For reference Figure 5 and Figure 6 The performance feature is a pattern performance feature. Optionally, Figure 5 Fig. 1 shows the first performance feature of the type marking area of the second trigger frame in one exemplary embodiment of the present application, Figure 6 Fig. 2 shows the second performance feature of the type marking area of the frame before the second trigger frame in one exemplary embodiment of the present application, Figure 7 Fig. 3 shows the first performance feature of the type marking area of the frame after the second trigger frame in one exemplary embodiment of the present application, wherein Figure 5 the first performance feature displayed by the type marking area 501 of the second trigger frame in Fig. 1 is a full-black rectangle, Figure 6 the first performance feature displayed by the type marking area 601 of the frame before the second trigger frame in Fig. 2 is a rectangle with diagonal lines, Figure 7 the first performance feature displayed by the type marking area 701 of the frame after the second trigger frame in Fig. 3 is a full-black rectangle.
[0134] S3: replace the second trigger frame, the frame before the second trigger frame, and the frame after the second trigger frame in the to-be-processed video frame list with the set second trigger frame, the set frame before the second trigger frame, and the set frame after the second trigger frame in sequence to generate a second video.
[0135] In one embodiment, the terminal generates the multi-segment second video, and the code controls at least one of a playing order, a playing time and a playing frequency of the multi-segment second video.
[0136] In one embodiment, the first video is recorded by a camera. Optionally, the camera records the first video in response to the terminal being connected with the camera. Optionally, the terminal stores a code for controlling the camera, and the code sets a control parameter of the camera, which at least includes at least one of opening, closing, aperture size, shutter speed and sensitivity.
[0137] In one embodiment, in response to the terminal storing the code, the code controls the second video to start playing when the point-reading device starts reading the second video, and the code controls the camera to open.
[0138] In summary, the above method obtains the second video by dotting the to-be-processed video, and sets the type calibration region on the second video frame, thereby reducing the similarity between the second trigger frame of the second video and the previous frame of the second trigger frame, and increasing the similarity between the second trigger frame of the second video and the next frame of the second trigger frame.
[0139] To obtain the second video, the method shown in Figure 8 is performed. In one embodiment, Figure 8 a flowchart for obtaining the second video according to one exemplary embodiment of the present application is shown, that is, the step 320 can include the following steps:
[0140] Step 801, recording a to-be-processed video;
[0141] In one embodiment, the terminal records a basic interaction video in which a finger points to a teaching material and the finger stops moving as the to-be-processed video. It is worth noting that the data obtained by the point-reading device through the video and the data obtained through the real interaction process should not have a large gap, and therefore it is necessary to adjust appropriate video shooting angles, brightness and other variables according to the actual situation.
[0142] In one embodiment, the to-be-processed video is obtained by obtaining an existing video in a video library; in one embodiment, the to-be-processed video is obtained by obtaining a video uploaded by a user.
[0143] Step 802, frame division;
[0144] The terminal divides the recorded video into frames to obtain a frame list F. In one embodiment, the video is obtained by refreshing many pictures in a certain frequency and order, and the frame division is to extract the original constituent pictures from the video. Optionally, the to-be-processed video is divided into frames by using opencv. Optionally, the to-be-processed video is divided into frames by using ffmpeg.
[0145] Step 803, analyzing the material to obtain the second trigger frame in the material;
[0146] The terminal analyzes the obtained video to be processed, and determines the second trigger frame in the frame list through expert experience, denoted as imagei.
[0147] Step 804, dotting the second trigger frame and the frame after the second trigger frame;
[0148] The terminal dots the frame imagei in the frame list. When j
[0149] Through the above dotting processing, the dotting areas of the frames before and after the imagei frame are not the same.
[0150] Step 805, video synthesis of the frame after processing;
[0151] The terminal replaces the original video frame with the processed frame to obtain the second video.
[0152] Step 806, saving to the material library.
[0153] The terminal saves the obtained second video to the material library.
[0154] In summary, the above method obtains the second video by dotting the video to be processed, and sets the type calibration area on the frame of the second video, thereby reducing the similarity between the second trigger frame of the second video and the frame before the second trigger frame, and increasing the similarity between the second trigger frame of the second video and the frame after the second trigger frame.
[0155] To record the first video, the method shown in Figure 9 is executed. In one embodiment, Figure 9 a flowchart of recording the first video according to one example embodiment of the present application is shown, i.e., the above step 320 further includes the following steps:
[0156] Step 901, the first person places the display screen in a suitable position;
[0157] The first person places the display screen of the reading device in a suitable position, and optionally, the display screen of the reading device faces the display screen of the terminal.
[0158] Step 902, reading the second video in the material library;
[0159] The terminal reads the second video in the material library.
[0160] Step 903, automatically opening the camera to start recording;
[0161] The terminal automatically opens the camera to start recording through the program.
[0162] Step 904, automatically playing the material;
[0163] The terminal automatically plays the material through the program.
[0164] Step 905, ending recording;
[0165] In response to the terminal stopping playing the material through the program, or the terminal closing the camera through the program, or the first person closing the point reading device, the terminal ends recording.
[0166] Step 906, saving the recorded video.
[0167] The terminal saves the video obtained by recording.
[0168] In summary, the above method realizes the opening of the camera, the automatic playing of the material, and the final recording of the first video through the terminal control without the need for human participation. The above method not only saves manpower investment, but also greatly improves the efficiency of recording the first video.
[0169] To determine the reaction speed of the point reading device, the flow chart of the reaction speed determination method of the point reading device of one exemplary embodiment of the present application is executed as shown in Figure 10 As shown in Figure 10 The method comprises:
[0170] Step 1001, reading the first video;
[0171] The terminal reads the recorded video as the first video.
[0172] Step 1002, adding a timestamp to the first video;
[0173] The terminal adds a timestamp to the first video. The purpose of this step is to add a timestamp to each frame of the video, which facilitates subsequent positioning to the key frame to obtain the accurate time point of the frame. The starting time of the timestamp is 0.000s. Optionally, the terminal adds a timestamp to the first video through ffmpeg.
[0174] Step 1003, dividing the first video into frames;
[0175] The terminal performs frame splitting on the first video. In an embodiment, the first video is refreshed by a plurality of pictures at a certain frequency and order, and the frame splitting is to extract the original constituent pictures from the first video. The frame splitting is performed on the first video by using frame splitting software opencv or ffmpeg.
[0176] Step 1004, positioning the first trigger frame;
[0177] The terminal positions the first trigger frame on the first video.
[0178] Step 1005, OCR (Optical Character Recognition) identifies the first trigger frame timestamp to obtain the expected reaction time point;
[0179] The terminal identifies the first trigger frame timestamp of the first video by using OCR technology to obtain the expected reaction time point T1.
[0180] Step 1006, positioning the response frame;
[0181] The terminal positions the response frame on the first video.
[0182] Step 1007, OCR identifies the response frame timestamp to obtain the actual reaction time point;
[0183] The terminal identifies the response frame timestamp of the first video by using OCR technology to obtain the actual reaction time point T2.
[0184] Step 1008, calculating the reaction time consumption;
[0185] The terminal calculates the difference T2-T1 between the actual reaction time point and the expected reaction time point, and takes the difference as the reaction time consumption of the point reading device.
[0186] Step 1009, outputting the result.
[0187] The terminal takes the difference T2-T1 between the actual reaction time point and the expected reaction time point as the reaction speed of the point reading device, and outputs the reaction speed.
[0188] In summary, by image processing the first video recording the interaction process of the user and the point reading device, the terminal obtains the video frame when the point reading device receives the point reading operation and the video frame when the point reading device starts to respond to the point reading operation, and determines the reaction speed of the point reading device based on the two video frames. The reaction speed determination method of the point reading device adopts the image processing technology, so that the manual intervention is not required in the process of determining the reaction speed of the point reading device, and the full automation is realized in data acquisition and analysis, which not only saves the manpower investment, but also greatly improves the efficiency of determining the reaction speed of the point reading device, and the finally determined reaction speed of the point reading device eliminates the human-induced error.
[0189] To realize the positioning of the first trigger frame of the first video, in one embodiment, Figure 11 The method flowchart for acquiring the first trigger frame on the first video in one exemplary embodiment of the application is shown, that is, the step 1004 comprises the following steps:
[0190] Step 1101, start;
[0191] The terminal receives the instruction of starting to position the first trigger frame of the first video.
[0192] Step 1102, set i = 1, the similarity threshold value t, the frame list F, the length N, and the positioning pos = null.
[0193] The terminal pre-sets the parameters of the first video.
[0194] Step 1103, i = i + 1 and i <= N-1.
[0195] The terminal judges the condition i = i + 1 and i <= N-1, if yes, it goes to step 1104, and if no, it goes to step 1107.
[0196] Step 1104, acquire three frames f(i-1), fi, f(i+1), and calculate the similarity of the type marking area of the first trigger frame in [f(i-1), fi] and [fi, f(i+1)], which are sim1 and sim2 in turn.
[0197] The terminal calculates three frames f(i-1), fi, f(i+1), and calculates the similarity of the type marking area of the first trigger frame in [f(i-1), fi] and [fi, f(i+1)], which are sim1 and sim2 in turn.
[0198] Step 1105, sim1 < t and sim2 >= t.
[0199] The terminal judges the sim1 < t and sim2 >= t judgment condition, if yes, enters step 1106, if no, enters step 1103.
[0200] Step 1106, positioning pos = i;
[0201] The terminal positions the first trigger frame of the first video.
[0202] Step 1107, end.
[0203] The terminal executes the instruction of ending positioning the first trigger frame of the first video.
[0204] In summary, by calculating the similarity of adjacent frames of the first video, when the similarity of the current frame and the previous frame is less than the first threshold value, and the similarity of the current frame and the next frame is greater than or equal to the first threshold value, the current frame is determined as the first trigger frame of the first video. The above method realizes the full-automatic determination of the first trigger frame of the first video, without the need for human participation, not only saving the human input, but also greatly improving the efficiency of determining the first trigger frame.
[0205] To realize the positioning of the response frame of the first video, in one embodiment, Figure 12 The method flow chart for acquiring the response frame on the first video of one exemplary embodiment of the application is shown, that is, the above-mentioned step 1006 includes the following steps:
[0206] Step 1201, start;
[0207] The terminal receives the instruction of starting positioning the response frame of the first video.
[0208] Step 1202, set i = N, the similarity threshold value is p, the frame list is F, the length is N, and the positioning pos = null;
[0209] The terminal pre-sets the parameters of the first video.
[0210] Step 1203, i = i-1 and i > 2;
[0211] The terminal judges the i = i-1 and i > 2 judgment condition, if yes, enters step 1204, if no, enters step 1207.
[0212] Step 1204, acquire three frames of f(i-1), fi, f(i+1), and calculate the similarity of the midpoint reading response area in [f(i-1), fi], [fi, f(i+1)], which are sim3 and sim4 in turn;
[0213] The terminal calculates f(i-1), f i, f(i+1) three frames, and calculates the similarity of the midpoint reading response area in [f(i-1), f i] and [f i, f(i+1)], which are sim3 and sim4 in turn.
[0214] In step 1205, sim3
[0215] The terminal judges the sim3
[0216] In step 1206, the position pos is i.
[0217] The terminal positions the response frame of the first video.
[0218] In step 1207, the process ends.
[0219] The terminal executes the instruction of ending the positioning of the response frame of the first video.
[0220] It is worth noting that the frame list of the first trigger frame positioned by the terminal is arranged in chronological order, that is, Figure 11 The method of positioning the trigger frame shown is to judge in chronological order; the frame list of the response frame positioned by the terminal is arranged in reverse chronological order, that is, Figure 12 The method of positioning the response frame shown is to judge in reverse chronological order.
[0221] In one embodiment, the similarity threshold t of the trigger frame positioned is the same as or different from the similarity threshold p of the response frame positioned.
[0222] In summary, by calculating the similarity of adjacent frames of the first video, when the similarity between the current frame and the previous frame is less than the second threshold, and the similarity between the current frame and the next frame is greater than or equal to the second threshold, the current frame is determined as the first trigger frame of the first video. The above method realizes the full-automatic determination of the response frame of the first video, without the need for human participation, which not only saves manpower investment, but also greatly improves the efficiency of determining the response frame.
[0223] Figure 13 is a structural block diagram of a reaction speed determination device of a point reading equipment provided by an exemplary embodiment of the present application, as Figure 13 The device comprises:
[0224] The acquisition module 1301 is configured to acquire a first video, the first video being obtained by recording a process of responding to a point reading operation of the point reading equipment;
[0225] The processing module 1302 is configured to extract a first trigger frame and a response frame of the first video from the first video based on similarity between adjacent video frames, where the first trigger frame is a video frame when the point reading device recognizes a point reading operation, and the response frame is a video frame when the point reading device starts to respond to the point reading operation.
[0226] The determining module 1303 is configured to determine a reaction speed of the point reading device based on the first trigger frame and the response frame.
[0227] In an optional embodiment, the processing module 1302 is further configured to perform a frame operation on the first video with the time stamp to obtain a first video frame list.
[0228] In an optional embodiment, the processing module 1302 is further configured to extract the first trigger frame from the first video frame list based on similarity of a point reading operation region in adjacent video frames, where the point reading operation region is a region for recognizing the point reading operation.
[0229] In an optional embodiment, the processing module 1302 is further configured to extract the response frame from the first video frame list based on similarity of a point reading response region in adjacent video frames, where the point reading response region is a region for responding to the point reading operation.
[0230] In an optional embodiment, the first video frame list contains N video frames.
[0231] In an optional embodiment, the processing module 1302 is further configured to calculate a first similarity between the point reading operation region in the i-th frame and the i-1-th frame of the first video frame list, and a second similarity between the point reading operation region in the i-th frame and the i+1-th frame of the first video frame list.
[0232] In an optional embodiment, the processing module 1302 is further configured to determine the i-th frame as the first trigger frame in response to the first similarity being less than a first threshold value and the second similarity being greater than or equal to the first threshold value.
[0233] where N is a positive integer greater than 3, i is a positive integer not greater than N-2, and i is greater than or equal to 2.
[0234] In an optional embodiment, the point reading operation region includes a point reading operation sub-region and a type marking region, where the type marking region is a region for marking a video frame type of a current video frame by using a visual feature marked by a human being in advance, and the video frame type includes at least one of a frame before the point reading operation, the trigger frame of the point reading operation, and a frame after the point reading operation.
[0235] In an optional embodiment, the processing module 1302 is further configured to calculate a first similarity between the type-labeled region of the i-th frame of the first video frame list and the type-labeled region of the i-1-th frame, and a second similarity between the type-labeled region of the i-th frame of the first video and the type-labeled region of the i+1-th frame.
[0236] In an optional embodiment, the processing module 1302 is further configured to calculate a third similarity between the point-and-read response region in the m-th frame and the point-and-read response region in the m-1-th frame of the first video frame list, and a fourth similarity between the point-and-read response region in the m-th frame and the point-and-read response region in the m+1-th frame of the first video frame list.
[0237] In an optional embodiment, the processing module 1302 is further configured to determine the m-th frame as the response frame in response to the third similarity being less than the second threshold value and the fourth similarity being greater than or equal to the second threshold value.
[0238] wherein N is a positive integer greater than 3, m is a positive integer not greater than N-1, and m is greater than or equal to 3.
[0239] In an optional embodiment, the processing module 1302 is further configured to obtain the first video with timestamps by setting timestamps for the first video.
[0240] In an optional embodiment, the determining module 1303 is further configured to obtain a first timestamp of the first trigger frame and a second timestamp of the response frame.
[0241] In an optional embodiment, the determining module 1303 is further configured to determine the reaction speed of the point-and-read device based on a difference between the second timestamp and the first timestamp.
[0242] In an optional embodiment, the obtaining module 1301 is further configured to control the camera to record the human-computer interaction process of the point-and-read device, and take the recorded video as the first video.
[0243] In an optional embodiment, the obtaining module 1301 is further configured to control the display device to play the second video in the point-and-read operation region, the second video being a video recorded by recording the human-computer interaction process of the point-and-read device, and the second video having a type-labeled region pre-labeled thereon.
[0244] In an optional embodiment, the obtaining module 1301 is further configured to control the camera to record the second video, and the process of the point-and-read device responding to the point-and-read operation played in the second video, and take the recorded video as the first video.
[0245] In an optional embodiment, the obtaining module 1301 is further configured to perform a frame operation on the to-be-processed video to obtain a to-be-processed video frame list, wherein the to-be-processed video displays a process in which a user points and reads a target word.
[0246] In an optional embodiment, the acquisition module 1301 is further configured to, in response to the calibration operation, set the second trigger frame in the list of video frames to be processed to display the first performance feature, the frame before the second trigger frame to display the second performance feature, and the frame after the second trigger frame to display the first performance feature.
[0247] In an optional embodiment, the acquisition module 1301 is further configured to replace the second trigger frame, the frame before the second trigger frame, and the frame after the second trigger frame in the list of video frames to be processed with the set second trigger frame, the set frame before the second trigger frame, and the set frame after the second trigger frame in sequence to generate the second video.
[0248] In an optional embodiment, the second trigger frame is in a type calibration area.
[0249] In an optional embodiment, the acquisition module 1301 is further configured to, in response to the calibration operation, set the type calibration area of the second trigger frame to display the first performance feature, the type calibration area of the frame before the second trigger frame to display the second performance feature, and the type calibration area of the frame after the second trigger frame to display the first performance feature.
[0250] In an optional embodiment, the performance feature includes at least one of a pattern performance feature, a color performance feature, and a text performance feature.
[0251] To sum up, the above device obtains the video frame when the point reading device recognizes the point reading operation and the video frame when the point reading device starts to respond to the point reading operation by image processing the first video recording the interaction process between the user and the point reading device, and determines the reaction speed of the point reading device based on the two video frames. The reaction speed determination device of the point reading device adopts image processing technology, so that manual intervention is not required in the process of determining the reaction speed of the point reading device, and full automation is realized in data acquisition and analysis results. Not only the human input is saved, but also the efficiency of determining the reaction speed of the point reading device is greatly improved, and the reaction speed of the point reading device determined finally eliminates the error introduced by human.
[0252] Figure 14A structural block diagram of an electronic device 1400 is shown, which is provided by an example embodiment of the present application. The electronic device 1400 can be a portable mobile terminal, such as a smartphone, a tablet computer, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a notebook computer, or a desktop computer. The electronic device 1400 can also be referred to as a user equipment, a portable terminal, a laptop terminal, a desktop terminal, or other names.
[0253] Generally, the electronic device 1400 includes a processor 1401 and a memory 1402.
[0254] The processor 1401 can include one or more processing cores, such as a 4-core processor, an 8-core processor, or the like. The processor 1401 can be implemented in the form of at least one of a DSP (Digital Signal Processing), an FPGA (Field-Programmable Gate Array), and a PLA (Programmable Logic Array). The processor 1401 can also include a main processor and a co-processor. The main processor is a processor for processing data in an awake state, also referred to as a CPU (Central Processing Unit). The co-processor is a low-power processor for processing data in a standby state. In some embodiments, the processor 1401 can be integrated with a GPU (Graphics Processing Unit) for rendering and drawing content to be displayed on a display screen. In some embodiments, the processor 1401 can further include an AI (Artificial Intelligence) processor for processing machine learning-related computing operations.
[0255] The memory 1402 can include one or more computer-readable storage media, which can be non-transitory. The memory 1402 can also include a high-speed random access memory, and a non-volatile memory such as one or more disk storage devices, flash storage devices. In some embodiments, the non-transitory computer-readable storage medium in the memory 1402 is used to store at least one instruction for being executed by the processor 1401 to implement an image inpainting method provided by a method embodiment of the present application.
[0256] In some embodiments, the electronic device 1400 can further include a peripheral device interface 1403 and at least one peripheral device. The processor 1401, the memory 1402 and the peripheral device interface 1403 can be connected through a bus or a signal line. Each peripheral device can be connected to the peripheral device interface 1403 through a bus, a signal line or a circuit board. Specifically, the peripheral device includes at least one of a radio frequency circuit 1404, a display screen 1405, a camera assembly 1406, an audio circuit 1407, a positioning assembly 1408 and a power supply 1409.
[0257] The peripheral device interface 1403 can be used to connect at least one peripheral device related to input / output to the processor 1401 and the memory 1402. In some embodiments, the processor 1401, the memory 1402 and the peripheral device interface 1403 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1401, the memory 1402 and the peripheral device interface 1403 can be implemented on a separate chip or circuit board, and the present embodiment is not limited in this regard.
[0258] The radio frequency circuit 1404 is used to receive and transmit RF (Radio Frequency, radio frequency) signals, also known as electromagnetic signals. The radio frequency circuit 1404 communicates with a communication network and other communication devices through electromagnetic signals. The radio frequency circuit 1404 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals into electrical signals. Optionally, the radio frequency circuit 1404 includes an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and the like. The radio frequency circuit 1404 can communicate with other terminals through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to the World Wide Web, a metropolitan area network, an intranet, various generations of mobile communication networks (2G, 3G, 4G and 5G), a wireless local area network and / or a WiFi (Wireless Fidelity, wireless fidelity) network. In some embodiments, the radio frequency circuit 1404 can also include NFC (Near Field Communication, near field communication) related circuit, and the present application is not limited in this regard.
[0259] The display screen 1405 is configured to display a UI (User Interface). The UI can include graphics, text, icons, video, and any combination thereof. When the display screen 1405 is a touch display screen, the display screen 1405 is further configured to capture touch signals on or above the surface of the display screen 1405. The touch signals can be input to the processor 1401 as control signals for processing. In this case, the display screen 1405 can also be configured to provide virtual buttons and / or virtual keyboard, also known as soft buttons and / or soft keyboard. In some embodiments, the display screen 1405 can be one, arranged on the front panel of the electronic device 1400; in other embodiments, the display screen 1405 can be at least two, arranged on different surfaces of the electronic device 1400 or in a folding design; in other embodiments, the display screen 1405 can be a flexible display screen, arranged on a curved surface or a folding surface of the electronic device 1400. Even, the display screen 1405 can also be arranged in an irregular shape, i.e. a special-shaped screen. The display screen 1405 can be made of LCD (Liquid Crystal Display), OLED (Organic Light-Emitting Diode), etc.
[0260] The camera assembly 1406 is configured to capture images or videos. Optionally, the camera assembly 1406 includes a front camera and a rear camera. Typically, the front camera is arranged on the front panel of the terminal, and the rear camera is arranged on the back of the terminal. In some embodiments, the rear camera is at least two, which is any one of a main camera, a depth-of-field camera, a wide-angle camera, and a long-focus camera, to realize the background blur function of the main camera and the depth-of-field camera, the panoramic shooting and VR (Virtual Reality) shooting function of the main camera and the wide-angle camera, or other fusion shooting functions. In some embodiments, the camera assembly 1406 can further include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. The dual-color temperature flash refers to the combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.
[0261] The audio circuit 1407 can include a microphone and a speaker. The microphone is used to collect sound waves of a user and an environment, and convert the sound waves into an electrical signal input to the processor 1401 for processing, or input to the radio frequency circuit 1404 to realize voice communication. For the purpose of stereo sound collection or noise reduction, the microphone can be multiple, respectively arranged at different parts of the electronic device 1400. The microphone can also be an array microphone or an omnidirectional collection type microphone. The speaker is used to convert the electrical signal from the processor 1401 or the radio frequency circuit 1404 into sound waves. The speaker can be a conventional diaphragm speaker, or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, not only can the electrical signal be converted into a sound wave audible to humans, but also can be converted into a sound wave inaudible to humans for ranging purposes. In some embodiments, the audio circuit 1407 can also include a headphone jack.
[0262] The positioning component 1408 is used to position the current geographic location of the electronic device 1400 to realize navigation or LBS (Location Based Service, Location Based Service). The positioning component 1408 can be a positioning component based on the U.S. GPS (Global Positioning System, Global Positioning System), China's Beidou system or Russia's Galileo system.
[0263] The power supply 1409 is used to supply power to each component in the electronic device 1400. The power supply 1409 can be alternating current, direct current, disposable battery or rechargeable battery. When the power supply 1409 includes a rechargeable battery, the rechargeable battery can be a wired charging battery or a wireless charging battery. The wired charging battery is a battery charged through a wired line, and the wireless charging battery is a battery charged through a wireless coil. The rechargeable battery can also be used to support fast charging technology.
[0264] In some embodiments, the electronic device 1400 further includes one or more sensors 1410. The one or more sensors 1410 include but are not limited to: an acceleration sensor 1411, a gyroscope sensor 1412, a pressure sensor 1413, a fingerprint sensor 1414, an optical sensor 1415 and a proximity sensor 1416.
[0265] The acceleration sensor 1411 can detect the acceleration size in three coordinate axes of the coordinate system established by the electronic device 1400. For example, the acceleration sensor 1411 can be used to detect the components of gravitational acceleration in three coordinate axes. The processor 1401 can control the display screen 1405 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor 1414. The acceleration sensor 1414 can also be used for gaming or user motion data collection.
[0266] The gyroscope sensor 1412 can detect the body direction and rotation angle of the electronic device 1400, and can collect 3D motions of a user with respect to the electronic device 1400 in cooperation with the acceleration sensor 1411. The processor 1401 can implement the following functions according to data collected by the gyroscope sensor 1412: motion sensing (e.g., changing a UI according to a tilt operation of a user), image stabilization during photographing, game control, and inertial navigation.
[0267] The pressure sensor 1413 is disposed at a side bezel of the electronic device 1400 and / or a lower layer of the display 1405. When the pressure sensor 1413 is disposed at the side bezel of the electronic device 1400, a grip signal of a user with respect to the electronic device 1400 can be detected, and left / right hand recognition or a shortcut operation can be performed by the processor 1401 according to the grip signal collected by the pressure sensor 1413. When the pressure sensor 1413 is disposed at the lower layer of the display 1405, a pressure operation of a user with respect to the display 1405 can be detected by the processor 1401, and an operable control on a UI can be controlled according to the pressure operation. The operable control includes at least one of a button control, a scroll bar control, an icon control, and a menu control.
[0268] The fingerprint sensor 1414 is used to collect a fingerprint of a user, and the processor 1401 can identify an identity of the user according to the fingerprint collected by the fingerprint sensor 1414, or the fingerprint sensor 1414 can identify the identity of the user according to the collected fingerprint. When the identity of the user is identified as a trusted identity, the processor 1401 can authorize the user to perform a related sensitive operation, and the sensitive operation includes unlocking a screen, viewing encrypted information, downloading software, payment, and changing a setting. The fingerprint sensor 1414 can be disposed at a front surface, a back surface, or a side surface of the electronic device 1400. When a physical button or a manufacturer's logo is disposed on the electronic device 1400, the fingerprint sensor 1414 can be integrated with the physical button or the manufacturer's logo.
[0269] The optical sensor 1415 is used to collect an ambient light intensity. In an embodiment, the processor 1401 can control a display brightness of the display 1405 according to the ambient light intensity collected by the optical sensor 1415. Specifically, when the ambient light intensity is high, the display brightness of the display 1405 is increased, and when the ambient light intensity is low, the display brightness of the display 1405 is decreased. In another embodiment, the processor 1401 can dynamically adjust a photographing parameter of the camera assembly 1406 according to the ambient light intensity collected by the optical sensor 1415.
[0270] The proximity sensor 1416, also referred to as a distance sensor, is usually arranged on the front panel of the electronic device 1400. The proximity sensor 1416 is used to collect the distance between the user and the front of the electronic device 1400. In an embodiment, when the proximity sensor 1416 detects that the distance between the user and the front of the electronic device 1400 gradually decreases, the display screen 1405 is switched from the bright screen state to the screen-off state under the control of the processor 1401; when the proximity sensor 1416 detects that the distance between the user and the front of the electronic device 1400 gradually increases, the display screen 1405 is switched from the screen-off state to the bright screen state under the control of the processor 1401.
[0271] Those skilled in the art can understand that the structure shown in the above embodiments is not a limitation on the electronic device 1400, and the electronic device 1400 can include more or fewer components than those shown in the figure, or combine certain components, or adopt a different arrangement of components. Figure 14 Those skilled in the art can understand that the structure shown in the above embodiments is not a limitation on the electronic device 1400, and the electronic device 1400 can include more or fewer components than those shown in the figure, or combine certain components, or adopt a different arrangement of components.
[0272] The present application also provides a computer readable storage medium, the storage medium stores at least one instruction, at least one program, a code set or an instruction set, the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to realize the reaction speed determination method of the point reading device provided by the above-mentioned method embodiment.
[0273] The present application provides a computer program product or computer program, the computer program product or computer program includes computer instructions, the computer instructions are stored in a computer readable storage medium. The processor of the computer device reads the computer instructions from the computer readable storage medium, and the processor executes the computer instructions, so that the computer device executes the reaction speed determination method of the point reading device provided by the above-mentioned method embodiment.
[0274] The above-mentioned embodiment serial number of the present application is only for description, not representing the advantages and disadvantages of the embodiments.
[0275] Those skilled in the art can understand that all or part of the steps of the above-mentioned embodiments can be completed by hardware, or by program to instruct related hardware, the program can be stored in a computer readable storage medium, and the above-mentioned storage medium can be read only memory, disk or optical disk, etc.
[0276] The above-mentioned only for optional embodiments of the present application, and does not limit the present application, any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application, should be included in the protection scope of the present application.
Claims
1. A method for determining the reaction speed of a point reading device, characterized in that: The method comprises: Controlling the display device to play a second video in the point reading operation area, where the second video is a video recorded of a human-computer interaction process of the point reading device, and a type calibration area is pre-marked on the second video, where the type calibration area is an area that uses pre-manually calibrated visual features to indicate the video frame type of the current video frame, where the video frame type includes at least one of: a preceding frame of the point reading operation, a triggering frame of the point reading operation, and a succeeding frame of the point reading operation; controlling the camera to record the second video and the process of the point reading device responding to the point reading operation played in the second video, and using the recorded video as the first video; Extracting a first trigger frame and a response frame of the first video from the first video based on similarity between adjacent video frames, wherein the first trigger frame is a video frame when the point reading device recognizes the point reading operation, and the response frame is a video frame when the point reading device begins to respond to the point reading operation; Based on the first trigger frame and the response frame, a reaction speed of the point reading device is determined.
2. The method according to claim 1, characterized in that The extracting, from the first video based on the similarity between adjacent video frames, a first trigger frame and a response frame of the first video includes: Performing a frame operation on the first video with the timestamp to obtain a first video frame list; Extracting the first trigger frame from the first video frame list based on similarity of the point reading operation areas in the adjacent video frames, the point reading operation area being an area for identifying the point reading operation; The response frame is extracted from the first video frame list based on the similarity of the point reading response areas in the adjacent video frames, where the point reading response area is an area for responding to the point reading operation.
3. The method according to claim 2, characterized in that The first video frame list includes N video frames; The extracting the first trigger frame from the first video frame list based on the similarity of the point reading operation areas in the adjacent video frames includes: Calculating a first similarity between the touch-reading operation areas in the i-th frame and the i-1-th frame in the first video frame list, and a second similarity between the touch-reading operation areas in the i-th frame and the i+1-th frame in the first video frame list; In response to the first similarity being less than a first threshold and the second similarity being greater than or equal to the first threshold, determining the i-th frame as the first triggering frame; Wherein, N is a positive integer greater than 3, i is a positive integer not greater than N-2, and i is greater than or equal to 2.
4. The method according to claim 3, characterized in that The point reading operation area includes: a point reading operation sub-area and the type calibration area; The calculating of obtaining a first similarity between the point reading operation areas in the i-th frame and the i-1-th frame in the first video frame list, and a second similarity between the point reading operation areas in the i-th frame and the i+1-th frame in the first video frame list, includes: A first similarity between the type calibration area of the i-th frame and the type calibration area of the i-1-th frame of the first video frame list, and a second similarity between the type calibration area of the i-th frame and the type calibration area of the i+1-th frame of the first video are calculated.
5. The method according to claim 2, characterized in that The first video frame list includes N video frames; The extracting the response frame from the first video frame list based on the similarity of the point reading response areas in the adjacent video frames includes: Calculating a third similarity between the point reading response areas in the m-th frame and the m-1-th frame in the first video frame list, and a fourth similarity between the point reading response areas in the m-th frame and the m+1-th frame in the first video frame list; In response to the third similarity being less than a second threshold and the fourth similarity being greater than or equal to the second threshold, determining the mth frame as the response frame; Wherein, N is a positive integer greater than 3, m is a positive integer not greater than N-1, and m is greater than or equal to 3.
6. The method according to any one of claims 1 to 5, characterized in that: The method further comprises: By setting a timestamp on the first video, a first video with a timestamp is obtained; The determining the response speed of the point reading device based on the first trigger frame and the response frame includes: Obtaining a first timestamp of the first trigger frame and a second timestamp of the response frame; A response speed of the point reading device is determined based on a difference between the second timestamp and the first timestamp.
7. The method according to any one of claims 1 to 5, characterized in that: The second video is obtained in the following manner: Divide the video to be processed into frames to obtain a list of video frames to be processed, wherein the video to be processed shows a process in which the user clicks on the target word; In response to the calibration operation, setting the second trigger frame in the list of video frames to be processed to display the first expression feature, the frame before the second trigger frame to display the second expression feature, and the frame after the second trigger frame to display the first expression feature; The second trigger frame after being set, the previous frame after being set, and the next frame after being set will replace the second trigger frame, the previous frame after being set, and the next frame after being set in the list of video frames to be processed in sequence to generate the second video.
8. The method according to claim 7, characterized in that The second trigger frame contains the type calibration area; In response to the calibration operation, setting the second trigger frame in the list of video frames to be processed to display the first expression feature, the frame before the second trigger frame to display the second expression feature, and the frame after the second trigger frame to display the first expression feature includes: In response to the calibration operation, the type calibration area of the second trigger frame is set to display the first performance feature, the type calibration area of the previous frame of the second trigger frame is set to display the second performance feature, and the type calibration area of the next frame of the second trigger frame is set to display the first performance feature.
9. The method according to claim 7, characterized in that The expression feature includes at least one of a pattern expression feature, a color expression feature, and a text expression feature.
10. A device for determining the reaction speed of a point reading device, characterized in that: The device comprises: an acquisition module, configured to control the display device to play a second video in the point reading operation area, the second video being a video recorded from a human-computer interaction process of the point reading device, and the second video being pre-marked with a type calibration area, the type calibration area being an area that uses pre-manually calibrated visual features to indicate the video frame type of the current video frame, the video frame type including at least one of: a preceding frame of the point reading operation, a triggering frame of the point reading operation, and a succeeding frame of the point reading operation; controlling a camera to record the second video and the process of the point reading device responding to the point reading operation played in the second video, and using the recorded video as the first video; a processing module, configured to extract a first trigger frame and a response frame from the first video by performing image processing on the first video, wherein the first trigger frame is a video frame when the point reading device receives a point reading operation, and the response frame is a video frame when the point reading device begins to respond to the point reading operation; A determination module is used to determine the response speed of the point reading device based on the first trigger frame and the response frame.
11. The device according to claim 10, characterized in that The acquisition module is further configured to perform a frame operation on the video to be processed to obtain a list of video frames to be processed, wherein the video to be processed shows a process in which the user clicks on the target word; The acquisition module is further configured to, in response to a calibration operation, set a second trigger frame in the list of video frames to be processed to display a first expression feature, a frame preceding the second trigger frame to display a second expression feature, and a frame following the second trigger frame to display the first expression feature; The acquisition module is also used to replace the second trigger frame, the previous frame of the second trigger frame and the next frame of the second trigger frame in the list of video frames to be processed with the set second trigger frame, the previous frame of the second trigger frame and the next frame of the second trigger frame in sequence to generate the second video.
12. A computer device, characterized in that: The computer device includes: a processor and a memory, the memory stores a computer program, and the computer program is loaded and executed by the processor to implement the method for determining the reaction speed of the point reading device according to any one of claims 1 to 9.
13. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which is loaded and executed by a processor to implement the method for determining the reaction speed of a point reading device according to any one of claims 1 to 9.
Citation Information
Patent Citations
Video analysis system and method
CN104240224A
Method, apparatus and device for testing reaction time of terminal user interface
CN105302701A