An end-to-end method, device, and equipment for detecting time delay
By recording and identifying time series of clicks and response indicator images, the accuracy of end-to-end time delay detection is solved, achieving a better user experience.
Patent Information
- Application Number
- CN202310104727.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-19
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2043-01-19
AI Technical Summary
The prior art is difficult to accurately detect end-to-end time delays, affecting the user experience.
By recording to generate videos to be parsed with click indication images and response indication images, these images are identified to generate time series and match click and response time series to determine the time delay.
Accurate and stable statistics of end-to-end time delays are achieved, improving user experience.
Smart Images

Figure CN116132710B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of Internet technologies, and in particular, to a method, apparatus, and device for detecting end-to-end time delay. Background Art
[0002] With the development of mobile Internet, the applications of cloud computers and cloud phones are becoming more and more widespread. In this scenario, end-to-end time delay is an important performance metric.
[0003] When a user executes an instruction on a local device, the response time of the local device is generally a few milliseconds and can be almost ignored. However, in the case involving the cloud, the general process is as follows: the user inputs an instruction on the local device → the cloud responds to generate feedback → the cloud returns the feedback to the local device → the local device renders and displays the feedback, and the overall time consumed in the entire link is the end-to-end time delay.
[0004] Based on this, an accurate detection scheme for end-to-end time delay is needed. Summary of the Invention
[0005] Embodiments of this specification provide a method, apparatus, device, and storage medium for detecting end-to-end time delay to solve the following technical problem: an accurate detection scheme for end-to-end time delay is needed.
[0006] To solve the above technical problem, one or more embodiments of this specification are implemented as follows:
[0007] In a first aspect, an embodiment of this specification provides a method for detecting end-to-end time delay, including: recording and generating a video to be parsed that includes a click indication image and a response indication image, where the click indication image is generated on a device side based on a click instruction, and the response indication image is generated by the cloud in response to the click instruction and sent to the device side for display; identifying the click indication image in the video to be parsed to generate a click time series; identifying the response indication image in the video to be parsed to generate a response time series; matching the click time series and the response time series to determine the time delay.
[0008] In a second aspect, an end-to-end time delay detection device provided by an embodiment of this specification includes: recording and generating a video to be parsed that contains a click indication image and a response indication image, where the click indication image is generated on a device side based on a click instruction, and the response indication image is generated by a cloud in response to the click instruction and sent to the device side for display; a click indication image recognition module that recognizes the click indication image in the video to be parsed and generates a click time series; a response indication image recognition module that recognizes the response indication image in the video to be parsed and generates a response time series; and a determination module that matches the click time series and the response time series to determine the time delay.
[0009] In a third aspect, one or more embodiments of this specification provide an electronic device, including:
[0010] at least one processor; and,
[0011] a memory communicatively connected to the at least one processor; where,
[0012] the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method as described in the first aspect.
[0013] In a fourth aspect, an embodiment of this specification provides a non-volatile computer storage medium storing computer-executable instructions, and when a computer reads the computer-executable instructions in the storage medium, the instructions cause one or more processors to execute the method as described in the first aspect.
[0014] One or more of the above technical solutions adopted by one or more embodiments of this specification can achieve the following beneficial effects: recording and generating a video to be parsed that contains a click indication image and a response indication image, where the click indication image is generated on a device side based on a click instruction, and the response indication image is generated by a cloud in response to the click instruction and sent to the device side for display; recognizing the click indication image in the video to be parsed and generating a click time series; recognizing the response indication image in the video to be parsed and generating a response time series; matching the click time series and the response time series to determine the time delay, thereby accurately and stably calculating the end-to-end time delay and improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] To more clearly illustrate the technical solutions in the embodiments of this specification or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0016] Figure 1a It is a schematic flowchart of a method for detecting end-to-end time delay provided by an embodiment of this specification;
[0017] Figure 1b It is a schematic architecture diagram of a system provided by an embodiment of this specification;
[0018] Figure 2 It is a schematic diagram of displaying a click indication image on a device end provided by an embodiment of this specification;
[0019] Figure 3 It is a schematic diagram of three images obtained by screen recording provided by an embodiment of this specification;
[0020] Figure 4 It is a schematic diagram of normal and abnormal situations provided by an embodiment of this specification;
[0021] Figure 5 It is a schematic flowchart of a specific embodiment provided by an embodiment of this specification;
[0022] Figure 6 It is a schematic structural diagram of a device for detecting end-to-end time delay provided by an embodiment of this specification;
[0023] Figure 7 It is a schematic structural diagram of an electronic device provided by an embodiment of this specification. Detailed implementation manners
[0024] Embodiments of this specification provide a method, a device, a device, and a storage medium for detecting end-to-end time delay.
[0025] To enable those skilled in the art to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in the embodiments of this specification in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only some embodiments of this application, rather than all embodiments. Based on the embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.
[0026] In the first aspect, as Figure 1a shown, Figure 1aThe flowchart shows a method for detecting end-to-end time delay provided by an embodiment of this specification, including the following steps:
[0027] S101: Record and generate a video to be parsed that includes a click indication image and a response indication image.
[0028] The click indication image is generated on the local device (hereinafter referred to as the device side) based on a click instruction. The click indication image may contain a first identifier to indicate that the image is a click indication image. For example, the first identifier may be a click icon (such as a cross pointer, a hand icon, or a mouse pointer, etc.); the identifier may also be a specified color. For example, while keeping the background color of the device side unchanged, when a click occurs, the display color of the device side is adjusted to another color.
[0029] The click instruction can be generated by a tester's manual click operation or an automatically executed click based on a script. Generally, the click instruction is issued multiple times to improve the accuracy of the detection result. Correspondingly, multiple click indication images will be generated.
[0030] In the case of a network connection between the device side and the cloud, the click instruction will be uploaded to the cloud. The cloud will then respond based on the instruction, generate a corresponding response indication image, and the response indication image can be returned to the device side in various forms (such as a video stream or an encoded file) and rendered and displayed by the device side. As Figure 1b shown, Figure 1b The schematic diagram of the architecture of a system provided by an embodiment of this specification. Through this architecture, the cloud phone simulates the environment of the local phone side, so that the generated video stream can be displayed locally.
[0031] The response indication image also contains a second identifier similar to the first identifier. That is, the second identifier can also be an indication icon, a number, a color, etc. Among them, the second identifier should be different from the first identifier.
[0032] By keeping the normal operation of the display screen of the device side, using a script to generate simulated click instructions, and at the same time, receiving and displaying the response indication image returned by the cloud. At this time, the phenomenon of alternating click indication images and response indication images will occur on the display screen of the device side.
[0033] At this time, the device can record this phenomenon to obtain a video to be parsed that includes click indication images and response indication images. The click indication images and response indication images in the recorded video to be parsed will appear in the form of "frame images", and the background image in the video to be parsed after removing the click indication images and response indication images may be an image without information.
[0034] For example, assume that the duration of a video to be parsed is 1000 seconds, and 1 second contains 24 frames. Then, the video to be parsed actually contains 24,000 frame images. Assume that the number of click instructions issued by the script is 1000. Then, it may contain 1000 non - consecutive click - indication images and 1000 non - consecutive response - indication images. For example, the Nth click - indication image is the 800th frame in the video to be parsed, and the (N + 1)th click - indication image is the 825th frame.
[0035] S103. Identify the click - indication images in the video to be parsed and generate a click time series.
[0036] All the frame images contained in the image to be parsed can be identified one by one. By identifying the first identifier contained in the video to be parsed, the click - indication images are determined and a click time series is generated.
[0037] The click times contained in the generated click time series can be represented in terms of "frames" or in real time. When represented in terms of frames, the click time series can be a digital sequence such as (1, 25, 51, 77, 100,...) indicating the frame numbers of the click - indication images in the image to be parsed. When represented in terms of time points, the click time series is a real - time series accurate to milliseconds such as (17:34:24:674, 17:34:26:587,...), and the real time is the time point at which the true image is displayed on the device side.
[0038] S105. Identify the response - indication images in the video to be parsed and generate a response time series.
[0039] Similar to the method of generating the click time series described above, all the frame images contained in the image to be parsed can also be identified one by one. By identifying the second identifier contained in the video to be parsed, the response - indication images are determined and a click time series is generated. The representation method of the click time in the generated click time series should be the same as that of the click time series.
[0040] S107. Match the click time series and the response time series to determine the time delay.
[0041] Assume that both the click time series and the response time series have 1000 elements, that is, 1000 clicks occur and 1000 responses are correspondingly generated. Then, the two sequences can be directly subtracted to obtain a delay sequence containing 1000 time lengths.
[0042] For example, assume that both the click time series and the response time series are represented by frames. Then, the resulting delay series obtained after subtraction may be a delay series containing 1000 frame differences, such as (30, 33, 50, 27,...). For each frame, relevant statistics can be performed on these 1000 frame differences, and its statistical value can be used to characterize the time delay.
[0043] Record and generate a video to be parsed that contains click indication images and response indication images. Among them, the click indication images are generated on the device side based on click instructions, and the response indication images are generated by the cloud in response to the click instructions and sent to the device side for display; identify the click indication images in the video to be parsed to generate a click time series; identify the response indication images in the video to be parsed to generate a response time series; match the click time series and the response time series to determine the time delay, thereby accurately and stably calculating the end-to-end time delay and improving the user experience.
[0044] In one implementation, the first click identifier can be in the form of a click icon. That is, generating on the device side based on a click instruction can be receiving a click instruction and generating a click indication image containing a click icon on the device side. For example, when the device side is an Android mobile phone, a function "Pointer Location" is provided in the developer options of the mobile phone. When this function is turned on, a click icon (default is a cross, which can also be modified to other forms of click icons) will be displayed at the click position every time a click occurs. At the same time, the real click time when the click icon appears on the device side can also be provided, thereby generating a click indication image containing a cross. As Figure 2 shown Figure 2 is a schematic diagram of displaying a click indication image on the device side provided by an embodiment of this specification. In this schematic diagram, both a click icon and the real click time are provided in the display interface of the device side.
[0045] At this time, when subsequently identifying the click indication images in the video to be parsed, the frame images containing a click icon (for example, a cross) can be determined as click indication images by identifying the video to be parsed frame by frame. Moreover, the time contained in this image can also be obtained and determined as the real click image.
[0046] In one implementation, the cloud can count the received click instructions and determine the serial number corresponding to the received click instruction. For example, a "touch counter" can be integrated in the cloud. Each time a click time sent by the client is received, the counter is incremented by 1, and a response indication image containing this serial number is generated.
[0047] Furthermore, when identifying the response indication image in the video to be parsed, it is only necessary to identify the serial number in the response indication image. By identifying the serial number included in the image (for example, using the OCR algorithm), not only can the response indication image be conveniently identified, but also the identified serial number can be used as the corresponding response serial number, facilitating the matching between the subsequent two sequences.
[0048] As Figure 3 shown, Figure 3 FIG. 1 is a schematic diagram of three images obtained by screen recording provided by an embodiment of the present specification. The first image is actually the background image displayed on the device after the 5th click, which includes the serial number 5 but does not include the click icon (i.e., does not include the cross); the second image is the image generated after the 5th click on the device is clicked again (i.e., the 6th click), which includes the serial number 5 and a cross-shaped click icon; the third image is the response indication image generated by the cloud and returned to the device for display, which includes the serial number 6.
[0049] In addition, these three images are actually extracted from the video to be parsed, and they may not be adjacent. At the same time, since the video to be parsed is obtained based on the screen recording function, the display time (i.e., the recording time) of the image can also be included in the image. For example, the display times T1, T2, and T3 are shown in the upper right corner of the image.
[0050] In one embodiment, when the cloud generates the response indication image, in addition to displaying the serial number, the color of the screen can also be changed to generate a response indication image that includes the serial number and has a different color from the previous response indication image. For example, if the Nth response indication image is green, the color of the N + 1 response indication image is blue, the color of the N + 2 response indication image is red, and the color of the N + 3 response indication image is green, and so on in a cycle.
[0051] Correspondingly, when identifying the response indication image in the video to be parsed, in addition to identifying the serial number included in the image, the serial number in the image frame with color change in the video to be parsed can also be identified by determining the image frame with color change.
[0052] As described above, when the image is displayed on the device, the click indication image and the response indication image are actually alternately displayed. However, since the click indication image is only one frame, while the response indication image will be displayed as the background image all the time. For example, in Figure 3In the schematic diagram shown, the first image is actually the response indication image for the 5th click. There may be many frames between it and the second image until the 6th click instruction is received and the click indication image will be displayed. At the same time, there are usually also many frames between the second image and the third image, but at this time, the first image is still displayed on the screen.
[0053] In other words, in the video to be parsed, if each frame is recognized, in fact, most of them are response indication images recognized. By marking with image colors and comparing the frame images with color changes before and after, the effective response indication images can be accurately indicated, avoiding the need to recognize each frame by digital serial numbers and improving efficiency.
[0054] After obtaining the click time series and the response time series, the two can be compared to determine the time delay. Under normal circumstances, usually when the user clicks N times, there should be N corresponding responses. At the same time, for the a-th click (1 ≤ a ≤ N) in the response series, the a-th response should be between the a-th click and the (a + 1)-th click. Under normal circumstances, the a-th click and the a-th response form a valid response pair, and a valid response pair represents the end-to-end delay time generated at the time of this click.
[0055] For multiple valid response pairs, statistics can be performed to determine the end-to-end time delay. For example, statistics such as the maximum delay, minimum delay, average delay, etc. can be calculated, and also statistics such as the proportion of those falling within a certain delay range can be calculated. For example, the proportion of delays between [70ms, 100ms], etc.
[0056] In one implementation, if the click time series and the response time series are represented by the number of frames, the time delay between the valid response pairs can be determined according to the frame rate of the device at the time of video recording and the difference in the number of frames.
[0057] For example, a high-performance high-speed camera or the screen recording function of the device can be used to generate the video to be parsed containing click indication images and response indication images. During recording, the corresponding frame rate of the device can be determined. For example, for a device with a 60HZ frame rate, the time resolution can only reach 16.67ms, a 90HZ mobile phone can reach 11.11ms, and a 120Hz can reach 8.3ms. That is, assuming the frame rate of the device at the time of recording is f, the resolution between frames is 1 / f. If the difference in the number of frames between a valid response pair is M, then the time delay corresponding to this valid response pair is M / f. In practical applications, devices supporting high frame rates can be used to effectively improve the accuracy.
[0058] Other situations other than normal situations should be regarded as abnormal situations. In abnormal situations, invalid responses are included. For example, Figure 4 as shown Figure 4 in the figure is a schematic diagram of normal and abnormal situations provided by the embodiments of this specification. The abnormal situations include out-of-order, downstream frame loss, and upstream frame loss.
[0059] Since the click instruction is generated locally on the device side, it can be ensured that the number and order of the recognized click indication images are exactly the same as the number and order of the click instructions. That is, in normal situations, only the response time series needs to be compared.
[0060] For example, in the out-of-order situation as shown Figure 4 in the figure, actually the a-th response occurs after the (a + 1)-th click (normally it should occur between the a-th click and the (a + 1)-th click). This situation is usually caused by network fluctuations and should be excluded.
[0061] Another example is that in the downstream frame loss situation as shown Figure 4 in the figure, actually the (a + 1)-th response is not generated or lost during network transmission, but other responses are still normal. At this time, the (a + 1)-th click and response can be excluded from the two sequences for subsequent matching, which does not affect the result.
[0062] Another example is that in the upstream frame loss situation as shown Figure 4 in the figure, actually due to network transmission problems, the click instruction is lost during upstream (but it is not known which click instruction is lost during upstream). Eventually, the number of responses on the cloud does not match the number of clicks on the device side. That is, the number of click times in the click time series is less than the number of click instructions. At this time, since it is no longer possible to identify which click on the device side the response on the cloud should match, it belongs to invalid experimental materials, and another video to be parsed needs to be re-recorded for parsing.
[0063] The following gives a specific embodiment to illustrate the solution of this application. As shown Figure 5 in the figure Figure 5 is a schematic flowchart of a specific embodiment provided by the embodiments of this specification. Before detection, environmental preparation is first carried out, and this preparation includes two aspects: the device side and the cloud side. On the local device, the click pointer display in the mobile phone function is turned on, the touch counter is turned on in the cloud, and the corresponding cloud phone is applied for, and then detection is entered, and the high-fps screen recording function is turned on.
[0064] Furthermore, start video recording on the local device, use a script to simulate clicking a predetermined number of times (e.g., 1000 times), end the screen recording after the clicking is completed, generate a video to be parsed, and perform video parsing. The video parsing includes using image parsing to identify the click indication image (by identifying the cross pointer), and using OCR to identify the response indication image (by identifying the serial number), and then by matching clicks and responses, the elapsed time is statistically calculated.
[0065] In a second aspect, based on the same idea, the embodiments of this specification also provide the corresponding apparatus and device for the above method, as Figure 6 and Figure 7 shown.
[0066] Figure 6 FIG. is a schematic structural diagram of an end-to-end time delay detection apparatus provided by the embodiments of this specification. The apparatus includes:
[0067] A recording module 601 that records and generates a video to be parsed including a click indication image and a response indication image. Among them, the click indication image is generated on the device side based on a click instruction, and the response indication image is generated by the cloud in response to the click instruction and sent to the device side for display;
[0068] A click indication image recognition module 603 that recognizes the click indication image in the video to be parsed and generates a click time series;
[0069] A response indication image recognition module 605 that recognizes the response indication image in the video to be parsed and generates a response time series;
[0070] A determination module 607 that matches the click time series and the response time series to determine the time delay.
[0071] Optionally, the apparatus further includes a click indication image generation module 609 that receives a click instruction and generates a click indication image including a click icon on the device side; correspondingly, the click indication image recognition module 603 recognizes the click icon in the video to be parsed.
[0072] Optionally, the cloud counts the received click instructions to determine the serial number corresponding to the received click instruction; in response to the click instruction, generates a response indication image including the serial number; correspondingly, the response indication image recognition module 605 recognizes the serial number in the response indication image.
[0073] Optionally, the cloud generates a response indication image that includes the serial number digits and has a color different from that of the previous response indication image; correspondingly, the response indication image recognition module 605 determines the image frames in the video to be parsed that have a color change, and recognizes the serial number digits in the image frames with the color change.
[0074] Optionally, the determination module 607 determines the valid response pairs of the response time series with respect to the click time series, and counts the time delay for determining the valid response pairs.
[0075] Optionally, the determination module 607 determines the frame number difference between valid response pairs, and determines the time delay of the valid response pairs according to the frame rate of the device side during video recording and the frame number difference.
[0076] Optionally, if the number of click times in the click time series is less than the number of click instructions, the recording module 601 records another video to be parsed.
[0077] Optionally, the recording module 601 records and generates a video to be parsed that includes a click indication image and a response indication image through the screen recording function of the device side.
[0078] In a third aspect, as Figure 7 shown, Figure 7 is a schematic structural diagram of an electronic device provided by an embodiment of this specification. The device includes:
[0079] At least one processor; and,
[0080] A memory communicatively connected to the at least one processor; wherein,
[0081] The memory stores instructions executable by the at least one processor. The instructions are executed by the at least one processor so that the at least one processor can execute the method described in the first aspect.
[0082] In a fourth aspect, based on the same idea, an embodiment of this specification further provides a non-volatile computer storage medium corresponding to the above method, storing computer-executable instructions. When a computer reads the computer-executable instructions in the storage medium, the instructions cause one or more processors to execute the method described in the first aspect.
[0083] In the 1990s, it was obvious to distinguish whether an improvement to a technology was a hardware improvement (e.g., improvement to circuit structures such as diodes, transistors, switches, etc.) or a software improvement (improvement to method flows). However, with the development of technology, many method flow improvements today can be regarded as direct improvements to hardware circuit structures. Almost all designers obtain the corresponding hardware circuit structure by programming the improved method flow into the hardware circuit. Therefore, it cannot be said that an improvement to a method flow cannot be implemented with a hardware entity module. For example, a Programmable Logic Device (PLD) (such as a Field Programmable Gate Array (FPGA)) is such an integrated circuit whose logical function is determined by the user's programming of the device. Designers can program themselves to "integrate" a digital system onto a PLD, without having to ask a chip manufacturer to design and fabricate a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly implemented using "logic compiler" software, which is similar to the software compiler used in program development and writing. The original code before compilation also has to be written in a specific programming language, which is called Hardware Description Language (HDL), and there is not only one type of HDL, but many types, such as ABEL (Advanced Boolean Expression Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. Currently, the most commonly used ones are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art should also be clear that by simply performing some logical programming on the method flow with the above-mentioned several hardware description languages and programming it into the integrated circuit, it is easy to obtain the hardware circuit that implements the logical method flow.
[0084] The controller can be implemented in any suitable manner. For example, the controller can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, an application specific integrated circuit (ASIC), a programmable logic controller, and an embedded microcontroller. Examples of the controller include, but are not limited to, the following microcontrollers: ARC 625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art also know that, in addition to implementing the controller in the form of pure computer-readable program code, it is entirely possible to make the controller implement the same function in the form of logic gates, switches, application specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or the structures within the hardware component.
[0085] The systems, devices, modules, or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0086] For the convenience of description, when describing the above devices, they are described separately as various units according to their functions. Of course, when implementing this specification, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0087] Those skilled in the art should understand that the embodiments of this specification can be provided as a method, a system, or a computer program product. Therefore, the embodiments of this specification can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of this specification can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.
[0088] This specification is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the specification. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate means for implementing the functions specified in one or more flows of the flowchart and / or one or more blocks of the block diagram.
[0089] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a particular manner, such that the instructions stored in the computer-readable memory produce a manufacture including instruction means that implement the functions specified in one or more flows of the flowchart and / or one or more blocks of the block diagram.
[0090] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows of the flowchart and / or one or more blocks of the block diagram.
[0091] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0092] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash RAM. The memory is an example of computer-readable media.
[0093] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0094] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0095] This specification may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This specification may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0096] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device, equipment, and non-volatile computer storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0097] The above description has been made of specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0098] The above is only one or more embodiments of this specification and is not intended to limit this specification. For those skilled in the art, various changes and modifications can be made to one or more embodiments of this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of one or more embodiments of this specification shall be included within the scope of the claims of this specification.
Claims
1. A method for detecting end-to-end time delay, which is applied to the scenario of cloud computers or cloud mobile phones. The method includes: Recording and generating a video to be parsed that contains click indication images and response indication images. Among them, the click indication images are generated on the device side based on click instructions, and the response indication images are generated by the cloud in response to the click instructions and sent to the device side for display. The click instructions are multiple times; Identifying the click indication images in the video to be parsed and generating a click time series; Identifying the response indication images in the video to be parsed and generating a response time series; Matching the click time series and the response time series to determine the time delay; Among them, the click indication images are generated on the device side based on click instructions, including: Receiving a click instruction and generating a click indication image containing a click icon on the device side; Correspondingly, identifying the click indication images in the video to be parsed includes: identifying the click icons in the video to be parsed; The response indication images are generated by the cloud in response to the click instructions, including: The cloud counts the received click instructions and determines the serial number corresponding to the received click instruction; In response to the click instruction, generating a response indication image containing the serial number; Correspondingly, identifying the response indication images in the video to be parsed includes: identifying the serial number in the response indication image.
2. The method according to claim 1, wherein Generating a response indication image containing the serial number includes: Generating a response indication image that contains the serial number and has a different color from the previous response indication image; Correspondingly, identifying the response indication images in the video to be parsed includes: determining the image frames with color changes in the video to be parsed and identifying the serial number in the image frames with color changes.
3. The method according to claim 1, wherein Matching the click time series and the response time series to determine the time delay includes: Determining the valid response pairs of the response time series for the click time series, and counting the valid response pairs to determine the time delay.
4. The method according to claim 3, wherein, Determining the valid response pairs of the response time series for the click time series, and counting the valid response pairs to determine the time delay, including: Determining the frame number difference between the valid response pairs, and determining the time delay of the valid response pairs according to the frame rate of the device side during video recording and the frame number difference.
5. According to the method described in claim 1, before matching the click time series and the response time series to determine the time delay, the method further includes: If the number of click times in the click time series is less than the number of click instructions, re-recording another video to be parsed.
6. The method according to claim 1, wherein, Recording and generating a video to be parsed that contains click indication images and response indication images, including: Recording and generating a video to be parsed that contains click indication images and response indication images through the screen recording function of the device side.
7. An end-to-end time delay detection device, which is applied to the scenario of cloud computers or cloud mobile phones. The device includes: Recording module, which records and generates a video to be parsed containing click indication images and response indication images. Among them, the click indication images are generated on the device side based on click instructions, and the response indication images are generated by the cloud in response to the click instructions and sent to the device side for display. The click instructions are multiple times; Click indication image recognition module, which recognizes the click indication images in the video to be parsed and generates a click time series; Response indication image recognition module, which recognizes the response indication images in the video to be parsed and generates a response time series; Determination module, which matches the click time series and the response time series to determine the time delay; Among them, the click indication images are generated on the device side based on click instructions, including: Receiving a click instruction and generating a click indication image containing a click icon on the device side; Correspondingly, recognizing the click indication images in the video to be parsed includes: recognizing the click icons in the video to be parsed; The response indication images are generated by the cloud in response to the click instructions, including: The cloud counts the received click instructions and determines the serial number corresponding to the received click instruction; In response to the click instruction, generating a response indication image containing the serial number; Correspondingly, recognizing the response indication images in the video to be parsed includes: recognizing the serial number in the response indication image.
8. An electronic device, including: At least one processor; And, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Method and device for determining response time
CN108900776A
Method and device for determining response time
JP2020030811A