Method and device for recognizing voice response time
By recording video on the target device screen and using OCR technology to identify characters in the frame image, the problem of low voice response time recognition efficiency is solved, and automated and efficient voice response time determination is achieved.
Patent Information
- Application Number
- CN202110524623.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-13
- Publication Date
- 2025-08-12
- Estimated Expiration
- 2041-05-13
AI Technical Summary
In the prior art, the recognition efficiency of voice response time is low, and it is necessary to manually analyze video frame by frame to determine the time points of the first and last characters of the voice command.
By acquiring the display screen of the target device, the first and last characters in the frame image are identified using optical character recognition (OCR) technology, and the voice response time is determined based on their timestamps.
It realizes automated and accurate determination of voice response time, improves recognition efficiency, reduces manual intervention, and saves recognition time.
Smart Images

Figure CN115346558B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computers, and in particular to a method and device for recognizing speech response time. Background Art
[0002] In related technologies, voice response time is usually used as one of the key indicators to measure the quality of voice recognition. Voice response time usually refers to the time from when a voice command is issued to when the voice command is recognized. The voice response time can be determined by the time when the first and last characters corresponding to the recognized voice command are displayed on the screen of the electronic device. At present, when a comparative analysis of voice response time is required, it is usually necessary to record multiple videos, and then manually analyze the videos frame by frame, and record the time points corresponding to the first and last characters of the voice command to analyze the response speed. However, the above method will result in low recognition efficiency of voice response time.
[0003] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention
[0004] The embodiments of the present invention provide a method and apparatus for recognizing speech response time, so as to at least solve the technical problem of low recognition efficiency of speech response time in the related art.
[0005] According to one aspect of an embodiment of the present invention, a method for identifying voice response time is provided, comprising: obtaining a target video obtained by recording a display screen of a target device, wherein the target video includes a picture displayed on the display screen when a target voice command is input to the target device; performing character recognition on frame images in the target video to obtain a first frame image in which the first character of first information appears, and a second frame image in which the last character of the first information appears, wherein the first information is part or all of the information obtained by the target device performing voice recognition on the target voice command and displayed on the display screen; the first information is the same as the second information, and the second information is the information represented by the target voice command; determining a first timestamp corresponding to the first frame image and a second timestamp corresponding to the second frame image, wherein the first timestamp and the second timestamp are used to determine the voice response time of the target device.
[0006] According to another aspect of an embodiment of the present invention, a device for identifying voice response time is also provided, including: an acquisition unit, used to acquire a target video obtained by recording the display screen of a target device, wherein the target video includes a picture displayed on the display screen when a target voice command is input to the target device; an identification unit, used to perform character recognition on the frame images in the target video, to obtain a first frame image in which the first character in the first information appears, and a second frame image in which the last character in the first information appears, wherein the first information is part or all of the information obtained by the target device performing voice recognition on the target voice command and displayed on the display screen; a determination unit, used to determine a first timestamp corresponding to the first frame image and a second timestamp corresponding to the second frame image, wherein the first timestamp and the second timestamp are used to determine the voice response time of the target device.
[0007] According to another aspect of the embodiments of the present invention, a computer-readable storage medium is provided, in which a computer program is stored. The computer program is configured to execute the above-mentioned method for recognizing speech response time when running.
[0008] According to another aspect of an embodiment of the present invention, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the above-mentioned voice response time recognition method through the computer program.
[0009] In an embodiment of the present application, character recognition is performed on frame images in a target video to obtain a first frame image in which the first character of the first message appears, and a second frame image in which the last character of the first message appears. The target device's voice response time is then determined based on the timestamps corresponding to the first and second frames. This eliminates manual verification of the voice response time, significantly improves the recognition efficiency of the target device's voice response time, saves recognition time, and resolves the technical issue of low voice response time recognition efficiency in related technologies. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0011] Figure 1 is a schematic diagram of an application environment of an optional method for recognizing speech response time according to an embodiment of the present application;
[0012] Figure 2is a schematic diagram of an application environment of another optional method for recognizing speech response time according to an embodiment of the present application;
[0013] Figure 3 is a flow chart of an optional method for recognizing voice response time according to an embodiment of the present application;
[0014] Figure 4A is an image display schematic diagram of another optional voice response time recognition method according to an embodiment of the present application;
[0015] Figure 4B is a schematic diagram of an image display of an optional method for recognizing voice response time according to an embodiment of the present application;
[0016] Figure 5 This is a schematic diagram of an interface display of another optional method for recognizing voice response time according to an example of the present application;
[0017] Figure 6 1 is a flow chart of an optional method for recognizing speech response time according to an embodiment of the present application;
[0018] Figure 7 is a flowchart of another optional method for recognizing speech response time according to an embodiment of the present application;
[0019] Figure 8 is a flowchart of another optional method for recognizing speech response time according to an embodiment of the present application;
[0020] Figure 9 is a flowchart of another optional method for recognizing speech response time according to an embodiment of the present application;
[0021] Figure 10 is a flowchart of another optional method for recognizing voice response time according to an embodiment of the present application;
[0022] Figure 11 1 is a schematic diagram of an interface display of another optional method for recognizing voice response time according to an embodiment of the present application;
[0023] Figure 12 1 is a schematic diagram of an interface display of another optional method for recognizing voice response time according to an embodiment of the present application;
[0024] Figure 13 is a flowchart of another optional method for recognizing voice response time according to an embodiment of the present application;
[0025] Figure 14 is a flowchart of another optional method for recognizing voice response time according to an embodiment of the present application;
[0026] Figure 15 is a flowchart of another optional method for recognizing voice response time according to an embodiment of the present application;
[0027] Figure 16 is a schematic structural diagram of an optional device for recognizing sound response time according to an embodiment of the present application;
[0028] Figure 17 It is a schematic structural diagram of an optional electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0029] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0030] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0031] According to one aspect of the embodiment of the present application, a method for identifying a speech response time is provided. Optionally, as an optional implementation, the above-mentioned recognition of a speech response time can be applied to, but not limited to, Figure 1The application environment shown in FIG. This application environment includes a terminal device 102 for human-computer interaction with a user, a network 104, and a server 106. Terminal device 102 may include, but is not limited to, in-vehicle electronic devices, handheld terminals, wearable devices, portable devices, and the like. User 108 can interact with terminal device 102, which runs a voice response time recognition application client. Terminal device 102 includes a human-computer interaction screen 1022, a processor 1024, and a memory 1026. Human-computer interaction screen 1022 is used to display images recorded from the target device's display screen. The processor 1024 is used to obtain a target video obtained by recording the display screen of the target device, wherein the target video includes the picture displayed on the display screen when the target voice command is input to the target device; perform character recognition on the frame images in the target video to obtain a first frame image in which the first character in the first information appears, and a second frame image in which the last character in the first information appears, wherein the first information is all or part of the information obtained by the target device performing voice recognition on the target voice command and displayed on the display screen; determine the voice response time of the target device according to the first timestamp corresponding to the first frame image and the second timestamp corresponding to the second frame image; the memory 1026 is used to store the target video, and store the picture displayed on the display screen when the target voice command is input to the target device.
[0032] The specific process is as follows: Assume that Figure 1 As shown, a terminal device 102 is running a voice response time recognition application client. User 108 operates a human-computer interaction screen 1022 to manage and operate a virtual character. In step S102, a target video is obtained by recording the display screen of the target device. The target video includes the image displayed on the display screen when a target voice command is input into the target device. Then, step S104 is executed to transmit the target video to server 106 via network 104. Upon receiving the request, server 106 executes steps S106 and S108 to perform character recognition on the frames in the target video, obtaining a first frame image in which the first character of a first message appears, and a second frame image in which the last character of the first message appears. The first message is all or part of the message obtained by the target device through voice recognition of the target voice command and displayed on the display screen. A first timestamp corresponding to the first frame image and a second timestamp corresponding to the second frame image are determined. The first and second timestamps are used to determine the voice response time of the target device. And as in step S110, the terminal device 102 is notified via the network 104 and the determined voice response time is returned.
[0033] As another optional implementation, the above-mentioned voice response time recognition method of the present application can be applied to Figure 2 In the application environment shown. Figure 2 As shown, human-computer interaction can be performed between user 202 and user device 204. User device 204 includes memory 206 and processor 208. In this embodiment, user device 204 can refer to, but is not limited to, executing the operations performed by terminal device 102 to obtain the voice response time of the target device.
[0034] Optionally, in this embodiment, the terminal device 102 and the user device 204 may include but are not limited to at least one of the following: a mobile phone (such as an Android phone, an iOS phone, etc.), a laptop computer, a tablet computer, a PDA, an MID (Mobile Internet Devices), a PAD, a desktop computer, a smart TV, etc. The target client may be a video client, an instant messaging client, a browser client, an education client, etc. The network 104 may include but is not limited to: a wired network, a wireless network, wherein the wired network includes: a local area network, a metropolitan area network and a wide area network, and the wireless network includes: Bluetooth, WIFI and other networks that implement wireless communication. The server may be a single server, or a server cluster consisting of multiple servers, or a cloud server. The above is only an example and is not limited to this in this embodiment.
[0035] Alternatively, as an optional implementation, Figure 3 As shown, the above-mentioned method for recognizing speech response time includes:
[0036] S302, obtaining a target video obtained by recording a display screen of a target device, wherein the target video includes an image displayed on the display screen when a target voice command is input to the target device;
[0037] S304: Performing character recognition on the frame images in the target video to obtain a first frame image in which the first character of the first information appears, and a second frame image in which the last character of the first information appears, wherein the first information is part or all of the information obtained by the target device performing voice recognition on the target voice command and displayed on the display screen, the first information is the same as the second information, and the second information is the information represented by the target voice command;
[0038] It can be understood that in some examples, if the start and end time points of the target device's recognition of the target voice instruction are accurate, then when the target device accurately recognizes the content of the target voice instruction, the target information recognized by the target device is the second information. In this case, the first information can be all the information obtained by the target device through voice recognition of the target voice instruction and displayed on the display screen. Figure 4A As shown, the first information can be all text information obtained by the target device 400 performing voice recognition on the target voice command, and displayed on the screen of the target device 400, such as the text information displayed in the voice command display box 402, that is, the second information ("I want to watch XXX's movie").
[0039] In other examples, the start and end time points of the target device's recognition of the target voice command may be inaccurate. The start time point of the recognition may be too early (in this case, the target device may have recognized other characters before recognizing the character that is the same as the first character of the second information), or the end time point of the recognition may be too late (in this case, after the recognized characters include the second information, the target device may have recognized other characters). Figure 4B As shown, when the target device 400 recognizes the target voice instruction, the entire information displayed on the display screen is "I want to watch XXX's movie at this time". Then, when the target device accurately recognizes the content of the target voice instruction, the target information recognized by the target device includes the second information as well as other information ("at this time"). At this time, the first information is only part of the information obtained by the target device 400 through voice recognition of the target voice instruction and displayed on the display screen. Figure 4B As shown, the first information can be the partial text information obtained by the target device 400 performing voice recognition on the target voice command, and displayed on the screen of the target device 400, such as the partial text information displayed in the voice command display box 402, that is, the second information ("I want to watch XXX's movie").
[0040] S306: Determine a first timestamp corresponding to the first frame image and a second timestamp corresponding to the second frame image, wherein the first timestamp and the second timestamp are used to determine a voice response time of the target device.
[0041] In step S302, in actual application, the display screen of the target device may be recorded through devices such as mobile phones, laptops, tablet computers, PDAs, MIDs (Mobile Internet Devices), PADs, and desktop computers. The target device may include electronic devices such as televisions, mobile phones, and laptops, which are not limited here. In this embodiment, for example, Figure 4AAs shown, the target video can be obtained by recording the display screen of the target device 400 through the mobile device 404. The above target video includes the picture displayed on the display screen when a voice command (the text "I want to watch a movie of XXX" in the voice command display box 402) is input to the target device 400.
[0042] Optionally, in one embodiment, the target device can be multiple electronic devices; for example, as Figure 5 shown, the target video is obtained by recording the display screens of the first target device 502 and the second target device 504 through the mobile device 500. The target video includes the display screens of both the first target device 502 and the second target device 504.
[0043] In step S304, in practical applications, performing character recognition on the frame images in the target video may include, but is not limited to, recognizing the text characters in the frame images through optical character recognition (OCR). As Figure 4A shown, in this embodiment, the first information can be the text information obtained by performing voice recognition on the target voice command by the target device 400 and displayed on the screen of the target device 400, such as the text information displayed in the voice command display box 402 (such as "I want to watch a movie of XXX"). As Figure 6 shown, after performing OCR recognition on multiple frame images in the target video, the characters included in each frame image can be obtained, and the frame image where the first character appears is recorded as the first frame image, and the frame image where the last character appears is recorded as the second frame image. In Figure 6 it, the first frame image is the frame image corresponding to 1.2 s in the target video, and this image includes the first character "I" of the first information; the second frame image is the frame image corresponding to the 6th s in the target video, and this image includes the last character "movie" of the first information.
[0044] In step S306, in practical applications, according to the first timestamp corresponding to the first frame image and the second timestamp corresponding to the above second frame image, determine the voice response time of the above target device. As Figure 6 shown, the timestamp corresponding to the first frame image is 1.2 s, and the timestamp corresponding to the first frame image is 6 s. The voice response time of the target device can be obtained through the time difference between the above two timestamps; for example, in this embodiment, the voice command "I want to watch a movie of XXX" contains 9 characters, and the response time of the above characters displayed on the screen of the target device is 4.8 s, then the time for one character response output can be obtained as 0.53 s.
[0045] In one or more embodiments, as Figure 6As shown, the above-mentioned method for identifying voice response time includes: after multiple frames of images in the target video are recognized by OCR, the text information display area in each frame image and the timestamp corresponding to each frame image are obtained, and the characters corresponding to the frame number, frame timestamp and text information display area in each frame image are recorded in a preset table.
[0046] like Figure 7 As shown, after OCR recognition, the text information display area 702 in the image of the first frame is obtained. After the displayed text information is empty, the frame number of the image of the first frame, the frame timestamp 0.1s, and the corresponding text information are recorded.
[0047] like Figure 8 As shown, after OCR recognition, the text information display area 802 in the image of the second frame is obtained. After the displayed text information is empty, the frame number of the image of the second frame, the frame timestamp 1.2s, and the corresponding text information "I" are recorded.
[0048] like Figure 9 As shown, after OCR recognition, the text information display area 902 in the image of the 12th frame is obtained. After the displayed text information is empty, the frame number of the image of the 12th frame, the frame timestamp 2s, and the corresponding text information "I think" are recorded.
[0049] like Figure 10 As shown, after OCR recognition, the text information display area 1002 in the image of the 20th frame is obtained. After the displayed text information is empty, the frame number of the image of the 12th frame, the frame timestamp 6s, and the corresponding text information "I want to watch XXX's movie" are recorded.
[0050] In one or more embodiments, Figure 11 As shown, it may include but is not limited to identifying the text information display area 1102 in the frame image in the target device through OCR to obtain the position of the displayed text information corresponding to the voice command.
[0051] In one or more embodiments, Figure 12 As shown, the above-mentioned voice response time recognition method includes: simultaneously recognizing the voice response time for different target devices, wherein the text information display area 1202a of target device 1102, the text information display area 1204a of target device 1204, and the text information display area 1206a of target device 1206 all display text information corresponding to the voice command "I want to watch XXX's movie." Target device 1202 displays the display screen of the target application, while target devices 1204 and 1206 display the display screens of the television screen.
[0052] In the embodiment of the present application, character recognition is performed on frame images in a target video to obtain a first frame image in which the first character of the first message appears, and a second frame image in which the last character of the first message appears. The target device's voice response time is then determined based on the timestamps corresponding to the first and second frames. This eliminates manual verification of the voice response time, significantly improves the recognition efficiency of the target device's voice response time, saves recognition time, and resolves the technical issue of low voice response time recognition efficiency in related technologies.
[0053] In one or more embodiments, step S304, performing character recognition on the frame images in the target video to obtain the first frame image in which the first character in the first information appears, and the second frame image in which the last character in the first information appears, includes: performing character recognition on a preset target display area in the frame images in the target video, respectively identifying the first character and the last character and determining the first frame image in which the first character appears and the second frame image in which the last character appears, wherein the target display area is the area in which the first information is displayed.
[0054] In this embodiment, if Figure 4A As shown, the target display area may include but is not limited to the area corresponding to the voice instruction display frame 402. The first information (I want to watch XXX's movie) is displayed in the target display area.
[0055] Through one or more embodiments provided by the present application, by performing character recognition on a preset target display area in a frame image in the target video, the first character and the last character are respectively identified, and the first frame image in which the first character appears and the second frame image in which the last character appears are determined. There is no need to perform character recognition on the entire screen of the frame image of the target video, which narrows the scope of character recognition and further improves the efficiency of voice command text information recognition.
[0056] In one or more embodiments, the above-mentioned character recognition is performed on the preset target display area in the frame image in the target video, and the above-mentioned first character and the above-mentioned last character are respectively recognized, and before the above-mentioned first frame image in which the above-mentioned first character appears and the above-mentioned second frame image in which the above-mentioned last character appears, it also includes: determining a preset area in the above-mentioned target device, wherein the above-mentioned preset area is used to display information obtained by the above-mentioned target device performing voice recognition on the input voice command; and determining the above-mentioned preset area as the above-mentioned target display area.
[0057] For example, the image size of the frame image in the target video is obtained; a preset area in the frame image is determined according to the image size, and the range of the preset area is marked, such as by adding a rectangular or elliptical border; and the preset area is determined as the target display area.
[0058] By determining a preset area for displaying the information obtained by the target device through voice recognition of the input voice command, the present application can accurately obtain the characters displayed by the voice command on the target device, thereby improving the efficiency of voice command text information recognition.
[0059] In one or more embodiments, the above-mentioned character recognition is performed on a preset target display area in a frame image in a target video, the first character is recognized, and the first frame image in which the first character appears is determined, including: determining whether the character recognized in the current frame image includes the first character in the second information, wherein the frame image in the target video includes the current frame image; when the recognized character includes the first character in the second information and the first character in the second information is not recognized in the frame image before the current frame image, the character in the recognized character that is the same as the first character in the second information is determined as the first character in the first information, and the current frame image is determined as the first frame image.
[0060] In this embodiment, if Figure 6 As shown, if the current frame in the target video is the frame image corresponding to the first frame, and no characters are recognized in the target display area 602 in the frame image corresponding to the first frame, when the next frame image of the first frame, that is, the target display area 604 of the image corresponding to the second frame recognizes characters (characters "I"), then the first character (character "I") of the above-mentioned recognized characters in the image corresponding to the second frame in the target video can be determined as the first character (character "I"), and the image corresponding to the second frame in the target video can be determined as the first frame image.
[0061] Through one or more embodiments provided by the present application, when the recognized characters include the first character in the second information and the first character in the second information is not recognized in the frame image before the current frame image, the character in the recognized characters that is the same as the first character in the second information is determined as the first character in the first information, and the current frame image is determined as the first frame image. The first character of the voice command in the target display area can be accurately obtained, and the frame image corresponding to the first character can be obtained without manual labeling, which further improves the efficiency of voice command text information recognition.
[0062] In one or more embodiments, the character recognition of the preset target display area in the frame image of the target video is performed to recognize the tail character, and the second frame image in which the tail character appears is determined, including: when the characters recognized in the target display area in N consecutive frame images in the target video are the same and the characters included in the first information and the second information are the same, determining the last character recognized in one of the N consecutive frame images as the tail character, and determining the frame image in which the tail character first appears as the second frame image; wherein, the second information is the information represented by the target voice command, and N is a natural number greater than or equal to 2.
[0063] In this embodiment, as Figure 6 shown, when the characters recognized in the target display area in the 30th to 33rd frames of the target video are the same, such as all being "I want to watch XXX movies", the last character recognized in one of the 4 consecutive frame images can be determined as the tail character (such as the character "movie"), and the frame image corresponding to the 30th frame, in which the tail character first appears, is determined as the second frame image.
[0064] Through one or more embodiments provided by the present application, when it is determined that the characters recognized in multiple consecutive frame images do not change, the frame image in which the tail character first appears is used as the first frame image, so that the tail character of the voice command in the target display area can be accurately obtained, and the frame image corresponding to the tail character can be obtained, without manual annotation, further improving the efficiency of recognizing the text information of the voice command.
[0065] In one or more embodiments, the character recognition of the preset target display area in the frame image of the target video is performed to recognize the tail character, and the second frame image in which the tail character appears is determined, including: determining whether the characters recognized in the current frame image include the second information, where the frame images in the target video include the current frame image; when the recognized characters include the second information and the tail character in the second information is not recognized in the frame image before the current frame image, determining the character in the recognized characters that is the same as the tail character in the second information as the tail character in the first information, and determining the current frame image as the second frame image.
[0066] In this embodiment, as Figure 6As shown, when the characters recognized in the current frame image (such as the 30th frame image) include the above-mentioned second information ("I want to watch XXX movies"), and the tail character ("movie") in the above-mentioned second information is not recognized in the frame image (the 29th frame image) before the 30th frame image, the character in the recognized characters that is the same as the tail character in the above-mentioned second information is determined as the tail character in the above-mentioned first information, that is, the character "movie" is determined as the tail character in the above-mentioned first information, and the 30th frame image is determined as the second frame image.
[0067] Through one or more embodiments provided by the present application, when the recognized characters include the above-mentioned second information and the tail character in the above-mentioned second information is not recognized in the frame image before the current frame image, the tail character and the second frame image can be obtained, so that the tail character of the voice command in the target display area can be accurately obtained, and the frame image corresponding to the tail character can be obtained, without manual annotation, further improving the efficiency of recognizing the text information of the voice command. In one or more embodiments, the method for recognizing the above-mentioned voice response time further includes: when the first character recognized in the above-mentioned target video is the same as the first character in the second information, obtaining the first timestamp corresponding to the first frame image, where the above-mentioned second information is the true information represented by the above-mentioned target voice command, that is, if the target device recognizes information that is exactly the same as the second information, then the recognition of the target voice command by the target device is accurate; when the above-mentioned tail character is recognized in the above-mentioned target video and the first information is the same as the above-mentioned second information, obtaining the second timestamp corresponding to the second frame image.
[0068] In this embodiment, as Figure 6 shown, when the first character (such as the character "I") recognized by OCR in the target video is the same as the first character in the second information ("I want to watch XXX movies"), the obtained first timestamp corresponding to the first frame image is 1.2s; when the tail character (such as the character "movie") is recognized by OCR in the target video, and the first information that appears in the Nth frame of the target recognition, and the characters included in the target voice command are all "I want to watch XXX movies", the obtained second timestamp corresponding to the Nth frame image is 6s.
[0069] Through one or more embodiments provided by the present application, when the first character recognized in the above-mentioned target video is the same as the first character in the second information, obtaining the first timestamp corresponding to the first frame image; when the above-mentioned tail character is recognized in the above-mentioned target video and the first information is the same as the above-mentioned second information, obtaining the second timestamp corresponding to the second frame image, the first timestamp corresponding to the first character and the second timestamp corresponding to the tail character can be accurately obtained, further improving the efficiency of recognizing the text information of the voice command.
[0070] In one or more embodiments, the method for identifying the above-mentioned voice response time further includes: when the first character recognized in the above-mentioned current frame is different from the first character in the second information and no character is recognized in the frame image before the above-mentioned current frame image, stopping character recognition for the frame images in the above-mentioned target video; wherein, the frame images in the above-mentioned target video include the above-mentioned current frame image. For example, when the first character in the second information is "I" and the first character recognized in the current frame of the target video is "Luo", stop character recognition for the frame images in the above-mentioned target video.
[0071] Through one or more embodiments provided by the present application, when the first character recognized in the above-mentioned target video is different from the first character in the second information, stopping character recognition for the frame images in the above-mentioned target video can save system computing resources and improve the efficiency of identifying voice command text information.
[0072] In one or more embodiments, after step S306, the method for identifying the above-mentioned voice response time further includes: determining the start response time of the above-mentioned target device measured based on the above-mentioned target video based on the above-mentioned first timestamp; recording the time interval between the above-mentioned first timestamp and the above-mentioned second timestamp corresponding to the above-mentioned target video, and the above-mentioned start response time corresponding to the above-mentioned target video, wherein the above-mentioned voice response time includes the above-mentioned time interval and the above-mentioned start response time.
[0073] In the actual process of detecting the voice response time, since the target device will go through the recognition and calculation of the processor after receiving the voice command, there is usually a time difference between when the target device receives the voice command and when the first character of the voice command appears; in the present application, as Figure 6 shown, for example, after starting recording, when the target device receives the voice command, at 0.1s, the first frame image in the target video does not show the character information corresponding to the voice command, and the second frame at 1.2s shows the first character corresponding to the voice command. The start response time of the above-mentioned target device can be 0.1s. The above-mentioned voice response time includes the time interval between the first timestamp and the second timestamp (i.e., 4.8s) and the above-mentioned start response time of 0.1s. The above-mentioned voice response time is 5.9s.
[0074] Through one or more embodiments provided herein, the voice response time can be determined by the time interval between the first and second timestamps corresponding to the target video, and the start response time corresponding to the target video. For example, in one embodiment, the time interval between the first and second timestamps corresponding to the target video, and the start response time corresponding to the target video, can be recorded, and the sum of the time interval and the start response time can be determined as the target device's voice response time for the target video. This allows for automated and relatively accurate acquisition of the time from when a voice command is issued to when the voice command completes a response on the target device, improving the efficiency of voice response time recognition.
[0075] In addition, in some embodiments, the voice response time can be determined solely by the start response time. For example, in one embodiment, when the recognized characters include the second information and the last character of the second information is not recognized in the frame image before the current frame image, the start response time can be obtained as the voice response time to represent the speed from the issuance of the voice command to the start of the target device's response. This allows the speed from the issuance of the voice command to the start of the voice command response on the target device to be automatically and more accurately obtained.
[0076] In one or more embodiments, the determining of the start response time of the target device measured based on the target video based on the first timestamp includes: determining the start response time of the target device measured based on the target video based on the first timestamp and the target time of inputting the target voice command to the target device; or determining the first timestamp as the start response time voice response time. Figure 6 For example, after starting to record a target video and receiving a voice command, the target device may not display the character information corresponding to the voice command in the first frame of the target video at 0.1 seconds. However, the first character corresponding to the voice command is displayed in the second frame at 1.2 seconds. The start response time of the target device may be 0.1 seconds. Furthermore, to improve the detection efficiency of the voice recognition time, the first timestamp 1.2 may be determined as the start response time or the voice response time.
[0077] In one or more embodiments, determining the voice response time of the target device according to the first timestamp corresponding to the first frame image and the second timestamp corresponding to the second frame image includes: determining the voice response time as the time interval between the first timestamp and the second timestamp; Figure 6 As shown, the time interval of 4.8s between the first timestamp, ie, the time 1.2s corresponding to the image of the second frame, and the second timestamp, ie, the time 6s corresponding to the image of the Nth frame, is determined as the time of the voice response.
[0078] Optionally, in one embodiment, a determination is made as to whether the first information and the second information are identical, wherein the second information is the information represented by the target voice command; if the first information and the second information are identical, the voice response time is determined to be the time interval between the first timestamp and the second timestamp. For example, if it is determined that both the first information and the second information contain the text "I want to watch XXX's movie," the time interval between the first timestamp (the time when the first character appears) and the second timestamp (the time when the last character appears) is determined to be the voice response time.
[0079] Through one or more embodiments provided herein, by determining the voice response time as the time interval between the first and second timestamps, or determining whether the first and second information are identical, and if the first and second information are identical, determining the voice response time as the time interval between the first and second timestamps, the time when the voice command is displayed on the target device can be accurately obtained, thereby improving the efficiency of voice command text information recognition.
[0080] In one or more embodiments, the voice response time of the target device is determined based on the first timestamp corresponding to the first frame image and the second timestamp corresponding to the second frame image, including: when the target video includes multiple videos, the time interval between the first timestamp and the second timestamp obtained in each video is determined as the candidate voice response time corresponding to each video; and the average value of the candidate voice response times corresponding to each video is determined as the voice response time of the target device.
[0081] In this embodiment, if Figure 4A As shown, when the target video includes multiple videos, that is, the target device 400 is recorded multiple times, and the average value of the candidate voice response time corresponding to each video is determined as the voice response time of the above-mentioned target device, for example, the target video includes 5 videos, and the average value of the candidate voice response time for each video is 5s, then the average value can be used as the voice response time of the target device 400.
[0082] Through one or more embodiments provided herein, the average of the candidate voice response times corresponding to each video is determined as the voice response time of the target device. This avoids errors caused by single recognition, accurately obtains the time when the voice command is displayed on the target device, and improves the efficiency of voice command text information recognition.
[0083] In one or more embodiments, the voice response time recognition method further includes: timestamping the acquired target video using a video processing tool, where the video processing tool includes, but is not limited to, OpenCV or FFmpeg. The technical means provided herein can accurately capture the time stamp of the image corresponding to the voice command on the target device, thereby improving the efficiency of voice command text information recognition.
[0084] It should be noted that for the aforementioned method embodiments, for simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that the present invention is not limited by the order of the actions described, because according to the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present invention.
[0085] According to another aspect of the embodiment of the present invention, the above-mentioned method for recognizing speech response time includes the following steps: Figure 13 The steps shown are: step S1302, recording batch videos through a mobile phone; step S1304, finding the text recognition display area of the target device through a video processing tool; step S1306, processing the acquired video and adding a timestamp through the video processing tool; step S1308, inputting the text recognition display areas of multiple videos in sequence, recording the start and end time points of the text recognition display areas through the OCR algorithm, and then obtaining the voice speed.
[0086] In one or more embodiments, Figure 14 As shown, the above-mentioned method for identifying voice response time also includes the following steps: Step S1402, processing the target video and adding a timestamp, that is, detecting the text area of each device and adding a timestamp to the recorded video through video processing tools; then determining the text display area in the target video. Step S1404, framing the video through software such as opencv or ffmpeg. Step S1406, performing text recognition processing on the screen area of the text in the framed image; that is, recognizing the characters in the framed image through the text recognition module. Step S1408, respectively recording the time points when the first and last characters in the framed image appear.
[0087] Optionally, in one embodiment, multiple target devices may be detected through steps S1402 to S1408, and the recognition data of each device may be aggregated to calculate the average time points for first and last word recognition for each device. The recognition response speed of each device may be compared based on the average time points for the first and last words.
[0088] In one or more embodiments, the above step S1404 specifically includes the following processes: step S1502, pre-processing the framed image, including methods such as image smoothing, layout analysis, and tilt correction; step S1504, detecting the text area in the processed image, that is, finding the text area containing characters; step S1506, performing image binarization processing on the above text area; step S1508, performing character segmentation on the binarized text area or text line to separate individual characters. Step S1510, converting the input character dot matrix image into text for text processing; during the text processing process, first compare the first character of the voice command with the recognized character to see if they are consistent. If they are consistent, record the time point of the first character recognition; then compare the recognized text with the entire voice command to see if they are consistent. If they are consistent, record the time point as the last character recognition time point.
[0089] In an embodiment of the present invention, character recognition is performed on frame images in a target video to obtain a first frame image in which the first character of the first message appears, and a second frame image in which the last character of the first message appears. The target device's voice response time is then determined based on the timestamps corresponding to the first and second frames. This eliminates manual verification of the voice response time, significantly improves the efficiency of voice response time recognition on the target device, saves recognition time, and resolves the technical issue of low voice response time recognition efficiency in related technologies.
[0090] According to another aspect of the embodiment of the present invention, there is also provided a speech response time recognition device for implementing the above-mentioned speech response time recognition method. Figure 16 As shown, the device includes:
[0091] An acquiring unit 1602 is configured to acquire a target video obtained by recording a display screen of a target device, wherein the target video includes a picture displayed on the display screen when a target voice command is input to the target device;
[0092] a recognition unit 1604 configured to perform character recognition on the frame images in the target video to obtain a first frame image in which the first character of the first information appears, and a second frame image in which the last character of the first information appears, wherein the first information is information obtained by the target device performing voice recognition on the target voice command and displayed on the display screen;
[0093] The determining unit 1606 is configured to determine a voice response time of the target device according to a first timestamp corresponding to the first frame image and a second timestamp corresponding to the second frame image.
[0094] In step S302, in practical applications, it may include, but is not limited to, recording the display screen of the target device through devices such as mobile phones, laptop computers, tablet computers, personal digital assistants, MID (Mobile Internet Devices), PADs, desktop computers, etc. The target device may include electronic devices such as televisions, mobile phones, laptop computers, etc., which are not limited herein. In this embodiment, for example, as Figure 4A or Figure 4B shown, the display screen of the target device 400 can be recorded by the mobile device 404 to obtain a target video. The above target video includes the画面displayed on the display screen when a voice command (the text "I want to watch a movie of XXX" in the voice command display box 402) is input to the target device 400.
[0095] In one or more embodiments, the target device can be multiple electronic devices; for example, as Figure 5 shown, the display screens of the first target device 502 and the second target device 504 are recorded by the mobile device 500 to obtain a target video, and the target video includes the display screens of both the first target device 502 and the second target device 504.
[0096] In one or more embodiments, performing character recognition on the frame images in the target video may include recognizing the text characters in the frame images through optical character recognition (OCR). As Figure 4A shown, in this embodiment, the first information can be obtained by performing voice recognition on the target voice command by the target device 400 and is the text information (such as "I want to watch a movie of XXX") displayed on the screen of the target device 400, such as the text information displayed in the voice command display box 402. As Figure 6 shown, after performing OCR recognition on multiple frame images in the target video, the characters included in each frame image can be obtained, and the frame image where the first character appears is recorded as the first frame image, and the frame image where the last character appears is recorded as the second frame image. In Figure 6 , the first frame image is the frame image corresponding to 1.2 s in the target video, and this image includes the first character "I" of the first information; the second frame image is the frame image corresponding to the 6th s in the target video, and this image includes the last character "movie" of the first information.
[0097] In one or more embodiments, according to the first timestamp corresponding to the first frame image and the second timestamp corresponding to the second frame image, determine the voice response time of the above target device; as Figure 6As shown, the timestamp corresponding to the first frame image is 1.2s, and the timestamp corresponding to the second frame image is 6s. The voice response time of the target device can be obtained by the direct time difference between the above two timestamps; for example, in this embodiment, the voice instruction "I want to watch XXX's movie" contains 9 characters, and the response time of the above characters displayed on the screen of the target device is 4.8s, so it can be obtained that 1.875 characters can be responded and output per second.
[0098] In the embodiment of the present application, character recognition is performed on frame images in a target video to obtain a first frame image in which the first character of the first message appears, and a second frame image in which the last character of the first message appears. The target device's voice response time is then determined based on the timestamps corresponding to the first and second frames. This eliminates manual verification of the voice response time, significantly improves the recognition efficiency of the target device's voice response time, saves recognition time, and resolves the technical issue of low voice response time recognition efficiency in related technologies.
[0099] According to another aspect of the embodiment of the present invention, an electronic device for implementing the above-mentioned voice response time recognition method is also provided. The electronic device may be Figure 1 The terminal device or server shown in FIG. This embodiment is described by taking the electronic device as a server as an example. Figure 17 As shown, the electronic device includes a memory 1702 and a processor 1704. The memory 1702 stores a computer program, and the processor 1704 is configured to execute the steps in any of the above method embodiments through the computer program.
[0100] Optionally, in this embodiment, the electronic device may be located in at least one network device among a plurality of network devices of a computer network.
[0101] Optionally, in this embodiment, the processor may be configured to execute the following steps through a computer program:
[0102] S1, obtaining a target video obtained by recording a display screen of a target device, wherein the target video includes a picture displayed on the display screen when a target voice command is input to the target device;
[0103] S2, performing character recognition on the frame images in the target video to obtain a first frame image in which the first character of the first information appears, and a second frame image in which the last character of the first information appears, wherein the first information is information obtained by the target device performing voice recognition on the target voice command and displayed on the display screen;
[0104] S3, determining the voice response time of the target device according to the first timestamp corresponding to the first frame image and the second timestamp corresponding to the second frame image.
[0105] Alternatively, those skilled in the art will appreciate that Figure 17 The structure shown is for illustration only, and the electronic device may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (MID), a PAD, or other terminal devices. Figure 17 It does not limit the structure of the electronic device. For example, the electronic device may also include Figure 17 More or fewer components (such as network interfaces, etc.) as shown in, or with Figure 17 Different configurations shown.
[0106] Among them, the memory 1702 can be used to store software programs and modules, such as the program instructions / modules corresponding to the method and device for recognizing voice response time in the embodiment of the present invention. The processor 1704 executes various functional applications and data processing by running the software programs and modules stored in the memory 1702, that is, realizing the above-mentioned method for recognizing voice response time. The memory 1702 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 1702 may further include a memory remotely located relative to the processor 1704, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. Among them, the memory 1702 can be used specifically but not limited to target video and other information. As an example, such as Figure 17 As shown, the memory 1702 may include, but is not limited to, the acquisition unit 1602, the recognition unit 1604, and the determination unit 1606 in the voice response time recognition device. In addition, it may also include, but is not limited to, other module units in the voice response time recognition device, which will not be repeated in this example.
[0107] Optionally, the transmission device 1706 is used to receive or send data via a network. Specific examples of the network may include a wired network and a wireless network. In one embodiment, the transmission device 1706 includes a network interface controller (NIC), which can be connected to other network devices and a router via a network cable to communicate with the Internet or a local area network. In one embodiment, the transmission device 1706 is a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0108] In addition, the electronic device further includes: a display 1708 for displaying the target video; and a connection bus 1710 for connecting various module components in the electronic device.
[0109] In other embodiments, the terminal device or server may be a node in a distributed system, wherein the distributed system may be a blockchain system, and the blockchain system may be a distributed system formed by connecting multiple nodes through network communication. The nodes may form a peer-to-peer (P2P) network, and any computing device, such as a server, terminal, or other electronic device, may become a node in the blockchain system by joining the peer-to-peer network.
[0110] In one or more embodiments, the present application further provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the above-described method for recognizing speech response time. The computer program is configured to execute the steps of any of the above-described method embodiments when executed.
[0111] Optionally, in this embodiment, the computer-readable storage medium may be configured to store a computer program for performing the following steps:
[0112] S1, obtaining a target video obtained by recording a display screen of a target device, wherein the target video includes a picture displayed on the display screen when a target voice command is input to the target device;
[0113] S2, performing character recognition on the frame images in the target video to obtain a first frame image in which the first character of the first information appears, and a second frame image in which the last character of the first information appears, wherein the first information is information obtained by the target device performing voice recognition on the target voice command and displayed on the display screen;
[0114] S3, determining the voice response time of the target device according to the first timestamp corresponding to the first frame image and the second timestamp corresponding to the second frame image.
[0115] Alternatively, in this embodiment, a person skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be performed by instructing hardware related to the terminal device through a program. The program may be stored in a computer-readable storage medium, which may include a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk. The serial numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the merits or demerits of the embodiments.
[0116] If the integrated units in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in the above-mentioned computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the existing technology, or all or part of the technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes a number of instructions for causing one or more computer devices (such as personal computers, servers, or network devices) to execute all or part of the steps of the methods described in various embodiments of the present invention.
[0117] In the above embodiments of the present invention, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0118] In the several embodiments provided in this application, it should be understood that the disclosed client can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, and can be electrical or other forms.
[0119] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0120] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0121] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.
Claims
1. A method for recognizing speech response time, characterized in that: include: Acquire a target video obtained by recording a display screen of a target device, wherein the target video includes a picture displayed on the display screen when a target voice command is input to the target device; Performing character recognition on the frame image in the target video to obtain first information, wherein the first information is all information displayed on the display screen by the target device; comparing the first information with second information actually represented by the target voice command to obtain a comparison result, and if the comparison result indicates that the first information and the second information are identical, determining a first frame image in which the first character of the first information appears and a second frame image in which the last character of the first information appears; Determine a first timestamp corresponding to the first frame of image and a second timestamp corresponding to the second frame of image; Determining a start response time of the target device based on the first timestamp, where the start response time is used to indicate a time when the target device starts to display characters; A voice response time of the target device is determined based on a time interval between the first timestamp and the second timestamp and the start response time.
2. The method according to claim 1, characterized in that When the comparison result indicates that the first information and the second information are identical, determining a first frame image in which the first character in the first information appears and a second frame image in which the last character in the first information appears, includes: Character recognition is performed on a preset target display area in the frame image in the target video, the first character and the last character are respectively identified, and the first frame image in which the first character appears and the second frame image in which the last character appears are determined, wherein the target display area is an area for displaying the first information.
3. The method according to claim 2, characterized in that The method further comprises: performing character recognition on a preset target display area in a frame image in the target video, respectively recognizing the first character and the last character, and determining before the first frame image in which the first character appears and the second frame image in which the last character appears: Determining a preset area in the target device, wherein the preset area is used to display information obtained by the target device performing voice recognition on an input voice command; The preset area is determined as the target display area.
4. The method according to claim 1, wherein When the comparison result indicates that the first information and the second information are identical, determining a first frame image in which the first character in the first information appears and a second frame image in which the last character in the first information appears includes: When the recognized character includes the first character in the second information and the first character in the second information is not recognized in the frame image before the current frame image, the character in the recognized character that is the same as the first character in the second information is determined as the first character in the first information, and the current frame image is determined as the first frame image.
5. The method according to claim 1, wherein When the comparison result indicates that the first information and the second information are identical, determining a first frame image in which the first character in the first information appears and a second frame image in which the last character in the first information appears, further comprising: When the recognized characters include the second information and the last character in the second information is not recognized in the frame image before the current frame image, the character in the recognized characters that is the same as the last character in the second information is determined as the last character in the first information, and the current frame image is determined as the second frame image.
6. The method according to claim 1, wherein The method further comprises: When the first character recognized in the current frame image is different from the first character in the second information and no character is recognized in the frame image before the current frame image, stop character recognition on the frame images in the target video, wherein the frame images in the target video include the current frame image.
7. The method according to claim 1, characterized in that The voice response time is used to represent the voice response speed of the target device.
8. The method according to claim 7, characterized in that The determining the voice response time of the target device based on the time interval between the first timestamp and the second timestamp and the start response time includes: Recording the time interval between the first timestamp and the second timestamp corresponding to the target video, and the start response time corresponding to the target video; The sum of the time interval corresponding to the target video and the start response time is determined as the voice response time of the target device to the target voice command.
9. The method according to claim 7, characterized in that The determining, based on the first timestamp, a start response time of the target device includes: determining, based on the first timestamp and a target time at which the target voice command is input to the target device, the start response time of the target device measured based on the target video; or The first timestamp is determined as the start response time.
10. A device for recognizing speech response time, characterized in that: include: an acquisition unit, configured to acquire a target video obtained by recording a display screen of a target device, wherein the target video includes a picture displayed on the display screen when a target voice command is input to the target device; a recognition unit configured to perform character recognition on the frame images in the target video to obtain first information, wherein the first information is all information displayed on the display screen of the target device; compare the first information with second information actually represented by the target voice command to obtain a comparison result, and if the comparison result indicates that the first information and the second information are the same, determine a first frame image in which the first character of the first information appears, and a second frame image in which the last character of the first information appears; a determining unit, configured to determine a first timestamp corresponding to the first frame of image and a second timestamp corresponding to the second frame of image; The device is also used to, after determining a first timestamp corresponding to the first frame image and a second timestamp corresponding to the second frame image, determine the start response time of the target device based on the first timestamp, the start response time being used to indicate the time when the target device starts displaying characters; and determine the voice response time of the target device based on the time interval between the first timestamp and the second timestamp, and the start response time.
11. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein the program is executed by a processor to perform the method according to any one of claims 1 to 9.
12. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the method according to any one of claims 1 to 9 are implemented.
13. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 9 through the computer program.
Citation Information
Patent Citations
Speech recognition test method, device and system
CN110335590A
Text information extraction method and device, electronic equipment and storage medium
CN112101353A
Voice instruction execution time recognition method and device and electronic equipment
CN115348476A