Method for real-time processing of dictation content and related products thereof
By acquiring and processing dictation images in real time, the problem of difficulty in associating dictation results with audio tasks in the prior art is solved, real-time recognition and correction with high accuracy is achieved, and the user experience and efficiency of dictation technology is improved.
Patent Information
- Application Number
- CN202111478647.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-06
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2041-12-06
AI Technical Summary
The existing dictation technology cannot recognize and correct the dictation results corresponding to each audio task in real time, resulting in unsatisfactory recognition and correction accuracy, and the association between the dictation results and audio cannot be traced back, and it is easy to have mismatch.
By obtaining dictation images in real time during the audio broadcast task, identifying dictation results and correcting them, using the time and spatial characteristics of the image information for correlation processing, including real-time acquisition of writing trajectory and timing position information of the target part, and using text detection and OCR technology for accurate identification.
Real-time correlation processing between audio tasks and dictation results is realized, the accuracy of recognition and correction is improved, mismatch is avoided, and user experience and dictation efficiency is improved.
Smart Images

Figure CN114187592B_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of information processing technology. More specifically, embodiments of the present invention relate to a method for real-time processing of dictation content, an apparatus for real-time processing of dictation content, a device for executing the aforementioned method, and a computer-readable storage medium. Background Art
[0002] This section is intended to provide background or context for embodiments of the present invention as recited in the claims. The description herein may include concepts that could be explored, but not necessarily concepts that have been previously conceived or explored. Therefore, unless otherwise indicated herein, the material described in this section is not prior art with respect to the specification and claims of this application and is not admitted to be prior art by inclusion in this section.
[0003] With the development of a new wave of artificial intelligence, AI (particularly deep learning) is profoundly impacting many aspects of our lives, production, and learning. Computer vision is considered a highly successful application of deep learning. Technologies such as facial recognition, smart security, and optical character recognition (OCR) have brought significant convenience and safety. In particular, the application of deep learning to intelligent learning hardware (such as smart learning lamps or tablets) can provide students with rich and effective learning assistance features (such as dictation training).
[0004] The dictation technology used by smart learning hardware in related technologies usually involves taking a photo of the completed workbook, then recognizing all the new words in the dictation results, and finally correcting all the dictation recognition content. More specifically, it involves the following process:
[0005] 1. The intelligent learning hardware needs to announce all the new words in the dictation vocabulary list at once, and the user writes the dictation results based on the voice announcement.
[0006] 2. After all the new words in this dictation are dictated, the camera will be turned on to take a photo of the dictation results, which will then be input into the text detection and recognition algorithm model to identify the content of the user's dictation.
[0007] 3. Match the recognized dictation results with the list of new words broadcast by the intelligent learning hardware for correction.
[0008] It can be seen that the existing dictation technology cannot accurately correspond the audio of unfamiliar words in the broadcast with the dictation results. It relies on taking pictures and recognizing the dictation book together after all unfamiliar words have been dictated. This method loses the time information of the dictation results, resulting in the inability to trace back to which broadcast audio corresponds to a certain dictation result. In addition, if there is an inclusion relationship between two dictations of unfamiliar words, there may be incorrect matches between the two dictation results and the two broadcast audios, which will interfere with the correction and ultimately affect the correction result. Summary of the Invention
[0009] The known recognition and correction effects of dictation content are not ideal, which is a very troublesome process.
[0010] For this reason, there is a great need for an improved solution and related products for real-time processing of dictation content, which can recognize and correct the dictation results corresponding to each audio task in real time, thereby effectively improving the accuracy of recognition and correction.
[0011] In this context, embodiments of the present invention are expected to provide a solution and related products for real-time processing of dictation content.
[0012] In the first aspect of the embodiments of the present invention, a method for real-time processing of dictation content is provided, including: during the broadcast of one or more audio tasks in the dictation content, obtaining in real time the dictation image corresponding to each audio task; recognizing from the dictation image the dictation result corresponding to the audio task; and correcting the recognized dictation result.
[0013] In an embodiment of the present invention, obtaining in real time the dictation image corresponding to each audio task includes: within a predetermined time after the broadcast of each audio task, obtaining in real time the image information of the content presented on the output medium through the input medium.
[0014] In another embodiment of the present invention, recognizing from the dictation image the dictation result corresponding to the audio task includes: according to the image information, obtaining the writing trajectory of the target part, where the target part is the part where the input medium contacts the output medium; according to the writing trajectory of the target part, extracting the area to be recognized from the image information; and recognizing the dictation result from the area to be recognized.
[0015] In yet another embodiment of the present invention, obtaining the writing trajectory of the target part according to the image information includes: obtaining the temporal position information of the target part in the image information; and determining the writing trajectory according to the temporal position information and the image information.
[0016] In yet another embodiment of the present invention, obtaining the timing position information of the target part in the image information includes: extracting an image of the target part from the image information; determining whether the target part is in a writing state according to the image of the target part; and obtaining the timing position information of the target part in the writing state in the image information.
[0017] In one embodiment of the present invention, the image information includes multiple frames of pictures. Extracting the image of the target part from the image information and determining whether the target part is in a writing state includes: determining whether the target part is in a writing state according to the image of the target part extracted from any one frame of the pictures; or extracting the images of the target part from consecutive multiple frames of pictures; forming the extracted images into video stream data; and determining whether the target part is in a writing state according to the video stream data.
[0018] In another embodiment of the present invention, it further includes: broadcasting the next audio task according to the correction result of the dictation result.
[0019] In yet another embodiment of the present invention, broadcasting the next audio task according to the correction result of the dictation result includes: determining whether the dictation result matches the reference information; in response to the dictation result matching the reference information, performing the operation of broadcasting the next audio task; or in response to the dictation result not matching the reference information, repeating the operations of recognizing and correcting the dictation result within the predetermined time, and when the current time is greater than the predetermined time, performing the operation of broadcasting the next audio task.
[0020] In the second aspect of the embodiment of the present invention, there is provided a device for real-time processing of dictation content, including: an audio broadcasting unit configured to broadcast one or more audio tasks in the dictation content; an image acquisition unit configured to, during the process of the audio broadcasting unit broadcasting one or more audio tasks in the dictation content, obtain in real time the dictation images corresponding to each audio task; and a processing unit connected to the audio broadcasting unit and the image acquisition unit and configured to: recognize the dictation result corresponding to the audio task from the dictation images; and correct the recognized dictation result.
[0021] In one embodiment of the present invention, the image acquisition unit is specifically configured to: within the predetermined time after the audio broadcasting unit finishes broadcasting each audio task, obtain in real time the image information of the content presented on the output medium through the input medium.
[0022] In another embodiment of the present invention, the processing unit includes: a trajectory acquisition unit configured to acquire a writing trajectory of a target part according to the image information, where the target part is the part where the input medium contacts the output medium; a region extraction unit configured to extract a region to be recognized from the image information according to the writing trajectory of the target part; and a content recognition unit for recognizing the dictation result from the region to be recognized.
[0023] In yet another embodiment of the present invention, the trajectory acquisition unit includes: a position acquisition unit configured to acquire the sequential position information of the target part in the image information; and a trajectory determination unit configured to determine the writing trajectory according to the sequential position information and the image information.
[0024] In yet another embodiment of the present invention, the position acquisition unit is specifically configured to: extract an image of the target part from the image information; determine whether the target part is in a writing state according to the image of the target part; and acquire the sequential position information of the target part in the writing state in the image information.
[0025] In an embodiment of the present invention, where the image information includes multiple frames of pictures, the position acquisition unit is specifically configured to: determine whether the target part is in a writing state according to the image of the target part extracted from any one frame of the pictures; or extract the images of the target part from consecutive multiple frames of pictures; form the extracted images into video stream data; and determine whether the target part is in a writing state according to the video stream data.
[0026] In another embodiment of the present invention, the processing unit is further configured to: trigger the audio broadcast unit to broadcast the next audio task according to the correction result of the dictation result.
[0027] In yet another embodiment of the present invention, the processing unit is specifically configured to: determine whether the dictation result matches the reference information; in response to the dictation result matching the reference information, trigger the audio broadcast unit to perform the operation of broadcasting the next audio task; or in response to the dictation result not matching the reference information, repeat the recognition and correction operations of the dictation result within the predetermined time, and when the current time is greater than the predetermined time, trigger the audio broadcast unit to perform the operation of broadcasting the next audio task.
[0028] In a third aspect of the embodiments of the present invention, a device is provided, including: a processor; and a memory storing computer instructions for real-time processing of dictation content, which, when run by the processor, cause the device to execute the methods described in multiple embodiments above and below.
[0029] In a fourth aspect of the embodiments of the present invention, a computer-readable storage medium is provided, containing program instructions for real-time processing of dictation content, which, when executed by a processor, cause the device to execute the methods described in multiple embodiments above and below.
[0030] The solution for real-time processing of dictation content and its related products according to the embodiments of the present invention can realize the recognition and correction of the dictation results of each audio task by obtaining in real time the dictation images corresponding to each audio task in the dictation content. It can be seen that the solution of the present invention can perform associated processing on the audio task and its dictation result through real-time image acquisition and recognition processing, thereby effectively improving the accuracy of recognition and correction. In some embodiments of the present invention, when collecting the dictation image, the image information presented on the output medium by the input medium can be collected in real time within a predetermined time after each audio task is broadcast, so as to accurately match the audio task and its dictation result based on the recognition of the time information and spatial information of the image, thereby effectively avoiding mis-matching and greatly improving the recognition and correction accuracy.
[0031] In some other embodiments of the present invention, the tracking technology of the writing trajectory can also be used to lock the area to be recognized in the image, which can exclude other interference information that may exist in the image to the greatest extent, thereby further improving the recognition accuracy of the dictation result. In addition, in some other embodiments of the present invention, by obtaining the temporal position information of the target part and judging its writing state, not only can the accurate tracking of the user's real writing trajectory be realized, but also the limitation of fixed writing tools (such as touch pads, etc.) can be eliminated to improve the user experience during dictation. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features and advantages of the exemplary embodiments of the present invention will become readily understood. In the drawings, several embodiments of the present invention are shown by way of illustration and not limitation, in which:
[0033] Figure 1 A block diagram of an exemplary computing system 100 suitable for implementing the embodiments of the present invention is schematically shown;
[0034] Figure 2Schematically shows a flowchart of a method for real-time processing of dictation content according to an embodiment of the present invention;
[0035] Figure 3 Schematically shows a flowchart of a method for recognizing dictation results from a dictation image according to an embodiment of the present invention;
[0036] Figure 4 Schematically shows a flowchart of a method for real-time processing of dictation content according to another embodiment of the present invention;
[0037] Figure 5 Schematically shows a method for recognizing dictation results from a dictation image according to another embodiment of the present invention;
[0038] Figure 6 Schematically shows a schematic diagram of a device for real-time processing of dictation content according to an embodiment of the present invention;
[0039] Figure 7 Schematically shows a schematic diagram of a device for real-time processing of dictation content according to another embodiment of the present invention; and
[0040] Figure 8 Schematically shows a schematic block diagram of a device according to an embodiment of the present invention.
[0041] In the drawings, the same or corresponding reference numerals denote the same or corresponding parts. Detailed implementation manners
[0042] The principles and spirit of the present invention will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are given only to enable those skilled in the art to better understand and then implement the present invention, and are not intended to limit the scope of the present invention in any way. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to be able to fully convey the scope of the present disclosure to those skilled in the art.
[0043] Figure 1 Shows a block diagram of an exemplary computing system 100 suitable for implementing the embodiments of the present invention. As Figure 1As shown, the computing system 100 may include: a central processing unit (CPU) 101, a random access memory (RAM) 102, a read-only memory (ROM) 103, a system bus 104, a hard disk controller 105, a keyboard controller 106, a serial interface controller 107, a parallel interface controller 108, a display controller 109, a hard disk 110, a keyboard 111, a serial external device 112, a parallel external device 113, and a display 114. Among these devices, those coupled to the system bus 104 are the CPU 101, the RAM 102, the ROM 103, the hard disk controller 105, the keyboard controller 106, the serial controller 107, the parallel controller 108, and the display controller 109. The hard disk 110 is coupled to the hard disk controller 105, the keyboard 111 is coupled to the keyboard controller 106, the serial external device 112 is coupled to the serial interface controller 107, the parallel external device 113 is coupled to the parallel interface controller 108, and the display 114 is coupled to the display controller 109. It should be understood that Figure 1 The structure block diagram described above is only for illustrative purposes and is not a limitation on the scope of the present invention. In some cases, certain devices may be added or reduced according to specific circumstances.
[0044] Those skilled in the art know that the embodiments of the present invention can be implemented as a system, a method, or a computer program product. Therefore, the present disclosure can be specifically implemented in the following forms: completely hardware, completely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software, generally referred to herein as "circuit", "module", "unit", or "system". In addition, in some embodiments, the present invention can also be implemented in the form of a computer program product in one or more computer-readable media, which contain computer-readable program codes.
[0045] Any combination of one or more computer-readable media can be adopted. The computer-readable media can be computer-readable signal media or computer-readable storage media. The computer-readable storage media can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive examples) of the computer-readable storage media can include, for example: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage media can be any tangible medium that contains or stores a program, which can be used by or in combination with an instruction execution system, apparatus, or device.
[0046] A computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0047] The program code contained on a computer-readable medium may be transmitted by any appropriate medium, including but not limited to wireless, wire, optical fiber cable, RF, etc., or any suitable combination of the above.
[0048] The computer program code for performing the operations of the present invention may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., connected through the Internet using an Internet service provider).
[0049] The embodiments of the present invention will be described below with reference to the flowchart of the method and the block diagram of the device (or system) of the embodiments of the present invention. It should be understood that each block of the flowchart and / or block diagram, and the combination of blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine. These computer program instructions are executed by the computer or other programmable data processing device, resulting in a device that implements the functions / operations specified in the blocks of the flowchart and / or block diagram.
[0050] These computer program instructions may also be stored in a computer-readable medium that can cause a computer or other programmable data processing device to work in a specific manner. In this way, the instructions stored in the computer-readable medium produce a product that includes an instruction device for implementing the functions / operations specified in the blocks of the flowchart and / or block diagram.
[0051] Computer program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other devices, so that a series of operation steps are performed on the computer, other programmable data processing apparatus, or other devices to generate a computer-implemented process, thereby enabling the instructions executed on the computer or other programmable apparatus to provide a process for implementing the functions / operations specified in the blocks of the flowchart and / or block diagram.
[0052] According to an embodiment of the present invention, a method for real-time processing of dictation content and related products thereof are provided. In addition, the number of any elements in the drawings is for illustration rather than limitation, and any naming is only for distinction and does not have any limiting meaning.
[0053] Next, the principles and spirit of the present invention will be explained in detail by referring to several representative embodiments of the present invention. Summary of the Invention
[0055] The inventors of the present invention have found that the existing dictation technologies are less user-friendly, and the recognition and correction accuracy are not ideal. For example, most of the existing dictation technologies uniformly obtain and recognize the content presented on the writing book after playing all the audio at once. Specifically, the user needs to use intelligent learning hardware to play all the new words in the current dictation vocabulary list at once, and trigger the camera to collect the dictation results after the dictation is completed, and then match and correct the dictation results with the new word list played by the intelligent learning hardware. This process not only requires the user to have relatively high dictation skills to avoid missing words, but also cannot trace back the correlation between the dictation results and the audio. For example, when there are "pencil" and "pencil case" in the audio task, the audio of "pencil" may be matched with the dictation result of "pencil case", or the audio of "pencil case" may be matched with the dictation result of "pencil", thus affecting the accuracy of recognition and correction.
[0056] Based on this, the inventors have found that if the accuracy of recognition and correction is to be ensured, the key problem lies in how to associate the played audio with the real-time dictation results. Specifically, the dictation images corresponding to each audio task can be obtained in real time, so as to associate them with the audio tasks based on the recognition of the dictation images, thereby improving the recognition and correction accuracy.
[0057] After introducing the basic principles of the present invention, various non-limiting embodiments of the present invention will be specifically introduced below.
[0058] Exemplary Method
[0059] Next, refer to Figure 2A method for real-time processing of dictation content according to an exemplary embodiment of the present invention will be described. It should be noted that the embodiments of the present invention can be applied to any applicable scenario.
[0060] Figure 2 A flowchart of a method 200 for real-time processing of dictation content according to an embodiment of the present invention is schematically shown. As Figure 2 shown, at step S201, during the broadcast of one or more audio tasks in the dictation content, a dictation image corresponding to each audio task can be obtained in real time. It should be noted that in the context of the present invention, the dictation content can include any audio-visual content that can be used for dictation training (such as audio-visual containing various language vocabularies). The dictation content can be preset by the system or customized by the user. And for the number of the aforementioned audio tasks, it can be specifically determined according to the dictation content and dictation requirements (such as difficulty level), etc.
[0061] Next, at step S202, the dictation result corresponding to the audio task can be recognized from the dictation image. In some embodiments, the recognition of the dictation image can be implemented by using a general text detection and algorithm recognition model (such as OCR technology). The dictation image here is collected in real time for each audio task to associate the dictation image with the audio task based on the time information of the dictation image, so as to realize the association between the audio task and the recognized dictation result. It should be noted that the description of the recognition process of the dictation image here is only an exemplary illustration, and the solution of the present invention is not limited thereto.
[0062] Next, at step S203, the recognized dictation result can be corrected. Specifically, in some embodiments, the dictation result can be matched with pre-stored reference information to implement the correction operation of the dictation result. Thus, through the dictation images corresponding to each audio task in the dictation content obtained in real time, the associated processing of the audio task and its dictation result is realized, thereby effectively improving the accuracy of recognition and correction.
[0063] The following further illustrates some possible exemplary implementation manners of Figure 2 each step.
[0064] In some embodiments, real-time acquisition of the aforementioned dictation image specifically involves acquiring, in real time, image information of the content presented on the output medium via the input medium within a predetermined time after each audio task is broadcast. It should be noted that the predetermined time can be adjusted according to the actual application scenario. The aforementioned output medium may include any medium capable of presenting handwriting (e.g., a paper dictation book, a desktop, a drawing board, an electronic touchpad, etc.), and the input medium may include any medium capable of cooperating with the output medium to present handwriting (e.g., an ordinary pen, an electronic stylus, a finger, etc.).
[0065] In addition, the inventors also found that when existing dictation technology uses a unified photo recognition of all content, if there is text in the dictation book before the dictation, it will also be recognized, and these background texts will be considered as the user's dictation content. In particular, if the background text happens to be in the list of new words, then this background text will cause a mismatch during the correction stage, resulting in correction errors. Based on this, the inventors also found that if the accuracy of recognition and correction is to be improved, the key to processing dictation content lies in how to identify the dictation results. For example, the dictation results can be accurately identified by tracking the user's writing trajectory.
[0066] Specifically, Figure 3 Detailed steps for obtaining dictation results in some embodiments are shown. Figure 3 As shown, at step S301, the writing trajectory of the target part can be obtained based on the aforementioned image information. Specifically, in some embodiments, the temporal position information of the target part in the image information can be obtained, and then the writing trajectory can be determined based on the aforementioned temporal position information and image information. It will be understood that the target part here is the part where the input medium contacts the output medium, such as a pen tip or fingertip. In addition, the acquisition of temporal position information will be described later.
[0067] Next, at step S302, the area to be identified can be extracted from the image information according to the writing trajectory of the aforementioned target part. For example, the image information includes multiple frames of pictures, and the corresponding area can be deducted from the last frame of the picture according to the writing trajectory as the area to be identified. Then, at step S303, the dictation result can be identified from the aforementioned area to be identified. For example, text detection technology can be used to identify the text box in the area to be identified, and then the content in the text box can be identified based on OCR technology as the dictation result. It should be noted that the description of the recognition process of the dictation result here is only an exemplary description, and the solution of the present invention is not limited to this.
[0068] Furthermore, in some embodiments, for the aforementioned temporal position information, specifically, an image of the target part (such as an image of the pen tip) can be extracted from the aforementioned image information, and then based on the image of the target part, it can be determined whether the target part is in a writing state, and the temporal position information of the target part in the writing state in the image information can be obtained. Specifically, in some actual application scenarios, the aforementioned image information may include multiple frames of pictures. An image of the target part can be extracted from any frame of the pictures (for example, any frame of the pictures can be input into a predetermined pen tip detection model to implement the image extraction operation), so as to determine whether the target part is in a writing state based on the extracted image, thereby achieving fast recognition of the writing state.
[0069] Alternatively, in some other embodiments, images of the target part can also be extracted from consecutive multiple frames of pictures to form video stream data, and based on this video stream data, it can be determined whether the target part is in a writing state (for example, the video stream data can be input into a predetermined pen tip classification model to implement the judgment of the writing state), thereby achieving accurate recognition of the writing state.
[0070] It should be noted that in the context of the present invention, the pen tip detection model and the pen tip classification model can be neural network models in computer vision technology (such as RCNN model, SVM model, etc.). Additionally, when the input medium is a fingertip or other medium, corresponding detection and classification models can be selected for processing. The solution of the present invention locks the area to be recognized in the image through the writing trajectory tracking technology, which can largely exclude other possible interference information in the image, thereby further improving the recognition accuracy of the dictation result.
[0071] In some other embodiments, after completing the correction operation of the dictation result corresponding to any audio task, the next audio task can be broadcast according to the correction result of the dictation result. Specifically, in some embodiments, it can be determined whether the dictation result matches the reference information, and when it is determined that the dictation result matches the reference information, the next audio task can be broadcast, which is beneficial to improving the dictation efficiency. When it is determined that the dictation result does not match the reference information, the recognition and correction operations of the dictation result can be repeatedly executed within a predetermined time to meet the correction requirements of the user's correction result, making the entire dictation process more in line with the actual needs. Then, when the current time is greater than the predetermined time, the operation of broadcasting the next audio task can be continued to achieve reasonable regulation of the entire dictation process.
[0072] Figure 4 Schematically shows a flowchart of a method 400 for real-time processing of dictation content according to another embodiment of the present invention. It can be understood that, Figure 4 can be combined with the foregoing Figure 2 and Figure 3An exemplary implementation of the described steps. Therefore, the detailed descriptions of the respective steps in the foregoing in conjunction with Figure 2 and Figure 3 are equally applicable to the following.
[0073] As Figure 4 shown, at step S401, audio broadcast can be performed. Specifically, after the intelligent learning hardware (such as an intelligent learning desk lamp, a tablet, etc.) enters the dictation mode, each time the intelligent learning hardware can broadcast the audio of a new word, allowing the user to write it in the dictation book. At the same time, a time t (such as 12 s) is set and the timing starts.
[0074] Next, at step S402, real-time dictation recognition can be performed. Specifically, starting from the timing, the camera in the intelligent learning hardware can collect the process of the user writing in real time and recognize the dictation result corresponding to the current new word in real time. There are various recognition methods for the dictation result. Figure 5 shows the specific steps of recognizing the dictation result in some embodiments.
[0075] As Figure 5 shown, at step S501, pen tip detection can be performed. For example, the i-th frame of the picture can be input into the pen tip detection model to obtain the rectangular frame of the position of the pen tip in the current frame. Next, at step S502, pen tip classification can be performed. For example, the picture of the pen tip can be cropped from the current frame picture according to the foregoing rectangular frame, and then input into the pen tip classification model to obtain whether the pen tip state is in the writing state. Alternatively, multiple frames of pictures before and after can also be collected to form a video stream data, and the video stream data can be input into the pen tip classification model to obtain whether the pen tip state is in the writing state.
[0076] If it is determined at step S502 that the pen tip is in the writing state, pen tip tracking can be performed at step S503. For example, the i-th frame of the picture and the pen tip position information can be added to a pen tip tracking trajectory in the writing state to form a writing trajectory. Then, at step S504, the writing area can be cropped from the picture according to the foregoing writing trajectory and input into the text detection model to obtain the text box in the writing area. In this process, the text detection model can output the position information of all texts in the writing area. Finally, at step S505, the picture within the foregoing text box can be cropped for picture pushing and pushed into OCR for recognition. Thus, the recognition of the dictation result is completed. <(
[0077] After completing the recognition of the dictation result corresponding to the current new word, continue Figure 4, at step S403, it is possible to check whether the foregoing dictation result is correct. Specifically, it is possible to determine whether the recognized dictation result matches (e.g., is consistent with) the reference new word. If it is determined to match, step S404 is executed, and the current dictation result can be recorded (e.g., record the current dictation result as "written correctly +1"), and after the recording is completed, return to continue the broadcast and dictation of the next new word. If it is determined not to match (e.g., is inconsistent), step S405 is executed.
[0078] At step S405, it is possible to determine whether it has timed out. For example, when it is determined that it has not timed out, return to execute step S402. When it is determined that the time t has been reached, the current dictation result can be recorded (e.g., record the current dictation result as "written wrongly +1"), and after the recording is completed, return to continue the broadcast and dictation of the next new word.
[0079] Through the solution of the present invention, based on the processing of the image information of the user's dictation process, the precise writing area corresponding to each new word can be obtained, and based on the recognition of the writing area and the correction of the recognition result, a real-time dictation function with high efficiency and good effect can be realized. Specifically, by dynamically recognizing each dictation result to obtain time and space information, the one-to-one correspondence between the broadcast audio and the writing content can be achieved, effectively avoiding mis-matching, thereby improving the correction efficiency and accuracy. In addition, each frame of the obtained picture can be combined with the pen tip position information to form a real writing pen tip movement trajectory (i.e., the writing trajectory). The whole process does not need to rely on hardware such as a touchpad, enabling the user to freely choose writing tools (such as a drawing board, a dictation book, a tablet, etc.). Furthermore, the means of picture classification or video classification can be used to determine whether the pen tip is in the writing state, providing a pre-judgment premise for performing pen tip trajectory tracking, so that the user's real and precise writing trajectory can be obtained more accurately.
[0080] Exemplary Device
[0081] After introducing the method of the exemplary embodiment of the present invention, next, refer to Figures 6 to 8 to describe the related products for real-time processing of dictation content of the exemplary embodiment of the present invention.
[0082] Figure 6 Schematically shows a schematic diagram of a device 600 for real-time processing of dictation content according to an embodiment of the present invention. As Figure 6As shown, the device 600 may include an audio broadcast unit 601, an image acquisition unit 602, and a processing unit 603. Among them, the audio broadcast unit 601 may be configured to broadcast one or more audio tasks in the dictation content. In practical applications, the audio broadcast unit may be a speaker or other audio-video playback APP. The aforementioned image acquisition unit 602 may be configured to, during the process of the audio broadcast unit broadcasting one or more audio tasks in the dictation content, acquire in real time the dictation images corresponding to each of the audio tasks. Here, the image acquisition unit 602 may be a camera. In practical applications, the image acquisition unit 602 may be integrated with other units in a device, or may be set separately (when set separately, it may communicate and interact with other units through wired or wireless communication methods).
[0083] The aforementioned processing unit 603 is connected to the audio broadcast unit 601 and the image acquisition unit 602, and is configured to identify the dictation results corresponding to the audio tasks from the dictation images, and correct the identified dictation results. In practical applications, the processing unit may be a CPU or CPU+GPU, etc., to support the processing operations of the dictation images. This device can support the real-time acquisition and recognition processing of images, and can associate the audio tasks and their dictation results, thereby effectively improving the accuracy of recognition and correction.
[0084] Figure 7 Schematically shown is a schematic diagram of a device 700 for real-time processing of dictation content according to another embodiment of the present invention. It should be noted that the device 700 can be understood as a further refinement and expansion of the functions of the Figure 6 device 600. Therefore, the relevant descriptions of the device in the foregoing in combination with Figure 6 also apply to the following text.
[0085] As Figure 7 shown, the device 700 may include an audio broadcast unit 701, an image acquisition unit 702, and a processing unit 703 (which may include a trajectory acquisition unit 703-1, a region extraction unit 703-2, and a content recognition unit 703-3). Among them, the trajectory acquisition unit 703-1 may further include a position acquisition unit and a trajectory determination unit.
[0086] Regarding the audio broadcast unit 701 and the image acquisition unit 702, they may have Figure 6The functions and configurations of the described audio broadcast unit and image acquisition unit. Further, the image acquisition unit 702 can specifically obtain, in real time within a predetermined time after the audio broadcast unit finishes broadcasting each audio task, the image information of the content presented on the output medium (such as a dictation book, a drawing board, an electronic touch screen, etc.) through the input medium (such as a pen, a finger, etc.). The descriptions of the input medium and the output medium here are only exemplary. For example, the input medium and the output medium can include other media that can cooperate with each other to present writing.
[0087] The foregoing trajectory acquisition unit 703-1 can be configured to obtain the writing trajectory of the target part (such as the pen tip or the fingertip, etc.) according to the image information. Specifically, in some embodiments, the position acquisition unit and the trajectory determination unit can be combined to obtain the writing trajectory by means of the sequential position information of the target part in the image information and the image information. Then, the region extraction unit 703-2 extracts the region to be recognized from the image information, and the content recognition unit 703-3 recognizes the dictation result from the region to be recognized. Specifically, reference can be made to the recognition process of the dictation result described in the foregoing in combination with Figure 5 which will not be elaborated here.
[0088] In addition, the processing unit 703 can also be configured to trigger the audio broadcast unit to broadcast the next audio task according to the correction result of the foregoing dictation result. Specifically, when it is determined that the dictation result is correct, the audio broadcast unit can be directly triggered to broadcast the next audio task. When it is determined that the dictation result is incorrect, the recognition and correction operations of the dictation result can be selectively repeated according to whether the current time exceeds the predetermined time. Thus, a reasonable management of the entire dictation process can be achieved to meet the actual requirements.
[0089] Figure 8 Schematically shows a schematic block diagram of a device 800 according to an embodiment of the present invention. As Figure 8 shown, the device 800 can include a processor 801 and a memory 802. The memory 802 stores computer instructions for real-time processing of dictation content. When the computer instructions are run by the processor 801, the device 800 is made to execute the method described in the foregoing in combination with Figures 2 to 4 which. For example, in some embodiments, the device 800 can execute the broadcast of audio tasks, the real-time acquisition of dictation images, the recognition and correction of dictation results, etc. Based on this, the accuracy of the recognition and correction of dictation content can be effectively improved through the device 800.
[0090] In some implementation scenarios, the device 800 may include an integrated device with audio and video broadcasting functions and image acquisition functions (such as a smart learning desk lamp or a tablet, etc.), or may also be an improved split device (such as a terminal with an audio broadcasting function + a terminal with a camera function). The solution of the present invention does not limit the structural design that the device 800 can have.
[0091] It should be noted that although several devices or sub-devices for real-time processing of dictation content are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present invention, the features and functions of the two or more devices described above can be embodied in one device. Conversely, the features and functions of one device described above can be further divided and embodied by multiple devices.
[0092] In addition, although the operations of the method of the present invention are described in a specific order in the drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. On the contrary, the steps depicted in the flowchart can be changed in the order of execution. Additionally or alternatively, some steps can be omitted, multiple steps can be combined into one step for execution, and / or one step can be decomposed into multiple steps for execution.
[0093] The use of the verbs "comprise", "include" and their inflectional forms mentioned in the application document does not exclude the existence of elements or steps other than those recited in the application document. The article "a" or "an" before an element does not exclude the existence of multiple such elements.
[0094] Although the spirit and principles of the present invention have been described with reference to several specific embodiments, it should be understood that the present invention is not limited to the specific embodiments disclosed, and the division of each aspect does not mean that the features in these aspects cannot be combined for benefit. This division is only for the convenience of expression. The present invention aims to cover various modifications and equivalent arrangements included within the spirit and scope of the appended claims. The scope of the appended claims is construed in the broadest sense so as to encompass all such modifications and equivalent structures and functions.
Claims
1. A method for real-time processing of dictation content, characterized in that: include: During the process of broadcasting one or more audio tasks in the dictation content, acquiring a dictation image corresponding to each of the audio tasks in real time; identifying a dictation result corresponding to the audio task from the dictation image, so as to associate the audio task with the dictation result corresponding to the audio task; as well as Correct the recognized dictation results; Acquiring the dictation image corresponding to each audio task in real time includes: within a predetermined time after each of the audio tasks is played, acquiring in real time image information of the content presented on the output medium via the input medium; The method further includes: announcing the next audio task according to the correction result of the dictation result; The step of broadcasting the next audio task according to the correction result of the dictation result includes: determining whether the dictation result matches the reference information; In response to the dictation result matching the reference information, performing an operation of broadcasting the next audio task; or In response to the dictation result not matching the reference information, repeatedly performing the recognition and correction operations on the dictation result within the predetermined time, and performing the operation of broadcasting the next audio task when the current time is greater than the predetermined time; The step of identifying a dictation result corresponding to the audio task from the dictation image includes: acquiring a writing trajectory of a target portion according to the image information, wherein the target portion is a portion where the input medium contacts the output medium; extracting a region to be identified from the image information according to the writing trajectory of the target part; and The dictation result is identified from the to-be-identified area.
2. The method according to claim 1, characterized in that Wherein obtaining the writing trajectory of the target part according to the image information includes: Acquiring temporal position information of the target part in the image information; and The writing trajectory is determined according to the temporal position information and the image information.
3. The method according to claim 2, characterized in that Acquiring the temporal position information of the target part in the image information includes: extracting an image of the target part from the image information; determining whether the target part is in a writing state according to the image of the target part; and The temporal position information of the target part in the writing state in the image information is obtained.
4. The method according to claim 3, characterized in that The image information includes multiple frames of pictures, and extracting the image of the target part from the image information and determining whether the target part is in a writing state includes: determining whether the target part is in a writing state according to an image of the target part extracted from any frame of the picture; or Extracting an image of the target part from multiple consecutive frames of images; Combining the extracted images into video stream data; and Determine whether the target part is in a writing state according to the video stream data.
5. A device for real-time processing of dictation content, characterized in that: include: an audio announcement unit configured to announce one or more audio tasks in the dictation content; An image acquisition unit is configured to acquire, in real time, a dictation image corresponding to each audio task during a process in which the audio broadcast unit broadcasts one or more audio tasks in the dictation content; as well as A processing unit connected to the audio broadcast unit and the image acquisition unit and configured to: identifying a dictation result corresponding to the audio task from the dictation image, so as to associate the audio task with the dictation result corresponding to the audio task; Correct the recognized dictation results; The image acquisition unit is specifically configured as follows: within a predetermined time after the audio broadcast unit finishes broadcasting each of the audio tasks, acquiring in real time image information of the content presented on the output medium via the input medium; The processing unit is further configured to: triggering the audio broadcast unit to broadcast the next audio task according to the correction result of the dictation result; The processing unit is specifically configured to: determining whether the dictation result matches the reference information; In response to the dictation result matching the reference information, triggering the audio broadcast unit to perform an operation of broadcasting the next audio task; or In response to the dictation result not matching the reference information, repeatedly performing the recognition and correction operations on the dictation result within the predetermined time, and triggering the audio broadcast unit to perform the operation of broadcasting the next audio task when the current time is greater than the predetermined time; The processing unit comprises: a trajectory acquisition unit configured to acquire a writing trajectory of a target portion according to the image information, wherein the target portion is a portion where the input medium contacts the output medium; an area extraction unit configured to extract an area to be identified from the image information according to the writing trajectory of the target part; and A content recognition unit is used to recognize the dictation result from the area to be recognized.
6. The device according to claim 5, characterized in that The trajectory acquisition unit includes: a position acquisition unit configured to acquire time-series position information of the target site in the image information; and A trajectory determination unit is configured to determine the writing trajectory according to the temporal position information and the image information.
7. The device according to claim 6, characterized in that The position acquisition unit is specifically configured to: extracting an image of the target part from the image information; determining, based on the image of the target part, whether the target part is in a writing state; as well as The temporal position information of the target part in the writing state in the image information is obtained.
8. The device according to claim 7, characterized in that The image information includes multiple frames of pictures, and the position acquisition unit is specifically configured to: determining whether the target part is in a writing state according to an image of the target part extracted from any frame of the picture; or Extracting an image of the target part from multiple consecutive frames of images; Combining the extracted images into video stream data; as well as Determine whether the target part is in a writing state according to the video stream data.
9. A device, characterized in that include: processor; as well as A memory storing computer instructions for processing dictation content in real time, wherein when the computer instructions are executed by the processor, the device executes the method according to any one of claims 1 to 4.
10. A computer-readable storage medium, characterized in that The device comprises program instructions for processing dictation content in real time, and when the program instructions are executed by a processor, the method according to any one of claims 1 to 4 is implemented.
Citation Information
Patent Citations
Dictation and reading method and electronic equipment
CN111081083A
Dictation answer acquisition method, family education equipment and storage medium
CN111081103A
Writing behavior recognition method and device based on artificial intelligence
CN112001236A