Assisted reading method, computer device and storage apparatus
By judging the stability of the pointer in the target frame image in the assisted reading method, the problem of the impact of finger position changes on the accuracy of assisted reading is solved, thus improving the accuracy of assisted reading.
Patent Information
- Application Number
- CN202011613152.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-30
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2040-12-30
AI Technical Summary
In existing assisted reading methods, changes in the user's finger position lead to a decrease in the accuracy of assisted reading.
By obtaining the similarity between the reference frame image and the target frame image, the stability of the indicator in the target frame image is determined, and it is decided whether to perform assisted reading.
It improves the accuracy of assisted reading and reduces erroneous assisted reading operations.
Smart Images

Figure CN112668491B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of assisted reading, and in particular to an assisted reading method, a computer device and a storage device. BACKGROUND
[0002] With the rapid development of artificial intelligence technology, people's daily life, learning and work are becoming more and more fast and convenient. For example, when a user is practicing or training reading, he or she may often encounter reading content that he or she does not understand or recognize. Assisted reading can effectively help the user correctly understand the reading content and learn more related knowledge in the reading content, so that the user can complete the reading process independently without relying on the assistance of others.
[0003] In the current assisted reading method, the user usually points to the area to be read with a finger, and then assisted reading is performed based on the indication position of the finger on the book to be read. However, the user's finger can change dynamically. If the position of the user's finger changes during the process of assisted reading with the finger, the accuracy of assisted reading will be affected. SUMMARY
[0004] The technical problem solved by the present application is to provide an assisted reading method, a computer device and a storage device, which can improve the accuracy of assisted reading.
[0005] To solve the above problem, the first aspect of the present application provides an assisted reading method, which comprises: acquiring an image after a reference frame image as a target frame image, wherein the reference frame image and the target frame image are images collected at different times on a reading area, and the reference frame image contains a preset indicator; determining the stability of the indicator of the target frame image based on the similarity of the reference frame image and the target frame image in a preset area; wherein the preset area in the reference frame image includes at least part of the preset indicator; and determining whether to perform assisted reading on the reading area based on the stability of the indicator of the target frame image.
[0006] To solve the above problem, the second aspect of the present application provides a computer device, which comprises a memory and a processor coupled to each other. The memory stores program data, and the processor is configured to execute the program data to implement any step of the above assisted reading method.
[0007] To solve the above problem, the third aspect of the present application provides a storage device, which stores program data capable of being executed by a processor. The program data is used to implement any step of the above assisted reading method.
[0008] The above scheme acquires an image located after a reference frame image as the target frame image. The reference frame image contains preset indicators. Based on the similarity between the reference frame image and the target frame image acquired at different times within a preset region of the area to be read, and since the preset region in the reference frame image includes at least part of the preset indicators, the stability of the indicators in the target frame image can be determined based on the similarity, thus improving the accuracy of judging the stability of the indicators in the target frame image. Furthermore, determining whether to perform assisted reading on the area to be read based on the stability of the indicators in the target frame image can reduce erroneous assisted reading operations, thereby improving the accuracy of assisted reading. Attached Figure Description
[0009] To more clearly illustrate the technical solutions in this application, the accompanying drawings required in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Among them:
[0010] Figure 1 This is a flowchart illustrating the first embodiment of the reading assistance method of this application;
[0011] Figure 2 This application Figure 1 A flowchart of an embodiment preceding step S11;
[0012] Figure 3 This is a schematic diagram of an embodiment of the indicator detection results of this application;
[0013] Figure 4 This application Figure 1 A flowchart illustrating an embodiment of step S12;
[0014] Figure 5 This application Figure 1 A flowchart illustrating an embodiment of step S13;
[0015] Figure 6 This is a flowchart illustrating the second embodiment of the auxiliary reading method of this application;
[0016] Figure 7 This is a schematic diagram of the structure of an embodiment of the computer device of this application;
[0017] Figure 8 This is a schematic diagram of the structure of an embodiment of the storage device of this application. Detailed Implementation
[0018] With reference to the drawings and the embodiments of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments in the present application, all the other embodiments obtained by those skilled in the art without creative work fall within the scope of the present application.
[0019] The terms "first", "second" in the present application are only for descriptive purpose, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined with "first", "second" can explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of "multiple" is at least two, for example, two, three, etc., unless otherwise specifically limited. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, the process, method, system, product or equipment including a series of steps or units is not limited to the listed steps or units, but can optionally include steps or units not listed, or can optionally include other steps or units inherent to the process, method, product or equipment.
[0020] In the present application, referring to "embodiments" means that the specific features, structures or properties described in combination with the embodiments can be included in at least one embodiment of the present application. The phrase appears at various places in the specification does not necessarily refer to the same embodiment, nor is it mutually exclusive or alternative to other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0021] Please refer to Figure 1 , Figure 1 is a flowchart of the first embodiment of the auxiliary reading method of the present application, which comprises the following steps:
[0022] S11: Acquire an image located after the reference frame image as a target frame image, wherein the reference frame image and the target frame image are images of the reading area collected at different time, and the reference frame image contains a preset indicator.
[0023] Real-time image acquisition of the reading area by a camera to acquire multiple frames of images collected in real time. The image containing the preset indicator in the multiple frames of images can be used as the reference frame image, and the image located after the reference frame image can be acquired as the target frame image. Wherein, the reference frame image and the target frame image are images of the reading area collected at different time, and the reference frame image and the target frame image are two consecutive images.
[0024] In addition, the preset indicator can be a pre-set indicator, for example, the preset indicator is a finger, a pen, a finger-like object or a pen-like object, and the position of the content to be read by the user is determined by the indication of the preset indicator. The present application does not limit the preset indicator.
[0025] S12: determining the stability of the indicator in the target frame image based on the similarity of the reference frame image and the target frame image in a preset area; wherein the preset area in the reference frame image includes at least part of the preset indicator.
[0026] In the preset area in the reference frame image, at least the preset indicator is included, and based on the similarity of the reference frame image and the target frame image in a preset area, the difference between the target frame image and the reference frame image can be determined. Since the content of the book to be read is unchanged, the indicator can be dynamically changed, so that the similarity of the reference frame image and the target frame image in a preset area can represent the difference of the preset area indicator in the two images. Therefore, based on the similarity of the reference frame image and the target frame image in a preset area, the stability of the indicator in the target frame image can be determined to determine whether the pointing reading operation is performed.
[0027] Optionally, the preset area can be the same preset area, so that the stability of the indicator in the target frame image is determined based on the similarity of the reference frame image and the target frame image in the same preset area.
[0028] S13: determining whether to assist reading in the reading area based on the stability of the indicator in the target frame image.
[0029] Based on the stability of the indicator in the target frame image, it is determined whether the pointing reading operation is performed to determine whether to assist reading in the reading area. If the indicator in the target frame image is in an unstable state, it can be determined that the user does not perform the pointing reading operation, and the reading area can not be assisted to read, thereby reducing the operation of incorrect auxiliary reading. If the indicator in the target frame image is in a stable state, it can be determined that the user does not perform the pointing reading operation, and the reading area can be assisted to read based on the position of the indicator.
[0030] In the embodiment, the image located after the reference frame image is obtained as the target frame image, the reference frame image contains the preset indicator, and the similarity of the reference frame image and the target frame image in a preset area is determined based on the reference frame image and the target frame image collected at different time points in the reading area. Since the preset area in the reference frame image includes at least part of the preset indicator, the stability of the indicator in the target frame image can be determined according to the similarity, which can improve the accuracy of the judgment of the stability of the indicator in the target frame image. In addition, based on the stability of the indicator in the target frame image, it is determined whether to assist reading in the reading area, which can reduce the operation of incorrect auxiliary reading, thereby improving the accuracy of the auxiliary reading.
[0031] In some embodiments, referring to Figure 2 Before step S11 in the above embodiments, the following steps can also be included:
[0032] S21: performing indicator detection on at least one frame of image collected for the reading area to obtain an indicator detection result for each frame of image, wherein the indicator detection result includes whether the image contains a preset indicator and detection information of the preset indicator contained.
[0033] Optionally, before step S21, at least one of the following preprocessing steps can also be included:
[0034] Perspective transformation is performed on the image. Specifically, when the image of the reading area is collected by the camera, the camera has a perspective effect on the image collected for the reading area due to the fact that the camera and the desktop of the reading area usually form a certain angle. Perspective transformation is performed on the image to obtain a perspective-transformed image.
[0035] The image is scaled to a preset size. The perspective-transformed image can be scaled to the preset size at a constant ratio to obtain a scaled image. The preset size is obtained through experiments on the accuracy of indicator detection, and the size of the preset size is not limited in the present application. In addition, if the image is not subjected to perspective transformation or the camera and the desktop of the reading area do not form a certain angle, the image collected for the reading area can also be scaled to the preset size to obtain a scaled image.
[0036] At least one frame of image collected for the reading area is subjected to indicator detection to obtain an indicator detection result for each frame of image. Specifically, after the image is scaled to the preset size, the image is subjected to feature extraction by using an indicator detection model and is subjected to discrimination based on the extracted features to obtain the indicator detection result for the image, which can reduce the computational amount of the indicator detection model. The indicator detection model is a model trained based on the features of sample indicators, and the features of the indicators can be contour features, area features, morphological features, etc.
[0037] The indicator detection result includes whether the image contains a preset indicator and detection information of the preset indicator contained. The detection information of the preset indicator includes at least one of position information, direction information, width, and proportion of the preset indicator in the image. Referring to Figure 3Taking a hand as an example of the preset indicator A, a frame of image C collected from the reading area B is used for detecting the indicator A. The position information p of the preset indicator A in the image includes at least one of the horizontal coordinate p_x and the vertical coordinate p_y of the pointing position of the indicator. The direction information includes at least one of the direction vector d of the indicator, the projection d_x of the direction vector of the indicator on the horizontal coordinate X axis, the projection d_y of the direction vector of the indicator on the vertical coordinate Y axis, and the angle a between the direction vector d of the indicator and the positive direction of the horizontal coordinate X axis. The width t of the indicator A is the width of the width direction of the indicator A. For example, when the indicator is a hand, the width t of the indicator is the width of the pointing hand. When the indicator is a pen, the width t of the indicator A is the width of the pen. The ratio r of the indicator can be the area ratio of the indicator A in the image. The ratio r of the indicator A can also be the ratio of the number of pixels occupied by the indicator A to the total number of pixels in the image.
[0038] In step S21, if the indicator detection result is that the preset indicator exists, step S22 is performed; otherwise, the above step S21 is continuously performed.
[0039] S22: determining whether the image has an abnormality based on the detection information of the preset indicator.
[0040] Optionally, if the ratio of the preset indicator in the image is greater than a first preset ratio or less than a second preset ratio, it is determined that the image has an abnormality. For example, in the process of auxiliary reading by a user, the camera continuously collects images of the reading area, and detects the indicator. In the dynamic process of the pointing action of the preset indicator, if the preset indicator is too close to the camera, so that the ratio of the preset indicator in the image is greater than the first preset ratio, it is determined that the image has an abnormality. In addition, if part of the preset indicator exists at the edge position of the detection image, so that the ratio of the preset indicator in the image is less than the second preset ratio, the specified action of the indicator can not be clear, and it is determined that the image has an abnormality.
[0041] Optionally, if the width of the preset indicator in the image is greater than a preset width ratio, it is determined that the image has an abnormality. For example, taking a hand as an example of the preset indicator, if the user uses the palm or the fist to read by pointing, the width of the hand in the image is greater than the preset width ratio, the pointing range is wider, and it is difficult to accurately determine the pointing position of the hand. At this time, the direction vector of the indicator is unreliable, and it is determined that the image has an abnormality.
[0042] S23: filtering the image with an abnormality.
[0043] The detection information of the preset indicator is used to determine whether the image has an abnormality, so as to filter the abnormal image, so that the auxiliary reading is not performed on the abnormal image, unnecessary subsequent processing of the image is reduced, and the response speed of the auxiliary reading is accelerated.
[0044] In some embodiments, referring to Figure 4 For step S12 in the above embodiments, the following steps can also be included:
[0045] S121: Obtain the position of the preset indicator in the reference frame image as the latest position of the preset indicator.
[0046] Wherein, the position of the preset indicator is the coordinate position p of the preset indicator in the image, the position of the preset indicator in the reference frame image is obtained as the latest position of the preset indicator, and the latest position is a global variable.
[0047] S122: Cut a first preset area containing the latest position from the reference frame image, and cut a second preset area containing the latest position from the target frame image.
[0048] The first preset area containing the latest position is cut from the reference frame image with the position of the preset indicator as the center, and the first preset area can be a rectangular image block with a width of w and a height of h. Wherein, the auxiliary reading scene can be pre-set, such as reading words, reading sentences, reading paragraphs, reading articles, etc., and each set of auxiliary reading scene corresponds to a first preset area.
[0049] The image located in the next frame of the reference frame image as the target frame image, the preset indicator detection can not be performed on the target frame image, and the second preset area containing the latest position is cut from the target image frame according to the latest position of the preset indicator in the reference frame image, and the cutting method of the second preset area can refer to the first preset area.
[0050] Optionally, the second preset area is the same as the first preset area.
[0051] S123: Obtain the similarity between the first preset area and the second preset area.
[0052] The first hash sequence of the first preset area and the second hash sequence of the second preset area are obtained by using the difference hash algorithm. The Hamming distance between the first hash sequence and the second hash sequence is obtained as the similarity. Wherein, the first hash sequence is a global variable, and the second hash sequence is a local variable. The smaller the Hamming distance, the more similar the first preset area and the second preset area can be considered, and the larger the Hamming distance, the less similar the first preset area and the second preset area can be considered.
[0053] S124: Determine the indicator stability of the target frame image based on the similarity.
[0054] Wherein, the indicator stability includes an unstable state of the indicator and a stable state of the indicator.
[0055] If the similarity is not less than the first preset threshold, for example, the Hamming distance is not less than the first preset threshold, it is determined that the first preset region and the second preset region are not similar, which indicates that the position of the preset indicator in the target frame image has changed, and it is determined that the target frame image is in the unstable state of the indicator.
[0056] If the similarity is less than the first preset threshold, for example, the Hamming distance is less than the first preset threshold, it is determined that the first preset region and the second preset region are similar, which indicates that the position of the preset indicator in the target frame image is relatively close, and it is determined that the target frame image is in the stable state of the indicator.
[0057] In the embodiment, by intercepting the first preset region containing the latest position from the reference frame image and the second preset region containing the latest position from the target frame image in the case of slow movement of the preset indicator, and determining the stable state of the indicator of the target frame image based on the similarity, the misjudgment of the stable state of the preset indicator can be reduced to reduce unnecessary auxiliary reading.
[0058] In some embodiments, referring to Figure 5 For step S13 in the above embodiment, based on the stable state of the indicator of the target frame image, it is determined whether to perform auxiliary reading on the to-be-read region, and the following steps can also be included:
[0059] S131: Determine whether the indicator of the target frame image is in a stable state.
[0060] If it is determined that the target frame image is in an unstable state of the indicator, step S132 is performed.
[0061] If it is determined that the target frame image is in a stable state of the indicator, step S133 is performed.
[0062] S132: Do not perform auxiliary reading on the to-be-read region.
[0063] When it is determined that the target frame image is in an unstable state of the indicator, it indicates that the pointing operation of the preset indicator is not completed at this time; and not performing auxiliary reading on the to-be-read region can reduce unnecessary auxiliary reading and reduce the amount of calculation of auxiliary reading.
[0064] When it is determined that the target frame image is in a stable state of the indicator, it indicates that the preset indicator has not left the existing pointing position; and not performing auxiliary reading on the to-be-read region can reduce unnecessary auxiliary reading and reduce the amount of calculation of auxiliary reading.
[0065] S133: Determine whether the target frame image is in an initial stable state of the indicator or a continuous stable state of the indicator.
[0066] The indicator stable state includes an indicator initial stable state and an indicator continuous stable state. When it is determined that the target frame image is in the indicator stable state, it can be further determined whether the target frame image is in the indicator initial stable state or the indicator continuous stable state.
[0067] Specifically, it is determined whether the front preset number of frame images of the target frame image is in the indicator stable state, for example, whether the front frame image of the target frame image is in the indicator stable state. If yes, it is determined that the target frame image is in the indicator continuous stable state. If no, it is determined that the target frame image is in the indicator initial stable state.
[0068] In step S133, if it is determined that the target frame image is in the indicator initial stable state, step S132 is performed.
[0069] If it is determined that the target frame image is in the indicator continuous stable state, step S135 is performed.
[0070] S134: performing auxiliary reading on the to-be-read region.
[0071] When it is determined that the target frame image is in the indicator initial stable state, auxiliary reading is performed on the to-be-read region. It can be detected whether there is a pre-stored audio resource associated with the position of the preset indicator in the target frame image. If yes, auxiliary reading is performed based on the pre-stored audio resource. If no, text recognition is performed on a third preset region containing the preset indicator in the target frame image, and auxiliary reading is performed based on the text recognition result.
[0072] In the embodiment, when it is determined that the target frame image is in the indicator unstable state, auxiliary reading is not performed on the to-be-read region, which can reduce unnecessary auxiliary reading and reduce the computational amount of auxiliary reading. In addition, when it is determined that the target frame image is in the indicator unstable state, it is further determined whether the target frame image is in the indicator initial stable state or the indicator continuous stable state. When the target frame image is in the indicator continuous stable state, auxiliary reading is not performed on the to-be-read region, which can further reduce unnecessary auxiliary reading and reduce the computational amount of auxiliary reading.
[0073] In some embodiments, the auxiliary reading performed on the to-be-read region in the above embodiments can include detecting whether there is a pre-stored audio resource associated with the position of the preset indicator in the target frame image. Specifically, a reading resource database supporting auxiliary reading can be established in the system, and books supporting auxiliary reading are pre-stored, so that audio resources corresponding to the books are pre-stored. By detecting whether there is a pre-stored audio resource associated with the position of the preset indicator in the target frame image, it can be determined whether the preset indicator in the target frame image is a book supporting auxiliary reading. If yes, auxiliary reading is performed based on the pre-stored audio resource associated with the position of the preset indicator in the target frame image.
[0074] If not, text recognition is performed on a third preset area containing the preset indicator in the target frame image, and auxiliary reading is performed based on the text recognition result. The third preset area can be a third preset area of a certain size obtained by centering on the position of the indicator. Auxiliary reading scenarios such as reading words, reading sentences, reading paragraphs, and reading articles can be preset, and each set of auxiliary reading scenarios corresponds to a third preset area. The text recognition is performed by recognizing the to-be-read content pointed to by the preset indicator in the third preset area through an OCR (Optical Character Recognition) technology, and the corresponding audio resource is searched for in the Internet or an auxiliary reading system based on the text recognition result, and the auxiliary reading is performed based on the audio resource.
[0075] In this embodiment, it is detected whether there is a pre-stored audio resource associated with the position of the preset indicator in the target frame image. When there is no pre-stored audio resource, text recognition is performed on a third preset area containing the preset indicator in the target frame image, and auxiliary reading is performed based on the text recognition result. The to-be-read book that is not supported by the auxiliary reading can be read with the aid of the auxiliary reading, so that the demand for auxiliary reading is better met. In addition, for the to-be-read content that is not supported by the auxiliary reading, the third preset area containing the preset indicator is subjected to text recognition, which is equivalent to text recognition of the entire content of the target frame image, so that the amount of calculation for word recognition can be reduced and the efficiency of auxiliary reading can be improved.
[0076] Please refer to Figure 6 , Figure 6 is a flowchart of a second embodiment of the auxiliary reading method of the present application, which includes the following steps:
[0077] S31: Acquire an image located after a reference frame image as a target frame image, wherein the reference frame image and the target frame image are images of a to-be-read region acquired at different moments, and the reference frame image contains a preset indicator.
[0078] S32: Determine the indicator stability of the target frame image based on the similarity of the reference frame image and the target frame image in a preset area, wherein the preset area in the reference frame image includes at least part of the preset indicator.
[0079] S33: Determine whether to perform auxiliary reading on the to-be-read region based on the indicator stability of the target frame image.
[0080] The specific implementation process of steps S31 to S33 in this embodiment can refer to the implementation process of steps S11 to S13 in the above embodiments, which will not be described here again.
[0081] If the target frame image is in the indicator unstable state, step S33 is performed.
[0082] If the target frame image is in the indicator stable state, step S35 to step 37 are performed.
[0083] S34: The target frame image or the nth frame image after the target frame image is taken as a new reference frame image, and the subsequent steps of taking the image after the reference frame image as the target frame image are re-executed, wherein n is a positive integer.
[0084] When the target frame image is in the indicator unstable state, it indicates that the position information of the preset indicator has changed, and the latest position of the preset indicator needs to be updated. At this time, the target frame image or the nth frame image after the target frame image can be taken as a new reference frame image, and the subsequent steps of taking the image after the reference frame image as the target frame image are re-executed, wherein n is a positive integer.
[0085] Specifically, the target frame image is subjected to indicator detection to obtain an indicator detection result of the target frame image, wherein the indicator detection result includes whether the image contains the preset indicator and detection information of the preset indicator contained.
[0086] If the target frame image contains the preset indicator, the target frame image is taken as a new reference frame image, the position of the preset indicator of the new reference frame image is updated to the latest position, the next frame image of the new reference frame image is taken as a new target frame image, and the subsequent steps of taking the image after the reference frame image as the target frame image are re-executed.
[0087] If the target frame image does not contain the preset indicator, the next frame image of the target frame image is subjected to preset indicator detection, the nth frame image after the target frame image containing the preset indicator is taken as a new reference frame image, the latest position of the preset indicator is updated, the image after the new reference frame image is taken as a new target frame image, and the subsequent steps of taking the image after the reference frame image as the target frame image are re-executed.
[0088] Optionally, the detection information of the target frame image containing the preset indicator determines that the image does not exist abnormally.
[0089] S35: The number of continuous frames of the target frame image in the indicator stable state compared with the same reference frame image is counted.
[0090] If the target frame image is in the persistent stable state of the indicator, the number of continuous frames of the target frame image compared with the same reference frame image and in the persistent stable state of the indicator is counted. For example, when the first frame image after the reference frame image is the target frame image, the target frame image is in the stable state of the indicator, when the second frame image after the reference frame image is the target frame image, the target frame image is also in the indicator, and the persistent stable state of the indicator is determined, and the number of continuous frames of the target frame image compared with the same reference frame image and in the persistent stable state of the indicator is recorded as 2.
[0091] In step S36, it is determined whether the number of continuous frames is greater than the second preset threshold.
[0092] In step S36, if the number of continuous frames is greater than the second preset threshold, step S34 is performed, the target frame image or the nth frame image after the target frame image is taken as a new reference frame image, and the subsequent steps of acquiring the image after the reference frame image as the target frame image and the like are re-executed. The target frame image can be the target frame image when the number of continuous frames is greater than the second preset threshold.
[0093] In step S36, if the number of continuous frames is not greater than the second preset threshold, step S37 is performed.
[0094] In step S37, the nth frame image after the target frame image is taken as a new target frame image, and the subsequent steps of determining the stable state of the indicator of the target frame image based on the similarity of the reference frame image and the target frame image in a preset region and the like are re-executed.
[0095] When the number of continuous frames is not greater than the second preset threshold, the image after the target frame image can be continuously processed, the nth frame image after the target frame image is taken as a new target frame image, for example, the next frame image of the target frame image is taken as a new target frame image, and the subsequent steps of determining the stable state of the indicator of the target frame image based on the similarity of the reference frame image and the target frame image in a preset region and the like are re-executed.
[0096] In the embodiment, when the target frame image is in the persistent stable state of the indicator, the number of continuous frames of the target frame image compared with the same reference frame image and in the persistent stable state of the indicator is counted, if the number of continuous frames is greater than the second preset threshold, the target frame image or the nth frame image after the target frame image is taken as a new reference frame image, and the subsequent steps of acquiring the image after the reference frame image as the target frame image and the like are re-executed. The accuracy of the auxiliary reading can be improved by reducing the misjudgment of the preset indicator of the target frame image and the preset stability of the indicator.
[0097] For the above-mentioned embodiments, the present application provides a computer device, please refer to Figure 7 , Figure 7is a structural schematic diagram of an embodiment of a computer device of the present application. The computer device 100 comprises a memory 101 and a processor 102, wherein the memory 101 and the processor 102 are coupled to each other, the memory 101 stores program data, and the processor 102 is configured to execute the program data to implement the steps of any of the above-mentioned auxiliary reading method embodiments.
[0098] In this embodiment, the processor 102 can also be referred to as a CPU (Central Processing Unit). The processor 102 can be an integrated circuit chip with processing capability. The processor 102 can also be a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), a FPGA (Field Programmable Gate Array) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor can be a microprocessor or the processor 102 can also be any conventional processor.
[0099] The specific implementation of this embodiment can refer to the implementation process of the above-mentioned embodiments, which will not be repeated here.
[0100] For the method of the above-mentioned embodiments, it can be implemented in the form of a computer program, and thus the present application proposes a storage device. Please refer to Figure 8 , Figure 8 is a structural schematic diagram of an embodiment of a storage device of the present application. The storage device 200 stores program data 201 capable of being executed by a processor, and the program data can be executed by the processor to implement the steps of any of the above-mentioned auxiliary reading method embodiments.
[0101] The specific implementation of this embodiment can refer to the implementation process of the above-mentioned embodiments, which will not be repeated here.
[0102] The storage device 200 of this embodiment can be a U disk, a mobile hard disk, a ROM (Read-Only Memory), a RAM (Random Access Memory), a magnetic disk or an optical disk, etc. which can store program data, or it can also be a server storing the program data. The server can send the stored program data to other devices for execution, or it can also execute the stored program data by itself.
[0103] In several embodiments provided in the present application, it should be understood that the disclosed methods and devices can be implemented in other manners. For example, the division of the apparatus embodiments described above is merely an example, and the division of the modules or units can be different, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electrical, mechanical or other forms.
[0104] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they can be located in one place or distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the present embodiment.
[0105] In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0106] If the integrated unit is realized in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage device, which is a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part of the prior art that contributes to the technical solutions or all or part of the technical solutions can be embodied in the form of a software product, which is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to perform all or part of the steps of the methods of the various embodiments of the present application.
[0107] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present application can be realized by general computing devices, which can be concentrated on a single computing device or distributed on a network composed of a plurality of computing devices, and optionally, they can be realized by program codes executable by computing devices, so that they can be stored in storage devices and executed by computing devices, or they can be made into individual integrated circuit modules, or a plurality of modules or steps can be made into a single integrated circuit module. Thus, the present application is not limited to any specific combination of hardware and software.
Claims
1. An auxiliary reading method, characterized by, The method comprises: indicating acquisition of at least one frame of image of a to-be-read region to obtain an indicator detection result of each frame of image, wherein the indicator detection result comprises whether the image contains a preset indicator and detection information of the preset indicator contained; if the indicator detection result is that the preset indicator is contained, determining whether the image is abnormal based on the detection information of the preset indicator; wherein if a proportion of the preset indicator in the image is greater than a first preset proportion or less than a second preset proportion, it is determined that the image is abnormal; if a width of the preset indicator in the image is greater than a preset width proportion, it is determined that the image is abnormal; filtering the image with abnormality; acquiring an image after a reference frame image as a target frame image, wherein the reference frame image and the target frame image are images of the to-be-read region acquired at different time points, and the reference frame image contains a preset indicator; determining an indicator stability condition of the target frame image based on a similarity of the reference frame image and the target frame image in a preset region; wherein the preset region in the reference frame image comprises at least part of the preset indicator, and the preset region contains a latest position of the preset indicator in the reference frame image, comprising: acquiring a position of the preset indicator in the reference frame image as the latest position of the preset indicator; cutting a first preset region containing the latest position from the reference frame image, and cutting a second preset region containing the latest position from the target frame image; acquiring a similarity between the first preset region and the second preset region; and determining the indicator stability condition of the target frame image based on the similarity; determining whether to perform auxiliary reading on the to-be-read region based on the indicator stability condition of the target frame image; wherein auxiliary reading is performed on the to-be-read region only in the case that the indicator stability condition of the target frame image is in an initial stable state of the indicator; wherein the indicator stability condition comprises an unstable state of the indicator and a stable state of the indicator, the stable state of the indicator comprises an initial stable state of the indicator and a sustained stable state of the indicator; and determining whether to perform auxiliary reading on the to-be-read region based on the indicator stability condition of the target frame image comprises: if the target frame image is in the unstable state of the indicator, auxiliary reading is not performed on the to-be-read region; if the target frame image is in the stable state of the indicator, it is determined whether the target frame image is in the initial stable state of the indicator or in the sustained stable state of the indicator; if the target frame image is in the initial stable state of the indicator, auxiliary reading is performed on the to-be-read region; if the target frame image is in the sustained stable state of the indicator, auxiliary reading is not performed on the to-be-read region.
2. The method of claim 1, wherein, The acquisition of the similarity between the first preset region and the second preset region comprises: acquiring a first hash sequence of the first preset region and a second hash sequence of the second preset region by using a difference hash algorithm respectively; obtaining a Hamming distance between the first hash sequence and the second hash sequence as the similarity.
3. The method of claim 1, wherein, The method further comprises: if the similarity is not less than a first preset threshold, determining that the target frame image is in an indicator unstable state; if the similarity is less than the first preset threshold, determining that the target frame image is in an indicator stable state.
4. The method of claim 1, wherein, The method further comprises: if the target frame image is in the indicator unstable state, taking the target frame image or an nth frame image after the target frame image as a new reference frame image, and re-executing the obtaining of the image after the reference frame image as a target frame image and the subsequent steps; wherein the n is a positive integer; and / or, if the target frame image is in the indicator stable state, counting a continuous frame number of the target frame image which is compared with the same reference frame image and is in the indicator stable state; if the continuous frame number is greater than a second preset threshold, taking the target frame image or an nth frame image after the target frame image as a new reference frame image, and re-executing the obtaining of the image after the reference frame image as a target frame image and the subsequent steps; if the continuous frame number is not greater than the second preset threshold, taking an nth frame image after the target frame image as a new target frame image, and re-executing the determining of the indicator stable state of the target frame image based on the similarity between the reference frame image and the target frame image in a preset region and the subsequent steps. The method further comprises:
5. The method of claim 1, wherein, if the target frame image contains the preset indicator, taking the target frame image as a new reference frame image; if the target frame image does not contain the preset indicator, taking the nth frame image after the target frame image which contains the preset indicator as a new reference frame image. The detection information of the preset indicator comprises at least one of position information, direction information, width and proportion of the preset indicator in the image.
6. The method of claim 5, wherein, Before the step of performing indicator detection on the at least one frame image collected from the reading area to obtain the indicator detection result of each frame image, the method further comprises at least one of the following preprocessing steps: perspective transformation is performed on the image; scaling the image to a preset size; 7. The method of claim 1, wherein, The method further comprises:
8. The method of claim 1, wherein, if the target frame image contains the preset indicator, taking the target frame image as a new reference frame image; if the target frame image does not contain the preset indicator, taking the nth frame image after the target frame image which contains the preset indicator as a new reference frame image. The image is subjected to feature extraction by a pointer detection model and discrimination based on the extracted features, to obtain a pointer detection result of the image.
9. The method of claim 1, wherein, The assisting reading of the to-be-read region comprises: detecting whether there is a pre-stored audio resource associated with the position of the preset pointer in the target frame image; if there is, assisting reading based on the pre-stored audio resource; and / or, if there is not, performing text recognition on a third preset region containing the preset pointer in the target frame image, and assisting reading based on the text recognition result.
10. A computer device, comprising: The memory and the processor are coupled with each other, the memory stores program data, and the processor is configured to execute the program data to implement the steps of the method in any one of claims 1 to 9.
11. A memory device, comprising: The memory stores program data capable of being executed by the processor, and the program data is used to implement the steps of the method in any one of claims 1 to 9.
Citation Information
Patent Citations
Method and device for detecting stability of image frame in video stream
CN110049309A
Auxiliary reading method and device, electronic equipment and storage medium
CN111539405A