Interactive reading method, device and system and computer readable storage medium

CN120239844APending Publication Date: 2025-07-01BOE TECHNOLOGY GROUP CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380011530.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-10-31
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

When existing picture book robots read through fingertips, the number of words per single time is limited, which makes it time-consuming and labor-consuming when reading large paragraphs of text.

Method used

The image acquisition device collects images of the user's fingers and book content, recognizes the finger pointing position and movement trajectory, judges it as a circle reading or point reading gesture, determines the corresponding book content and reports it.

Benefits of technology

It supports quick selection of large paragraphs of text, solves the problem of time-consuming and labor-intensive reading of single-word dots, and has high interactions and no special gesture activation is required.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120239844A_ABST
    Figure CN120239844A_ABST
Patent Text Reader

Abstract

The invention discloses an interactive reading method, device and system and a computer readable storage medium, and the method comprises the steps: obtaining an image collected by an image collection device, and the image collected by the image collection device comprises the finger and book content of a user; according to the image acquired by the image acquisition device, identifying a finger point position and a finger point moving track of the user; judging whether the action of the user is a circle reading gesture or a point reading gesture according to the finger point position and the finger point moving track of the user; and determining book content corresponding to the circle reading gesture or the point reading gesture, and broadcasting the book content corresponding to the circle reading gesture or the point reading gesture.
Need to check novelty before this filing date? Find Prior Art

Description

Interactive reading method, device, system and computer-readable storage medium Technical Field

[0001] The embodiments of the present disclosure relate to, but are not limited to, the field of target recognition technology, and in particular to an interactive reading method, device, system, and computer-readable storage medium. Background Art

[0002] Most picture book robots currently on the market interact with users through fingertip reading, which limits the number of words that can be selected at a time. This is time-consuming and labor-intensive when users need to read large sections of text.

[0003] Summary of the Invention

[0004] The following is a summary of the subject matter described in detail herein. This summary is not intended to limit the scope of the claims.

[0005] The embodiment of the present disclosure provides an interactive reading system, comprising: an image acquisition device and an interactive reading device, wherein the interactive reading device is configured to:

[0006] Acquiring an image captured by an image capturing device, where the image captured by the image capturing device includes a user's finger and book content;

[0007] Identifying the user's finger position and finger movement trajectory based on the image captured by the image capture device;

[0008] Determine whether the user's action is a circle reading gesture or a point reading gesture based on the user's finger position and finger movement trajectory;

[0009] The book content corresponding to the circle reading gesture or the point reading gesture is determined, and the book content corresponding to the circle reading gesture or the point reading gesture is announced.

[0010] The present disclosure also provides an interactive reading method, including:

[0011] Acquiring an image captured by an image capturing device, where the image captured by the image capturing device includes a user's finger and book content;

[0012] Identifying the user's finger position and finger movement trajectory based on the image captured by the image capture device;

[0013] Determine whether the user's action is a circle reading gesture or a point reading gesture based on the user's finger position and finger movement trajectory;

[0014] Determine the book content corresponding to the circle reading gesture or the point reading gesture, and read out the book content corresponding to the circle reading gesture or the point reading gesture.

[0015] An embodiment of the present disclosure also provides an interactive reading device, comprising a memory; and a processor connected to the memory, wherein the memory is used to store instructions, and the processor is configured to execute the steps of the interactive reading method described in any embodiment of the present disclosure based on the instructions stored in the memory.

[0016] An embodiment of the present disclosure further provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the interactive reading method described in any embodiment of the present disclosure is implemented.

[0017] Other aspects will become apparent upon reading and understanding the drawings and detailed description.

[0018] Summary of the Figures

[0019] The accompanying drawings are intended to provide a further understanding of the technical solutions of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the technical solutions of the present disclosure and do not constitute a limitation of the technical solutions of the present disclosure. The shapes and sizes of the components in the drawings do not reflect the actual scale and are intended only to illustrate the contents of the present disclosure.

[0020] FIG1 is a flow chart of an interactive reading method provided by an exemplary embodiment of the present disclosure;

[0021] FIG2 is a schematic diagram of an application scenario of an interactive reading method provided by an exemplary embodiment of the present disclosure;

[0022] FIG3 is a schematic diagram of an image captured by an image capture device provided by an exemplary embodiment of the present disclosure;

[0023] FIG4A is a schematic diagram of the positions of convex points, concave points, and center of gravity of a hand provided by an exemplary embodiment of the present disclosure;

[0024] FIG4B is a schematic diagram of a clockwise angle between a line connecting a finger point and the center of gravity of a hand connected domain and a first direction in an image captured by the image capture device in FIG3 ;

[0025] FIG5A is a flow chart of another interactive reading method provided by an exemplary embodiment of the present disclosure;

[0026] FIG5B is a flow chart of another interactive reading method provided by an exemplary embodiment of the present disclosure;

[0027] FIG6 is a schematic structural diagram of an interactive reading device provided by an exemplary embodiment of the present disclosure.

[0028] FIG7 is a schematic structural diagram of another interactive reading device provided by an exemplary embodiment of the present disclosure.

[0029] Details

[0030] To make the objectives, technical solutions and advantages of the present disclosure more clearly understood, the embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. It should be noted that, unless there is a conflict, the embodiments and features in the embodiments of the present disclosure can be combined with each other in any manner.

[0031] Unless otherwise defined, the technical or scientific terms used in the embodiments of the present disclosure should have the ordinary meaning understood by people with ordinary skills in the field to which the present disclosure belongs. The words "first", "second" and similar words used in the embodiments of the present disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. The words "include" or "comprising" and similar words mean that the elements or objects preceding the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects.

[0032] As shown in FIG1 , an embodiment of the present disclosure provides an interactive reading method, including:

[0033] Step 101: Acquire an image captured by an image acquisition device, where the image captured by the image acquisition device includes a user's finger and book content;

[0034] Step 102: Identify the user's finger position and finger movement trajectory based on the image captured by the image capture device;

[0035] Step 103: Determine whether the user's action is a circle reading gesture or a point reading gesture based on the user's finger position and finger movement trajectory;

[0036] Step 104: Determine the book content corresponding to the circle reading gesture or the point reading gesture, and announce the book content corresponding to the circle reading gesture or the point reading gesture.

[0037] The interactive reading method provided by the embodiments of the present disclosure identifies the user's finger position and finger movement trajectory from the image captured by the image acquisition device; judges the user's action as a circle reading gesture or a point reading gesture based on the identified finger position and finger movement trajectory; then determines the book content corresponding to the circle reading gesture or the point reading gesture, and broadcasts the book content corresponding to the circle reading gesture or the point reading gesture, supports text selection by circling or pointing with the finger, and solves the problem of time-consuming and labor-intensive point reading of single words; and the finger circle reading interaction does not require special gestures for pre-activation, and the interaction is more natural.

[0038] The finger point position described in the embodiment of the present disclosure is the fingertip position. When determining the movement trajectory of the user's finger point, the movement trajectory of the finger point farthest from the center of gravity of the hand can be determined as the movement trajectory of the user's finger point.

[0039] In some exemplary embodiments, as shown in Figure 2 , the image acquisition device may include: a camera or other image acquisition devices. In the embodiment of the present disclosure, the user's finger and the book are both located within the target acquisition area of ​​the image acquisition device.

[0040] In some exemplary embodiments, the method further comprises: pre-processing the image captured by the image capture device.

[0041] Exemplarily, the preprocessing may include any one or more of the following operations: image denoising, image enhancement, image resizing, image rotation and flipping, image translation and affine transformation, image segmentation, image background removal, etc.

[0042] Among them, the image denoising operation can reduce the noise of the image through filtering and other methods; the image enhancement operation can improve the detail information of the image by adjusting the image parameters such as contrast and brightness; the image resizing operation can adjust the size of the image through operations such as scaling and cropping to adapt to different application scenarios and needs; the image translation and affine transformation operations are used to adjust the position and size of the image, such as translation, scaling, and distortion; the image segmentation operation is to divide the image into multiple sub-regions for better subsequent processing; the image background removal operation removes background information in the image to better detect and recognize targets.

[0043] As shown in FIG2 , the image acquisition device is generally placed diagonally above a person's hand for shooting. The effect of the acquired image after being flipped 180 degrees is shown in FIG3 .

[0044] In some exemplary embodiments, identifying the user's finger position and finger movement trajectory based on the image captured by the image capture device includes:

[0045] Segmenting the user's hands from the image captured by the image acquisition device using a hand segmentation model to obtain one or more hand connected domains;

[0046] Determine one or more salient point positions contained in each hand connected domain, and select one or more finger point positions from the salient point positions;

[0047] Detecting whether the number of finger points in each hand connected domain is less than or equal to a first preset number;

[0048] Select a hand connected domain with the number of finger points less than or equal to a first preset number as the target hand, determine the hand detection frame corresponding to the target hand and the finger points bound to the hand detection frame, and calculate the movement trajectory of the finger points bound to the hand detection frame through a tracking algorithm.

[0049] In the disclosed embodiment, after acquiring the image captured by the image acquisition device, a pre-trained lightweight hand segmentation model can be used to segment the hand from the entire image. Among them, the hand segmentation model can use lightweight hand segmentation models such as PIDNet (a real-time semantic segmentation model guided by an attention mechanism) and PP-LiteSeg (a real-time semantic segmentation model using a codec architecture). The hand segmentation model displays the hand segmentation result in the captured image through a mask to obtain one or more hand connected domains (usually one hand corresponds to one hand connected domain range).

[0050] In some exemplary embodiments, when there are multiple hand connected domains in which the number of finger points is less than or equal to a first preset number, the hand connected domain with the number of finger points less than or equal to the first preset number and the largest area can be selected as the target hand.

[0051] When multiple hand connected domains are obtained, the number of finger points in the hand connected domains can be detected one by one in the order of the area of ​​the hand connected domains from large to small. When it is detected that the number of finger points in a certain hand connected domain is less than or equal to the first preset number, the hand connected domain is determined as the target hand, the finger point positions and finger point movement trajectories of the target hand are identified, and the user's action is judged as a circle reading gesture or a point reading gesture based on the identified finger point positions and finger point movement trajectories of the target hand. The hand connected domains with an area smaller than the currently detected hand connected domain in the current frame image can no longer undergo the above-mentioned interactive reading detection, that is, they can be reported according to the interactive reading operation corresponding to the currently detected target hand. In other exemplary embodiments, all hand connected domains with a number of finger points less than or equal to the first preset number can also be regarded as target hands, that is, the number of selected target hands is greater than 1.

[0052] As shown in Figure 4A, by detecting the contour of each connected domain of the hand, multiple salient point coordinates, multiple concave point coordinates, and a center of gravity coordinate are obtained. One or more candidate finger points are selected from the multiple salient points, and then one or more final finger points are selected from the candidate finger points. In the disclosed embodiment, not all salient points are candidate finger points. In Figure 3, only the salient point corresponding to the index finger is a candidate finger point. The salient points corresponding to other fingers are not candidate finger points because they are curled up or blocked.

[0053] In some exemplary embodiments, candidate finger points may be selected in the following manner:

[0054] The distance between the candidate finger point and the center of gravity of the hand connected domain is greater than or equal to a first preset threshold.

[0055] In some exemplary embodiments, the first preset threshold=ratio_thresh1*sqrt(contour_area1), where ratio_thresh1 is between 0 and 1, and contour_area1 is the area of ​​the hand connected region.

[0056] Exemplarily, ratio_thresh1 may be 0.5. However, the embodiment of the present disclosure does not limit this, and the value of ratio_thresh1 may be adjusted as needed.

[0057] In some exemplary embodiments, as shown in FIG4B , the final finger point may be selected as follows:

[0058] The clockwise angle α between the final line connecting the finger point and the center of gravity of the hand connected domain and the first direction X is between angle_th1 and angle_th2, where angle_th1 is a first preset angle, angle_th2 is a second preset angle, and the first direction X is the row direction of the text content in the image captured by the image capture device.

[0059] For example, angle_th1 can be 45°, and angle_th2 can be 135°. This means that the clockwise angle between the line connecting all candidate finger points and the center of gravity and the first direction is calculated, and the candidate finger point whose clockwise angle between the line and the first direction falls within [45°, 135°] is selected as the final finger point. However, the disclosed embodiments are not limited to this, and the values ​​of angle_th1 and angle_th2 can be adjusted as needed. The white dot at the top of the index finger in Figure 3 is the calculated final finger point position. In Figure 4B, the second direction Y is perpendicular to the first direction X.

[0060] In some exemplary embodiments, the first preset number may be 2.

[0061] In this disclosed embodiment, the algorithm detects the number of finger points within a single hand connected domain. When the number of finger points within a single hand connected domain is greater than or equal to three, it determines that the hand corresponding to that single hand connected domain has not performed any interactive reading operations. At this point, if there are other hand connected domains, the number of finger points in these other hand connected domains is detected again. If the number of finger points in all hand connected domains is greater than or equal to three, the algorithm acquires the next image frame and continues detecting interactive reading for the next frame.

[0062] In some exemplary embodiments, determining a hand detection frame corresponding to a target hand includes:

[0063] The minimum value of the horizontal and vertical coordinates of all contour points in the hand connected domain of the target hand is used as the coordinate of the lower left corner point of the hand detection frame corresponding to the target hand, and the maximum value of the horizontal and vertical coordinates of all contour points is used as the coordinate of the upper right corner point of the hand detection frame corresponding to the target hand. The coordinates of the lower left corner point and the upper right corner point are used as the diagonal coordinates to construct a rectangular hand detection frame;

[0064] The finger point with the largest vertical coordinate among all the finger points is recorded as the finger point bound to the hand detection frame. As shown in FIG4B , the coordinates of the top finger point are selected as the coordinates of the finger point bound to the hand detection frame.

[0065] In some exemplary embodiments, calculating the movement trajectory of the finger points bound to the hand detection frame using a tracking algorithm includes:

[0066] When the distance between the hand detection frame of the new tracking serial number (ID) and the hand detection frame of the most recent tracking ID is less than a first preset distance, the new tracking ID and the most recent tracking ID are merged. The first preset distance is ratio_thresh2*sqrt(contour_area2), where ratio_thresh2 is between 0 and 1, and contour_area2 is the area of ​​the hand connected domain corresponding to the new tracking ID.

[0067] In the disclosed embodiment, after obtaining the hand detection frame, a tracking algorithm is used to establish the movement trajectory of the finger points bound to the hand detection frame. Exemplarily, the tracking algorithm may be the SORT (Simple Online and Realtime Tracking) tracking algorithm. The SORT tracking algorithm is a simple online real-time multi-target tracking algorithm that primarily uses Kalman filtering to propagate target objects into future frames and then uses IoU as a metric to establish relationships and achieve multi-target tracking.

[0068] As the hand moves, motion blur may cause poor hand segmentation results in some frames, making it impossible to identify the target hand and causing tracking to be lost. To address this issue, the disclosed embodiment calculates the Euclidean distance between the coordinates of the finger point bound to the hand detection frame corresponding to the new tracking ID and the coordinates of the finger point bound to the hand detection frame in the last frame of the most recent tracking ID each time a new tracking ID appears. If the distance value is less than a first preset threshold, it is considered that the new tracking ID has lost tracking due to motion blur, and the trajectory of the new tracking ID is merged with the most recent tracking ID.

[0069] In some exemplary embodiments, ratio_thresh2 may be 0.5. However, the embodiment of the present disclosure is not limited thereto, and the value of ratio_thresh2 may be adjusted as needed.

[0070] In the embodiment of the present disclosure, the first preset distance may be set to 0.5*sqrt(contour_area2), where contour_area2 is the area of ​​the hand connected domain corresponding to the new tracking ID.

[0071] Since the image acquisition device (which can be the built-in camera of the reading robot) shoots the hand from an oblique angle above, there will be a distortion effect of "large at the top and small at the bottom". The present disclosure sets the first preset distance to ratio_thresh2*sqrt(contour_area2). Contour_area2 will change according to the size of the hand connected domain detected each time. That is, the first preset distance uses an adaptive threshold setting, which can avoid the problem of fixed threshold being unsuitable due to the change in the area of ​​the hand in the shooting picture when the hand moves up and down.

[0072] In some exemplary embodiments, determining whether to start an interactive reading operation based on the user's finger position and finger movement trajectory includes:

[0073] Determine the tracking ID of the target hand and the finger point position corresponding to the tracking ID in each frame image;

[0074] Determine whether the distances of the finger point positions corresponding to the most recent N image frames relative to the reference finger point position do not exceed a second preset distance, where the most recent N image frames include the current image frame and N-1 image frames before the current image frame, where N is a natural number greater than 1;

[0075] When the distances between the finger point positions corresponding to the latest N frames of image and the reference finger point position do not exceed a second preset distance, it is determined that the interactive reading operation is started.

[0076] In some exemplary embodiments, the reference finger point position may be the finger point position corresponding to the Nth frame image before the current frame image, however, embodiments of the present disclosure are not limited thereto. For example, the reference finger point position may also be the finger point position corresponding to the frame image before the current frame image for which the finger point position is to be determined.

[0077] The interactive reading method of the disclosed embodiment does not require the user to use specific gestures to activate the interactive reading function when performing finger interaction. Instead, the user only needs to pause their finger briefly when completing the circle reading or point reading operation, which is in line with user usage habits. Specifically, the algorithm determines whether the Euclidean distance of the finger point position coordinates corresponding to the tracking ID established by the hand detection frame of the current frame does not exceed a second preset distance in the N historical frames. If the distance does not exceed the second preset distance, it is considered that the finger is almost stationary (i.e., it is paused briefly), and then the "circle reading" or "point reading" interactive operation is initiated for the user.

[0078] In some exemplary embodiments, the second preset distance may be a distance of m1 pixels, where m1 is a natural number between 5 and 15. For example, m1 may be 10. However, the embodiment of the present disclosure is not limited to this, and the value of m1 may be adjusted as needed.

[0079] In some exemplary embodiments, N may be a natural number between 5 and 10.

[0080] Exemplarily, N can be 8. However, the embodiment of the present disclosure does not limit this, and the value of N can be adjusted as needed.

[0081] In some exemplary embodiments, determining whether the user's action is a circle reading gesture or a point reading gesture includes:

[0082] Determine the tracking ID of the target hand and the finger point position corresponding to the tracking ID in each frame image;

[0083] Detect whether the finger point positions corresponding to some or all of the frames before the current frame image form a closed figure;

[0084] When the finger point positions corresponding to part or all of the frame images form a closed figure, the encircled area of ​​the closed figure is calculated;

[0085] When the circled area of ​​the closed figure is greater than or equal to a first preset area, determining that the user's action is a circle reading gesture;

[0086] When the finger point positions corresponding to the partial or full frame images do not form a closed figure, or when the finger point positions corresponding to the partial or full frame images form a closed figure but the circled area of ​​the closed figure is smaller than the first preset area, the user's action is determined to be a point-reading gesture.

[0087] In the embodiment of the present disclosure, since the finger point positions corresponding to multiple image frames are discrete points, when detecting whether the finger point positions corresponding to some or all of the frame images before the current frame image constitute a closed figure, the constituted closed figure can be an approximate closed figure, that is, as long as the distance between the starting point and the end point is less than the third preset distance, it can be considered that all points between the starting point and the end point approximately constitute a closed figure.

[0088] In the embodiment of the present disclosure, when calculating the encircled area of ​​a closed figure, it can be the area of ​​a closed polygon composed of multiple finger point positions, or it can be the area of ​​the minimum circumscribed rectangle of a closed figure composed of multiple finger point positions. The embodiment of the present disclosure does not limit this.

[0089] In some exemplary embodiments, detecting whether finger point positions corresponding to a portion of or all of the frames before the current frame form a closed graph includes:

[0090] Starting from the first frame image before the first reference image frame, the distance between the finger point position corresponding to the M frames image before the first reference image frame and the second reference finger point position is detected one by one, where the first reference image frame is the Nth frame image before the current frame image, and the second reference finger point position is the finger point position corresponding to any frame image between the first reference image frame and the current frame image, where N and M are both natural numbers greater than 1;

[0091] When the distance between the finger point position corresponding to the Jth frame image before the first reference image frame and the second reference finger point position does not exceed a third preset distance, determining that the finger point positions corresponding to all frame images between the Jth frame image before the first reference image frame and the current frame image form a closed figure, where J is a natural number less than or equal to M;

[0092] The finger point position corresponding to the J-th frame image before the first reference image frame is determined as the circle selection starting point, and the second reference finger point position is determined as the circle selection end point.

[0093] Exemplarily, M may be 100. However, the embodiment of the present disclosure does not limit this, and the value of M may be adjusted as needed.

[0094] Exemplarily, the third preset distance can be a distance of m2 pixels (pix), where m2 is a natural number between 5 and 15. Exemplarily, m2 can be 10. However, the embodiment of the present disclosure is not limited to this, and the value of m2 can be adjusted as needed.

[0095] Exemplarily, the first preset area may be an area of ​​m3 pixels, where m3 is a natural number between 150 and 250. Exemplarily, m3 may be 200. However, the embodiment of the present disclosure does not impose a limitation on this, and the value of m3 may be adjusted as needed.

[0096] In the disclosed embodiment, for a tracking ID that triggers an interactive reading operation, the coordinate motion trajectory of its finger point is retraced point by point; if the Euclidean distance between a point in the historical trajectory and the last stop point of the tracking ID does not exceed a third preset distance (for example, the third preset distance may be 10 pix), the historical point is considered to be the starting point of a finger circle selection, and the last stop point is the end point of the circle selection; the area of ​​the polygon enclosed by all the track points between the circle selection start point and the circle selection end point in the captured image is calculated. If the area of ​​the polygon is greater than or equal to a first preset area (for example, the first preset area may be 200 pix), the area of ​​the current circle selection is considered to be large enough, and the "circle reading" interactive operation is enabled for the user. When the number of track points retraced exceeds M (for example, M may be 100) points, and no circle selection start point that meets the requirements for triggering the "circle reading" interaction has been found or the area of ​​the enclosed polygon is less than 200 pix, the "point reading" interactive operation is enabled for the user, and the coordinates of the finger point in the last frame are used as the target point of the "point reading" operation.

[0097] In the embodiment of the present disclosure, since the coordinates of the finger points in multiple frames of images are not continuous, when determining the closed figure enclosed by all the trajectory points between the circle starting point and the circle end point in the captured image, adjacent trajectory points and the circle starting point and the circle end point can be connected by line segments, thereby obtaining a closed figure enclosed by all the trajectory points between the circle starting point and the circle end point in the captured image.

[0098] When the "Point Reading" interactive operation is enabled, optical character recognition (OCR) is performed on the captured image. The coordinates of the center point of each word OCR detection frame are calculated, and the word detection frame closest to the "Point Reading" target point is selected for text recognition. The recognition results are synthesized into a TTS (Text To Speech) voice signal and played for the user.

[0099] Text-to-speech (TSS) technology converts text into sound, offering a variety of voice options for various scenarios and languages. It supports customizable parameters like volume and speaking rate, making pronunciation more professional and tailored to specific scenarios. Speech synthesis is widely used in business scenarios such as intelligent customer service, audio reading, news broadcasting, and human-computer interaction.

[0100] In some exemplary embodiments, when it is determined that the user's action is a circle reading gesture, determining the book content corresponding to the circle reading gesture includes:

[0101] Perform text detection on the captured image to obtain one or more line text detection frames;

[0102] Determine the circumscribed rectangle of the closed figure as the selected rectangle;

[0103] Calculate the IOU (Intersection over Union) score between each line of text detection box and the selected rectangle;

[0104] The line of text corresponding to the line text detection box whose calculated IOU score is greater than or equal to the second preset threshold is determined as the book content corresponding to the circle reading gesture (ie, the circled text).

[0105] In other embodiments, when determining the book content corresponding to the circle reading gesture, the IOU score between each line text detection box and the closed graphic can also be directly calculated, and the line text corresponding to the line text detection box whose calculated IOU score is greater than or equal to the second preset threshold is determined as the book content corresponding to the circle reading gesture (i.e., the circled text).

[0106] When the "Circle Reading" interactive operation is turned on, line text detection is performed on the captured image. The same processing flow as the previous steps is used to obtain the circumscribed rectangle of the circled polygon as the circled rectangle, and the IOU score of each line text detection box and the circled rectangle is calculated using the following formula, where area ocr Indicates the area of ​​the line text detection box, area poly Indicates the area of ​​the selected rectangle, area ocr ∩area poly Indicates the intersection area of ​​the line text detection box and the selected rectangle. IOU=(area ocr ∩area poly ) / area ocr .

[0107] When the IOU score calculated for a line of text detection boxes is greater than or equal to a second preset threshold, the line of text is considered to be selected by the user. The second preset threshold can be between 0 and 1, and for example, the second preset threshold can be 0.9. After obtaining all the circled line text boxes, text recognition is performed in order from top to bottom of the vertical coordinate, and the recognition results are synthesized into a TTS voice signal and played for the user.

[0108] The interactive reading method provided by the present disclosure calculates the IOU score between the finger-circled area and the OCR text detection frame, selects the appropriate text frame for recognition and voice signal synthesis, improves the accuracy of circling, and ensures user experience.

[0109] In other exemplary embodiments, when it is determined that the user's action is a circle reading gesture, determining the book content corresponding to the circle reading gesture includes:

[0110] Perform single-character detection on the captured image to obtain the detection frame position information of each character;

[0111] Based on the detection frame position information of each character and the position information of the closed figure, it is determined whether each character is located within the closed figure, and the characters located within the closed figure are determined as the book content corresponding to the circle reading gesture.

[0112] In the embodiment of the present disclosure, single-word detection is performed on the captured image after the algorithm determines that the user's action is a circle reading gesture, or it can be performed when the algorithm determines that the user's action is a circle reading gesture but the algorithm detects that no entire line of text is selected (that is, the calculated IOU of each line of text detection box and the circled rectangle is less than the preset second threshold). The embodiment of the present disclosure does not impose any restrictions on this.

[0113] In some exemplary embodiments, when the calculated IOU between each line of text detection box and the circled rectangle is less than a preset second threshold, determining the book content corresponding to the circle reading gesture includes:

[0114] Select one or more lines of text as candidate processing text based on the calculated IOU;

[0115] Perform single-word detection on the candidate processing text to obtain the detection box position information of each character in the candidate processing text;

[0116] Based on the detection box position information of each character in the candidate processing text and the position information of the circled rectangle (or closed figure), it is determined whether each character in the candidate processing text is located within the circled rectangle (or closed figure), and the characters located within the circled rectangle (or closed figure) are determined as the book content corresponding to the circle reading gesture.

[0117] In the embodiment of the present disclosure, the candidate processing text can be one line or multiple lines, and the embodiment of the present disclosure does not limit this. Optionally, when the selection rectangle includes multiple lines of text, only the line of text with the largest calculated IOU is selected as the candidate processing text, that is, when the calculated IOU of each line of text detection box and the selection rectangle is less than the preset second threshold, regardless of whether the actual selection rectangle includes one line or multiple lines of text, only the line of text with the largest IOU is selected as the candidate processing text.

[0118] In the embodiment of the present disclosure, determining whether each character in the candidate processing text is located within the enclosed rectangle based on the detection box position information of each character in the candidate processing text and the position information of the enclosed rectangle may include:

[0119] Record the horizontal and vertical coordinates of the upper left corner point of the detection box of a single character as (left_x1, left_y1), and the horizontal and vertical coordinates of the lower right corner point as (right_x1, right_y1); record the horizontal and vertical coordinates of the upper left corner point of the selected rectangle as (left_x2, left_y2), and the horizontal and vertical coordinates of the lower right corner point as (right_x2, right_y2). When the condition (left_x1>left_x2)&&(left_y1<left_y2)&&(right_x1<right_x2)&&(right_y1<right_y2) is satisfied, it is determined that the character is located within the selected rectangle and is the selected character.

[0120] In the embodiments of the present disclosure, the first direction X can be set as the X-axis direction, and the second direction Y can be set as the Y-axis direction. Correspondingly, the coordinate position corresponding to the X-axis of each pixel point is the horizontal coordinate, and the coordinate position corresponding to the Y-axis of each pixel point is the vertical coordinate. When all the selected characters in a row are determined, the TTS voices of all the selected characters in that row can be played for the user in ascending order of the horizontal coordinates of the upper left corner points of the detection boxes. If the selected characters are located in multiple rows, the TTS voices of all the selected characters in each row can be broadcast in sequence according to the row order. When broadcasting the TTS voices of all the selected characters in each row, they are broadcast in ascending order of the horizontal coordinates of the upper left corner points of the detection boxes in sequence.

[0121] In some exemplary embodiments, as shown in FIGS. 5A and 5B, an interactive reading method according to an embodiment of the present disclosure includes:

[0122] Step 1: Use the camera built into the reading robot to collect images in real time. Generally, the camera is placed obliquely above the human hand for shooting, and the obtained image after a 180-degree flip is shown in FIG. 3. After obtaining the image, use the pre-trained lightweight hand segmentation model to segment the hand from the entire image. The hand segmentation model can be a lightweight hand segmentation model such as PIDNet or PP-LiteSeg. As shown in FIG. 3, the hand segmentation result of the captured image is displayed in the form of a mask.

[0123] Step 2: Process the acquired segmentation mask to obtain multiple salient point coordinates and a centroid coordinate for each connected domain. Select salient point coordinates whose distance from the centroid coordinate is greater than or equal to ratio_thresh1*sqrt(contour_area1) as candidate finger point coordinates, where contour_area1 is the area of ​​the connected domain and ratio_thresh1 can be a real number between 0 and 1, for example, 0.5. Determine the direction of the line connecting all candidate finger points and the centroid point, and select the candidate finger point whose clockwise angle between the connecting line direction and the horizontal direction falls within [angle_th1, angle_th2] as the final finger point, where angle_th1 and angle_th2 can be 45 and 135, respectively. As shown in Figure 3, the white dot on the index finger is the final selected finger point coordinate.

[0124] Step 3: For a connected domain in a single-frame image, if the final number of finger points screened out is ≥3, it is determined that the hand corresponding to the connected domain has not performed any interactive reading action; for all connected domains with the final number of finger points screened out being less than 3, the connected domain with the largest area is selected as the hand connected domain of the target hand.

[0125] Step 4: For the hand connected domain of the selected target hand, calculate the minimum value of the horizontal and vertical coordinates of all contour points in the hand connected domain of the target hand as the coordinate of the lower left corner point of the hand detection frame, and the maximum value as the coordinate of the upper right corner point of the detection frame, to obtain the hand detection frame corresponding to the target hand connected domain. Then, select the finger point closest to the camera (i.e., the finger point with the largest vertical coordinate) and record it as the finger point bound to the hand detection frame of the target hand. For the shooting effect in Figure 3, the coordinates of the top finger point are selected as the coordinates of the finger point bound to the hand detection frame.

[0126] Step 5: After obtaining the hand detection frame, use the tracking algorithm to establish the movement trajectory of the finger points bound to the target hand. Exemplarily, the SORT tracking algorithm can be used for target tracking. Since the hand is moving, the hand segmentation results of some frames may be poor due to motion blur, and the target hand cannot be screened out and the tracking is lost. In order to solve this problem, the present disclosure calculates the Euclidean distance between the coordinates of the finger points bound to the hand detection frame corresponding to the new tracking ID and the coordinates of the finger points bound to the hand detection frame of the last frame of the most recent tracking ID each time a new tracking ID appears. If the distance value is less than a pre-set threshold, it is considered that the new tracking ID is a loss of tracking due to motion blur, and the new tracking ID is merged with the most recent tracking ID in history. Exemplarily, the pre-set threshold can be set to 0.5*sqrt(contour_area2), where contour_area2 is the area of ​​the hand connected domain corresponding to the latest tracking ID. Since the built-in camera of the reading robot shoots the hands from an oblique angle above, there will be a distortion effect of "bigger at the top and smaller at the bottom". The present disclosure uses adaptive threshold settings to solve the problem of fixed threshold being unsuitable due to the change in area in the shooting picture when the hands move up and down.

[0127] Step 6: For the tracking ID established by the hand detection frame in the current frame, calculate the Euclidean distance of the finger coordinate movement of the ID in the historical N frames. If the distance does not exceed 10pix, it is considered that the finger is almost stationary, and the user initiates a "circle reading" or "touch reading" interaction. When performing finger interaction, the present disclosure does not require the use of specific gestures to activate the function. The user only needs to pause the finger briefly when completing the interaction, which is in line with user usage habits.

[0128] Step 7: After the algorithm determines that the user has initiated a "circle read" or "touch read" interaction, the algorithm needs to determine whether the interaction is a "circle read" or "touch read" operation. Specifically, for the tracking ID that triggered the interaction, the coordinate movement trajectory of its finger point is retraced point by point; if the Euclidean distance between a point in the historical trajectory and the last point of the tracking ID does not exceed 10pix, the historical point is considered to be the starting point of a finger circle selection; the area of ​​the polygon enclosed by all the track points from the starting point to the last point in the captured image is calculated. If the area of ​​the polygon is greater than 200pix, the area of ​​this circle selection is considered to be large enough, and the "circle read" interaction is enabled for the user. If the number of track points retraced exceeds 100 points and no circle selection starting point that meets the requirements for triggering the "circle read" interaction is found or the area of ​​the enclosed polygon is less than or equal to 200pix, the "touch read" interaction is enabled for the user, and the finger point coordinates of the last frame are used as the target point for "touch read".

[0129] Step 8: When the "Point Reading" interactive operation is enabled, single-word OCR detection is performed on the captured image. The coordinates of the center point of each single-word detection frame are calculated, and the single-word detection frame closest to the "Point Reading" target point is selected for text recognition. The recognition results are synthesized into a TTS voice signal and played for the user.

[0130] Step 9: When the "Circle Reading" interactive operation is turned on, perform line text detection on the captured image. Use the same processing flow as in step 4 to obtain the circumscribed rectangle of the circled polygon, and use the following formula to calculate the IOU score between each line text detection box and the circled rectangle, where area ocr ∩area poly Indicates the area intersection of the line text detection box and the selected rectangle. IOU=(area ocr ∩area poly ) / area ocr ;

[0131] When the calculated IOU score for a line of text detection is greater than 0.9, the line of text is considered to be selected by the user. After obtaining all the selected line text boxes, text recognition is performed in order from top to bottom of the vertical coordinate, and the recognition results are synthesized into a TTS voice signal and played for the user.

[0132] As shown in FIG6 , the embodiment of the present disclosure further provides an interactive reading device, comprising: an acquisition module 601 , an identification module 602 , a determination module 603 , and a reporting module 604 , wherein:

[0133] An acquisition module 601 is configured to acquire an image captured by an image acquisition device, where the image captured by the image acquisition device includes a user's finger and book content;

[0134] The recognition module 602 is configured to recognize the user's finger position and finger movement trajectory based on the image captured by the image capture device;

[0135] The determination module 603 is configured to determine whether the user's action is a circle reading gesture or a point reading gesture based on the user's finger position and finger movement trajectory;

[0136] The announcement module 604 is configured to determine the book content corresponding to the circle reading gesture or the point reading gesture, and announce the book content corresponding to the circle reading gesture or the point reading gesture.

[0137] In some exemplary embodiments, the recognition module 602 recognizes the user's finger position and finger movement trajectory based on the image captured by the image capture device, including:

[0138] Segmenting the user's hands from the image captured by the image acquisition device using a hand segmentation model to obtain one or more hand connected domains;

[0139] Determine one or more salient point positions included in each hand connected domain, and select one or more finger point positions from the salient point positions;

[0140] Detecting whether the number of finger points in each hand connected domain is less than or equal to a first preset number;

[0141] Select a hand connected domain with the number of finger points less than or equal to a first preset number as the target hand, determine the hand detection frame corresponding to the target hand and the finger points bound to the hand detection frame, and calculate the movement trajectory of the finger points bound to the hand detection frame through a tracking algorithm.

[0142] In some exemplary embodiments, when there are multiple hand connected domains in which the number of finger points is less than or equal to a first preset number, the recognition module 602 is further configured to: select the hand connected domain in which the number of finger points is less than or equal to the first preset number and the area is the largest as the target hand.

[0143] In some exemplary embodiments, selecting one or more finger point positions from the salient point positions includes: selecting one or more candidate finger points from the salient points, wherein the distance between the candidate finger points and the center of gravity of the hand connected domain is greater than or equal to a first preset threshold; and selecting one or more final finger points from the candidate finger points.

[0144] In some exemplary embodiments, the distance between the candidate finger point and the center of gravity of the hand connected domain is greater than or equal to a first preset threshold, where the first preset threshold = ratio_thresh1*sqrt(contour_area1), where ratio_thresh1 is between 0 and 1, and contour_area1 is the area of ​​the hand connected domain.

[0145] In some exemplary embodiments, when selecting one or more final finger points from the candidate finger points, the clockwise angle between the line connecting the final finger point and the center of gravity of the hand connected domain and the horizontal direction is between angle_th1 and angle_th2, where angle_th1 is a first preset angle and angle_th2 is a second preset angle.

[0146] In some exemplary embodiments, the recognition module 602 calculates the movement trajectory of the finger point bound to the hand detection frame using a tracking algorithm, including:

[0147] When the distance between the hand detection frame of the new tracking ID and the hand detection frame of the most recent tracking ID is less than a first preset distance, the new tracking ID and the most recent tracking ID are merged. The first preset distance is ratio_thresh2*sqrt(contour_area2), where ratio_thresh2 is between 0 and 1, and contour_area2 is the area of ​​the hand connected domain corresponding to the new tracking ID.

[0148] In some exemplary embodiments, the determination module 603 is further configured to determine whether to start an interactive reading operation based on the user's finger point position and finger point movement trajectory.

[0149] In some exemplary embodiments, the determination module 603 determines whether to start the interactive reading operation based on the user's finger position and finger movement trajectory, including:

[0150] Determine the tracking ID of the target hand and the finger point position corresponding to the tracking ID in each frame of the image;

[0151] Determining whether the distances of the finger point positions corresponding to the most recent N image frames relative to the reference finger point position do not exceed a second preset distance, where the most recent N image frames include the current frame image and the N-1 frame image before the current frame image, and the reference finger point position is the finger point position corresponding to the Nth frame image before the current frame image, where N is a natural number greater than 1;

[0152] When the distances between the finger point positions corresponding to the latest N frames of image and the reference finger point position do not exceed a second preset distance, it is determined that the interactive reading operation is started.

[0153] In some exemplary embodiments, the determining module 603 determines whether the user's action is a circle reading gesture or a point reading gesture, including:

[0154] Determine the tracking ID of the target hand and the finger point position corresponding to the tracking ID in each frame of the image;

[0155] Detect whether the finger point positions corresponding to some or all of the frames before the current frame image form a closed figure;

[0156] When the finger point positions corresponding to part or all of the frame images form a closed figure, the encircled area of ​​the closed figure is calculated;

[0157] When the circled area of ​​the closed figure is greater than or equal to a first preset area, determining that the user's action is a circle reading gesture;

[0158] When the finger point positions corresponding to the partial or full frame images do not form a closed figure, or when the finger point positions corresponding to the partial or full frame images form a closed figure but the circled area of ​​the closed figure is smaller than the first preset area, the user's action is determined to be a point-reading gesture.

[0159] In some exemplary embodiments, when it is determined that the user's action is a circle reading gesture, the broadcast module 604 determines the book content corresponding to the circle reading gesture, including: performing text detection on the captured image; determining the circumscribed rectangle of the closed figure as a circle selection rectangle; calculating the IOU of each line text detection box and the circle selection rectangle; and determining the line text corresponding to the line text detection box whose calculated IOU is greater than or equal to a preset second threshold as the book content corresponding to the circle reading gesture.

[0160] In other exemplary embodiments, when it is determined that the user's action is a circle reading gesture, the broadcast module 604 determines the book content corresponding to the circle reading gesture, including: performing single-word detection on the captured image to obtain detection box position information of each character; based on the detection box position information of each character and the position information of the closed figure, determining whether each character is located within the closed figure, and determining the characters located within the closed figure as the book content corresponding to the circle reading gesture.

[0161] The interactive reading method and device of the disclosed embodiments can be applied to picture book robots. This disclosed embodiment supports the rapid selection of large paragraphs of text by circling with a finger, solving the time-consuming and labor-intensive problem of single-word reading. Furthermore, without requiring special gesture activation, a paragraph of text can be circled for OCR recognition and aloud while reading a picture book, resulting in a highly natural and real-time interactive experience.

[0162] The present disclosure also provides an interactive reading system, including an image acquisition device and an interactive reading device, wherein the interactive reading device is configured to:

[0163] Acquiring an image captured by the image capture device, where the image captured by the image capture device includes a user's finger and book content;

[0164] Identifying the user's finger position and finger movement trajectory based on the image captured by the image capture device;

[0165] Determining whether the user's action is a circle reading gesture or a point reading gesture based on the user's finger position and finger movement trajectory;

[0166] The book content corresponding to the circle reading gesture or the point reading gesture is determined, and the book content corresponding to the circle reading gesture or the point reading gesture is announced.

[0167] In the embodiment of the present disclosure, how the interactive reading device specifically performs interactive reading can refer to the embodiment of the interactive reading method described above, and will not be repeated here.

[0168] An embodiment of the present disclosure also provides an interactive reading device, comprising a memory; and a processor connected to the memory, wherein the memory is used to store instructions, and the processor is configured to execute the steps of the interactive reading method described in any embodiment of the present disclosure based on the instructions stored in the memory.

[0169] As shown in FIG7 , in one example, an interactive reading device may include: a processor 710, a memory 720, a bus system 730, and a transceiver 740. The processor 710, the memory 720, and the transceiver 740 are connected via the bus system 730. The memory 720 is configured to store instructions, and the processor 710 is configured to execute the instructions stored in the memory 720 to control the transceiver 740 to transmit and receive signals. Specifically, under the control of the processor 710, the transceiver 740 may acquire an image captured by an image acquisition device, where the image captured by the image acquisition device includes a user's finger and book content. The processor 710 may identify the user's finger position and finger movement trajectory based on the image captured by the image acquisition device; determine whether the user's action is a circle reading gesture or a point reading gesture based on the user's finger position and finger movement trajectory; determine the book content corresponding to the circle reading gesture or the point reading gesture; and announce the book content corresponding to the circle reading gesture or the point reading gesture.

[0170] It should be understood that the processor 710 may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.

[0171] The memory 720 may include a read-only memory and a random access memory, and provides instructions and data to the processor 710. A portion of the memory 720 may also include a non-volatile random access memory. For example, the memory 720 may also store information about the device type.

[0172] In addition to the data bus, the bus system 730 may also include a power bus, a control bus, a status signal bus, etc. However, for the sake of clarity, various buses are labeled as the bus system 730 in FIG.

[0173] During implementation, the processing performed by the processing device can be completed by the hardware integrated logic circuit in the processor 710 or by instructions in the form of software. That is, the method steps of the embodiment of the present disclosure can be embodied as being executed by a hardware processor, or being executed by a combination of hardware and software modules in the processor. The software module can be located in a storage medium such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 720, and the processor 710 reads the information in the memory 720 and completes the steps of the above method in combination with its hardware. To avoid repetition, it will not be described in detail here.

[0174] The present disclosure also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the interactive reading method described in any of the embodiments of the present disclosure. The method for driving interactive reading by executing executable instructions is substantially the same as the interactive reading method described in the aforementioned embodiments of the present disclosure and is not further described here.

[0175] In some possible implementations, various aspects of the interactive reading method provided by the present disclosure may also be implemented in the form of a program product, which includes program code. When the program product is run on a computer device, the program code is used to enable the computer device to execute the steps of the interactive reading method according to various exemplary implementations of the present disclosure described above in this specification. For example, the computer device may execute the interactive reading method recorded in the embodiments of the present disclosure.

[0176] The program product may employ any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.

[0177] It will be appreciated by those skilled in the art that all or some of the steps, systems, and functional modules / units in the methods disclosed above may be implemented as software, firmware, hardware, and appropriate combinations thereof. In hardware implementations, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed by several physical components in cooperation. Some or all components may be implemented as software executed by a processor, such as a digital signal processor or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or temporary medium). As is well known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable, and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, it is well known to those skilled in the art that communication media generally embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.

[0178] It should be noted that the above-described embodiments or implementations are merely illustrative and not restrictive. Therefore, the present disclosure is not limited to what is specifically shown and described herein. Various modifications, substitutions, or omissions may be made to the forms and details of the implementations without departing from the scope of the present disclosure.

Claims

1. An interactive reading system, comprising: An image acquisition device and an interactive reading device, wherein the interactive reading device is configured as follows: Acquire an image captured by the image acquisition device, wherein the image captured by the image acquisition device includes a user's finger and book content; According to the image captured by the image acquisition device, identifying the user's finger position and finger movement trajectory; According to the finger position and finger movement trajectory of the user, determining whether the user's action is a circle reading gesture or a point reading gesture; The book content corresponding to the circle reading gesture or the point reading gesture is determined, and the book content corresponding to the circle reading gesture or the point reading gesture is announced.

2. The interactive reading system according to claim 1, wherein: The interactive reading device is also configured as: Whether to start the interactive reading operation is determined according to the finger position and the finger movement trajectory of the user.

3. The interactive reading system according to claim 2, wherein: The determining whether to start the interactive reading operation according to the finger position and the finger movement trajectory of the user includes: Determine the tracking ID of the target hand and the finger point position corresponding to the tracking ID in each frame image; Determine whether the distances of the finger point positions corresponding to the latest N frame images relative to the first reference finger point position do not exceed a second preset distance, the latest N frame images include the current frame image and N-1 frame images before the current frame image, where N is a natural number greater than 1; When the distances between the finger point positions corresponding to the most recent N frames of images and the first reference finger point position do not exceed a second preset distance, it is determined that the interactive reading operation is started.

4. The interactive reading system according to claim 3, wherein: The first reference finger point position is a finger point position corresponding to the Nth frame image before the current frame image.

5. The interactive reading system according to claim 1, wherein: The determining whether the user's action is a circle reading gesture or a point reading gesture includes: Determine the tracking ID of the target hand and the finger point position corresponding to the tracking ID in each frame image; Detect whether the finger point positions corresponding to the partial or all frame images before the current frame image form a closed figure; When the finger point positions corresponding to some or all frame images form a closed figure, the encircled area of ​​the closed figure is calculated; When the circled area of ​​the closed figure is greater than or equal to a first preset area, determining that the user's action is a circle reading gesture; When the finger point positions corresponding to the partial frames or all frame images do not form a closed figure, or when the finger point positions corresponding to the partial frames or all frame images form a closed figure but the circled area of ​​the closed figure is smaller than the first preset area, the user's action is determined to be a point-reading gesture.

6. The interactive reading system according to claim 5, wherein: The detecting whether the finger point positions corresponding to the partial or all frame images before the current frame image form a closed figure includes: Starting from the first frame image before the first reference image frame, the M frames before the first reference image frame are detected one by one. The distance between the finger point position corresponding to the frame image and the second reference finger point position, the first reference image frame is the Nth frame image before the current frame image, the second reference finger point position is the finger point position corresponding to any frame image between the first reference image frame and the current frame image, and N and M are both natural numbers greater than 1; When the distance between the finger point position corresponding to the Jth frame image before the first reference image frame and the second reference finger point position does not exceed the third preset distance, it is determined that the finger point positions corresponding to all frame images between the Jth frame image before the first reference image frame and the current frame image constitute a closed figure, where J is a natural number less than or equal to M; the finger point position corresponding to the Jth frame image before the first reference image frame is determined as the starting point of the circle selection, and the second reference finger point position is determined as the end point of the circle selection.

7. The interactive reading system according to claim 5, wherein: When it is determined that the user's action is a circle reading gesture, determining the book content corresponding to the circle reading gesture includes: Perform text detection on the captured image to obtain one or more line text detection frames; Determine the circumscribed rectangle of the closed figure as the enclosed rectangle; Calculate the IOU between each line text detection box and the circled rectangle; The line of text corresponding to the line text detection box whose calculated IOU is greater than or equal to the preset second threshold is determined as the book content corresponding to the circle reading gesture.

8. The interactive reading system according to claim 5, wherein: When it is determined that the user's action is a circle reading gesture, determining the book content corresponding to the circle reading gesture includes: Perform single-word detection on the captured image to obtain the detection frame position information of each character; According to the detection frame position information of each character and the position information of the closed figure, it is determined whether each character is located in the closed figure, and the characters located in the closed figure are determined as the book content corresponding to the circle reading gesture.

9. The interactive reading system according to claim 1, wherein: According to the image captured by the image acquisition device, identifying the user's finger position and finger movement trajectory includes: Segmenting the user's hands from the image captured by the image acquisition device using a hand segmentation model to obtain one or more hand connected domains; Determine one or more salient point positions included in each of the hand connected domains, and select one or more finger point positions from the salient point positions; Detecting whether the number of finger points in each of the hand connected domains is less than or equal to a first preset number; A hand connected domain with a number of finger points less than or equal to a first preset number is selected as the target hand, a hand detection frame corresponding to the target hand and the finger points bound to the hand detection frame are determined, and a movement trajectory of the finger points bound to the hand detection frame is established through a tracking algorithm.

10. The interactive reading system according to claim 9, wherein: When there are multiple hand connected domains in which the number of finger points is less than or equal to a first preset number, selecting a hand connected domain in which the number of finger points is less than or equal to the first preset number as a target hand includes: The hand connected domain with the largest area and the number of finger points being less than or equal to the first preset number is selected as the target hand.

11. The interactive reading system according to claim 9, wherein: The selecting one or more finger point positions from the salient point positions comprises: Selecting one or more candidate finger points from the salient points, where the distance between the candidate finger points and the center of gravity of the hand connected domain is greater than or equal to a first preset threshold; One or more final finger points are selected from the candidate finger points.

12. The interactive reading system according to claim 11, wherein: The clockwise angle between the line connecting the final finger point and the center of gravity of the hand connected domain and the first direction is between angle_th1 and angle_th2, where angle_th1 is a first preset angle, angle_th2 is a second preset angle, and the first direction is the row direction of the text content in the image captured by the image acquisition device.

13. The interactive reading system according to claim 11, wherein: The step of establishing the movement trajectory of the finger points bound to the hand detection frame by using a tracking algorithm includes: When the distance between the hand detection frame of the new tracking ID and the hand detection frame of the most recent tracking ID is less than a first preset distance, the new tracking ID and the most recent tracking ID are track-merged.

14. The interactive reading system according to claim 13, wherein: The first preset distance is ratio_thresh2*sqrt(contour_area2), wherein ratio_thresh2 is between 0 and 1, and contour_area2 is the area of ​​the hand connected domain corresponding to the new tracking ID.

15. An interactive reading method, comprising: Acquire an image captured by an image acquisition device, wherein the image captured by the image acquisition device includes a user's finger and book content; According to the image captured by the image acquisition device, identifying the user's finger position and finger movement trajectory; According to the finger position and finger movement trajectory of the user, determining whether the user's action is a circle reading gesture or a point reading gesture; The book content corresponding to the circle reading gesture or the point reading gesture is determined, and the book content corresponding to the circle reading gesture or the point reading gesture is announced.

16. The interactive reading method according to claim 15, further comprising: Whether to start the interactive reading operation is determined according to the finger position and the finger movement trajectory of the user.

17. An interactive reading device, comprising a memory; and a processor connected to the memory, wherein the memory is used to store instructions, and the processor is configured to execute the steps of the interactive reading method according to any one of claims 15 to 16 based on the instructions stored in the memory.

18. A computer-readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, the interactive reading method according to any one of claims 15 to 16 is implemented.