Information processing method and apparatus for paper-screen interaction, storage medium, and computer device
By acquiring image frame sequences and recognizing the pointing position and changes in pointing of valid gestures, the problem of low information processing efficiency and accuracy in paper screen interaction is solved, achieving efficient and accurate text content transmission and improving the user experience.
Patent Information
- Application Number
- PCT/CN2024/121461
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-07
- Filing Date
- 2024-09-26
- Publication Date
- 2026-02-12
AI Technical Summary
In existing paper-screen interaction processes, information processing relies on manual operation, resulting in low interaction efficiency and low text recognition accuracy, making it difficult to achieve efficient information exchange and high-precision text content recognition.
By acquiring image frame sequences, the pointing position and pointing changes of valid gestures are identified. The mirror imaging principle and pre-trained image processing module are used to perform mirror rotation and region cropping. Combined with the gesture recognition module and text detection module, the area to be interacted with is determined and the text content is transmitted to the electronic device.
It improves the operational flexibility and text recognition accuracy of paper-screen interaction, enhances the user experience, and achieves efficient information exchange and high-precision text content recognition.
Smart Images

Figure CN2024121461_12022026_PF_FP_ABST
Abstract
Description
Paper screen interaction information processing method and device, storage medium and computer equipment
[0001] Cross-reference to related applications
[0002] This application claims priority to Chinese Patent Application No. 202411080520.8, filed August 7, 2024, the entire contents of which are incorporated herein by reference for all purposes. TECHNICAL FIELD
[0003] The present application relates to the technical field of information processing, in particular to a paper screen interaction information processing method, device, storage medium and computer equipment. BACKGROUND
[0004] With the rapid development of information technology, intelligent hardware has become an indispensable tool in the field of learning and education. They provide an interactive learning experience, allowing users to more easily access knowledge. However, despite the increasing power of intelligent hardware, offline learning remains the main scenario for learning, offline learning and online learning based on intelligent hardware complement each other, and paper screen interaction becomes the mainstream learning solution. Due to the limiting factors of physical conditions, offline paper materials and intelligent hardware are separated in learning scenarios, making it difficult for intelligent hardware and offline paper materials to interact in the paper screen interaction process.
[0005] The existing paper screen interaction process mainly relies on touch screens or other forms of direct input, which is not intuitive and convenient when dealing with paper documents. For example, when using a learning machine, if a user wants to refer to the content on a book or notes, they often need to perform multiple steps of manual operation to input or scan the paper material, which not only wastes time but also is prone to errors, making the paper screen interaction process inefficient for text information exchange. In addition, the learning machine often needs physical contact or specific markers between the paper and the device when recognizing the content of the paper material, which limits the user's freedom of operation and makes the paper screen interaction process less accurate for text content recognition.
[0006] SUMMARY
[0007] In view of the above, the present application provides a paper screen interaction information processing method, device, storage medium and computer equipment, the main purpose of which is to solve the problem that the existing paper screen interaction information processing relies on manual operation, making it difficult to achieve efficient information exchange and high-precision text content recognition.
[0008] According to a first aspect of the present application, a paper screen interaction information processing method is provided, comprising:
[0009] obtaining a sequence of image frames to be processed, the sequence of image frames comprising at least one image frame output by a target object in a paper screen interaction process;
[0010] When it is detected that there is a valid gesture in the target image frame, a pointing position of the valid gesture in the target image frame is recognized to obtain a pointing position of the valid gesture;
[0011] If the pointing position of the valid gesture moves, a pointing change of the valid gesture in the target image frame is recognized to obtain a pointing change of the valid gesture;
[0012] According to the pointing change, a to-be-interacted region in the paper region is determined, and text content of the to-be-interacted region is transmitted to an electronic device.
[0013] Further, the obtaining of the image frame sequence to be processed comprises
[0014] According to a mirror imaging principle, an image preview stream output by an image acquisition device is received in a paper screen interaction process, the image preview stream being a mirror-inverted image frame sequence collected by using the image acquisition device by a target object in the paper screen interaction process;
[0015] Each image frame in the image preview stream is subjected to mirror rotation and region cropping by a pre-trained image processing module, so as to form an image frame sequence to be processed by the mirror rotation and the region cropping.
[0016] Further, after the obtaining of the image frame sequence to be processed, the method further comprises:
[0017] Each image frame in the image frame sequence is subjected to gesture recognition by a pre-trained gesture recognition module, so as to determine whether there is a gesture object in a current image frame;
[0018] If there is a gesture object in the current image frame, whether there is a valid gesture in the current image frame is detected according to the gesture object;
[0019] If there is no gesture object in the current image frame, a next image frame in the image frame sequence is subjected to gesture recognition.
[0020] Further, the detecting of whether there is a valid gesture in the current image frame according to the gesture object comprises:
[0021] The gesture object in the current image frame is subjected to fingertip recognition by a pre-trained fingertip recognition module, so as to determine whether there is fingertip information in the gesture object;
[0022] If the fingertip information meets a set recognition condition, it is determined that there is a valid gesture in the current image frame;
[0023] If the fingertip information does not meet the set recognition condition, the next image frame in the image frame sequence is subjected to gesture recognition.
[0024] Further, the fingertip position recognition on the effective gesture in the target image frame comprises:
[0025] The continuous image frames within the preset time after the target image frame is intercepted from the image frame sequence;
[0026] The fingertip position recognition on the effective gesture in the continuous image frames is performed by using the pre-trained position recognition module, so as to determine the pointing position of the effective gesture according to the fingertip positions of the effective gesture within the preset time.
[0027] Further, the determination of the pointing position of the effective gesture according to the fingertip positions of the effective gesture within the preset time comprises:
[0028] If the fingertip positions of the effective gesture within the preset time move, it is determined that the pointing position of the effective gesture moves; otherwise, it is determined that the pointing position of the effective gesture does not move.
[0029] Correspondingly, after the pointing position of the effective gesture in the target image frame is determined when it is detected that the effective gesture exists in the target image frame, the method further comprises:
[0030] If the pointing position of the effective gesture does not move, the region to be interacted is determined in the paper region according to the pointing position, so as to transmit the text content of the region to be interacted to the electronic device.
[0031] Further, the determination of the region to be interacted in the paper region according to the pointing position, so as to transmit the text content of the region to be interacted to the electronic device, comprises:
[0032] The image region within the upper rectangle corresponding to the fingertip in the paper region is intercepted as the region to be interacted according to the pointing position, so as to call the text detection module to perform text detection on the region to be interacted.
[0033] If no text is detected in the region to be interacted, the region to be interacted is enlarged, and the text detection on the region to be interacted after the region is enlarged is repeated.
[0034] If no text is detected in the region to be interacted after the region to be interacted is enlarged for a preset number of times, the next image frame in the image frame sequence is subjected to gesture recognition.
[0035] If text is detected in the to-be-interacted region or text is detected after the to-be-interacted region is enlarged for a preset number of times, the text content of the to-be-interacted region is transmitted to the electronic device.
[0036] Further, the pointing change recognition on the valid gesture in the target image frame comprises:
[0037] The fingertip movement information of the valid gesture in a preset time is acquired according to the pointing position of the valid gesture;
[0038] The fingertip trajectory recognition is performed on the fingertip movement information by using a pre-trained trajectory recognition module, so as to determine the pointing change of the valid gesture according to the recognized fingertip movement trajectory and fingertip movement angle.
[0039] Further, the determination of the pointing change of the valid gesture according to the recognized fingertip movement trajectory and fingertip movement angle comprises:
[0040] The start point coordinate and end point coordinate of the valid gesture are acquired according to the recognized fingertip movement trajectory and fingertip movement angle;
[0041] The pointing change of the valid gesture is determined according to the start point coordinate and end point coordinate of the valid gesture;
[0042] Correspondingly, the determination of the to-be-interacted region in the paper region according to the pointing change, so as to transmit the text content of the to-be-interacted region to the electronic device, comprises:
[0043] The movement trajectory and movement angle of the valid gesture are obtained by connecting the start point coordinate and end point coordinate of the valid gesture according to the pointing change;
[0044] The to-be-interacted region is determined in the paper region according to the movement trajectory and movement angle of the valid gesture, so as to call a text detection module to perform text detection on the to-be-interacted region;
[0045] If text is not detected in the to-be-interacted region, the gesture recognition is performed on a next image frame in the image frame sequence;
[0046] If text is detected in the to-be-interacted region, the text content of the to-be-interacted region is transmitted to the electronic device.
[0047] Further, if the start point coordinate and end point coordinate of the valid gesture are the same, the determination of the to-be-interacted region in the paper region according to the movement trajectory and movement angle of the valid gesture comprises:
[0048] If the movement track of the effective gesture is greater than the length change of the coordinate horizontal axis and the movement angle presents irregular change, and the track points corresponding to the movement track are uniformly distributed around the center, then the track points are connected in time sequence to form a circular border to determine the to-be-interacted region in the paper region.
[0049] If the start point coordinate and the end point coordinate of the effective gesture are different, then the determining the to-be-interacted region in the paper region according to the movement track and the movement angle of the effective gesture comprises:
[0050] If the movement track of the effective gesture is consistent with the length change of the coordinate horizontal axis, then the start point and the end point of the effective gesture are taken as a diagonal line of a rectangular border to determine the to-be-interacted region in the paper region.
[0051] If the movement track of the effective gesture is greater than the length change of the coordinate horizontal axis and the movement angle is kept at a preset angle, then the start point and the end point of the effective gesture are taken as a diagonal line of a rectangular border to determine the to-be-interacted region in the paper region.
[0052] According to a second aspect of the present application, an information processing device for paper screen interaction is provided, comprising:
[0053] An acquisition unit is configured to acquire a to-be-processed image frame sequence, the image frame sequence comprising at least one image frame output by a target object in a paper screen interaction process;
[0054] A first recognition unit is configured to, when detecting that an effective gesture exists in a target image frame, perform pointing position recognition on the effective gesture in the target image frame to obtain a pointing position of the effective gesture.
[0055] A second recognition unit is configured to, if the pointing position of the effective gesture moves, perform pointing change recognition on the effective gesture in the target image frame to obtain a pointing change of the effective gesture.
[0056] A first determination unit is configured to determine a to-be-interacted region in a paper region according to the pointing change, to transmit text content of the to-be-interacted region to an electronic device.
[0057] Further, the acquisition unit is specifically configured to: utilize a mirror imaging principle to receive an image preview stream output by an image acquisition device in a paper screen interaction process, the image preview stream being a mirror-inverted image frame sequence collected by the target object using the image acquisition device in the paper screen interaction process; perform mirror rotation and region cropping on each image frame in the image preview stream through a pre-trained image processing module, to form a to-be-processed image frame sequence from the image frames obtained through the mirror rotation and the region cropping.
[0058] Further, the apparatus further comprises: a third identification unit, configured to, after the image frame sequence to be processed is acquired, perform gesture identification on each image frame in the image frame sequence by a pre-trained gesture identification module to determine whether a gesture object exists in a current image frame; a detection unit, configured to, if the gesture object exists in the current image frame, detect whether a valid gesture exists in the current image frame according to the gesture object; and the third identification unit is further configured to, if the gesture object does not exist in the current image frame, perform gesture identification on a next image frame in the image frame sequence.
[0059] Further, the detection unit is specifically configured to: perform fingertip identification on the gesture object in the current image frame by a pre-trained fingertip identification module to determine whether fingertip information exists in the gesture object; if the fingertip information meets a set identification condition, it is determined that the valid gesture exists in the current image frame; and if the fingertip information does not meet the set identification condition, gesture identification is performed on a next image frame in the image frame sequence.
[0060] Further, the first identification unit is specifically configured to: intercept continuous image frames within a preset time after the target image frame in the image frame sequence; and perform fingertip position identification on the valid gesture in the continuous image frames by a pre-trained position identification module to determine a pointing position of the valid gesture according to fingertip positions of the valid gesture within the preset time.
[0061] Further, the first identification unit is specifically further configured to: if the fingertip positions of the valid gesture within the preset time move, it is determined that the pointing position of the valid gesture moves; otherwise, it is determined that the pointing position of the valid gesture does not move.
[0062] Correspondingly, the apparatus further comprises: a second determination unit, configured to, after the pointing position of the valid gesture is determined by performing the pointing position identification on the valid gesture in the target image frame when it is detected that the valid gesture exists in the target image frame, if the pointing position of the valid gesture does not move, determine a to-be-interacted region in the paper region according to the pointing position to transmit text content in the to-be-interacted region to an electronic device.
[0063] Further, the second determining unit is specifically configured to: intercept an image region in a top rectangle corresponding to a fingertip in the paper region as a to-be-interacted region according to the pointing position, so as to call a text detection module to perform text detection on the to-be-interacted region; if no text is detected in the to-be-interacted region, performing region expansion on the to-be-interacted region, and then repeatedly performing text detection on the to-be-interacted region after the region expansion; if no text is detected after the to-be-interacted region is expanded for a preset number of times, performing gesture recognition on a next image frame in the image frame sequence; if text is detected in the to-be-interacted region or the text is detected after the to-be-interacted region is expanded for the preset number of times, transmitting text content of the to-be-interacted region to an electronic device.
[0064] Further, the second identifying unit is specifically configured to: acquire fingertip movement information of the effective gesture within a preset time according to the pointing position of the effective gesture; perform fingertip trajectory identification on the fingertip movement information within a preset time by using a pre-trained trajectory identification module, so as to determine a pointing change of the effective gesture according to the identified fingertip movement trajectory and fingertip movement angle.
[0065] Further, the second identifying unit is specifically configured to: acquire a start point coordinate and an end point coordinate of the effective gesture according to the identified fingertip movement trajectory and fingertip movement angle; determine a pointing change of the effective gesture according to the start point coordinate and the end point coordinate of the effective gesture.
[0066] Correspondingly, the first determining unit is specifically configured to: connect the start point coordinate and the end point coordinate of the effective gesture according to the pointing change, so as to obtain a movement trajectory and a movement angle of the effective gesture; determine a to-be-interacted region in the paper region according to the movement trajectory and the movement angle of the effective gesture, so as to call a text detection module to perform text detection on the to-be-interacted region; if no text is detected in the to-be-interacted region, perform gesture recognition on a next image frame in the image frame sequence; if text is detected in the to-be-interacted region, transmit text content of the to-be-interacted region to an electronic device.
[0067] Further, if the start point coordinate and the end point coordinate of the effective gesture are the same, the first determining unit is further configured to: if the movement trajectory of the effective gesture is greater than the length change of the coordinate horizontal axis and the movement angle presents irregular change, and the trajectory points corresponding to the movement trajectory are uniformly distributed around the center, then connect the trajectory points in time sequence to form a circular border to determine the to-be-interacted region in the paper region; if the start point coordinate and the end point coordinate of the effective gesture are different, the first determining unit is further configured to: if the movement trajectory of the effective gesture is consistent with the length change of the coordinate horizontal axis, then take the start point and the end point of the effective gesture as a rectangular border to determine the to-be-interacted region in the paper region; if the movement trajectory of the effective gesture is greater than the length change of the coordinate horizontal axis and the movement angle is kept at a preset angle, then take the start point and the end point of the effective gesture as the diagonal lines of a region to determine the to-be-interacted region in the paper region.
[0068] According to a third aspect of the present application, a computer device is provided, comprising a memory and a processor, the memory stores a computer program, and the processor implements the steps of the method of the first aspect when executing the computer program.
[0069] According to a fourth aspect of the present application, a readable storage medium is provided, which stores a computer program, and the computer program implements the steps of the method of the first aspect when executed by a processor.
[0070] By means of the above technical solution, compared with the prior art which relies on manual operation to realize paper-screen interaction information processing, the paper-screen interaction information processing method, device, storage medium and computer device provided by the present application can obtain an image frame sequence to be processed, the image frame sequence comprising at least one image frame output by a target object in a paper-screen interaction process, when detecting that an effective gesture exists in a target image frame, perform pointing position recognition on the effective gesture in the target image frame to obtain the pointing position of the effective gesture, if the pointing position of the effective gesture moves, perform pointing change recognition on the effective gesture in the target image frame to obtain the pointing change of the effective gesture, and determine a to-be-interacted region in a paper region according to the pointing change to transmit the text content of the to-be-interacted region to an electronic device. The entire process is based on the detection of the existence of an effective gesture in a target image frame, and the paper-screen interaction information processing is realized through a simple gesture, which increases the operation flexibility of the user, and through the pointing position recognition and the pointing change recognition of the effective gesture, the user can define a to-be-interacted region with higher precision in the paper region through the effective gesture, which improves the text recognition accuracy and increases the operation experience of the user in the paper-screen interaction process.
[0071] The above description is only a summary of the technical solutions of the present application. In order to enable one of ordinary skill in the art to better understand and thus implement the technical solutions of the present application, the following will give a specific implementation manner of the present application in accordance with the contents of the description, and in order to enable the above and other purposes, features and advantages of the present application to be more apparent, the following will give a specific implementation manner of the present application. BRIEF DESCRIPTION OF DRAWINGS
[0072] The accompanying drawings, which are included to provide a further understanding of the present application, form a part of the present application and are configured to explain the present application, and do not constitute improper limitations on the present application. In the drawings:
[0073] Fig. 1 is a flow chart of a paper screen interactive information processing method provided by an embodiment of the present application;
[0074] Fig. 2 is a flow chart of another paper screen interactive information processing method provided by an embodiment of the present application;
[0075] Fig. 3 is a flow chart of a specific implementation manner of step 202 in Fig. 2;
[0076] Fig. 4 is a flow chart of a specific implementation manner of step 103 in Fig. 1;
[0077] Fig. 5 is a flow chart of a specific implementation manner of step 104 in Fig. 1;
[0078] Fig. 6a is a schematic diagram of a gesture interactive scene of pointing to a horizontal line provided by an embodiment of the present application;
[0079] Fig. 6b is a schematic diagram of a gesture interactive scene of pointing to a slider provided by an embodiment of the present application;
[0080] Fig. 6c is a schematic diagram of a gesture interactive scene of pointing to a circle provided by an embodiment of the present application;
[0081] Fig. 7 is a flow chart of another paper screen interactive information processing method provided by an embodiment of the present application;
[0082] Fig. 8 is a schematic diagram of a structure of a paper screen interactive information processing device provided by an embodiment of the present application;
[0083] Fig. 9 is a schematic diagram of a device structure of a computer device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0084] The content of the present application will now be discussed with reference to several exemplary embodiments. It should be understood that the discussion of these embodiments is only for the purpose of enabling one of ordinary skill in the art to better understand and thus implement the content of the present application, and is not intended to imply any limitation on the scope of the present application.
[0085] As used herein, the term "includes" and its variants are to be read as open-ended terms that mean "includes, but is not limited to." The term "based on" is to be construed as "based at least in part on." The terms "one embodiment" and "an embodiment" are to be read as "at least one embodiment." The term "another embodiment" is to be read as "at least one other embodiment."
[0086] With the development of intelligent hardware and optical character recognition technology, the related technology in the information processing process of paper screen interaction has derived intelligent hardware based on the assistance of a mirror. Using the principle of optical imaging reflection, the mirror collects the paper content on the plane, and the camera collects the mirror image in real time to complete the linkage and collection of multiple learning scenarios. However, this method has the following shortcomings in data collection and user experience: in the information collection process, the paper screen interaction information of most products only supports the collection of the entire screen desktop, and cannot locate and identify the required content; in the information recognition process, the optical character recognition technology has high requirements for the external environment, and in actual application, the accuracy and robustness of the optical character recognition technology need to be improved due to factors such as light, paper texture, desktop cleanliness, and handwriting clarity, which will affect the user experience.
[0087] For the above-mentioned few products that use object recognition technology to complete accurate positioning of information, for example, a single finger point is used to recognize the text near the fingertip, but it is difficult to accurately realize information recognition and information processing for complex gesture scenarios. This limitation requires the user to provide complex interaction operations, especially in paper screen interaction scenarios such as labeling, selecting multiple text paragraphs, or multi-task operations.
[0088] To solve this problem, the embodiment provides a paper screen interaction information processing method, as shown in FIG. 1, which includes the following steps:
[0089] 101, obtaining a sequence of image frames to be processed.
[0090] In this embodiment, the sequence of image frames to be processed is a sequence of multiple frames of images formed by a paper screen interaction image output by an image acquisition device in real time. The paper screen interaction image is equivalent to an image obtained by an electronic device through an image acquisition device to assist in collecting the image of a target object in a paper screen interaction process. The image acquisition device can be a camera, a video camera, or other devices with image acquisition functions. The image acquisition device can be embedded into the electronic device as a functional module of the electronic device, or it can be a functional module connected externally to the electronic device.
[0091] Among them, the sequence of image frames includes at least one image frame output by the target object in the paper screen interaction process. Considering that the paper screen interaction collects mirror images, the image frames are usually preprocessed.
[0092] Specifically in the paper screen interaction process, considering the continuity of the image frames, the image acquisition device can transmit the acquired original images to the electronic device in the form of a preview stream in real time. For each original image in the preview stream, the following method can be used for preprocessing: first, mirror-rotating the original image to obtain a mirror-rotated original image. Due to the acquisition angle, the visible area is usually trapezoidal. Then, performing rectangular cropping on the mirror-rotated original image to obtain a preprocessed image frame.
[0093] 102. When it is detected that there is a valid gesture in the target image frame, performing pointing position recognition on the valid gesture in the target image frame to obtain a pointing position of the valid gesture.
[0094] Considering that each image frame in the image frame sequence is acquired around the paper screen interaction process, the corresponding image frame can display a paper corresponding text area and a target object operating on the text area. For the case where the target object does not operate on the text area, the image frame displays an image of the paper corresponding text area. For the case where the target object operates on the text area, the image displays an image of the target object operating on the paper corresponding text area by a gesture.
[0095] The valid gesture is an operating gesture of the target object on the paper corresponding text area. The number of fingertips and the pointing direction of the fingertips can be determined on the basis of detecting the gesture in the image frame. In general, at least one fingertip pointing to the paper corresponding text area can be determined as detecting a valid gesture in the image frame, for example, there is a gesture of the target object in the image, the finger position of the target object is in the paper corresponding text area, and there is a pointing gesture. It is determined that a valid gesture is detected in the target image frame.
[0096] In this step, the target object is usually at least one operating user of the paper screen interaction. For a paper screen interaction scene of one operating user, the operating gesture of the operating user is taken as a valid gesture. For a paper screen interaction scene of multiple operating users, the operating gestures of the multiple operating users are all taken as valid gestures. The operating gesture of a target operating user can also be selected as a valid gesture. If there are multiple valid gestures, the valid gesture that is shielded can be ignored.
[0097] It should be noted that the paper area includes different target objects, such as a character object, an image object, or a blank object. Accordingly, the pointing position corresponding to the valid gesture in the target image frame can correspond to different target objects in the paper area. It can be understood that the pointing position of the valid gesture can be used to select a corresponding position area as a to-be-interacted area to identify the target object in the to-be-interacted area.
[0098] In an implementable manner, the fingertip position in the target image frame can be recognized by using an existing image processing method, and the coordinate information of the recognized fingertip position can be used as the pointing position of the effective gesture. It can be understood that, in general, the pointing position of the effective gesture can be fixed or changed during the paper screen interaction. After the target image frame in which the pointing position is recognized, continuous image frames after the target image frame are continuously acquired. If the fingertip position does not displace within a preset time, it is indicated that the effective gesture is a gesture of pointing behavior, and the coordinate information of the recognized fingertip position can be used as the pointing position of the effective gesture. If the fingertip position displaces within the preset time, it is indicated that the effective gesture is not a gesture of pointing behavior, and the movement information of the recognized fingertip position can be recorded as the pointing position of the effective gesture.
[0099] 103. If the pointing position of the effective gesture moves, pointing change recognition is performed on the effective gesture in the target image frame to obtain the pointing change of the effective gesture.
[0100] In the embodiment, if the pointing position of the effective gesture moves, it is indicated that the effective gesture has a movement characteristic, and the pointing change recognition can be performed on the effective gesture in the target image frame. Specifically, the coordinate points corresponding to each pointing position can be connected according to the pointing positions of the effective gesture in the continuous target image frames to obtain the pointing change of the effective gesture. Here, the pointing change of the effective gesture can be understood as a trajectory formed by the movement of the fingertip of the finger of the target object in the paper region, which can be a pointing horizontal change, a pointing slider change, a pointing circle change, and the like. In order to facilitate the paper screen information transmission, different gesture interaction types can be defined according to the movement characteristic of the effective gesture, and different gesture interaction types have different pointing changes.
[0101] Specifically, for the gesture interaction type of the pointing horizontal change, the sliding characteristic of the effective gesture is the horizontal sliding of the pointing position from one point to another point. The left end point of the pointing position sliding is used as the starting point, the right end point of the pointing position sliding is used as the ending point, and the horizontal trajectory composed of the starting point and the ending point is used as the pointing change of the effective gesture. For the gesture interaction type of the pointing slider change, the sliding characteristic of the effective gesture is the oblique sliding of the pointing position from one point to another point. At this time, the oblique trajectory composed of the starting point and the ending point is used as the pointing change of the effective gesture. For the gesture interaction type of the pointing circle change, the sliding characteristic of the effective gesture is the arc sliding of the pointing position from one point to another point. At this time, the arc trajectory composed of the starting point and the ending point is used as the pointing change of the effective gesture.
[0102] 104. A to-be-interacted region is determined in the paper region according to the pointing change, so as to transmit the text content of the to-be-interacted region to an electronic device.
[0103] In the embodiment, for different gesture interaction types, the pointing change forms a corresponding track line or track area in the paper region. If the pointing change forms a track line in the paper region, it indicates that the gesture interaction type is a pointing horizontal stroke change or a pointing slider change. The corresponding pointing change is presented as a track line in the paper region. For the pointing horizontal stroke change, the text region above the two end points of the track line can be selected as the to-be-interacted region. For the pointing slider change, the text region formed by the track line as a diagonal line can be selected as the to-be-interacted region. If the pointing change forms a track area in the paper region, it indicates that the gesture interaction type is a pointing circle drawing change. The corresponding pointing change is presented as a track area in the paper region. It should be noted that the track area can be a fully closed area, indicating that the starting point and the ending point of the pointing change correspond to overlapping pointing positions. In this case, the track area can be selected as the to-be-interacted region. The track area can also be a semi-closed area, indicating that the starting point and the ending point of the pointing change do not correspond to overlapping pointing positions, and the distance between the two pointing positions is less than a set distance. In general, the starting point and the ending point of the pointing change are close to each other, and the track points are uniformly distributed from the center point. In this case, the track area is selected as the to-be-interacted region.
[0104] In the embodiment, the pointing position of the effective gesture is recognized, and the pointing change of the effective gesture is recognized based on the movement of the pointing position. The text content near the fingertip of the target object can be accurately recognized. The pointing position and the pointing change are used to enrich the text region of the paper screen interaction. The text region can be a text line or a text block. Single characters, single words, or entire text content can be efficiently and accurately recognized. The information recognition speed and accuracy of the paper screen interaction are greatly improved.
[0105] Further, the text recognition technology can be used to recognize the text in the to-be-interacted region in the paper region to obtain the text content of the to-be-interacted region, and transmit the text content to the electronic device. It can be understood that, in order to improve the accuracy of text recognition in the to-be-interacted region, the text recognition technology can be used to pre-recognize the text in the paper region to obtain the text content of the paper region. Then, the text content of the to-be-interacted region is corrected in combination with the text content of the paper region, to ensure the information accuracy of the paper screen interaction.
[0106] The paper screen interaction process in the embodiment is an information interaction between a paper material and an electronic device. The process mainly improves the efficiency of the paper screen interaction through the pointing position recognition and the pointing change recognition of the effective gesture. The text content in the paper material that is accurately positioned can be transmitted to the electronic device in a more efficient manner to improve the user's interaction experience. The information processing process of the paper screen interaction can be applied in the fields of education, intelligent office, remote conference, virtual reality, and augmented reality.
[0107] Specifically in the field of education, the embodiments of the present application can be applied to an intelligent teaching auxiliary system. By recognizing the pointing position and pointing change of the gestures of students in the teaching material / teaching aid area, the intelligent teaching device can provide instant explanation, data supplement or interactive practice of the relevant content, thereby realizing personalized tutoring. In addition, the technology can also be configured for automatic correction, development of intelligent interactive teaching materials, etc.
[0108] Specifically in the field of intelligent office, the embodiments of the present application can be applied to file management and conference display. By recognizing the pointing position and pointing change of the gestures of office workers in the specific area of the file, the detailed information or edited content of the corresponding electronic version of the specific area of the file can be quickly read and retrieved, thereby improving work efficiency. At the same time in the conference or speech, the speaker can inform the presentation operation through gestures, such as page turning or highlighting, etc., so that the information transmission in the paper screen interaction process is more smooth and accurate.
[0109] Specifically in the field of remote conference, the embodiments of the present application can capture the gestures of the participants through remote cameras. By recognizing the pointing position and pointing change of the gestures of the participants in the paper area, the text content required by the participants can be shared to other participants, so that the text content can be displayed on the screens of multiple participants in real time, thereby improving the efficiency of remote communication and making remote collaboration more intuitive and personalized.
[0110] Specifically in the field of virtual reality and augmented implementation, the user can interact with virtual objects through gestures in the embodiments of the present application. For example, in the augmented reality textbook, the user points to a formula or table, and the system can display the related three-dimensional model or animation on the screen, providing a more vivid teaching experience. In addition, designers or engineers can interact with paper design drawings through gestures, and the system can convert the content on the design drawings into three-dimensional models and display and modify them in the virtual reality environment.
[0111] The paper screen interaction information processing method provided in the embodiments of the present application, compared with the paper screen interaction information processing method realized by manual operation in the prior art, acquires an image frame sequence to be processed, the image frame sequence includes at least one image frame output by a target object in a paper screen interaction process, detects whether there is an effective gesture in the target image frame, performs pointing position recognition on the effective gesture in the target image frame to obtain a pointing position of the effective gesture, performs pointing change recognition on the effective gesture in the target image frame to obtain a pointing change of the effective gesture if the pointing position of the effective gesture moves, and determines a region to be interacted in a paper region according to the pointing change to transmit text content in the region to be interacted to an electronic device. The whole process realizes paper screen interaction information processing through a simple gesture on the basis of detecting that there is an effective gesture in the target image frame, increases the operation flexibility of a user, and enables the user to define a region to be interacted in the paper region with higher precision through the effective gesture by performing pointing position recognition and pointing change recognition on the effective gesture, thereby improving the text recognition accuracy and increasing the operation experience of the user in the paper screen interaction process.
[0112] In an actual application scenario, considering the particularity of paper screen interaction, the image frame sequence to be processed is an original image collected in real time by an image collection module of a hardware device and an auxiliary mirror, the original image has a mirror image feature, and a visible region corresponding to the original image is not a standard shape, so that the original image needs to be preprocessed. Specifically, in the process of acquiring the image frame sequence to be processed, the mirror imaging principle can be used to receive an image preview stream output by an image collection device in a paper screen interaction process, the image preview stream is a mirror-inverted image frame sequence collected by the image collection device in the paper screen interaction process, and then a pre-trained image processing module is used to perform mirror rotation and region cropping on each image frame in the image preview stream, so as to form the image frame sequence to be processed.
[0113] FIG. 2 is a flowchart of another paper screen interaction information processing method provided in the embodiments of the present application. In the above embodiments, the effective gesture is not a simple gesture object, but an expected interaction action generated by the target object through the gesture object. Considering the recognition process of the effective gesture, after step 101, the method further includes the following steps as shown in FIG. 2.
[0114] 201. Perform gesture recognition on each image frame in the image frame sequence by using a pre-trained gesture recognition module to determine whether there is a gesture object in the current image frame.
[0115] 202. If there is a gesture object in the current image frame, detect whether there is an effective gesture in the current image frame according to the gesture object.
[0116] 203. If there is no gesture object in the current image frame, performing gesture recognition on a next image frame in the image frame sequence.
[0117] It can be understood that the paper screen interaction can make the image frame in the image frame sequence have a gesture object, which can be recognized from the image frame by a pre-trained gesture recognition module. The gesture object label is set in advance, the image frame sequence to be processed is input into the gesture recognition model, and the corresponding output is whether there is a gesture object in the current image frame and the position of the gesture object. If there are multiple gesture objects in the current image frame, the position of each gesture object is marked in the current image frame. For the case that there is no gesture object in the current image frame, gesture recognition is performed on a next image frame in the image frame sequence. For the case that there is a gesture object in the current image frame, it is further detected whether there is a valid gesture in the current image frame.
[0118] In some possible implementations, specifically, before performing gesture recognition on each image frame in the image frame sequence to determine whether there is a gesture object in the current image frame, the gesture recognition module can be trained first. The trained model is used as the gesture recognition module, a sample set of different hand images is obtained in advance, and then the sample set of hand images is extracted and classified in the training process to obtain hand image features. After each image frame in the image frame sequence is input into the gesture recognition module, the gesture recognition module can extract hand image features from the current image frame. If the hand image features are extracted from the current image frame, it indicates that there is a gesture object in the current image frame. If the hand image features are not extracted from the current image frame, it indicates that there is no gesture object in the current image frame.
[0119] Considering that the gesture object can present different features under different backgrounds, illumination conditions and angles, first, a deep convolutional network can be used to extract image features of the gesture object. Through the stacking of multiple convolutional layers and pooling layers, local features and global features of the image can be effectively extracted. Second, the idea of transfer learning can be introduced to fine-tune and transfer the pre-trained neural network on the basis of the recognition task of the gesture object. In this way, the performance of gesture object recognition can be improved by using existing large-scale data sets and model optimization. In addition, for the similarity between different gesture objects, time information can be introduced for recognition and classification.
[0120] Specifically, as shown in FIG. 3, the step 202 includes the following steps:
[0121] 301. Performing fingertip recognition on the gesture object in the current image frame by a pre-trained fingertip recognition module to determine whether there is fingertip information in the gesture object.
[0122] 302、if the fingertip information meets the set recognition condition, it is determined that the valid gesture exists in the current image frame.
[0123] 303、if the fingertip information does not meet the set recognition condition, gesture recognition is performed on the next image frame in the image frame sequence.
[0124] In some possible implementation manners, specifically, before fingertip recognition is performed on the gesture object in the current image frame, the model training process of the gesture recognition module can also be referred to, so that the gesture object in the current image frame can be subjected to feature extraction by the fingertip recognition module. If fingertip features are extracted from the gesture object, it is indicated that the fingertip information exists in the gesture object. If no fingertip features are extracted from the gesture object, it is indicated that the fingertip information does not exist in the gesture object.
[0125] The fingertip information herein can include but is not limited to a fingertip orientation, a finger type, a number of fingertips, and the like. The set recognition condition can be determined according to a paper screen interaction requirement. For example, the paper screen interaction requirement is that a single finger is oriented to a paper region, that is, the set recognition condition is that the fingertip is oriented to the paper region, the number of fingertips is 1, and the finger type is an index finger. For another example, the paper screen interaction requirement is that two fingers are oriented to the paper region, that is, the set recognition condition is that the fingertip is oriented to the paper region, the number of fingertips is 2, and the finger type is an index finger and a thumb.
[0126] In an actual application scenario, the pointing position of the valid gesture is not fixed. In order to ensure the accuracy of the pointing position recognition, specifically, as shown in FIG. 4, step 103 includes the following steps:
[0127] 401、after the target image frame is intercepted from the image frame sequence, continuous image frames in a preset time are intercepted.
[0128] 402、the fingertip position of the valid gesture in the continuous image frames is recognized by using a pre-trained position recognition module, so as to determine the pointing position of the valid gesture according to the fingertip positions of the valid gesture in the preset time.
[0129] In this embodiment, the fingertip position of the valid gesture in the continuous image frames in the preset time after the target image frame can be used as the pointing position of the valid gesture. For example, if the fingertip positions of the valid gesture in the continuous image frames in 0.25 seconds do not change, it is indicated that the fingertip position of the valid gesture does not change, and the fingertip position can be used as the pointing position of the valid gesture.
[0130] Specifically, if the fingertip positions of the valid gesture in the preset time move, it is determined that the pointing of the valid gesture moves; otherwise, it is determined that the pointing position of the valid gesture does not move.
[0131] Correspondingly, after the pointing position of the effective gesture is recognized in the target image frame in which the effective gesture is detected, and the pointing position of the effective gesture does not move, a region to be interacted is determined in the paper region according to the pointing position, so as to transmit the text content of the region to be interacted to the electronic device. For the case that the pointing position of the effective gesture does not move, it indicates that the target user has a pointing behavior in the paper region, which can be any target object in the paper region, such as a blank part, a text part, an image part, etc. Further, an image region in an upper rectangle corresponding to the fingertip in the paper region can be intercepted as the region to be interacted according to the pointing position, so as to call a text detection module to perform text detection on the region to be interacted. If no text is detected in the region to be interacted, the region to be interacted is enlarged, and the text detection on the region to be interacted after the enlargement is repeated. If no text is detected in the region to be interacted after the region to be interacted is enlarged for a preset number of times, gesture recognition is performed on a next image frame in the image frame sequence. If text is detected in the region to be interacted or the text is detected in the region to be interacted after the region to be interacted is enlarged for a preset number of times, the text content of the region to be interacted is transmitted to the electronic device.
[0132] For example, when the effective gesture is detected in the target image frame, and the effective gesture corresponds to a pointing behavior, a rectangular region above the fingertip is first intercepted as a region to be interacted, the size of the region is 100*50, a text detection module is called to perform text detection on the region to be interacted, if no text is detected in the region to be interacted, the region to be interacted is enlarged, the size of the region is 200*100, the text detection module is continuously called to perform text detection on the region to be interacted after the enlargement, the number of times of enlargement of the region to be interacted is limited to 3, if no text is detected after the region to be interacted is enlarged for 3 times, it indicates that the pointing behavior of the target object corresponds to a blank region in the paper region, gesture recognition is performed on a next image frame in the image frame sequence, if text is detected in the region to be interacted or in the region to be interacted after any expansion, the detected text content is transmitted to the electronic device. It can be understood that the region to be interacted is expanded by the automatic expansion technology, which is compatible with different text recognition ranges and improves the text recognition accuracy, the region to be interacted after the enlargement has higher compatibility, and the problems of finger offset and short distance of the target object are solved.
[0133] In an actual application scenario, the pointing change of the effective gesture is based on the movement of the pointing position. In order to ensure the accuracy of the pointing change recognition, specifically, as shown in FIG. 5, step 103 includes the following steps:
[0134] 501. Obtain fingertip movement information of the effective gesture in a preset time according to the pointing position of the effective gesture.
[0135] 502、using the pre-trained trajectory recognition module to perform fingertip trajectory recognition in the range of the fingertip movement information, to determine the pointing change of the effective gesture according to the recognized fingertip movement trajectory and fingertip movement angle.
[0136] It can be understood that if the pointing position of the effective gesture moves, it means that the effective gesture corresponds to a movement trajectory in the paper area. Specifically, the starting point coordinates and the end point coordinates of the effective gesture can be obtained according to the recognized fingertip movement trajectory and fingertip movement angle, and the pointing change of the effective gesture can be determined according to the starting point coordinates and the end point coordinates of the effective gesture.
[0137] Correspondingly, in the process of determining the to-be-interacted area in the paper area according to the pointing change, and transmitting the text content of the to-be-interacted area to the electronic device, the starting point coordinates and the end point coordinates of the effective gesture can be connected according to the pointing change to obtain the movement trajectory and the movement angle of the effective gesture, the to-be-interacted area can be determined in the paper area according to the movement trajectory and the movement angle of the effective gesture, the text detection module can be called to perform text detection on the to-be-interacted area, if no text is detected in the to-be-interacted area, the next image frame in the image frame sequence is subjected to gesture recognition, if text is detected in the to-be-interacted area, the text content of the to-be-interacted area is transmitted to the electronic device.
[0138] If the starting point coordinates and the end point coordinates of the effective gesture are the same, specifically in the process of determining the to-be-interacted area in the paper area according to the movement trajectory and the movement angle of the effective gesture, if the movement trajectory of the effective gesture is greater than the length change of the coordinate horizontal axis and the movement angle presents irregular change, and the trajectory points corresponding to the movement trajectory are uniformly distributed around the center, the trajectory points are connected in time sequence to form a circular frame to determine the to-be-interacted area in the paper area; if the starting point coordinates and the end point coordinates of the effective gesture are different, in the process of determining the to-be-interacted area in the paper area according to the movement trajectory and the movement angle of the effective gesture, if the movement trajectory of the effective gesture is consistent with the length change of the coordinate horizontal axis, the starting point and the end point of the effective gesture are taken as a rectangular frame to determine the to-be-interacted area in the paper area; if the movement trajectory of the effective gesture is greater than the length change of the coordinate horizontal axis and the movement angle is kept at a preset angle, the starting point and the end point of the effective gesture are taken as the diagonals of the area to determine the to-be-interacted area in the paper area.
[0139] As a first implementation scenario, when it is detected that there is an effective gesture in the target image frame, and the effective gesture corresponds to a pointing change, the fingertip movement trajectory and the movement angle are recorded and saved, and when the finger position remains unchanged, the fingertip starting position and the fingertip ending position corresponding to the pointing change are obtained, it is judged whether the fingertip starting position and the fingertip ending position overlap in a preset range, if not, it is indicated that the pointing change of the effective gesture forms a trajectory line, and the change length formed by the fingertip starting position and the fingertip ending position on the horizontal axis is further calculated, if the trajectory line is consistent with the change length formed on the horizontal axis, it is indicated that the trajectory line is generated by the effective gesture through the pointing horizontal change, and the gesture interaction scene schematic diagram is shown in FIG. 6a, at this time, the fingertip starting position and the fingertip ending position are taken as two edges of a to-be-interacted region to form a rectangular frame, a text region corresponding to the rectangular frame is intercepted, a text detection module is called to perform text detection on the to-be-interacted region, and the recognized text content is transmitted to an electronic device.
[0140] As a second implementation scenario, when it is detected that there is an effective gesture in the target image frame, and the effective gesture corresponds to a pointing change, the fingertip movement trajectory and the movement angle are recorded and saved, it is judged whether the fingertip starting position and the fingertip ending position overlap in a preset range, if not, it is indicated that the pointing change of the effective gesture forms a trajectory line, and the change length formed by the fingertip starting position and the fingertip ending position on the horizontal axis is further calculated, if the trajectory line is inconsistent with the change length formed on the horizontal axis, and the trajectory line is greater than the change length formed on the horizontal axis, it is indicated that the trajectory line is generated by the effective gesture through the pointing slider change, and the gesture interaction scene schematic diagram is shown in FIG. 6b, at this time, the fingertip starting position and the fingertip ending position are taken as diagonal lines of a to-be-interacted region to form a rectangular frame, a text region corresponding to the rectangular frame is intercepted, a text detection module is called to perform text detection on the to-be-interacted region, and the recognized text content is transmitted to an electronic device.
[0141] As a third implementation scenario, when it is detected that there is an effective gesture in the target image frame, and the effective gesture corresponds to a pointing change, the fingertip movement trajectory and the movement angle are recorded and saved, it is judged whether the fingertip starting position and the fingertip ending position overlap in a preset range, if they overlap, it is indicated that the pointing change of the effective gesture is a trajectory circle, and the change length of the fingertip starting position and the fingertip ending position on the horizontal axis is further calculated, if the trajectory line is greater than the change length formed on the horizontal axis, and the trajectory angle changes obviously, the trajectory points are uniformly distributed around the center, it is indicated that the trajectory circle is generated by the effective gesture through the pointing circle-drawing change, and the gesture interaction scene schematic diagram is shown in FIG. 6c, at this time, the trajectory points form a circular frame in time sequence, a text region corresponding to the circular frame is intercepted, a text detection module is called to perform text detection on the to-be-interacted region, and the recognized text content is transmitted to an electronic device.
[0142] In the education application scenario, the information processing flowchart of the specific paper screen interaction is shown in FIG. 7. In FIG. 7, when the target object starts the information processing flow of the paper screen interaction through the hardware device, the preview stream corresponding to the data acquisition device is acquired in real time, the current image frame in the preview stream is read, it is detected whether there is a gesture object in the current image frame, if no gesture object is detected, the next image frame is detected for the gesture object, if the gesture object is detected, it is detected whether there is a fingertip in the current image frame, if there is a fingertip, it is judged whether the number of fingertips is 1 according to the fingertip information, if the number of fingertips is greater than or equal to 2, it is indicated that the current gesture object is an invalid gesture, the next image frame is detected for the gesture object, if the number of fingertips is 1, it is judged whether there is displacement of the fingertip position in 0.25s according to the current image frame, if there is no displacement of the fingertip position, a rectangular region near the fingertip is positioned as a to-be-interacted region, the to-be-interacted region is cropped, if there is displacement of the fingertip position, the fingertip position of the current image frame is recorded, the track point and the track angle are preserved, it is judged whether the finger moves out, if the finger does not move out, the fingertip position of the image frame is continuously recorded to update the track point and the track angle, if the finger moves out, the start point coordinates and the end point coordinates of the fingertip are compared, the track length and the angle feature are extracted, it is judged whether the track length is close to the coordinate horizontal axis change length, if the track length is close to the coordinate horizontal axis change length, it is determined that the interaction gesture is a horizontal swipe, two edges corresponding to the fingertip start point and the fingertip end point are selected to form a rectangular region as a to-be-interacted region, if the track length is inconsistent with the coordinate horizontal axis change length, it is judged whether the track angle is close to 45 degrees, if the track angle is close to 45 degrees, it is determined that the interaction gesture is a sliding block, the diagonal line of the rectangle is selected as the connection between the fingertip start point and the fingertip end point to form a rectangular region as a to-be-interacted region, if the track angle is not close to 45 degrees, it is judged whether the track points are uniformly distributed around the center, if not, it is indicated that the current fingertip recognition is invalid, the next image frame is detected for the gesture object, if yes, it is determined that the interaction gesture is a pointing circle, the figure formed by the surrounding track is taken as a to-be-interacted region, then the to-be-interacted region is cropped, and the to-be-interacted region after cropping is further subjected to text recognition through a text recognition technology to determine whether there is text in the to-be-interacted region, if there is no text in the to-be-interacted region, the next image frame is detected for the gesture object, if there is text in the to-be-interacted region, the text content is subjected to error correction processing through natural language processing, and the processed text content is transmitted to an electronic device.
[0143] Further, as a specific implementation of the methods of FIGS. 1-7, the embodiment of the present application provides an information processing device for paper screen interaction, as shown in FIG. 8, which comprises an acquisition unit 61, a first recognition unit 62, a second recognition unit 63, and a first determination unit 64.
[0144] The acquisition unit 61 is configured to acquire an image frame sequence to be processed, the image frame sequence comprising at least one image frame output by a target object in a paper screen interaction process;
[0145] The first identification unit 62 is configured to, when detecting that there is a valid gesture in a target image frame, perform pointing position identification on the valid gesture in the target image frame to obtain a pointing position of the valid gesture.
[0146] The second identification unit 63 is configured to, if the pointing position of the valid gesture moves, perform pointing change identification on the valid gesture in the target image frame to obtain a pointing change of the valid gesture.
[0147] The first determination unit 64 is configured to determine a region to be interacted in a paper region according to the pointing change, so as to transmit text content of the region to be interacted to an electronic device.
[0148] The paper screen interaction information processing device provided by the embodiment of the present application, compared with the paper screen interaction information processing mode realized by manual operation in the prior art, acquires an image frame sequence to be processed, the image frame sequence comprising at least one image frame output by a target object in a paper screen interaction process, when detecting that there is a valid gesture in a target image frame, performs pointing position identification on the valid gesture in the target image frame to obtain a pointing position of the valid gesture, if the pointing position of the valid gesture moves, performs pointing change identification on the valid gesture in the target image frame to obtain a pointing change of the valid gesture, and determines a region to be interacted in a paper region according to the pointing change, so as to transmit text content of the region to be interacted to an electronic device. The whole process is realized by a simple gesture on the basis of detecting that there is a valid gesture in a target image frame, increases the operation flexibility of a user, and through the pointing position identification and the pointing change identification of the valid gesture, can enable the user to define a region to be interacted in a paper region with higher precision through the valid gesture, improves the text recognition accuracy, and increases the operation experience of the user in the paper screen interaction process.
[0149] In an actual application scenario, the acquisition unit is specifically configured to: utilize a mirror imaging principle to receive an image preview stream output by an image acquisition device in a paper screen interaction process, the image preview stream being a mirror-inverted image frame sequence collected by the target object using the image acquisition device in the paper screen interaction process; and perform mirror rotation and region cropping on each image frame in the image preview stream through a pre-trained image processing module, so as to form the image frame sequence to be processed from the image frame obtained through the mirror rotation and the region cropping.
[0150] In an actual application scenario, the apparatus further includes a third identification unit configured to, after the image frame sequence to be processed is acquired, perform gesture identification on each image frame in the image frame sequence by using a pre-trained gesture identification module to determine whether a gesture object exists in a current image frame; a detection unit configured to, if the gesture object exists in the current image frame, detect whether a valid gesture exists in the current image frame according to the gesture object; and the third identification unit is further configured to, if the gesture object does not exist in the current image frame, perform gesture identification on a next image frame in the image frame sequence.
[0151] In an actual application scenario, the detection unit is specifically configured to perform fingertip identification on the gesture object in the current image frame by using a pre-trained fingertip identification module to determine whether fingertip information exists in the gesture object; if the fingertip information meets a set identification condition, it is determined that a valid gesture exists in the current image frame; and if the fingertip information does not meet the set identification condition, gesture identification is performed on a next image frame in the image frame sequence.
[0152] In an actual application scenario, the first identification unit is specifically configured to, after the target image frame in the image frame sequence is intercepted, continuously capture image frames within a preset time; and perform fingertip position identification on the valid gesture in the continuous image frames by using a pre-trained position identification module to determine a pointing position of the valid gesture according to fingertip positions of the valid gesture within the preset time.
[0153] In an actual application scenario, the first identification unit is specifically further configured to, if the fingertip positions of the valid gesture within the preset time move, it is determined that the pointing position of the valid gesture moves; otherwise, it is determined that the pointing position of the valid gesture does not move.
[0154] Correspondingly, the apparatus further includes a second determination unit configured to, after the pointing position of the valid gesture is determined by performing pointing position identification on the valid gesture in the target image frame when it is detected that the valid gesture exists in the target image frame, if the pointing position of the valid gesture does not move, determine a region to be interacted in the paper region according to the pointing position, and transmit text content in the region to be interacted to an electronic device.
[0155] In an actual application scenario, the second determining unit is specifically configured to: intercept an image region in a top rectangle corresponding to a fingertip in the paper region according to the pointing position as a to-be-interacted region to call a text detection module to perform text detection on the to-be-interacted region; if no text is detected in the to-be-interacted region, performing region expansion on the to-be-interacted region and then repeatedly performing text detection on the to-be-interacted region after the region expansion; if no text is detected after the to-be-interacted region is expanded for a preset number of times, performing gesture recognition on a next image frame in the image frame sequence; if text is detected in the to-be-interacted region or text is detected after the to-be-interacted region is expanded for a preset number of times, transmitting text content of the to-be-interacted region to an electronic device.
[0156] In an actual application scenario, the second identifying unit is specifically configured to: acquire fingertip movement information of the effective gesture within a preset time according to the pointing position of the effective gesture; perform fingertip trajectory identification in a range of the fingertip movement information by using a pre-trained trajectory identification module to determine pointing changes of the effective gesture according to an identified fingertip movement trajectory and a fingertip movement angle.
[0157] In an actual application scenario, the second identifying unit is specifically configured to: acquire a start point coordinate and an end point coordinate of the effective gesture according to the identified fingertip movement trajectory and the fingertip movement angle; determine pointing changes of the effective gesture according to the start point coordinate and the end point coordinate of the effective gesture.
[0158] Correspondingly, the first determining unit is specifically configured to: connect the start point coordinate and the end point coordinate of the effective gesture according to the pointing changes to obtain a movement trajectory and a movement angle of the effective gesture; determine a to-be-interacted region in the paper region according to the movement trajectory and the movement angle of the effective gesture to call a text detection module to perform text detection on the to-be-interacted region; if no text is detected in the to-be-interacted region, perform gesture recognition on a next image frame in the image frame sequence; if text is detected in the to-be-interacted region, transmit text content of the to-be-interacted region to an electronic device.
[0159] In an actual application scenario, if the start point coordinate and the end point coordinate of the effective gesture are the same, the first determining unit is further configured to: if the movement trajectory of the effective gesture is greater than the length change of the coordinate horizontal axis and the movement angle presents irregular change, and the trajectory points corresponding to the movement trajectory are uniformly distributed around the center, then connect the trajectory points in time sequence to form a circular border to determine the to-be-interacted region in the paper region; if the start point coordinate and the end point coordinate of the effective gesture are different, the first determining unit is further configured to: if the movement trajectory of the effective gesture is consistent with the length change of the coordinate horizontal axis, then take the start point and the end point of the effective gesture as a rectangular border to determine the to-be-interacted region in the paper region; if the movement trajectory of the effective gesture is greater than the length change of the coordinate horizontal axis and the movement angle is kept at a preset angle, then take the start point and the end point of the effective gesture as the diagonal lines of a region to determine the to-be-interacted region in the paper region.
[0160] It should be noted that other corresponding descriptions of the functions of the paper screen interaction information processing device provided in the embodiment can refer to the corresponding descriptions in FIGS. 1-7, which will not be described here.
[0161] Based on the above method as shown in FIGS. 1-7, correspondingly, the embodiment of the present application further provides a storage medium having a computer program stored thereon, which is executed by a processor to implement the above paper screen interaction information processing method as shown in FIGS. 1-7.
[0162] Based on such understanding, the technical solution of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a U disk, a mobile hard disk, etc.), and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the method described in each embodiment scenario of the present application.
[0163] Based on the above method as shown in FIGS. 1-7, and the virtual device embodiment as shown in FIG. 8, in order to achieve the above purpose, the embodiment of the present application further provides a paper screen interaction information processing entity device, which can be a computer, a smart phone, a tablet computer, a smart watch, a server, or a network device, etc. The entity device includes a storage medium and a processor; the storage medium is configured to store a computer program; the processor is configured to execute the computer program to implement the above paper screen interaction information processing method as shown in FIGS. 1-7.
[0164] Optionally, the entity device can further include a user interface, a network interface, a camera, a radio frequency (RF) circuit, a sensor, an audio circuit, a WI-FI module, and the like. The user interface can include a display, an input unit such as a keyboard, and the like. Optionally, the user interface can further include a USB interface, a card reader interface, and the like. The network interface can optionally include a standard wired interface, a wireless interface (such as a WI-FI interface), and the like.
[0165] In the example embodiment, referring to FIG. 9, the entity device includes a communication bus, a processor, a memory, and a communication interface, and can further include an input / output interface and a display device, wherein the respective functional units can communicate with each other through the bus. The memory stores a computer program, and the processor is configured to execute the program stored in the memory to perform the paper-screen interaction information processing method in the above embodiment.
[0166] Those skilled in the art can understand that the structure of the entity device for paper-screen interaction information processing provided in the embodiment does not constitute a limitation on the entity device, and can include more or fewer components, or combine certain components, or different component arrangements.
[0167] The storage medium can further include an operating system and a network communication module. The operating system is a program for managing hardware and software resources of the entity device for paper-screen interaction information processing, and supports the running of the information processing program and other software and / or programs. The network communication module is configured to realize communication between components in the storage medium, and communication with other hardware and software in the information processing entity device.
[0168] Through the above description of the embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software and necessary general hardware platforms, or by hardware. By applying the technical solutions of the present application, compared with the current existing mode, the present application realizes paper-screen interaction information processing through simple gestures on the basis of detecting that the target image frame has an effective gesture, increases the operation flexibility of the user, and through pointing position recognition and pointing change recognition of the effective gesture, can enable the user to define a higher-precision to-be-interacted region in the paper region through the effective gesture, improve the text recognition accuracy, and increase the operation experience of the user in the paper-screen interaction process.
[0169] Those skilled in the art can understand that the modules or flows in the drawings are not necessarily required for implementing the present application. Those skilled in the art can understand that the modules in the devices in the implementation scenarios can be distributed in the devices in the implementation scenarios according to the description of the implementation scenarios, or can be changed to be located in one or more devices different from the implementation scenarios. The modules in the above implementation scenarios can be combined into one module, or can be further split into multiple sub-modules.
[0170] The above application numbers are only for description, and do not represent the advantages and disadvantages of the implementation scenarios. The above disclosure is only some specific implementation scenarios of the present application, but the present application is not limited thereto, and any variations that can be thought of by those skilled in the art should fall within the protection scope of the present application.
Claims
1. An information processing method of paper screen interaction, wherein, The method comprises the following steps: acquiring a sequence of image frames to be processed, the sequence of image frames comprising at least one image frame output by a target object in a paper screen interaction process; when detecting that there is a valid gesture in a target image frame, performing pointing position recognition on the valid gesture in the target image frame to obtain a pointing position of the valid gesture; if the pointing position of the valid gesture moves, performing pointing change recognition on the valid gesture in the target image frame to obtain a pointing change of the valid gesture; determining a region to be interacted in a paper region according to the pointing change, and transmitting text content of the region to be interacted to an electronic device.
2. The method of claim 1, wherein, The acquiring of the sequence of image frames to be processed comprises using a mirror imaging principle, receiving an image preview stream output by an image acquisition device in the paper screen interaction process, the image preview stream being a sequence of image frames collected by the target object using the image acquisition device in the paper screen interaction process, the image frames being mirror-inverted images; performing mirror rotation and region cropping on each image frame in the image preview stream by using a pre-trained image processing module, so as to form a sequence of image frames to be processed.
3. The method of claim 1, wherein, After the acquiring of the sequence of image frames to be processed, the method further comprises performing gesture recognition on each image frame in the sequence of image frames by using a pre-trained gesture recognition module, so as to determine whether there is a gesture object in a current image frame; if there is a gesture object in the current image frame, detecting whether there is a valid gesture in the current image frame according to the gesture object; if there is no gesture object in the current image frame, performing gesture recognition on a next image frame in the sequence of image frames.
4. The method of claim 3, wherein, The detecting whether there is a valid gesture in the current image frame according to the gesture object comprises performing fingertip recognition on the gesture object in the current image frame by using a pre-trained fingertip recognition module, so as to determine whether there is fingertip information in the gesture object; if the fingertip information meets a set recognition condition, determining that there is a valid gesture in the current image frame; if the fingertip information does not meet the set recognition condition, performing gesture recognition on a next image frame in the sequence of image frames.
5. The method of claim 1, wherein, The performing of the pointing position recognition on the valid gesture in the target image frame to obtain the pointing position of the valid gesture comprises intercepting continuous image frames within a preset time after the target image frame in the sequence of image frames; performing fingertip position recognition on the valid gesture in the continuous image frames by using a pre-trained position recognition module, so as to determine the pointing position of the valid gesture according to fingertip positions of the valid gesture within the preset time.
6. The method of claim 5, wherein, The determining of the pointing position of the valid gesture according to the fingertip positions of the valid gesture within the preset time comprises if the fingertip positions of the valid gesture within the preset time move, determining that the pointing position of the valid gesture moves; otherwise, determining that the pointing position of the valid gesture does not move; correspondingly, after the detecting that there is a valid gesture in the target image frame, the method further comprises If the pointing position of the effective gesture does not move, a region to be interacted with is determined in the paper region according to the pointing position, so as to transmit text content of the region to be interacted with to an electronic device.
7. The method of claim 6, wherein, The determining of the region to be interacted with in the paper region according to the pointing position to transmit text content of the region to be interacted with to an electronic device comprises: According to the pointing position, an image region in a rectangle above a fingertip in the paper region is intercepted as the region to be interacted with, so as to call a text detection module to perform text detection on the region to be interacted with; If no text is detected in the region to be interacted with, the region to be interacted with is enlarged, and then the text detection is repeated on the region to be interacted with after the enlargement; If no text is detected in the region to be interacted with after the region to be interacted with is enlarged for a preset number of times, gesture recognition is performed on a next image frame in the image frame sequence; If text is detected in the region to be interacted with or the text is detected in the region to be interacted with after the region to be interacted with is enlarged for a preset number of times, text content of the region to be interacted with is transmitted to an electronic device.
8. The method of claim 1, wherein, The pointing change recognition of the effective gesture in the target image frame comprises: Obtaining fingertip movement information of the effective gesture within a preset time according to the pointing position of the effective gesture; Using a pre-trained trajectory recognition module to perform fingertip trajectory recognition in the range of the fingertip movement information, so as to determine the pointing change of the effective gesture according to the recognized fingertip movement trajectory and fingertip movement angle.
9. The method of claim 8, wherein, The determination of the pointing change of the effective gesture according to the recognized fingertip movement trajectory and fingertip movement angle comprises: According to the recognized fingertip movement trajectory and fingertip movement angle, obtaining the start point coordinates and end point coordinates of the effective gesture; According to the start point coordinates and end point coordinates of the effective gesture, determining the pointing change of the effective gesture; Accordingly, the determining of the region to be interacted with in the paper region according to the pointing change to transmit text content of the region to be interacted with to an electronic device comprises: According to the start point coordinates and end point coordinates of the effective gesture connected by the pointing change, obtaining a movement trajectory and a movement angle of the effective gesture; According to the movement trajectory and the movement angle of the effective gesture, determining a region to be interacted with in the paper region to call a text detection module to perform text detection on the region to be interacted with; If no text is detected in the region to be interacted with, gesture recognition is performed on a next image frame in the image frame sequence; If text is detected in the region to be interacted with, text content of the region to be interacted with is transmitted to an electronic device.
10. The method of claim 9, wherein, If the start point coordinates and the end point coordinates of the effective gesture are the same, the determining of the region to be interacted with in the paper region according to the movement trajectory and the movement angle of the effective gesture comprises: If the movement trajectory of the effective gesture is greater than the length change of the coordinate horizontal axis and the movement angle presents irregular change, and the trajectory points corresponding to the movement trajectory are uniformly distributed around the center, a circular frame formed by connecting the trajectory points in time sequence is determined as the region to be interacted with in the paper region; If the start point coordinate and the end point coordinate of the effective gesture are different, the determining the to-be-interacted region in the paper region according to the moving track and the moving angle of the effective gesture comprises: If the moving track of the effective gesture is consistent with the length change of the coordinate horizontal axis, the start point and the end point of the effective gesture are determined as the rectangular frame in the paper region to determine the to-be-interacted region. If the moving track of the effective gesture is greater than the length change of the coordinate horizontal axis and the moving angle is kept at a preset angle, the start point and the end point of the effective gesture are determined as the diagonal line of the region in the paper region to determine the to-be-interacted region.
11. An information processing apparatus of paper screen interaction, wherein, Comprise: An acquisition unit configured to acquire a to-be-processed image frame sequence, the image frame sequence comprising at least one image frame output by a target object in a paper screen interaction process; A first identification unit configured to, when detecting that an effective gesture exists in a target image frame, perform pointing position identification on the effective gesture in the target image frame to obtain a pointing position of the effective gesture; A second identification unit configured to, if the pointing position of the effective gesture moves, perform pointing change identification on the effective gesture in the target image frame to obtain a pointing change of the effective gesture; A first determination unit configured to determine a to-be-interacted region in a paper region according to the pointing change, to transmit text content of the to-be-interacted region to an electronic device.
12. A computer device comprising a memory and a processor, the memory storing a computer program, wherein, The processor executes the computer program to realize the steps of the paper screen interaction information processing method in any one of claims 1 to 10.
13. A computer readable storage medium having stored thereon a computer program, wherein, The computer program is executed by the processor to realize the steps of the paper screen interaction information processing method in any one of claims 1 to 10.
Citation Information
Patent Citations
Character recognition method and device, electronic equipment and storage medium
CN112036315A
Information acquisition method and device, electronic equipment and storage medium
CN112817445A
Text content recognition method and device, computer equipment and storage medium
CN115131693A
Method for recognizing touch-to-read text, and electronic device
WO2022194180A1