Point reading system and device

By combining gesture recognition and text recognition technologies, the problems of high price, complex operation and limited content of existing educational software are solved, realizing low-cost and high-quality educational services suitable for the learning needs of children aged 3-6.

CN117315676BActive Publication Date: 2025-11-28UNIV OF SCI & TECH BEIJING
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310747391.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-25
Publication Date
2025-11-28
Estimated Expiration
2043-06-25

AI Technical Summary

Technical Problem

Existing online educational software is expensive, complex to operate, and has limited content, making it difficult to meet the learning needs of children aged 3-6.

Method used

This invention provides a point-and-read system that employs a gesture recognition module, a target detection module, and a text extraction module, combined with artificial intelligence technology, to achieve intelligent recognition of gestures and text. It supports camera and video recognition modes, reduces dependence on specific teaching materials, and increases content richness and ease of operation.

Benefits of technology

It reduces the cost of learning with the device, improves the quality and accessibility of education, enhances children's interest in learning, is highly applicable, can be used directly, and does not require the purchase of related teaching materials.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117315676B_ABST
    Figure CN117315676B_ABST
Patent Text Reader

Abstract

The application provides a point reading system and device, and belongs to the technical field of computers.The point reading system provided by the application has multiple functions, can realize detection and identification of objects through target detection technology, can detect and identify text through text processing, and then displays the identified content to the user, realizes intelligent detection and identification, and can be used without specific teaching materials, is rich in content, breaks the content limitation, and the point reading system can detect the gesture operation of the user, so as to provide the gesture operation of the user to realize the point reading service, is simple and flexible to operate, reduces the use difficulty, adds the interestingness, and also can play a better guiding role.In addition, the point reading system has good expansibility, can be directly used, does not need to purchase related teaching materials, has good applicability, improves the education quality, has a relatively high popularization rate, well balances the uneven resource allocation and the uneven education quality, and greatly reduces the cost.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer, in particular to a pointing reading system and device. BACKGROUND

[0002] With the development of computer technology, more and more electronic systems are applied in various fields to replace manual work and bring convenience to people's life and learning. Among them, the pointing reading system is an electronic system, and people can learn through the pointing reading system, especially children aged 3-6. 3-6 years old is the key period of children's cognitive formation, and education at this stage has a great promoting effect on children's future development. Nowadays, the demand for systems to assist children in independent learning is increasing.

[0003] Through investigation, it is found that the quality of current online auxiliary education software is uneven, and some better auxiliary education software on the market has the following pain points:

[0004] (1) Expensive, a common physical education pointing machine is thousands of yuan, and the pointing machine can only be used for specially designed books or professional equipment, so the starting cost is high.

[0005] (2) Complex operation, most of the current education pointing machines have complex functions, which increases the difficulty of using the equipment and increases the cost of learning the additional equipment, causing children to develop a learning aversion.

[0006] (3) Limited content, children can only learn according to the existing data, courses and books in the pointing machine system database, which limits the richness of children's learning content. The limitation of learning content will inevitably lead to the difficulty of stimulating children's interest in learning. SUMMARY

[0007] The present application provides a pointing reading system and device, which reduces cost, is simple to operate, has rich content, improves education quality and has high popularization rate. The technical solution is as follows:

[0008] On the one hand, a pointing reading system is provided, which comprises:

[0009] A gesture recognition module is configured to detect hand key points in a plurality of video frames in real time to obtain hand skeleton information in each video frame, determine a recognition mode based on the hand skeleton information of the plurality of video frames, and determine a to-be-recognized region selected by a finger in the plurality of video frames based on the recognition mode and the hand skeleton information in the plurality of video frames.

[0010] A target detection module is configured to perform target detection on the to-be-recognized region to obtain a first detection result, and perform recognition on an object in the to-be-recognized region in response to the first detection result indicating that the to-be-recognized region contains the object.

[0011] a text extraction and recognition module, configured to perform text detection on the to-be-recognized region to obtain a second detection result, and in response to the second detection result indicating that the to-be-recognized region contains text, perform text recognition on the text;

[0012] a display module, configured to display a recognition result of the target detection module and / or the text extraction and recognition module.

[0013] In some embodiments, the hand skeleton information includes a number of hands, a type of hand, and coordinates of skeleton nodes of the hand, and the type of hand includes a left hand and a right hand.

[0014] The determination of the recognition mode based on the hand skeleton information of the plurality of video frames includes:

[0015] The determination of the recognition mode is based on the number of hands and / or the type of hand in the hand skeleton information.

[0016] The determination of the to-be-recognized region selected by the fingers based on the recognition mode and the hand skeleton information in the plurality of video frames includes:

[0017] In response to the recognition mode being a single-hand mode, and based on the coordinates of the skeleton nodes of the hand in the plurality of continuous video frames, it is determined that a moving distance of a target finger within a first target time length is less than a first target distance, a candidate box is displayed at the coordinates of the target finger, the candidate box is enlarged as the coordinates of the target finger change, and when it is determined again based on the coordinates of the skeleton nodes of the hand in the plurality of continuous video frames that the moving distance of the target finger within the first target time length is less than the first target distance, a region in the candidate box is determined as the to-be-recognized region selected by the fingers.

[0018] In response to the recognition mode being a double-hand mode, and based on the coordinates of the skeleton nodes of the hand in the plurality of continuous video frames, it is determined that a moving distance of target fingers of the two hands within a second target time length is less than a second target distance, and a to-be-recognized region selected by the fingers is determined based on a line connecting the coordinates of the target fingers of the two hands as a diagonal line.

[0019] In some embodiments, the text extraction and recognition module is further configured to perform text recognition on an uploaded document to obtain a recognition result, and the display module is further configured to display the recognition result.

[0020] In some embodiments, the point-reading system further includes a mode selection module, which provides a camera recognition mode and an uploaded video recognition mode. In response to a selection instruction for the camera recognition mode, the camera is activated to capture multiple video frames. In response to a selection instruction for the uploaded video recognition mode, the display module displays a video upload page. The gesture recognition module is used to perform real-time detection of the multiple video frames captured by the camera or to perform real-time detection of multiple video frames in a video uploaded to the video upload page.

[0021] In some embodiments, the point-reading system further includes a query module, which is used to perform an online query based on the recognition results of the target detection module and / or the text extraction and recognition module to obtain query results; the display module is also used to display the query results.

[0022] In some embodiments, the query module is used to query at least one of the following: the Chinese name, Chinese pinyin, English name, English phonetic symbols, and Chinese definition of the object, based on the recognition result of the target detection module; the query module is also used to query at least one of the following: the Chinese name, Chinese pinyin, English name, English phonetic symbols, Chinese definition, example sentences, and allusions of the text, based on the recognition result of the text extraction and recognition module.

[0023] In some embodiments, the point-reading system further includes a voice broadcast module, which is configured to perform at least one of the following:

[0024] The recognition results of the target detection module and / or the text extraction and recognition module are broadcast aloud via voice.

[0025] The query results obtained from the query module will be read aloud via voice.

[0026] It can read aloud the text content of uploaded documents.

[0027] In some embodiments, the voice broadcasting module and the gesture recognition module operate on different threads.

[0028] On the one hand, a reading device is provided, which includes electronic equipment and a 45-degree reflector;

[0029] The electronic device is equipped with a camera, and the 45-degree reflector is mounted above the camera to assist the camera in capturing text content laid flat on the table and the user's finger.

[0030] The electronic device is equipped with a reading system, which includes multiple functional modules for implementing the reading function based on the user's gesture operation.

[0031] On one hand, a computer-readable storage medium is provided, wherein at least one computer program is stored in the computer-readable storage medium, the at least one computer program being loaded and executed by a processor to realize the functions of the point-reading system described above.

[0032] On one hand, a computer program product or computer program is provided, the computer program product or computer program comprising one or more lines of program code stored in a computer-readable storage medium. One or more processors of an electronic device read the one or more lines of program code from the computer-readable storage medium, and the one or more processors execute the one or more lines of program code, causing the electronic device to perform any of the functions of the aforementioned point-and-read system.

[0033] The point-reading system provided by this invention has multiple functions. It can detect and recognize objects through object detection technology and text through text processing, then display the recognized content to the user, achieving intelligent detection and recognition without the need for specific textbooks. The rich content breaks through content limitations. Furthermore, the system has a gesture recognition module that can detect user gestures, providing point-reading services through gesture operations. The operation is simple and flexible, reducing the difficulty of use. For children, it adds fun and provides better guidance. In addition, the system has excellent scalability; it can be used directly without purchasing related textbooks, making it highly applicable, improving educational quality, and increasing its popularity. This effectively balances the uneven distribution of resources and inconsistent educational quality, significantly reducing costs. Attached Figure Description

[0034] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0035] Figure 1 This is a schematic diagram of the structure of a reading device provided in an embodiment of the present invention;

[0036] Figure 2 This is a schematic diagram of the structure of a point-reading system provided in an embodiment of the present invention;

[0037] Figure 3 This is a schematic diagram illustrating an application scenario of a point-reading system provided in an embodiment of the present invention;

[0038] Figure 4is a point reading system application scene real scene drawing provided by an embodiment of the application;

[0039] Figure 5 is a terminal interface drawing in an upload video recognition mode provided by an embodiment of the application;

[0040] Figure 6 is a terminal interface drawing in a camera recognition mode provided by an embodiment of the application;

[0041] Figure 7 is a terminal interface drawing in a camera recognition mode provided by an embodiment of the application;

[0042] Figure 8 is a terminal interface drawing in a single hand mode provided by an embodiment of the application;

[0043] Figure 9 is a terminal interface drawing in a single hand mode provided by an embodiment of the application;

[0044] Figure 10 is a terminal interface drawing in a double hand mode provided by an embodiment of the application;

[0045] Figure 11 is a terminal interface drawing in a double hand mode provided by an embodiment of the application;

[0046] Figure 12 is a terminal interface drawing in a double hand mode provided by an embodiment of the application;

[0047] Figure 13 is a terminal interface drawing in a double hand mode provided by an embodiment of the application;

[0048] Figure 14 is a terminal interface drawing in a single hand mode provided by an embodiment of the application;

[0049] Figure 15 is a terminal interface drawing in a single hand mode provided by an embodiment of the application;

[0050] Figure 16 is a terminal interface drawing in a double hand mode provided by an embodiment of the application;

[0051] Figure 17 is a terminal interface drawing in a double hand mode provided by an embodiment of the application. DETAILED DESCRIPTION

[0052] In order to make the objects, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions of the embodiments of the present application will be described clearly and completely below with reference to the drawings of the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the described embodiments of the present application, all other embodiments obtained by a person of ordinary skill in the art without creative effort belong to the scope of protection of the present application.

[0053] Unless otherwise defined, technical terms or scientific terms used in the present application shall have the usual meaning understood by a person of ordinary skill in the art to which the present application belongs. The terms "first", "second" and similar words used in the present application do not represent any order, number or importance, but are only used to distinguish different components. Similarly, the terms "one", "an" or "the" and similar words do not represent a quantity limitation, but represent the existence of at least one. The terms "including" or "containing" and similar words mean that the elements or objects before the word cover the elements or objects listed after the word and their equivalents, without excluding other elements or objects. The terms "connected" or "connected" and similar words are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect.

[0054] Here, the point reading system provided by the present application is first introduced briefly.

[0055] The point reading system provided by the present application can be named iReader. The iReader project is a smart point reading system for preschool children. Of course, it is also suitable for young children entering school, and even people of any age who want to learn. The iReader project has the characteristics of application portability, function scalability and multi-sensory interaction, and integrates the AI design prototype realized by a plurality of technologies such as target detection, text recognition (Optical Character Recognition, OCR), gesture recognition and voice broadcast in the field of artificial intelligence (Artificial Intelligence, AI). It can be used as a bottom module applied to various terminal devices with a camera, and supports the high-level design calling of applications such as Android, IOS and Web (WorldWide Web, Global Wide Area Network, also known as World Wide Web). Among them, IOS is a mobile operating system developed by Apple Inc.

[0056] Children have strong curiosity and desire for knowledge about the outside world, and this age characteristic determines that early reading needs to meet the interest of children's learning style. "Interest" has become the basic premise of children's reading. iReader, in order to solve the above problems, combined with the leading AI algorithm to build a smart reading system for preschool children, which has the advantages of application portability, function scalability, multi-sensory interaction, etc., and truly realizes "where you don't know, point where you want to know".

[0057] The project uses machine learning and other artificial intelligence technologies, which are also the current popular research direction. On the one hand, with the further maturity of artificial intelligence technology and the increasing investment of government and industry, the cloudization of artificial intelligence application will continue to accelerate, and the global artificial intelligence industry size will enter a period of rapid growth in the next 10 years. Let the machine itself recognize children's actions will also be an inevitable trend.

[0058] Among them, AI is to use digital computers or digital computer controlled machine simulation, extension and expansion of human intelligence, perception of the environment, knowledge acquisition and use of knowledge to obtain the best results of theory, method, technology and application system. In other words, artificial intelligence is a comprehensive technology of computer science, which tries to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines, so that machines have the functions of perception, reasoning and decision-making.

[0059] Artificial intelligence technology is a comprehensive discipline, involving a wide range of fields, both hardware and software technologies. Artificial intelligence basic technology generally includes such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction system, mechatronics and other technologies. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology and machine learning / deep learning and other several directions.

[0060] Computer Vision (CV) Computer vision is a science that studies how to make machines "see". More specifically, it refers to using cameras and computers to replace human eyes to identify, track and measure targets, and further process graphics so that the computer processing becomes images that are more suitable for human observation or transmission to instruments for detection. As a scientific discipline, computer vision researches related theories and technologies, and attempts to establish artificial intelligence systems that can obtain information from images or multidimensional data. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, and other technologies. It also includes common face recognition, fingerprint recognition and other biometric identification technologies.

[0061] The key technologies of speech technology include automatic speech recognition technology (ASR) and speech synthesis technology (TTS) as well as voiceprint recognition technology. Letting computers be able to hear, see, speak and feel is the future direction of human-computer interaction, and voice has become one of the most promising human-computer interaction methods in the future.

[0062] Nature Language processing (NLP) is an important direction in the field of computer science and artificial intelligence. It studies various theories and methods that can realize effective communication between people and computers using natural language. Natural language processing is a science that integrates linguistics, computer science and mathematics. Therefore, the research in this field will involve natural language, i.e. the language used in daily life, so it is closely related to the study of linguistics. Natural language processing technology usually includes text processing, semantic understanding, machine translation, robot question and answer, knowledge graph and other technologies.

[0063] Figure 1 is a structural schematic diagram of a point reading device provided by an embodiment of the application, referring to Figure 1 , the point reading device comprises an electronic device 101 and a 45-degree mirror 102.

[0064] Among them, the electronic device 101 is installed with a camera 103, and the 45-degree mirror 102 is installed above the camera, used to assist the camera 103 to collect the text content and the user's fingers laid on the desktop.

[0065] The electronic device 101 is installed with a point reading system, and the point reading system comprises a plurality of function modules, which are used to realize the point reading function based on the user's gesture operation.

[0066] In some embodiments, the electronic device 101 can be a computer, a tablet, a mobile terminal device with a camera, such as an Android, an IOS, a Harmony, a Web, and a computer terminal, and the like Internet of Things device with a camera.

[0067] Figure 2 is a structural schematic diagram of a point reading system provided by an embodiment of the application, referring to Figure 2 The point reading system can include a gesture recognition module 201, a target detection module 202, a text extraction and recognition module 203, and a display module 204.

[0068] The gesture recognition module 201 is configured to detect hand key points in a plurality of video frames in real time to obtain hand skeleton information in each video frame, determine a recognition mode based on the hand skeleton information of the plurality of video frames, and determine a to-be-recognized region selected by a finger in the plurality of video frames based on the recognition mode and the hand skeleton information in the plurality of video frames.

[0069] The target detection module 202 is configured to perform target detection on the to-be-recognized region to obtain a first detection result, and perform recognition on an object in the to-be-recognized region in response to the first detection result indicating that the to-be-recognized region contains the object.

[0070] The text extraction and recognition module 203 is configured to perform text detection on the to-be-recognized region to obtain a second detection result, and perform recognition on text in the to-be-recognized region in response to the second detection result indicating that the to-be-recognized region contains the text.

[0071] The display module 204 is configured to display a recognition result of the target detection module and / or the text extraction and recognition module.

[0072] The point reading system provided by the application has multiple functions, can realize detection and recognition of an object through a target detection technology, can realize detection and recognition of text through text processing, and further displays the recognized content to a user, realizes intelligent detection and recognition, and does not require specific teaching materials, is rich in content, breaks the content limitation, and has a gesture recognition module, which can detect a gesture operation of the user to provide a gesture operation service for the user, is simple and flexible to operate, reduces the use difficulty, is interesting for children, and can also play a better guiding role. In addition, the point reading system has good expansibility, can be directly used, does not require purchase of related teaching materials, has good applicability, improves the education quality, has a relatively high popularization rate, well balances the uneven resource allocation and the uneven education quality, and greatly reduces the cost.

[0073] The modular design and rich functions of the point reading system will be described in detail below.

[0074] The point reading system provided by the application adopts modular design, and can realize target detection, text recognition, gesture recognition and other functions. In some embodiments, the point reading system can also realize voice broadcast function.

[0075] For the gesture recognition module 201, the gesture recognition module 201 is used for gesture recognition, recognizing the user's gesture, determining the corresponding instruction according to the user's gesture, and then providing the point reading service for the user according to the instruction. Based on computer vision, a gesture operation based point reading function is provided for the user, which is simple and flexible to operate, reduces the operation difficulty, and even preschool children can easily use it, and also improves the interest.

[0076] Among them, the multiple video frames detected by the gesture recognition module 201 in real time can be real-time video frames collected by the camera, or video frames in the video uploaded in advance.

[0077] In some embodiments, the point reading system further comprises a mode selection module, and the mode selection module is used for providing a camera recognition mode and an uploaded video recognition mode. In the embodiment of the application, two modes are provided, which can be selected by the user.

[0078] The camera recognition mode means that the camera can be started, and the user's behavior can be detected in real time. If the user has gesture operation related to the instruction, it can be detected, and the instruction can be automatically executed to realize the effect of controlling the point reading function through gesture operation. Specifically, in response to the selection instruction of the camera recognition mode, the camera is started to collect multiple video frames.

[0079] The uploaded video recognition mode means that the user can upload a video, and then recognize the uploaded video to realize point reading. The uploaded video can be recorded in advance by the user, or can be downloaded through networking, and the embodiment of the application does not limit it. Specifically, in response to the selection instruction of the uploaded video recognition mode, the display module 204 is used for displaying a video upload page.

[0080] In the above two modes, the gesture recognition module is used for real-time detection of the multiple video frames collected by the camera or real-time detection of the multiple video frames in the video uploaded by the video upload page, so as to perform subsequent recognition steps.

[0081] For example, as Figure 3As shown, a schematic diagram of an application scenario of a point reading system is provided, which can be referred to as iReader, and the main interface of the point reading system can be built by using PyQt5. PyQt5 is a Python language implementation based on the graphical program framework Qt5, which is composed of a group of Python modules. Qt is a cross-platform C++ development library, which is mainly used to develop a graphical user interface (GUI). The point reading system can be mounted on a terminal device with a camera, and the terminal device is installed with an identification camera, and a 45-degree mirror is installed above the camera, which can make the identification object flat on the table for identification. In this way, the user can place the identification content on the table, such as the identification content placement area shown in the figure, which can place the identification content, for example, some reading materials, some books, etc. The point reading system can be provided with an identification content display area on the screen, which is used to display and broadcast the identification content.

[0082] In some embodiments, the gesture recognition module can be implemented based on the MediaPipe industrial-level framework of Google, or based on other gesture recognition related frameworks, and the embodiments of the present application do not limit this.

[0083] In some embodiments, the hand skeleton information includes the number of hands, the type of hands, and the coordinates of the skeleton nodes of the hands, and the type of hands includes left hands and right hands. Specifically, the captured video frame can be used to initialize the solutions.hands.Hands() object of MediaPipe, and the video frame (frame) can be processed by calling the process method of the object, so as to obtain the hand skeleton information in the frame, including the number of hands, left hands or right hands, the coordinates of the skeleton nodes of each hand, etc.

[0084] Correspondingly, when determining the recognition mode based on the hand skeleton information of the plurality of video frames, the gesture recognition module can determine the recognition mode based on the number of hands and / or the type of hands in the hand skeleton information.

[0085] In some embodiments, in response to the number of hands in the hand skeleton information being one, and / or the type of hands being left hands or right hands, it is determined that the recognition mode is a single-hand mode. In response to the number of hands in the hand skeleton information being two, or the type of hands being left hands and right hands, it is determined that the recognition mode is a double-hand mode.

[0086] After the gesture recognition module determines the recognition mode, the gesture recognition module can further determine a region to be recognized selected by a finger based on the recognition mode and the hand skeleton information, that is, the gesture recognition module can further analyze the intention of the user in the current gesture operation and determine the region to be recognized by the user. Specifically, the determination of the region to be recognized can be different in different recognition modes, that is, the user can use different operation modes to draw the region to be recognized.

[0087] In the first mode, in response to the recognition mode being the single-hand mode and based on the coordinates of the skeleton nodes of the hand in the plurality of continuous video frames, it is determined that a moving distance of a target finger of the user within a first target time length is less than a first target distance, a candidate box is displayed at the coordinates of the target finger, the candidate box is enlarged as the coordinates of the target finger change, and the region in the candidate box is determined as the region to be recognized selected by the finger when it is determined again based on the coordinates of the skeleton nodes of the hand in the plurality of continuous video frames that the moving distance of the target finger within the first target time length is less than the first target distance.

[0088] In the first mode, if the operation is a single-hand operation, the user can first stop the finger at a location, the candidate box appears, and then the finger moves, the candidate box is enlarged according to the movement of the finger, and the region in the candidate box is determined as the region to be recognized when the finger stops at the end point. It should be noted that the single-hand operation does not necessarily draw a closed box, but can also determine the range to be enclosed by the user by starting the finger movement track, and the range is determined as the region to be recognized.

[0089] The target finger can be any finger of the user or a specific finger of the user, for example, the target finger can be the index finger. The embodiments of the present application do not limit this.

[0090] The first target distance and the first target time length can be set by a person skilled in the art according to requirements. The first target time length can be set as a fixed value or a range. For example, the first target time length can be 0.3-0.5 seconds (s).

[0091] For example, in one specific example, the index finger is identified, the first target duration is 0.3-0.5 seconds, and the coordinates (x1, y1) of the captured single-hand index finger bone node are used. When the hand is continuously identified in multiple frames, the distance value of the front and back coordinates of the two hands of the two consecutive frames is less than a certain set threshold (anti-shake), for example, the two frames of the right hand coordinates are (x1, y1) and (x1', y1'), when (x1-x1')^2+(x2-x2')^2<t, it is considered that the finger is fixed. Wherein, t is the set threshold, that is, the first target duration. If the left hand is detected, the determination method is the same, and details are not described here. When the finger is identified as not moving for more than 0.3-0.5 seconds, the box selection mode is entered, and the moving finger can box select objects or text. After the box selection is completed, the finger is not moved for more than 0.3-0.5 seconds, and the box selection mode is exited. The content of the rectangular region or irregular region formed by the two points before and after the box selection is cropped out.

[0092] For example, as shown in Figure 4 , an application scenario real scene diagram of a point reading system is provided, which is realized by a computer plus a 45° mirror (serial number 1). The specific implementation process can include the following steps:

[0093] Step 1: Install the 45° mirror (serial number 1) on the computer to facilitate the placement of the content to be identified flat and facilitate the identification operation.

[0094] Step 2: Place the children's picture book (serial number 3) under the computer camera (serial number 2).

[0095] Step 3: The identification area (serial number 4) displays the content collected by the computer camera (serial number 2). The user places his hand under the computer camera (serial number 2), and the identification area (serial number 4) displays the frame of the identified hand and identifies the content in the identification area.

[0096] Step 4: The user performs a gesture operation, and the box selection area content, that is, the to-be-identified area, is displayed in the box selection content (serial number 6), and its content is identified as text below. Only the case where the box selection content is displayed in the upper right corner of the identification area (serial number 4) is described here. The box selection content can also be displayed in other positions, which can be set by relevant technical personnel according to requirements, and the embodiments of the present application are not limited thereto.

[0097] For the identification area (serial number 4), it can be divided into multiple sub-areas, each of which is used to display different content. For example, the upper left corner can display the current frame rate, the number of palms, and the identification mode of the entire video identification. The upper right corner can display a preview of the cropped to-be-identified area and the identification result. In addition to the areas displayed in the two corners for displaying the content collected by the camera and the basic information of the hand skeleton.

[0098] Step 5: For the content in the box (No. 6), the target detection module 202 and the text extraction and recognition module 203 can also identify the content to obtain the identified content, which can be displayed in the identified content display page (No. 5). For example, based on the text extraction and recognition module 203, the display module 204 can display the Chinese, Chinese pinyin, English, English standard, Chinese interpretation, example sentences, and allusion information of the recognition result in the identified content display page (No. 5), so as to facilitate the child to better extend the learning. Based on the target detection module 202, the display module 204 can display the Chinese, Chinese pinyin, English, English standard, and Chinese interpretation information of the recognition result in the identified content display page (No. 5).

[0099] In one specific example, the point reading system can be called iReader, and the prototype design interface of the iReader is the interface for calling the camera to operate and recognize. The left upper corner of the page mainly displays some basic information of the operation, which is convenient for debugging and visual display of the program running process. The right upper corner mainly displays the thumbnail of the main area of the gesture box selection and the result or label of the recognition thereof. Other areas are used for gesture operation. The information of the gesture part of the whole panel mainly includes the hand skeleton, the orange box of the main area of the palm, the current coordinate of the index finger, and the index finger staying still to start timing. The information of the gesture part of the whole panel mainly includes the hand skeleton, the main area of the palm, the current coordinate of the index finger, and the index finger staying still to start timing.

[0100] In the second way, in response to the recognition mode being the double-hand mode, and based on the coordinates of the skeleton nodes of the hands in the continuous multiple video frames, it is determined that the moving distance of the target fingers of the double hands within the second target time length is less than the second target distance. The line between the coordinates of the target fingers of the double hands is taken as a diagonal line to determine the to-be-recognized area selected by the fingers.

[0101] In the second way, if it is a double-hand operation, the two fingers can actually determine a rectangle through one up and one down or one left and one right, so as to select the box selection position. In this way, the double hands do not need to move, and only need to stay at different positions for a period of time to determine that it is not a false touch but an operation intended for recognition.

[0102] The second target distance and the second target time length can be set by a person skilled in the art according to requirements. The second target time length can be set as a fixed value or a range. For example, the second target time length can be 0.3-0.5 seconds (s).

[0103] In some embodiments, the first target distance and the second target distance can be the same or different, and the first target time length and the second target time length can be the same or different, which can be set by a person skilled in the art according to requirements, and the embodiments of the present application do not limit this.

[0104] The target finger can be any finger of the user or a specific finger of the user, for example, the target finger can be the index finger. The embodiments of the present application do not make any limitation in this regard.

[0105] For example, in a specific example, taking the index finger and a second target duration of 0.3-0.5 seconds as an example, the coordinates (x1, y1) and (x2, y2) of the captured index finger knuckle nodes of the two hands are used. When a plurality of frames are continuously identified, the distance values of the front and rear coordinates of the two hands of the two consecutive frames are less than a certain set threshold (anti-shake), for example, the two frame coordinates of the right hand are (x1, y1) and (x1', y1'), when (x1-x1')^2+(x2-x2')^2<t, it is considered that the right hand finger is fixed. The same left hand determination process is the same. When the finger is fixed for about 0.3-0.5 seconds, the content in the selected range of (x1, y1) and (x2, y2) is cropped out, and it is noted that the two points are generally framed on the diagonal of the rectangle.

[0106] For the target detection module 202, the task of object detection is to find all the targets (objects) of interest in the image, determine their categories and positions, and is one of the core problems in the field of computer vision. In the embodiments of the present application, the target detection module 202 has the function of target detection.

[0107] In some embodiments, the target detection module 202 is developed and implemented based on the PaddlePaddle framework and PaddleDetection of Baidu, where PaddleDetection is an end-to-end development kit for target detection based on PaddlePaddle.

[0108] In some embodiments, the target detection module 202 can be implemented by a model, and the main model thereof can adopt a Cascade R-CNN (Cascade Region-Convolutional Neural Networks) structure, which can be roughly divided into a backbone network result ResNet (Residual Network) 101 for extracting features, a multi-scale pyramid result FPN (Feature Pyramid Networks) layer, and an FPN RPN Head region generation network module. Among them, RPN (Region Proposal Network) and Head are the head. The target of FPN is to use the hierarchical semantic features brought by the convolutional network itself to construct a feature pyramid. The target detection module 202 adopts a lightweight model, which can reduce the load and ensure that the result can be normally output under the central processing unit (CPU).

[0109] Specifically, the target detection module 202 can input the to-be-recognized region obtained by the gesture recognition module 201 into the target detection model, perform target detection on the to-be-recognized region by the target detection model to obtain a first detection result, perform recognition on the object in response to the first detection result indicating that the to-be-recognized region contains the object, and output a recognition result.

[0110] For the target detection module 202, the first detection result obtained by performing target detection on the to-be-recognized region is used to indicate whether the to-be-recognized region contains an object. If the to-be-recognized region contains an object, the target detection module 202 is further used to perform recognition on the object. If the to-be-recognized region does not contain an object, the target detection module 202 does not need to perform object recognition.

[0111] For example, if the to-be-recognized region contains an elephant, the elephant can be detected and the position of the elephant can be determined during target detection, and then it is recognized that it is an elephant.

[0112] In one specific example of the present application, the entire model of target detection can currently recognize 676 categories of objects, and more powerful target detection models can be obtained by training more data according to some specific scenes.

[0113] For the text extraction and recognition module 203, the text detection is responsible for detecting the position of the text, and the text recognition is responsible for recognizing the text content in the position.

[0114] In some embodiments, the text extraction and recognition module 203 can also be implemented based on a text detection model and a text recognition model. The text recognition model can be developed and implemented based on the PaddlePaddle framework and PaddleOCR of Baidu. For the text detection model, a super-lightweight model can be designed, and support can be provided for Chinese and English, and even other languages, to realize multi-language text detection. For the text recognition model, a super-lightweight model can also be designed, and recognition of numbers can be added on the basis of Chinese and English, or even other languages, to reduce the load, so as to ensure that the CPU can also normally run out of results.

[0115] Specifically, the text extraction and recognition module 203 can input the to-be-recognized region obtained by the gesture recognition module 201 into a text detection model, perform text detection on the to-be-recognized region by the text detection model to obtain a second detection result, in response to the second detection result indicating that the to-be-recognized region contains text, input the detected text region in the to-be-recognized region into a text recognition model, perform recognition on the text by the text recognition model, and output a recognition result.

[0116] For the text extraction and recognition module 203, the text extraction and recognition module 203 can perform text detection on the to-be-recognized region to obtain a second detection result, and the second detection result is used to indicate whether the to-be-recognized region contains text. If the to-be-recognized region contains text, the text extraction and recognition module 203 can further recognize the text. If the to-be-recognized region does not contain text, the text extraction and recognition module 203 does not need to perform text recognition.

[0117] In some embodiments, in addition to the above two cases of containing an object and containing text, the to-be-recognized region can also contain both an object and text. For this case, how to recognize and display can be set by a person skilled in the art.

[0118] In some embodiments, a display priority can be set. In response to the to-be-recognized region containing an object and text, the module with a higher priority in the target detection module 202 and the text extraction and recognition module 203 is preferentially recognized according to the display priority, and the display module 204 preferentially displays the recognition result obtained by recognition.

[0119] In some other embodiments, the target detection module 202 and the text extraction and recognition module 203 can also simultaneously perform recognition, and then the display module 204 preferentially displays the recognition result of the module with a higher display priority.

[0120] In another embodiment, for the above case, the skilled in the art can also set to only recognize and display the text when both the object and the text are contained. Specifically, in response to the fact that the to-be-recognized region contains both the object and the text, the text extraction and recognition module 203 recognizes the text in the to-be-recognized region according to the display priority, and the display module 204 displays the recognition result of the text extraction and recognition module 203.

[0121] The above provides several possible implementation manners, and the embodiments of the present application do not limit which manner is specifically used.

[0122] In some embodiments, the above target detection module 202 and the text extraction and recognition module 203 can be implemented by a model. For example, the target detection module 202 can be implemented by any target detection model, and the text extraction and recognition module 203 can be implemented by any text detection and text recognition model. The embodiments of the present application do not limit the structure and type of the specific model.

[0123] In some embodiments, the text extraction and recognition module 203 also has a document recognition function, that is, it can directly recognize long text or even files in file format. Specifically, the text extraction and recognition module 203 is used to perform text recognition on an uploaded document to obtain a recognition result, and the display module 204 is also used to display the recognition result.

[0124] The language used in the document is not limited in the document recognition function, and the text extraction and recognition module can recognize multiple languages.

[0125] In some embodiments, in addition to being able to recognize existing content, the point reading system can also be connected to query more content for display. Specifically, the point reading system further includes a query module, the query module is used to perform network query based on the recognition result of the target detection module and / or the text extraction and recognition module to obtain a query result, and the display module is also used to display the query result.

[0126] In some embodiments, the query content of the object and the text can also be different. For example, the object can query at least one of the Chinese name, Chinese pinyin, English name, English phonetic symbol, and Chinese explanation of the object. The text can query at least one of the Chinese, Chinese pinyin, English, English phonetic symbol, Chinese explanation, example sentence, and allusion of the text.

[0127] Specifically, the query module is configured to query at least one of a Chinese name, a Chinese pinyin, an English name, an English phonetic symbol, and a Chinese explanation of the object based on the identification result of the target detection module; and the query module is configured to query at least one of Chinese, Chinese pinyin, English, English phonetic symbol, Chinese explanation, example sentence, and allusion of the text based on the identification result of the text extraction and identification module.

[0128] In some embodiments, for pinyin query, a PyPinyin package provided by Python is used for identification. For translation content, a crawler technology can be used to obtain content of a translation software (e.g., Youdao translation, Baidu translation, etc.), including Chinese-English translation, phonetic symbol acquisition, etc. For explanation example sentence allusion, a Baidu API (Application Programming Interface) query interface can be called to obtain.

[0129] In some embodiments, in addition to displaying the identification result and the query result, the point reading system can also have a voice broadcast function, and can perform voice broadcast on the identification result and / or the query result. Specifically, the point reading system further includes a voice broadcast module, and the voice broadcast module is configured to perform at least one of the following steps one to three.

[0130] Step one, voice broadcast is performed on the identification result of the target detection module and / or the text extraction and identification module.

[0131] Step two, voice broadcast is performed on the query result obtained by the query module.

[0132] Step three, voice broadcast is performed on the text content in the uploaded document.

[0133] In some embodiments, the voice broadcast module can use a pyttsx3 module to perform reading operation, or can use other reading software to implement, and the embodiments of the present application do not limit this.

[0134] In some embodiments, the point reading system can support loading of a voice package, and the point reading system can add a Baidu voice package, a Microsoft voice package, etc. to achieve more diversified voice broadcast.

[0135] In some embodiments, in order to avoid thread confliction and blocking problem of the voice broadcast module and the gesture recognition module, the voice broadcast module and the gesture recognition module work based on different threads. That is, the voice broadcast module can work based on a separate thread, so as not to block the identification process.

[0136] In some embodiments, the background of the point reading system can record the use data of the user, further optimize and update the point reading system based on the use data of the user, or analyze and summarize the situation of the user. For example, a relevant query API interface can be called, and the pinyin and phonetic symbols are returned when identifying Chinese characters and English, and there is a translation function, and the user learning record is tracked, which is convenient for students to review and parents to supervise. Even access to knowledge dialogue question and answer system, gesture access to browser engine search function, make inquiry more portable, rich and popular, and the application population can be popularized to the general public.

[0137] It should be noted that the point reading system provided by the present application has the advantages of strong expansibility, strong interactivity and strong portability, and the advantages will be described in detail below.

[0138] Strong expansibility: The modular design makes the entire application encapsulated as a general kernel prototype by Python modularization, and the content expansion is very strong. The function of the application program can be quickly expanded at any time according to new ideas and inspirations, and the module encapsulation can be designed on various mobile terminal devices to carry corresponding high-level products, such as Android, IOS, Harmony, Web and computer terminal and other Internet of Things devices containing cameras. At the same time, the function expansion is strong, and new functions can be encapsulated into API interfaces at any time for easy calling by actual applications. For example, the knowledge question and answer system is expanded. At the same time, the application population can be expanded, not only for young children, but also for foreign language learning of teenagers, etc. Perhaps the foreign language API such as phonetic symbol and translation can be expanded to realize foreign language learning function applicable to more people.

[0139] Strong interactivity: The gesture recognition is used to control the content that the user actually wants to learn and understand, which increases the "interest" of children's learning. A humanized, intelligent and emotional human-computer interaction mode is adopted, which integrates vision, hearing, touch and other multiple senses. The interactive instruction is simple and easy to use, which makes the human-computer interaction more natural, easy and efficient. Most operations can be completed by the index finger, and the recognition result is displayed in the form of text or voice, which fully improves the reading experience and interest of children.

[0140] Strong portability: the children's education machines existing in the current market are all built on the equipment specific to manufacturers. Users need to purchase additional equipment such as tablets, home education machines to experience assisted education, or need to specially purchase special books, and the additional equipment also has the characteristics of inconvenient to carry. The wisdom point reading system based on computer vision application provided by the present application does not need to purchase additional point reading equipment, and through the mobile terminal equipment with a carrying camera such as computer, tablet and mobile phone commonly used by the public, cooperating with a very cheap 45-degree mirror, the learning can be more portable, so that the user can use the point reading application for learning at any time and anywhere. And the point reading system provided by the present application can be loaded on a mobile terminal, and can be packaged as a web application, so that the user can log in to the web version of the point reading system on the mobile terminal to enjoy the point reading service.

[0141] Through investigation, it is found that the quality of the current online assisted education software is uneven, and some better assisted education software on the market has the problem of high price. The iReader is a software product, which greatly reduces the cost of children's education. Moreover, the point reading machine can only be used for specially designed books or professional equipment, and after purchasing the point reading machine, the specially designed books need to be purchased, and the education content of children is limited and the cost is high. The point reading system provided by the present application does not need to purchase books, and all books in daily life can be used in the software, which further reduces the cost of education and improves the purchase desire of the consumer group.

[0142] For the point reading system provided by the present application, the related technical personnel has tested and analyzed the point reading system, and the test environment, test process and test results are described below.

[0143] Test environment:

[0144] The computer model used is XiaoXin-15IIL 2020; Windows 10 Home Chinese Edition.

[0145] The processor is: processor: Intel(R) Core(TM) i7-1065G7 CPU@1.30GHz 1.50GHz.

[0146] The running speed test is as follows:

[0147] The recognition result duration is about 1 minute (min) for Chinese characters and about 2 minutes (min) for objects.

[0148] The gesture recognition delay is about 0.15 seconds (s).

[0149] The running speed result shows that the GPU processor is faster and the CPU processor is slower, but both are normal indicators and have no effect on user experience.

[0150] The results of the various functional tests are shown in Table 1.

[0151] Table 1

[0152] Test function Result Video upload recognition Normal Two-hand skeleton recognition Normal, delay within normal range Text recognition Normal Object recognition Both single and double hands are normal Point-to-read recognition Both single and double hands are normal

[0153] The test results are explained in detail below.

[0154] like Figure 5 As shown, in the video upload recognition mode, the uploaded video content can be as follows: Figure 5 The image shown is a promotional image for a street dance competition. In the video, the user's finger is on the English text below the word "dance" in the competition. The gesture recognition module can mark the position of the hand skeleton nodes in the video and display a preview image of the area where the finger is located in the upper right corner, along with the object recognition result "None" and the text recognition content "NCECONTEST". Because it is an uploaded video recognition mode, it did not recognize the single-hand or two-hand mode.

[0155] like Figure 6 As shown, in camera recognition mode, selecting "Open Camera" on the right will activate the camera. In this case, the left side of the screen will display the camera's view. If you place your hands under the camera, as shown... Figure 7 As shown, the camera captures the user's hands and displays the recognition effect of the hand skeleton.

[0156] like Figure 8 As shown, if a user performs a single-handed gesture using their right hand, the top left corner displays the current camera frame rate as 4, the number of hands as 1, and the recognition mode as single-handed. After the user's index finger stops moving from one point, a candidate box appears. Then, as the user moves their index finger, the movement trajectory is displayed in the image, and the candidate box expands accordingly. Once the candidate box has completely enclosed the "waiting for the rabbit to run into the tree stump" image, the user's index finger stops moving. At this point, the following can be displayed: Figure 9 As shown, the recognition results are displayed in the upper right corner of the camera's acquisition area. Here, a preview image of the candidate box, which is the area to be recognized, is displayed, along with the object recognition result "Whiteboard" and the text recognition result "Waiting for a rabbit to run into a tree stump". The right side of the screen displays the Chinese, pinyin, English, phonetic symbols, Chinese definition, example sentences, and allusion of "Waiting for a rabbit to run into a tree stump", as well as the Chinese, pinyin, English, phonetic symbols, and Chinese definition of the object "Whiteboard".

[0157] like Figure 10As shown, if the user uses both hands to perform gesture operation, the upper left corner can display that the current frame rate of camera capture is 4, the number of palms is 2, and the recognition mode is double (double-hand mode). The user's left index finger and right index finger are respectively placed at the lower left corner and the upper right corner of "Zhu Shu Daiyu" and do not move, and then the upper right corner of the camera capture area displays the preview of the region to be recognized and the object recognition result "Whiteboard", as well as the text recognition result "Zhu Shu Daiyu". As shown in Figure 11 , based on the recognition result, a query can be performed, and the right window displays the Chinese, pinyin, English, phonetic alphabet, Chinese interpretation, example sentence, and allusion of "Zhu Shu Daiyu", as well as the Chinese, pinyin, English, phonetic alphabet, and Chinese interpretation of the object "Whiteboard".

[0158] The above several examples take recognizing Chinese text as an example. If English text is recognized, as shown in Figure 12 , if the user uses both hands to perform gesture operation, the upper left corner can display that the current frame rate of camera capture is 4, the number of palms is 2, and the recognition mode is double (double-hand mode). The user's left index finger and right index finger are respectively placed at the lower left corner and the upper right corner of "Earth" and do not move. As shown in Figure 13 , after the result is recognized, the upper right corner of the camera capture area is refreshed to display the preview of the region to be recognized and the object recognition result "None", as well as the text recognition result "Earth", and the right window displays the Chinese, pinyin, English, phonetic alphabet, Chinese interpretation, and example sentence of "Earth".

[0159] As shown in Figure 14 , if the user uses the right hand to perform single-hand gesture operation, the upper left corner can display that the current frame rate of camera capture is 3, the number of palms is 1, and the recognition mode is single (single-hand mode). The user wants to recognize the animal in the figure, and then the index finger stops at a point, and a candidate box appears. Then the user moves the index finger, and the moving track is displayed in the figure, and the candidate box is enlarged. When the candidate box frames the animal to be recognized, the user's index finger stops. Further, as shown in Figure 15 , the upper right corner of the camera capture area displays the preview of the region to be recognized and the object recognition result "Zebra", and the right side area of the screen displays the Chinese, pinyin, English, phonetic alphabet, and Chinese interpretation of "Zebra".

[0160] As shown in Figure 16 , if the user uses both hands to perform gesture operation, the upper left corner can display that the current frame rate of camera capture is 4, the number of palms is 2, and the recognition mode is double (double-hand mode). The user wants to recognize the animal in the figure, and the left index finger and right index finger are respectively placed at the lower left corner and the upper right corner of the animal and do not move. As shown in Figure 17As shown, the upper right corner of the camera capture area displays a preview of the region to be identified and the object recognition result "Elephant", and the right window displays the Chinese, pinyin, English, phonetic alphabet, Chinese interpretation of "Elephant".

[0161] In an example embodiment, a computer readable storage medium is also provided, for example, a memory including at least one computer program executable by a processor to perform the functions of the point reading system in the above embodiments. For example, the computer readable storage medium is a Read-Only Memory (ROM), a Random Access Memory (RAM), a Compact Disc Read-Only Memory (CD-ROM), a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0162] In an example embodiment, a computer program product or computer program is also provided, the computer program product or the computer program including one or more program codes stored in a computer readable storage medium. One or more processors of an electronic device read the one or more program codes from the computer readable storage medium, and the one or more processors execute the one or more program codes to cause the electronic device to perform the above multi-objective wild horse optimization method of parameters.

[0163] In some embodiments, the computer program related to the embodiments of the present application can be deployed to execute on one computer device, or on multiple computer devices located in one place, or on multiple computer devices distributed in multiple places and interconnected through a communication network, which can constitute a blockchain system.

[0164] Those of ordinary skill in the art understand that all or part of the steps of the above embodiments can be implemented by hardware, or by programs instructing related hardware, which are stored in a computer readable storage medium, and the storage medium mentioned above is a Read-Only Memory, a magnetic disk or an optical disk, etc.

[0165] The above description is only an optional embodiment of the present application and does not limit the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A point-and-read system, characterized in that, The point-reading system includes: A gesture recognition module is used to detect key hand points in multiple video frames in real time to obtain hand skeleton information in each video frame. The hand skeleton information includes the number of hands, the hand type, and the coordinates of the hand skeleton nodes. The hand type includes left and right hands. A recognition mode is determined based on the number of hands and / or the hand type in the hand skeleton information. In response to the recognition mode being a single-hand module, and based on the coordinates of the hand skeleton nodes in multiple consecutive video frames, if the movement distance of the target finger within a first target duration is determined to be less than a first target distance, a candidate box is displayed at the coordinates of the target finger. As the coordinates of the target finger change, the candidate box is expanded until it is again determined based on the coordinates of the hand skeleton nodes in multiple consecutive video frames that the movement distance of the target finger within the first target duration is less than the first target distance. The area within the candidate box is then used as the area to be recognized for finger selection. In response to the recognition mode being a two-hand mode, and based on the coordinates of the hand skeleton nodes in multiple consecutive video frames, if the movement distance of the target fingers of both hands within a second target duration is determined to be less than a second target distance, the area to be recognized for finger selection is determined using the line connecting the coordinates of the target fingers of both hands as a diagonal. The target detection module is used to perform target detection on the region to be identified to obtain a first detection result, and in response to the first detection result indicating that the region to be identified contains an object, to identify the object; The text extraction and recognition module is used to perform text detection on the region to be recognized to obtain a second detection result, and in response to the second detection result indicating that the region to be recognized contains text, to recognize the text; The display module is used to display the recognition results of the target detection module and / or the text extraction and recognition module.

2. The system according to claim 1, characterized in that, The text extraction and recognition module is also used to perform text recognition on the uploaded document to obtain the recognition result; the display module is also used to display the recognition result.

3. The system according to claim 1, characterized in that, The point-reading system also includes a mode selection module, which provides a camera recognition mode and an uploaded video recognition mode. In response to the selection instruction for the camera recognition mode, the camera is activated to capture multiple video frames. In response to the selection instruction for the uploaded video recognition mode, the display module displays a video upload page. The gesture recognition module is used to detect multiple video frames captured by the camera in real time or to detect multiple video frames in a video uploaded to the video upload page in real time.

4. The system according to claim 1, characterized in that, The point-reading system also includes a query module, which is used to perform online queries based on the recognition results of the target detection module and / or the text extraction and recognition module to obtain query results; the display module is also used to display the query results.

5. The system according to claim 4, characterized in that, The query module is used to query at least one of the following: Chinese name, Chinese pinyin, English name, English phonetic symbols, and Chinese definition of the object, based on the recognition results of the target detection module; the query module is also used to query at least one of the following: Chinese name, Chinese pinyin, English name, English phonetic symbols, Chinese definition, example sentences, and allusions of the text, based on the recognition results of the text extraction and recognition module.

6. The system according to any one of claims 1-5, characterized in that, The point-reading system further includes a voice broadcast module, which is used to perform at least one of the following: The recognition results of the target detection module and / or the text extraction and recognition module are broadcast aloud via voice. The query results obtained from the query module will be read aloud via voice. It can read aloud the text content of uploaded documents.

7. The system according to claim 6, characterized in that, The voice broadcast module and the gesture recognition module operate on different threads.

8. A reading device, characterized in that, The reading device includes electronic equipment and a 45-degree reflector; The electronic device is equipped with a camera, and the 45-degree reflector is mounted above the camera to assist the camera in capturing text content laid flat on the table and the user's finger. The electronic device is equipped with a reading system, the reading system comprising: A gesture recognition module is used to detect key hand points in multiple video frames in real time to obtain hand skeleton information in each video frame. The hand skeleton information includes the number of hands, the hand type, and the coordinates of the hand skeleton nodes. The hand type includes left and right hands. A recognition mode is determined based on the number of hands and / or the hand type in the hand skeleton information. In response to the recognition mode being a single-hand module, and based on the coordinates of the hand skeleton nodes in multiple consecutive video frames, if the movement distance of the target finger within a first target duration is determined to be less than a first target distance, a candidate box is displayed at the coordinates of the target finger. As the coordinates of the target finger change, the candidate box is expanded until it is again determined based on the coordinates of the hand skeleton nodes in multiple consecutive video frames that the movement distance of the target finger within the first target duration is less than the first target distance. The area within the candidate box is then used as the area to be recognized for finger selection. In response to the recognition mode being a two-hand mode, and based on the coordinates of the hand skeleton nodes in multiple consecutive video frames, if the movement distance of the target fingers of both hands within a second target duration is determined to be less than a second target distance, the area to be recognized for finger selection is determined using the line connecting the coordinates of the target fingers of both hands as a diagonal. The target detection module is used to perform target detection on the region to be identified to obtain a first detection result, and in response to the first detection result indicating that the region to be identified contains an object, to identify the object; The text extraction and recognition module is used to perform text detection on the region to be recognized to obtain a second detection result, and in response to the second detection result indicating that the region to be recognized contains text, to recognize the text; The display module is used to display the recognition results of the target detection module and / or the text extraction and recognition module.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one computer program, which is loaded and executed by a processor to implement the functions of the point-reading system as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Real-time three-dimensional double-hand gesture recognition method and system based on binocular vision

    CN103927016A

  • DISTINGUISHING BETWEEN ONE-HANDED AND TWO-HANDED GESTURE SEQUENCES IN VIRTUAL, AUGMENTED, AND MIXED REALITY (xR) APPLICATIONS

    US20190384405A1