Multi-finger word lookup method and apparatus, computer device, and storage medium

By determining the order of word search for multiple fingers and processing it in turn, the problem of failure and low accuracy of recognition when word searching for multiple fingers is solved, and a higher text recognition accuracy and user experience is achieved.

WO2025123307A1PCT designated stage expired Publication Date: 2025-06-19GUANGZHOU XIBEISI INTELLIGENT TECHNOLOGY CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2023/139014
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-15
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

The existing finger word search methods cannot effectively deal with the situation where multiple fingers look up words in succession, resulting in failed recognition or low text recognition accuracy.

Method used

By obtaining the video to be detected, the word search order of multiple fingers is determined, and the images to be recognized are obtained in this order, pre-processing and text recognition are performed, and the recognition results are output.

Benefits of technology

This has achieved improved text recognition accuracy when looking up words with multiple fingers, and improved user human-computer interaction experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2023139014_19062025_PF_FP_ABST
    Figure CN2023139014_19062025_PF_FP_ABST
Patent Text Reader

Abstract

The embodiments of the present application relate to the technical field of artificial intelligence, in particular to a multi-finger word lookup method and apparatus, a computer device and a storage medium. The method comprises: acquiring a video to be tested; determining that there are at least two fingers performing word lookup actions in said video; determining the word lookup sequence of the corresponding fingers, and, according to the word lookup sequence, sequentially acquiring images to be identified; preprocessing said images to obtain text region identification images; performing text identification on the text region identification images, and, according to the word lookup sequence, sequentially outputting text identification results pointed to by the fingertips. The method can improve the accuracy of text identification when multiple fingers perform word lookups sequentially.
Need to check novelty before this filing date? Find Prior Art

Description

Multi-finger word search method, device, computer equipment and storage medium Technical Field

[0001] The embodiments of the present application relate to the field of artificial intelligence technology, and in particular to a multi-finger word search method, apparatus, computer equipment, and storage medium. Background Art

[0002] With the rapid development of artificial intelligence technology, more and more people are choosing computers for online learning. Finger search, as an intelligent application of human-computer interaction, has been widely used in educational scenarios. Finger search means that when a user places their hand under the camera and points their fingertips at a word or sentence, the application can return the corresponding content. However, existing finger search methods are based on the processing logic of single-finger search, and do not take into account the situation where multiple fingers can search for words in sequence. If the user searches for words with multiple fingers in sequence, recognition failure or low text recognition accuracy may occur.

[0003] Summary of the Invention

[0004] The embodiments of the present application aim to provide a multi-finger word search method, apparatus, computer device and storage medium to solve the technical problem of recognition failure or low text recognition accuracy when multiple fingers search for words one after another.

[0005] In a first aspect, a multi-finger word-searching method is provided, the method comprising: obtaining a video to be detected; determining whether at least two fingers in the video to be detected have word-searching actions; determining a word-searching order for the corresponding fingers, and sequentially obtaining images to be recognized according to the word-searching order; pre-processing the images to be recognized to obtain text area recognition images; performing text recognition on the text area recognition images, and sequentially outputting text recognition results of the fingertips pointing to the text according to the word-searching order.

[0006] In a second aspect, a multi-finger word-searching device is provided, which includes: an acquisition module for acquiring a video to be detected; a palm detection module for determining whether at least two fingers in the video to be detected have a word-searching action; the palm detection module is also used to determine the word-searching order of the corresponding fingers, and to sequentially acquire images to be recognized according to the word-searching order; a text processing module for pre-processing the image to be recognized to obtain a text area recognition image, and performing text recognition on the text area recognition image; and an output module for sequentially outputting text recognition results of the fingertips pointing to the fingertips according to the word-searching order.

[0007] In a third aspect, a computer device is provided, comprising a camera, a memory, and a processor, wherein the camera is used to obtain a video to be detected, the memory and the camera are connected to the processor, and the processor is used to execute one or more computer programs stored in the memory. When the processor executes the one or more computer programs, the computer device implements the method described in the first aspect.

[0008] According to a fourth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the processor executes the method according to the first aspect.

[0009] The embodiments of the present application can achieve the following technical effects: a computer device obtains a video to be detected, and when it is detected that at least two fingers in the video to be detected have a word-looking action, determines the word-looking order of the corresponding fingers, and obtains images to be recognized in sequence according to the word-looking order, pre-processes the images to be recognized to obtain text area recognition images, performs text recognition based on the text area recognition images, and outputs recognition results in sequence according to the word-looking order. By realizing multi-target tracking when multiple fingers are looking up words, obtaining images to be recognized based on the finger states, and outputting recognition results in sequence according to the word-looking order, the text recognition accuracy during multi-finger word-looking can be guaranteed, and the user's human-computer interaction experience can be improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments of the present application. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0011] FIG1A is a schematic diagram showing the effect of a conventional single-finger word search;

[0012] FIG1B is a schematic diagram of the effect of multi-finger word search provided by an embodiment of the present application;

[0013] FIG2 is a schematic diagram of a curved text image provided by an embodiment of the present application;

[0014] FIG3 is a flow chart of a multi-finger word search method according to an embodiment of the present application;

[0015] FIG4 is a schematic diagram of a first text kernel map provided in an embodiment of the present application;

[0016] FIG5 is a schematic diagram of a first text segmentation graph provided in an embodiment of the present application;

[0017] FIG6 is a schematic diagram of an initial text line graph provided in an embodiment of the present application;

[0018] FIG7 is a schematic diagram of a third text segmentation graph provided in an embodiment of the present application;

[0019] FIG8 is a schematic diagram of a text region recognition image provided by an embodiment of the present application;

[0020] FIG9 is a schematic diagram of the structure of a multi-finger word-searching device provided in an embodiment of the present application.

[0021] FIG10 is a schematic structural diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0022] In order to make the purpose, technical solutions and advantages of this application more clearly understood, the present application is further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0023] It should be noted that, if there is no conflict, the various features in the embodiments of the present application can be combined with each other and are all within the scope of protection of the present application. In addition, although the functional modules are divided in the device schematic and the logical order is shown in the flow chart, in some cases, the steps shown or described can be performed in a different order than the module division in the device or the order in the flow chart. Furthermore, the words "first", "second", "third", etc. used in this application do not limit the data and execution order, but only distinguish between the same items or similar items with basically the same functions and effects.

[0024] To make it easier to understand this application, we first introduce the existing finger search scheme. Finger search is an intelligent application of human-computer interaction. When the user puts his hand under the camera and extends his fingertip to point at a word or sentence, the corresponding word (or sentence) interpretation, pronunciation and other content can be displayed on the device. The existing finger search method is usually based on the processing logic of single-finger search. When searching, when the user's fingertip is stationary, text detection and recognition will be performed based on the area where the fingertip is located at the current moment, so as to obtain the word corresponding to the fingertip and its related content, and return it to the application interface. For example, as shown in Figure 1A, the user puts a book under the camera of the computer device. When the user points his finger to a certain place in the book and keeps the fingertip in a stable state within a certain time range, the camera captures this scene, and the computer device can start the finger search function. According to the position pointed by the fingertip, the corresponding text box area is obtained for text detection and recognition, and the word corresponding to the user's fingertip and its related content are returned on the application screen.

[0025] Based on the processing logic of single-finger word search, if different fingers of the user (different fingers of the same palm, or fingers of different palms) appear in the recognition range at the same time or nearly at the same time, the existing solution will generally randomly select a finger or default to the first finger that appears in the recognition range for word search recognition and output the word search result, and will not analyze all the fingers that enter the recognition range. This method cannot well reflect the user's word search intention. It can be imagined that in the single-finger word search solution, the computer device will only output one word search result. If the method of outputting the word search result corresponding to the finger with the longest stable time is selected according to the time of finger word search, although it can represent the user's intention to a certain extent, it requires the user to be highly cooperative. It is difficult to require some younger users to have a high degree of cooperation. Such users are usually the key group for finger word search. Therefore, this method is usually not chosen in existing solutions.

[0026] There is a solution in the related art that can determine the user's search finger by analyzing the priority of finger search and output the search result corresponding to the finger. For example, the search result with the highest completeness of the finger entering the recognition area can be output, or the priority order can be preset in advance as: index finger > middle finger > ring finger > little finger > thumb. If it is determined that there is an index finger in the recognition area, the search result of the index finger is output. However, during the research and development process, the inventor found that if the priority construction logic is directly applied to the result output of multi-finger search (that is, the search result of the finger with the highest priority is output first, and then the search result of the finger with the second highest priority is output), although the result output of multi-finger search is achieved, it still cannot better reflect the user's true search intention, because in the multi-finger search process, the words the user wants to search usually have a certain logical connection, and the finger with the highest completeness of the recognition area is not necessarily the search finger that the user wants to output first, and its accuracy and precision are low. The preset finger priority method requires the user to remember the finger priority, which adds an extra burden to the user and the user experience is poor.

[0027] In addition, text detection and recognition in existing solutions can be done in two ways: the first is to crop the corresponding text box area in the original image based on the position of the fingertip for word recognition; the second is to directly obtain the foreground area in the original text image for word recognition without cropping. However, if the text image collected by the first method is curved (as shown in Figure 2), the position of other text areas may be included during cropping, which may affect the recognition speed and accuracy to a certain extent; in the second method, since recognition is performed directly without cropping, a large amount of edge information and noise must be processed during the recognition process, which may also reduce the accuracy and speed of text recognition to a certain extent.

[0028] To address the above issues, this application creatively proposes a novel technical solution that reconstructs the processing logic for finger search based on the multi-finger word search scenario. Specifically, through multi-target tracking, the current state of each finger is determined. When multiple fingers are clearly searching for a word, the search is performed in the order in which the fingertips are stable. Furthermore, during the text recognition process, optimizations are made for text image detection and recognition, effectively ensuring the speed and accuracy of text recognition.

[0029] For example, please refer to FIG1B, which is a schematic diagram of the effect of a multi-finger word search provided by an embodiment of the present application. In FIG1B, the computer device can collect the image of the current book in real time through the camera and save it as video stream data. If it is detected that there is a finger in the word search action, the word search function can be started (or the user can also turn on the word search function in advance). Furthermore, palm detection is performed on the single-frame image in the video stream data, and the palm images of all palms in the single-frame image are obtained. Based on the tracking algorithm and other methods, a matching relationship between the palms of the previous and next frames in the video stream is established to determine whether it is the same palm, thereby determining the movement trajectory of each palm. Then, it is determined whether each palm has a word search action. When there is a word search action, the key points of the palm are estimated to determine the fingertip position. If the fluctuation range of the fingertip position is small within the preset time range, it can be considered that the fingertip of the palm is in a stable state. Then, the image at the current moment can be obtained for text detection and recognition, and the word recognition result of the corresponding position of the fingertip is determined and output on the current application interface.

[0030] Based on the above method, the computer device can determine the order in which the fingertips of multiple fingers are in a stable state, obtain corresponding images in sequence according to the order, and output the word search results in sequence. In this way, multiple fingers can be used to search for words in sequence. Compared with the existing technology that only considers the situation when searching for words with a single finger, the multi-finger word search scheme shown in this application fully considers the association relationship between the fingertips in the multi-finger word search process (such as the need to first identify whether the previous and next frames are the same palm), the order of finger word search (such as determining the order of word search result output according to the moment when the fingertips are in a stable state), and other factors, and reconstructs the processing logic of multi-finger word search, so that the accuracy of word search recognition is effectively improved in the application scenario of multi-finger word search, and there is no need for users to remember the priority order of fingers, which reduces the user's memory burden and improves the user experience.

[0031] Preferably, in the process of acquiring the image at the current moment for text detection and recognition, the computer device can first intercept the target area at the current moment (such as the ROI area), and perform text segmentation on the target area to obtain the text core map and text segmentation map corresponding to the fingertip position, as well as other text areas, and then set the values ​​of the other text areas to the values ​​of the background area. Furthermore, the text line of the text can be obtained based on the text core map, and the text height of each position in the text line can be determined based on the text segmentation area, the text height is smoothed, and the text height of the smoothed text line is obtained. Then, from left to right, a circle is drawn for each text position in turn. The diameter of the circle can be the text height of the corresponding position, and the interior of the circle is all set as the foreground. Then, the union of the reconstructed area and the original area is taken to obtain a new segmentation map. Furthermore, a new text area is obtained based on the new segmentation map, text recognition is performed on the text area, and the position of each character is obtained. Based on the character position, the word corresponding to the fingertip is obtained.

[0032] It can be seen that compared with the existing technology of first cropping and then recognizing text, or directly recognizing text, the above-mentioned text detection and recognition method optimizes the original text area by performing a series of processes such as cropping, segmenting, smoothing, and reconstructing the acquired image, removes the original noise and edge information, and obtains a new text area. Text recognition is performed based on the new text area, which can effectively ensure the accuracy of text recognition. Especially when the text is curved text, the method shown in this application will be more applicable.

[0033] The following is a detailed description of the multi-finger word search method described in this application. Please refer to Figure 3, which shows the multi-finger word search method described in this application applied to a computer device. The computer device can be a smart learning device, a computer, a tablet device, a smartphone, or other device with a finger word search application or function, and this application does not limit this. The method includes:

[0034] S301: Obtain a video to be detected.

[0035] In a feasible implementation, the computer device can capture images of the current lens shooting area in real time through its own camera or a connected camera. When a book or paper with text or a terminal with text on its screen (such as an e-book, a web page interface, etc.) appears in the current lens shooting area, video recording can be started and used as the video to be detected.

[0036] S302: Determine whether at least two fingers in the video to be detected have a word-looking action.

[0037] It should be noted that the at least two fingers may be at least two different fingers of the same palm, or at least one finger of different palms. That is, in this embodiment, the at least two fingers performing word-searching actions may include the following situations: different fingers of the same palm perform word-searching actions one after another, or at least one finger of different palms perform word-searching actions one after another.

[0038] In one embodiment, before determining whether at least two fingers in the video to be detected are performing word-looking actions, the method further includes:

[0039] S3021: Perform palm detection on the video to be detected.

[0040] In this embodiment, palm detection can determine whether the palms in the video are the same palms, and can track the same palms and determine the movement trajectory of the same palms, which can lay the foundation for the subsequent positioning of multiple fingertips and determining the order of finger word search.

[0041] In one embodiment, the S3021 step includes: obtaining each single-frame image in the video to be detected; identifying whether each single-frame image contains a palm image; if so, determining the matching relationship between the palms in the previous and next single-frame images based on a tracking algorithm, and determining whether the palms in the video to be detected are the same palms based on the matching relationship; and determining the movement trajectory of the same palm.

[0042] In this embodiment, the computer device can perform palm detection when confirming that the finger word lookup function is turned on. The finger word lookup function can be turned on by default, or can be turned on by the user according to the user's needs. This application does not impose any restrictions on this.

[0043] In this embodiment, the computer device can identify whether each single frame of the video stream to be detected contains a palm image based on an image recognition algorithm. For example, a classifier can be trained to identify hands in an image. The training data set should contain palm image samples in different postures and conditions. Alternatively, deep learning models such as convolutional neural networks can be used for training. These models take images as input and extract palm features in multiple convolution and pooling layers. These features are then classified to determine whether there is a palm in the image. Of course, the above method is only an example, and the computer device can also identify the palm image based on other methods. This application does not impose any restrictions on this.

[0044] In this embodiment, the tracking algorithm may be the DEEPSORT algorithm, which can be used to establish a matching relationship between palms in previous and next frames in a video stream, thereby determining whether they are the same palm and further determining the motion trajectory of each identical palm. For example, if the similarity between the matching relationships between palms in previous and next frames in a video stream exceeds a preset range, the palms can be determined to be the same; otherwise, they can be determined to be different palms.

[0045] In this embodiment, the matching relationship between the palms of the previous and next frames can be obtained based on multiple feature values ​​such as the shape, size, skin, and color of the palms.

[0046] As a feasible implementation, if the palm images in the video to be detected are all of the same palm, then the determined motion trajectory of the same palm may be one. If the palm images in the video to be detected are not the same palm, assuming there is palm A and palm B, the motion trajectory of palm A and the motion trajectory of palm B can be determined separately. In other words, when the palm images are of different palms, the number of determined motion trajectories can be multiple, and the specific number can be consistent with the number of different palms.

[0047] S3022: If there is a palm image in the video to be detected and the palm has a word-looking gesture, determine the movement trajectory of the palm and estimate key points of the palm.

[0048] It should be noted that the word-searching gesture refers to a gesture by which the user instructs the computer device to search for a word at a certain location, such as the gesture shown in FIG1B .

[0049] As a feasible implementation, a computer device performs key point estimation on the palm of a hand, aiming to locate key points on the palm, thereby determining the positions of fingertips among numerous key points. Key point estimation can use a convolutional neural network (CNN)-based method, such as the hourglass network or OpenPose, and this application does not impose any restrictions on this.

[0050] S3023. Determine the fingertip position based on the key point estimation result.

[0051] As a feasible implementation method, the key point estimation results may include key point positions such as finger joints, palm centers, and fingertips. The computer device can determine the fingertip positions from the key point positions, thereby more accurately capturing information such as the fingertip movement amplitude and trajectory.

[0052] As a feasible implementation method, after determining the position of the fingertip, the fingertip may continue to shake, which may indicate that the user has not yet determined the word to be queried. Therefore, the computer device can further execute the S3024 step to determine whether the fingertip is stable based on the shaking amplitude of the fingertip within a fixed time, thereby improving the accuracy of text recognition.

[0053] S3024: If the amplitude change of the fingertip position within the preset time range does not exceed the preset amplitude range, it is determined that the finger has a word-looking action, and the fingertip position of the finger having the word-looking action is located.

[0054] For example, after locating the fingertip, if the computer device determines that the amplitude of the fingertip changes within a preset time range within the preset amplitude range, the fingertip may be considered stable at this time; otherwise, the fingertip is not stable. If the fingertip is stable, it indicates that the user has determined the word to be searched, and the computer device may determine that the finger is performing a word-searching action and locate the fingertip position at the current moment.

[0055] Furthermore, after locating the fingertip position of the finger that is performing the word-searching action, the computer device may obtain an image of the moment when the fingertip is determined to be in a stable state as the image to be recognized. If multiple fingers are performing the word-searching action, step S303 may be executed.

[0056] S303: Determine the word search order corresponding to the finger, and sequentially obtain the images to be recognized according to the word search order.

[0057] In this embodiment, each finger can search for words in sequence, and the computer device can locate the image at the corresponding moment according to the sequence of searching for words and use it as the image to be recognized.

[0058] In some feasible implementations, each finger may also perform word search at the same time. If the word search time is the same, the computer device may obtain the image to be recognized at the same time.

[0059] In one embodiment, determining the word-searching order of the corresponding fingers and sequentially acquiring the images to be recognized according to the word-searching order includes: determining the time sequence of the word-searching actions of each finger; determining the word-searching order of the corresponding fingers according to the time sequence, and sequentially acquiring the images to be recognized according to the word-searching order.

[0060] In this embodiment, the computer device can determine the time corresponding to the finger's word-searching action and arrange the fingers in chronological order to determine the time sequence of each finger. For example, assuming that the time corresponding to the finger A's word-searching action is determined to be 10:00:00 and the time corresponding to the finger B's word-searching action is determined to be 10:00:05, the time sequence of the fingers can be determined to be finger A > finger B, and the word-searching order of the fingers can be further determined to be finger A > finger B, and the images to be recognized can be acquired in order of the word-searching order.

[0061] It should be noted that after the computer device obtains the images to be recognized in sequence according to the word search order, it can output the finger word search recognition results in sequence according to the word search order, so that multiple fingers can search for words in sequence. Compared with the existing technology that only considers the situation when searching for words with a single finger, the method shown in this application reconstructs the processing logic of multi-finger word search based on the sequence of finger word search, so that the accuracy of word search recognition in the application scenario of multi-finger word search is effectively improved.

[0062] S304: Preprocess the image to be recognized to obtain a text region recognition image.

[0063] In this embodiment, preprocessing of the image to be recognized may include steps such as text segmentation, text cropping, and text scaling. The text region recognition image may be the text image of the row (or column) or n rows (or columns) corresponding to the fingertip performing the word lookup action, where n is a positive integer greater than or equal to 2.

[0064] In one embodiment, the preprocessing of the image to be identified to obtain the text area recognition image includes: intercepting a target area corresponding to a fingertip in the image to be identified; and preprocessing the target area to obtain the text area recognition image.

[0065] It should be noted that the target region may be a ROI (region of interest).

[0066] For example, the computer device can extract a square (or other circular shape) based on the coordinates of the fingertip, with the coordinates as the center point, and use the square (or other circular shape) as the ROI area. The range of the ROI area can be preset by the user based on experience, or automatically determined by a text detection algorithm, wherein the text detection algorithm includes but is not limited to a text detection method based on deep learning, a text detection method based on connection proposals, etc.

[0067] In one embodiment, preprocessing the target area to obtain a text area recognition image includes:

[0068] S3041: Perform text segmentation on the target area to obtain a first text core map, a first text segmentation map, and other text areas.

[0069] In this embodiment, the text region recognition image may be a text region recognition image containing curved text, or may be a text region recognition image in a normal text state.

[0070] As a feasible implementation, the computer device can perform text segmentation on the image to be recognized based on a progressive expansion network (PSENet). PSENet is a text detection method based on semantic segmentation. By reducing the size of the original text line and then gradually expanding it, it can detect text of any shape and close distance. Based on the PSENet algorithm, text segmentation can be performed on the recognition image to obtain a first text kernel map, a first text segmentation map, and other text area maps corresponding to the fingertip.

[0071] Among them, the first text kernel map can be a text kernel map, and the kernel can represent the central area of ​​the text. It can be understood that a first text kernel map can contain multiple texts, which can also include multiple kernels. The first text kernel map can be composed according to the multiple kernels. The first text segmentation map can include a background area and a main text area. The other text area map refers to an area containing text other than the main text area. For example, the first text kernel map can be as shown in Figure 4, where the black part is the background area and the white part is the kernel area. The first text segmentation map can be as shown in Figure 5.

[0072] As a feasible implementation, the computer device can determine whether the area containing text is the primary text area or another text area based on the position of the fingertip. For example, the text closest to the fingertip is the primary text. Based on the text arrangement relationship or text association relationship, the text corresponding to n rows (or n columns, or other arrangements) of the primary text can be determined as the primary text area, where n is a positive integer greater than or equal to 1. Text areas other than the primary text area can be determined as another text area.

[0073] S3042: Set the other text areas as background areas, and update the first text segmentation map based on the setting result to obtain a second text segmentation map.

[0074] It should be noted that the second text segmentation graph refers to a text segmentation graph obtained by updating the first text segmentation graph.

[0075] In this embodiment, the other text area is set as the background area by erasing the text in the other text area and then setting the area to be consistent with the background area of ​​the first text segmentation image. This method can reduce the interference level of text recognition in the later stage and improve the accuracy and efficiency of text recognition.

[0076] S3043: Perform text smoothing processing based on the second text segmentation map and the first text kernel map to obtain a target text line map.

[0077] For example, the text smoothing method may be mean filtering.

[0078] In one embodiment, the text smoothing is performed based on the second text segmentation map and the first text kernel map to obtain a target text line map, including: obtaining the text line of the text to be recognized according to the first text kernel map, and determining the text height of the text to be recognized according to the second text segmentation map; obtaining an initial text line map according to the text line and the text height; smoothing the text height, and using the smoothed initial text line map as the target text line map.

[0079] For example, the computer device may obtain text lines of the text based on the first text kernel map, and determine the text height of each position in the text line based on the second text segmentation map, and form an initial text line map based on the text lines and text heights (for example, as shown in FIG6 ). Furthermore, the text heights in the initial text line map are smoothed (such as by mean filtering, etc.) to obtain the text heights of the smoothed text lines, and the initial text line map is updated based on this to obtain the target text line map.

[0080] S3044 , performing circle drawing processing on each position in the target text line graph in sequence, and constructing a third text segmentation graph according to the circle drawing processing results.

[0081] It should be noted that the circle drawing process referred to in this application refers to a process of encircling each text in a text line drawing with a closed figure. Among them, the closed figure is more commonly a circle, but in some feasible implementations, it can also be a square, an ellipse, an irregular polygon, etc.

[0082] In one embodiment, the method of sequentially drawing a circle for each position in the target text line graph and constructing a third text segmentation map based on the circle drawing results includes: sequentially drawing a circle with a preset diameter for each position in the target text line graph, and setting the image inside the circle as the foreground area, wherein the preset diameter is the text height at the corresponding position; and replacing the main text area in the second text segmentation map with the foreground area to generate a third text segmentation map.

[0083] For example, the computer device may draw a circle for each position from left to right, with the diameter of the circle being the height of the text at the corresponding position, and the entire interior of the circle being set as the foreground area. The main text area in the second text segmentation map is replaced with the foreground area to obtain a third text segmentation map. The third text segmentation map may be shown in FIG7 .

[0084] S3045: Obtain a text region recognition image according to the third text segmentation image.

[0085] For example, the text area recognition image may be shown in FIG8 . It can be seen that, compared with FIG2 , the main text area in FIG8 is clearer, and noise and interference are reduced.

[0086] S305: Perform text recognition on the text area recognition image, and output text recognition results pointed by the fingertip in sequence according to the word search order.

[0087] It can be understood that if the word search order of each finger is the same, the text recognition results can be output simultaneously.

[0088] As a feasible implementation, when outputting text recognition results, the computer device can differentiate the output results based on the order of word search. For example, the transparency of the output results with words that appear earlier in the search order can be reduced, the display color can be lighter, or the display range can be reduced, creating a gradient effect to allow users to more intuitively view the recognition results.

[0089] In one embodiment, the text recognition is performed on the text area recognition image, and the text recognition results pointed to by the fingertip are output in sequence according to the word search order, including: performing text recognition on the text area recognition image to obtain text recognition results of all texts in the text area recognition image; obtaining the text position of each text; determining the text position corresponding to the fingertip from the text positions of the each text, and outputting the text recognition results of the corresponding text positions in sequence according to the word search order.

[0090] For example, when performing text recognition, the computer device may first identify all text in the text area recognition image, and then output the text recognition result of the text pointed to by the fingertip position based on the fingertip position.

[0091] For another example, the text recognition result may include at least one or more of the text itself, the text interpretation, the text pinyin (or phonetic symbol), and the text search, and this application does not impose any restrictions on this.

[0092] It can be seen that through the multi-fingertip word-searching method shown in the embodiment of the present application, the computer device obtains the video to be detected. When it is detected that at least two fingers in the video to be detected have the word-searching action, the word-searching order of the corresponding fingers is determined, and the image to be recognized is obtained in sequence according to the word-searching order. After pre-processing the image to be recognized, a text area recognition image is obtained. Text recognition is performed based on the text area recognition image, and the recognition results are output in sequence according to the word-searching order. A multi-target processing logic is constructed when multiple fingers perform word-searching. At the same time, the text recognition process is optimized to make the text recognition process more suitable for the case of curved text, thereby ensuring the text recognition accuracy during multi-finger word-searching and improving the user's human-computer interaction experience.

[0093] It should be noted that, in each of the above-mentioned embodiments, there is not necessarily a certain order between the above-mentioned steps. A person skilled in the art can understand, based on the description of the embodiments of this application, that in different embodiments, the above-mentioned steps may have different execution orders, that is, they may be executed in parallel, or may be executed interchangeably, etc.

[0094] As another aspect of the present invention, an embodiment of the present invention provides a multi-finger word search device. The multi-finger word search device may be a software module comprising a plurality of instructions stored in a memory. A processor may access the memory and execute the instructions to implement the multi-finger word search method described in each of the above embodiments.

[0095] In some embodiments, the multi-finger word search device can also be constructed by hardware devices. For example, the multi-finger word search device can be constructed by one or more chips, and the chips can work in coordination with each other to complete the multi-finger word search method described in the above embodiments. For another example, the multi-finger word search device can also be constructed by various logic devices, such as a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a single-chip microcomputer, an ARM (Acorn RISC Machine) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination of these components.

[0096] Please refer to FIG9, which is a schematic diagram of the structure of a multi-finger word search device provided in an embodiment of the present application. The multi-finger word search device 90 shown in FIG9 includes:

[0097] The acquisition module 901 is used to acquire the video to be detected.

[0098] The palm detection module 902 is used to determine whether at least two fingers in the video to be detected have a word-looking action.

[0099] The palm detection module 902 is further configured to determine a word search order corresponding to the fingers, and sequentially obtain images to be recognized according to the word search order.

[0100] The text processing module 903 is configured to pre-process the image to be recognized to obtain a text region recognition image, and perform text recognition on the text region recognition image.

[0101] The output module 904 is used to output the text recognition results pointed by the fingertip in sequence according to the word search order.

[0102] In one possible design, before determining whether at least two fingers in the video to be detected are performing word-searching gestures, the palm detection module 902 is further configured to: perform palm detection on the video to be detected; if a palm image is present in the video to be detected and the palm is performing word-searching gestures, perform key point estimation on the palm; and determine the fingertip positions based on the key point estimation results.

[0103] In one possible design, when the palm detection module 902 is used to perform palm detection on the video to be detected, it is specifically used to: obtain each single-frame image in the video to be detected; identify whether the each single-frame image contains a palm image; if so, determine the matching relationship between the palms in the previous and next single-frame images based on a tracking algorithm, determine whether the palms in the video to be detected are the same palm according to the matching relationship; and determine the movement trajectory of the same palm.

[0104] In one possible design, the tracking algorithm is the DEEPSORT algorithm.

[0105] In one possible design, the palm detection module 902 is used to determine the fingertip position based on the key point estimation result, and is also used to: if the amplitude change of the fingertip position within a preset time range does not exceed a preset amplitude range, then determine that the finger has a word-looking action, and locate the fingertip position of the finger that has a word-looking action.

[0106] In one possible design, the palm detection module 902 is used to determine the word-searching order of the corresponding fingers, and to obtain the images to be recognized in sequence according to the word-searching order. Specifically, it is used to: determine the time sequence of the word-searching actions of each finger; determine the word-searching order of the corresponding fingers according to the time sequence, and to obtain the images to be recognized in sequence according to the word-searching order.

[0107] In one possible design, when the text processing module 903 is used to preprocess the image to be identified to obtain a text area recognition image, it is specifically used to: intercept the target area corresponding to the fingertip in the image to be identified; and obtain a text area recognition image after preprocessing the target area.

[0108] In a possible design, the target region is a ROI region.

[0109] In one possible design, when the text processing module 903 is used to pre-process the target area to obtain a text area recognition image, it is specifically used to: perform text segmentation on the target area to obtain a first text kernel map, a first text segmentation map, and other text area maps; set the other text area as the background area, and update the first text segmentation map based on the setting result to obtain a second text segmentation map, wherein the first text segmentation map includes the background area and the main text area; perform text smoothing based on the second text segmentation map and the first text kernel map to obtain a target text line map; circle each position in the target text line map in turn, and construct a third text segmentation map according to the circle drawing result; and obtain a text area recognition image based on the third text segmentation map.

[0110] In one possible design, the text processing module 903 is used to perform text smoothing based on the second text segmentation map and the first text kernel map, and when obtaining the target text line map, it is specifically used to: obtain the text line of the text to be recognized according to the first text kernel map, and determine the text height of the text to be recognized according to the second text segmentation map; obtain the initial text line map according to the text line and the text height; smooth the text height, and use the smoothed initial text line map as the target text line map.

[0111] In one possible design, the text processing module 903 is used to draw a circle for each position in the target text line graph in turn. When constructing the third text segmentation map based on the circle drawing results, it is specifically used to: draw a circle with a preset diameter for each position in the target text line graph in turn, and set the image inside the circle as the foreground area, and the preset diameter is the text height at the corresponding position; replace the main text area in the second text segmentation map with the foreground area to generate a third text segmentation map.

[0112] In one possible design, when the output module 9 is configured to sequentially output the text recognition results for the text pointed to by the fingertip in the search order, it is specifically configured to: perform text recognition on the text region recognition image to obtain text recognition results for all text in the text region recognition image; obtain the text position of each text; determine the text position corresponding to the fingertip from the text positions of each text, and sequentially output the text recognition results for the corresponding text positions in the search order.

[0113] It can be seen that the above-mentioned device obtains the video to be detected, and when it is detected that at least two fingers in the video to be detected have the action of looking up words, it determines the word-looking order of the corresponding fingers, and obtains the images to be recognized in sequence according to the word-looking order, pre-processes the images to be recognized to obtain the text area recognition image, performs text recognition based on the text area recognition image, and outputs the recognition results in sequence according to the word-looking order, realizes multi-target tracking when multiple fingers are looking up words, obtains the images to be recognized based on the finger states, and outputs the recognition results in sequence according to the word-looking order, ensures the text recognition accuracy when multiple fingers are looking up words, and improves the user's human-computer interaction experience.

[0114] It should be noted that the multi-finger word search device described above can execute the multi-finger word search method provided in the embodiments of this application, and has the corresponding functional modules and beneficial effects of the execution method. For technical details not fully described in the embodiments of the multi-finger word search device, please refer to the multi-finger word search method provided in the embodiments of this application.

[0115] 10 is a schematic diagram of the structure of a computer device according to an embodiment of the present application. The computer device 100 includes a camera 1001, one or more processors 1002, and a memory 1003. The memory 1003 is connected to the one or more processors, for example, via a bus.

[0116] The camera 1001 is used to obtain the video to be detected.

[0117] The processor 1002 is configured to support the computer device in executing the corresponding functions of the method in the above method embodiment. The processor can be a central processing unit (CPU), a network processor (NP), a hardware chip, or any combination thereof. The above hardware chip can be an application specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The above PLD can be a complex programmable logic device (CPLD), a field programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0118] Memory 1003 is used to store program code, etc. Memory 1003 may include volatile memory (VM), such as random access memory (RAM); non-volatile memory (NVM), such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid-state drive (SSD); or a combination of the aforementioned types of memory.

[0119] The memory 1003 can be used to store non-volatile software programs, non-volatile computer executable programs, and modules, such as the program instructions / modules corresponding to the multi-finger word search method in the embodiments of the present application. The processor executes the non-volatile software programs, instructions, and modules stored in the memory to perform various functional applications and data processing of the multi-finger word search method and the multi-finger word search device, thereby realizing the functions of the multi-finger word search method and the various modules or units of the multi-finger word search device provided in the above method embodiments.

[0120] The memory may include a program storage area and a data storage area. The program storage area may store an operating system and application programs required for at least one function. The data storage area may store data generated based on the use of the multi-finger word search device. In some embodiments, the memory may optionally include a memory remote from the processor, and such remote memory may be connected to the multi-finger word search device via a network. Examples of such networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0121] The one or more modules are stored in the memory, and when executed by the one or more processors, the multi-finger word search method in any of the above method embodiments is executed, for example, the method steps described in the above method embodiments are executed to realize the functions of the modules described in the above device embodiments.

[0122] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a computer, the computer executes the method as described in the above embodiment.

[0123] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing related hardware through a computer program. The program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0124] The above disclosure is only a preferred embodiment of the present application, and certainly cannot be used to limit the scope of rights of the present application. Therefore, equivalent changes made according to the claims of the present application are still within the scope covered by the present application.

Claims

1. A multi-finger word lookup method, characterized in that, The method includes: Obtain the video to be detected; Determine that there are at least two fingers making word query actions in the video to be detected; Determine the word query order of the corresponding fingers, and sequentially obtain the images to be recognized according to the word query order; Preprocess the images to be recognized to obtain text region recognition images; Perform text recognition on the text region recognition images, and sequentially output the text recognition results pointed to by the fingertips of the fingers according to the word query order.

2. The method according to claim 1, characterized in that, Before determining that there are at least two fingers making word query actions in the video to be detected, it further includes: Perform palm detection on the video to be detected; If there is a palm image in the video to be detected and the palm has a word query gesture, perform key point estimation on the palm; Determine the fingertip positions according to the key point estimation results.

3. The method according to claim 2, characterized in that, The performing palm detection on the video to be detected includes: Obtain each single-frame picture in the video to be detected; Identify whether the palm image is included in each single-frame picture; If included, determine the matching relationship of the palms in the front and back single-frame pictures based on the tracking algorithm, and determine whether the palms in the video to be detected are the same palms according to the matching relationship; Determine the motion trajectory of the same palm.

4. The method according to claim 3, characterized in that, The tracking algorithm is the DEEPSORT algorithm.

5. The method according to any one of claims 2 to 4, characterized in that, After determining the fingertip positions according to the key point estimation results, it further includes: If the amplitude change of the fingertip positions within a preset time range does not exceed the preset amplitude range, determine that the finger has a word query action, and locate the fingertip positions of the fingers with word query actions.

6. The method according to claim 1, characterized in that, The determining the word query order of the corresponding fingers and sequentially obtaining the images to be recognized according to the word query order includes: Determine the time sequence of each finger having a word query action; Determine the word query order of the corresponding fingers according to the time sequence, and sequentially obtain the images to be recognized according to the word query order.

7. The method according to claim 1, characterized in that, The preprocessing the images to be recognized to obtain text region recognition images includes: Intercept the target region corresponding to the fingertip in the image to be recognized; Preprocess the target region to obtain a text region recognition image.

8. The method according to claim 7, characterized in that, The target region is the ROI region.

9. The method according to claim 7 or 8, characterized in that, The preprocessing the target region to obtain a text region recognition image includes: Perform text segmentation on the target region to obtain a first text kernel map, a first text segmentation map, and other text regions; Set the other text regions as background regions, and update the first text segmentation map based on the setting result to obtain a second text segmentation map, where the first text segmentation map includes the background region and the main text region; Perform text smoothing processing based on the second text segmentation map and the first text kernel map to obtain a target text line map; Perform circle drawing processing on each position in the target text line map in sequence, and construct a third text segmentation map according to the circle drawing processing results; Obtain a text region recognition image according to the third text segmentation map.

10. The method according to claim 9, wherein, The performing text smoothing processing based on the second text segmentation map and the first text kernel map to obtain a target text line map includes: Obtain the text line of the text to be recognized according to the first text kernel map, and determine the text height of the text to be recognized according to the second text segmentation map; Obtain an initial text line map based on the text line and the text height; Perform smoothing processing on the text height, and use the smoothed initial text line map as the target text line map.

11. The method according to claim 9, wherein, Successively perform circle drawing processing on each position in the target text line map, and construct a third text segmentation map according to the circle drawing processing result, including: Successively perform circle drawing processing on each position in the target text line map with a preset diameter, and set the image inside the circle as the foreground area, and the preset diameter is the text height of the corresponding position; Replace the main text area in the second text segmentation map with the foreground area to generate a third text segmentation map.

12. The method according to claim 1, wherein, Perform text recognition on the text area recognition image, and sequentially output the text recognition results pointed to by the fingertips of the fingers according to the word query order, including: Perform text recognition on the text area recognition image to obtain the text recognition results of all texts in the text area recognition image; Obtain the text positions of each text; Determine the text position corresponding to the fingertips of the fingers from the text positions of each text, and sequentially output the text recognition results of the corresponding text positions according to the word query order.

13. A multi - finger word - lookup device, wherein, The multi-finger word query device includes: An acquisition module for acquiring a video to be detected; A palm detection module for determining that there are at least two fingers in the video to be detected performing word query actions; The palm detection module is further configured to determine the word query order of the corresponding fingers, and sequentially acquire the images to be recognized according to the word query order; A text processing module for preprocessing the image to be recognized to obtain a text area recognition image, and performing text recognition on the text area recognition image; An output module for sequentially outputting the text recognition results pointed to by the fingertips of the fingers according to the word query order.

14. A computer device, wherein, Including a camera, a memory and a processor, the camera is used to acquire a video to be detected, the memory and the camera are connected to the processor, and the processor is used to execute one or more computer programs stored in the memory. When the processor executes the one or more computer programs, the computer device realizes the method according to any one of claims 1-12.

15. A computer - readable storage medium, wherein, The computer-readable storage medium stores a computer program, the computer program includes program instructions, and when the program instructions are executed by the processor, the processor executes the method according to any one of claims 1-12.

Citation Information

Patent Citations

  • Test question shooting method and device with multiple designated positions, electronic equipment and storage medium

    CN111711758A

  • Intent detection with computing device

    CN111736702A

  • Text content recognition method and device, computer equipment and storage medium

    CN115131693A

  • Word searching method and device based on gestures and computer readable storage medium

    CN115273220A

  • Text recognition method, electronic equipment and storage medium

    CN116824588A