Test method, device and equipment based on machine vision and real-time text recognition

By using machine vision and real-time text recognition testing methods, the characters scanned by the dictionary pen are automatically compared with the sample characters, solving the problem of relying on manual verification for dictionary pen testing and improving testing efficiency and accuracy.

CN117011851BActive Publication Date: 2026-02-10EEASY TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310820823.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-05
Publication Date
2026-02-10
Estimated Expiration
2043-07-05

AI Technical Summary

Technical Problem

In the current technology, the testing of dictionary pens relies on manual verification, which is inefficient and costly. Furthermore, the diverse test samples require a high level of knowledge from personnel, making verification errors prone to occur.

Method used

The test method, based on machine vision and real-time text recognition, extracts a sample dictionary from the sample text image using a first camera, obtains the start and end positions of the scanning of the dictionary pen under test, and compares the scanned character set with the sample character set to automatically determine the test result.

Benefits of technology

It eliminates the need for manual verification, improves testing efficiency and accuracy, simplifies the test preparation process, and enhances the flexibility and automation of testing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117011851B_ABST
    Figure CN117011851B_ABST
Patent Text Reader

Abstract

The application provides a test method, device and equipment based on machine vision and real-time text recognition, which comprises the following steps: extracting words from a first image of sample text to obtain a sample dictionary through a first camera; obtaining a test video stream through the first camera, determining the start and end positions of a scanned word dictionary pen according to the test video stream; determining a sample character set from the sample dictionary according to the start and end positions, wherein the sample character set comprises at least one sample character; obtaining a scanned character set scanned by the second camera, wherein the scanned character set comprises at least one scanned character; and determining a target test result by comparing the scanned character set and the sample character set. According to the technical scheme of the embodiment of the application, sample information can be obtained through machine vision as a content reference for testing, and the test result can be automatically determined by comparing the characters scanned by the scanned word dictionary pen, so that manual checking is not needed, and the test efficiency and accuracy are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image recognition testing technology, and in particular to a testing method, apparatus, and equipment based on machine vision and real-time text recognition. Background Technology

[0002] Currently, image recognition technology has been applied to dictionary pens. By using a camera embedded in the pen, text content is extracted from a scanned area and then read aloud. Before leaving the factory, dictionary pens undergo reliability testing to ensure accurate recognition of the extracted content. In related technologies, dictionary pen testing primarily relies on manual verification. Furthermore, test samples need to cover text characters from Chinese and English textbooks, as well as various languages ​​and scripts. The test texts are highly diverse, requiring a high level of knowledge from the testers. Manual verification is also inefficient, costly, and prone to errors. Summary of the Invention

[0003] This invention aims to at least solve one of the technical problems existing in the prior art. To this end, this invention proposes a testing method, apparatus, and device based on machine vision and real-time text recognition, which can automatically complete the verification of text recognition and improve the testing efficiency of dictionary pens.

[0004] In a first aspect, embodiments of the present invention provide a testing method based on machine vision and real-time text recognition, applied to a testing device. The testing device is communicatively connected to a dictionary pen under test. The testing device includes a first camera, and the dictionary pen under test is equipped with a second camera. The testing method based on machine vision and real-time text recognition includes:

[0005] The first image of the sample text is obtained by extracting text from the first image using the first camera to obtain a sample dictionary. The sample dictionary includes multiple sample characters recorded in the sample text and the character position of each sample character.

[0006] The test video stream is acquired through the first camera, and the scanning start and end positions of the dictionary pen under test are determined based on the test video stream, wherein the test video stream records the scanning process of the dictionary pen under test on the sample text;

[0007] A sample character set is determined from the sample dictionary based on the scan start and end positions, the sample character set including at least one of the sample characters;

[0008] The dictionary pen under test obtains a scanned character set by scanning through the second camera, the scanned character set including at least one scanned character;

[0009] The target test result is determined by comparing the scanned character set and the sample character set.

[0010] According to some embodiments of the present invention, the character position is the character sample coordinate, and the step of extracting text from the first image to obtain a sample dictionary includes:

[0011] Perform text recognition in the text region of the first image, and acquire and record each sample character obtained from the text recognition;

[0012] A coordinate system is established based on the first image to determine the character sample coordinates of each sample character;

[0013] The sample dictionary is generated based on all the sample characters and the corresponding coordinates of the sample characters.

[0014] According to some embodiments of the present invention, the scan start and end positions include scan start coordinates and scan stop coordinates, and determining the scan start and end positions of the dictionary pen under test based on the test video stream includes:

[0015] The scan start time and scan stop time of the dictionary pen under test are determined based on the test video stream;

[0016] A second image is acquired from the test video stream based on the scan start time, and a third image is acquired from the test video stream based on the scan stop time;

[0017] Pen tip recognition is performed on the second image and the third image respectively. The pen tip coordinates identified in the second image are determined as the scan start coordinates, and the pen tip coordinates identified in the third image are determined as the scan stop coordinates.

[0018] According to some embodiments of the present invention, the shooting direction of the first camera is perpendicular to the sample text, and the dictionary pen under test is further provided with an indicator light. The step of determining the scanning start time and scanning stop time of the dictionary pen under test based on the test video stream includes:

[0019] The lighting and turning-off times of the indicator lights are detected from the test video stream. The lighting time is determined as the scan start time, and the turning-off time is determined as the scan stop time.

[0020] Alternatively, distance detection can be performed on the image frames of the test video stream to determine the target distance between the dictionary pen under test and the sample file in the image frames. The time corresponding to the first detection of the target distance being less than a preset distance threshold in the image frame is determined as the scan start time, and the time corresponding to the first detection of the target distance being greater than the preset distance threshold in the image frame after the scan start time is determined as the scan stop time.

[0021] According to some embodiments of the present invention, determining the target test result by comparing the scanned character set and the sample character set includes:

[0022] The character scan coordinates corresponding to each scanned character are determined based on the test video stream;

[0023] Based on the principle of pairing the character scan coordinates, multiple comparison groups are determined, and each comparison group includes one sample character and one scan character.

[0024] Determine the character alignment result for each alignment group, and obtain the target test result based on all the character alignment results.

[0025] According to some embodiments of the present invention, obtaining the target test result based on all the character comparison results includes:

[0026] When at least one of the character comparison results indicates a text error, the character scan coordinates corresponding to the comparison group indicating the text error are determined as abnormal coordinates, and the target test result is obtained based on all the character comparison results and the abnormal coordinates.

[0027] Alternatively, when the number of characters in the scanned character set is greater than the number of characters in the sample character set, abnormal characters and abnormal position information are determined based on the scanned character set, and the target test result is obtained based on all the character comparison results, the abnormal characters, and the abnormal position information. The abnormal characters are the scanned characters located between two adjacent comparison groups, and the position information of the abnormal characters is determined based on the character sample coordinates corresponding to the two adjacent comparison groups.

[0028] Alternatively, when the number of characters in the scanned character set is less than the number of characters in the sample character set, the missing characters and missing coordinates are determined based on the sample character set, and the target test result is obtained based on all the character comparison results, the missing characters, and the missing coordinates, wherein the missing characters are the sample characters that are not paired with the scanned characters, and the missing coordinates are the character sample coordinates corresponding to the missing characters.

[0029] According to some embodiments of the present invention, the method further includes:

[0030] When multiple scan start and end positions are determined based on the test video stream, scan data is obtained from the dictionary pen under test, and the scan data includes multiple scanned characters;

[0031] Based on each pair of scan start times and scan stop times, the scan data is divided into multiple scan character sets;

[0032] Multiple sample character sets are obtained from the sample dictionary according to each scan start and end position, and the scan character set corresponding to each sample character set is determined according to the time order;

[0033] The grouping test results are determined based on the comparison of the paired scanned character sets and the sample character sets;

[0034] The target test result is determined based on all the group test results.

[0035] Secondly, embodiments of the present invention provide a testing apparatus based on machine vision and real-time text recognition, including at least one control processor and a memory for communicatively connecting to the at least one control processor; the memory stores instructions executable by the at least one control processor, which, when executed by the at least one control processor, enables the at least one control processor to perform the testing method based on machine vision and real-time text recognition as described in the first aspect above.

[0036] Thirdly, embodiments of the present invention provide a testing device, including a testing apparatus based on machine vision and real-time text recognition as described in the second aspect above.

[0037] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing computer-executable instructions for performing the test method based on machine vision and real-time text recognition as described in the first aspect above.

[0038] The testing method based on machine vision and real-time text recognition according to embodiments of the present invention has at least the following beneficial effects: A sample dictionary is obtained by extracting text from a first image of sample text using a first camera. The sample dictionary includes multiple sample characters recorded in the sample text and the position of each sample character. A test video stream is acquired using the first camera, and the start and end positions of the scanning of the dictionary pen under test are determined based on the test video stream, wherein the test video stream records the scanning process of the dictionary pen under test on the sample text. A sample character set is determined from the sample dictionary based on the start and end positions of the scanning, and the sample character set includes at least one sample character. A scanned character set obtained by the dictionary pen under test through a second camera is acquired, and the scanned character set includes at least one scanned character. The target test result is determined by comparing the scanned character set and the sample character set. According to the technical solution of the present invention, sample information can be obtained through machine vision as a reference for the test content, and the test result can be automatically determined by comparing it with the characters scanned by the dictionary pen under test, without the need for manual verification, thus improving testing efficiency and accuracy. Attached Figure Description

[0039] Figure 1This is a schematic diagram of an implementation environment provided in one embodiment of the present invention;

[0040] Figure 2 This is a flowchart of a testing method based on machine vision and real-time text recognition provided in one embodiment of the present invention;

[0041] Figure 3 This is a flowchart of generating a sample dictionary provided in another embodiment of the present invention;

[0042] Figure 4 This is a flowchart for determining the start and end positions of a scan, provided in another embodiment of the present invention;

[0043] Figure 5 This is a flowchart for determining the start and end times of a scan, provided in another embodiment of the present invention;

[0044] Figure 6 This is a flowchart of character comparison provided in another embodiment of the present invention;

[0045] Figure 7 This is a flowchart of a positioning test error provided in another embodiment of the present invention;

[0046] Figure 8 This is a flowchart of performing multiple sets of tests provided in another embodiment of the present invention;

[0047] Figure 9 This is a structural diagram of a testing device based on machine vision and real-time text recognition provided in another embodiment of the present invention. Detailed Implementation

[0048] Embodiments of the present invention are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and should not be construed as limiting the present invention.

[0049] In the description of this invention, it should be understood that the orientation descriptions, such as up, down, front, back, left, right, etc., are based on the orientation or positional relationship shown in the accompanying drawings. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limiting this invention.

[0050] In the description of this invention, "several" means one or more, "more than" means two or more, "greater than," "less than," and "exceeding" are understood to exclude the stated number, while "above," "below," and "within" are understood to include the stated number. The use of "first" and "second" in the description is merely for distinguishing technical features and should not be construed as indicating or implying relative importance, or implicitly indicating the number of indicated technical features, or implicitly indicating the order of the indicated technical features.

[0051] In the description of this invention, unless otherwise explicitly defined, terms such as "set up," "install," and "connect" should be interpreted broadly, and those skilled in the art can reasonably determine the specific meaning of the above terms in this invention in conjunction with the specific content of the technical solution.

[0052] This invention provides a testing method, apparatus, and device based on machine vision and real-time text recognition. The testing method includes: extracting text from a first image of sample text using a first camera to obtain a sample dictionary, the sample dictionary including multiple sample characters recorded in the sample text and the position of each sample character; acquiring a test video stream using the first camera, and determining the start and end positions of the scanning of the dictionary pen under test based on the test video stream, wherein the test video stream records the scanning process of the dictionary pen under test on the sample text; determining a sample character set from the sample dictionary based on the scan start and end positions, the sample character set including at least one sample character; acquiring a scanned character set obtained by the dictionary pen under test through a second camera, the scanned character set including at least one scanned character; and determining the target test result by comparing the scanned character set and the sample character set. According to the technical solution of this invention, sample information can be obtained through machine vision as a test content reference, and the test result can be automatically determined by comparing it with the characters scanned by the dictionary pen under test, eliminating the need for manual verification and improving testing efficiency and accuracy.

[0053] First, an exemplary implementation environment for the present invention will be described. This example is not intended to limit the various structures used for testing, but rather to describe a specific implementation environment in which the technical solution of the present invention can be executed. (Refer to...) Figure 1 , Figure 1 The schematic diagram of the implementation environment provided by the present invention includes a test device 10, a dictionary pen under test 20, and a sample text 30.

[0054] For example, the test equipment 10 can adopt such as Figure 1The document scanner shown has a first camera 11 horizontally positioned on the upper side of the testing device 10. Sample text 30 is placed horizontally on a platform below the testing device 10. If the first camera 11 captures the sample text 30 at an angle, the characters in the first image will be distorted, increasing the complexity of text recognition and coordinate positioning. Therefore, in this embodiment, the first camera 11 is positioned horizontally with its shooting angle perpendicular to the platform, allowing it to capture the sample text 30 from a top-down perspective, improving the standardization of the content in the first image and increasing the efficiency of image recognition. The testing device 10 also includes a processing device 12, which can be a common computer or server capable of image recognition and data storage.

[0055] For example, the sample text 30 may be a paper book, a printed document, or an electronic document displayed through an electronic terminal. This embodiment does not limit the specific form.

[0056] For example, the dictionary pen under test 20 includes a second camera 21 and an indicator light 22. When the dictionary pen under test 20 is started, the indicator light 22 is lit. The indicator light 22 is used to provide supplementary light for the camera 21. The second camera 21 acquires images in real time and extracts text information from the images. When it passes over the sample text 30, it can achieve a scanning effect and extract the scanned characters from the scanning path. After the dictionary pen under test 20 stops working, the indicator light 22 is turned off and the second camera 21 switches to standby mode.

[0057] For example, to improve the stability of data transmission, the dictionary pen under test 20 can communicate with... (The sentence is incomplete and requires more context to translate accurately.)

[0058] It is worth noting that the dictionary pen 20 under test usually scans text line by line. Therefore, those skilled in the art are motivated to adjust the shooting range of the second camera 21 according to actual needs to ensure that multiple lines of text or less than one line of text are not scanned. The configuration of the second camera 21 will not be described in detail in this embodiment.

[0059] The following is based on the appendix Figure 1 The implementation environment shown will be used to further illustrate the control method of this embodiment of the invention.

[0060] Reference Figure 2 , Figure 2 The flowchart illustrates a testing method based on machine vision and real-time text recognition, as provided in this embodiment of the invention. The testing method includes, but is not limited to, the following steps:

[0061] S21, using the first camera to extract text from the first image of the sample text to obtain a sample dictionary, the sample dictionary includes multiple sample characters recorded in the sample text and the character position of each sample character;

[0062] S22, acquire the test video stream through the first camera, and determine the scanning start and end positions of the dictionary pen under test based on the test video stream. The test video stream records the scanning process of the dictionary pen under test on the sample text.

[0063] S23, determine the sample character set from the sample dictionary based on the scan start and end positions, the sample character set including at least one sample character;

[0064] S24, Obtain the scanned character set obtained by the dictionary pen under test through the second camera, the scanned character set includes at least one scanned character;

[0065] S25, determine the target test result by comparing the scanned character set and the sample character set.

[0066] It should be noted that the testing of the dictionary pen relies on the comparison between the scanned content and the actual content. In related technologies, manual verification is inefficient. Some technologies use electronic devices for automatic comparison, but to obtain sample text for reference, testers need to input the content of paper books into electronic documents or download data from the internet and organize it into a preset format. Then, the sample text is pre-entered into the relevant electronic device. In other words, these technologies require a significant amount of preprocessing work to obtain the actual content for comparison. This embodiment utilizes machine vision to automatically acquire and extract the sample text. A first image of the sample text is captured by a first camera, and each character of the sample text can be extracted through simple image recognition. Furthermore, each sample character can be located using simple positioning technology to establish a sample dictionary. Using the sample dictionary as the actual content carrier, after acquiring the scanned content, the content at the corresponding position in the sample dictionary can be retrieved for comparison, effectively improving testing efficiency.

[0067] It should be noted that the first image can be an image of the entire sample text or an image of a specific area. For example, taking a book as an example, a page of the sample text includes both text and illustration areas. Since dictionary pens are typically used for text recognition, after capturing the image with the first camera, the text area can be cropped to obtain the first image. Distinguishing between text and illustration areas is a common technique in image recognition and will not be elaborated upon here. Alternatively, if a page of the sample text consists entirely of text, the entire page can be photographed to obtain the first image. Those skilled in the art will have the incentive to adjust the first image acquisition strategy according to the actual form of the sample text, and will not be limited here.

[0068] It should be noted that after text extraction through image recognition, the processing device will only store a set of multiple sample characters. Since the characters scanned by the dictionary pen may be identical (e.g., multiple identical characters in the same line), the sample characters cannot be located simply by text comparison. It is necessary to ensure the uniqueness of each sample character. Therefore, in this embodiment, after recognizing the sample characters in the first image, the position information of each sample character is also obtained. The position information is then associated with the sample characters to ensure the uniqueness of each element in the generated sample dictionary, thus ensuring the accuracy of the comparison. It is understood that the position information can be in the form of coordinates, such as establishing a coordinate system for the first image to determine the coordinates of each character; the position information can also be in the form of relative positions, such as using row and column information as the position information of each character. This embodiment does not limit the specific form of the position information.

[0069] It should be noted that the test video stream can be started after the sample dictionary is obtained from the first image. To ensure that the test video stream can include the complete scanning process of the dictionary pen under test, the test video stream can be started first, and then the dictionary pen under test can be operated to scan. This allows the start and end positions of the scanning of the dictionary pen under test to be located based on the test video stream. With the sample dictionary having position information, the sample character corresponding to each scanned character can be determined, ensuring the accuracy of character comparison and avoiding comparison with incorrect sample characters, which would affect the test results.

[0070] It should be noted that the start and end positions of the scan are not limited to the same line. The dictionary pen under test can scan multiple lines of characters in the sample text, and the start and end positions of the scan can be determined by the start and end positions of the scan. This not only improves the flexibility of the test, but also improves the test results by scanning more characters.

[0071] It is understandable that the technique of extracting scanned characters by a dictionary pen equipped with a second camera during the scanning process is well known to those skilled in the art, and the specific text extraction techniques of the scanned character set will not be elaborated on here.

[0072] It should be noted that after obtaining the scanned character set, since the start and end positions of the scan and the position of each character in the first image are known, the position information of each scanned character can be determined. By comparing the scanned characters at the same position with the sample characters, it can be determined whether the scanned characters are correct, thereby obtaining the target test result of this scan test.

[0073] The technical solution of this embodiment enables the automatic generation of a sample dictionary based on machine vision-captured first images, eliminating the need for testers to perform preprocessing tasks such as inputting electronic documents into the test equipment, thus improving the efficiency of preparing basic test data. By locating the real-time scanning position of the dictionary pen under test based on machine vision, the scanning range can be arbitrarily determined, increasing the flexibility of the test. By matching the sample dictionary with the corresponding sample character set from the scanned character set, automatic text comparison is achieved, enabling real-time acquisition of test results for each scan, thereby improving test efficiency and accuracy.

[0074] In another embodiment, the character position is the character sample coordinate, referring to... Figure 3 , Figure 2 Step S21 in the illustrated embodiment also includes, but is not limited to, the following steps:

[0075] S31, perform text recognition in the text region of the first image, and obtain and record each sample character obtained from the text recognition;

[0076] S32, Establish a coordinate system based on the first image and determine the character sample coordinates of each sample character;

[0077] S33, Generate a sample dictionary based on all sample characters and their corresponding character sample coordinates.

[0078] It should be noted that, according to Figure 2 The description of the illustrated embodiment, which identifies text regions, is a conventional technique in the art and will not be elaborated upon here.

[0079] It should be noted that this embodiment establishes a coordinate system based on the first image, and uses the coordinate values ​​as the position information of each sample character, which can improve the positioning accuracy. For example, the coordinates of the center point or a certain corner of each sample character can be used as the character sample coordinates. This embodiment does not impose any limitations on this, as long as the position of each sample character can be represented by the character sample coordinates.

[0080] It should be noted that after obtaining the coordinates of the character samples, a mapping relationship between each sample character and the character samples is established and saved to the sample dictionary. This allows the sample characters with the same position as the scanned character to be retrieved from the sample dictionary by coordinate matching after the scanned character set is obtained, thereby improving the accuracy of text comparison and testing.

[0081] In another embodiment, the scan start and end positions include scan start coordinates and scan stop coordinates, as shown in the reference. Figure 4 , Figure 2 Step S22 of the illustrated embodiment also includes, but is not limited to, the following steps:

[0082] S41, Determine the scan start time and scan stop time of the dictionary pen under test based on the test video stream;

[0083] S42. Obtain a second image from the test video stream according to the scanning start time, and obtain a third image from the test video stream according to the scanning stop time;

[0084] S43. Perform pen tip recognition on the second image and the third image respectively, determine the pen tip coordinates recognized in the second image as the scanning start coordinates, and determine the pen tip coordinates recognized in the third image as the scanning stop coordinates.

[0085] It should be noted that since the test video stream records the complete process of scanning by the tested dictionary pen, the scanning start time and the scanning stop time can be determined according to the test video stream. When the time is determined, the image at the scanning start time is obtained from the test video stream as the second image, and the image at the scanning stop time is obtained as the third image. The scanning start and stop positions can be determined by recognizing the second image and the third image, so as to automatically determine the scanning start and stop positions.

[0086] It should be noted that after obtaining the second image and the third image, the pen tip positions of the tested dictionary pen can be determined to correspond to the scanning start and end positions respectively. Since the above embodiments use coordinates to represent the position information, the start and stop positions can be determined by performing target detection on the pen tip in this embodiment. The algorithm for performing target detection on a pen tip with a specific shape is well-known to those skilled in the art and will not be elaborated here.

[0087] It should be noted that since the main test objective of the tested dictionary pen is to judge whether the character recognition is correct, the scanning direction usually conforms to the reading direction, that is, from left to right, and after completing a line, scan again from the next line. Therefore, after determining the scanning start coordinates and the scanning stop coordinates, the sample characters in the sample dictionary with coordinate values within this interval can be extracted into the sample character set. For example, the scanning start coordinates are (a, b), the scanning stop coordinates are (c, d), and the character sample coordinates (x, y) of the sample character set satisfy: y = b, x ≥ a, or satisfy b < y < d, x ∈ N, where N is the number of characters in each line; or satisfy y = d, x ≤ c. Of course, the scanning direction can also be adjusted according to actual needs. After determining the scanning start point coordinates and the scanning stop coordinates, the scanning direction can be determined by a common trajectory recognition method, and the character scanning coordinates of the scanned characters can be determined in turn according to the scanning direction, and then compared with the corresponding sample characters.

[0088] It should be noted that since there will be a certain positional deviation in each scanning action, the coordinate matching in this embodiment can adopt the nearest matching method. For example, the scanning start coordinate can be the coordinate of the sample character closest to the pen tip position. After obtaining the character scanning coordinate, the character sample coordinate closest to the character scanning coordinate is determined as the paired sample character. This will not be repeated later.

[0089] Additionally, in one embodiment, reference is made to Figure 5 , Figure 4 Step S41 in the illustrated embodiment also includes, but is not limited to, the following steps:

[0090] S51, detect the light-on and light-off times of the indicator lights from the test video stream, determine the light-on time as the scan start time, and determine the light-off time as the scan stop time;

[0091] S52, perform distance detection on the image frames of the test video stream, determine the target distance between the dictionary pen under test and the sample file in the image frame, determine the time corresponding to the first image frame in which the target distance is less than the preset distance threshold as the scan start time, and determine the time corresponding to the first image frame in which the target distance is greater than the preset distance threshold after the scan start time as the scan stop time.

[0092] It should be noted that the reference Figure 1 As shown, the dictionary pen 20 under test typically has an indicator light 22 near the camera to provide supplementary lighting. The light illuminates when scanning begins and turns off after scanning is complete. Therefore, the moment the light illuminates can be determined from the test video stream as the start time of scanning, and the moment the light turns off can be determined as the stop time. Knowing these times, corresponding image frames are extracted from the video stream for image recognition, thus automatically determining the start and end times without requiring manual input from the tester, improving automation and accuracy.

[0093] It should be noted that for dictionary pens without indicator lights, the start and end times can be determined by measuring the distance between the measured location and the sample file in step S52. Target detection and distance recognition of two objects in the image are techniques well known to those skilled in the art, and will not be elaborated upon here.

[0094] Additionally, in one embodiment, reference is made to Figure 6 , Figure 2 Step S25 of the illustrated embodiment also includes, but is not limited to, the following steps:

[0095] S61, determine the character scanning coordinates corresponding to each scanned character based on the test video stream;

[0096] S62, based on the principle of pairing character scan coordinates, multiple comparison groups are determined, each comparison group including a sample character and a scan character;

[0097] S63, determine the character alignment results for each alignment group, and obtain the target test result based on all character alignment results.

[0098] It should be noted that the character scanning coordinates of each scanned character can be determined by the character sample coordinates of the sample characters between the start and end positions as described in the above embodiments, or by detecting the position of each character individually from the test video stream. With an image available, those skilled in the art will be motivated to choose the method of determining the character scanning coordinates according to actual needs, and no further limitations will be made here.

[0099] It should be noted that after determining the character scanning coordinates, the sample characters corresponding to the coordinates of the matching or closest character samples can be paired to obtain comparison groups. Based on each comparison group, the character comparison is performed to determine whether the scanned character is correct, and the target test result is obtained, thereby realizing automated testing without the need for manual verification by testers.

[0100] Additionally, in one embodiment, reference is made to Figure 7 , Figure 2 Step S25 of the illustrated embodiment also includes, but is not limited to, the following steps:

[0101] S71, when at least one character comparison result indicates a text error, the character scan coordinates corresponding to the comparison group whose character comparison result indicates a text error are determined as abnormal coordinates, and the target test result is obtained based on all character comparison results and abnormal coordinates;

[0102] S72, when the number of characters in the scanned character set is greater than the number of characters in the sample character set, abnormal characters and abnormal position information are determined based on the scanned character set, and the target test result is obtained based on all character comparison results, abnormal characters and abnormal position information. Among them, abnormal characters are scanned characters located between two adjacent comparison groups, and the position information of abnormal characters is determined based on the coordinates of the character samples corresponding to the two adjacent comparison groups.

[0103] S73, when the number of characters in the scanned character set is less than the number of characters in the sample character set, the missing characters and missing coordinates are determined according to the sample character set. The target test result is obtained based on all character comparison results, missing characters and missing coordinates. Among them, the missing characters are the sample characters of the unpaired scanned characters, and the missing coordinates are the character sample coordinates corresponding to the missing characters.

[0104] It should be noted that when errors are found in the scanned characters, alignment and localization are necessary to facilitate the subsequent output of test analysis reports. Errors encountered during scanning can include character recognition errors, extra characters, missing characters, etc. Related technologies primarily rely on manual verification; this embodiment uses an automatic comparison method for judgment, as detailed below:

[0105] For example, when the comparison result indicates a text error, that is, the scanned character is different from the sample character, the scanned character can be determined to be an erroneous character, and its coordinates can be recorded as abnormal coordinates in the target test result. Testers can locate the scanned character and the sample character based on the abnormal coordinates and conduct subsequent analysis on the scanning error.

[0106] For example, when the number of characters in the scanned character set is greater than the number of characters in the sample character set, it can be determined that an error of multiple characters has occurred during the scanning process. This is likely due to the scanning speed being too fast or the font distance between two adjacent sample characters being too close. Abnormal characters are usually located between two correct scanned characters. In this embodiment, the comparison group is determined by coordinate matching. Therefore, the correct scanned characters to the left and right of the extra abnormal character will match the corresponding sample characters. A one-to-one correspondence principle can be adopted to prevent the extra abnormal character from matching the corresponding sample character, thereby realizing the identification of abnormal characters. The coordinates of the characters in the two adjacent comparison groups are used to determine their positions, allowing testers to determine the cause of the extra characters from the sample characters.

[0107] For example, if the number of characters in the scanned character set is less than the number of characters in the sample character set, it can be determined that a missing character error has occurred during the scanning process. For example, the blurry handwriting of the sample character may cause recognition failure. Therefore, referring to the above description, since the principle of one-to-one correspondence of coordinates is adopted, there may be cases where the sample character cannot be matched with the scanned character. This character can be identified as a missing character, and the coordinates of the corresponding sample character are determined as the missing coordinates. They are recorded together in the target test results so that the tester can determine the cause of the omission.

[0108] Additionally, in one embodiment, reference is made to Figure 8 The method in this embodiment also includes, but is not limited to, the following steps:

[0109] S81, determine multiple scan start and end positions based on the test video stream, and obtain scan data from the dictionary pen under test. The scan data includes multiple scanned characters.

[0110] S82 divides the scan data into multiple scan character sets based on each pair of scan start and scan stop times;

[0111] S83, obtain multiple sample character sets from the sample dictionary according to each scan start and end position, and determine the scan character set corresponding to each sample character set according to the time order;

[0112] S84, determine the group test results based on the comparison of the paired scan character sets and sample character sets;

[0113] S85, determine the target test result based on all group test results.

[0114] It should be noted that, in order to improve the accuracy of the test, multiple tests can be performed on the sample text, such as multiple scans. According to the description of the above embodiment, the scan start time and scan stop time can be determined by the identification of the indicator light or the distance threshold. Therefore, if multiple tests are performed, multiple scan start times and multiple scan stop times can be identified in the test video stream. By pairing them in chronological order, multiple pairs of scan start times and scan stop times can be obtained. For example, the first scan start time and the first scan stop time can be determined as the first pair of start and stop times, and the scan start times and scan stop times after the first scan stop time can be determined as the second pair of start and stop times, and so on.

[0115] It should be noted that the scanning positions of multiple scans can overlap. Therefore, the scan data needs to be divided into multiple scan character sets according to the start and end times. The test method of this embodiment is repeated for each scan character set to obtain the corresponding sample character set, thereby obtaining multiple group test results. Then, the final target test result is obtained based on the multiple group test results. For example, the mean accuracy method can be used, or a weighted sum can be performed according to the size of the scan data, etc. There are no further limitations here.

[0116] like Figure 9 As shown, Figure 9 This is a structural diagram of a testing device based on machine vision and real-time text recognition provided in one embodiment of the present invention. The present invention also provides a testing device based on machine vision and real-time text recognition, comprising:

[0117] The processor 901 can be implemented using a general-purpose central processing unit (CPU), microprocessor, application specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application.

[0118] The memory 902 can be implemented as a read-only memory (ROM), static storage device, dynamic storage device, or random access memory (RAM). The memory 902 can store the operating system and other applications. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 902 and is called and executed by the processor 901 to execute the machine vision and real-time text recognition testing method of the embodiments of this application.

[0119] The input / output interface 903 is used to implement information input and output;

[0120] The communication interface 904 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.).

[0121] Bus 905 transmits information between various components of the device (e.g., processor 901, memory 902, input / output interface 903, and communication interface 904);

[0122] The processor 901, memory 902, input / output interface 903, and communication interface 904 are connected to each other within the device via bus 905.

[0123] This application also provides a testing device, including the testing apparatus based on machine vision and real-time text recognition as described above.

[0124] This application embodiment also provides a storage medium, which is a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the above-described test method based on machine vision and real-time text recognition.

[0125] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof. The device embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separate, and may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0126] It will be understood by those skilled in the art that all or some of the steps and systems in the methods disclosed above can be implemented as software, firmware, hardware, and suitable combinations thereof. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. Furthermore, as is known to those skilled in the art, communication media typically include computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.

[0127] The above provides a detailed description of the preferred embodiments of the present invention. However, the present invention is not limited to the above embodiments. Those skilled in the art can make various equivalent modifications or substitutions without departing from the spirit of the present invention. All such equivalent modifications or substitutions are included within the scope defined by the claims of the present invention.

Claims

1. A testing method based on machine vision and real-time text recognition, characterized in that, The testing equipment is used in a test setup, which is communicatively connected to the dictionary pen under test. The test setup includes a first camera, and the dictionary pen under test is equipped with a second camera. The testing method based on machine vision and real-time text recognition includes: The first image of the sample text is obtained by extracting text from the first image using the first camera to obtain a sample dictionary. The sample dictionary includes multiple sample characters recorded in the sample text and the character position of each sample character. The test video stream is acquired through the first camera, and the scanning start and end positions of the dictionary pen under test are determined based on the test video stream, wherein the test video stream records the scanning process of the dictionary pen under test on the sample text; A sample character set is determined from the sample dictionary based on the scan start and end positions, the sample character set including at least one of the sample characters; The dictionary pen under test obtains a scanned character set by scanning through the second camera, the scanned character set including at least one scanned character; The target test result is determined by comparing the scanned character set and the sample character set.

2. The testing method based on machine vision and real-time text recognition according to claim 1, characterized in that, The character position is the character sample coordinate, and the step of extracting text from the first image to obtain a sample dictionary includes: Perform text recognition in the text region of the first image, and acquire and record each sample character obtained from the text recognition; A coordinate system is established based on the first image to determine the character sample coordinates of each sample character; The sample dictionary is generated based on all the sample characters and the corresponding coordinates of the sample characters.

3. The testing method based on machine vision and real-time text recognition according to claim 2, characterized in that, The scan start and end positions include scan start coordinates and scan stop coordinates. Determining the scan start and end positions of the dictionary pen under test based on the test video stream includes: The scan start time and scan stop time of the dictionary pen under test are determined based on the test video stream; A second image is acquired from the test video stream based on the scan start time, and a third image is acquired from the test video stream based on the scan stop time; Pen tip recognition is performed on the second image and the third image respectively. The pen tip coordinates identified in the second image are determined as the scan start coordinates, and the pen tip coordinates identified in the third image are determined as the scan stop coordinates.

4. The testing method based on machine vision and real-time text recognition according to claim 3, characterized in that, The first camera's shooting direction is perpendicular to the sample text. The dictionary pen under test is also equipped with an indicator light. Determining the scanning start and stop times of the dictionary pen under test based on the test video stream includes: The lighting and turning-off times of the indicator lights are detected from the test video stream. The lighting time is determined as the scan start time, and the turning-off time is determined as the scan stop time. Alternatively, distance detection can be performed on the image frames of the test video stream to determine the target distance between the dictionary pen under test and the sample file in the image frames. The time corresponding to the first detection of the target distance being less than a preset distance threshold in the image frame is determined as the scan start time, and the time corresponding to the first detection of the target distance being greater than the preset distance threshold in the image frame after the scan start time is determined as the scan stop time.

5. The testing method based on machine vision and real-time text recognition according to claim 3, characterized in that, The step of determining the target test result by comparing the scanned character set and the sample character set includes: The character scan coordinates corresponding to each scanned character are determined based on the test video stream; Based on the principle of pairing the character scan coordinates, multiple comparison groups are determined, and each comparison group includes one sample character and one scan character. Determine the character alignment result for each alignment group, and obtain the target test result based on all the character alignment results.

6. The testing method based on machine vision and real-time text recognition according to claim 5, characterized in that, The step of obtaining the target test result based on all the character comparison results includes: When at least one of the character comparison results indicates a text error, the character scan coordinates corresponding to the comparison group indicating the text error are determined as abnormal coordinates, and the target test result is obtained based on all the character comparison results and the abnormal coordinates. Alternatively, when the number of characters in the scanned character set is greater than the number of characters in the sample character set, abnormal characters and abnormal position information are determined based on the scanned character set, and the target test result is obtained based on all the character comparison results, the abnormal characters, and the abnormal position information. The abnormal characters are the scanned characters located between two adjacent comparison groups, and the position information of the abnormal characters is determined based on the character sample coordinates corresponding to the two adjacent comparison groups. Alternatively, when the number of characters in the scanned character set is less than the number of characters in the sample character set, the missing characters and missing coordinates are determined based on the sample character set, and the target test result is obtained based on all the character comparison results, the missing characters, and the missing coordinates, wherein the missing characters are the sample characters that are not paired with the scanned characters, and the missing coordinates are the character sample coordinates corresponding to the missing characters.

7. The testing method based on machine vision and real-time text recognition according to claim 4, characterized in that, The method further includes: When multiple scan start and end positions are determined based on the test video stream, scan data is obtained from the dictionary pen under test, and the scan data includes multiple scanned characters; Based on each pair of scan start times and scan stop times, the scan data is divided into multiple scan character sets; Multiple sample character sets are obtained from the sample dictionary according to each scan start and end position, and the scan character set corresponding to each sample character set is determined according to the time order; The grouping test results are determined based on the comparison of the paired scanned character sets and the sample character sets; The target test result is determined based on all the group test results.

8. A testing device based on machine vision and real-time text recognition, characterized in that, It includes at least one control processor and a memory for communicatively connecting to the at least one control processor; the memory stores instructions executable by the at least one control processor, which, when executed by the at least one control processor, enable the at least one control processor to perform the test method based on machine vision and real-time text recognition as described in any one of claims 1 to 7.

9. A testing device, characterized in that, The test apparatus based on machine vision and real-time text recognition as described in claim 8.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions for causing a computer to perform the test method based on machine vision and real-time text recognition as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Learning scanning pen based on OCR recognition algorithm

    CN111950542A

  • Dictionary pen testing device

    CN215932635U