Data annotation method, computing device and storage medium

By using the initial model to determine the labeling information of the second training data in the ancient book text recognition system, the problem of high-cost labeling in ancient book text recognition is solved, and automatic labeling and efficient labeling are realized.

CN114973220BActive Publication Date: 2025-08-19ALIBABA GROUP HOLDING LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110215063.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-25
Publication Date
2025-08-19
Estimated Expiration
2041-02-25

AI Technical Summary

Technical Problem

The recognition of ancient text requires a large amount of labeling data, which is expensive and requires professionals to mark it, which has become the main bottleneck.

Method used

By obtaining the first training data with the labeling information, training at least two preset models is obtained, and then determining the labeling information of the second training data is reduced according to the initial model, the labeling requirement for the second training data is reduced.

Benefits of technology

It reduces the cost of data labeling, improves labeling efficiency, reduces dependence on professionals, and realizes automatic labeling of ancient texts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114973220B_ABST
    Figure CN114973220B_ABST
Patent Text Reader

Abstract

The present application provides a data annotation method, computing device, and storage medium. In the present application, first training data and second training data are obtained, wherein the first training data has annotation information for determining the true result corresponding to the first training data, and the second training data does not have annotation information. At least two preset models are trained based on the first training data to obtain corresponding initial models. The training result corresponding to the second training data is determined based on the initial model, and the annotation information corresponding to the second training data is determined based on the training result. Thus, by annotating part of the training data, namely the first training data, other training data can be automatically annotated, thereby reducing data annotation, reducing the cost of data annotation, and improving annotation efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a data annotation method, a method for recognizing ancient text, a computing device, and a storage medium. Background Art

[0002] With the development of object recognition, we can begin to face more complex recognition scenarios, such as the recognition of complex text. However, for complex recognition scenarios, such as the recognition of single words in text, especially the recognition of ancient texts, deep learning-based recognition requires a large amount of annotated data, which is very costly. Summary of the Invention

[0003] Various aspects of the present application provide a data annotation method, an ancient text recognition method, a computing device, and a storage medium, which can reduce data annotation and lower annotation costs.

[0004] An embodiment of the present application provides a data labeling method, including: obtaining first training data and second training data, where the first training data has labeling information for determining a true result corresponding to the first training data, and the second training data does not have labeling information; training at least two preset models based on the first training data to obtain corresponding initial models; determining a training result corresponding to the second training data based on the initial model, and determining labeling information corresponding to the second training data based on the training result.

[0005] An embodiment of the present application also provides a data annotation method, including: obtaining first image training data with ancient book text and second image training data with ancient book text, the first image training data having the annotation position of the ancient book text in the corresponding image, which is used to determine the actual annotation position corresponding to the first image training data, and the second image training data does not have the annotation position; training at least two preset target detection models according to the first image training data to obtain a corresponding initial target detection model; determining the training result corresponding to the second image training data according to the initial target detection model, and determining the annotation position corresponding to the second image training data according to the training result.

[0006] An embodiment of the present application also provides a method for recognizing ancient text, including: receiving a recognition request, obtaining a picture to be recognized in the recognition request, wherein the picture to be recognized contains ancient text; determining the marking position of the ancient text in the picture to be recognized through a preset target detection model; recognizing the ancient text according to the marking position, and returning it.

[0007] An embodiment of the present application also provides a method for recognizing ancient text, including: providing an ancient text recognition interface; selecting a picture to be recognized in response to a user's selection operation of the picture to be recognized; sending a recognition request to a recognition device in response to the user's recognition operation, wherein the recognition request carries the picture to be recognized, and the picture to be recognized contains ancient text, so that the recognition device determines the marked position of the ancient text in the picture to be recognized through a preset target detection model, and recognizes the ancient text based on the marked position; and receiving the recognized ancient text.

[0008] An embodiment of the present application also provides an information identification method, including: receiving a call request for a recognition service, and using the processing resources corresponding to the recognition service to implement the following steps; obtaining a picture to be identified in the call request, wherein the picture to be identified has an identification object; determining the marked position of the identification object in the picture to be identified through a preset target detection model; identifying the object based on the marked position, and returning the identified object.

[0009] An embodiment of the present application also provides a computing device, including: a memory and a processor; the memory is used to store a computer program; the processor executes the computer program to: obtain first training data and second training data, the first training data has labeling information, which is used to determine the true result corresponding to the first training data, and the second training data does not have labeling information; train at least two preset models according to the first training data to obtain corresponding initial models; determine the training result corresponding to the second training data according to the initial model, and determine the labeling information corresponding to the second training data according to the training result; train the initial model based on the second training data with labeling information to obtain a trained model.

[0010] An embodiment of the present application also provides a computing device, including: a memory, a processor, and a communication component; the memory is used to store a computer program; the processor is used to execute the computer program, so as to: determine the annotation position of the ancient text in the image to be identified through a preset target detection model; identify the ancient text according to the annotation position, and return it; the communication component is used to receive an identification request and obtain the image to be identified in the identification request, wherein the image to be identified contains ancient text.

[0011] An embodiment of the present application also provides a computing device, comprising: a memory, a processor, and a communication component; the memory is used to store a computer program; the processor is used to execute the computer program to: provide an interface for recognizing ancient text; select the picture to be recognized in response to a user's selection operation of the picture to be recognized; the communication component is used to send a recognition request to a recognition device in response to the user's recognition operation, wherein the recognition request carries the picture to be recognized, and the picture to be recognized contains ancient text, so that the recognition device determines the annotation position of the ancient text in the picture to be recognized through a preset target detection model; recognizes the ancient text according to the annotation position; and receives the recognized ancient text.

[0012] An embodiment of the present application also provides a computing device, comprising: a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program, so as to: obtain first image training data with ancient book text and second image training data with ancient book text, the first image training data having the annotated position of the ancient book characters in the ancient book text in the corresponding image, and being used to determine the actual annotated position corresponding to the first image training data, and the second image training data does not have the annotated position; training at least two preset target detection models according to the first image training data to obtain a corresponding initial target detection model; determining a training result corresponding to the second image training data according to the initial target detection model, and determining the annotated position corresponding to the second image training data according to the training result.

[0013] An embodiment of the present application further provides a computer-readable storage medium storing a computer program. When the computer program is executed by one or more processors, the one or more processors implement the steps in the above method.

[0014] In an embodiment of the present application, first training data and second training data are obtained, wherein the first training data has annotation information for determining the true result corresponding to the first training data, and the second training data does not have annotation information; at least two preset models are trained based on the first training data to obtain corresponding initial models; the training result corresponding to the second training data is determined based on the initial model, and the annotation information corresponding to the second training data is determined based on the training result. In this way, by annotating part of the training data, i.e., the first training data, the other training data (i.e., the second training data) can be automatically annotated, thereby reducing the data annotation, reducing the data annotation cost, and improving the annotation efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0016] Figure 1 A schematic structural diagram of a data annotation system according to an exemplary embodiment of the present application;

[0017] Figure 2 A flowchart of a data annotation method according to an exemplary embodiment of the present application is shown;

[0018] Figure 3 A flowchart of a data annotation method according to an exemplary embodiment of the present application is shown;

[0019] Figure 4 A flowchart of a method for recognizing ancient text according to an exemplary embodiment of the present application is shown;

[0020] Figure 5 A flowchart of a method for recognizing ancient text according to an exemplary embodiment of the present application is shown;

[0021] Figure 6 A schematic structural diagram of a data annotation device provided in an exemplary embodiment of the present application;

[0022] Figure 7 A schematic structural diagram of a data annotation device provided in an exemplary embodiment of the present application;

[0023] Figure 8 A schematic diagram of the structure of an ancient book character recognition device provided by an exemplary embodiment of the present application;

[0024] Figure 9 A schematic diagram of the structure of an ancient book character recognition device provided by an exemplary embodiment of the present application;

[0025] Figure 10 A schematic diagram of the structure of a computing device provided for an exemplary embodiment of the present application;

[0026] Figure 11 A schematic diagram of the structure of a computing device provided for an exemplary embodiment of the present application;

[0027] Figure 12 A schematic diagram of the structure of a computing device provided for an exemplary embodiment of the present application;

[0028] Figure 13 A schematic diagram of the structure of a computing device provided as an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0029] To make the purpose, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the specific embodiments of this application and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0030] With the practical application of OCR (Optical Character Recognition) technology, we can now handle text in more complex scenarios. One of the challenges is recognizing individual characters in ancient books. Ancient book recognition has high practical value, but deep learning-based ancient book recognition requires a large amount of annotated data, which is very costly and a major bottleneck in this area.

[0031] At present, the recognition of ancient text requires a large amount of annotated data, and due to the particularity of ancient books, it requires professional annotators who study ancient text to carry out the annotation, and the annotation cost is extremely high.

[0032] Therefore, the embodiments proposed in this application can reduce data annotation and lower annotation costs.

[0033] Figure 1 This is a structural diagram of a data annotation system provided by an exemplary embodiment of the present application. Figure 1 As shown, the system 100 may include: a first device 101 and a second device 102 .

[0034] The first device 101 can be a device with certain computing capabilities, capable of sending data to the second device 102 and receiving data sent by the second device 102. The basic structure of the first device 101 can include: at least one processor. The number of processors can depend on the configuration and type of the device with certain computing capabilities. The device with certain computing capabilities can also include memory, which can be volatile, such as RAM, or non-volatile, such as read-only memory (ROM), flash memory, etc., or can include both types. The memory typically stores an operating system (OS), one or more application programs, and may also store program data. In addition to the processing unit and memory, the device with certain computing capabilities also includes some basic configurations, such as a network card chip, an I / O bus, a display component, and some peripheral devices. Optionally, some peripheral devices may include, for example, a keyboard, a stylus, etc. Other peripheral devices are well known in the art and are not described in detail here. Optionally, the first device 101 can be a smart terminal, such as a mobile phone, a desktop computer, a laptop, a tablet computer, etc.

[0035] Second device 102 refers to a device that can provide computing and processing services in a network virtual environment. It can also be a device that uses the network to acquire data and perform model training. Physically, second device 102 can be any device that can provide computing services, respond to service requests, acquire data, and perform model training. For example, it can be a cloud server, cloud host, virtual center, conventional server, etc. Second device 102 primarily comprises a processor, hard drive, memory, system bus, etc., similar to a general computer architecture.

[0036] Specifically, the second device 102 obtains first training data and second training data, where the first training data has labeling information for determining a true result corresponding to the first training data, and the second training data does not have labeling information; trains at least two preset models based on the first training data to obtain a corresponding initial model; determines a training result corresponding to the second training data based on the initial model, and determines the labeling information corresponding to the second training data based on the training result.

[0037] The training result may refer to an output result corresponding to the second training data according to the initial model, for example, it may refer to the annotated position of the ancient text in the picture in the second training data.

[0038] In addition, the second device 102 trains the initial model based on the second training data with labeled information to obtain a trained model.

[0039] In addition, the second device 102 obtains a trained initial model based on the second training data, obtains third training data, and the third training data does not have labeling information; determines the training results corresponding to the third training data according to the trained initial model, and determines the labeling information corresponding to the third training data according to the training results; trains the trained initial model based on the third training data with labeling information to obtain a trained model.

[0040] It should be understood that the second training data and the third training data are essentially similar, both being training data and lacking annotated information. However, the third training data and the second training data differ in data content. For example, the second training data could be images from ancient book A, while the third training data could be images from ancient book B.

[0041] Specifically, the second device 102 obtains pictures corresponding to the ancient text, and uses the pictures with the text-annotated positions as the first training data; and uses the pictures without the text-annotated positions as the second training data.

[0042] Specifically, the second device 102 inputs the pictures corresponding to the ancient text into the preset model respectively, so that the preset model is trained based on the pictures and the corresponding text annotation positions to obtain at least two corresponding initial models.

[0043] Specifically, the second device 102 inputs the pictures corresponding to the ancient characters into at least two initial models respectively, and obtains at least two annotation positions of the ancient characters in the pictures respectively; for the corresponding ancient characters, determines whether the overlap of the at least two annotation positions is greater than the overlap threshold, and determines the number of ancient characters whose overlap is greater than the overlap threshold; and determines the final annotation position of the ancient characters in the picture based on the number of ancient characters.

[0044] Specifically, the second device 102 determines the proportion of the number of ancient book characters to the total number of ancient book characters in the image. If the proportion is greater than the proportion threshold, the final annotation position is determined according to the annotation position obtained by the initial model.

[0045] Specifically, the second device 102 determines the proportion of the number of ancient characters in the total number of ancient characters in the picture. If the proportion is less than the proportion threshold, the final annotation position of the ancient characters with an overlap greater than the overlap threshold is determined based on the annotation position obtained by the initial model; the ancient characters in the picture and the final annotation position of the corresponding ancient characters are provided for display; and the final annotation position of other ancient characters is determined based on the user's annotation operation.

[0046] In addition, the second device 102 determines the annotation positions of the ancient characters in the image to be identified that contains the ancient characters through at least one trained model; and identifies the ancient characters based on the annotation positions.

[0047] Specifically, the first device 101 sends a recognition request containing an image to be recognized, wherein the image contains ancient text. After the second device 102 recognizes the ancient text, it can return the recognized ancient text to the first device 101 for display. This allows the user to save the ancient text and facilitate subsequent editing.

[0048] In addition, the second device 102 may receive a call request for the recognition service and use the processing resources corresponding to the recognition service to implement the following steps: obtaining a picture to be recognized in the call request, wherein the picture to be recognized contains an object to be recognized; determining the marked position of the object to be recognized in the picture to be recognized through a preset target detection model; identifying the object according to the marked position, and returning the recognized object.

[0049] The second device 102 determines a calling interface of the identification service and a calling address of the identification service; and provides the identification service according to the calling interface and the calling address.

[0050] The user can obtain information about the calling interface and the calling address. In addition, the second device 102 can also determine the request method and return type of the identification service. It can also determine the corresponding request instance and provide this information to the outside, such as the request method, return type, and request instance. This information corresponds to or is associated with the called service, that is, the identification service, and the corresponding identification service can be located through this information. The second device 102 can record the above information in the form of an interface document and provide it to the corresponding user.

[0051] Based on this information, the user can send a call request from the first device 101 to the second device 102, which carries the image to be identified. After receiving the call request, the second device 102 calls the recognition service based on the call address and call interface in the call request to identify the image to be identified. The recognition service can be implemented using corresponding object detection models and recognition models. The corresponding recognition result is then returned to the first device 101 through the second device 102 for display, i.e., the identified object, such as ancient text.

[0052] In the recognition scenario of ancient text in the embodiment of the present application, user 103 can log in to the ancient text recognition interface of second device 102, such as cloud server, through the web browser of first device 101, such as computer. The computer displays the interface to user 103 for viewing. User 103 can operate on the interface, such as the user can select the image to be recognized on the interface to upload, and then the user can click the recognition button to recognize (or directly trigger the subsequent recognition after uploading the image to be recognized). The computer can respond to the user's recognition operation and send a recognition request to the cloud server. The recognition request carries the image to be recognized, and the image to be recognized contains ancient text. After receiving the recognition request, the cloud server inputs the image to be recognized into a preset target detection model, determines the position of each ancient text in the image to be recognized, and then can recognize the ancient text in the corresponding position through the preset recognition model to recognize the corresponding ancient text. Then, the cloud server returns the recognized ancient text to the computer so that the computer can display it. The user can store or directly edit the displayed ancient text.

[0053] In addition, the situation of calling the recognition service by the above-mentioned calling address and calling interface can also be: for example, according to the foregoing, the user can input calling information through a web browser according to the form of a request instance, such as the name of the preset API (Application Programming Interface) interface (such as the ancient book text recognition API interface), the calling address, such as the URL (Uniform Resource Locator) address, the corresponding request method, such as the POST method, the return type, such as the JSON type and the image to be recognized. Then, the user triggers the call request, and the computer responds to the user's triggering operation through the browser and sends the above information to the cloud server through the call request. After receiving it, the cloud server allows the use of the corresponding recognition service according to the calling address and the name of the preset API interface in the call request, and calls the corresponding recognition service, and inputs the request parameters, such as the image to be recognized, into the target detection model and the recognition model in the corresponding recognition service (wherein, the annotation information of the target detection model, such as the annotation position and the image to be recognized, can be input into the recognition model), and obtains the text in the final image to be recognized, such as the ancient book text in the ancient book image. The ancient text is then returned to the user's computer browser in the form of JSON for display.

[0054] It should be noted that the image to be identified can be in the form of an image corresponding to the text in an ancient book.

[0055] In addition, the training process for the above-mentioned preset target detection model is as follows: the second device 102, such as a cloud server, can first obtain the first training data. The first training data can be multiple pictures corresponding to ancient texts, and each ancient text in each picture is marked with its corresponding position, that is, the marked position. Then, the two preset models, such as SSD (Single Shot MultiBox Detector, a target detection model) and CenterNet (a target detection model), are trained using the first training data. For example, the first training data is input into the two preset models for initial training to obtain an initial model. Then the cloud server obtains the second training data. The second training data is multiple pictures corresponding to ancient texts and has no marked data. The second training data is input into the initial model respectively to obtain the corresponding results. Since there are two initial models, two results can be obtained for the same picture. The result refers to the position marking of the ancient text in the picture, that is, the marked position. For each image, the overlap of the marked positions of each ancient text in the image is determined based on the two results, and then it is determined whether the overlap is greater than a threshold value. The number of ancient texts corresponding to the overlap greater than the threshold value is determined, and the proportion of this number to all the ancient texts in the image is determined. If the proportion is greater than the threshold value, the image can be automatically marked based on the above results, such as by marking the positions of each ancient text in the image based on the results of one of the models. If the proportion is less than a preset value, it is necessary to manually mark the positions of each ancient text in the image. When manually marking, the image can be first displayed using a display device, such as a display computer, for manual marking. In order to reduce the cost of manual marking, it is also possible to mark the corresponding positions of ancient texts with an overlap greater than the threshold value using the results obtained by the model during the display, while no position marking is performed for ancient texts with an overlap less than the threshold value. The marking personnel can be prompted to mark the positions of ancient texts without marked positions. In this way, the marked positions of each ancient text in the image are determined. Then, the above two models are used through the second training data. The cloud server can also continue to acquire third training data, which can also be multiple images corresponding to ancient texts and without annotations. This data is processed in a similar manner to the second training data and will not be further described here. Training is terminated until the trained model meets the conditions of the final training iteration, generating the final preset object detection model, which can be used to determine the location of ancient texts in images.

[0056] It should be noted that in addition to the above-mentioned recognition of ancient texts, it is also possible to query ancient texts and play the ancient texts' corresponding audio. The specific application scenario process is similar to the above-mentioned recognition process and will not be repeated here. In addition to being applied to the recognition of ancient texts, the embodiments of this application can also be applied to the recognition of texts in other languages. The process is similar to that described above and will not be repeated here.

[0057] In addition to being applicable to the above-mentioned text scenarios, the embodiments of the present application can also be applied to other application scenarios, such as the recognition and labeling of product promotion images in e-commerce scenarios and the labeling and recognition of TCM medical records in medical scenarios. Specific examples are as follows:

[0058] For the application scenario of identifying and annotating product promotional images in e-commerce scenarios, for example, a user can upload product promotional images from online store A to the recognition interface provided by the computer's web browser through a cloud server in the manner described above. In response to the user's recognition operation, the computer can send a recognition request to the cloud server, which carries a product promotional image to be identified, and the image contains the corresponding product. After receiving the recognition request, the cloud server inputs the image to be identified into a preset target detection model to determine the location of at least one product in the image to be identified. The product at the corresponding location can then be identified using the preset recognition model. The cloud server then returns the identified product name to the computer for display, and the user can store or edit the displayed product name. In addition, the specific method for determining the product location is similar to the method for determining the location of ancient text described above, and will not be repeated here.

[0059] For the annotation and recognition of TCM medical records in medical scenarios, for example, the doctor can upload a picture of the medical record to the recognition interface provided by the computer's web browser through the cloud server in the manner described above. In response to the doctor's recognition operation, the computer can send a recognition request to the cloud server, and the recognition request carries a picture of the medical record to be recognized, and the picture has corresponding case text. After receiving the recognition request, the cloud server inputs the picture to be recognized into a preset target detection model, determines the position of at least one text in the picture to be recognized, and then recognizes the text at the corresponding position through the preset recognition model to identify the corresponding text. The cloud server then returns the recognized text to the computer for display, and the doctor can store or edit it according to the displayed text. In addition, the specific method of determining the position of the text is similar to the method of determining the position of the ancient text described above, and will not be repeated here.

[0060] For the recognition and annotation of textbooks, student assignments, test papers and other content in educational scenarios, there are:

[0061] Specifically, a recognition request is received, and a picture to be recognized in the recognition request is obtained, where the picture to be recognized contains text to be recognized; the text to be recognized refers to text used to detect learning status or to transmit knowledge; the annotation position of the text to be recognized in the picture to be recognized is determined; and the corresponding text to be recognized is identified according to the annotation position.

[0062] The text may include characters, numbers, and symbols, etc. For the text to be recognized, determining the marking position thereof may refer to marking the position of each character, each number, and / or each symbol in the text.

[0063] The text to be recognized can be textbooks, student homework, test papers, etc. in educational scenarios. The corresponding image is the image to be recognized.

[0064] Furthermore, with regard to annotation, specifically, training data for a first text image and training data for a second text image are obtained, wherein the training data for the first text image has annotation information for determining the true result corresponding to the training data for the first text image, and the training data for the second text image does not have annotation information; at least two preset models are trained based on the training data for the first text image to obtain corresponding initial models; the training result corresponding to the training data for the second text image is determined based on the initial model, and the annotation information corresponding to the training data for the second text image is determined based on the training result. Text refers to text used to detect learning conditions or to transmit knowledge, and the image corresponding to the text is the text image.

[0065] The annotation information refers to the position of each data in the training data in the corresponding text image, for example, the position of each character in the Chinese answer in the test paper in the text image.

[0066] For example, in the manner described above, educators, such as teachers, can upload images of student assignments to the recognition interface provided by a computer's web browser via a cloud server. In response to the teacher's recognition operation, the computer can send a recognition request to the cloud server, which carries an image of the student assignment to be identified, containing the corresponding assignment content. After receiving the recognition request, the cloud server inputs the image of the student assignment to be identified into a pre-set object detection model, determines the position of each word, number, and / or symbol in the image of the student assignment to be identified, and then uses the pre-set recognition model to recognize the words, numbers, and / or symbols at the corresponding positions, ultimately identifying the corresponding assignment content. The cloud server then returns the identified assignment content to the computer for display, allowing the computer to store or edit the displayed text. Furthermore, the cloud server can compare the identified assignment content with pre-set assignment content to determine its accuracy, then proofread the assignment and send the proofreading results to the computer for display to assist the teacher, or directly help the teacher critique the assignment, and can also assign grades according to grading rules.

[0067] In addition, the specific method of determining the position of text, numbers and / or symbols is similar to the method of determining the position of ancient text described above, and will not be repeated here.

[0068] In the above embodiment, the first device 101 and the second device 102 are connected to a network, which may be a wireless connection. If the first device 101 and the second device 102 are connected for communication, the network standard of the mobile network may be any one of 2G (GSM), 2.5G (GPRS), 3G (WCDMA, TD-SCDMA, CDMA2000, UTMS), 4G (LTE), 4G+ (LTE+), WiMax, 5G, etc.

[0069] The data labeling process is described in detail below in conjunction with the method embodiment.

[0070] Figure 2 This is a flow chart of a data annotation method according to an exemplary embodiment of the present application. The method 200 provided in the embodiment of the present application is executed by a computing device, such as a cloud server. The method 200 includes the following steps:

[0071] 201: Acquire first training data and second training data.

[0072] The first training data has labeling information, which is used to determine the true result corresponding to the first training data, and the second training data does not have labeling information.

[0073] 202: Train at least two preset models according to the first training data to obtain corresponding initial models.

[0074] 203: Determine a training result corresponding to the second training data according to the initial model, and determine labeling information corresponding to the second training data according to the training result.

[0075] The following is a detailed explanation of the above steps:

[0076] 201: Acquire first training data and second training data.

[0077] The first training data has annotation information for determining the true result corresponding to the first training data, and the second training data does not have annotation information. The first training data and the second training data can be text training data, i.e., first text training data and second text training data. Its specific implementation form can be a picture. For example, an ancient book can be stored or presented in the form of a picture, then the picture contains a corresponding ancient book fragment, and the ancient book fragment can also contain ancient book text. In addition, the picture does not necessarily need to be an ancient book fragment of the same ancient book, but can also be an ancient book fragment of different ancient books. The picture can also be ancient book text without the ancient book source.

[0078] The annotation information may be the position annotation of the ancient text in the image, or it may be a real result. The annotation information of the first training data is manually annotated.

[0079] It should be noted that, in addition to ancient text in images, training data can also include other training data, such as other objects to be recognized, ordinary text, different animals of the same type (such as cats, etc.). The corresponding annotation information can be the annotation location or other real results, such as the name of the object.

[0080] For example, as described above, the cloud server may obtain preset first training data and second training data from other servers, where the first training data has labeling information and the second training data has no labeling information.

[0081] Specifically, the first training data and the second training data include: obtaining pictures corresponding to ancient texts, and using pictures with text-annotated locations in the pictures as first training data; and using pictures without text-annotated locations in the pictures as second training data.

[0082] For example, as described above, a cloud server can obtain training data, namely, multiple images, from other servers or other service nodes. These images can contain content from ancient texts. Images with annotated locations of the ancient texts can serve as the first training data, while images without annotated locations can serve as the second training data. The annotated locations can be the corresponding coordinates of each ancient text in the image. These coordinates can be used to extract the ancient text from the image to obtain the target ancient text, and can be the coordinates of the circumscribed rectangular box of the ancient text.

[0083] 202: Train at least two preset models according to the first training data to obtain corresponding initial models.

[0084] Pre-configured models can be determined and configured based on different needs. For example, the object detection model, SSD model, and CenterNet model can be used. In addition to these two object detection models, other object detection models can also be used, such as FCOS (Fully Convolutional One-Stage Object Detection) and YOLO (an object detection model or algorithm). These will not be discussed in detail here.

[0085] In addition to the two preset models mentioned above, there can also be three, four, or five preset models, such as multiple preset object detection models. The training process for multiple preset models is similar to the training process for two preset models, and the training process for two preset models can be used as an example to illustrate:

[0086] For example, as described above, after acquiring the first training data, the cloud server inputs the first training data into the SSD model and CenterNet model, respectively, to perform initial training on these two models. This results in the initial SSD model and CenterNet model. This initial training process iterates the model training based on the annotated positions of the first training data, thereby training the model parameters.

[0087] Specifically, at least two preset models are trained according to the first training data to obtain corresponding initial models, including: inputting pictures corresponding to ancient text into the preset models respectively, so that the preset models are trained based on the pictures and the corresponding text annotation positions to obtain corresponding at least two initial models.

[0088] Since it has been explained in the previous article, I will not repeat it here. I will just explain that you can input the pictures into the SSD model and the CenterNet model respectively to get the output results, and then initially train the two models based on the annotation positions and the output results.

[0089] 203: Determine a training result corresponding to the second training data according to the initial model, and determine labeling information corresponding to the second training data according to the training result.

[0090] The training result refers to the output result corresponding to the second training data according to the initial model, and may refer to the annotated position of the ancient text in the picture in the second training data.

[0091] For example, according to the foregoing, the cloud server can input the second training data into the initial model SSD model and the CenterNet model respectively to obtain the corresponding output results, that is, the labeling information of the second training data.

[0092] Specifically, the training results corresponding to the second training data are determined according to the initial model, and the annotation information corresponding to the second training data is determined according to the training results, including: inputting the pictures corresponding to the ancient characters into at least two initial models respectively, and obtaining at least two annotation positions of the ancient characters in the pictures respectively; for the corresponding ancient characters, determining whether the overlap of the at least two annotation positions is greater than the overlap threshold, and determining the number of ancient characters whose overlap is greater than the overlap threshold; and determining the final annotation position of the ancient characters in the picture according to the number of ancient characters.

[0093] Among them, the overlap degree IoU (Intersection over Union) can also be called "intersection over union ratio".

[0094] Specifically, the final annotation position of the ancient characters in the image is determined based on the number of ancient characters, including: determining the proportion of the number of ancient characters to the total number of ancient characters in the image; if the proportion is greater than a first proportion threshold, determining the final annotation position based on the annotation position obtained by the initial model.

[0095] For example, according to the above, the cloud server can input the pictures in the second training data into the initial model SSD model and the CenterNet model respectively to obtain the coordinate position of each ancient text in each picture in the corresponding picture. Then the overlap of the coordinate positions of the relative ancient texts in the same picture can be determined, that is, the overlap of the boxes represented by the coordinate positions, and the overlap of each ancient text can be determined. If the overlap is greater than the overlap threshold, such as 70%, the number of ancient texts greater than the overlap threshold can be counted for the same picture, and the number of ancient texts less than or equal to the overlap threshold can be counted for the same picture. Then determine the proportion of the number of ancient texts greater than the overlap threshold in the same picture to the total number of all ancient texts in the picture. If the proportion is greater than the first weight threshold, such as 95% (it can be determined that the picture is above 95 points), the output result of any initial model can be used as the annotation position of each ancient text in the picture. The annotation position of the ancient text can also be determined based on the mean of the two output results of the two initial models.

[0096] Specifically, the final marking position of the ancient characters in the picture is determined based on the number of ancient characters, including: determining the proportion of the number of ancient characters to the total number of ancient characters in the picture; if the proportion is less than a second proportion threshold, determining the final marking position of the ancient characters whose overlap is greater than the overlap threshold based on the marking position obtained by the initial model, and the second proportion threshold is less than the first proportion threshold; providing the ancient characters in the picture and the final marking position of the corresponding ancient characters for display; and determining the final marking position of other ancient characters based on the user's marking operation.

[0097] For example, as mentioned above, if the proportion is less than the second proportion threshold, such as 90% (it can be determined that the corresponding image is less than 90 points), then manual annotation is required. The cloud server can then display the image to be annotated to the annotator through a display device, such as a computer, so that the annotator can annotate it. After the annotation is completed, the cloud server can receive the annotated image returned by the display device, thereby determining the annotated location of each ancient text in the image.

[0098] To reduce the labeling workload for annotators, the cloud server can also determine the annotated locations of ancient texts in the same image that are greater than a threshold overlap, for display. This determination method is the same as described above and will not be repeated here. Then, in the same image, ancient texts that are less than or equal to the overlap threshold are not annotated, and the image is displayed on a display device, allowing annotators to only annotate the unannotated ancient texts.

[0099] In addition, the second proportion threshold can also be equal to the first proportion threshold. If the proportion is less than or equal to the second proportion threshold, the ancient text in the picture and the final annotation position of the corresponding ancient text are provided for display; according to the user's annotation operation, the final annotation position of other ancient texts is determined.

[0100] Alternatively, if the second weight threshold is not equal to the first weight threshold, the cloud server can discard the proportion between the two (e.g., 90-95%, then between 90-95 points), or manually mark it according to the above method, which will not be repeated here.

[0101] Thus, the cloud server can obtain the annotation information of the second training data. It should be noted that pictures with a score of 95 or above (i.e., a proportion greater than 95%) do not need to be annotated, and pictures with a score of less than 90 (i.e., a proportion less than 90%) can be annotated by annotators, which can greatly reduce the annotation of single word boxes (i.e., annotation information). Moreover, after a certain number of model iterations, the model accuracy may be high enough, and the proportion of pictures with a score of less than 90 can be reduced to less than 10%. The number of single ancient text characters that need to be annotated in each picture is about 10%, so the total annotation cost can be reduced to less than 1% of the original cost.

[0102] In addition, for the same picture, there can be two output results corresponding to the ancient text (i.e., the output results of the above two models), or multiple output results (corresponding to the output results of multiple models), to determine whether the output results meet the above threshold (such as the overlap threshold), and whether the output results are marked with the corresponding complete ancient text. If there are half-marked ancient texts, this needs to be avoided.

[0103] In addition, the method 200 further includes: training the initial model based on second training data with labeled information to obtain a trained model.

[0104] For example, according to the foregoing, the cloud server can input the second training data with the marked position into the initial model SSD model and the CenterNet model respectively, so as to train the SSD model and the CenterNet model to obtain the trained model.

[0105] In addition, when training the initial model, the second training data with labeled information can be selected to select training data that meets the preset conditions, such as training data that meets the preset scores, such as pictures with a score greater than 90, thereby improving the accuracy and model capabilities of the model.

[0106] In addition, based on the second training data, a trained initial model is obtained. The method 200 also includes: obtaining third training data, the third training data does not have labeling information; determining the training results corresponding to the third training data based on the trained initial model, and determining the labeling information corresponding to the third training data based on the training results; training the trained initial model based on the third training data with labeling information to obtain a trained model.

[0107] After the initial model is trained using the second training data, the trained model can be further trained using other training data, which may also be unlabeled. Labeling information for the other training data is determined using the method described above, allowing the model to be trained using the other training data until model training is complete. This will not be further described here.

[0108] It should be noted that, through subsequent model training, such as fine-tuning the corresponding model with the second training data and the third training data, the model iteration cost and iteration efficiency are greatly reduced. In addition, the various steps of the embodiment of the present application can be modularized and operated by non-technical personnel, which is simple and effective.

[0109] Two object detection models, SSD and CenterNet, were used simultaneously to detect individual characters in ancient books. The final bounding box for each character was a combination of the bounding boxes output by these two trained models (i.e., the detection boxes represented by the labeled locations). Furthermore, during pre-labeling, the overlap between the SSD and CenterNet detection boxes was used to determine the percentage of each image. This combined the advantages of both object detection methods.

[0110] It should be noted that the only place where manual intervention is required in the present embodiment is the process of supplementing the annotation of images below 90 points. After the process of the present embodiment is established, there is no need for algorithm researchers to participate in the iterative process of the target detection model for ancient texts, and the detection effect of ancient text characters can be actively learned and iterated.

[0111] In addition to the above-mentioned manual intervention, manual intervention can also be performed in the following ways.

[0112] Specifically, the method 200 also includes: providing annotation information corresponding to the second training data for display; modifying the annotation information corresponding to the second training data according to the user's annotation operation; and training the initial model based on the second training data with the modified annotation information to obtain a trained model.

[0113] For example, as described above, the cloud server can display the annotated locations of the ancient text in the image in the second training data to the annotator via a display device, such as a computer, so that the annotator can review the annotated locations and determine whether the annotated locations are correct. If incorrect, the annotator can directly modify the annotated locations and then return the modified image to the cloud server. The cloud server can receive the modified image returned by the display device, determine the annotated locations of each ancient text in the image, and then train the initial model to obtain a trained model.

[0114] After the model is trained, recognition can be performed based on the trained model.

[0115] Specifically, the method 200 further includes: determining the annotation positions of the ancient characters in the image to be identified having the ancient characters through at least one trained model; and identifying the ancient characters according to the annotation positions.

[0116] Among them, the at least one trained model can be a trained SSD model, CenterNet model or other target detection model.

[0117] For example, as described above, a cloud server can obtain an image to be recognized from another server or service node. It can then input the image to be recognized into a trained SSD model, obtaining the model's output—the annotated locations of each ancient text character in the image to be recognized. Based on these annotated locations, the cloud server can then use a pre-set recognition model to identify the ancient text. The cloud server can then store the recognized ancient text.

[0118] It should be noted that annotation can also be performed using more than two trained models, and an average result, that is, an average annotation position corresponding to multiple annotation positions, can be determined based on the output results of multiple models.

[0119] In addition, the method 200 further includes: receiving a recognition request, obtaining a picture to be recognized in the recognition request, wherein the picture to be recognized contains ancient text.

[0120] For example, as described above, a cloud server can receive a recognition request sent by a computer. The computer can be a user's computer, and the user can log in to the ancient text recognition interface provided by the cloud server through the browser installed on the computer. Then, based on this interface, the user uploads an image to be recognized, which contains ancient text. The computer can respond to the user's recognition operation, such as the user clicking the recognition button on the recognition interface, and send a recognition request to the cloud server. After receiving the recognition request, the cloud server sends the image to be recognized to the trained SSD model or CenterNet model based on the image to be recognized carried in the recognition request, obtains the corresponding annotation position, and then recognizes the ancient text through the preset recognition model, and returns it to the user's computer for display.

[0121] In addition, the method 200 also includes: storing the identified ancient text in order; receiving a query request for querying the ancient text, and searching for the corresponding ancient text, the ancient text segment where the ancient text is located, and the corresponding ancient book identifier from the stored ancient text according to the query information in the query request; and sending the queried ancient text, the ancient text segment where the ancient text is located, and the corresponding ancient book identifier.

[0122] For example, as mentioned above, in addition to the recognition interface that the cloud server can provide, in addition to the recognition of ancient texts, it can also provide a query function. It can also be implemented through other interfaces, such as through a query interface. The query interface is provided in a similar manner to the above-mentioned recognition interface, and will not be repeated here. The user can enter query information on the interface, such as the ancient text to be queried. The computer can respond to the user's query operation, such as clicking the query button on the interface. The computer sends a query request to the cloud server. After receiving the query request, the cloud server searches for the corresponding ancient text from the stored ancient text (which can be obtained through the above-mentioned trained model) based on the ancient text to be queried in the request, and then selects the ancient text segment where the ancient text is located and the name of the ancient book to which the ancient text belongs. Then, the cloud server returns the queried ancient text, the ancient text segment where it is located, and the name of the ancient book to which it belongs to the computer for display.

[0123] In addition, the method 200 also includes: converting the recognized ancient text into corresponding voices in sequence and storing them; receiving a playback request to play the voices corresponding to the ancient texts; searching for the corresponding voices from the stored voices according to the playback information in the playback request, and providing the voices for playback.

[0124] For example, as mentioned above, the cloud server can also convert the recognized ancient text into ancient audio through a TTS text-to-speech model and save it. A playback interface can be provided to the user's computer. In response to the user's playback operation, the computer sends a playback request to the cloud server. Upon receiving the playback request, the cloud server searches for the corresponding ancient audio from the stored audio based on the playback information, such as the name of the ancient book, and returns the audio stream of the audio to the computer for playback.

[0125] Based on the above similar inventive concept, Figure 3 A flow chart of a data annotation method provided by another exemplary embodiment of the present application is shown. The method 300 provided in the embodiment of the present application is executed by a server, such as a cloud server, and the method 300 includes the following steps:

[0126] 301: Acquire first image training data containing ancient text and second image training data containing ancient text.

[0127] The first image training data has the annotated positions of the ancient text in the corresponding image, which is used to determine the real annotated positions corresponding to the first image training data, and the second image training data does not have an annotated position.

[0128] 302: Train at least two preset object detection models according to the first image training data to obtain corresponding initial object detection models.

[0129] 303: Determine a training result corresponding to the second image training data according to the initial object detection model, and determine a labeling position corresponding to the second image training data according to the training result.

[0130] Since the specific implementation of steps 301-303 has been described in detail above, they will not be repeated here. It is only noted that the first image training data and the second image training data may refer to image training data. The preset object detection model may be an SSD model, a CenterNet model, or the like.

[0131] In addition, the method 300 further includes: training the initial target detection model based on the second image training data with the marked positions to obtain a trained target detection model.

[0132] Since this has been explained in the previous article, I will not repeat it here.

[0133] In addition, for the contents not described in detail in the present method 300 , reference may also be made to the steps in the above-mentioned method 200 .

[0134] Based on the above similar inventive concept, Figure 4A flow chart of a method for recognizing ancient text provided by another exemplary embodiment of the present application is shown. The method 400 provided in the embodiment of the present application is executed by a server, such as a cloud server, and the method 400 includes the following steps:

[0135] 401: Receive the recognition request and obtain the image to be recognized in the recognition request.

[0136] Among them, the picture to be identified contains ancient text.

[0137] 402: Determine the labeling position of the ancient text in the image to be recognized by using a preset object detection model.

[0138] 403: Identify the ancient text based on the marked position and return it.

[0139] Since the specific implementation of steps 401-403 has been described in detail above, they will not be repeated here.

[0140] In addition, the method 400 also includes: receiving a query request for querying ancient texts, and searching for corresponding ancient texts, ancient text segments where the ancient texts are located, and corresponding ancient book identifiers from stored ancient texts based on query information in the query request, wherein the stored ancient texts are identified and stored from pictures with ancient texts through a preset target detection model; and sending the queried ancient texts, the ancient text segments where the ancient texts are located, and the corresponding ancient book identifiers.

[0141] Since this has been explained in the previous article, I will not repeat it here.

[0142] In addition, for the contents not described in detail in the present method 400 , reference may also be made to the steps in the above-mentioned method 200 .

[0143] Based on the above similar inventive concept, Figure 5 A flowchart of a method for recognizing ancient text provided by another exemplary embodiment of the present application is shown. The method 500 provided in the embodiment of the present application is executed by a smart terminal, such as a computer, and the method 500 includes the following steps:

[0144] 501: Provide an interface for recognizing ancient texts.

[0145] 502: In response to the user's operation of selecting a picture to be identified, select a picture to be identified.

[0146] 503: In response to the user's recognition operation, a recognition request is sent to the recognition device. The recognition request carries the image to be recognized, and the image to be recognized contains ancient text, so that the recognition device can determine the annotation position of the ancient text in the image to be recognized through a preset target detection model, and recognize the ancient text based on the annotation position.

[0147] 504: Receiving the recognized ancient text.

[0148] Since the specific implementation of steps 501-504 has been described in detail above, they will not be repeated here. It is only explained that the identification device can be the server mentioned above.

[0149] In addition, the method 500 also includes: providing a query interface for ancient text; determining the ancient text in response to the user's ancient text determination operation; sending a query request for querying the ancient text to a query device in response to the user's ancient text query operation, so that the query device searches for the corresponding ancient text, the ancient text segment where the ancient text is located, and the corresponding ancient book identifier from the stored ancient text according to the ancient text in the query request, wherein the stored ancient text is identified and stored from pictures with ancient text through a preset target detection model; receiving the queried ancient text, the ancient text segment where the ancient text is located, and the corresponding ancient book identifier.

[0150] Since this has been explained above, it will not be repeated here. It is only explained that the query device can be the server mentioned above, and the query device can be the same as the identification device, that is, the same server, or different from the identification device, that is, different servers.

[0151] In addition, for the contents not described in detail in the present method 500 , reference may also be made to the steps in the above-mentioned method 200 .

[0152] Figure 6 This is a schematic diagram of the structural framework of a data annotation device provided in an exemplary embodiment of the present application. The device 600 can be applied to a server. The device 600 includes: an acquisition module 601, a training module 602, and a determination module 603. The functions of each module are described in detail below:

[0153] The acquisition module 601 is used to acquire first training data and second training data. The first training data has labeling information and is used to determine the true result corresponding to the first training data. The second training data does not have labeling information.

[0154] The training module 602 is used to train at least two preset models according to the first training data to obtain corresponding initial models.

[0155] The determination module 603 is configured to determine a training result corresponding to the second training data according to the initial model, and determine labeling information corresponding to the second training data according to the training result.

[0156] In addition, the training module 602 is further configured to train the initial model based on the second training data with labeled information to obtain a trained model.

[0157] In addition, based on the second training data, a trained initial model is obtained, and the acquisition module 601 is also used to: obtain third training data, and the third training data does not have labeling information; the determination module 603 is also used to determine the training results corresponding to the third training data based on the trained initial model, and determine the labeling information corresponding to the third training data based on the training results; the training module 602 is also used to train the trained initial model based on the third training data with labeling information to obtain a trained model.

[0158] Specifically, the acquisition module 601 includes: an acquisition unit for acquiring pictures corresponding to ancient texts, and using pictures with text-annotated positions in the pictures as first training data; and using pictures without text-annotated positions in the pictures as second training data.

[0159] Specifically, the training module 602 is specifically used to: input the pictures corresponding to the ancient text into the preset model respectively, so that the preset model is trained based on the pictures and the corresponding text annotation positions to obtain at least two corresponding initial models.

[0160] Specifically, the determination module 603 includes: an input unit, used to input the pictures corresponding to the ancient characters into at least two initial models respectively, and obtain at least two annotation positions of the ancient characters in the pictures respectively; a determination unit, used to determine whether the overlap of the at least two annotation positions is greater than the overlap threshold for the corresponding ancient characters, and determine the number of ancient characters whose overlap is greater than the overlap threshold; a determination unit, used to determine the final annotation position of the ancient characters in the picture according to the number of ancient characters.

[0161] Specifically, the determination unit is used to determine the proportion of the number of ancient book characters to the total number of ancient book characters in the image. If the proportion is greater than a first proportion threshold, the final annotation position is determined according to the annotation position obtained by the initial model.

[0162] Specifically, the determination unit is used to determine the proportion of the number of ancient characters in the total number of ancient characters in the picture. If the proportion is less than the second proportion threshold, the final marking position of the ancient characters with an overlap greater than the overlap threshold is determined based on the marking position obtained by the initial model, and the second proportion threshold is less than the first proportion threshold; the ancient characters in the picture and the final marking position of the corresponding ancient characters are provided for display; and the final marking position of other ancient characters is determined based on the user's marking operation.

[0163] In addition, the determination module 603 is also used to determine the annotation position of the ancient text in the image to be identified with the ancient text through at least one trained model; the device 600 can also include: a recognition module, used to identify the ancient text according to the annotation position.

[0164] In addition, the device 600 may also include: a storage module for storing the identified ancient text in sequence; a search module for receiving a query request for querying ancient text, and searching for the corresponding ancient text, the ancient text segment where the ancient text is located, and the corresponding ancient book identifier from the stored ancient text according to the query information in the query request; and a sending module for sending the queried ancient text, the ancient text segment where the ancient text is located, and the corresponding ancient book identifier.

[0165] In addition, the storage module is also used to convert the recognized ancient text into corresponding speech in sequence and store it; the device 600 also includes: a receiving module for receiving a playback request and playing the speech corresponding to the ancient text; a providing module for searching for the corresponding speech from the stored speech according to the playback information in the playback request and providing the speech for playback.

[0166] In addition, the receiving module is further used to: receive a recognition request, obtain the image to be recognized in the recognition request, and the image to be recognized contains ancient text.

[0167] Figure 7 A schematic diagram of the structural framework of a data annotation device provided by another exemplary embodiment of the present application is shown. The device 700 can be applied to a server. The device 700 includes: an acquisition module 701, a training module 702, and a determination module 703. The functions of each module are described in detail below:

[0168] Acquisition module 701 is used to obtain first image training data with ancient book text and second image training data with ancient book text. The first image training data has the marked positions of the ancient book characters in the ancient book text in the corresponding image, and is used to determine the actual marked positions corresponding to the first image training data. The second image training data does not have a marked position.

[0169] The training module 702 is configured to train at least two preset object detection models according to the first image training data to obtain corresponding initial object detection models.

[0170] The determination module 703 is configured to determine a training result corresponding to the second image training data according to the initial object detection model, and determine a labeling position corresponding to the second image training data according to the training result.

[0171] In addition, the training module 702 is further configured to train the initial target detection model based on the second image training data with the marked positions to obtain a trained target detection model.

[0172] It should be noted that for some contents not mentioned in the apparatus 700 , reference may be made to the contents of the above-mentioned apparatus 600 .

[0173] Figure 8 A schematic diagram of the structural framework of an ancient text recognition device provided by another exemplary embodiment of the present application is shown. The device 800 can be applied to a server. The device 800 includes: a receiving module 801, a determination module 802, and a recognition module 803. The functions of each module are described in detail below:

[0174] The receiving module 801 is configured to receive a recognition request and obtain a picture to be recognized in the recognition request, wherein the picture to be recognized contains ancient text.

[0175] The determination module 802 is used to determine the annotation position of the ancient text in the image to be recognized by using a preset object detection model.

[0176] The recognition module 803 is used to recognize the ancient text according to the marked position and return it.

[0177] In addition, the receiving module 801 is also used to: receive a query request for querying ancient text, and according to the query information in the query request, search for the corresponding ancient text, the ancient text segment where the ancient text is located, and the corresponding ancient book identifier from the stored ancient text, wherein the stored ancient text is identified and stored from pictures with ancient text through a preset target detection model; the device 800 is also used to: a sending module for sending the queried ancient text, the ancient text segment where the ancient text is located, and the corresponding ancient book identifier.

[0178] It should be noted that for some contents not mentioned in the apparatus 800 , reference may be made to the contents of the above-mentioned apparatus 600 .

[0179] Figure 9 A schematic diagram of the structural framework of an ancient text recognition device provided by another exemplary embodiment of the present application is shown. The device 900 can be applied to a smart terminal. The device 900 includes: a providing module 901, a selecting module 902, a sending module 903, and a receiving module 904. The functions of each module are described in detail below:

[0180] A module 901 is provided for providing an interface for recognizing ancient text.

[0181] The selection module 902 is configured to select a picture to be identified in response to a user's selection operation of the picture to be identified.

[0182] The sending module 903 is used to send a recognition request to the recognition device in response to the user's recognition operation. The recognition request carries the picture to be recognized, and the picture to be recognized contains ancient text, so that the recognition device can determine the annotation position of the ancient text in the picture to be recognized through a preset target detection model, and recognize the ancient text based on the annotation position.

[0183] The receiving module 904 is used to receive the recognized ancient text.

[0184] In addition, the providing module 901 is also used to: provide a query interface for ancient text; the device 900 also includes: a determination module, used to determine the ancient text in response to the user's ancient text determination operation; the sending module 903 is also used to respond to the user's ancient text query operation and send a query request for querying the ancient text to the query device, so that the query device can search for the corresponding ancient text, the ancient text segment where the ancient text is located, and the corresponding ancient book identifier from the stored ancient text according to the ancient text in the query request, wherein the stored ancient text is identified and stored from pictures with ancient text through a preset target detection model; the receiving module 904 is also used to receive the queried ancient text, the ancient text segment where the ancient text is located, and the corresponding ancient book identifier.

[0185] It should be noted that for some contents not mentioned in the apparatus 900 , reference may be made to the contents of the above-mentioned apparatus 600 .

[0186] The above describes Figure 6 The internal functions and structure of the device 600 shown, in one possible design, Figure 6 The structure of the apparatus 600 shown can be implemented as a computing device, such as a cloud server. Figure 10 As shown, the device 1000 may include: a memory 1001, a processor 1002;

[0187] The memory 1001 is used to store computer programs.

[0188] Processor 1002 is configured to execute a computer program to: obtain first training data and second training data, where the first training data has annotation information and is used to determine a true result corresponding to the first training data, and the second training data does not have annotation information; train at least two preset models based on the first training data to obtain corresponding initial models; determine a training result corresponding to the second training data based on the initial model, and determine the annotation information corresponding to the second training data based on the training result; and train the initial model based on the second training data having the annotation information to obtain a trained model.

[0189] In addition, the processor 1002 is further configured to train the initial model based on the second training data having labeled information to obtain a trained model.

[0190] In addition, based on the second training data, a trained initial model is obtained, and the processor 1002 is also used to: obtain third training data, where the third training data does not have labeling information; determine the training results corresponding to the third training data based on the trained initial model, and determine the labeling information corresponding to the third training data based on the training results; and train the trained initial model based on the third training data with labeling information to obtain a trained model.

[0191] Specifically, the processor 1002 is specifically configured to: obtain pictures corresponding to ancient texts, and use pictures with text-annotated locations in the pictures as first training data; and use pictures without text-annotated locations in the pictures as second training data.

[0192] Specifically, the processor 1002 is specifically used to: input the pictures corresponding to the ancient text into the preset model respectively, so that the preset model is trained based on the pictures and the corresponding text annotation positions to obtain at least two corresponding initial models.

[0193] Specifically, processor 1002 is specifically used to: input the pictures corresponding to the ancient characters into at least two initial models respectively, and obtain at least two annotation positions of the ancient characters in the pictures respectively; a determination unit is used to determine whether the overlap of the at least two annotation positions is greater than the overlap threshold for the corresponding ancient characters, and determine the number of ancient characters whose overlap is greater than the overlap threshold; a determination unit is used to determine the final annotation position of the ancient characters in the picture according to the number of ancient characters.

[0194] Specifically, the processor 1002 is specifically configured to determine the proportion of the number of ancient book characters to the total number of ancient book characters in the image; if the proportion is greater than a first proportion threshold, determine the final annotation position based on the annotation position obtained by the initial model.

[0195] Specifically, processor 1002 is specifically used to: determine the proportion of the number of ancient characters in the total number of ancient characters in the picture; if the proportion is less than a second proportion threshold, determine the final annotation position of the ancient characters whose overlap is greater than the overlap threshold based on the annotation position obtained by the initial model, and the second proportion threshold is less than the first proportion threshold; provide the ancient characters in the picture and the final annotation position of the corresponding ancient characters for display; determine the final annotation position of other ancient characters based on the user's annotation operation.

[0196] In addition, the processor 1002 is further configured to determine the annotated positions of the ancient characters in the image to be identified containing the ancient characters through at least one trained model; the device 600 may further include: a recognition module configured to identify the ancient characters based on the annotated positions.

[0197] In addition, the processor 1002 is also used to store the identified ancient texts in sequence; the search module is used to receive a query request for querying ancient texts, and according to the query information in the query request, search for the corresponding ancient texts, the ancient text fragments where the ancient texts are located, and the corresponding ancient book identifiers from the stored ancient texts; the sending module is used to send the queried ancient texts, the ancient text fragments where the ancient texts are located, and the corresponding ancient book identifiers.

[0198] In addition, processor 1002 is also used to convert the recognized ancient text into corresponding speech in sequence and store it; receive a playback request to play the speech corresponding to the ancient text; provide a module to search for the corresponding speech from the stored speech according to the playback information in the playback request, and provide the speech for playback.

[0199] In addition, the processor 1002 is further configured to receive a recognition request and obtain a picture to be recognized in the recognition request, where the picture to be recognized contains ancient text.

[0200] In addition, an embodiment of the present invention provides a computer storage medium, which, when a computer program is executed by one or more processors, causes the one or more processors to implement Figure 2 The present invention provides steps of a data labeling method in a method embodiment.

[0201] The above describes Figure 7 The internal functions and structure of the device 700 shown, in one possible design, Figure 7 The structure of the device 700 shown can be implemented as a computing device, such as a cloud server. Figure 11 As shown, the device 1100 may include: a memory 1101 and a processor 1102 .

[0202] The memory 1101 is used to store computer programs.

[0203] Processor 1102 is used to execute the computer program to: obtain first image training data with ancient text and second image training data with ancient text, the first image training data having the annotated positions of ancient text characters in the ancient text in the corresponding image, and used to determine the actual annotated positions corresponding to the first image training data, and the second image training data does not have an annotated position; train at least two preset target detection models based on the first image training data to obtain a corresponding initial target detection model; determine a training result corresponding to the second image training data based on the initial target detection model, and determine the annotated position corresponding to the second image training data based on the training result.

[0204] In addition, the processor 1102 is further configured to train the initial target detection model based on the second image training data with the marked positions to obtain a trained target detection model.

[0205] It should be noted that for some contents not mentioned in the device 1100, reference may be made to the contents of the above-mentioned device 1000.

[0206] In addition, an embodiment of the present invention provides a computer storage medium, which, when a computer program is executed by one or more processors, causes the one or more processors to implement Figure 3 The present invention provides steps of a data labeling method in a method embodiment.

[0207] The above describes Figure 8 The internal functions and structure of the device 800 shown, in one possible design, Figure 8 The structure of the device 800 shown can be implemented as a computing device, such as a cloud server. Figure 12 As shown, the device 1200 may include: a memory 1201 , a processor 1202 , and a communication component 1203 .

[0208] The memory 1201 is used to store computer programs.

[0209] The processor 1202 is configured to execute the computer program to: determine the annotated position of the ancient text in the image to be identified by using a preset target detection model; identify the ancient text according to the annotated position, and return the result.

[0210] The communication component 1203 is used to receive a recognition request and obtain a picture to be recognized in the recognition request, where the picture to be recognized contains ancient text.

[0211] In addition, processor 1202 is also used to: receive a query request for querying ancient texts, and search for corresponding ancient texts, ancient text segments where the ancient texts are located, and corresponding ancient book identifiers from stored ancient texts based on query information in the query request, wherein the stored ancient texts are identified and stored from pictures with ancient texts through a preset target detection model; and send the queried ancient texts, the ancient text segments where the ancient texts are located, and the corresponding ancient book identifiers.

[0212] It should be noted that for some contents not mentioned in the device 1200, reference may be made to the contents of the above-mentioned device 1000.

[0213] In addition, an embodiment of the present invention provides a computer storage medium, which, when a computer program is executed by one or more processors, causes the one or more processors to implement Figure 4 The present invention provides steps of a method for recognizing characters in ancient books in a method embodiment.

[0214] The above describes Figure 9 The internal functions and structure of the device 900 shown, in one possible design, Figure 9 The structure of the device 900 shown can be implemented as a computing device, such as a cloud server. Figure 13 As shown, the device 1300 may include: a memory 1301 , a processor 1302 , and a communication component 1303 .

[0215] The memory 1301 is used to store computer programs.

[0216] The processor 1302 is configured to execute the computer program to: provide an interface for recognizing ancient text; and select a picture to be recognized in response to a user's operation of selecting a picture to be recognized.

[0217] Communication component 1303 is used to respond to the user's recognition operation and send a recognition request to the recognition device. The recognition request carries the image to be recognized, and the image to be recognized contains ancient text, so that the recognition device can determine the annotation position of the ancient text in the image to be recognized through a preset target detection model, and recognize the ancient text based on the annotation position; and receive the recognized ancient text.

[0218] It should be noted that for some contents not mentioned in the device 1300, reference may be made to the contents of the above-mentioned device 1000.

[0219] In addition, an embodiment of the present invention provides a computer storage medium, which, when a computer program is executed by one or more processors, causes the one or more processors to implement Figure 5 The present invention provides steps of a method for recognizing characters in ancient books in a method embodiment.

[0220] Based on the similar inventive concept described above, another exemplary embodiment of the present application provides an information identification method, including:

[0221] Receive the call request of the recognition service and use the processing resources corresponding to the recognition service to implement the following steps;

[0222] Get the image to be identified in the call request, where the image to be identified contains an identification object.

[0223] The preset target detection model is used to determine the label position of the identification object in the image to be identified.

[0224] According to the marked position, the object is identified and the identified object is returned.

[0225] Since the specific implementation of the above steps has been described in detail above, it will not be repeated here. It is only explained that the execution subject of this method can be the server mentioned above.

[0226] In addition, for the contents not described in detail in this method, reference can also be made to the various steps in the above-mentioned contents.

[0227] Based on this, an exemplary embodiment of the present application provides an information identification device. The device can be applied to a server. The device includes: a determination module and a provision module. The functions of each module are described in detail below:

[0228] The receiving module is used to receive the call request of the identification service and implement the following steps using the processing resources corresponding to the identification service;

[0229] The acquisition module is used to obtain the image to be identified in the call request, where the image to be identified has an identification object.

[0230] The determination module is used to determine the marked position of the identification object in the image to be identified through a preset target detection model.

[0231] The recognition module is used to identify objects based on the marked locations and return the identified objects.

[0232] It should be noted that for some contents not mentioned in the above device, reference can be made to the contents of the device described above.

[0233] Therefore, in one possible design, the structure of the apparatus can be implemented as a computing device, such as a cloud server. The device may include: a memory, a processor, and a communication component.

[0234] Memory for storing computer programs.

[0235] The communication component is used to receive the call request of the identification service.

[0236] The processor is configured to execute a computer program for: utilizing processing resources corresponding to the recognition service to implement the following steps: obtaining an image to be recognized in a call request, wherein the image to be recognized has an object to be recognized; determining a marked position of the object to be recognized in the image to be recognized using a preset target detection model; recognizing the object based on the marked position, and returning the recognized object.

[0237] It should be noted that for some contents not mentioned in this device, please refer to the contents of the device described above.

[0238] In addition, an embodiment of the present invention provides a computer storage medium. When a computer program is executed by one or more processors, the computer storage medium causes the one or more processors to implement the steps of an information identification method in the above method embodiment.

[0239] In addition, in some of the processes described in the above embodiments and the accompanying drawings, multiple operations that appear in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this article or may be executed in parallel. The sequence numbers of the operations, such as 201, 202, 203, etc., are only used to distinguish between different operations, and the sequence numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this article are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.

[0240] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, i.e., they may be located in one location or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of the present embodiment. Persons of ordinary skill in the art will be able to understand and implement the present invention without inventive effort.

[0241] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by adding a necessary general hardware platform, and of course can also be implemented by a combination of hardware and software. Based on this understanding, the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a computer product. The present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0242] The present invention is described with reference to flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as a combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable multimedia data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable multimedia data processing device generate instructions for implementing the process in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0243] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable multimedia data processing device to operate in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0244] These computer program instructions can also be loaded onto a computer or other programmable multimedia data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable device to implement the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0245] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.

[0246] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0247] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.

[0248] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A data labeling method, characterized in that: include: Obtaining first training data and second training data, wherein the first training data has annotation information and is used to determine a true result corresponding to the first training data, and the second training data does not have annotation information; obtaining images corresponding to ancient text, and using the images with locations where the text is annotated as the first training data; and using the images without locations where the text is annotated as the second training data; Training at least two preset models according to the first training data to obtain corresponding initial models; Input the pictures corresponding to the ancient characters into at least two initial models respectively, and obtain at least two marking positions of the ancient characters in the pictures respectively; for the corresponding ancient characters, determine whether the overlap of the at least two marking positions is greater than an overlap threshold, and determine the number of ancient characters with an overlap greater than the overlap threshold; determine the proportion of the number of ancient characters to the total number of ancient characters in the picture; if the proportion is greater than a first proportion threshold, determine the final marking position according to the marking position obtained by the initial model; if the proportion is less than a second proportion threshold, determine the final marking position of the ancient characters with an overlap greater than the overlap threshold according to the marking position obtained by the initial model, and the second proportion threshold is less than the first proportion threshold; provide the ancient characters in the picture and the final marking position of the corresponding ancient characters for display; determine the final marking position of other ancient characters according to the user's marking operation.

2. The method according to claim 1, characterized in that The method further comprises: The initial model is trained based on the second training data with labeled information to obtain a trained model.

3. The method according to claim 2, characterized in that Based on the second training data, a trained initial model is obtained, and the method further includes: Acquire third training data, where the third training data does not have labeling information; Determining a training result corresponding to the third training data according to the trained initial model, and determining labeling information corresponding to the third training data according to the training result; The trained initial model is trained based on the third training data with labeled information to obtain a trained model.

4. The method according to claim 1, wherein The step of training at least two preset models according to the first training data to obtain corresponding initial models includes: The pictures corresponding to the ancient text are respectively input into the preset model, so that the preset model is trained based on the pictures and the corresponding text annotation positions to obtain at least two corresponding initial models.

5. The method according to claim 2 or 3, characterized in that The method further comprises: Determining, by using at least one trained model, the marking positions of the ancient characters in the image to be recognized that contains the ancient characters; The ancient book characters are identified according to the marked positions.

6. The method according to claim 5, characterized in that The method further comprises: Store the recognized ancient texts in order; receiving a query request for searching for ancient text, and searching, based on query information in the query request, for corresponding ancient text, the ancient text segment where the ancient text is located, and the corresponding ancient book identifier from stored ancient text; The queried ancient book characters, the ancient book character segments where the ancient book characters are located, and the corresponding ancient book identifiers are sent.

7. The method according to claim 5, characterized in that The method further comprises: Convert the recognized ancient text into corresponding pronunciation in sequence and store it; Receive playback requests to play the corresponding audio of ancient texts; According to the play information in the play request, the corresponding voice is searched from the stored voices, and the voice is provided for play.

8. The method according to claim 1, characterized in that The method further comprises: Providing annotation information corresponding to the second training data for display; Modifying the annotation information corresponding to the second training data according to the user's annotation operation; The initial model is trained based on the second training data with the modified labeling information to obtain a trained model.

9. A data labeling method, characterized in that: include: Obtaining first image training data having ancient text and second image training data having ancient text, wherein the first image training data has annotated positions of ancient text characters in the ancient text in corresponding images, and determining true annotated positions corresponding to the first image training data, and the second image training data does not have the annotated positions; Training at least two preset object detection models according to the first image training data to obtain corresponding initial object detection models; Inputting images corresponding to ancient characters into at least two initial object detection models, respectively, to obtain at least two annotated positions of the ancient characters in the images, the images being first image training data and second image training data; determining, for the corresponding ancient characters, whether a degree of overlap of the at least two annotated positions is greater than a degree of overlap threshold, and determining the number of ancient characters whose degree of overlap is greater than the degree of overlap threshold; Determining the proportion of the number of ancient book characters to the total number of ancient book characters in the image; If the weight is greater than a first weight threshold, determining a final annotation position according to the annotation position obtained by the initial object detection model; If the weight is less than a second weight threshold, determining the final marking position of the ancient text with a degree of overlap greater than the overlap threshold based on the marking position obtained by the initial object detection model, and the second weight threshold is less than the first weight threshold; The ancient text in the picture and the final marking position of the corresponding ancient text are provided for display; and the final marking position of other ancient texts is determined according to the user's marking operation.

10. The method according to claim 9, characterized in that The method further comprises: The initial object detection model is trained based on the second image training data having the marked positions to obtain a trained object detection model.

11. A method for recognizing characters in ancient books, characterized in that: include: Receiving a recognition request, obtaining a picture to be recognized in the recognition request, wherein the picture to be recognized contains ancient text; Determining the annotated position of the ancient text in the image to be identified by using a preset target detection model; wherein the target detection model is trained according to the method of claim 10; According to the marked position, the ancient book text is identified and returned.

12. The method according to claim 11, characterized in that The method further comprises: receiving a query request for searching for ancient characters, and searching, based on query information in the query request, for corresponding ancient characters, ancient character segments containing the ancient characters, and corresponding ancient book identifiers from stored ancient characters, wherein the stored ancient characters are identified from images containing the ancient characters using a preset object detection model and stored; The queried ancient book characters, the ancient book character segments where the ancient book characters are located, and the corresponding ancient book identifiers are sent.

13. A method for recognizing characters in ancient books, characterized in that: include: Provide an interface for recognizing ancient texts; In response to a user's selection operation of a picture to be identified, selecting the picture to be identified; In response to a user's recognition operation, a recognition request is sent to a recognition device, wherein the recognition request carries an image to be recognized, wherein the image to be recognized contains ancient text, so that the recognition device determines the annotated location of the ancient text in the image to be recognized using a preset object detection model, and recognizes the ancient text based on the annotated location; wherein the object detection model is trained according to the method of claim 10; Receive the recognized ancient text.

14. The method according to claim 13, characterized in that The method further comprises: Provides a query interface for ancient texts; In response to a user's operation of determining an ancient character, determining the ancient character; In response to a user's query operation for ancient characters, a query request for searching for the ancient characters is sent to a query device, so that the query device searches for the corresponding ancient characters, the ancient character segment in which the ancient characters are located, and the corresponding ancient book identifier from stored ancient characters according to the ancient characters in the query request, wherein the stored ancient characters are recognized from images containing the ancient characters using a preset object detection model and stored; The queried ancient book characters, the ancient book character segments where the ancient book characters are located, and the corresponding ancient book identifiers are received.

15. A method for identifying information, characterized in that: include: Receive a call request for an identification service and use the processing resources corresponding to the identification service to implement the following steps; Obtaining a picture to be identified in a call request, wherein the picture to be identified contains an identification object; Determining the marked position of the identification object in the image to be identified by using a preset target detection model; wherein the target detection model is trained according to the method of claim 10; The object is identified according to the marked position, and the identified object is returned.

16. A computing device comprising: Memory, processor; The memory is used to store computer programs; The processor executes the computer program to: Obtaining first training data and second training data, wherein the first training data has annotation information and is used to determine a true result corresponding to the first training data, and the second training data does not have annotation information; obtaining images corresponding to ancient text, and using the images with locations where the text is annotated as the first training data; and using the images without locations where the text is annotated as the second training data; Training at least two preset models based on the first training data to obtain corresponding initial models; training the initial models based on the second training data with labeled information to obtain trained models; Input the pictures corresponding to the ancient characters into at least two initial models respectively, and obtain at least two marking positions of the ancient characters in the pictures respectively; for the corresponding ancient characters, determine whether the overlap of the at least two marking positions is greater than an overlap threshold, and determine the number of ancient characters with an overlap greater than the overlap threshold; determine the proportion of the number of ancient characters to the total number of ancient characters in the picture; if the proportion is greater than a first proportion threshold, determine the final marking position according to the marking position obtained by the initial model; if the proportion is less than a second proportion threshold, determine the final marking position of the ancient characters with an overlap greater than the overlap threshold according to the marking position obtained by the initial model, and the second proportion threshold is less than the first proportion threshold; provide the ancient characters in the picture and the final marking position of the corresponding ancient characters for display; determine the final marking position of other ancient characters according to the user's marking operation.

17. A computing device comprising: memory, processors, and communication components; The memory is used to store computer programs; The processor is configured to execute the computer program to: Determining the annotated position of the ancient text in the image to be identified by using a preset target detection model, wherein the target detection model is trained according to the method of claim 10; According to the marked position, the ancient book characters are identified and returned; The communication component is used to receive a recognition request and obtain a picture to be recognized in the recognition request, wherein the picture to be recognized contains ancient text.

18. A computing device comprising: memory, processors, and communication components; The memory is used to store computer programs; The processor is configured to execute the computer program to: Provide an interface for recognizing ancient texts; In response to a user's selection operation of a picture to be identified, selecting the picture to be identified; The communication component is configured to send a recognition request to a recognition device in response to a user's recognition operation, wherein the recognition request carries an image to be recognized, and the image to be recognized contains ancient text, so that the recognition device determines the annotated position of the ancient text in the image to be recognized using a preset object detection model; According to the marked position, the ancient text is recognized; wherein the target detection model is trained according to the method of claim 10; Receive the recognized ancient text.

19. A computing device comprising: Memory, processor; The memory is used to store computer programs; The processor is configured to execute the computer program to: Obtaining first image training data having ancient text and second image training data having ancient text, wherein the first image training data has annotated positions of ancient text characters in the ancient text in corresponding images, and determining true annotated positions corresponding to the first image training data, and the second image training data does not have the annotated positions; Training at least two preset target detection models based on the first image training data to obtain corresponding initial target detection models; wherein the initial target detection models are trained according to the method of claim 9; A training result corresponding to the second image training data is determined according to the initial object detection model, and the annotation position corresponding to the second image training data is determined according to the training result.

20. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by one or more processors, the one or more processors are caused to implement the steps of the method according to any one of claims 1 to 15.

Citation Information

Patent Citations

  • Method and device for recognizing characters in image

    CN102855480A

  • Method and system for text to speech conversion

    CN103098124A