Ancient character recognition and translation method and device, electronic equipment and storage medium

Through the method of identifying and generating prompt sentences in visual models, and combining large language models for ancient text translation, the problem of difficult to recognize and translate ancient text in the existing technology is solved, and efficient and intelligent ancient text recognition and translation are achieved.

CN119992560APending Publication Date: 2025-05-13SHANDONG INSPUR SCI RES INST CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510004758.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-02
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing technology is difficult to meet the needs of large-scale ancient characters recognition and translation, especially when complex symbol systems, ancient writing forms and modern languages ​​are different.

Method used

By preprocessing the original image, using a visual model to recognize ancient characters, and generating prompt statement propt, combining a large language model for translation, realizing the recognition and translation of ancient characters.

Benefits of technology

It improves the digital efficiency of ancient text recognition and translation, enhances the accuracy and intelligence of large-language model translation, and meets the needs of large-scale ancient text recognition and translation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992560A_ABST
    Figure CN119992560A_ABST
Patent Text Reader

Abstract

The invention provides an ancient character recognition and translation method and device, electronic equipment and a storage medium, and belongs to the technical field of education demonstration, and the method comprises the steps: carrying out the preprocessing of an original image, and obtaining an ancient character image; the ancient character image is input into a visual model, a character recognition result output by the visual model and a corresponding prompt statement prompt are obtained, and the prompt statement prompt comprises categories of ancient characters in the ancient character image and background knowledge of the ancient characters; and taking the prompt statement prompt as a prompt word, and inputting the prompt word and the character recognition result into a pre-trained large language model to obtain an ancient character translation result output by the large language model. According to the method, feature extraction is performed on the ancient character image through the visual model, and the extracted character is translated by using the large language model, so that the digitization efficiency of ancient character recognition and translation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of educational demonstration technology, and in particular to an ancient character recognition and translation method, device, electronic equipment and storage medium. Background Art

[0002] As human society enters the information and digital age, the protection and inheritance of cultural heritage has become a focus of attention. As an important part of cultural heritage, ancient characters carry rich historical information and the evolution of civilization. However, due to the passage of time, the research and interpretation of ancient characters face many challenges, mainly reflected in their complex symbol system, ancient writing form, and huge differences from modern languages.

[0003] The recognition and translation of ancient characters usually require manual interpretation by experts in the field of linguists, archaeologists, etc. This process is not only time-consuming and laborious, but also limited by the knowledge and experience of individual experts, and it is difficult to meet the needs of large-scale digitization of documents and cultural relics. Although traditional text recognition technologies, such as optical character recognition (OCR) technology, have achieved certain results in processing modern fonts and structured texts, they often seem powerless when faced with ancient characters. Ancient characters have diverse fonts, complex and irregular writing forms, and are especially difficult to recognize on damaged cultural relics. In addition, the semantics of ancient characters usually need to be deeply interpreted in combination with historical background, context and semiotics knowledge, which makes it difficult to meet the needs of ancient character translation by relying solely on image processing technology.

[0004] Therefore, how to meet the needs of large-scale ancient character recognition and translation has become a technical problem that needs to be solved urgently. Summary of the invention

[0005] The present invention provides an ancient character recognition and translation method, device, electronic device and storage medium, which are used to solve the defect that the prior art is difficult to meet the needs of ancient character recognition and translation in large-scale documents and cultural relics.

[0006] The present invention provides an ancient character recognition and translation method, comprising the following steps: Preprocess the original image to obtain the ancient text image; Input the ancient character image into the visual model to obtain the character recognition result output by the visual model and the corresponding prompt sentence prompt, wherein the prompt sentence prompt includes the category of the ancient characters in the ancient character image and the background knowledge of the ancient characters; The prompt sentence prompt is used as a prompt word, and the prompt word and the text recognition result are input into a pre-trained large language model to obtain an ancient text translation result output by the large language model.

[0007] According to an ancient character recognition and translation method provided by the present invention, the visual model includes a linear prediction head and a decoder prediction head parallel to the linear prediction head; The linear prediction head is used to identify and obtain the text recognition result, and the text recognition result includes the modern sentence with the maximum probability corresponding to the ancient text and the probability of the modern sentence; The decoder prediction head is used to generate a prompt sentence corresponding to the ancient characters.

[0008] According to an ancient character recognition and translation method provided by the present invention, the linear prediction head includes a multi-layer perceptron MLP and a softmax activation function, the multi-layer perceptron MLP is used to convert high-dimensional features into low-dimensional vectors, and the softmax activation function is used to convert the low-dimensional vectors into probability distributions.

[0009] According to an ancient character recognition and translation method provided by the present invention, before inputting the ancient character image into a visual model to obtain a character recognition result output by the visual model and a corresponding prompt sentence prompt, the method further includes: Annotate the ancient character images in the ancient character image collection; Based on the ancient text image set, determining a supervised dataset for training a visual model and a text dataset for fine-tuning a large language model; Based on the supervised data set, training the visual model to be trained to obtain the visual model; Based on the text dataset, the large language model is fine-tuned to obtain the pre-trained large language model.

[0010] According to an ancient character recognition and translation method provided by the present invention, after training the visual model to be trained based on the supervised data set to obtain the visual model, the method further comprises: The recognition accuracy of the visual model is determined based on the following formula: ; ; in, represents the recall rate of the visual model, Indicates the number of positive samples correctly identified by the visual model, Indicates the number of negative samples that are incorrectly identified by the visual model, represents the accuracy of the visual model, Indicates the number of positive samples that are incorrectly identified by the visual model.

[0011] According to an ancient character recognition and translation method provided by the present invention, after fine-tuning the large language model based on the text data set to obtain the pre-trained large language model, the method further includes: The translation accuracy of the large language model is determined based on the following formula: ; in, Indicates the quality of bilingual translation. Indicates the length of the reference translation. Indicates the length of the candidate translation, represents the maximum n-gram order, Indicates The precision of n-gram of order, Indicates The weight of the n-gram of order.

[0012] The present invention also provides an ancient character recognition and translation device, comprising the following modules: A preprocessing module is used to preprocess the original image to obtain an ancient text image; A text recognition module, used for inputting the ancient text image into a visual model, obtaining a text recognition result output by the visual model, and a corresponding prompt sentence prompt, wherein the prompt sentence prompt includes the category of the ancient text in the ancient text image and the background knowledge of the ancient text; The translation module is used to use the prompt sentence prompt as a prompt word, input the prompt word and the text recognition result into a pre-trained large language model, and obtain the ancient text translation result output by the large language model.

[0013] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, any one of the above-mentioned ancient character recognition and translation methods is implemented.

[0014] The present invention also provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the ancient character recognition and translation method as described above is implemented.

[0015] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the ancient character recognition and translation method described above is implemented.

[0016] The ancient character recognition and translation method, device, electronic device and storage medium provided by the present invention obtain an ancient character image by preprocessing the original image; input the ancient character image into a visual model to obtain a character recognition result output by the visual model, and a corresponding prompt sentence prompt, wherein the prompt sentence prompt includes the category of the ancient characters in the ancient character image and the background knowledge of the ancient characters; the prompt sentence prompt is used as a prompt word, and the prompt word and the character recognition result are input into a pre-trained large language model to obtain an ancient character translation result output by the large language model. This solution extracts features from the ancient character image through a visual model, and uses a large language model to translate the extracted characters, which greatly improves the digital efficiency of ancient character recognition and translation; in addition, the visual model automatically generates prompt words for the large language model, which further improves the accuracy and intelligence of the large language model translation, meeting the needs of large-scale ancient character recognition and translation. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0018] Figure 1 This is one of the flow charts of the ancient character recognition and translation method provided by the present invention.

[0019] Figure 2 This is the second flow chart of the ancient character recognition and translation method provided by the present invention.

[0020] Figure 3 It is a schematic diagram of the structure of the ancient character recognition and translation device provided by the present invention.

[0021] Figure 4 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0022] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0023] It should be noted that in the description of the embodiments of the present invention, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "include one..." do not exclude the existence of other identical elements in the process, method, article or device including the elements. The orientation or position relationship indicated by the terms "upper", "lower" and the like is based on the orientation or position relationship shown in the drawings, which is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention. Unless otherwise clearly specified and limited, the terms "installed", "connected" and "connected" should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, or it can be a connection between two elements. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0024] The terms "first", "second", etc. in this application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present application can be implemented in an order other than those illustrated or described here, and the objects distinguished by "first", "second", etc. are generally of one type, and the number of objects is not limited. For example, the first object can be one or more. In addition, "and / or" represents at least one of the connected objects, and the character " / " generally indicates that the objects associated with each other are in an "or" relationship.

[0025] Combine the following Figure 1-Figure 4 The invention describes an ancient Chinese character recognition and translation method, device, electronic device and storage medium provided by an embodiment of the invention.

[0026] Figure 1 This is one of the flow charts of the ancient character recognition and translation method provided by the present invention, such as Figure 1 As shown, the method includes the following: S110, preprocessing the original image to obtain an ancient character image; S120, inputting the ancient character image into a visual model, obtaining a character recognition result output by the visual model and a corresponding prompt sentence prompt, wherein the prompt sentence prompt includes the category of the ancient character in the ancient character image and background knowledge of the ancient character; S130, taking the prompt sentence prompt as a prompt word, inputting the prompt word and the text recognition result into a pre-trained large language model, and obtaining an ancient text translation result output by the large language model.

[0027] It should be noted that the execution entity of the ancient character recognition and translation method provided in the embodiment of the present application can be a server, a computer device, such as a mobile phone, a tablet computer, a laptop computer, an ultra-mobile personal computer (UMPC), a netbook or a personal digital assistant (PDA), etc.

[0028] In S110 , the image preprocessing method includes but is not limited to steps such as resizing, cropping, denoising, and repairing.

[0029] In the embodiment of the present invention, in order to realize the integrated processing of ancient character recognition and translation, the optimized visual model is connected to the fine-tuned large language model through a channel. Specifically, the prompt sentence generated by the visual model is used as the prompt word of the large language model, and the spliced ​​modern sentence is generated as the input of the language model. This connection method enables the large language model to directly receive and process the recognition results from the visual model, thereby avoiding information loss and misunderstanding. Through this integrated method, the model can achieve integrated reasoning.

[0030] In recent years, deep learning techniques, especially Visual Transformer (ViT) and Large Language Model (LLM), have made breakthrough progress in image recognition and natural language processing. These models have powerful feature extraction and pattern recognition capabilities, and can capture key information in complex image and language data. However, despite the outstanding performance of ViT and LLM in their respective fields, the research on combining these two technologies for ancient text recognition and translation is relatively limited. The image processing of ancient texts and the conversion of modern languages ​​is essentially a cross-modal problem, involving the whole process from the extraction of visual information to semantic understanding. Therefore, it is difficult to cope with this complex task by relying solely on a single modality model. Through a multi-stage system, the visual model is used for image feature extraction of ancient texts, and the large language model is used to translate and polish the extracted text, which can effectively overcome the limitations of the single-modality method.

[0031] The ancient character recognition and translation method provided by the embodiment of the present invention obtains an ancient character image by preprocessing the original image; the ancient character image is input into the visual model to obtain the character recognition result output by the visual model, and the corresponding prompt sentence prompt, the prompt sentence prompt includes the category of the ancient character in the ancient character image and the background knowledge of the ancient character; the prompt sentence prompt is used as a prompt word, the prompt word and the character recognition result are input into the pre-trained large language model, and the ancient character translation result output by the large language model is obtained. This solution extracts features from the ancient character image through the visual model, and uses the large language model to translate the extracted characters, which greatly improves the digital efficiency of ancient character recognition and translation; in addition, the prompt words of the large language model are automatically generated through the visual model, which further improves the accuracy and intelligence of the large language model translation, meeting the needs of large-scale ancient character recognition and translation.

[0032] In an optional embodiment, the visual model includes a linear prediction head, and a decoder prediction head in parallel with the linear prediction head; The linear prediction head is used to identify and obtain the text recognition result, and the text recognition result includes the modern sentence with the maximum probability corresponding to the ancient text and the probability of the modern sentence; The decoder prediction head is used to generate a prompt sentence corresponding to the ancient characters.

[0033] Specifically, the ViT visual framework is modified to add a new prediction head of the decoder architecture in parallel with the original linear prediction head, that is, the output of the visual model , The decoder prediction head outputs the prompt sentence corresponding to the ancient characters. The decoder prediction head autoregressively generates the prompt sentence of the large language model. , the prompt statement can include but is not limited to the category of ancient characters and related background knowledge; is the output of the linear prediction head, which converts the feature vector into a category probability distribution through linear transformation.

[0034] The ancient character recognition and translation method provided by the embodiment of the present invention adds a decoder prediction head, which decodes the encoding vector through the prediction head of the decoder architecture to obtain the prompt sentence corresponding to the ancient character, and automatically generates the prompt word of the large language model without relying on manual input, thereby further improving the accuracy and automation of model prediction.

[0035] Furthermore, the linear prediction head includes a multi-layer perceptron MLP and a softmax activation function, wherein the multi-layer perceptron MLP is used to convert high-dimensional features into low-dimensional vectors, and the softmax activation function is used to convert the low-dimensional vectors into probability distributions.

[0036] Specifically, the linear prediction head adopts the MLP (Multi-Layer Perceptron) structure and uses softmax ( ,in, Represents the original output vector The elements, yes The exponential function value of Yes all The sum of the exponential values ​​of is used for normalization to ensure that the sum of the output probabilities is 1) as an activation function so that the model can provide confidence for each output result. , It is the high-dimensional feature output by ViT, which is converted into a low-dimensional vector using MLP and the result with the highest probability is used as the prediction result of the model.

[0037] Here, MLP is a classic feedforward neural network model, which is widely used in tasks such as classification and regression. By stacking multiple hidden layers, MLP is able to learn complex patterns in the input data.

[0038] Based on any of the above embodiments, before inputting the ancient character image into the visual model to obtain the character recognition result output by the visual model and the corresponding prompt sentence prompt, the method further includes: Annotate the ancient character images in the ancient character image collection; Based on the ancient text image set, determining a supervised dataset for training a visual model and a text dataset for fine-tuning a large language model; Based on the supervised data set, training the visual model to be trained to obtain the visual model; Based on the text dataset, the large language model is fine-tuned to obtain the pre-trained large language model.

[0039] In the specific implementation process, various ancient text image resources including ancient cultural relics and ancient books are collected. To ensure the accuracy of the recognition results, these images need to be manually annotated and interpreted by experts in related fields. At the same time, a vocabulary library for the interpretation of ancient characters should also be established to provide a vocabulary library foundation for the large language model. In addition, it is necessary to collect text content from ancient books extensively, especially texts covering various historical periods, different calligraphy styles and diverse forms of expression, to ensure the diversity and representativeness of the data. On this basis, supervised data sets for training visual models are sorted out separately. and text datasets for fine-tuning large language models .

[0040] In an embodiment of the present invention, the visual model is systematically trained and optimized based on the prepared data set. After the training, the visual model can effectively learn complex font features and detail changes from ancient text images.

[0041] In the embodiment of the present invention, the language model is fine-tuned through the collected ancient texts and QA question-answering data containing historical background, so that it can effectively process and understand the output results from the visual model. In particular, it can capture the information in the prompt sentence provided by the visual model and further analyze this information so as to generate a more contextual and logical explanation and provide relevant historical information. Through this process, the language model will have a stronger ability to understand ancient characters and express modern language.

[0042] Furthermore, after training the visual model to be trained based on the supervised data set to obtain the visual model, the method further includes: The recognition accuracy of the visual model is determined based on the following formula: ; ; in, represents the recall rate of the visual model, Indicates the number of positive samples correctly identified by the visual model, Indicates the number of negative samples that are incorrectly identified by the visual model, represents the accuracy of the visual model, Indicates the number of positive samples that are incorrectly identified by the visual model.

[0043] The ancient character recognition and translation method provided by the embodiment of the present invention has a recall rate ( ) is the proportion of all positive samples identified by the model, and the accuracy ( ) is the proportion of samples predicted to be positive that are actually positive, which is calculated by the recall rate ( )、Accuracy( ) and multiple evaluation criteria to conduct comprehensive and rigorous testing on the model to ensure its recognition and accuracy on different ancient text images.

[0044] Furthermore, after fine-tuning the large language model based on the text dataset to obtain the pre-trained large language model, the method further includes: The translation accuracy of the large language model is determined based on the following formula: ; in, Indicates the quality of bilingual translation. Indicates the length of the reference translation. represents the length of the candidate translation (i.e. the length of the generated translation), represents the maximum n-gram order, Indicates The precision of n-gram of order, Indicates The weight of the n-gram of order.

[0045] Based on any of the above embodiments, in order to enhance the user experience and facilitate the user to intuitively understand the processing process, a system page containing a visual operation interface is created. In this page, the user can simultaneously view the original ancient text image, the recognition results of the visual model, and the interpretation and translation results of the large language model. By displaying these information side by side, the user can more clearly understand the processing logic of the model and evaluate and provide feedback on the final translation results.

[0046] Figure 2 This is the second flow chart of the ancient character recognition and translation method provided by the present invention. For ease of understanding, the following is combined with Figure 2 The preferred ancient Chinese character recognition and translation method of the present invention is described.

[0047] like Figure 2 As shown, first, the visual model and the large language model are pre-trained and fine-tuned based on the collected training data, and the trained visual model is stored separately. With large language models When the model needs to be used for ancient character recognition and translation, the collected images will be denoised so that the model can better process the image information. It consists of an embedding layer, a backbone network, and two prediction heads. The embedding layer represents the image as a high-dimensional information feature, the backbone network processes and extracts the high-dimensional information, and the MLP-based linear prediction head recognizes the ancient characters in the image. Output, including the modern sentence with the highest probability corresponding to the ancient characters and its probability ; Generate prompts corresponding to ancient characters based on the decoder's prediction head , including ancient character categories, historical background, etc. In many cases, the modern sentences output by the linear prediction head are classical Chinese or difficult to understand, so cross-modal operations are required to convert the output results of the visual model Input to language model In for The prompt words increase the model's understanding ability. Based on the prompt words, the input is further understood and translated, and finally easy-to-understand and fluent text is provided In addition, as shown in the figure, the original image, the result of the visual model and the final result are displayed simultaneously in one interface so that users can compare and consult, providing personalized operations.

[0048] In summary, the ancient character recognition and translation method provided by the present invention realizes the self-prompt function on the basis of the combination of the visual model and the language model, and automatically provides the language model with prompt sentences based on the optimized visual model architecture, thereby increasing the model's understanding ability and prediction ability. Cross-modal processing from ancient character images to modern sentences is realized to ensure that the system can not only recognize ancient characters, but also accurately translate. Secondly, the method provides a credibility score for each character during the recognition process, so that the user can intuitively understand the model's recognition confidence for each word, thereby improving the interpretability of the translation result. In addition, the language model can give feedback to the visual model based on the confidence level, thereby increasing the predictive performance and generalization ability of the visual model. The design of end-to-end training enables the model to optimize the entire recognition and translation process and capture the complex relationship between ancient characters and modern languages. At the same time, the model has good scalability and can adapt to different types of ancient characters and language translation needs by adjusting the training data and model structure. The system also designs a visual operation page to facilitate users to view the original image, recognition results and translation results at the same time, thereby improving the user experience. Finally, the method performs fine-grained processing at the character level to ensure higher recognition accuracy in complex ancient character images.

[0049] The following is a description of the ancient character recognition and translation device provided in an embodiment of the present application. The ancient character recognition and translation device described below and the ancient character recognition and translation method described above can be referenced to each other.

[0050] Figure 3 Schematic diagram of the structure of the ancient Chinese character recognition and translation device provided by the present invention. Figure 3 As shown, the ancient character recognition and translation device may include but is not limited to; The preprocessing module 310 is used to preprocess the original image to obtain an ancient character image; The text recognition module 320 is used to input the ancient text image into the visual model to obtain the text recognition result output by the visual model and the corresponding prompt sentence prompt, wherein the prompt sentence prompt includes the category of the ancient text in the ancient text image and the background knowledge of the ancient text; The translation module 330 is used to use the prompt sentence prompt as a prompt word, input the prompt word and the text recognition result into a pre-trained large language model, and obtain the ancient text translation result output by the large language model.

[0051] It should be noted that the ancient character recognition and translation device provided in the embodiment of the present invention can execute the ancient character recognition and translation method described in any of the above embodiments during specific operation, which will not be elaborated in this embodiment.

[0052] Figure 4 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 4 As shown, the electronic device may include: a processor 410, a communication interface 420, a memory 430 and a communication bus 440, wherein the processor 410, the communication interface 420 and the memory 430 communicate with each other through the communication bus 440. The processor 410 may call the logic instructions in the memory 430 to execute the ancient character recognition and translation method, which includes: preprocessing the original image to obtain an ancient character image; inputting the ancient character image into a visual model to obtain a character recognition result output by the visual model and a corresponding prompt sentence prompt, wherein the prompt sentence prompt includes the category of the ancient character in the ancient character image and the background knowledge of the ancient character; using the prompt sentence prompt as a prompt word, inputting the prompt word and the character recognition result into a pre-trained large language model, and obtaining an ancient character translation result output by the large language model.

[0053] In addition, the logic instructions in the above-mentioned memory 430 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0054] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the ancient character recognition and translation methods provided by the above methods, which include: preprocessing the original image to obtain an ancient character image; inputting the ancient character image into a visual model to obtain a character recognition result output by the visual model, and a corresponding prompt sentence prompt, wherein the prompt sentence prompt includes the category of the ancient characters in the ancient character image and the background knowledge of the ancient characters; using the prompt sentence prompt as a prompt word, inputting the prompt word and the character recognition result into a pre-trained large language model, and obtaining the ancient character translation result output by the large language model.

[0055] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the ancient character recognition and translation method provided by the above-mentioned methods, the method comprising: preprocessing the original image to obtain an ancient character image; inputting the ancient character image into a visual model to obtain a character recognition result output by the visual model, and a corresponding prompt sentence prompt, wherein the prompt sentence prompt includes the category of the ancient characters in the ancient character image and the background knowledge of the ancient characters; using the prompt sentence prompt as a prompt word, inputting the prompt word and the character recognition result into a pre-trained large language model, and obtaining an ancient character translation result output by the large language model.

[0056] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0057] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0058] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for ancient character recognition and translation, characterized in that: include: Preprocess the original image to obtain the ancient text image; Input the ancient character image into the visual model to obtain the character recognition result output by the visual model and the corresponding prompt sentence prompt, wherein the prompt sentence prompt includes the category of the ancient characters in the ancient character image and the background knowledge of the ancient characters; The prompt sentence prompt is used as a prompt word, and the prompt word and the text recognition result are input into a pre-trained large language model to obtain an ancient text translation result output by the large language model.

2. The ancient character recognition and translation method according to claim 1, characterized in that: The visual model includes a linear prediction head, and a decoder prediction head in parallel with the linear prediction head; The linear prediction head is used to identify and obtain the text recognition result, and the text recognition result includes the modern sentence with the maximum probability corresponding to the ancient text and the probability of the modern sentence; The decoder prediction head is used to generate a prompt sentence corresponding to the ancient characters.

3. The ancient character recognition and translation method according to claim 2, characterized in that: The linear prediction head includes a multi-layer perceptron MLP and a softmax activation function, wherein the multi-layer perceptron MLP is used to convert high-dimensional features into low-dimensional vectors, and the softmax activation function is used to convert the low-dimensional vectors into probability distributions.

4. The ancient character recognition and translation method according to any one of claims 1 to 3, characterized in that: Before inputting the ancient character image into the visual model to obtain the character recognition result output by the visual model and the corresponding prompt sentence prompt, the method further includes: Annotate the ancient character images in the ancient character image collection; Based on the ancient text image set, determining a supervised dataset for training a visual model and a text dataset for fine-tuning a large language model; Based on the supervised data set, training the visual model to be trained to obtain the visual model; Based on the text dataset, the large language model is fine-tuned to obtain the pre-trained large language model.

5. The ancient character recognition and translation method according to claim 4, characterized in that: After the visual model to be trained is trained based on the supervised data set to obtain the visual model, the method further includes: The recognition accuracy of the visual model is determined based on the following formula: ; ; in, represents the recall rate of the visual model, Indicates the number of positive samples correctly identified by the visual model, Indicates the number of negative samples that are incorrectly identified by the visual model, represents the accuracy of the visual model, Indicates the number of positive samples that are incorrectly identified by the visual model.

6. The ancient character recognition and translation method according to claim 4, characterized in that: After fine-tuning the large language model based on the text dataset to obtain the pre-trained large language model, the method further includes: The translation accuracy of the large language model is determined based on the following formula: ; in, Indicates the quality of bilingual translation. Indicates the length of the reference translation. Indicates the length of the candidate translation, represents the maximum n-gram order, Indicates The precision of n-gram of order, Indicates The weight of the n-gram of order.

7. An ancient character recognition and translation device, characterized in that: include: A preprocessing module is used to preprocess the original image to obtain an ancient text image; A text recognition module, used for inputting the ancient text image into a visual model, obtaining a text recognition result output by the visual model, and a corresponding prompt sentence prompt, wherein the prompt sentence prompt includes the category of the ancient text in the ancient text image and the background knowledge of the ancient text; The translation module is used to use the prompt sentence prompt as a prompt word, input the prompt word and the text recognition result into a pre-trained large language model, and obtain the ancient text translation result output by the large language model.

8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the ancient character recognition and translation method as described in any one of claims 1 to 6 is implemented.

9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the ancient character recognition and translation method as described in any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the ancient character recognition and translation method as described in any one of claims 1 to 6 is implemented.

Citation Information

Cited By

  • Qinxiang character text recognition method

    CN120564207A