Ticket recognition methods, devices, electronic equipment and computer-readable storage media

By using a pre-defined recognizer trained based on the Prompt paradigm and leveraging randomly encoded latent information and sample ticket information, the problem of low accuracy and insufficient flexibility of ticket classification models in the case of few samples is solved, achieving higher recognition accuracy and flexibility.

CN116740748BActive Publication Date: 2026-04-03BOE TECHNOLOGY GROUP CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-21
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing ticket classification models have low classification accuracy and lack flexibility in classification mechanisms when there are few samples.

Method used

A pre-defined recognizer trained based on the Prompt paradigm is used to determine hidden information and sample ticket information through random encoding, thereby improving recognition accuracy and enhancing flexibility.

Benefits of technology

The accuracy of ticket recognition is improved in cases with few samples, and the Prompt paradigm template is determined by randomly encoded latent vectors, achieving greater recognition flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116740748B_ABST
    Figure CN116740748B_ABST
Patent Text Reader

Abstract

This application provides a ticket recognition method, apparatus, electronic device, and computer-readable storage medium, relating to the field of computer technology. The method includes: acquiring target ticket information of a ticket to be recognized; processing the target ticket information using a preset recognizer to obtain a predicted ticket category of the ticket to be recognized; wherein the preset recognizer is trained based on sample recognition information; the sample recognition information includes latent information and sample ticket information of sample tickets; the latent information is determined by randomly encoding the recognition prompt statement. This application achieves improved recognition accuracy by recognizing tickets using a preset recognizer trained based on the Prompt paradigm in cases with few samples. Furthermore, by determining the Prompt paradigm template using randomly encoded latent vectors, this application achieves greater flexibility in the Prompt paradigm compared to a Prompt paradigm template determined by fixed encoding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and more specifically, to a ticket recognition method, apparatus, electronic device, and computer-readable storage medium. Background Technology

[0002] With the development of information technology, people's work and life involve various tickets and certificates. For example, tickets and certificates can include vouchers, documents, receipts, invoices, tickets, coupons, certificates, documents, and so on.

[0003] For numerous tickets and receipts, management can be achieved through classification and archiving. In related technologies, ticket classification can be implemented using classification models; however, with a small sample size, the classification accuracy of these models is low, and their classification mechanisms lack flexibility. Summary of the Invention

[0004] The purpose of this application is to address at least one of the aforementioned technical deficiencies, particularly the low classification accuracy and lack of flexibility in the classification mechanism of the classification model.

[0005] According to one aspect of this application, a ticket recognition method is provided, the method comprising:

[0006] Obtain the target ticket information for the ticket to be identified;

[0007] The target ticket information is identified by a preset recognizer to obtain the predicted ticket category of the ticket to be identified.

[0008] The preset recognizer is obtained by training based on sample recognition information;

[0009] The sample identification information includes hidden information and sample ticket information of the sample ticket; the hidden information is determined by randomly encoding the identification prompt statement.

[0010] Optionally, obtaining the target ticket information of the ticket to be identified includes:

[0011] Image recognition is performed on the multimedia file of the ticket to be identified to determine the text information in the multimedia file;

[0012] The text information is processed to obtain the target ticket information of the ticket to be identified.

[0013] Optionally, before obtaining the target ticket information of the ticket to be identified, the method further includes:

[0014] Obtain a training sample set; the training sample set includes sample data of the sample tickets;

[0015] The sample identification information is determined based on the hidden information and the sample ticket information in the sample data;

[0016] The sample identification information is input into the initial identification model to obtain the predicted sample category corresponding to each sample ticket;

[0017] The training loss value is determined based on the predicted sample category, the reference ticket category, and the prediction accuracy corresponding to the implicit information.

[0018] Based on the training loss value, the initial recognition model is trained to obtain the preset recognizer that meets the training conditions.

[0019] Optionally, determining the training loss value includes:

[0020] A first loss value is determined based on the predicted sample category and the reference ticket category;

[0021] The second loss value is determined based on the prediction accuracy corresponding to the implicit information.

[0022] The training loss value includes the first loss value and the second loss value.

[0023] Optionally, obtaining the training sample set includes:

[0024] Determine the preset information collection template corresponding to each of the sample tickets;

[0025] Obtain the sample ticket information that corresponds to the preset information collection template from the sample tickets.

[0026] Optionally, the step of processing the target ticket information using a preset recognizer to obtain the predicted ticket category of the ticket to be recognized includes:

[0027] The target ticket information is identified and processed by a preset recognizer to obtain the ticket category code of the target ticket information;

[0028] The predicted ticket category is determined based on the ticket category code.

[0029] Optionally, determining the predicted ticket category based on the ticket category code includes:

[0030] Calculate the similarity between the ticket category code and the preset category code;

[0031] The predicted ticket category is determined based on the preset category code with the highest similarity.

[0032] Optionally, the preset recognizer is trained based on the Prompt paradigm;

[0033] The sample identification information is determined based on the Prompt paradigm template.

[0034] According to one aspect of this application, a ticket recognition method is provided, the method comprising:

[0035] Receive the target ticket information of the ticket to be identified through the page recognition template;

[0036] Displays the predicted ticket category of the ticket to be identified;

[0037] The predicted ticket category is obtained by identifying the target ticket information through a preset recognizer.

[0038] The preset recognizer is trained based on sample recognition information;

[0039] The sample identification information includes hidden information and sample ticket information of the sample ticket; the hidden information is determined by randomly encoding the identification prompt statement.

[0040] According to another aspect of this application, a ticket recognition device is provided, the device comprising:

[0041] The acquisition module is used to acquire the target ticket information of the ticket to be identified;

[0042] The identification module is used to identify the target ticket information through a preset recognizer to obtain the predicted ticket category of the ticket to be identified.

[0043] The preset recognizer is obtained by training based on sample recognition information;

[0044] The sample identification information includes hidden information and sample ticket information of the sample ticket; the hidden information is determined by randomly encoding the identification prompt statement.

[0045] According to another aspect of this application, an electronic device is provided, the electronic device comprising:

[0046] A memory, a processor, and a computer program stored in the memory, characterized in that the processor executes the computer program to implement the steps of the ticket recognition method according to any one of the first aspects of this application. For example, a third aspect of this application provides a computing device including: a processor, a memory, a communication interface, and a communication bus, wherein the processor, the memory, and the communication interface communicate with each other via the communication bus;

[0047] The memory is used to store at least one executable instruction that causes the processor to perform the operation corresponding to the ticket recognition method shown in the first aspect of this application.

[0048] According to another aspect of this application, a computer-readable storage medium is provided, on which a computer program is stored, characterized in that, when the computer program is executed by a processor, it implements the steps of the ticket recognition method according to any one of the first aspects of this application.

[0049] For example, in a fourth aspect of the present application, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the ticket recognition method shown in the first aspect of the present application.

[0050] According to one aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in various alternative implementations of the first aspect described above.

[0051] The beneficial effects of the technical solution provided in this application are:

[0052] In this embodiment, target ticket information is obtained, and a preset recognizer is used to process the target ticket information to obtain a predicted ticket category. The preset recognizer is trained based on sample recognition information, which includes latent information and sample ticket information. The latent information is determined by randomly encoding the recognition prompt statement. This application achieves improved recognition accuracy by using a preset recognizer trained on the Prompt paradigm to identify tickets with few samples. Furthermore, by determining the Prompt paradigm template using randomly encoded latent vectors, this application achieves greater flexibility in the Prompt paradigm compared to a fixed-encoding template. Attached Figure Description

[0053] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below.

[0054] Figure 1 A schematic diagram of the system architecture of a ticket recognition method provided in this application embodiment;

[0055] Figure 2 One of the flowcharts for a ticket recognition method provided in this application embodiment;

[0056] Figure 3 A schematic diagram of a ticket for a ticket recognition method provided in this application embodiment;

[0057] Figure 4 This is a flowchart illustrating the text recognition process in a ticket recognition method provided in an embodiment of this application.

[0058] Figure 5 A schematic diagram of a classification algorithm framework for a ticket recognition method provided in an embodiment of this application;

[0059] Figure 6 A vector representation diagram of a ticket recognition method provided in an embodiment of this application;

[0060] Figure 7 One of the schematic diagrams of a model framework for a ticket recognition method provided in an embodiment of this application;

[0061] Figure 8 A second schematic diagram of the model framework for a ticket recognition method provided in an embodiment of this application;

[0062] Figure 9 This is one of the application scenario diagrams of a ticket recognition method provided in the embodiments of this application;

[0063] Figure 10 This is a second schematic diagram illustrating an application scenario of a ticket recognition method provided in this application embodiment;

[0064] Figure 11 This is the third schematic diagram illustrating an application scenario of a ticket recognition method provided in this application embodiment;

[0065] Figure 12 This is one of the schematic diagrams illustrating the implementation process of a ticket recognition method provided in this application embodiment;

[0066] Figure 13 A second schematic diagram illustrating the implementation process of a ticket recognition method provided in this application embodiment;

[0067] Figure 14 A second schematic flowchart illustrating a ticket recognition method provided in this application embodiment;

[0068] Figure 15 This is a schematic diagram of the structure of a ticket recognition device provided in an embodiment of this application;

[0069] Figure 16 This is a schematic diagram of the structure of an electronic device for ticket recognition provided in an embodiment of this application. Detailed Implementation

[0070] The embodiments of this application are described below with reference to the accompanying drawings. It should be understood that the embodiments described below with reference to the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions of the embodiments of this application.

[0071] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the terms “comprising” and “including” as used in embodiments of this application mean that the corresponding feature can be implemented as the presented feature, information, data, step, operation, element, and / or component, but do not exclude implementation as other features, information, data, step, operation, element, component, and / or combinations thereof supported by the art. It should be understood that when we say that an element is “connected” or “coupled” to another element, the one element can be directly connected or coupled to the other element, or it can mean that the one element and the other element establish a connection relationship through an intermediate element. Furthermore, “connected” or “coupled” as used herein can include wireless connection or wireless coupling. The term “and / or” as used herein indicates at least one of the items defined by the term; for example, “A and / or B” can be implemented as “A,” or as “B,” or as “A and B.”

[0072] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.

[0073] At least some of the content of the ticket recognition method provided in this application involves fields such as machine learning in the field of artificial intelligence, as well as various fields of cloud technology, such as cloud computing, cloud services, and related data computing and processing fields in the field of big data.

[0074] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.

[0075] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.

[0076] Natural Language Processing (NLP) is an important field within computer science and artificial intelligence. It studies the theories and methods for enabling effective communication between humans and computers using natural language. NLP is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field involves natural language—the language people use in daily life—and thus it has a close relationship with linguistic research. NLP techniques typically include text processing, semantic understanding, machine translation, question answering, and knowledge graphs.

[0077] Machine Learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instructional learning.

[0078] With the research and advancement of artificial intelligence (AI) technology, AI is being studied and applied in various fields, such as smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, autonomous driving, drones, robots, smart healthcare, and intelligent recognition. It is believed that with the development of technology, AI will be applied in more fields and play an increasingly important role.

[0079] The ticket recognition method in this application can be applied to technical fields such as intelligent recognition.

[0080] To further illustrate the technical solutions provided in the embodiments of this application, a detailed description is provided below in conjunction with the accompanying drawings and specific implementation methods. Although the embodiments of this application provide method operation steps as shown in the following embodiments or drawings, the method may include more or fewer operation steps based on conventional or non-inventive methods. In steps where there is no logically necessary causal relationship, the execution order of these steps is not limited to the execution order provided in the embodiments of this application.

[0081] First, combine Figure 1 This is a system architecture diagram of the ticket recognition method provided in the embodiments of this application. The system may include a server 101 and a terminal cluster, wherein the server 101 can be considered as the backend server for ticket recognition.

[0082] The terminal cluster may include: terminal 102, terminal 103, terminal 104, ..., wherein each terminal may have a client installed that supports ticket recognition. Communication connections may exist between the terminals; for example, there may be a communication connection between terminal 102 and terminal 103, and a communication connection between terminal 103 and terminal 104.

[0083] Meanwhile, server 101 can provide services to the terminal cluster through communication connection function. Any terminal in the terminal cluster can have a communication connection with server 101. For example, terminal 102 has a communication connection with server 101, and terminal 103 has a communication connection with server 101. The above-mentioned communication connection is not limited to the connection method. It can be directly or indirectly connected through wired communication, or directly or indirectly connected through wireless communication, or through other methods.

[0084] The network for the aforementioned communication connection can be a wide area network (WAN), a local area network (LAN), or a combination of both. This application does not impose any restrictions on this.

[0085] The ticket recognition method of this application can be executed on the server side or the terminal side, and the execution subject is not limited in this application embodiment. During the ticket recognition process, target ticket information of the ticket to be recognized is obtained, and the target ticket information is processed by a preset recognizer to obtain the predicted ticket category of the ticket to be recognized. The preset recognizer is trained based on sample recognition information; the sample recognition information includes latent information and sample ticket information of sample tickets; the latent information is determined by randomly encoding the recognition prompt statement. This application achieves recognition of tickets to be recognized using a preset recognizer trained based on the Prompt paradigm even with few samples, improving the accuracy of recognition. Furthermore, by determining the Prompt paradigm template through randomly encoded latent vectors, this application achieves greater flexibility in the Prompt paradigm compared to a Prompt paradigm template determined by fixed encoding.

[0086] Therefore, the method provided in this application embodiment can be executed by a computer device, which includes, but is not limited to, a terminal (including the aforementioned user terminal) or a server (including the aforementioned server 101). The aforementioned server can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The aforementioned terminal can be a smartphone, tablet computer, laptop computer, desktop computer, smart speaker, smartwatch, etc., but is not limited to these. The terminal and server can be directly or indirectly connected via wired or wireless communication, which is not limited herein.

[0087] Of course, the methods provided in the embodiments of this application are not limited to... Figure 1 The application scenarios shown can also be used in other possible scenarios, and this application does not impose any limitations. Figure 1 The functions that each device in the application scenario shown can achieve will be described in subsequent method embodiments, and will not be elaborated on here.

[0088] This application provides a possible implementation, which can be executed by any electronic device. Optionally, any electronic device can be a server device with ticket recognition capabilities, or a device or chip integrated into these devices. Figure 2 As shown, this is one of the flowcharts of a ticket recognition method provided in this application embodiment. The method includes the following steps:

[0089] Step S201: Obtain the target ticket information of the ticket to be identified.

[0090] Optionally, the embodiments of this application can be applied to application scenarios that classify tickets to be identified.

[0091] Specifically, the ticket to be identified may include a ticket to be identified; wherein, the ticket may include a voucher document; for example, the ticket may include a receipt, invoice, ticket, voucher, certificate, document, etc.; in this embodiment of the application, the ticket includes but is not limited to the above examples. In actual implementation scenarios, the ticket may be a document containing graphics or text, etc., and this application does not limit it in this regard.

[0092] The target ticket information includes the ticket information in the ticket to be identified; taking an invoice as an example, the invoice information may include invoice code, invoice number, verification code, invoice date, buyer's name, goods name, goods amount, goods unit price, goods quantity, etc. As an example, the ticket to be identified and the target ticket information can be referenced. Figure 3 As shown.

[0093] Optionally, the target ticket information may include all or a portion of the ticket information in the ticket to be identified.

[0094] In some alternative implementations, the target ticket information can be extracted by performing optical character recognition (OCR) on the multimedia file of the ticket to be identified. For example, the multimedia file can be an image of the ticket to be identified. First, text detection can be performed on the multimedia file to extract the text information. Then, the text information can be recognized to obtain the target ticket information.

[0095] In cases where the target ticket information is a portion of the ticket information to be identified, after detecting the text of the ticket to be identified, a portion of the information can be extracted as the target ticket information. For example, information with a higher preset weight coefficient can be extracted as the target ticket information. For instance, the invoice code, invoice number, verification code, buyer's name, goods name, goods amount, unit price, and quantity of goods can be extracted as the target ticket information.

[0096] Step S202: The target ticket information is identified by a preset recognizer to obtain the predicted ticket category of the ticket to be identified.

[0097] The preset recognizer is obtained by training based on sample recognition information;

[0098] The sample identification information includes hidden information and sample ticket information of the sample ticket; the hidden information is determined by randomly encoding the identification prompt statement.

[0099] In this embodiment of the application, the target ticket information can be identified by a preset recognizer to obtain the predicted ticket category of the ticket to be identified.

[0100] The predicted ticket category refers to the predicted category of the ticket. For example, in actual implementation scenarios, the predicted ticket category may include value-added tax invoices, scenic spot tickets, coupons, graduation certificates, degree certificates, English certificates, etc.

[0101] Optionally, the preset recognizer in this application can be trained based on the Prompt paradigm. The sample recognition information can be determined through the Prompt paradigm template. Since the Prompt paradigm performs well in model training scenarios with few or no samples, it can improve the prediction accuracy of the preset recognizer.

[0102] Specifically, the Prompt paradigm performs well in training scenarios with few or no samples. In related technologies, ticket classification is typically a supervised training task. For example, in practical applications, ticket information x is used as input data, and ticket type y is used as the model's output data to train the model P(y|x; θ), where θ is the maximum likelihood parameter, which represents the probability of accurate prediction. However, the aforementioned methods in these technologies require high-quality labeled data, resulting in high initial costs.

[0103] The training method based on the Prompt paradigm in this application can solve the above problems. The Prompt paradigm guides the model to model the text itself and predicts the ticket type based on the maximum likelihood parameter (probability), thereby reducing the dependence of the training model on a large amount of supervised data. The Prompt paradigm can be mapped by the function f. prompt The expression, i.e., x′=f prompt (x; θ); where x′ represents the expression function mapped by the prompt function; f prompt θ represents the prompt function; x represents the ticket information; θ represents the maximum likelihood parameter.

[0104] In practical implementation scenarios, the training method based on the Prompt paradigm can train the model by using the Prompt paradigm template as input. The Prompt paradigm template, as described in this embodiment, is the sample identification information, including identification prompt statements and sample ticket information. Optionally, in this embodiment, the identification prompt statements can be randomly encoded to obtain latent information (i.e., latent vectors). The Prompt paradigm template is then constructed using the latent information and sample ticket information. Thus, by determining the Prompt paradigm template through randomly encoded latent vectors, this application achieves greater flexibility in the Prompt paradigm compared to a Prompt paradigm template determined by fixed encoding.

[0105] For example, the Prompt paradigm template can be illustrated with the following example:

[0106] Taking the identification of a ticket type as an example, the Prompt paradigm template can be represented as: "This is a [Z] ticket, [X: ticket information]". Here, "This is a [Z] ticket" is the identification prompt; [Z] represents the ticket type to be predicted; and X represents the ticket information. It should be noted that during the prediction process, the identification prompt in the above Prompt paradigm template can be randomly encoded to obtain implicit information; that is, the Prompt paradigm template is trained based on the implicit information and the ticket information.

[0107] As an example, the recognition prompts in the Prompt paradigm template can include various expressions, such as: "The following is a [MASK] ticket," "The following ticket is taken from [MASK] ticket," "The following ticket should belong to [MASK] ticket," "Classify the following ticket into [MASK] ticket," etc. Different expressions have varying effects on model performance. This application can represent the recognition prompts using latent information (latent vectors), that is, by combining the characteristics of various expressions in the form of random encoding, making the expression more flexible.

[0108] Furthermore, during training, the downstream task can be adjusted to one of the two main tasks in the pre-trained model: the MLM task and the NSP task. Based on the input format of the MLM task, special characters such as "[MASK]" can be used to fill the [Z] position in the Prompt paradigm template. This results in "This is a [MASK] type ticket, [X: ticket text]", where "[MASK]" represents the ticket category that the MLM task needs to predict.

[0109] In the Prompt paradigm template, "[MASK]" is typically a single word in English; however, the smallest granularity of Chinese segmentation is a single character. Therefore, in the recognition prompt statements, Chinese and English are not aligned; for example, "This is a communication ticket" should correspond to "This is a [MASK][MASK] ticket," and "This is a bank draft ticket" should correspond to "This is a [MASK][MASK][MASK][MASK] ticket." Since the number of "[MASK]" in these two recognition prompt statements is inconsistent, this application embodiment can map the ticket type to be predicted (i.e., [Z]) to a single letter; for example, "communication" type can be represented by "a," "ID card" type by "b," "bank draft" type by "c," and so on. Here, a single letter does not represent any actual meaning, but only indicates the symbol of that category, thereby improving the model's computational speed.

[0110] In this embodiment, target ticket information is obtained, and a preset recognizer is used to process the target ticket information to obtain a predicted ticket category. The preset recognizer is trained based on sample recognition information, which includes latent information and sample ticket information. The latent information is determined by randomly encoding the recognition prompt statement. This application achieves improved recognition accuracy by using a preset recognizer trained on the Prompt paradigm to identify tickets with few samples. Furthermore, by determining the Prompt paradigm template using randomly encoded latent vectors, this application achieves greater flexibility in the Prompt paradigm compared to a fixed-encoding template.

[0111] In one embodiment of this application, obtaining the target ticket information of the ticket to be identified includes:

[0112] The multimedia file of the ticket to be identified is subjected to text recognition to determine the text information in the multimedia file;

[0113] The text information is processed to obtain the target ticket information of the ticket to be identified.

[0114] In some optional implementations, the target ticket information can be extracted by performing optical character recognition (OCR) on the multimedia file of the ticket to be identified. As an example, the multimedia file can be an image of the ticket to be identified. First, text detection can be performed on the multimedia file to extract the text information. Then, the text information can be recognized to obtain the target ticket information. The OCR recognition process can be found in [link to OCR documentation]. Figure 4 As shown.

[0115] In one embodiment of this application, before obtaining the target ticket information of the ticket to be identified, the method further includes:

[0116] Obtain a training sample set; the training sample set includes sample data of the sample tickets;

[0117] The sample identification information is determined based on the hidden information and the sample ticket information in the sample data;

[0118] The sample identification information is input into the initial identification model to obtain the predicted sample category corresponding to each sample ticket;

[0119] The training loss value is determined based on the predicted sample category, the reference ticket category, and the prediction accuracy corresponding to the implicit information.

[0120] Based on the training loss value, the initial recognition model is trained to obtain the preset recognizer that meets the training conditions.

[0121] Training conditions may include reaching a preset number of iterations, convergence of the training loss value, and the training loss value being less than a preset value.

[0122] This application embodiment may also include a step of training the preset recognizer (the algorithm framework of the preset recognizer can be combined with...) Figure 5 As shown), specifically, sample data of the sample ticket can be obtained first, and then the sample identification information can be determined based on the hidden information and the sample ticket information in the sample data. The sample identification information is then encoded to obtain the input vector of the input model.

[0123] In practical implementation scenarios, the input vector can take the form T = {P} 0:n ,y,K i ,X i}; combination Figure 6 As shown, P 0:n y represents 0-n characters (the characters are indefinite); y represents the label category of the sample ticket (the [MASK] label can be used to mask the real label); K i X represents the vector representation of the i-th keyword, where the keyword is a key word in the prompt statement, such as "ticket" or "receipt"; i This represents the vector representation of the i-th character in the sample ticket information (i ranges from 0 to m). As an example, the encoded input vector can be E(T) = {h0, h1, ..., h...} n ,e(y),e(K i ),e(X i )}. Where h i Let represent the latent vector of the i-th character (where i ranges from 0 to n), and let e(j) represent the vector of the j-th character (where j can take values ​​from y and k). i X i ).

[0124] Optionally, a lightweight CNN+ELMO model can be used to encode the recognition prompts to obtain vector representations of the interrelationships between the text. The lightweight CNN+ELMO model can be found in [reference needed]. Figure 7 As shown.

[0125] In the lightweight CNN+ELMO model encoding process, word vector construction is dynamic. The CNN obtains word vectors by performing convolution operations on characters. During convolution, multiple convolution kernels (n, token_num, embd_dim) can be constructed, where n represents the number of convolution kernels, token_num represents the character span, and embd_dim represents the vector dimension of the character. The word vectors are then transformed into vectors using softmax, where the vectors (embeddings) of the characters, including the sentence length, are randomly initialized.

[0126] Furthermore, the vector representation of each character in the sample ticket information can be represented by an L-layer bidirectional LSTM language model, which can be expressed as: L represents the number of layers; R k This represents the vector representation of each character in the sample ticket information.

[0127] By default, word embeddings are input to the 0th layer of the LSTM, and all other layers are represented by two vectors, i.e., the two vectors can be represented as h. k , This represents the forward representation of the k-th character at the j-th level. This represents the backward representation of the k-th character at level j. The two vectors are the concatenation of the forward and backward representations. For example, and Both are s-dimensional column vectors; concatenating them results in a 2s-dimensional column vector. Following this vector concatenation method, a large number of concatenated vectors will be obtained. Therefore, the more concatenated vectors and the deeper the layers, the better for eliminating word ambiguity. Furthermore, since the weights of each layer of vectors are different, adjustment coefficients can be added to each layer. Therefore, the total vector E(h) of the k-th word across all layers k ;Θ task ) can be represented as Θ task The parameter space representing the coefficient adjustment; h k This represents the concatenation of forward and backward representations; h_{k,j} represents h. k Vector representation in layer j.

[0128] Furthermore, latent vectors can be trained and optimized using LSTM models, which are time-series models that effectively integrate features from preceding and following contexts. This model can be divided into two time-series directions: a forward language model and a backward language model. The forward language model predicts the following context given the preceding context, while the backward language model predicts the preceding context given the following context.

[0129] Among them, the ELMO model in the LSTM model is a language model based on forward and backward LSTM. The ELMO model can be represented as:

[0130] Where p represents the probability that the Prompt paradigm template prediction is accurate; e i Let represent the vector of the i-th word in the Prompt paradigm template. The loss function of this model is:

[0131]

[0132] Where p represents the probability that the Prompt paradigm template predicts accurately, and Θ represents the parameter space. The parameters represent the feedforward language model. Θ represents the parameters of the backward language model. emb Represents the mapping layer parameters, e i This represents the vector of the i-th word in the Prompt paradigm template.

[0133] Then, the input vector corresponding to the sample identification information is input into the initial identification model to obtain the predicted sample category for each sample ticket. The initial identification model can be a BERT model.

[0134] The optimal latent vector can be obtained by optimizing the loss function of the entire model. For example, the optimal latent vector can be expressed as:

[0135] in, Let represent the latent vector; Losd(Bert(y,K,X)) represents the loss function of the BERT model; the BERT model is trained using the MLM task. L(p|Θ) represents the loss function that optimizes the latent vector.

[0136] The loss function implemented in this application may include two parts: a first loss function (i.e., a first loss value) determined according to the predicted sample category and the reference ticket category, namely Loss(Bert(y,K,X)) as described above; and a second loss function (i.e., a second loss value) determined according to the prediction accuracy corresponding to the latent information, namely L(p|Θ) as described above.

[0137] In this embodiment, the predicted sample category can be output via hash encoding. Specifically, after the Prompt paradigm template is input into the BERT model, a high-dimensional vector embedding of [MASK] (i.e., the predicted ticket category) can be obtained. The high-dimensional vector is E = {e i |i=1,…n}; among them, e i Represents vector elements. The resulting vector M is obtained by pairing the embeddings in pairs, where M = {(e...} k1,e k1 ,y k )|e k1 ∈E,e k2 ∈E}, e k1 This represents the hash code of the predicted [MASK] for the first ticket, e k2 This represents the hash code of the predicted [MASK] for the second ticket. (y can be used...) k =1 indicates that the two tickets belong to the same category, y k =0 indicates that the two tickets belong to different categories. The embedding is input into a two-layer linear network, using ReLU as the activation function. Then, the mean and standard deviation of the [MASK] floating-point representation are calculated, and the embedding is standardized. The standardized embedding can be represented as e. k (p t ):

[0138]

[0139] Among them, e k (p t ) represents the hash code for predicting the sample class; median represents the preset value, which, as an example, can be the median or mean in the embedding vector; p t This represents the value at the t-th position in the embedding.

[0140] Furthermore, the loss function for a two-layer linear network can be found in [reference needed]. Figure 8 , Figure 8 The loss of a twin-tower structure is shown, where the loss can be represented as:

[0141]

[0142] Where M represents the number of vector pairs obtained by pairwise pairing of the embedding; e k1 This represents the hash code of the predicted [MASK] for the first ticket; e k2 This represents the hash code of the predicted [MASK] for the second ticket; y k =1 indicates that the two tickets belong to the same category; y k =0 indicates that the two tickets are of different categories.

[0143] In one embodiment of this application, obtaining the training sample set includes:

[0144] Determine the preset information collection template corresponding to each of the sample tickets;

[0145] Obtain the sample ticket information that corresponds to the preset information collection template from the sample tickets.

[0146] Optional, combined Figures 9 to 11 As shown in the embodiments of this application, it can be achieved through... Figure 9 and Figure 10 The page shown displays a pre-built, customizable invoice recognition template, also known as a preset information collection template. This template includes pre-defined information collection items. For example, an invoice image can be uploaded via the web interface as the preset information collection template, and the image can be labeled with the information to be collected, such as invoice code, invoice number, verification code, invoice date, etc. Furthermore, as... Figure 11 As shown, a preset information collection template can be built by uploading multiple images.

[0147] In one embodiment of this application, the step of identifying the target ticket information using a preset recognizer to obtain the predicted ticket category of the ticket to be identified includes:

[0148] The target ticket information is identified and processed by a preset recognizer to obtain the ticket category code of the target ticket information;

[0149] The predicted ticket category is determined based on the ticket category code.

[0150] In this embodiment, a preset recognizer can output the ticket category code of the target ticket information, wherein the ticket category code can be a hash code. The hash method is an effective nearest neighbor retrieval method in the field of information retrieval, aiming to map data into binary hash codes while preserving the similarity of the data itself. Furthermore, the hash method has unique advantages: a small amount of binary code can represent high-dimensional space vectors, and binary codes can accelerate computation.

[0151] Then, the predicted ticket category is determined based on the ticket category code. Specifically, the similarity between the ticket category code and a preset category code can be calculated; the predicted ticket category is determined based on the preset category code with the highest similarity. For example, the Hamming distance between the ticket category code and the preset category code can be calculated, and the category corresponding to the preset category code with the smallest Hamming distance can be taken as the predicted ticket category.

[0152] In this embodiment, target ticket information is obtained, and a preset recognizer is used to process the target ticket information to obtain a predicted ticket category. The preset recognizer is trained based on sample recognition information, which includes latent information and sample ticket information. The latent information is determined by randomly encoding the recognition prompt statement. This application achieves improved recognition accuracy by using a preset recognizer trained on the Prompt paradigm to identify tickets with few samples. Furthermore, by determining the Prompt paradigm template using randomly encoded latent vectors, this application achieves greater flexibility in the Prompt paradigm compared to a fixed-encoding template.

[0153] Combination Figure 12 and Figure 13 The overall implementation process and application scenarios of the embodiments of this application are described below:

[0154] First, a ticket template is built, which is a preset information collection template, and a classifier (i.e., a preset recognizer) is trained; then, users can upload tickets, and the classifier will identify the ticket category.

[0155] The identification method of this application embodiment can be applied to various scenarios. For example, it can be specifically applied to recruitment scenarios, where applicants' identification documents may include resumes, ID cards, academic certificates, English proficiency certificates, etc. When classifying applicants' identification documents, a document information template (i.e., a preset information collection template) is first constructed, training data is uploaded, and the model undergoes automated training to obtain a classifier. Then, the identification documents to be identified are received and classified through the classifier's API interface.

[0156] like Figure 14 As shown, this is one of the flowcharts of a ticket recognition method provided in this application embodiment. The method includes the following steps:

[0157] Step S1401: Receive the target ticket information of the ticket to be identified through the page recognition template;

[0158] Step S1402: Display the predicted ticket category of the ticket to be identified;

[0159] The predicted ticket category is obtained by identifying the target ticket information through a preset recognizer.

[0160] The preset recognizer is trained based on sample recognition information;

[0161] The sample identification information includes hidden information and sample ticket information of the sample ticket; the hidden information is determined by randomly encoding the identification prompt statement.

[0162] Specifically, the embodiments of this application can be executed via a client, which may have a recognition system for ticket identification installed. The target ticket information to be identified can be received through the page recognition template of the recognition system.

[0163] The ticket to be identified may include a ticket to be identified; wherein, the ticket may include a voucher document; for example, the ticket may include a receipt, invoice, ticket, voucher, certificate, document, etc.; in this embodiment of the application, the ticket includes but is not limited to the above examples. In actual implementation scenarios, the ticket may be a document containing graphics or text, etc., and this application does not limit it in this regard.

[0164] The target ticket information includes the ticket information in the ticket to be identified; taking an invoice as an example, the invoice information may include invoice code, invoice number, verification code, invoice date, buyer's name, goods name, goods amount, goods unit price, goods quantity, etc. As an example, the ticket to be identified and the target ticket information can be referenced. Figure 3 As shown.

[0165] Optionally, the target ticket information may include all or a portion of the ticket information in the ticket to be identified.

[0166] In some alternative implementations, the target ticket information can be extracted by performing optical character recognition (OCR) on the multimedia file of the ticket to be identified. For example, the multimedia file can be an image of the ticket to be identified. First, text detection can be performed on the multimedia file to extract the text information. Then, the text information can be recognized to obtain the target ticket information.

[0167] In cases where the target ticket information is a portion of the ticket information to be identified, after detecting the text of the ticket to be identified, a portion of the information can be extracted as the target ticket information. For example, information with a higher preset weight coefficient can be extracted as the target ticket information. For instance, the invoice code, invoice number, verification code, buyer's name, goods name, goods amount, unit price, and quantity of goods can be extracted as the target ticket information.

[0168] In this embodiment of the application, the target ticket information can be identified by a preset recognizer to obtain the predicted ticket category of the ticket to be identified.

[0169] The predicted ticket category refers to the predicted category of the ticket. For example, in actual implementation scenarios, the predicted ticket category may include value-added tax invoices, scenic spot tickets, coupons, graduation certificates, degree certificates, English certificates, etc.

[0170] Optionally, the preset recognizer in this application can be trained based on the Prompt paradigm. The sample recognition information can be determined through the Prompt paradigm template. Since the Prompt paradigm performs well in model training scenarios with few or no samples, it can improve the prediction accuracy of the preset recognizer.

[0171] Specifically, the Prompt paradigm performs well in training scenarios with few or no samples. In related technologies, ticket classification is typically a supervised training task. For example, in practical applications, ticket information x is used as input data, and ticket type y is used as the model's output data to train the model P(y|x; θ), where θ is the maximum likelihood parameter, which represents the probability of accurate prediction. However, the aforementioned methods in these technologies require high-quality labeled data, resulting in high initial costs.

[0172] The training method based on the Prompt paradigm in this application can solve the above problems. The Prompt paradigm guides the model to model the text itself and predicts the ticket type based on the maximum likelihood parameter (probability), thereby reducing the dependence of the training model on a large amount of supervised data. The Prompt paradigm can be mapped by the function f. prompt The expression, i.e., x′=f prompt (x; θ); where x′ represents the expression function mapped by the prompt function; f prompt θ represents the prompt function; x represents the ticket information; θ represents the maximum likelihood parameter.

[0173] In practical implementation scenarios, the training method based on the Prompt paradigm can train the model by using the Prompt paradigm template as input. The Prompt paradigm template, as described in this embodiment, is the sample identification information, including identification prompt statements and sample ticket information. Optionally, in this embodiment, the identification prompt statements can be randomly encoded to obtain latent information (i.e., latent vectors). The Prompt paradigm template is then constructed using the latent information and sample ticket information. Thus, by determining the Prompt paradigm template through randomly encoded latent vectors, this application achieves greater flexibility in the Prompt paradigm compared to a Prompt paradigm template determined by fixed encoding.

[0174] Furthermore, during training, the downstream task can be adjusted to one of the two main tasks in the pre-trained model: the MLM task and the NSP task. Based on the input format of the MLM task, special characters such as "[MASK]" can be used to fill the [Z] position in the Prompt paradigm template. This results in "This is a [MASK] type ticket, [X: ticket text]", where "[MASK]" represents the ticket category that the MLM task needs to predict.

[0175] In the Prompt paradigm template, "[MASK]" is typically a single word in English; however, the smallest granularity of Chinese segmentation is a single character. Therefore, in the recognition prompt statements, Chinese and English are not aligned; for example, "This is a communication ticket" should correspond to "This is a [MASK][MASK] ticket," and "This is a bank draft ticket" should correspond to "This is a [MASK][MASK][MASK][MASK] ticket." Since the number of "[MASK]" in these two recognition prompt statements is inconsistent, this application embodiment can map the ticket type to be predicted (i.e., [Z]) to a single letter; for example, "communication" type can be represented by "a," "ID card" type by "b," "bank draft" type by "c," and so on. Here, a single letter does not represent any actual meaning, but only indicates the symbol of that category, thereby improving the model's computational speed.

[0176] In this embodiment, target ticket information is received, and a preset recognizer is used to process the target ticket information to obtain a predicted ticket category. The preset recognizer is trained based on sample recognition information, which includes latent information and sample ticket information. The latent information is determined by randomly encoding the recognition prompt statement. This application achieves improved recognition accuracy by using a preset recognizer trained on the Prompt paradigm to identify tickets with few samples. Furthermore, by determining the Prompt paradigm template using randomly encoded latent vectors, this application achieves greater flexibility in the Prompt paradigm compared to a fixed-encoding template.

[0177] This application provides a ticket recognition device, such as... Figure 15 As shown, the ticket recognition device 150 may include: an acquisition module 1501 and a recognition module 1502, wherein,

[0178] The acquisition module 1501 is used to acquire the target ticket information of the ticket to be identified;

[0179] The identification module 1502 is used to identify the target ticket information through a preset recognizer to obtain the predicted ticket category of the ticket to be identified.

[0180] The preset recognizer is obtained by training based on sample recognition information;

[0181] The sample identification information includes hidden information and sample ticket information of the sample ticket; the hidden information is determined by randomly encoding the identification prompt statement.

[0182] In one embodiment of this application, obtaining the target ticket information of the ticket to be identified includes:

[0183] Image recognition is performed on the multimedia file of the ticket to be identified to determine the text information in the multimedia file;

[0184] The text information is processed to obtain the target ticket information of the ticket to be identified.

[0185] In one embodiment of this application, before obtaining the target ticket information of the ticket to be identified, the method further includes:

[0186] Obtain a training sample set; the training sample set includes sample data of the sample tickets;

[0187] The sample identification information is determined based on the hidden information and the sample ticket information in the sample data;

[0188] The sample identification information is input into the initial identification model to obtain the predicted sample category corresponding to each sample ticket;

[0189] The training loss value is determined based on the predicted sample category, the reference ticket category, and the prediction accuracy corresponding to the implicit information.

[0190] Based on the training loss value, the initial recognition model is trained to obtain the preset recognizer that meets the training conditions.

[0191] In one embodiment of this application, determining the training loss value includes:

[0192] A first loss value is determined based on the predicted sample category and the reference ticket category;

[0193] The second loss value is determined based on the prediction accuracy corresponding to the implicit information.

[0194] The training loss value includes the first loss value and the second loss value.

[0195] In one embodiment of this application, obtaining the training sample set includes:

[0196] Determine the preset information collection template corresponding to each of the sample tickets;

[0197] Obtain the sample ticket information that corresponds to the preset information collection template from the sample tickets.

[0198] In one embodiment of this application, the step of identifying the target ticket information using a preset recognizer to obtain the predicted ticket category of the ticket to be identified includes:

[0199] The target ticket information is identified and processed by a preset recognizer to obtain the ticket category code of the target ticket information;

[0200] The predicted ticket category is determined based on the ticket category code.

[0201] In one embodiment of this application, determining the predicted ticket category based on the ticket category code includes:

[0202] Calculate the similarity between the ticket category code and the preset category code;

[0203] The predicted ticket category is determined based on the preset category code with the highest similarity.

[0204] The apparatus in this application embodiment can execute the method provided in this application embodiment, and the implementation principle is similar. The actions performed by each module in the apparatus of each embodiment of this application correspond to the steps in the method of each embodiment of this application. For detailed functional descriptions of each module of the apparatus, please refer to the descriptions in the corresponding methods shown above, which will not be repeated here.

[0205] In this embodiment, target ticket information is obtained, and a preset recognizer is used to process the target ticket information to obtain a predicted ticket category. The preset recognizer is trained based on sample recognition information, which includes latent information and sample ticket information. The latent information is determined by randomly encoding the recognition prompt statement. This application achieves improved recognition accuracy by using a preset recognizer trained on the Prompt paradigm to identify tickets with few samples. Furthermore, by determining the Prompt paradigm template using randomly encoded latent vectors, this application achieves greater flexibility in the Prompt paradigm compared to a fixed-encoding template.

[0206] This application provides an electronic device comprising: a memory and a processor; at least one program stored in the memory, which, when executed by the processor, can achieve the following compared to existing technologies: In this application, by acquiring target ticket information of a ticket to be identified, the target ticket information is identified using a preset recognizer to obtain a predicted ticket category of the ticket to be identified; wherein, the preset recognizer is trained based on sample recognition information; the sample recognition information includes latent information and sample ticket information of sample tickets; the latent information is determined by randomly encoding the recognition prompt statement. This application achieves the identification of tickets to be identified using a preset recognizer trained based on the Prompt paradigm in cases with few samples, improving the accuracy of identification. Furthermore, by determining the Prompt paradigm template through randomly encoded latent vectors, this application achieves greater flexibility in the Prompt paradigm compared to a Prompt paradigm template determined by fixed encoding.

[0207] In one alternative embodiment, an electronic device is provided, such as Figure 16 As shown, Figure 16 The illustrated electronic device 4000 includes a processor 4001 and a memory 4003. The processor 4001 and the memory 4003 are connected, for example, via a bus 4002. Optionally, the electronic device 4000 may further include a transceiver 4004, which can be used for data interaction between the electronic device and other electronic devices, such as sending and / or receiving data. It should be noted that in practical applications, the transceiver 4004 is not limited to one type, and the structure of the electronic device 4000 does not constitute a limitation on the embodiments of this application.

[0208] Processor 4001 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 4001 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.

[0209] Bus 4002 may include a pathway for transmitting information between the aforementioned components. Bus 4002 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. Bus 4002 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 16 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0210] The memory 4003 may be ROM (Read Only Memory) or other types of static storage devices capable of storing static information and instructions, RAM (Random Access Memory) or other types of dynamic storage devices capable of storing information and instructions, or EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto.

[0211] The memory 4003 stores application code (computer program) that executes the solution of this application, and its execution is controlled by the processor 4001. The processor 4001 executes the application code stored in the memory 4003 to implement the content shown in the foregoing method embodiments.

[0212] Electronic devices include, but are not limited to: mobile phones, laptops, multimedia players, desktop computers, etc.

[0213] This application provides a computer-readable storage medium storing a computer program that, when run on a computer, enables the computer to execute the corresponding content in the aforementioned method embodiments.

[0214] In this embodiment, target ticket information is obtained, and a preset recognizer is used to process the target ticket information to obtain a predicted ticket category. The preset recognizer is trained based on sample recognition information, which includes latent information and sample ticket information. The latent information is determined by randomly encoding the recognition prompt statement. This application achieves improved recognition accuracy by using a preset recognizer trained on the Prompt paradigm to identify tickets with few samples. Furthermore, by determining the Prompt paradigm template using randomly encoded latent vectors, this application achieves greater flexibility in the Prompt paradigm compared to a fixed-encoding template.

[0215] The terms "first," "second," "third," "fourth," "1," "2," etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in a sequence other than that shown in the figures or text.

[0216] It should be understood that although arrows indicate various operation steps in the flowcharts of this application's embodiments, the order in which these steps are implemented is not limited to the order indicated by the arrows. Unless explicitly stated herein, in some implementation scenarios of this application's embodiments, the implementation steps in each flowchart can be executed in other orders as required. Furthermore, some or all steps in each flowchart, based on the actual implementation scenario, may include multiple sub-steps or multiple stages. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage can also be executed at different times. In scenarios where execution times differ, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and this application's embodiments do not limit this.

[0217] The above description is only an optional implementation method for some implementation scenarios of this application. It should be noted that for those skilled in the art, other similar implementation methods based on the technical concept of this application without departing from the technical concept of this application also fall within the protection scope of the embodiments of this application.

Claims

1. A ticket recognition method, characterized in that, include: Obtain the target ticket information for the ticket to be identified; Before obtaining the target ticket information of the ticket to be identified, the method further includes: Obtain a training sample set; the training sample set includes sample data of sample tickets; The sample identification information is determined based on the hidden information and the sample ticket information in the sample data; The sample identification information is input into the initial identification model to obtain the predicted sample category corresponding to each sample ticket; The training loss value is determined based on the predicted sample category, the reference ticket category, and the prediction accuracy corresponding to the implicit information. Based on the training loss value, the initial recognition model is trained to obtain a preset recognizer that meets the training conditions; The target ticket information is identified by a preset recognizer to obtain the predicted ticket category of the ticket to be identified. The preset recognizer is obtained by training based on sample recognition information; The sample identification information includes hidden information and sample ticket information; the hidden information is determined by randomly encoding the identification prompt statement; the hidden information is based on the loss function of the BERT model. Confirmed, as shown in the following formula: argmin h Loss(Bert(y,K,X))- L(p|Θ); Wherein, represents implicit information; Loss(Bert(y,K,X)) This represents the loss function of the BERT model. L(p| Θ) The loss function representing the optimization of hidden information, the Loss(Bert(y,K,X)) The predicted sample category and the reference ticket category are determined accordingly. L(p|Θ) The accuracy is determined based on the prediction accuracy corresponding to the implicit information.

2. The ticket recognition method according to claim 1, characterized in that, The acquisition of the target ticket information for the ticket to be identified includes: The multimedia file of the ticket to be identified is subjected to text recognition to determine the text information in the multimedia file; The text information is processed to obtain the target ticket information of the ticket to be identified.

3. The ticket recognition method according to claim 1, characterized in that, Determining the training loss value includes: A first loss value is determined based on the predicted sample category and the reference ticket category; The second loss value is determined based on the prediction accuracy corresponding to the implicit information; The training loss value includes the first loss value and the second loss value.

4. The ticket recognition method according to claim 1, characterized in that, The acquisition of the training sample set includes: Determine the preset information collection template corresponding to each of the sample tickets; Obtain the sample ticket information that corresponds to the preset information collection template from the sample tickets.

5. The ticket recognition method according to claim 1, characterized in that, The step of identifying the target ticket information using a preset recognizer to obtain the predicted ticket category of the ticket to be identified includes: The target ticket information is identified and processed by a preset recognizer to obtain the ticket category code of the target ticket information; The predicted ticket category is determined based on the ticket category code.

6. The ticket recognition method according to claim 5, characterized in that, Determining the predicted ticket category based on the ticket category code includes: Calculate the similarity between the ticket category code and the preset category code; The predicted ticket category is determined based on the preset category code with the highest similarity.

7. The ticket recognition method according to claim 1, characterized in that, The preset recognizer is trained based on the Prompt paradigm. The sample identification information is determined based on the Prompt paradigm template.

8. A ticket recognition method, characterized in that, include: Receive the target ticket information of the ticket to be identified through the page recognition template; Before receiving the target ticket information of the ticket to be identified through the page recognition template, the method further includes: Obtain a training sample set; the training sample set includes sample data of sample tickets; The sample identification information is determined based on the hidden information and the sample ticket information in the sample data; The sample identification information is input into the initial identification model to obtain the predicted sample category corresponding to each sample ticket; The training loss value is determined based on the predicted sample category, the reference ticket category, and the prediction accuracy corresponding to the implicit information. Based on the training loss value, the initial recognition model is trained to obtain a preset recognizer that meets the training conditions; Displays the predicted ticket category of the ticket to be identified; The predicted ticket category is obtained by identifying the target ticket information through a preset recognizer. The preset recognizer is obtained by training based on sample recognition information; The sample identification information includes hidden information and sample ticket information; the hidden information is determined by randomly encoding the identification prompt statement; the hidden information is based on the loss function of the BERT model. Confirmed, as shown in the following formula: argmin h Loss(Bert(y,K,X))- L(p|Θ); Wherein, represents implicit information; Loss(Bert(y,K,X)) This represents the loss function of the BERT model. L(p| Θ) The loss function representing the optimization of hidden information, the Loss(Bert(y,K,X)) The predicted sample category and the reference ticket category are determined accordingly. L(p|Θ) The accuracy is determined based on the prediction accuracy corresponding to the implicit information.

9. A ticket recognition device, characterized in that, include: The acquisition module is used to acquire the target ticket information of the ticket to be identified; The acquisition module is also used to acquire a training sample set; The training sample set includes sample data of sample tickets; The sample identification information is determined based on the hidden information and the sample ticket information in the sample data; The sample identification information is input into the initial identification model to obtain the predicted sample category corresponding to each sample ticket; The training loss value is determined based on the predicted sample category, the reference ticket category, and the prediction accuracy corresponding to the implicit information. Based on the training loss value, the initial recognition model is trained to obtain a preset recognizer that meets the training conditions; The identification module is used to identify the target ticket information through a preset recognizer to obtain the predicted ticket category of the ticket to be identified; The preset recognizer is obtained by training based on sample recognition information; The sample identification information includes hidden information and sample ticket information; the hidden information is determined by randomly encoding the identification prompt statement; the hidden information is based on the loss function of the BERT model. Confirmed, as shown in the following formula: argmin h Loss(Bert(y,K,X))- L(p|Θ); Wherein, represents implicit information; Loss(Bert(y,K,X)) This represents the loss function of the BERT model. L(p| Θ) The loss function representing the optimization of hidden information, the Loss(Bert(y,K,X)) The predicted sample category and the reference ticket category are determined accordingly. L(p|Θ) The accuracy is determined based on the prediction accuracy corresponding to the implicit information.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the steps of the ticket recognition method according to any one of claims 1-8.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the ticket recognition method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Prompt-based building entity identification and classification method and system

    CN115859164A

  • Small sample relation classification method and system based on prompt learning, medium and electronic equipment

    CN115982363A

  • Image recognition method and device

    CN116189201A