Method for determining text label and related device

By splicing the target text in the target picture with the target text box feature vector and inputting a pre-trained label determination model, the problem of low efficiency and low accuracy of picture text label determination in the prior art is solved, and higher label determination accuracy is achieved.

CN114359913BActive Publication Date: 2025-06-06SHENZHEN IDEAMAKE SOFTWARE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210004883.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-04
Publication Date
2025-06-06
Estimated Expiration
2042-01-04

AI Technical Summary

Technical Problem

In the prior art, the extraction and recognition of picture text are two independent processes, resulting in low efficiency and low accuracy of labeling corresponding labels to the text in the picture.

Method used

By splicing the target text in the target picture with the target text box feature vector, input pre-trained labels to determine the model, and determine the label of the target text. The model is trained from multiple training texts, data spliced ​​by the feature vectors of the training text box, and a preset second label.

Benefits of technology

Improve the accuracy of determining text tags, and enhance the recognition ability of the tag determination model by combining the feature information of text and text boxes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114359913B_ABST
    Figure CN114359913B_ABST
Patent Text Reader

Abstract

The embodiment of the present application discloses a method for determining a text label and a related device, the method comprising: obtaining a target text in a target image and a target text box feature vector corresponding to the target text, inputting the target text and the target text box feature vector into a pre-trained label determination model, determining a first label corresponding to the target text, the label determination model is trained by data obtained by splicing a plurality of training texts, training text box feature vectors corresponding to the training texts, and a second label corresponding to the training text, the second label being a pre-set label. The present application splices the training text box feature vector and the training text, trains the label determination model through the spliced ​​training text box feature vector and the training text, so that the trained label determination model outputs the second label corresponding to the training text, thereby improving the accuracy of determining the text label.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a method for determining a text label and a related device. Background Art

[0002] With the development of science and technology, the scale of multimedia resources including text and pictures is getting larger and larger. Text retrieval has gradually become a research hotspot in the field of natural language processing, and many text retrieval methods based on optical character recognition (OCR) technology have been produced. This method recognizes the text content from the picture and then uses text retrieval technology to implement a text and picture retrieval system. With the existing picture extraction technology, after extracting the text from the picture, the text information is extracted, and the text is labeled according to the results of recognition and extraction. The extraction of picture text and the recognition of text are two independent processes, resulting in low efficiency and low accuracy in labeling the text in the picture. Summary of the invention

[0003] The embodiment of the present application provides a method and related device for determining a text label, which can determine the label of the target text in the target image after splicing the target text in the target image with the target text box feature vector, thereby improving the accuracy of determining the text label.

[0004] In a first aspect, an embodiment of the present application provides a method for determining a text label, the method comprising:

[0005] Get the target image;

[0006] Extracting target text in the target image and a target text box feature vector corresponding to the target text;

[0007] The target text and the target text box feature vector are input into a pre-trained label determination model to determine a first label corresponding to the target text. The label determination model is trained by data obtained by splicing multiple training texts, training text box feature vectors corresponding to the training texts, and a second label corresponding to the training text. The training text box feature vector includes the vertex coordinates of the area where the training text is located in the training image, and the ratio of the hypotenuse length of the area to the hypotenuse length of the training image. The second label is a pre-set label.

[0008] In a second aspect, an embodiment of the present application provides a device for determining a text label, the device comprising:

[0009] A first acquisition unit, used to acquire a target image;

[0010] An extraction unit, used to extract the target text in the target image and a target text box feature vector corresponding to the target text;

[0011] A first input unit is used to input the target text and the target text box feature vector into a pre-trained label determination model to determine a first label corresponding to the target text, wherein the label determination model is trained by data obtained by splicing multiple training texts, training text box feature vectors corresponding to the training texts, and a second label corresponding to the training text, wherein the training text box feature vector includes the vertex coordinates of the area where the training text is located in the training image, and the ratio of the hypotenuse length of the area to the hypotenuse length of the training image, and the second label is a pre-set label.

[0012] In a third aspect, an embodiment of the present application provides a terminal device, comprising a processor, a memory, a communication interface, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the processor, and the program includes instructions for executing some or all of the steps described in the method described in the first aspect above.

[0013] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium that stores a computer program for electronic data exchange, wherein the computer program enables a computer to execute some or all of the steps in the first aspect of the embodiment.

[0014] It can be seen that in this embodiment, the technical solution provided by this application, after obtaining the target image, extracts the target text in the target image and the target text box feature vector corresponding to the target text; inputs the target text and the target text box feature vector into a pre-trained label determination model to determine the first label corresponding to the target text. Among them, the label determination model is obtained by training a plurality of training texts, data obtained by splicing the training text box feature vectors corresponding to the training texts, and the second label corresponding to the training text. The training text box feature vector includes the vertex coordinates of the area where the training text is located in the training image, and the ratio of the hypotenuse length of the area to the hypotenuse length of the training image. The second label is a pre-set label. Through the training text, the vertex coordinates of the area where the training text is located, and the ratio of the hypotenuse length of the area where the training text is located to the hypotenuse length of the training image, it is ensured that the training text is the text corresponding to the second label, thereby improving the recognition accuracy of the label determination model. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0016] Figure 1 This is a schematic diagram of an application architecture provided by an embodiment of the present application;

[0017] Figure 2 It is a flowchart of a method for determining a text label provided in an embodiment of the present application;

[0018] Figure 3 is a schematic diagram of a method for determining a text label provided in an embodiment of the present application;

[0019] Figure 4 This is a schematic diagram of a label determination model training process provided by an embodiment of the present application;

[0020] Figure 5 It is a schematic diagram of the label determination model training process provided in an embodiment of the present application;

[0021] Figure 6 This is a block diagram of the functional units of a device for determining a text label provided in an embodiment of the present application;

[0022] Figure 7 It is a structural diagram of a terminal device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0023] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0024] The terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally includes steps or units that are not listed, or optionally includes other steps or units inherent to these processes, methods, products or devices.

[0025] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.

[0026] In order to facilitate understanding of the technical solution provided by this application, firstly, the relevant concepts involved in this application are explained.

[0027] First label: The first label is used to identify the target text information of the target text. For example, the content of the first label is: house area, and the recognized target text is: 80 square meters. After recognizing the target text, it is determined that the label corresponding to the target text should be house area, so the first label is added to the target text, indicating that the area of ​​the house is 80.

[0028] Second label: The second label is used to identify the training text information of the training text. For example, the second label is: apartment type name, and the recognized training text is: XX household. After recognizing the target text, it is determined that the label corresponding to the training text should be the apartment type name, and the first label is added to the training text, indicating that the apartment type name is XX household.

[0029] It can be understood that the first label and the second label are pre-set labels, which can be set according to actual needs. For example, they can also be project names, apartment sizes, etc., which are not specifically limited here.

[0030] A training text box feature vector includes the vertex coordinates of the training text area in the training image, and the ratio of the hypotenuse length of the training text area to the hypotenuse length of the training image. The training text box feature vector also includes other information obtained when extracting the training image, such as the relative confidence of the training text recognition confidence relative to the training text box area, etc. The label determination model is trained by using the vertex coordinates and ratio information in the training text box feature vector to improve the accuracy of the label determination model in determining text labels.

[0031] The target text box feature vector includes the vertex coordinates of the target text area in the target image, and the ratio of the hypotenuse length of the target text area to the hypotenuse length of the target image. The target text box feature vector also includes other information obtained when extracting the target image, such as the relative confidence of the target text recognition confidence relative to the target text box area, etc. The label of the target text is determined by the vertex coordinates and ratio information in the target text box feature vector, thereby improving the accuracy of the determination result.

[0032] Optical Character Recognition (OCR) refers to the process in which an electronic device (such as a scanner or digital camera) examines characters printed on paper, determines their shapes by detecting dark and light patterns, and then uses character recognition methods to translate the shapes into computer text; that is, for printed characters, an optical method is used to convert the text in a paper document into a black and white dot matrix image file, and recognition software is used to convert the text in the image into text format for further editing and processing by word processing software.

[0033] In an embodiment of the present application, a method for determining a text label is provided, the method comprising:

[0034] Acquire a target image; extract a target text in the target image and a target text box feature vector corresponding to the target text; input the target text and the target text box feature vector into a pre-trained label determination model to determine a first label corresponding to the target text, wherein the label determination model is trained by concatenating data of multiple training texts, training text box feature vectors corresponding to the training texts, and a second label corresponding to the training text, wherein the training text box feature vector includes the vertex coordinates of the area where the training text is located in the training image, and the ratio of the hypotenuse length of the area to the hypotenuse length of the training image, and the second label is a pre-set label.

[0035] By using the training text, the vertex coordinates of the area where the training text is located, and the ratio of the hypotenuse length of the area where the training text is located to the hypotenuse length of the training image, it is ensured that the training text is the label corresponding to the second label, thereby improving the accuracy of label determination model recognition.

[0036] See also Figure 1 , Figure 1 A schematic diagram of an application architecture provided for an embodiment of the present application includes a server 110 and a terminal device 120. The terminal device 120 may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, etc., but is not limited thereto. Various applications (Application, APP) may be installed on the terminal device 120, such as an OCR recognition program.

[0037] The server 110 can provide various network services for the terminal device 120. For different applications, the server 110 can be a corresponding background server. The server 110 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.

[0038] The terminal device 120 and the server 110 may be directly or indirectly connected via wired or wireless communication, which is not limited in this application. For example, the terminal device 120 and the server 110 are connected via the Internet to achieve mutual communication. Optionally, the above-mentioned Internet uses standard communication technology and / or protocols. The Internet is usually the Internet, but it can also be any network, including but not limited to a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a private network or any combination of a virtual private network.

[0039] It should be noted that the label determination method in the embodiment of the present application is mainly executed by the terminal device 120 side. For example, the label determination model is located on the terminal device 120. The user inputs the target image on the terminal device 120, and extracts the target text and the target text box feature vector of the target image. The terminal device 120 identifies the target text and the target text box feature vector through the label determination model, determines the label of the target text, and outputs the target text and the label corresponding to the target text. Specifically: the terminal device 120 obtains the target image; the target text and the target text box feature vector in the target image are input into the label determination model to obtain the first label corresponding to the target text. The target text and the target text box feature vector can be obtained through models such as OCR models or twin tower models. The target text box feature vector includes the vertex coordinates of the area where the target text is located in the target image, and the ratio of the hypotenuse length of the area where the target text is located to the hypotenuse length of the training image. The vertex coordinates are the coordinates corresponding to the target text. Different target texts have different vertex coordinates, and the vertex coordinates and other information such as the ratio are spliced ​​into the target text and then input into the label determination model to obtain the first label, thereby improving the accuracy of determination.

[0040] like Figure 1The application architecture shown is explained by taking the application on the terminal device 120 side as an example. Of course, the speech recognition method in the embodiment of the present application can also be executed by the server 110. For example, the label determination model is located on the server 110. The user inputs the target image on the terminal device 120 and needs to extract the target text and the target text box feature vector in the target image. The terminal device 120 sends the target text and the target text box feature vector to the server 110 and sends a confirmation request. After receiving the confirmation request, the server 110 determines the label of the target text for the target text and the target text box feature vector through the label determination model, and after determining the first label of the target text, outputs the first label and the target text. Specifically: the server 110 obtains the target text and the target text box feature vector, inputs the target text and the target text box feature vector into the label determination model, and obtains the first label.

[0041] The application architecture diagram in the embodiment of the present application is intended to more clearly illustrate the technical solution in the embodiment of the present application, and does not constitute a limitation on the technical solution provided in the embodiment of the present application. For other application architectures and applications, the technical solution provided in the embodiment of the present application is also applicable to similar problems.

[0042] Based on the above examples, please refer to Figure 2 , Figure 2 This is a flow chart of a method for determining a text label in an embodiment of the present application, and the method includes the following steps.

[0043] S210: Acquire a target image.

[0044] It can be understood that the target image in the embodiments of the present application can be different types of images, such as floor plans, engineering drawings, etc. The embodiments of the present application are mainly described using floor plans as an example.

[0045] S220: Extracting the target text in the target image and the target text box feature vector corresponding to the target text.

[0046] Among them, the target image includes multiple target texts, each of the multiple target texts has a corresponding target text box feature vector, and each target text in the target image and the target text box feature vector corresponding to the target text are extracted to prepare for the subsequent input label determination model.

[0047] The target text box feature vector includes the vertex coordinates of the area where the target text is located in the target image, and the ratio of the hypotenuse length of the area to the hypotenuse length of the target image. Among them, each target text and the target text box feature vector corresponding to each target text can be extracted through the OCR model. Among them, the OCR model includes a text extraction model and a text recognition model. The text extraction model can detect the text line in the target image through methods such as DB and PSENET. The text recognition model uses algorithms such as CRNN and RARE to recognize the text line and then recognize the specific text.

[0048] Specifically, after the terminal device 120 obtains the target image, the target text in the target image is obtained through the OCR model. When there are multiple target texts obtained, the multiple target texts are collected in the same list, and the information of the area where each target text is located is obtained, such as the four vertex coordinates of the area where the target text is located, the spacing distance between adjacent areas, the line information of the area where the target text is located, and the confidence of text recognition, etc., to form a vector table. Each vector in the vector table is processed to obtain the target text box feature vector. The specific processing includes: after obtaining the vertex coordinates of the target image, the length and width of the target image are calculated according to the vertex coordinates of the target image, and then the hypotenuse length of the target image is obtained. According to the vertex coordinates of the area where the target text is located, the length and width of the area where the target text is located are calculated, and then the hypotenuse length of the area where the target text is located is obtained. The hypotenuse length of the area where the target text is located is divided by the hypotenuse length of the target image to obtain the ratio of the hypotenuse length of the area where the target text is located to the hypotenuse length of the target image, and the target text box feature vector includes the ratio and the vertex coordinates. Obtaining the line information of the area where the target text is located includes: obtaining the thickness of the line in the area where the target text is located, the line color and other information. Furthermore, the symbol information in the area where the target text is located or in the set range of the target text can also be obtained. After the symbol is obtained, the corresponding symbol information is searched in the symbol set. For example, the symbol of a door is obtained, and the symbol is enlarged or reduced, and the same symbol is found in the symbol set, so that the corresponding information of the symbol is obtained as a door, and the symbol information of the door is output to determine the first label. The symbol set is a pre-set symbol information with different symbols corresponding to each of the different symbols.

[0049] Furthermore, the target text and the target text box feature vector may be preprocessed to remove repeated target text and target text box feature vectors.

[0050] S230: Input the target text and the target text box feature vector into a pre-trained label determination model to determine a first label corresponding to the target text, wherein the label determination model is trained by data obtained by splicing multiple training texts, training text box feature vectors corresponding to the training texts, and a second label corresponding to the training texts, wherein the training text box feature vector includes the vertex coordinates of the area where the training text is located in the training image, and the ratio of the hypotenuse length of the area to the hypotenuse length of the training image, and the second label is a pre-set label.

[0051] See also Figure 3 , Figure 3 It is a schematic diagram of a method for determining a text label provided by an embodiment of the present application. The acquired target text and target text box feature vector are input into a pre-trained label determination model to determine the first label corresponding to the target text, and the first label is a label that identifies the target text. For example, the target image is a floor plan, and the target text in the floor plan is extracted: 30 square meters. The target text is determined to be the floor area through the target text and the target text box feature vector, and the first label is determined, and the content of the first label is the floor area. Since different texts in the floor plan have their prescribed formats and position requirements, the first label corresponding to the target text is determined by the information obtained by splicing the target text box feature vector and the target text, thereby improving the accuracy of the determination.

[0052] The label determination model is obtained by splicing data of multiple training texts, training text box feature vectors corresponding to the training texts, and the second label corresponding to the training text. The training text box feature vector includes the vertex coordinates of the area where the training text is located in the training image, and the ratio of the hypotenuse length of the area to the hypotenuse length of the training image. The second label is a pre-set label. Different texts have their prescribed format and position requirements. Therefore, the label determination model is determined by training the spliced ​​training text box feature vector and the training text to improve the accuracy of the label determination model.

[0053] See also Figure 4 and Figure 5 , Figure 4 is a schematic diagram of a label determination model training process provided in an embodiment of the present application, Figure 5 It is a schematic diagram of the label determination model training process provided in an embodiment of the present application.

[0054] S410: Acquire a plurality of the training texts, the training text box feature vectors corresponding to the training texts, and the second labels corresponding to the training texts.

[0055] Before obtaining a plurality of training texts, a training picture is obtained, wherein the training picture contains at least one training text and a training text box feature vector corresponding to the training text, and a second label is pre-set for each training text.

[0056] Extract all the training texts in the training image: text1, text2, etc. to form a training text list [text1, text2…textn], and the information of the area where each training text is located, such as the four vertex coordinates of the area where the training text is located, and the confidence of text recognition, etc., to form a vector table [info1, info2…info3]. Process the information of the area where each training text is located in the vector table into a training text box feature vector, and form multiple training text box feature vectors [info_feat1, info_feat2, …, info_featm]. The specific processing process is: after obtaining the vertex coordinates of the training image, calculate the length and width of the training image according to the vertex coordinates of the training image, and then obtain the hypotenuse length of the training image; calculate the length and width of the area where the training text is located according to the vertex coordinates of the area where the training text is located, and then obtain the hypotenuse length of the area where the training text is located. Divide the hypotenuse length of the area where the training text is located by the hypotenuse length of the training image to obtain the ratio of the hypotenuse length of the area where the training text is located to the hypotenuse length of the training image. The training text box feature vector includes the vertex coordinates and the ratio.

[0057] Before concatenating the training text with the training text box feature vector, the method further includes: preprocessing the training text, the training text box feature vector and the second label to remove duplicate training text, the training text box feature vector and the second label.

[0058] S420: Concatenate the training text with the training text box feature vector to obtain a first matrix.

[0059] Among them, the splicing of the training text with the training text box feature vector to obtain the first matrix includes: converting the training text into text numbers through a text dictionary; converting the text numbers into text vectors through a semantic representation model to obtain a second matrix, and the second matrix includes the text vector; converting the training text box feature vector into a third matrix with a dimension of 1 and a fourth matrix with a dimension of k through a fully connected neural network, wherein k is a hyperparameter; determining a fifth matrix according to the third matrix and the second matrix, and determining a sixth matrix according to the fourth matrix and the second matrix; and determining the first matrix according to the fifth matrix and the sixth matrix.

[0060] Specifically, the training text is converted into text numbers through a text dictionary; each training text [text1, text2...textn] in the list is converted into a text number through a text dictionary.

[0061] The text numbers are converted into text vectors through the semantic representation model to obtain a second matrix, which includes the text vectors; each text number is converted into a text vector through the semantic representation model to obtain a second matrix [seq1 seq2…seq3]. The training text box feature vector is converted into a third matrix with a dimension of 1 and a fourth matrix with a dimension of k through a fully connected neural network, where k is a hyperparameter; the fifth matrix is ​​determined based on the third matrix and the second matrix, and the sixth matrix is ​​determined based on the fourth matrix and the second matrix; the first matrix is ​​determined based on the fifth matrix and the sixth matrix.

[0062] Among them, the training text box feature vectors [info_feat1, info_feat2, …, info_featm] are converted into a third matrix m with a dimension of 1 through a fully connected neural network. 1 and a fourth matrix m2 of dimension k, where k is a hyperparameter. A hyperparameter is an unknown variable, but it is different from a parameter in the training process. It is a parameter that can affect the parameters obtained by training. It needs to be manually input by the trainer and adjusted to optimize the effect of the training model. The fifth matrix is ​​determined according to the third matrix and the second matrix, and the sixth matrix is ​​determined according to the fourth matrix and the second matrix. Optionally, the fifth matrix is ​​determined according to the third matrix and the second matrix, and the sixth matrix is ​​determined according to the fourth matrix and the second matrix, including: splicing the third matrix before the second matrix to obtain the fifth matrix by the CONCATENATE function, and splicing the fourth matrix after the second matrix to obtain the sixth matrix. That is, the CONCATENATE method can be used to splice m1 before [seq1 seq2…seq3] to form the fifth matrix, and the CONCATENATE method can be used to splice m2 after [seq1seq2…seq3] to form the sixth matrix, and the first matrix [s1, s2,…, sn] is determined according to the fifth matrix and the sixth matrix.

[0063] In other embodiments, the third matrix may be concatenated before the second matrix to obtain a fifth matrix, and the fourth matrix may be concatenated after the second matrix to obtain a sixth matrix, by using a dimension expansion and element addition method.

[0064] Specifically, before determining the fifth matrix according to the third matrix and the second matrix, it also includes: processing the third matrix by applying an activation function to obtain a seventh matrix; determining the fifth matrix according to the third matrix and the second matrix includes: determining the fifth matrix according to the seventh matrix and the second matrix.

[0065] The m1 activation function is processed to obtain the seventh matrix. The activation function introduces nonlinear factors to the neuron, so that the neural network can arbitrarily approximate any nonlinear function, so that the neural network can be applied to many nonlinear models. The activation functions include Sigmoid function or Tanh function.

[0066] Among them, the seventh matrix is ​​spliced ​​in [seq 1 ,seq 2 …seq n ] before forming the fifth matrix.

[0067] S430: Input the first matrix into the label determination model to obtain a third label, and adjust the label determination model according to the difference between the third label and the second label until the training end condition is met to obtain the label determination model.

[0068] Specifically, after obtaining the first matrix [s1, s2, ..., sn], the first matrix is ​​input into the label determination model, the third label of each training text is obtained through the CRF layer, the third label is compared with the preset label, and the label determination model is adjusted according to the difference between the third label and the second label. If the third label is different from the second label, the label determination model is adjusted until the third label of each training text output is the same as the preset second label. Among them, the CRF layer can add some constraints to the final predicted label to ensure that the predicted label is legal.

[0069] After determining the first label corresponding to the target text, the target text corresponding to the first label is obtained; the target text and the first label are concatenated; and the concatenated first label and the target text are output.

[0070] Specifically, the above mainly introduces the scheme of the embodiment of the present application from the perspective of the execution process on the method side. It is understandable that in order to realize the above functions, the terminal device includes a hardware structure and / or software module corresponding to the execution of each function. It should be easily appreciated by those skilled in the art that, in combination with the units and algorithm steps of each example described in the embodiments provided herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present application.

[0071] The embodiment of the present application can divide the terminal device into functional units according to the above method example. For example, each functional unit can be divided according to each function, or two or more functions can be integrated into one processing unit. The above integrated unit can be implemented in the form of hardware or in the form of software functional units. It should be noted that the division of units in the embodiment of the present application is schematic and is only a logical function division. There may be other division methods in actual implementation.

[0072] See also Figure 6 , Figure 6 6 is a functional unit composition block diagram of a device for determining a text label provided in an embodiment of the present application, the device comprising: a first acquisition unit 610, an extraction unit 620 and a first input unit 630, wherein:

[0073] The first acquisition unit 610 is used to acquire a target image;

[0074] The extraction unit 620 is used to extract the target text in the target image and the target text box feature vector corresponding to the target text;

[0075] The first input unit 630 is used to input the target text and the target text box feature vector into a pre-trained label determination model to determine a first label corresponding to the target text. The label determination model is trained by data obtained by splicing multiple training texts, training text box feature vectors corresponding to the training texts, and a second label corresponding to the training text. The training text box feature vector includes the vertex coordinates of the area where the training text is located in the training image, and the ratio of the hypotenuse length of the area to the hypotenuse length of the training image. The second label is a pre-set label.

[0076] Furthermore, the device also includes:

[0077] A second acquisition unit, used for acquiring a plurality of the training texts, the training text box feature vectors corresponding to the training texts, and the second labels corresponding to the training texts;

[0078] A first concatenation unit, used for concatenating the training text with the training text box feature vector to obtain a first matrix;

[0079] The second input unit is used to input the first matrix into the label determination model to obtain a third label, and adjust the label determination model according to the difference between the third label and the second label until the training end condition is met to obtain the label determination model.

[0080] Furthermore, the device also includes:

[0081] A preprocessing unit is used to preprocess the training text, the training text box feature vector and the second label to remove repeated training text, the training text box feature vector and the second label.

[0082] Furthermore, the first splicing unit is also used for:

[0083] Convert the training text into text numbers through a text dictionary;

[0084] Convert the text numerals into text vectors through a semantic representation model to obtain a second matrix, wherein the second matrix includes the text vectors;

[0085] Converting the training text box feature vector into a third matrix with a dimension of 1 and a fourth matrix with a dimension of k through a fully connected neural network, wherein k is a hyperparameter;

[0086] Determine a fifth matrix according to the third matrix and the second matrix, and determine a sixth matrix according to the fourth matrix and the second matrix;

[0087] The first matrix is ​​determined according to the fifth matrix and the sixth matrix.

[0088] The first splicing unit is further used for:

[0089] The third matrix is ​​concatenated before the second matrix by using the CONCATENATE function to obtain a fifth matrix, and the fourth matrix is ​​concatenated after the second matrix to obtain a sixth matrix.

[0090] Furthermore, the device also includes:

[0091] an activation function processing unit, configured to process the third matrix by applying an activation function to obtain a seventh matrix;

[0092] A determining unit is used to determine a fifth matrix according to the seventh matrix and the second matrix.

[0093] Furthermore, the device also includes:

[0094] A third acquisition unit, configured to acquire the target text corresponding to the first tag;

[0095] A second splicing unit, used for splicing the target text and the first label;

[0096] An output unit is used to output the concatenated first label and the target text.

[0097] See also Figure 7 , Figure 7A terminal device provided in an embodiment of the present application includes: a processor, a memory, a transceiver, and one or more programs. The processor, the memory, and the transceiver are interconnected via a communication bus.

[0098] The processor may be one or more central processing units (CPUs). When the processor is a CPU, the CPU may be a single-core CPU or a multi-core CPU.

[0099] The one or more programs are stored in the memory and configured to be executed by the processor; the programs include instructions for performing the following steps:

[0100] Get the target image;

[0101] Extracting target text in the target image and a target text box feature vector corresponding to the target text;

[0102] The target text and the target text box feature vector are input into a pre-trained label determination model to determine a first label corresponding to the target text. The label determination model is trained by data obtained by splicing multiple training texts, training text box feature vectors corresponding to the training texts, and a second label corresponding to the training text. The training text box feature vector includes the vertex coordinates of the area where the training text is located in the training image, and the ratio of the hypotenuse length of the area to the hypotenuse length of the training image. The second label is a pre-set label.

[0103] It should be noted that the specific implementation process of the embodiments of the present application can refer to the specific implementation process described in the above method embodiments, which will not be repeated here.

[0104] An embodiment of the present application also provides a computer storage medium, wherein the computer storage medium stores a computer program for electronic data exchange, and the computer program enables a computer to execute part or all of the steps of any method recorded in the above method embodiments.

[0105] The present application also provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program, and the computer program is operable to cause a computer to execute some or all of the steps of any method described in the above method embodiment. The computer program product may be a software installation package.

[0106] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present application is not limited by the described order of actions, because according to the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present application.

[0107] In the several embodiments provided in the present application, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are only schematic, such as the division of the above-mentioned units, which is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical or other forms.

[0108] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the embodiments of the present application.

[0109] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0110] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a memory and includes several instructions for a computer device (which can be a personal computer, server or network device, etc.) to perform all or part of the steps of the various embodiments of the present application. The aforementioned memory includes: various media that can store program codes, such as USB flash drives, ROM, RAM, mobile hard disks, magnetic disks or optical disks.

[0111] A person skilled in the art may understand that all or part of the steps in the various methods of the above embodiments may be completed by instructing the relevant hardware through a program, and the program may be stored in a computer-readable memory, which may include a flash drive, ROM, RAM, a magnetic disk or an optical disk, etc.

[0112] The embodiments of the present application are introduced in detail above. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea. At the same time, for those skilled in the art, according to the idea of ​​the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A method for determining a text label, It is characterized in that The method comprises: Get the target image; Extracting target text in the target image and a target text box feature vector corresponding to the target text; The target text and the target text box feature vector are input into a pre-trained label determination model to determine a first label corresponding to the target text, wherein the label determination model is trained by the following steps: Acquire multiple training texts, training text box feature vectors corresponding to the training texts, and second labels corresponding to the training texts, wherein the training text box feature vectors include vertex coordinates of a region where the training texts are located in a training image, and a ratio of the hypotenuse length of the region to the hypotenuse length of the training image, and the second label is a preset label; The training text is concatenated with the training text box feature vector to obtain a first matrix, comprising: Convert the training text into text numbers through a text dictionary; Convert the text numerals into text vectors through a semantic representation model to obtain a second matrix, wherein the second matrix includes the text vectors; Converting the training text box feature vector into a third matrix with a dimension of 1 and a fourth matrix with a dimension of k through a fully connected neural network, wherein k is a hyperparameter; determining a fifth matrix according to the third matrix and the second matrix, and determining a sixth matrix according to the fourth matrix and the second matrix; Determining the first matrix according to the fifth matrix and the sixth matrix; The first matrix is ​​input into the label determination model to obtain a third label, and the label determination model is adjusted according to the difference between the third label and the second label until the training end condition is met to obtain the label determination model.

2. The method according to claim 1, It is characterized in that Before splicing the training text with the training text box feature vector, the method further includes: The training text, the training text box feature vector and the second label are preprocessed to remove repeated training text, the training text box feature vector and the second label.

3. The method according to claim 1, It is characterized in that The determining of the fifth matrix according to the third matrix and the second matrix, and the determining of the sixth matrix according to the fourth matrix and the second matrix comprises: The third matrix is ​​concatenated before the second matrix by using the CONCATENATE function to obtain the fifth matrix, and the fourth matrix is ​​concatenated after the second matrix to obtain the sixth matrix.

4. The method according to claim 1, It is characterized in that Before determining the fifth matrix according to the third matrix and the second matrix, the method further includes: Processing the third matrix by applying an activation function to obtain a seventh matrix; Determining the fifth matrix according to the third matrix and the second matrix comprises: The fifth matrix is ​​determined according to the seventh matrix and the second matrix.

5. The method according to claim 1, It is characterized in that After determining the first label corresponding to the target text, the method further includes: Acquire the target text corresponding to the first tag; concatenating the target text and the first label; The concatenated first label and the target text are output.

6. A device for determining a text label, It is characterized in that The device comprises: A first acquisition unit, used to acquire a target image; An extraction unit, used to extract a target text in the target image and a target text box feature vector corresponding to the target text; The first input unit is used to input the target text and the target text box feature vector into a pre-trained label determination model to determine a first label corresponding to the target text, wherein the label determination model is trained by the following steps: Acquire multiple training texts, training text box feature vectors corresponding to the training texts, and second labels corresponding to the training texts, wherein the training text box feature vectors include vertex coordinates of a region where the training texts are located in a training image, and a ratio of the hypotenuse length of the region to the hypotenuse length of the training image, and the second label is a preset label; The training text is concatenated with the training text box feature vector to obtain a first matrix, comprising: Convert the training text into text numbers through a text dictionary; Convert the text numerals into text vectors through a semantic representation model to obtain a second matrix, wherein the second matrix includes the text vectors; Converting the training text box feature vector into a third matrix with a dimension of 1 and a fourth matrix with a dimension of k through a fully connected neural network, wherein k is a hyperparameter; determining a fifth matrix according to the third matrix and the second matrix, and determining a sixth matrix according to the fourth matrix and the second matrix; Determining the first matrix according to the fifth matrix and the sixth matrix; The first matrix is ​​input into the label determination model to obtain a third label, and the label determination model is adjusted according to the difference between the third label and the second label until the training end condition is met to obtain the label determination model.

7. A terminal device, It is characterized in that The terminal device includes a processor, a memory, a communication interface, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the processor, and the programs include instructions for executing the steps in the method as described in any one of claims 1-5.

8. A computer-readable storage medium, It is characterized in that A computer program for electronic data exchange is stored, wherein the computer program enables a computer to execute the steps in the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Document layout analysis method and device, model training method and device, and equipment

    CN113378580A

  • Question splitting model training method, question splitting method and related device

    CN113762223A