Text annotation method, device and equipment

By combining eye tracking technology and optical character recognition technology, text annotation is automatically processed, and the problems of low efficiency, high cost and high subjectivity in the existing technology are solved, and efficient and low-cost text annotation are achieved.

CN112906683BActive Publication Date: 2025-05-02INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110180619.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-08
Publication Date
2025-05-02
Estimated Expiration
2041-02-08

AI Technical Summary

Technical Problem

The existing text annotation methods mainly rely on manual annotation, which is inefficient, high cost, and subjective, affecting the efficiency and accuracy of the annotation.

Method used

Using a method combining eye tracking technology and optical character recognition technology, the text to be marked is converted into the picture to be marked, and the eye tracking technology is used to obtain the salesperson's attention information on the picture, convert it into the character information of attention, and filter it through optical character recognition technology to automatically obtain the text's annotation information.

Benefits of technology

It realizes text automation and non-information labeling, improves text labeling efficiency, reduces cost, and reduces subjective impact.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112906683B_ABST
    Figure CN112906683B_ABST
Patent Text Reader

Abstract

The embodiments of this specification relate to the field of artificial intelligence technology, and disclose a text annotation method, device, and equipment, the method comprising: converting the text to be annotated into a picture to be annotated; using eye tracking technology to obtain the salesperson's attention image information on the picture to be annotated, the attention image information includes the salesperson's attention area and attention frequency on the picture to be annotated; according to the correspondence between the pixel points in the picture to be annotated and the characters of the text to be annotated, converting the attention image information into attention character information; based on the optical character recognition technology, the attention character information is screened to obtain the annotation information of the text to be annotated. No manual annotation is required, which realizes the automatic and imperceptible annotation of text and improves the efficiency of text annotation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and more particularly to a text annotation method, device and apparatus. Background Art

[0002] With the development of computer Internet technology, it is becoming more and more important to use computer technology to process natural language to facilitate people's work and life, such as intelligent conversation robots, which is a way to realize the intelligence of natural language. Text annotation can mark important information in natural language text to facilitate users to view and understand, or use the annotated text to provide a data basis for subsequent artificial intelligence dialogue.

[0003] At present, text annotation is mostly done manually. The annotator clicks and selects in the developed annotation system. This process requires specialized experts to spend time and attention on annotation, which results in low annotation efficiency and high annotation cost. In addition, this method of text annotation is subjective. Each annotator may annotate the same text with labels of different granularity or direction based on his or her own understanding, which affects the efficiency and accuracy of text annotation.

[0004] To address the above problems, no effective solution has been proposed yet. Summary of the invention

[0005] The purpose of the embodiments of this specification is to provide a text annotation method, device and equipment, which realizes automatic and non-sensitive annotation of text, improves the efficiency of text annotation and reduces the cost of text annotation.

[0006] On the one hand, an embodiment of the present specification provides a text annotation method, the method comprising:

[0007] Convert the text to be annotated into the image to be annotated;

[0008] Using eye tracking technology to obtain the salesperson's attention image information on the picture to be labeled, the attention image information includes the salesperson's attention area and attention frequency on the picture to be labeled;

[0009] Converting the image information into character information according to the correspondence between the pixels in the image to be annotated and the characters in the text to be annotated;

[0010] The attention character information is screened based on optical character recognition technology to obtain the annotation information of the text to be annotated.

[0011] Furthermore, the use of eye tracking technology to obtain the image information that the salesperson is interested in the picture to be labeled includes:

[0012] Using eye tracking technology to obtain the salesperson's gaze information on the image to be labeled;

[0013] According to the sight stop information, the number of times the salesperson pays attention to each pixel point in the image to be labeled is obtained;

[0014] Constructing an image matrix to be labeled according to the image to be labeled, wherein the elements in the image matrix to be labeled are the pixels of the image to be labeled;

[0015] The value of each element in the to-be-annotated image matrix is ​​set as the number of attentions of each element in the to-be-annotated image, and the to-be-annotated image matrix with the element values ​​determined is used as the attention image information.

[0016] Furthermore, the converting the image information of interest into character information of interest according to the correspondence between the pixel points in the image to be annotated and the characters of the text to be annotated includes:

[0017] Constructing a text matrix to be annotated according to the text to be annotated, wherein the elements in the text matrix to be annotated are characters of the text to be annotated;

[0018] According to the correspondence between the characters in the text matrix to be annotated and the pixels in the picture to be annotated, the value of each element in the image matrix to be annotated is converted into the value of each element in the text matrix to be annotated;

[0019] The text matrix to be marked with the determined element values ​​is used as the focus character information.

[0020] Furthermore, the filtering of the concerned character information based on the optical character recognition technology to obtain the annotation information of the text to be annotated includes:

[0021] Construct an optical character recognition model based on optical character recognition technology;

[0022] Annotating historical images of interest based on eye tracking technology, and converting the historical images into corresponding characters of interest;

[0023] Obtaining the historical sample confirmation labeling information that the salesperson labels the historical sample to be labeled;

[0024] Using the historical focused character information as model training input data of the optical character recognition model, using the historical sample confirmation annotation information as model training label data of the optical character recognition model, and performing model training on the optical character recognition model until the optical character recognition model meets the preset requirements;

[0025] The trained optical character recognition model is used to screen the concerned character information to obtain the annotation information of the text to be annotated.

[0026] Furthermore, the filtering of the concerned character information based on the optical character recognition technology to obtain the annotation information of the text to be annotated includes:

[0027] Acquire confirmation annotation information of the specified number of images to be annotated obtained by the salesperson performing text annotation on the specified number of images to be annotated corresponding to the text to be annotated;

[0028] Using the attention character information of the specified number of images to be labeled as the optimization training input data of the optical character recognition model, using the confirmed label information of the specified number of images to be labeled as the optimization training label data of the optical character recognition model, optimizing the optical character recognition model, and obtaining an optimized optical character recognition model;

[0029] The optimized optical character recognition model is used to mark the focus character information of the text to be marked, so as to obtain the marking information of the text to be marked.

[0030] Furthermore, the filtering of the concerned character information based on the optical character recognition technology to obtain the annotation information of the text to be annotated includes:

[0031] The attention character information is screened based on optical character recognition technology to obtain a labeling character matrix corresponding to the text to be labeled, wherein the element values ​​in the labeling character matrix represent the labeling frequency of each character in the text to be labeled;

[0032] Determine the two-dimensional image coordinates of the key marked area in the text to be marked according to the marked character matrix;

[0033] According to the correspondence between the text to be annotated and the picture to be annotated, converting the two-dimensional picture coordinates of the key annotated area into a one-dimensional character index;

[0034] The annotation information of the text to be annotated is obtained based on the one-dimensional character index of the key annotated area.

[0035] Furthermore, the two-dimensional image coordinates of the key marked area are determined using the following formula:

[0036] I1={(x 1_1 :x 1_2 ,y 1_1 :y 1_2 ),(x 2_1 :x 2_2 ,y 2_1 :y 2_2 )...(xn_1 :x n_2 ,y n_1 :y n_2 )}

[0037] Among them, I1 represents the set of two-dimensional image coordinates of the key marked area, x n_1 Indicates the starting point of the horizontal axis of the nth highlighted area, x n_2 Indicates the end point of the horizontal axis of the nth key marked area, y n_1 Indicates the vertical axis starting point of the nth key marked area, y n_2 Indicates the vertical axis end point of the nth highlighted area, (x n_1 :x n_2 ,y n_1 :y n_2 ) represents the two-dimensional image coordinates of the nth key marked area.

[0038] Furthermore, the two-dimensional image coordinates of the key marked area are converted into a one-dimensional character index using the following formula:

[0039] I2={(y 1_1 ×i+x 1_1 :y 1_2 ×i+x 1_2 ),(y 2_1 ×i+x 2_1 :y 2_2 ×i+x 2_2 )...(y n_1 ×i+x n_1 :y n_2 ×i+x n_2 )}

[0040] Where I2 represents the one-dimensional character index of the key marked area, i represents the number of columns of the marked character matrix, (y n_1 ×i+x n_1 :y n_2 ×i+x n_2 ) represents the one-dimensional character index of the nth highlighted area.

[0041] In another aspect, the present specification provides a text annotation device, the device comprising:

[0042] A text conversion module, used to convert the text to be annotated into the image to be annotated;

[0043] An eye tracking and annotation module, used to obtain the salesperson's attention image information of the picture to be annotated by using the eye tracking technology, wherein the attention image information includes the salesperson's attention area and attention frequency of the picture to be annotated;

[0044] A labeling information character conversion module, used for converting the focus image information into focus character information according to the correspondence between the pixel points in the to-be-labeled picture and the characters of the to-be-labeled text;

[0045] The annotation information screening module is used to screen the concerned character information based on the optical character recognition technology to obtain the annotation information of the text to be annotated.

[0046] Furthermore, the eye tracking annotation module is specifically used for:

[0047] Using eye tracking technology to obtain the salesperson's gaze information on the image to be labeled;

[0048] According to the sight stop information, the number of times the salesperson pays attention to each pixel point in the image to be labeled is obtained;

[0049] Constructing an image matrix to be labeled according to the image to be labeled, wherein the elements in the image matrix to be labeled are the pixels of the image to be labeled;

[0050] The value of each element in the to-be-annotated image matrix is ​​set as the number of attentions of each element in the to-be-annotated image, and the to-be-annotated image matrix with the element values ​​determined is used as the attention image information.

[0051] Furthermore, the annotation information screening module is specifically used for:

[0052] Constructing a text matrix to be annotated according to the text to be annotated, wherein the elements in the text matrix to be annotated are characters of the text to be annotated;

[0053] According to the correspondence between the characters in the text matrix to be annotated and the pixels in the picture to be annotated, the value of each element in the image matrix to be annotated is converted into the value of each element in the text matrix to be annotated;

[0054] The text matrix to be marked with the determined element values ​​is used as the focus character information.

[0055] On the other hand, an embodiment of the present specification provides a text annotation device, which is applied to a server. The device includes at least one processor and a memory for storing processor executable instructions. When the instructions are executed by the processor, the above-mentioned text annotation method is implemented.

[0056] The text annotation method, device and equipment provided in this specification combine eye tracking technology and optical character recognition technology to annotate text, and embed the annotation method into the business system, so as to realize automatic annotation of text in the process of business personnel handling related business. The annotation process is automated and non-sensitive, and no professional annotators are required, which reduces the cost and time of text annotation. In addition, the text annotation method in the embodiment of this specification transforms "annotating paragraphs with business value in text" into two tasks: "annotating areas with business value in images" + "recognizing text in images", and the cost of annotation is lower. As paperless office becomes a trend, more and more business-related documents are entered, presented and processed in electronic form, which makes it possible to use eye tracking for annotation. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] In order to more clearly illustrate the embodiments of this specification or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0058] Figure 1 It is a flowchart of a text annotation method embodiment provided in an embodiment of this specification;

[0059] Figure 2 is a schematic diagram of an interface for automatic text annotation in an embodiment of this specification;

[0060] Figure 3 is a structural schematic diagram of a text annotation device in one embodiment of this specification;

[0061] Figure 4 It is a hardware structure block diagram of a text annotation server in one embodiment of this specification. DETAILED DESCRIPTION

[0062] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of this specification.

[0063] In a scenario example provided in the embodiments of this specification, the text annotation method can be applied to a device for performing text annotation, and the device may include a server or a server cluster composed of multiple servers. The text annotation method can be integrated into a business system, and when a salesperson views or uses relevant information in a business system, the salesperson can obtain information that the salesperson focuses on based on the text annotation method provided in the embodiments of this specification. This information can be used as annotation information to achieve automatic annotation of text, and the annotation information can provide a data basis for subsequent business processing or intelligent business dialogues.

[0064] Figure 1 It is a flowchart of the text annotation method embodiment provided by the embodiment of this specification. Although this specification provides the method operation steps or device structure shown in the following embodiments or drawings, more or fewer operation steps or module units after partial merger may be included in the method or device based on routine or no creative labor. In the steps or structures where there is no necessary causal relationship logically, the execution order of these steps or the module structure of the device is not limited to the execution order or module structure shown in the embodiments or drawings of this specification. When the method or module structure is applied in an actual device, server or terminal product, it can be executed sequentially or in parallel according to the method or module structure shown in the embodiment or drawings (for example, a parallel processor or multi-threaded processing environment, or even a distributed processing, server cluster implementation environment).

[0065] A specific implementation example is Figure 1 As shown, in one embodiment of the text annotation method provided in this specification, the method can be applied to terminals such as servers, computers, tablet computers, and smart phones. The method may include the following steps:

[0066] Step 102: Convert the text to be annotated into a picture to be annotated.

[0067] In the specific implementation process, the text to be annotated can be understood as the text processed by the salesperson, or the specified text that needs to be annotated, and the embodiments of this specification do not make specific limitations. When the business personnel process text-related information, the annotation system embedded in the business system can be automatically opened. At this time, the annotation system will treat the text as a picture, such as: the text displayed on the current screen can be regarded as a picture as the picture to be annotated. When the salesperson scrolls the text on the screen, the text that has changed on the screen is used as a new picture. It can be seen that the text to be annotated can be converted into multiple pictures to be annotated, which can be determined according to actual needs, and the embodiments of this specification do not make specific limitations.

[0068] Step 104: using eye tracking technology to obtain the salesperson's attention image information on the picture to be labeled, the attention image information includes the salesperson's attention area and attention frequency on the picture to be labeled.

[0069] In the specific implementation process, eye tracking is a scientific application technology. In principle, eye tracking mainly studies the acquisition, modeling and simulation of eye movement information, and has a wide range of uses. In addition to infrared devices, the device that obtains eye movement information can also be an image acquisition device, or even a camera on a general computer or mobile phone, which can also achieve eye tracking with the support of software.

[0070] In the embodiments of this specification, eye tracking technology can be used to collect, model, and simulate eye movement information in advance, so as to establish a model for marking pictures according to the line of sight movement of business personnel during the office work process, which is recorded as M1. The established model M1 is then used to annotate the picture to be annotated after the text to be annotated is converted, and the attention image information of the picture to be annotated can be output. The attention image information may include the attention area and attention frequency of the picture to be annotated by the salesperson, such as: based on the eye tracking technology, the line of sight information of the salesperson when viewing the picture to be annotated corresponding to the text to be annotated can be obtained, that is, the area in the picture to be annotated where the salesperson's line of sight stays for a longer time can be obtained, that is, the area of ​​attention, and the frequency of the salesperson's attention to the attention area can be further obtained.

[0071] In some embodiments of this specification, the use of eye tracking technology to obtain the image information that the salesperson pays attention to the picture to be labeled includes:

[0072] Using eye tracking technology to obtain the salesperson's gaze information on the image to be labeled;

[0073] According to the sight stop information, the number of times the salesperson pays attention to each pixel point in the image to be labeled is obtained;

[0074] Constructing an image matrix to be labeled according to the image to be labeled, wherein the elements in the image matrix to be labeled are the pixels of the image to be labeled;

[0075] The value of each element in the to-be-annotated image matrix is ​​set as the number of attentions of each element in the to-be-annotated image, and the to-be-annotated image matrix with the element values ​​determined is used as the attention image information.

[0076] In the specific implementation process, a matrix of images to be annotated can be first constructed based on the images to be annotated, and the elements in the matrix of images to be annotated can represent the pixels of the images to be annotated, such as: based on the images presented on the current device screen, an a×b matrix P can be obtained, where a is the number of pixels in the horizontal direction of the image, b is the number of pixels in the vertical direction of the image, and each element in P corresponds to a pixel in the image. Eye tracking technology is used to obtain the line of sight information of the salesperson when viewing the images to be annotated, and the number of times the salesperson pays attention to each pixel in the images to be annotated can be obtained based on the obtained line of sight information. The value of each element in the matrix of images to be annotated is set to the number of times the salesperson pays attention to each element in the images to be annotated counted by the eye tracking technology, and each element in the matrix of images to be annotated with the element value determined can represent the pixel of the images to be annotated, and the element value can represent the frequency of attention to each pixel in the images to be annotated. Therefore, the matrix of images to be annotated can represent the image information of interest in the above-mentioned embodiments. For example: In an example of this specification, the matrix of images to be annotated can be expressed as:

[0077]

[0078] The larger the value in the above matrix P, the higher the frequency of attention, and the value of 0 can be understood as no attention. The matrix can be used to intuitively represent the key areas of attention in the image to be labeled and the degree of attention paid to each pixel.

[0079] It should be noted that when using eye tracking technology to count the number of times each pixel is paid attention to, the cumulative value of the number of times the salesperson pays attention to each pixel in the labeled image within the specified time can be counted, or the specified time can be divided into n times, and the number of times the salesperson pays attention to each pixel in the labeled image is counted n times to obtain n labeled image matrices, and then the element values ​​in these n labeled image matrices are accumulated to obtain the final labeled image matrix. For example: assuming that the specified time is 10 seconds, the salesperson's attention information to the labeled image is recorded once every 1 second, and 10 labeled image matrices can be obtained. The element values ​​in the 10 labeled image matrices are accumulated to obtain the final labeled image matrix. The statistical method of the frequency of attention to the pixel points of the image to be labeled is set according to actual needs, and the embodiments of this specification do not make specific limitations.

[0080] Step 106: convert the focus image information into focus character information according to the correspondence between the pixel points in the to-be-annotated picture and the characters in the to-be-annotated text.

[0081] In a specific implementation process, the image to be annotated is obtained by converting the text to be annotated. There is a correspondence between each pixel of the image to be annotated and each character of the text to be annotated. Based on the correspondence, the obtained image information of interest can be converted into character information of interest.

[0082] In some embodiments of the present specification, converting the focus image information into focus character information according to the correspondence between the pixel points in the image to be annotated and the characters of the text to be annotated includes:

[0083] Constructing a text matrix to be annotated according to the text to be annotated, wherein the elements in the text matrix to be annotated are characters of the text to be annotated;

[0084] According to the correspondence between the characters in the text matrix to be annotated and the pixels in the picture to be annotated, the value of each element in the image matrix to be annotated is converted into the value of each element in the text matrix to be annotated;

[0085] The text matrix to be marked with the determined element values ​​is used as the focus character information.

[0086] In the specific implementation process, a text matrix to be annotated can be constructed based on the text to be annotated, such as: constructing an m×n matrix G with all 0s, where m can be the length of each line of the text to be annotated, n is the number of lines of the text to be annotated, and each element in G represents a character in the text to be annotated.

[0087]

[0088] Each element in the image matrix to be annotated corresponds to a pixel point in the image. A mapping relationship can be established between each character in the text matrix to be annotated and an area in the image matrix to be annotated. Based on this mapping relationship, the values ​​of each element in the image matrix to be annotated can be converted into the values ​​of each element in the text matrix to be annotated. For example, the values ​​of the elements corresponding to the same character in the image matrix to be annotated are accumulated to obtain the values ​​of the elements in the text matrix to be annotated corresponding to the character, and then the text matrix to be annotated with values ​​is obtained, which can be expressed as the following G1:

[0089]

[0090] The value of each element in the text matrix to be annotated can represent the number of times each character in the text to be annotated is paid attention to. Based on the text matrix to be annotated, it can be intuitively indicated whether each character in the text to be annotated is a key focus object. The text matrix to be annotated can represent the focused character information in the above embodiment.

[0091] Step 108: Screen the concerned character information based on optical character recognition technology to obtain the annotation information of the text to be annotated.

[0092] In the specific implementation process, the eye tracking technology determines the annotation information of the text to be annotated based on the information of the salesperson's sight during the business processing. However, the annotation information obtained based on the eye tracking technology may contain some annotation focuses that are not required by the business. In the embodiment of this specification, after the attention character information of the text to be annotated is obtained based on the eye tracking technology, the optical payment recognition technology, namely the OCR (Optical Character Recognition) positioning technology, can be used to screen the obtained attention character information. Among them, the OCR positioning model can be constructed by OCR technology training in advance, and then the OCR positioning model can be used to screen and further confirm the attention character information obtained based on the eye tracking technology to improve the accuracy of text annotation.

[0093] At the same time, when using optical character recognition technology to filter the annotation information obtained by eye tracking technology, the focus character information is used instead of pixel information. The size of the focus character information will be much smaller than the matrix converted from the image sampled by the RGB model. In the matrix calculation process of deep learning, the resources consumed will be significantly less than the image represented by pixels, which improves data processing efficiency and thus improves the efficiency of text annotation.

[0094] In some embodiments of the present specification, the screening of the attention character information based on the optical character recognition technology to obtain the annotation information of the text to be annotated includes:

[0095] Construct an optical character recognition model based on optical character recognition technology;

[0096] Annotating historical images of interest based on eye tracking technology, and converting the historical images into corresponding characters of interest;

[0097] Obtaining the historical sample confirmation labeling information that the salesperson labels the historical sample to be labeled;

[0098] Using the historical focused character information as model training input data of the optical character recognition model, using the historical sample confirmation annotation information as model training label data of the optical character recognition model, and performing model training on the optical character recognition model until the optical character recognition model meets the preset requirements;

[0099] The trained optical character recognition model is used to screen the concerned character information to obtain the annotation information of the text to be annotated.

[0100] In the specific implementation process, in some embodiments of this specification, an optical character recognition model can be constructed based on OCR technology. The optical character recognition model can specifically use a neural network algorithm or other intelligent learning algorithm, which is not specifically limited in the embodiments of this specification. Then use eye tracking technology to annotate the historical samples to be annotated, obtain historical image information of interest, and convert the historical image information of interest into corresponding historical character information of interest. Then present the historical samples to be annotated to the salesperson for manual confirmation, obtain the historical sample confirmation annotation information corresponding to the historical annotated samples, use the obtained historical character information of interest as the model training input data of the optical character recognition model, and use the corresponding historical sample confirmation annotation information as the model training label data of the optical character recognition model, and perform model training on the optical character recognition model until the optical character recognition model meets the preset requirements and the optical character recognition model is trained. Then use the trained optical character recognition model to screen the character information of interest corresponding to the current text to be annotated to obtain the annotation information of the text to be annotated.

[0101] The method of converting the historical focused image information into the corresponding historical focused character information refers to the method of converting the focused image information into the focused character information in the above embodiment, which will not be described in detail here.

[0102] In addition, the embodiments of this specification can embed the text annotation system corresponding to the text annotation method into a relatively large-scale business system such as banking, finance, and medical treatment. In this way, relying on the characteristics of large data volume, large business volume, and a large number of business personnel with professional knowledge in banks, finance and other institutions, eye tracking can be applied to the annotation system with practical value. In the Internet industry or research institutions such as universities that dominate natural language processing, the business volume is small, and there is a lack of business experts who have sufficient understanding of the actual business field. Even if a corresponding system is developed for ordinary annotation personnel, it is not practical. In larger institutions such as banks, medical systems, and government agencies, there are a large number of business experts with professional knowledge, which is suitable for large-scale application of this non-sensitive annotation system. With the promotion of intelligent office and the maturity of eye tracking technology, a large-scale institution should be invented to quickly obtain a large amount of low-cost and high-quality annotation data that meets the business needs of the industry.

[0103] The text annotation method provided in the implementation of this specification combines eye tracking technology and optical character recognition technology to annotate text, and embeds the annotation method into the business system, so as to realize automatic annotation of text in the process of business personnel handling related business. The annotation process is automated and non-sensitive, and no professional annotators are required, which reduces the cost and time of text annotation. In addition, the text annotation method in the embodiment of this specification transforms "annotating paragraphs with business value in text" into two tasks: "annotating areas with business value in images" + "recognizing text in images", and the cost of annotation is lower. As paperless office becomes a trend, more and more business-related documents are entered, presented and processed in electronic form, which makes it possible to use eye tracking for annotation.

[0104] On the basis of the above embodiments, in some embodiments of this specification, the screening of the concerned character information based on the optical character recognition technology to obtain the annotation information of the text to be annotated includes:

[0105] Acquire confirmation annotation information of the specified number of images to be annotated obtained by the salesperson performing text annotation on the specified number of images to be annotated corresponding to the text to be annotated;

[0106] Using the attention character information of the specified number of images to be labeled as the optimization training input data of the optical character recognition model, using the confirmed label information of the specified number of images to be labeled as the optimization training label data of the optical character recognition model, optimizing the optical character recognition model, and obtaining an optimized optical character recognition model;

[0107] The optimized optical character recognition model is used to mark the focus character information of the text to be marked, so as to obtain the marking information of the text to be marked.

[0108] In the specific implementation process, when using OCR technology to filter the annotation information obtained by eye tracking technology, a specified number of images to be annotated corresponding to the text to be annotated can be presented to the salesperson for manual annotation, and the manual annotation results can be used to optimize the training of the optical character recognition model, and the optimized optical character recognition model can be used to filter the annotation information obtained by eye tracking technology.

[0109] For example: if 10 pictures to be annotated are obtained when the text to be annotated is converted into pictures to be annotated, the eye tracking technology can be used to obtain the image information corresponding to the 10 pictures to be annotated, and then the character information corresponding to the 10 pictures to be annotated can be obtained. Then, 2 pictures are randomly selected from the 10 pictures to be annotated and shown to the salesperson for manual annotation, and the confirmation annotation information corresponding to the 2 pictures to be annotated is obtained. The character information corresponding to the 2 pictures to be annotated (obtained by the eye tracking technology) is used as the optimized training input data of the optical character recognition model created in the above embodiment, and the confirmation annotation information corresponding to the 2 pictures to be annotated is used as the optimized training label data of the optical character recognition model, and the optical character recognition model is optimized to obtain the optimized optical character recognition model. The character information corresponding to the 10 pictures to be annotated is annotated using the optimized optical character recognition model, that is, the character information corresponding to the 10 pictures to be annotated is input into the optimized optical character recognition model, and the output result of the model is obtained, which is the annotation information of the text to be annotated.

[0110] When the embodiments of this specification use OCR technology to filter the annotation information identified by eye tracking technology, fewer recognition samples are extracted from the recognition results of the eye tracking technology, the model established by the OCR recognition technology is optimized, and then the optimized model is used to filter the recognition results of the eye tracking technology. This not only ensures the accuracy of text annotation, but also reduces the workload of the salesperson in the process of confirming the recognition results of the eye tracking technology, thereby improving the efficiency of text annotation.

[0111] In some embodiments of the present specification, the screening of the attention character information based on the optical character recognition technology to obtain the annotation information of the text to be annotated includes:

[0112] The attention character information is screened based on optical character recognition technology to obtain a labeling character matrix corresponding to the text to be labeled, wherein the element values ​​in the labeling character matrix represent the labeling frequency of each character in the text to be labeled;

[0113] Determine the two-dimensional image coordinates of the key marked area in the text to be marked according to the marked character matrix;

[0114] According to the correspondence between the text to be annotated and the picture to be annotated, converting the two-dimensional picture coordinates of the key annotated area into a one-dimensional character index;

[0115] The annotation information of the text to be annotated is obtained based on the one-dimensional character index of the key annotated area.

[0116] In the specific implementation process, referring to the records of the above embodiments, after the eye tracking technology obtains the concerned character information of the text to be annotated, the OCR technology is used to screen the obtained concerned information, and the annotated character matrix corresponding to the text to be annotated can be obtained, and the element values ​​in the annotated character matrix represent the annotation frequency of each character in the text to be annotated. The rows and columns of the annotated character matrix correspond to the text length and the number of text lines corresponding to each row of the text to be annotated. It can be understood from the records of the above embodiments that each element of the annotated character matrix is ​​actually associated with the text arrangement of the text to be annotated or the picture to be annotated, that is, the annotation result obtained by the eye tracking technology is a two-dimensional information corresponding to the text to be annotated or the picture to be annotated. According to the annotated character matrix, the two-dimensional image coordinates of the key annotated area in the text to be annotated can be determined, such as: the position coordinates of the elements in the annotated character matrix whose element values ​​are greater than the specified threshold in the matrix can be used as the two-dimensional image coordinates of the element, and so on, the two-dimensional image coordinates of each key annotated area in the annotated character matrix are obtained. Then, according to the corresponding relationship between the text to be annotated and the picture to be annotated, the obtained two-dimensional image coordinates of each key annotated area are converted into a one-dimensional character index, and the annotation information of the text to be annotated is obtained based on the obtained one-dimensional character index.

[0117] The one-dimensional character index can be understood as the character order of the characters in the key annotation area in the entire text to be annotated. For example, if the value of an element in the annotation character matrix is ​​greater than the specified threshold, this element is taken as the key annotation area, and the position of this element in the annotation character matrix is ​​obtained as the two-dimensional image coordinates of the element. Then, according to the correspondence between the text to be annotated and the image to be annotated, the order of this element in the entire text to be annotated is obtained. For example, if this element is the second element in the second row in the annotation character matrix, each row of the annotation character matrix has 10 elements, and the number of elements in each row of the annotation character matrix is ​​equal to the length of each row of the text to be annotated, then based on the alignment relationship between the text to be annotated, the image to be annotated, and the matrix, it can be determined that the order of this element in the entire text to be annotated should be 12 characters lower.

[0118] Based on the one-dimensional character index of the key annotation area in the text to be annotated, it is possible to quickly obtain which characters in the text to be annotated need to be annotated. Based on actual usage needs, the characters that need to be annotated in the text to be annotated can be annotated in a specified format according to the one-dimensional character index.

[0119] In some embodiments of this specification, the following formula may be used to determine the two-dimensional image coordinates of the key marked area:

[0120] I1={(x 1_1 :x 1_2 ,y 1_1 :y 1_2 ),(x 2_1:x 2_2 ,y 2_1 :y 2_2 )...(x n_1 :x n_2 ,y n_1 :y n_2 )}

[0121] Among them, I1 represents the set of two-dimensional image coordinates of the key marked area, x n_1 Indicates the starting point of the horizontal axis of the nth highlighted area, x n_2 Indicates the end point of the horizontal axis of the nth key marked area, y n_1 Indicates the vertical axis starting point of the nth key marked area, y n_2 Indicates the vertical axis end point of the nth highlighted area, (x n_1 :x n_2 ,y n_1 :y n_2 ) represents the two-dimensional image coordinates of the nth key marked area.

[0122] As can be seen from the above formula, in the embodiment of this specification, the key marking area can be determined according to the value of each element in the marking character matrix, such as: the element whose value is greater than a specified threshold is regarded as the key marking area. Based on the position of the elements in the key marking area in the marking character matrix, the two-dimensional image coordinates of the key marking area are obtained. For example, assuming that the marking character matrix is The element value greater than 2 is regarded as the key marked area, then the 4th to 5th elements in the first row of the matrix G3 can be regarded as the first key marked area. In the embodiment of this specification, when obtaining the position of the key marked area in the marked character matrix, the rows and columns of the marked character matrix are calculated from 0, then the two-dimensional image coordinates of the key marked area are (3:4, 0:0), and so on. The two-dimensional image coordinate set of the key marked area in G3 can be obtained as follows:

[0123] I1={(3:4, 0:0), (4:5, 1:1), (7:9, 3:3)}

[0124] In the above example, a highlighted area is located in the same row of the labeled character matrix, so the vertical axis starting point and ending point of the highlighted area are the same. According to actual usage needs, there may be highlighted areas in several consecutive rows, in which case the vertical axis starting point and ending point of the highlighted area will be different.

[0125] In some embodiments of this specification, the following formula is used to convert the two-dimensional image coordinates of the key marked area into a one-dimensional character index:

[0126] I2={(y 1_1 ×i+x 1_1 :y1_2 ×i+x 1_2 ),(y 2_1 ×i+x 2_1 :y 2_2 ×i+x 2_2 )...(y n_1 ×i+x n_1 :y n_2 ×i+x n_2 )}

[0127] Where I2 represents the one-dimensional character index of the key marked area, i represents the number of columns of the marked character matrix minus 1, (y n_1 ×i+x n_1 :y n_2 ×i+x n_2 ) represents the one-dimensional character index of the nth highlighted area.

[0128] Continuing with the labeled character matrix G3 in the above embodiment, where the number of columns i of the labeled character matrix G3 is 10, the one-dimensional character index of the highlighted labeled area can be obtained according to the two-dimensional image coordinate set of the highlighted labeled area in G3:

[0129] I2={(3:4), (14:15), (37:39)}

[0130] The one-dimensional character index of the highlighted area can be understood as the order of the highlighted characters in the entire text to be annotated. Based on the one-dimensional character index, the content that needs to be highlighted in the text to be annotated can be quickly searched.

[0131] Referring to the records in the above embodiments, it can be known that in the embodiments of this specification, when calculating the two-dimensional image coordinates and the one-dimensional character index, the rows and columns of the labeled character matrix are calculated starting from 0.

[0132] In the embodiments of the present specification, the text to be annotated is converted into a picture to be annotated, and the picture to be annotated is annotated using eye tracking technology and OCR technology to obtain a corresponding annotation result, which is two-dimensional annotation information aligned with the picture to be annotated. The two-dimensional annotation information is converted into a one-dimensional character index, which is more convenient for natural language processing and improves the efficiency of text annotation.

[0133] After obtaining the annotation information of the text to be annotated, the embodiment of this specification can also optimize the models of the eye tracking technology and the OCR technology based on the feedback of the salesperson on the annotation information, so that the text annotation results are more in line with business needs.

[0134] The following is a detailed description of the text annotation process in this application with a scenario example:

[0135] The annotation system is embedded in the business system, and relevant information such as business name and business content is collected from the business system and configured in the annotation system. This information can be used as the basis for subsequent sorting of annotation information and integration of tags.

[0136] When business personnel process text-related information, the annotation system is automatically turned on. At this time, the system will treat the entire text as a picture and use eye tracking technology to collect, model, and simulate eye movement information, thereby establishing a model (M1) for marking pictures based on the business personnel's line of sight during office work. The output image marking coordinates and marking frequency are used to record which areas on the picture have the business personnel's line of sight lingering, and how long the line of sight stays in these areas. Areas with long line of sight lingering time are the focus of business personnel and have higher business value.

[0137] An m×n matrix G with all zeros can be constructed based on the text to be annotated, where m is the length of each line of the text to be annotated, n is the number of lines of the text to be annotated, and each element in G represents a character in the text, as follows:

[0138]

[0139] At the same time, an a×b matrix P can be obtained from the image of the text to be annotated on the screen, where a is the number of pixels in the horizontal direction of the image, b is the number of pixels in the vertical direction of the image, and each element in P corresponds to a pixel point in the image. Each character in the matrix G and an area in the matrix P [q m ,q n ] can establish a mapping relationship. When business personnel process text, they will leave eye tracking records. Using the eye tracking model M1, the annotation information of the image to be annotated corresponding to the text to be annotated is obtained, that is, the attention image information P1. For example, P1 can be expressed as follows:

[0140]

[0141] Use softmax for each region in P1 and convert it into a mark M of whether each element in G is being paid attention to. M is 0 or 1. When a business person processes a picture to be labeled, he will leave n times of marking information. Then, the accumulated n times (M1, M2, ..., M n ) to obtain a natural number k, the larger the k, the more records there are, thereby obtaining G1, which is a matrix that records whether each character in the text to be annotated is concerned, in which the value of each element can represent the number of times the character at the corresponding coordinate in the text is concerned, such as G1 can be expressed as:

[0142]

[0143] G1 is a digital image information represented by a matrix. The difference is that matrix G1 stores the sampling information of text character granularity rather than pixel granularity. Use G1 to mark the original image and present it to the business personnel for confirmation. The confirmation result is recorded as G2:

[0144]

[0145] G1 is used as training data and G2 as labels, and sent to the deep learning model for training to obtain a model M2 that predicts important information areas based on eye tracking results. M2 can be used for subsequent optimization processes to reduce or eliminate the process of confirmation by business personnel. During the initial annotation, business personnel need a lot of confirmation, which is equivalent to OCR annotation. After obtaining M2, the number of business personnel confirmations can be gradually reduced, and the information of each confirmation is used as incremental training data for M2. For example, a small number of images to be annotated corresponding to the text to be annotated can be selected from the images to be annotated and presented to the salesperson for confirmation. M2 is optimized and trained based on the annotation information of the selected images to be annotated by the salesperson, and G1 is further annotated and screened using the optimized M2.

[0146] Since each element in the matrix G1 represents a character rather than a pixel, the size of the matrix G1 will be much smaller than the matrix converted from the image sampled by the RGB model. Therefore, in the matrix calculation process of deep learning, the resources consumed will be significantly less than the image represented by pixels. Use the optimized model M2 to predict G1 and obtain the annotation G3 of the important information area. In this process, G1 is the recognition result of the eye tracking technology, G2 is the annotation information confirmed by the business personnel, and G3 is the confirmed annotation recognition result of G1 by the OCR model. Finally, the model M2 can directly predict G1 to obtain G3, and the subsequent G2 can be used as incremental training data for M2.

[0147] Based on G3, the two-dimensional image coordinates I1 of the important n information areas are generated. At this time, the annotation information is the two-dimensional image coordinates that are closer to the OCR annotation, which is not convenient to use in natural language processing. According to the relevant information of the alignment between the document and the image, the n two-dimensional coordinates of the image of the important information area can be mapped into n one-dimensional character indexes I1, where the determination method of I1 and I2 is specifically described in the above embodiment, which will not be repeated here. In this way, the annotation information I2 of the text to be annotated is obtained, that is, the annotation result of "original text-key content".

[0148] The annotation results can be sampled and provided to business personnel for confirmation, and the system and model can be optimized based on the results confirmed by the business personnel. The following optimizations can be performed based on the feedback from business personnel: 1) Adjust the eye tracking threshold to obtain a more accurate business personnel focus area. 2) Optimize the OCR positioning model to obtain better annotation results.

[0149] Existing annotation methods are more suitable for image annotation, but because natural language corpus can be regarded as one-dimensional in existing annotation systems, annotators cover less information when annotating natural language data, and the annotation efficiency is lower. At the same time, because natural language annotation requires more business knowledge, it often requires business personnel to have a deeper business background and spend special attention, which is costly and inefficient. The text annotation method provided in the embodiment of this specification changes the annotation content: from the previous "annotating paragraphs with business value in the text" to "annotating areas with business value in the image" + "identifying text in the image", the cost of annotation is lower. As paperless office becomes a trend, more and more business-related documents are entered, presented and processed in electronic form, which makes it possible to use eye tracking for annotation.

[0150] Existing text annotation requires annotators to use interactive tools such as a mouse to click, drag, input, save, turn pages, etc. in the system, which is inconvenient to operate and requires business experts to spend time learning and familiarizing themselves with the annotation system and perform annotation actions. The time cost and labor cost of annotation are very high. The embodiments of this specification use eye tracking technology for the annotation system and embed it into the business office process, canceling the special annotation behavior, achieving non-sensing annotation, and reducing the annotation cost.

[0151] Furthermore, the text annotation method in the present embodiment has determined that the annotation object is a text in the form of an image, so character granularity information can be used for sampling during the annotation process, and the matrix after the image is digitized is smaller, which greatly reduces the amount of calculation of the neural network. The efficiency of annotation, training, and prediction will be higher, and a deeper and wider network structure can also be used to obtain better results.

[0152] In addition, the embodiments of this specification improve the existing annotation system by embedding the annotation system into the business system. When business personnel are reading business-related texts and working, the system automatically marks the areas on the screen where the business personnel are concentrating through eye-tracking related hardware. Figure 2 is a schematic diagram of the interface for automatic text annotation in the embodiment of this specification, such as Figure 2 As shown, the embodiments of the present specification can automatically convert such annotations into annotations for text. Of course, the form of annotations can be adjusted according to actual needs. There is no need for business experts to spend time on annotations, and there are no special annotation actions. This is because the annotation process is completely synchronized with the business personnel's office work. In this process, business personnel only need to work normally, and the system automatically collects relevant information to complete the natural language annotation, and the annotation cost is low.

[0153] In this specification, each embodiment of the above method is described in a progressive manner, and the same or similar parts between the embodiments can be referenced to each other, and each embodiment focuses on the differences from other embodiments. For relevant parts, refer to the partial description of the method embodiment.

[0154] Based on the above-mentioned text annotation method, one or more embodiments of this specification also provide a device for text annotation. The device may include a system (including a distributed system), software (application), module, component, server, client, etc. using the method of the embodiment of this specification and combined with the necessary implementation hardware. Based on the same innovative concept, the device in one or more embodiments provided by the embodiment of this specification is as follows. Since the implementation scheme and method of the device to solve the problem are similar, the implementation of the specific device of the embodiment of this specification can refer to the implementation of the aforementioned method, and the repetitions will not be repeated. As used below, the term "unit" or "module" can implement a combination of software and / or hardware of predetermined functions. Although the system and device described in the following embodiments are preferably implemented in software, the implementation of hardware, or a combination of software and hardware is also possible and conceived.

[0155] Figure 3 is a structural diagram of a text annotation device in one embodiment of this specification, such as Figure 3 As shown, the text annotation device provided in some embodiments of this specification can be applied to the server in the above embodiment, and specifically may include:

[0156] A text conversion module 31, used to convert the text to be annotated into a picture to be annotated;

[0157] An eye tracking and labeling module 32 is used to obtain the salesperson's attention image information of the picture to be labeled by using the eye tracking technology, wherein the attention image information includes the salesperson's attention area and attention frequency of the picture to be labeled;

[0158] The annotation information character conversion module 33 is used to convert the focus image information into focus character information according to the correspondence between the pixel points in the image to be annotated and the characters of the text to be annotated;

[0159] The annotation information screening module 34 is used to screen the attention character information based on the optical character recognition technology to obtain the annotation information of the text to be annotated.

[0160] The text annotation device provided in the embodiment of this specification combines eye tracking technology and optical character recognition technology to annotate text, and embeds the annotation method into the business system, so as to realize automatic annotation of text in the process of business personnel handling related business. The annotation process is automated and non-sensitive, and no professional annotators are required, which reduces the cost and time of text annotation. In addition, the text annotation method in the embodiment of this specification transforms "annotating paragraphs with business value in text" into two tasks: "annotating areas with business value in images" + "recognizing text in images", and the cost of annotation is lower. As paperless office becomes a trend, more and more business-related documents are entered, presented and processed in electronic form, which makes it possible to use eye tracking for annotation.

[0161] In some embodiments of this specification, the eye tracking annotation module is specifically used to:

[0162] Using eye tracking technology to obtain the salesperson's gaze information on the image to be labeled;

[0163] According to the sight stop information, the number of times the salesperson pays attention to each pixel point in the image to be labeled is obtained;

[0164] Constructing an image matrix to be labeled according to the image to be labeled, wherein the elements in the image matrix to be labeled are the pixels of the image to be labeled;

[0165] The value of each element in the to-be-annotated image matrix is ​​set as the number of attentions of each element in the to-be-annotated image, and the to-be-annotated image matrix with the element values ​​determined is used as the attention image information.

[0166] In some embodiments of this specification, the annotation information screening module is specifically used to:

[0167] Constructing a text matrix to be annotated according to the text to be annotated, wherein the elements in the text matrix to be annotated are characters of the text to be annotated;

[0168] According to the correspondence between the characters in the text matrix to be annotated and the pixels in the picture to be annotated, the value of each element in the image matrix to be annotated is converted into the value of each element in the text matrix to be annotated;

[0169] The text matrix to be marked with the determined element values ​​is used as the focus character information.

[0170] The text annotation device provided in the embodiment of the present specification can intuitively represent the focus area in the image to be annotated and the degree of focus of each pixel point in the form of a matrix.

[0171] It should be noted that the above-mentioned device may also include other implementation modes according to the description of the corresponding method embodiment. The specific implementation modes may refer to the description of the corresponding method embodiment, and will not be described one by one here.

[0172] The embodiment of this specification also provides a text annotation device, which is applied to a server. The device includes at least one processor and a memory for storing instructions executable by the processor. When the instructions are executed by the processor, the text annotation method in the above embodiment is implemented, such as:

[0173] Convert the text to be annotated into the image to be annotated;

[0174] Using eye tracking technology to obtain the salesperson's attention image information on the picture to be labeled, the attention image information includes the salesperson's attention area and attention frequency on the picture to be labeled;

[0175] Converting the image information into character information according to the correspondence between the pixels in the image to be annotated and the characters in the text to be annotated;

[0176] The attention character information is screened based on optical character recognition technology to obtain the annotation information of the text to be annotated.

[0177] It should be noted that the above-mentioned device may also include other implementation modes according to the description of the method embodiment. The specific implementation modes may refer to the description of the relevant method embodiment, and will not be described one by one here.

[0178] The methods or devices of the above embodiments provided in this specification can implement business logic through computer programs and record them on storage media, and the storage media can be read and executed by computers to achieve the effects of the solutions described in the embodiments of this specification.

[0179] The method embodiments provided in the embodiments of this specification can be executed in a mobile terminal, a computer terminal, a server or a similar computing device. Taking running on a server as an example, Figure 4 1 is a hardware structure block diagram of a text annotation server in one embodiment of this specification. The computer terminal may be the text annotation server or text annotation processing device in the above embodiment. Figure 4 The server 10 shown may include one or more (only one is shown in the figure) processors 100 (the processor 100 may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a non-volatile memory 200 for storing data, and a transmission module 300 for communication functions. Those skilled in the art will understand that Figure 4 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 4 More or fewer components shown in the figure may also include other processing hardware, such as a database or multi-level cache, GPU, or other hardware with Figure 4 Different configurations shown.

[0180] The non-volatile memory 200 can be used to store software programs and modules of application software, such as the program instructions / modules corresponding to the taxi data processing method in the embodiment of this specification. The processor 100 executes various functional applications and resource data updates by running the software programs and modules stored in the non-volatile memory 200. The non-volatile memory 200 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the non-volatile memory 200 may further include a memory remotely arranged relative to the processor 100, and these remote memories may be connected to the computer terminal via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0181] The transmission module 300 is used to receive or send data via a network. The specific example of the above network may include a wireless network provided by a communication provider of a computer terminal. In one example, the transmission module 300 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission module 300 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0182] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0183] The above-mentioned text annotation method or device provided in the embodiments of this specification can be implemented by a processor in a computer executing corresponding program instructions, such as using the C++ language of the Windows operating system to implement it on a PC, on a Linux system, or other systems such as Android and iOS system programming languages ​​to implement it on a smart terminal, as well as based on the processing logic of a quantum computer.

[0184] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referenced to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the hardware + program embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0185] Although one or more embodiments of the present specification provide method operation steps such as embodiments or flow charts, more or less operation steps may be included based on conventional or non-creative means. The order of steps listed in the embodiment is only one way of executing the order of many steps, and does not represent the only execution order. When the device or terminal product in practice is executed, it can be executed in sequence or in parallel according to the method shown in the embodiment or the accompanying drawings (for example, a parallel processor or a multi-threaded processing environment, or even a distributed resource data update environment). The term "include", "include" or any other variant thereof is intended to cover non-exclusive inclusion, so that the process, method, product or equipment including a series of elements includes not only those elements, but also includes other elements that are not explicitly listed, or also includes elements inherent to such process, method, product or equipment. In the absence of more restrictions, it is not excluded that there are other identical or equivalent elements in the process, method, product or equipment including the elements. The first, second, etc. words are used to represent the name, and do not represent any specific order.

[0186] For the convenience of description, the above devices are described in various modules according to their functions. Of course, when implementing one or more of the present specification, the functions of each module can be implemented in the same or more software and / or hardware, or the module implementing the same function can be implemented by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are only schematic. For example, the division of units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0187] Each embodiment in this specification is described in a progressive manner, and the same and similar parts between the embodiments can be referenced to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment. In the description of this specification, the description of the reference term "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of this specification. In this specification, the schematic representation of the above terms does not necessarily target the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples without contradiction.

[0188] The above are only examples of one or more embodiments of this specification and are not intended to limit one or more embodiments of this specification. For those skilled in the art, one or more embodiments of this specification may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this specification should be included in the scope of the claims.

Claims

1. A text annotation method, characterized in that: The method comprises: Convert the text to be annotated into the image to be annotated; Using eye tracking technology to obtain the salesperson's attention image information on the picture to be labeled, the attention image information includes the salesperson's attention area and attention frequency on the picture to be labeled; Converting the image information into character information according to the correspondence between the pixels in the image to be annotated and the characters in the text to be annotated; Screening the concerned character information based on optical character recognition technology to obtain the annotation information of the text to be annotated; Among them, the screening of the focus character information based on the optical character recognition technology to obtain the annotation information of the text to be annotated includes: constructing an optical character recognition model based on the optical character recognition technology; annotating the historical focus image information of the historical samples to be annotated based on the eye tracking technology, and converting the historical focus image information into corresponding historical focus character information; obtaining the historical sample confirmation annotation information annotated by the salesperson on the historical samples to be annotated; using the historical focus character information as the model training input data of the optical character recognition model, using the historical sample confirmation annotation information as the model training label data of the optical character recognition model, and model training the optical character recognition model until the optical character recognition model meets the preset requirements; using the trained optical character recognition model to screen the focus character information to obtain the annotation information of the text to be annotated.

2. The method according to claim 1, characterized in that The method of using the eye tracking technology to obtain the image information that the salesperson pays attention to the image to be labeled includes: Using eye tracking technology to obtain the salesperson's gaze information on the image to be labeled; According to the sight stop information, the number of times the salesperson pays attention to each pixel point in the image to be labeled is obtained; Constructing an image matrix to be labeled according to the image to be labeled, wherein the elements in the image matrix to be labeled are the pixels of the image to be labeled; The value of each element in the to-be-annotated image matrix is ​​set as the number of attentions of each element in the to-be-annotated image, and the to-be-annotated image matrix with the element values ​​determined is used as the attention image information.

3. The method according to claim 2, characterized in that The converting the focus image information into focus character information according to the correspondence between the pixel points in the to-be-annotated picture and the characters in the to-be-annotated text includes: Constructing a text matrix to be annotated according to the text to be annotated, wherein the elements in the text matrix to be annotated are characters of the text to be annotated; According to the correspondence between the characters in the text matrix to be annotated and the pixels in the picture to be annotated, the value of each element in the image matrix to be annotated is converted into the value of each element in the text matrix to be annotated; The text matrix to be marked with the determined element values ​​is used as the focus character information.

4. The method according to claim 1, characterized in that The step of screening the concerned character information based on the optical character recognition technology to obtain the annotation information of the text to be annotated includes: Acquire confirmation annotation information of the specified number of images to be annotated obtained by the salesperson performing text annotation on the specified number of images to be annotated corresponding to the text to be annotated; Using the attention character information of the specified number of images to be labeled as the optimization training input data of the optical character recognition model, using the confirmed label information of the specified number of images to be labeled as the optimization training label data of the optical character recognition model, optimizing the optical character recognition model, and obtaining an optimized optical character recognition model; The optimized optical character recognition model is used to mark the focus character information of the text to be marked, so as to obtain the marking information of the text to be marked.

5. The method according to claim 1, characterized in that The step of screening the concerned character information based on the optical character recognition technology to obtain the annotation information of the text to be annotated includes: The attention character information is screened based on optical character recognition technology to obtain a labeling character matrix corresponding to the text to be labeled, wherein the element values ​​in the labeling character matrix represent the labeling frequency of each character in the text to be labeled; Determine the two-dimensional image coordinates of the key marked area in the text to be marked according to the marked character matrix; According to the correspondence between the text to be annotated and the picture to be annotated, converting the two-dimensional picture coordinates of the key annotated area into a one-dimensional character index; The annotation information of the text to be annotated is obtained based on the one-dimensional character index of the key annotated area.

6. The method according to claim 5, characterized in that The following formula is used to determine the two-dimensional image coordinates of the key marked area: I1={(x 1_1 :x 1_2 ,and 1_1 :and 1_2 ),(x 2_1 :x 2_2 ,and 2_1 :and 2_2 )...(x n_1 :x n_2 ,and n_1 :and n_2 )} Among them, I1 represents the set of two-dimensional image coordinates of the key marked area, x n_1 Indicates the starting point of the horizontal axis of the nth highlighted area, x n_2 Indicates the end point of the horizontal axis of the nth highlighted area, y n_1 Indicates the vertical axis starting point of the nth key marked area, y n_2 Indicates the vertical axis end point of the nth highlighted area, (x n_1 :x n_2 ,y n_1 :y n_2 ) represents the two-dimensional image coordinates of the nth key marked area.

7. The method according to claim 6, characterized in that The following formula is used to convert the two-dimensional image coordinates of the key marked area into a one-dimensional character index: I2={(and 1_1 ×i+x 1_1 :and 1_2 ×i+x 1_2 ),(and 2_1 ×i+x 2_1 :and 2_2 ×i+x 2_2 )...(and n_1 ×i+x n_1 :and n_2 ×i+x n_2 )} Where I2 represents the one-dimensional character index of the key marked area, i represents the number of columns of the marked character matrix, (y n_1 ×i+x n_1 :y n_2 ×i+x n_2 ) represents the one-dimensional character index of the nth highlighted area.

8. A text annotation device, characterized in that: The device comprises: A text conversion module, used to convert the text to be annotated into the image to be annotated; An eye tracking and annotation module, used to obtain the salesperson's attention image information of the picture to be annotated by using the eye tracking technology, wherein the attention image information includes the salesperson's attention area and attention frequency of the picture to be annotated; A labeling information character conversion module, used for converting the focus image information into focus character information according to the correspondence between the pixel points in the to-be-labeled picture and the characters of the to-be-labeled text; A marking information screening module, used for screening the concerned character information based on optical character recognition technology to obtain the marking information of the text to be marked; Among them, the annotation information screening module is specifically used to: build an optical character recognition model based on optical character recognition technology; annotate historical image information of historical samples to be annotated based on eye tracking technology, and convert the historical image information into corresponding historical character information; obtain historical sample confirmation annotation information annotated by the salesperson on the historical samples to be annotated; use the historical character information as model training input data of the optical character recognition model, use the historical sample confirmation annotation information as model training label data of the optical character recognition model, and perform model training on the optical character recognition model until the optical character recognition model meets the preset requirements; use the trained optical character recognition model to screen the character information of interest to obtain the annotation information of the text to be annotated.

9. The device according to claim 8, characterized in that The eye tracking annotation module is specifically used for: Using eye tracking technology to obtain the salesperson's gaze information on the image to be labeled; According to the sight stop information, the number of times the salesperson pays attention to each pixel point in the image to be labeled is obtained; Constructing an image matrix to be labeled according to the image to be labeled, wherein the elements in the image matrix to be labeled are the pixels of the image to be labeled; The value of each element in the to-be-annotated image matrix is ​​set as the number of attentions of each element in the to-be-annotated image, and the to-be-annotated image matrix with the element values ​​determined is used as the attention image information.

10. The device according to claim 9, characterized in that The annotation information screening module is specifically used for: Constructing a text matrix to be annotated according to the text to be annotated, wherein the elements in the text matrix to be annotated are characters of the text to be annotated; According to the correspondence between the characters in the text matrix to be annotated and the pixels in the picture to be annotated, the value of each element in the image matrix to be annotated is converted into the value of each element in the text matrix to be annotated; The text matrix to be marked with the determined element values ​​is used as the focus character information.

11. A text annotation device, characterized in that: Applied to a server, the device includes at least one processor and a memory for storing processor executable instructions, and when the instructions are executed by the processor, the steps of any one of the methods of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Eye movement based interactive image retrieval method for extracting image area of interest

    CN105426399A

  • Automatic image statement marking method based on attention feedback mechanism

    CN108960338A