Information extraction method, apparatus, and system
Patent Information
- Application Number
- CA3140455
- Authority / Receiving Office
- CA · CA
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-11-25
- Filing Date
- 2021-11-25
- Publication Date
- 2026-08-11
- Estimated Expiration
- 2041-11-25
Abstract
Description
INFORMATION EXTRACTION METHOD, APPARATUS, AND SYSTEM Field
[0001] The present disclosure relates to computer technology, particularly to an information extraction method, apparatus, and system. Background
[0002] Information extraction is a kind of technology which converts the text information of natural language into key-value pairs, represents by data structuring to locate specific information in document with natural language. At present, information extraction method commonly uses automatic learning, the extraction model which is commonly used comprises: models derived from regular expression grammars, models derived from templates, and models based on structural comparison, models based on visual characteristics and so on. However, in the prior art, the information extraction methods of the above model that used for ordinary files and files with specific formats are the same, which makes it difficult to improve the accurate rate of information extraction. Invention Content
[0003] To solve the present technical problems, the present application provides an information extraction method, apparatus and system. The technical solutions are as following:
[0004] The first aspect is providing an information extraction method, comprising:
[0005] Obtaining text information of a file and position information of characters in the text information;
[0006] Constructing several sentence vectors according to the text information;
[0007] Classifying the sentence vectors in combination with location information, obtaining categories of the sentence vectors;
[0008] Generating structured string information according to the categories of the sentence vectors.
[0009] Furthermore, wherein classifying the sentence vectors in combination with location information, obtaining categories of the sentence vectors, comprising:
[0010] Representing sentence vectors as nodes, determining the position of characters contained in the text information corresponding to the sentence vector as edges to construct a image network;
[0011] Classifying the nodes in the image network by using the image network model, obtaining categories of the sentence vectors.
[0012] Furthermore, wherein generating structured string information according to the categories of the sentence vectors, comprising:
[0013] Splicing and combining the text information corresponding to sentence vectors of the same category according to the position information, generating the structured string information.
[0014] Furthermore, constructing several sentence vectors according to the text information, comprising:
[0015] Performing word segmentation processing on the text information, obtaining word segmentation;
[0016] Converting the word segmentation into word vector;
[0017] Constructing the sentence vector according to the word vector.
[0018] Furthermore, wherein converting the word segmentation into word vector, comprising: matching the word segmentation with the correspondingly word vector by using word vector model.
[0019] Furthermore, wherein constructing the sentence vector according to the word vector, comprising: using a bag of words model or a statistical model to process the word vector, constructing the sentence vector.
[0020] The second aspect is providing an information extraction apparatus, comprising:
[0021] A recognition module configured to obtain text information of a file and position information of characters in the text information;
[0022] A sentence vector construction module configured to construct several sentence vectors according to the text information;
[0023] A category recognition module configured to classify the sentence vectors in combination with location information and obtain categories of the sentence vectors;
[0024] A conversion module configured to generate structured string information according to the categories of the sentence vectors.
[0025] Furthermore, wherein the category recognition module, comprising:
[0026] An image construction module configured to represent sentence vectors as nodes, determining the position of characters contained in the text information corresponding to the sentence vector as edges to construct an image network;
[0027] A classification module configured to classifying the nodes in the image network by using the image network model, obtaining categories of the sentence vectors.
[0028] Furthermore, the conversion module specifically configured to splice and combine the text information corresponding to sentence vectors of the same category according to the position information, generating the structured string information.
[0029] Furthermore, the sentence vector construction module, comprising:
[0030] Word segmentation processing module configured to word segmentation processing on the text information, obtaining word segmentation;
[0031] Word vector obtaining module configured to convert the word segmentation into word vector;
[0032] Construction module configured to construct the sentence vector according to the word vector.
[0033] Furthermore, word vector obtaining module configured to match the word segmentation with the correspondingly word vector by using word vector model.
[0034] Furthermore, construction module configured to use a word bag model or a statistical model to process the word vector, constructing the sentence vector.
[0035] The third aspect is providing a computer system, comprising:
[0036] One or plural processors; and
[0037] A memory associated with one or plural processors, the memory is configured to store program commands, if the program commands are executed by one or plural processors, executing the information extraction method as mentioned in the aspect 1.
[0038] The beneficial effects brough by the technical solution provided in the implementation of the present invention are:
[0039] 1. The present invention aims at a file with a specific format, combining the position information of the characters in the text information to classify the constructed sentence vector in the text information, generating structured string according to the category of the sentence vector, so that when judging the sentence vector, referring to the two dimensions indicators of text and location information to ensure the accuracy of the classification, which is good for determining text information characteristics corresponding to the sentence vector according to the category of the sentence vector, thereby improving the accuracy of information extraction from files with specific formats;
[0040] 2. The present invention uses a image network model to extract structured information, which can be compared with a model based on template derivation to adapt to text information with different lengths, and which can effectively improve the accuracy, robustness and versatility of information extraction;
[0041] 3. When the present invention generates structured string information, splicing and combining the text information corresponding to sentence vectors of the same category according to the position information, ensuring the correctness of the splicing of the text information through the position information, making coherent semantics. Drawing Description
[0042] In order to describe the technical solutions clearer in the implementations of the present application or the prior art, the following are drawings that need to be used are briefly introduced. Obviously, the drawings in the following description are only some implementations of the application, for those of ordinary skill in the art, without creative work, they can also obtain other drawings based on these drawings.
[0043] Figure 1 is a process diagram of an information extraction method in implementation 1 of the present application;
[0044] Figure 2 is a structural diagram of an information extraction apparatus in implementation 2 of the present application;
[0045] Figure 3 is a structural diagram of computer system in implementation 3 of the present application. Specific implementation methods
[0046] The following will describe the technical solutions of the implementations in the present application with accompanying drawings, obviously the described implementations are only a part of the implementations in the present application. Based on the implementations in the present application, all other implementations obtained by those of ordinary skilled in the art will fall in the protection scope of the present application.
[0047] There is no information extraction method for files in specific formats in the existing information extraction technology, but we have found that the file with the specific format contains structural information themselves. If the format information can be combined with the text semantic information, and then extracting the combined information, which will be able to further improve the extracting accuracy of the file information with specific format. Therefore, to further improve the extracting accuracy of the file information with specific format, the format information of the file with specific format is combined with the semantic information, the present invention discloses an information extraction method, apparatus and system, the specific technical solutions are as follows:
[0048] As shown in Figure 1, an information extraction method includes:
[0049] S1, obtaining text information of a file and position information of characters in the text information;
[0050] In the above-mentioned, files mainly refer to files with specific formats, which specifically can be: business licenses, certificates, ID cards, receipts and so on. Text information mainly refers to characters, numbers, letters, special symbols and other characters in the files, in general, punctuation characters in the file are used as the basis for dividing sentences in the text information and are not contained in the text information.
[0051] In an implementation, step S1 is specifically using optical character recognition technology to obtain the text information in the file image and the position information of the characters in the text information in the file image.
[0052] Optical Character Recognition (OCR) technology comprises:
[0053] S11, obtaining the file image of the file and pre-processing the file image;
[0054] S12, recognizing the text direction in the file image;
[0055] S13, text detection;
[0056] S14, text recognition.
[0057] As the above-mentioned, the file image can be a photo of the file or a scanned copy of the file. Pre-processing the file image, the main purpose is to correct the imaging problems of the image, comprising: geometric transformation, blurriness removal, image enhancement, light correction and so on. Text detection is mainly to determine the text area in the image, the commonly used method is depth leaning model Faster R-CNN. Text recognition is mainly for recognizing a character or a string located by text detection, text detection is generally located by text lines. The position information of the character in Step S11 is generally the coordinate of the character line divided by the test detection process.
[0058] S2, constructing several sentence vectors according to the text information;
[0059] As the above-mentioned, since the number of words in each text line in the text information is not equal, therefore, it is necessary to construct a fixed-dimension sentence vector to represent test line, and the sentence vector is vectorized representation of a character line in the text information.
[0060] In an implementation, step S2 comprises:
[0061] S21, performing word segmentation processing on the text information to obtain word segmentation;
[0062] S22, converting the word segmentation into word vector;
[0063] S23, constructing a sentence vector according to the word vector.
[0064] As mentioned above, the word segmentation processing in step S21 can adopt the dictionary matching method, natural language model analysis method (NLP) in the prior art, one-element model method, N-element model method, etc. In step S22, the word segmentation is converted into word vector, by the matching method of the word vector model, which matches the correspondingly word vector with the segmentation word. Wherein the word vector model usually adopts Word2Vec which have completed training, Word2Vec takes a large text corpus as input to generate a vector space, each unique word in corpus is assigned a correspondingly vector in this space. In step S23, the bag of words model or statistical model can be used for processing word vector, constructing sentence vector. The bag of words model assumes that for a text, ignoring its word order, grammar, syntax and other elements, and treating it as just a collection of several words, the appearance of each word in the text is independent and does not depend on whether other words appear or not, a vector is constructed by word frequency. Statistical model such as TF- IDF, statistical-based co-occurrence matrix model, topic model and so on.
[0065] S3, classifying the sentence vectors in combination with location information, obtaining categories of the sentence vectors.
[0066] For the above-mentioned, the purpose of classifying sentence vectors is to determine whether the text information corresponding to different sentence vectors represents the same category of information, so that the correspondingly relationship between the type and the text information can be determined later. Specifically, according to different categories of sentence vectors included in different files, for example, for a business license, the categories of sentence vector can be: name, type, property, legal representative, establishment date, business period, business scope, etc.; for ID cards, the categories of sentence vector can be: name, gender, date of birth, address, ID number, etc. In general, these categories is generally the key of the structured character information, the text information corresponding to the sentence vector is usually the value of the structured character information.
[0067] In an implementation, step S3 comprises:
[0068] S31, representing the sentence vector as a node, representing the position information included in the text information corresponding to the sentence vector as edges to build a image network;
[0069] S32, using the image network model to classify the nodes in the image network to obtain the category of the sentence vector;
[0070] For the above-mentioned, since the sentence vector is converted by a line of characters in the text information, therefore, each sentence of file information and the position information of the characters in each sentence are included in the image network. The image network model uses neural network model which completes training by image network with classification marks. The image network model has a high inductive bias, so the number of samples required for its training is less than the general neural network model. The output is the probability of each node in different categories when classifying, judging the category of the node according to the probability, and then obtaining the category of the sentence vector. The present invention also considers the position information of the characters when classifying sentence vectors, then makes the sentence vector corresponding to the text information of same character type more accurate when classifying, for example, during the receipt information extraction process, representing the character of numeric type adopted by amounts and unit price, it is easy to be confused by the general information extraction metho, but judging the type by combing the position information will greatly improve the accuracy. In addition, there is no template rule for the image network model, comparing with the general template derivation model, it is more suitable for text information with different lengths and more flexible.
[0071] S4, generating structured string information according to the categories of the sentence vectors.
[0072] In an implementation, step S4 includes splicing and combining the text information corresponding to sentence vectors of the same category according to the position information, generating the structured string information.
[0073] For the above-mentioned, the splicing and combination of text information is carried out in the order of the coordinates, without considering the semantic, then ensuring the semantic coherence and smoothness of the text information corresponding to each sentence vector after the completion of the splicing. To be instructed, the structured representation of string information mainly refers to the output of string information in the form of key-value pairs (key = value).
[0074] As shown in Figure 2, based on the above-mentioned information extraction method, the present invention also provides an information extraction apparatus, comprising:
[0075] The recognition module 201 configured to obtain text information of a file and position information of characters in the text information.
[0076] For the above-mentioned, the file mainly refers to a file with a specific format, and the text information mainly refers to the words, numbers, letters, special symbols and so on, in general, punctuation characters in the file are used as the basis for dividing sentences in the text information and are not contained in the text information.
[0077] In an implementation, the recognition module 201, specifically using optical character recognition technology to obtain the text information in the file image and the position information of the characters in the text information in the file image.
[0078] The sentence vector construction module 202 configured to construct several sentence vectors according to the text information.
[0079] In an implementation, the sentence vector construction module 202, comprising:
[0080] The segmentation word processing module configured to perform word segmentation processing on the text information to obtain word segmentation.
[0081] The word vector obtaining module configured to convert the word segmentation into word vector.
[0082] The construction module configured to construct a sentence vector according to the word vector.
[0083] In an implementation, the word vector obtaining module, matching the word segmentation with the correspondingly word vector by using word vector model.
[0084] In an implementation, the construction module, using a bag of words model or a statistical model to process the word vector, constructing the sentence vector.
[0085] The category recognition module 203 configured to classify the sentence vectors in combination with location information and obtain categories of the sentence vectors.
[0086] In an implementation, the category recognition module 203, comprising:
[0087] An image construction module configured to represent sentence vectors as nodes, determining the position of characters contained in the text information corresponding to the sentence vector as edges to construct an image network.
[0088] A classification module configured to classifying the nodes in the image network by using the image network model, obtaining categories of the sentence vectors.
[0089] The conversion module 204 configured to generate structured string information according to the categories of the sentence vectors.
[0090] In an implementation, the conversion module 204 is specifically configured to splice and combine the text information corresponding to sentence vectors of the same category according to the position information, generating the structured string information.
[0091] According to the above-mentioned information extraction method, the present invention also provides a computer system, comprising:
[0092] One or plural processors; and
[0093] A memory associated with one or plural processors, the memory is configured to store program commands, if the program commands are executed by one or plural processors, executing the above- mentioned information extraction method.
[0094] Wherein, Figure 3 exemplarily shows the architecture of the computer system, which can specifically include a processor 310, video display adapter 311, disk driver 312, input / output interface 313, network interface 314, and memory 320. The above-mentioned processor 310, video display adapter 311, disk driver 312, input / output interface 313, network interface 314 and memory 320 can be connected through a communication bus 330.
[0095] Wherein, the processor 310 can be achieved by using a general CPU (Central Processing Unit), Microprocessor, Application Specific Integrated Circuit (ASIC), or one or more integrated circuits, which are used to execute some relative program to achieve the technical solutions provided in this application.
[0096] The memory 320 can adopt ROM (Read Only Memory), RAM (Random Access Memory), static storage devices and dynamic storage devices to achieve. The memory 320 can store operate system 321 used to control the running of the computer system 300, used to control the low-level operation of the computer system 300's Basic Input Output System (BIOS) 322. In addition, storing a web browser 323, data storage management 324, and device identity information processing system 325 and so on. The above-mentioned device identity information processing system 325 can be the specific application that implements the above-mentioned steps. To sum up, when achieving the technical solutions provided by this application through software or firmware, related program codes are stored in the memory 320 and executed by a processor 310.
[0097] Input / output interface 313 is used for connecting input / output modules to achieve the information input and output. Input / output module can be configured in the device as a component ( not shown in the figure), or it can be connected to the device to provide corresponding functions. Wherein, Input devices can include keyboards, mice, touch screens, microphones, various sensors, etc., and output devices can include monitors, speakers, vibrators, lights and so on.
[0098] The network interface 314 is used to connect a communication module (not shown in the figure) to achieve the communication interaction between this device and other devices. Wherein, the communication module can achieve communication through wired means (such as USB, network cable, etc.), or through wireless methods (such as mobile network, WIFI, Bluetooth, etc.) to achieve communication.
[0099] The bus 330 includes a path and transmits information among various components of the device (such as the processor 310, the video display adapter 311, the disk driver 312, the input / output interface 313, the network interface 314, and the memory 320).
[0100] In addition, the electronic device 300 also can obtain information with specific receiving conditions from the virtual resource object's receiving condition information database 341 for condition judgement and son on.
[0101] It should be noted that although the above device only shows the processor 310, the video display adapter 311, the disk driver 312, input / output interface 313, network interface 314, memory 320, bus 330, etc., but in the process of the specific implementation, the device may also include other essential components for normal operation. In addition, those skilled in the art can understand that the above apparatus can comprise only the essential components of the present application to achieve the implementation, but there is no need to contain all the components as shown in figure.
[0102] Known from the description of the above implementations that those skilled in the art can clearly understand that the application can be achieved with the help of software and essential general hardware platform. Based on this understanding, the essence of the technical solution of this application, or in other words, the part that contributes to the existing technology can be implemented in the form of a software product, the computer software product can be stored in storage media, such as ROM / RAM, magnetic disks, optical disks, etc., including several commands to make a computer device (can be a personal computer, a cloud server, or a network device, etc.) to execute the methods described in each implementation or some of the implementations of the present application.
[0103] The various implementations in this description are described in a progressive manner, the same and similar parts among the various implementations can be referred to each other separately, and each implementation focuses on the differences compared with the other implementations. Especially for the concern of the system or the system implementations, since it is basically similar to the implementation method, the description is relatively simple. For related details, please refer to the implementation method. The system and system implementations described in the above are only illustrative, and the units described by separate parts may or may not be physically separate, and the parts displayed as units may or may not be physical units, which means, it can be in one place, or it may be distributed to plural network units. Some or all the modules are selected according to actual needs to achieve the implementation's solution purpose. The ordinary skill in the art can understand and implement without creative work.
[0104] The beneficial effects of the technical solutions provided by the implementations of the present invention are:
[0105] 1. The present invention aims at a file with a specific format, combining the position information of the characters in the text information to classify the constructed sentence vector in the text information, generating structured string according to the category of the sentence vector, so that when judging the sentence vector, referring to the two dimensions indicators of text and location information to ensure the accuracy of the classification, which is good for determining text information characteristics corresponding to the sentence vector according to the category of the sentence vector, thereby improving the accuracy of information extraction from files with specific formats;
[0106] 2. The present invention uses a image network model to extract structured information, which can be compared with a model based on template derivation to adapt to text information with different lengths, and which can effectively improve the accuracy, robustness and versatility of information extraction;
[0107] 3. When the present invention generates structured string information, splicing and combining the text information corresponding to sentence vectors of the same category according to the position information, ensuring the correctness of the splicing of the text information through the position information, making coherent semantics.
[0108] All the above-mentioned optional technical solutions can be combined in any way to form an optional implementation of the present invention which will not be repeated here.
[0109] The above-mentioned are only preferred implementations of the present invention, but not used to restrict the present invention, anything in the spirits and the principles of the present invention, then any modifications, equivalent replacements, improvements shall be included in the protection scope of the present invention.
Claims
<pat:ClaimStatement>CLAIMS: An information extraction system, the system comprising: a processor, the processor comprising: a recognition module configured to obtain text information of a file and position information of characters in the text information; a sentence vector construction module configured to construct a plurality of sentence vectors according to the text information, wherein the sentence vectors are constructed according to a word vector, and wherein the sentence vectors are further constructed by matching a bag of words model or a statistical model to process the word vector; a category recognition module configured to classify the sentence vectors in combination with location information of the sentence vectors in text and obtain a categories of the sentence vectors; a conversion module configured to generate structured string information according to the categories of the sentence vectors; an image construction module configured to represent the sentence vectors as nodes, determining a position of characters contained in the text information corresponding to the sentence vector as edges to construct an image network; and a classification module configured to classifying the nodes in the image network by using an image network model, obtaining the categories of the sentence vectors.< / pat:ClaimStatement> <pat:Claims com:id="claims"> <pat:Claim com:id="CLM-00002"> <pat:ClaimNumber>2< / pat:ClaimNumber> <pat:ClaimText>2. The system of claim 1, wherein the conversion module is configured to splice and combine the text information corresponding to the sentence vectors of the same category according to the position information, generating the structured string information. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00003"> <pat:ClaimNumber>3< / pat:ClaimNumber> <pat:ClaimText>3. The system of claim 2, wherein the processor is configured to perform word segmentation processing on the text information. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00004"> <pat:ClaimNumber>4< / pat:ClaimNumber> <pat:ClaimText>4. The system of claim 3, wherein the processor is configured to obtain a word segmentation. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00005"> <pat:ClaimNumber>5< / pat:ClaimNumber> <pat:ClaimText>5. The system of claim 4, wherein the processor is configured to obtain the word segmentation into a word vector. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00006"> <pat:ClaimNumber>6< / pat:ClaimNumber> <pat:ClaimText>6. The system of claim 5, wherein the processor is configured to match the word segmentation with the corresponding word vector by using word vector model. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00007"> <pat:ClaimNumber>7< / pat:ClaimNumber> <pat:ClaimText>7. The system of claim 6, wherein files include business licenses, certificates, ID cards, and receipts. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00008"> <pat:ClaimNumber>8< / pat:ClaimNumber> <pat:ClaimText>8. The system of claim 7, wherein the text information includes characters, numbers, letters, and special symbols. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00009"> <pat:ClaimNumber>9< / pat:ClaimNumber> <pat:ClaimText>9. The system of claim 8, wherein punctuation characters are used as the basis for dividing sentences in the text information and are not contained in the text information. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00010"> <pat:ClaimNumber>10< / pat:ClaimNumber> <pat:ClaimText>10. The system of claim 9, wherein using optical character recognition technology to obtain the text information in a file image and the position information of the characters in the text information in the file image. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00011"> <pat:ClaimNumber>11< / pat:ClaimNumber> <pat:ClaimText>11. The system of claim 10, wherein the processor is configured to obtain the file image of the file and pre- processing the file image. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00012"> <pat:ClaimNumber>12< / pat:ClaimNumber> <pat:ClaimText>12. The system of claim 11, wherein the processor is configured to recognize text direction in the file image. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00013"> <pat:ClaimNumber>13< / pat:ClaimNumber> <pat:ClaimText>13. The system of claim 12, wherein the processor is configured to provide a text detection and a text recognition. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00014"> <pat:ClaimNumber>14< / pat:ClaimNumber> <pat:ClaimText>14. The system of claim 13, wherein the file image is a photo of the file and a scanned copy of the file. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00015"> <pat:ClaimNumber>15< / pat:ClaimNumber> <pat:ClaimText>15. The system of claim 14, wherein the processor is configured to pre-process the file image to correct imaging problems of the image. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00016"> <pat:ClaimNumber>16< / pat:ClaimNumber> <pat:ClaimText>16. The system of claim 15, wherein the processor is configured to provide geometric transformation, blurriness removal, image enhancement, and light correction. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00017"> <pat:ClaimNumber>17< / pat:ClaimNumber> <pat:ClaimText>17. The system of claim 16, wherein the text detection is performed to determine a text area in the image. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00018"> <pat:ClaimNumber>18< / pat:ClaimNumber> <pat:ClaimText>18. The system of claim 17, wherein the text recognition is performed for recognizing a character or a string located by the text detection. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00019"> <pat:ClaimNumber>19< / pat:ClaimNumber> <pat:ClaimText>19. The system of claim 18, wherein the text detection performed by text lines. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00020"> <pat:ClaimNumber>20< / pat:ClaimNumber> <pat:ClaimText>20. The system of claim 19, wherein the position information of the character is a coordinate of a character line divided by a text detection process. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00021"> <pat:ClaimNumber>21< / pat:ClaimNumber> <pat:ClaimText>21. The system of claim 20, wherein the processor is configured to construct the sentence vectors according to the text information. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00022"> <pat:ClaimNumber>22< / pat:ClaimNumber> <pat:ClaimText>22. The system of claim 21, wherein the processor is configured to construct a fixed-dimension sentence vector to represent test line. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00023"> <pat:ClaimNumber>23< / pat:ClaimNumber> <pat:ClaimText>23. The system of claim 22, wherein the sentence vector is vectorized representation of the character line in the text information. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00024"> <pat:ClaimNumber>24< / pat:ClaimNumber> <pat:ClaimText>24. The system of claim 23, wherein the processor is configured to perform the word segmentation processing on the text information to obtain the word segmentation. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00025"> <pat:ClaimNumber>25< / pat:ClaimNumber> <pat:ClaimText>25. The system of claim 24, wherein the word segmentation processing adopts a dictionary matching system. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00026"> <pat:ClaimNumber>26< / pat:ClaimNumber> <pat:ClaimText>26. The system of claim 25, wherein the word segmentation processing adopts a natural language model analysis system (NLP). < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00027"> <pat:ClaimNumber>27< / pat:ClaimNumber> <pat:ClaimText>27. The system of claim 26, wherein the word segmentation processing adopts a one-element model system. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00028"> <pat:ClaimNumber>28< / pat:ClaimNumber> <pat:ClaimText>28. The system of claim 27, wherein the word segmentation processing adopts a N-element model system. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00029"> <pat:ClaimNumber>29< / pat:ClaimNumber> <pat:ClaimText>29. The system of claim 28, wherein the word segmentation processing adopts the dictionary matching system. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00030"> <pat:ClaimNumber>30< / pat:ClaimNumber> <pat:ClaimText>30. The system of claim 29, wherein the word segmentation is converted into the word vector by a matching system of the word vector model, which matches a corresponding word vector with the segmentation word. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00031"> <pat:ClaimNumber>31< / pat:ClaimNumber> <pat:ClaimText>31. The system of claim 30, wherein the word vector model includes Word2Vec. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00032"> <pat:ClaimNumber>32< / pat:ClaimNumber> <pat:ClaimText>32. The system of claim 31, wherein the Word2Vec includes a large text corpus as input to generate a vector space. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00033"> <pat:ClaimNumber>33< / pat:ClaimNumber> <pat:ClaimText>33. The system of claim 32, wherein each unique word in corpus is assigned a corresponding vector in this space. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00034"> <pat:ClaimNumber>34< / pat:ClaimNumber> <pat:ClaimText>34. The system of claim 33, wherein the bag of words model includes a text, and discards word order, grammar, and syntax. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00035"> <pat:ClaimNumber>35< / pat:ClaimNumber> <pat:ClaimText>35. The system of claim 34, wherein the processor is configured to determine whether the text information corresponding to different sentence vectors represents the same category of information. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00036"> <pat:ClaimNumber>36< / pat:ClaimNumber> <pat:ClaimText>36. The system of claim 35, wherein the processor is configured to determine a corresponding relationship between a type and the text information. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00037"> <pat:ClaimNumber>37< / pat:ClaimNumber> <pat:ClaimText>37. The system of claim 36, wherein the processor is configured to determine the corresponding relationship according to different categories of the sentence vectors included in the files. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00038"> <pat:ClaimNumber>38< / pat:ClaimNumber> <pat:ClaimText>38. The system of claim 37, wherein the categories of sentence vector include name, type, property, legal representative, establishment date, business period, and business scope. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00039"> <pat:ClaimNumber>39< / pat:ClaimNumber> <pat:ClaimText>39. The system of claim 38, wherein the categories of the sentence vector include ID cards, name, gender, date of birth, address, and ID number. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00040"> <pat:ClaimNumber>40< / pat:ClaimNumber> <pat:ClaimText>40. The system of claim 39, wherein the processor is configured to generate the structured string information according to the categories of the sentence vectors. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00041"> <pat:ClaimNumber>41< / pat:ClaimNumber> <pat:ClaimText>41. The system of claim 40, wherein the splicing and combination of the text information is carried in the order of coordinates, without considering a semantic. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00042"> <pat:ClaimNumber>42< / pat:ClaimNumber> <pat:ClaimText>42. The system of claim 41, wherein the processor is configured to ensure a semantic coherence and smoothness of the text information corresponding to each sentence vector after the completion of the splicing. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00043"> <pat:ClaimNumber>43< / pat:ClaimNumber> <pat:ClaimText>43. The system of claim 42, wherein the structured representation of the string information mainly refers to the output of string information in a form of key-value pairs. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00044"> <pat:ClaimNumber>44< / pat:ClaimNumber> <pat:ClaimText>44. An information extraction apparatus, the apparatus comprising: a recognition module configured to obtain text information of a file and position information of characters in the text information; a sentence vector construction module configured to construct a plurality of sentence vectors according to the text information, further configured to construct the sentence vector according to a word vector, and further configured to use a bag of words model or a statistical model to process the word vector to construct the sentence vector; a category recognition module configured to classify the sentence vectors in combination with location information of the sentence vectors in text and obtain a categories of the sentence vectors; a conversion module configured to generate structured string information according to the categories of the sentence vectors; an image construction module configured to represent the sentence vectors as nodes, determining a position of characters contained in the text information corresponding to the sentence vector as edges to construct an image network; and a classification module configured to classifying the nodes in the image network by using an image network model, obtaining the categories of the sentence vectors. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00045"> <pat:ClaimNumber>45< / pat:ClaimNumber> <pat:ClaimText>45. The apparatus of claim 44, wherein the conversion module is configured to splice and combine the text information corresponding to the sentence vectors of the same category according to the position information, generating the structured string information. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00046"> <pat:ClaimNumber>46< / pat:ClaimNumber> <pat:ClaimText>46. The apparatus of claim 45, wherein the apparatus further comprises: performing a word segmentation processing on the text information. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00047"> <pat:ClaimNumber>47< / pat:ClaimNumber> <pat:ClaimText>47. The apparatus of claim 46, wherein the apparatus further comprises: obtaining a word segmentation. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00048"> <pat:ClaimNumber>48< / pat:ClaimNumber> <pat:ClaimText>48. The apparatus of claim 47, wherein the apparatus further comprises: converting the word segmentation into a word vector. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00049"> <pat:ClaimNumber>49< / pat:ClaimNumber> <pat:ClaimText>49. The apparatus of claim 48, wherein the apparatus further comprises: matching the word segmentation with the corresponding word vector by using word vector model. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00050"> <pat:ClaimNumber>50< / pat:ClaimNumber> <pat:ClaimText>50. The apparatus of claim 49, wherein files include business licenses, certificates, ID cards, and receipts. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00051"> <pat:ClaimNumber>51< / pat:ClaimNumber> <pat:ClaimText>51. The apparatus of claim 50, wherein the text information includes characters, numbers, letters, and special symbols. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00052"> <pat:ClaimNumber>52< / pat:ClaimNumber> <pat:ClaimText>52. The apparatus of claim 51, wherein punctuation characters are used as the basis for dividing sentences in the text information and are not contained in the text information. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00053"> <pat:ClaimNumber>53< / pat:ClaimNumber> <pat:ClaimText>53. The apparatus of claim 52, wherein using optical character recognition technology to obtain the text information in a file image and the position information of the characters in the text information in the file image. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00054"> <pat:ClaimNumber>54< / pat:ClaimNumber> <pat:ClaimText>54. The apparatus of claim 53, wherein the apparatus further comprises: obtaining the file image of the file and pre-processing the file image. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00055"> <pat:ClaimNumber>55< / pat:ClaimNumber> <pat:ClaimText>55. The apparatus of claim 54, wherein the apparatus further comprises: recognizing text direction in the file image. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00056"> <pat:ClaimNumber>56< / pat:ClaimNumber> <pat:ClaimText>56. The apparatus of claim 55, wherein the apparatus further comprises: a text detection and a text recognition. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00057"> <pat:ClaimNumber>57< / pat:ClaimNumber> <pat:ClaimText>57. The apparatus of claim 56, wherein the file image is a photo of the file and a scanned copy of the file. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00058"> <pat:ClaimNumber>58< / pat:ClaimNumber> <pat:ClaimText>58. The apparatus of claim 57, wherein the apparatus further comprises: pre-processing the file image to correct imaging problems of the image. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00059"> <pat:ClaimNumber>59< / pat:ClaimNumber> <pat:ClaimText>59. The apparatus of claim 58, wherein the apparatus further comprises: geometric transformation, blurriness removal, image enhancement, and light correction. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00060"> <pat:ClaimNumber>60< / pat:ClaimNumber> <pat:ClaimText>60. The apparatus of claim 59, wherein the text detection is performed to determine a text area in the image. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00061"> <pat:ClaimNumber>61< / pat:ClaimNumber> <pat:ClaimText>61. The apparatus of claim 60, wherein the text recognition is performed for recognizing a character or a string located by the text detection. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00062"> <pat:ClaimNumber>62< / pat:ClaimNumber> <pat:ClaimText>62. The apparatus of claim 61, wherein the text detection performed by text lines. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00063"> <pat:ClaimNumber>63< / pat:ClaimNumber> <pat:ClaimText>63. The apparatus of claim 62, wherein the position information of the character is a coordinate of a character line divided by a text detection process. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00064"> <pat:ClaimNumber>64< / pat:ClaimNumber> <pat:ClaimText>64. The apparatus of claim 63, wherein the apparatus further comprises: constructing the sentence vectors according to the text information. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00065"> <pat:ClaimNumber>65< / pat:ClaimNumber> <pat:ClaimText>65. The apparatus of claim 64, wherein the apparatus further comprises: constructing a fixed-dimension sentence vector to represent test line. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00066"> <pat:ClaimNumber>66< / pat:ClaimNumber> <pat:ClaimText>66. The apparatus of claim 65, wherein the sentence vector is vectorized representation of the character line in the text information. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00067"> <pat:ClaimNumber>67< / pat:ClaimNumber> <pat:ClaimText>67. The apparatus of claim 66, wherein the apparatus further comprises: performing the word segmentation processing on the text information to obtain the word segmentation. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00068"> <pat:ClaimNumber>68< / pat:ClaimNumber> <pat:ClaimText>68. The apparatus of claim 67, wherein the word segmentation processing adopts a dictionary matching apparatus. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00069"> <pat:ClaimNumber>69< / pat:ClaimNumber> <pat:ClaimText>69. The apparatus of claim 68, wherein the word segmentation processing adopts a natural language model analysis apparatus (NLP). < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00070"> <pat:ClaimNumber>70< / pat:ClaimNumber> <pat:ClaimText>70. The apparatus of claim 69, wherein the word segmentation processing adopts the one-element model apparatus. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00071"> <pat:ClaimNumber>71< / pat:ClaimNumber> <pat:ClaimText>71. The apparatus of claim 70, wherein the word segmentation processing adopts the N-element model apparatus. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00072"> <pat:ClaimNumber>72< / pat:ClaimNumber> <pat:ClaimText>72. The apparatus of claim 71, wherein the word segmentation processing adopts the dictionary matching apparatus. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00073"> <pat:ClaimNumber>73< / pat:ClaimNumber> <pat:ClaimText>73. The apparatus of claim 72, wherein the word segmentation is converted into the word vector by the matching apparatus of the word vector model, which matches the corresponding word vector with the segmentation word. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00074"> <pat:ClaimNumber>74< / pat:ClaimNumber> <pat:ClaimText>74. The apparatus of claim 73, wherein the word vector model includes Word2Vec. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00075"> <pat:ClaimNumber>75< / pat:ClaimNumber> <pat:ClaimText>75. The apparatus of claim 74, wherein the Word2Vec includes a large text corpus as input to generate a vector space. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00076"> <pat:ClaimNumber>76< / pat:ClaimNumber> <pat:ClaimText>76. The apparatus of claim 75, wherein each unique word in corpus is assigned a correspondingly vector in this space. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00077"> <pat:ClaimNumber>77< / pat:ClaimNumber> <pat:ClaimText>77. The apparatus of claim 76, wherein the bag of words model and statistical model is used for processing word vector and constructing sentence vector. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00078"> <pat:ClaimNumber>78< / pat:ClaimNumber> <pat:ClaimText>78. The apparatus of claim 77, wherein the bag of words model includes a text, and discards word order, grammar, and syntax. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00079"> <pat:ClaimNumber>79< / pat:ClaimNumber> <pat:ClaimText>79. The apparatus of claim 78, wherein the apparatus further comprises: classifying the sentence vectors in combination with location information. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00080"> <pat:ClaimNumber>80< / pat:ClaimNumber> <pat:ClaimText>80. The apparatus of claim 79, wherein the apparatus further comprises: obtaining categories of the sentence vectors. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00081"> <pat:ClaimNumber>81< / pat:ClaimNumber> <pat:ClaimText>81. The apparatus of claim 80, wherein the apparatus further comprises: determining whether the text information corresponding to different sentence vectors represents the same category of information. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00082"> <pat:ClaimNumber>82< / pat:ClaimNumber> <pat:ClaimText>82. The apparatus of claim 81, wherein the apparatus further comprises: determining a corresponding relationship between a type and the text information. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00083"> <pat:ClaimNumber>83< / pat:ClaimNumber> <pat:ClaimText>83. The apparatus of claim 82, wherein the apparatus further comprises: determining the corresponding relationship according to different categories of the sentence vectors included in the files. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00084"> <pat:ClaimNumber>84< / pat:ClaimNumber> <pat:ClaimText>84. The apparatus of claim 83, wherein the categories of sentence vector include name, type, property, legal representative, establishment date, business period, and business scope. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00085"> <pat:ClaimNumber>85< / pat:ClaimNumber> <pat:ClaimText>85. The apparatus of claim 84, wherein the categories of the sentence vector include ID cards, name, gender, date of birth, address, and ID number. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00086"> <pat:ClaimNumber>86< / pat:ClaimNumber> <pat:ClaimText>86. The apparatus of claim 85, wherein the apparatus further comprises: generating the structured string information according to the categories of the sentence vectors. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00087"> <pat:ClaimNumber>87< / pat:ClaimNumber> <pat:ClaimText>87. The apparatus of claim 86, wherein the splicing and combination of the text information is carried in the order of coordinates, without considering a semantic. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00088"> <pat:ClaimNumber>88< / pat:ClaimNumber> <pat:ClaimText>88. The apparatus of claim 87, wherein the apparatus further comprises: ensuring a semantic coherence and smoothness of the text information corresponding to each sentence vector after the completion of the splicing. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00089"> <pat:ClaimNumber>89< / pat:ClaimNumber> <pat:ClaimText>89. The apparatus of claim 88, wherein the structured representation of the string information mainly refers to the output of string information in a form of key-value pairs. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00090"> <pat:ClaimNumber>90< / pat:ClaimNumber> <pat:ClaimText>90. A computer readable physical memory having stored thereon a computer program executed by a computer configured to: obtain text information of a file and position information of characters in the text information; construct a plurality of sentence vectors according to the text information, construct the sentence vector according to the word vector, and use a bag of words model or a statistical model to process the word vector to construct the sentence vector; classify the sentence vectors in combination with location information of the sentence vectors in the text; obtain categories of the sentence vectors; generate structured string information according to the categories of the sentence vectors; represent the sentence vectors as nodes, determining a position of characters contained in the text information corresponding to the sentence vector as edges to construct an image network; and classify the nodes in the image network by using an image network model, obtaining the categories of the sentence vectors. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00091"> <pat:ClaimNumber>91< / pat:ClaimNumber> <pat:ClaimText>91. The memory of claim 90, wherein the conversion module is configured to splice and combine the text information corresponding to the sentence vectors of the same category according to the position information, generating the structured string information. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00092"> <pat:ClaimNumber>92< / pat:ClaimNumber> <pat:ClaimText>92. The memory of claim 91, wherein the memory further comprises: performing a word segmentation processing on the text information. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00093"> <pat:ClaimNumber>93< / pat:ClaimNumber> <pat:ClaimText>93. The memory of claim 92, wherein the memory further comprises: obtaining a word segmentation. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00094"> <pat:ClaimNumber>94< / pat:ClaimNumber> <pat:ClaimText>94. The memory of claim 93, wherein the memory further comprises: converting the word segmentation into a word vector. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00095"> <pat:ClaimNumber>95< / pat:ClaimNumber> <pat:ClaimText>95. The memory of claim 94, wherein the memory further comprises: matching the word segmentation with the corresponding word vector by using word vector model. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00096"> <pat:ClaimNumber>96< / pat:ClaimNumber> <pat:ClaimText>96. The memory of claim 95, wherein files include business licenses, certificates, ID cards, and receipts. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00097"> <pat:ClaimNumber>97< / pat:ClaimNumber> <pat:ClaimText>97. The memory of claim 96, wherein the text information includes characters, numbers, letters, and special symbols. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00098"> <pat:ClaimNumber>98< / pat:ClaimNumber> <pat:ClaimText>98. The memory of claim 97, wherein punctuation characters are used as the basis for dividing sentences in the text information and are not contained in the text information. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00099"> <pat:ClaimNumber>99< / pat:ClaimNumber> <pat:ClaimText>99. The memory of claim 98, wherein using optical character recognition technology to obtain the text information in a file image and the position information of the characters in the text information in the file image. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00100"> <pat:ClaimNumber>100< / pat:ClaimNumber> <pat:ClaimText>100. The memory of claim 99, wherein the memory further comprises: obtaining the file image of the file and pre-processing the file image. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00101"> <pat:ClaimNumber>101< / pat:ClaimNumber> <pat:ClaimText>101. The memory of claim 100, wherein the memory further comprises: recognizing text direction in the file image. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00102"> <pat:ClaimNumber>102< / pat:ClaimNumber> <pat:ClaimText>102. The memory of claim 101, wherein the memory further comprises: a text detection and a text recognition. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00103"> <pat:ClaimNumber>103< / pat:ClaimNumber> <pat:ClaimText>103. The memory of claim 102, wherein the file image is a photo of the file and a scanned copy of the file. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00104"> <pat:ClaimNumber>104< / pat:ClaimNumber> <pat:ClaimText>104. The memory of claim 103, wherein the memory further comprises: pre-processing the file image to correct imaging problems of the image. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00105"> <pat:ClaimNumber>105< / pat:ClaimNumber> <pat:ClaimText>105. The memory of claim 104, wherein the memory further comprises: geometric transformation, blurriness removal, image enhancement, and light correction. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00106"> <pat:ClaimNumber>106< / pat:ClaimNumber> <pat:ClaimText>106. The memory of claim 105, wherein the text detection is performed to determine a text area in the image. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00107"> <pat:ClaimNumber>107< / pat:ClaimNumber> <pat:ClaimText>107. The memory of claim 106, wherein the text recognition is performed for recognizing a character or a string located by the text detection. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00108"> <pat:ClaimNumber>108< / pat:ClaimNumber> <pat:ClaimText>108. The memory of claim 107, wherein the text detection performed by text lines. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00109"> <pat:ClaimNumber>109< / pat:ClaimNumber> <pat:ClaimText>109. The memory of claim 108, wherein the position information of the character is a coordinate of a character line divided by a text detection process. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00110"> <pat:ClaimNumber>110< / pat:ClaimNumber> <pat:ClaimText>110. The memory of claim 109, wherein the memory further comprises: constructing the sentence vectors according to the text information. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00111"> <pat:ClaimNumber>111< / pat:ClaimNumber> <pat:ClaimText>111. The memory of claim 110, wherein the memory further comprises: constructing a fixed-dimension sentence vector to represent test line. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00112"> <pat:ClaimNumber>112< / pat:ClaimNumber> <pat:ClaimText>112. The memory of claim 111, wherein the sentence vector is vectorized representation of the character line in the text information. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00113"> <pat:ClaimNumber>113< / pat:ClaimNumber> <pat:ClaimText>113. The memory of claim 112, wherein the memory further comprises: performing the word segmentation processing on the text information to obtain the word segmentation. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00114"> <pat:ClaimNumber>114< / pat:ClaimNumber> <pat:ClaimText>114. The memory of claim 113, wherein the word segmentation processing adopts a dictionary matching memory. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00115"> <pat:ClaimNumber>115< / pat:ClaimNumber> <pat:ClaimText>115. The memory of claim 114, wherein the word segmentation processing adopts a natural language model analysis memory (NLP). < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00116"> <pat:ClaimNumber>116< / pat:ClaimNumber> <pat:ClaimText>116. The memory of claim 115, wherein the word segmentation processing adopts the one-element model memory. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00117"> <pat:ClaimNumber>117< / pat:ClaimNumber> <pat:ClaimText>117. The memory of claim 116, wherein the word segmentation processing adopts the N-element model memory. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00118"> <pat:ClaimNumber>118< / pat:ClaimNumber> <pat:ClaimText>118. The memory of claim 117, wherein the word segmentation processing adopts the dictionary matching memory. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00119"> <pat:ClaimNumber>119< / pat:ClaimNumber> <pat:ClaimText>119. The memory of claim 118, wherein the word segmentation is converted into the word vector by the matching memory of the word vector model, which matches the corresponding word vector with the segmentation word. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00120"> <pat:ClaimNumber>120< / pat:ClaimNumber> <pat:ClaimText>120. The memory of claim 119, wherein the word vector model includes Word2Vec. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00121"> <pat:ClaimNumber>121< / pat:ClaimNumber> <pat:ClaimText>121. The memory of claim 120, wherein the Word2Vec includes a large text corpus as input to generate a vector space. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00122"> <pat:ClaimNumber>122< / pat:ClaimNumber> <pat:ClaimText>122. The memory of claim 121, wherein each unique word in corpus is assigned a corresponding vector in this space. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00123"> <pat:ClaimNumber>123< / pat:ClaimNumber> <pat:ClaimText>123. The memory of claim 122, wherein the bag of words model and statistical model is used for processing word vector and constructing sentence vector. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00124"> <pat:ClaimNumber>124< / pat:ClaimNumber> <pat:ClaimText>124. The memory of claim 123, wherein the bag of words model includes a text, and discards word order, grammar, and syntax. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00125"> <pat:ClaimNumber>125< / pat:ClaimNumber> <pat:ClaimText>125. The memory of claim 124, wherein the memory further comprises: determining whether the text information corresponding to different sentence vectors represents the same category of information. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00126"> <pat:ClaimNumber>126< / pat:ClaimNumber> <pat:ClaimText>126. The memory of claim 125, wherein the memory further comprises: determining a corresponding relationship between a type and the text information. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00127"> <pat:ClaimNumber>127< / pat:ClaimNumber> <pat:ClaimText>127. The memory of claim 126, wherein the memory further comprises: determining the corresponding relationship according to different categories of the sentence vectors included in the files. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00128"> <pat:ClaimNumber>128< / pat:ClaimNumber> <pat:ClaimText>128. The memory of claim 127, wherein the categories of sentence vector include name, type, property, legal representative, establishment date, business period, and business scope. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00129"> <pat:ClaimNumber>129< / pat:ClaimNumber> <pat:ClaimText>129. The memory of claim 128, wherein the categories of the sentence vector include ID cards, name, gender, date of birth, address, and ID number. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00130"> <pat:ClaimNumber>130< / pat:ClaimNumber> <pat:ClaimText>130. The memory of claim 129, wherein the memory further comprises: generating the structured string information according to the categories of the sentence vectors. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00131"> <pat:ClaimNumber>131< / pat:ClaimNumber> <pat:ClaimText>131. The memory of claim 130, wherein the splicing and combination of the text information is carried in the order of coordinates, without considering a semantic. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00132"> <pat:ClaimNumber>132< / pat:ClaimNumber> <pat:ClaimText>132. The memory of claim 131, wherein the memory further comprises: ensuring a semantic coherence and smoothness of the text information corresponding to each sentence vector after the completion of the splicing. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00133"> <pat:ClaimNumber>133< / pat:ClaimNumber> <pat:ClaimText>133. The memory of claim 132, wherein the structured representation of the string information mainly refers to the output of string information in a form of key-value pairs. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00134"> <pat:ClaimNumber>134< / pat:ClaimNumber> <pat:ClaimText>134. An information extraction method, the method comprising: obtaining text information of a file and position information of characters in the text information; constructing a plurality of sentence vectors according to the text information, constructing the sentence vector according to a word vector, and using a bag of words model or a statistical model to process the word vector to construct the sentence vector; classifying the sentence vectors in combination with location information of the sentence vectors in the text; obtaining a categories of the sentence vectors; generating structured string information according to the categories of the sentence vectors; representing the sentence vectors as nodes; determining a position of characters contained in the text information corresponding to the sentence vector as edges to construct an image network; and classifying the nodes in the image network by using an image network model, obtaining the categories of the sentence vectors. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00135"> <pat:ClaimNumber>135< / pat:ClaimNumber> <pat:ClaimText>135. The method of claim 134, wherein the method further comprises: splicing and combining the text information corresponding to the sentence vectors of the same category according to the position information, generating the structured string information. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00136"> <pat:ClaimNumber>136< / pat:ClaimNumber> <pat:ClaimText>136. The method of claim 135, wherein the method further comprises: performing a word segmentation processing on the text information. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00137"> <pat:ClaimNumber>137< / pat:ClaimNumber> <pat:ClaimText>137. The method of claim 136, wherein the method further comprises: obtaining a word segmentation. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00138"> <pat:ClaimNumber>138< / pat:ClaimNumber> <pat:ClaimText>138. The method of claim 137, wherein the method further comprises: converting the word segmentation into a word vector. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00139"> <pat:ClaimNumber>139< / pat:ClaimNumber> <pat:ClaimText>139. The method of claim 138, wherein the method further comprises: matching the word segmentation with the corresponding word vector by using word vector model. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00140"> <pat:ClaimNumber>140< / pat:ClaimNumber> <pat:ClaimText>140. The method of claim 139, wherein files include business licenses, certificates, ID cards, and receipts. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00141"> <pat:ClaimNumber>141< / pat:ClaimNumber> <pat:ClaimText>141. The method of claim 140, wherein the text information includes characters, numbers, letters, and special symbols. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00142"> <pat:ClaimNumber>142< / pat:ClaimNumber> <pat:ClaimText>142. The method of claim 141, wherein punctuation characters are used as the basis for dividing sentences in the text information and are not contained in the text information. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00143"> <pat:ClaimNumber>143< / pat:ClaimNumber> <pat:ClaimText>143. The method of claim 142, wherein using optical character recognition technology to obtain the text information in a file image and the position information of the characters in the text information in the file image. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00144"> <pat:ClaimNumber>144< / pat:ClaimNumber> <pat:ClaimText>144. The method of claim 143, wherein the method further comprises: obtaining the file image of the file and pre-processing the file image. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00145"> <pat:ClaimNumber>145< / pat:ClaimNumber> <pat:ClaimText>145. The method of claim 144, wherein the method further comprises: recognizing text direction in the file image. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00146"> <pat:ClaimNumber>146< / pat:ClaimNumber> <pat:ClaimText>146. The method of claim 145, wherein the method further comprises: a text detection and a text recognition. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00147"> <pat:ClaimNumber>147< / pat:ClaimNumber> <pat:ClaimText>147. The method of claim 146, wherein the file image is a photo of the file and a scanned copy of the file. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00148"> <pat:ClaimNumber>148< / pat:ClaimNumber> <pat:ClaimText>148. The method of claim 147, wherein the method further comprises: pre-processing the file image to correct imaging problems of the image. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00149"> <pat:ClaimNumber>149< / pat:ClaimNumber> <pat:ClaimText>149. The method of claim 148, wherein the method further comprises: geometric transformation, blurriness removal, image enhancement, and light correction. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00150"> <pat:ClaimNumber>150< / pat:ClaimNumber> <pat:ClaimText>150. The method of claim 149, wherein the text detection is performed to determine a text area in the image. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00151"> <pat:ClaimNumber>151< / pat:ClaimNumber> <pat:ClaimText>151. The method of claim 150, wherein the text recognition is performed for recognizing a character or a string located by the text detection. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00152"> <pat:ClaimNumber>152< / pat:ClaimNumber> <pat:ClaimText>152. The method of claim 151, wherein the text detection performed by text lines. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00153"> <pat:ClaimNumber>153< / pat:ClaimNumber> <pat:ClaimText>153. The method of claim 152, wherein the position information of the character is a coordinate of a character line divided by a text detection process. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00154"> <pat:ClaimNumber>154< / pat:ClaimNumber> <pat:ClaimText>154. The method of claim 153, wherein the method further comprises: constructing the sentence vectors according to the text information. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00155"> <pat:ClaimNumber>155< / pat:ClaimNumber> <pat:ClaimText>155. The method of claim 154, wherein the method further comprises: constructing a fixed-dimension sentence vector to represent test line. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00156"> <pat:ClaimNumber>156< / pat:ClaimNumber> <pat:ClaimText>156. The method of claim 155, wherein the sentence vector is vectorized representation of the character line in the text information. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00157"> <pat:ClaimNumber>157< / pat:ClaimNumber> <pat:ClaimText>157. The method of claim 156, wherein the method further comprises: performing the word segmentation processing on the text information to obtain the word segmentation. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00158"> <pat:ClaimNumber>158< / pat:ClaimNumber> <pat:ClaimText>158. The method of claim 157, wherein the word segmentation processing adopts a dictionary matching method. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00159"> <pat:ClaimNumber>159< / pat:ClaimNumber> <pat:ClaimText>159. The method of claim 158, wherein the word segmentation processing adopts a natural language model analysis method (NLP). < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00160"> <pat:ClaimNumber>160< / pat:ClaimNumber> <pat:ClaimText>160. The method of claim 159, wherein the word segmentation processing adopts the one-element model method. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00161"> <pat:ClaimNumber>161< / pat:ClaimNumber> <pat:ClaimText>161. The method of claim 160, wherein the word segmentation processing adopts the N-element model method. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00162"> <pat:ClaimNumber>162< / pat:ClaimNumber> <pat:ClaimText>162. The method of claim 161, wherein the word segmentation processing adopts the dictionary matching method. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00163"> <pat:ClaimNumber>163< / pat:ClaimNumber> <pat:ClaimText>163. The method of claim 162, wherein the word segmentation is converted into the word vector by the matching method of the word vector model, which matches the corresponding word vector with the segmentation word. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00164"> <pat:ClaimNumber>164< / pat:ClaimNumber> <pat:ClaimText>164. The method of claim 163, wherein the word vector model includes Word2Vec. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00165"> <pat:ClaimNumber>165< / pat:ClaimNumber> <pat:ClaimText>165. The method of claim 164, wherein the Word2Vec includes a large text corpus as input to generate a vector space. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00166"> <pat:ClaimNumber>166< / pat:ClaimNumber> <pat:ClaimText>166. The method of claim 165, wherein each unique word in corpus is assigned a corresponding vector in this space. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00167"> <pat:ClaimNumber>167< / pat:ClaimNumber> <pat:ClaimText>167. The method of claim 166, wherein the bag of words model and statistical model is used for processing word vector and constructing sentence vector. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00168"> <pat:ClaimNumber>168< / pat:ClaimNumber> <pat:ClaimText>168. The method of claim 167, wherein the bag of words model includes a text, and discards word order, grammar, and syntax. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00169"> <pat:ClaimNumber>169< / pat:ClaimNumber> <pat:ClaimText>169. The method of claim 168, wherein the method further comprises: determining whether the text information corresponding to different sentence vectors represents the same category of information. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00170"> <pat:ClaimNumber>170< / pat:ClaimNumber> <pat:ClaimText>170. The method of claim 169, wherein the method further comprises: determining a corresponding relationship between a type and the text information. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00171"> <pat:ClaimNumber>171< / pat:ClaimNumber> <pat:ClaimText>171. The method of claim 170, wherein the method further comprises: determining the corresponding relationship according to different categories of the sentence vectors included in the files. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00172"> <pat:ClaimNumber>172< / pat:ClaimNumber> <pat:ClaimText>172. The method of claim 171, wherein the categories of sentence vector include name, type, property, legal representative, establishment date, business period, and business scope. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00173"> <pat:ClaimNumber>173< / pat:ClaimNumber> <pat:ClaimText>173. The method of claim 172, wherein the categories of the sentence vector include ID cards, name, gender, date of birth, address, and ID number. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00174"> <pat:ClaimNumber>174< / pat:ClaimNumber> <pat:ClaimText>174. The method of claim 173, wherein the method further comprises: generating the structured string information according to the categories of the sentence vectors. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00175"> <pat:ClaimNumber>175< / pat:ClaimNumber> <pat:ClaimText>175. The method of claim 174, wherein the splicing and combination of the text information is carried in the order of coordinates, without considering a semantic. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00176"> <pat:ClaimNumber>176< / pat:ClaimNumber> <pat:ClaimText>176. The method of claim 175, wherein the method further comprises: ensuring a semantic coherence and smoothness of the text information corresponding to each sentence vector after the completion of the splicing. < / pat:ClaimText> < / pat:Claim> <pat:Claim com:id="CLM-00177"> <pat:ClaimNumber>177< / pat:ClaimNumber> <pat:ClaimText>177. The method of claim 176, wherein the structured representation of the string information mainly refers to the output of string information in a form of key-value pairs. < / pat:ClaimText> < / pat:Claim> < / pat:Claims>