Method, apparatus, device, and readable storage medium for identifying text content
By identifying quantity words and their associated related words in unstructured text, the problem of difficult to understand the actual meaning of quantity words in the prior art is solved, and more efficient information interaction and accurate quantity word recognition are achieved.
Patent Information
- Application Number
- CN202110578706.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-05-26
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-05-26
AI Technical Summary
In the prior art, extracting quantifier words in unstructured text based on rule matching cannot effectively understand the actual meaning of quantifier words, resulting in a low information interaction rate and the actual meaning of the quantifier words cannot be determined.
By obtaining the quantity words in the target text content, determining the relational words associated with the quantity words based on the context content, outputting the matching relationship between the quantity words and the relational words, using feature vectors and probability prediction to determine the starting and end characters of the relational words, identifying entities, attributes and qualifying relational words.
The information interaction rate is improved, allowing users to understand the actual meaning of quantifier words more deeply, and the recognition accuracy and recall rate of quantifier words are improved.
Smart Images

Figure CN113761126B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of computers, and particularly to a method, apparatus, device and readable storage medium for identifying text content. Background Art
[0002] In unstructured texts involved in various fields, there are some quantitative words with actual measurement values, which can provide us with some objective and important information, such as 180 cm, 90 kg, etc.
[0003] In the related art, a method based on rule matching is usually adopted to directly extract quantitative words from unstructured texts, and the method based on rule matching requires manually writing rules or regular expressions for quantitative word extraction.
[0004] However, in the actual application process, just knowing the quantitative words is far from enough to understand them. That is, obtaining a single quantitative word cannot further understand the quantitative word, the information interaction rate is low, and it is impossible to determine whether the extracted quantitative word has practical significance. Summary of the Invention
[0005] The embodiments of the present application provide a method, apparatus, device and readable storage medium for identifying text content, which improves the information interaction rate to a certain extent. The technical solutions are as follows:
[0006] On the one hand, a method for identifying text content is provided. The method includes:
[0007] Obtain target text content, where the target text content includes quantitative words;
[0008] Extract the quantitative words from the target text content;
[0009] Based on the context content of the quantitative words in the target text content, determine relationship words associated with the quantitative words from the target text content, where the relationship words are used to express the meaning of the quantitative words in the target text content;
[0010] Output the matching relationship between the quantitative words and the relationship words.
[0011] On the other hand, a device for identifying text content is provided. The device includes:
[0012] An obtaining module, configured to obtain target text content, where the target text content includes quantitative words;
[0013] An extraction module, configured to extract the quantitative words from the target text content;
[0014] A determination module, configured to determine, based on the context of the quantifier in the target text content, a relational word associated with the quantifier from the target text content, where the relational word is used to express the meaning of the quantifier in the target text content;
[0015] An output module, configured to output the matching relationship between the quantifier and the relational word.
[0016] In an optional embodiment, the determination module is further configured to encode the quantifier and the target text content to obtain a first feature vector corresponding to n characters in the target text content, where n is a positive integer; perform a relational probability prediction on the first feature vector of the i-th character to obtain a probability value that the i-th character belongs to the relational word, where 1 ≤ i ≤ n; and determine the relational word from the target text content based on the probability value.
[0017] In an optional embodiment, the probability value includes a first probability value and a second probability value, where the first probability value is used to represent the probability that the i-th character is the starting character of the relational word, and the second probability value is used to represent the probability that the i-th character is the ending character of the relational word;
[0018] The determination module is further configured to determine the first starting character of the relational word from the n characters based on the first probability value corresponding to the n characters; determine the first ending character of the relational word from the n characters based on the second probability value corresponding to the n characters; and obtain the characters from the start of the first starting character to the end of the first ending character as the relational word.
[0019] In an optional embodiment, the relational word includes at least one of an attribute relational word, an entity relational word, and a limiting relational word;
[0020] In response to the relational word including an attribute relational word, the first starting character includes an attribute starting character, and the first ending character includes an attribute ending character;
[0021] In response to the relational word including an entity relational word, the first starting character includes an entity starting character, and the first ending character includes an entity ending character;
[0022] In response to the relational word including a limiting relational word, the first starting character includes a limiting starting character, and the first ending character includes a limiting ending character.
[0023] In an alternative embodiment, the extraction module is further configured to encode the target text content to obtain second feature vectors corresponding to n characters in the target text content; predict the probability of the quantifier for the second feature vector corresponding to the i-th character to obtain the probability value that the i-th character belongs to the quantifier; and extract the quantifier from the target text content based on the probability values that the n characters belong to the quantifier.
[0024] In an alternative embodiment, the probability value that the i-th character belongs to the quantifier includes a third probability value and a fourth probability value. The third probability value represents the probability that the i-th character is the start character of the quantifier, and the fourth probability value is used to represent the probability that the i-th character is the end character of the quantifier.
[0025] The extraction module is further configured to determine a second start character of the quantifier from the n characters based on the third probability values corresponding to the n characters; determine a second end character of the quantifier from the n characters based on the fourth probability values corresponding to the n characters; and obtain the characters from the start of the second start character to the end of the second end character as the relationship word.
[0026] In an alternative embodiment, the apparatus further includes:
[0027] A classification module, configured to perform classification prediction on the quantifier to obtain a classification result corresponding to the quantifier. The classification result is used to represent the measurement value of the quantifier, and the classification result includes any one of a quantitative value, a range value, and an approximate value.
[0028] In an alternative embodiment, the output module is further configured to, in response to the last character of the quantifier appearing in a preset unit set, output the longest unit corresponding to the last character in the preset unit set, and determine the longest unit as the unit corresponding to the quantifier; or,
[0029] The output module is further configured to, in response to the last character not appearing in the preset unit set, perform a forward traversal operation on the quantifier, output the character after the first non-alphabetic character in the traversal process, and use the output character as the unit of the quantifier.
[0030] On the other hand, a computer device is provided. The computer device includes a processor and a memory. At least one instruction, at least one program, a code set, or an instruction set is stored in the memory. The at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the text content recognition method according to any one of the foregoing embodiments of the present application.
[0031] On the other hand, a computer-readable storage medium is provided, in which at least one instruction, at least one program, a code set or an instruction set is stored, and the at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by a processor to implement the text content recognition method as described in any one of the embodiments of the present application above.
[0032] On the other hand, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the text content recognition method as described in any one of the above embodiments.
[0033] The beneficial effects brought by the technical solutions provided in the embodiments of the present application at least include:
[0034] By determining the quantifiers in the target text content, combining the context content of the quantifiers, determining the relational words used to describe or limit the quantifiers from the target text content, and matching the relationship between the extracted relational words and the corresponding quantifiers, the information interaction rate is greatly improved, enabling users to have a further in-depth understanding of the quantifiers in combination with the relational words, and improving the efficiency of understanding the actual meaning expressed by the quantifiers. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0036] Figure 1 is a schematic diagram of an implementation environment provided by an exemplary embodiment of the present application;
[0037] Figure 2 is a flowchart of a text content recognition method provided by an exemplary embodiment of the present application;
[0038] Figure 3 is a flowchart of a method for extracting quantifiers provided by another exemplary embodiment of the present application;
[0039] Figure 4 is a flowchart of classifying quantifiers provided by another exemplary embodiment of the present application;
[0040] Figure 5 is a flowchart of identifying quantifier units provided by another exemplary embodiment of the present application;
[0041] Figure 6 It is a schematic diagram of the method for extracting quantifiers and their relationships provided by another exemplary embodiment of the present application;
[0042] Figure 7 It is a structural block diagram of the text content recognition device provided by an exemplary embodiment of the present application;
[0043] Figure 8 It is a structural block diagram of the text content recognition device provided by another exemplary embodiment of the present application;
[0044] Figure 9 It is a structural block diagram of the server provided by an exemplary embodiment of the present application. Detailed implementation manners
[0045] To make the objectives, technical solutions and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0046] First, a brief introduction to the application scenarios of the embodiments provided by the present application:
[0047] First, the present application can be applied to the medical scenario, specifically for the automated analysis of medical record documents / detection reports, automatically extracting the quantifiers corresponding to various detection indicators of patients and their corresponding units from the medical record documents / detection reports, and at the same time providing key information such as entity relationship words, attribute relationship words, and qualification relationship words for understanding the quantifiers. The embodiments of the present application can also convert the input unstructured text (medical record documents / detection reports) into structured table information to provide corresponding support for downstream medical record analysis work. Schematically, a certain hospital measures the pulse of inpatients and conducts statistics on the measurement results. The method provided by the present application can be used to extract the quantifiers and relationship words corresponding to the pulse indicators in the detection reports of each patient, and obtain the summary information shown in Table 1 below.
[0048] Table 1: Pulse summary of inpatients in a certain hospital
[0049] Name Test Items Test Results Zhang San Pulse 66 beats / min Li Si Pulse 76 beats / min Wang Wu Pulse 80 beats / min Wu Hai Pulse 70 beats / min
[0050] Second, it can be applied to the paper analysis scenario, extracting key indicator data from a large number of papers, literature and other materials, automatically analyzing the key indicators, being able to quickly determine data differences, and determining the subsequent research / experimental directions based on the differences.
[0051] Third, it can be applied to the scenario of report data analysis, extracting the experimental data in the report and comparing it with the established standard data and exclusion criteria to automatically determine whether the important data in the report meets the standards.
[0052] The above scenarios are only exemplary and can also be applied to other scenarios for analyzing text content. This application does not limit this.
[0053] Secondly, a brief introduction to the nouns involved in the embodiments of this application is given:
[0054] Artificial Intelligence (AI) uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, a theory, method, technology, and application system that perceives the environment, acquires knowledge, and uses knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.
[0055] Artificial intelligence technology is an interdisciplinary subject involving a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0056] The key technologies of speech technology include automatic speech recognition technology (Automatic Speech Recognition, abbreviated as ASR), text-to-speech technology (Text-To-Speech, abbreviated as TTS), and voiceprint recognition technology. Enabling the computer to listen, see, speak, and feel is the future development direction of human-computer interaction, and among them, speech has become one of the most promising human-computer interaction methods in the future.
[0057] Machine Learning (ML) is an interdisciplinary subject involving multiple fields such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration.
[0058] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, robots, smart healthcare, smart customer service, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0059] Unstructured text is the simplest form of file organization. Unstructured text organizes data into records in sequence and accumulates them for storage. It is a collection of ordered information and is measured in bytes. Since unstructured text has no specific structure, access to records can only be achieved through exhaustive search. In the embodiments of this application, the main focus is on identifying the quantifiers in unstructured text and the relational words corresponding to the quantifiers. Unstructured text can be medical record documents, experimental data, scientific research articles, etc., and this application does not impose any restrictions on it.
[0060] A quantifier is a quantifier with practical meaning that appears in unstructured text such as medical and scientific articles, for example, 180 cm, 90 kg. In the embodiments of this application, this quantifier is generally a character with unit description.
[0061] Relational words are used to express the actual meaning of quantifiers in unstructured text in combination with context information, including entity relational words, attribute relational words, and restrictive relational words. Among them, entity relational words refer to the entities described by quantifiers, which can be people, objects, etc.; attribute relational words refer to the attributes described by quantifiers, which can be height, weight, volume, etc.; restrictive relational words are expressions that help in understanding the quantifier in the context where the quantifier and the described entity appear. For example, for the average height of 180 cm, the quantifier is 180 cm, and its restrictive relational word is "average". It can be seen that the restrictive relational word is a way of restricting the quantifier, and the restrictive relational word is an important context information for understanding the quantifier.
[0062] In the embodiments of this application, the main task is to identify the quantifiers contained in unstructured text and the relational words corresponding to the quantifiers; this application can also be used to determine the relationship between the quantifier and the relational word, and is actually implemented as a relation extraction model. This relation extraction model includes an entity extraction model, an attribute extraction model, and a restrictive expression extraction model.
[0063] The entity extraction model is used to represent the relationship between the quantifier and the described entity / described attribute (entity relational word / attribute relational word), and can be implemented to output in the form of a triple. For example, (entity relational word a, entity extraction, quantifier a) is used to represent the specific meaning that the entity a is described by the quantifier a.
[0064] The attribute extraction model is used to represent that the entity being described has a certain attribute, and can be implemented to output in the form of a triple. For example, the specific meaning represented by (entity relation word b, attribute extraction, attribute relation word b) is that entity b has attribute b.
[0065] The restricted expression extraction model is used to represent the restriction relationship between a quantifier and a restrictive relation word, and can be implemented to output in the form of a triple. For example, the specific meaning represented by (restrictive relation word c, restricted expression extraction, quantifier c) is that restrictive relation word c restricts quantifier c.
[0066] Finally, the implementation environment provided by the embodiments of this application is described in combination with the above application scenarios and noun introductions.
[0067] Figure 1 is a schematic diagram of the implementation environment provided by an exemplary embodiment of this application. As Figure 1 shown, this implementation environment includes a terminal 110 and a server 120, and the terminal 110 and the server 120 are connected through a communication network 130;
[0068] An application program for identifying text content or a web page corresponding to the identified text is installed in the terminal 110. Optionally, after the user determines the target text content in the terminal 110, the target text content is selected as the content of the quantifier to be identified. The target text content can be online text, local text, audio content, etc. In some embodiments, the terminal 110 uploads the target text content to the server 120 through the communication network 130, or when the target text content is online text, the server 120 directly obtains the target text content from the server where the online text is located, or when the target text content is audio content, the terminal 110 converts the audio content into text content using ASR technology and uploads the converted text content to the server 120.
[0069] The server 120 obtains the target text content, determines the corresponding quantifier from the target text content, and determines the relation word matching the quantifier from the target text content based on the quantifier. After the server 120 extracts the quantifier and the relation word in the target text content, it feeds back the extracted relation word to the terminal 110 through the communication network 130, or feeds back the matching relationship between the extracted quantifier and the relation word to the terminal 110 through the communication network 130. In some embodiments, the server 120 can feed back the matching relationship between the quantifier and the relation word to the terminal 110 in the form of a table, a tuple, etc. The embodiments of this application do not limit the form of feedback.
[0070] It should be noted that the above-mentioned terminal 110 can be implemented as a mobile terminal such as a mobile phone, a tablet computer, a wearable device, a portable laptop computer, etc., or can be implemented as a terminal such as a desktop computer. The embodiments of the present application do not limit this.
[0071] The above-mentioned server 120 can be implemented as a single server or a server cluster composed of multiple servers. The above-mentioned server 120 can be implemented as a physical server or a cloud server. The embodiments of the present application do not limit this.
[0072] Combined with the above implementation environment, the text content recognition method involved in the embodiments of the present application will be described. Figure 2 It is a flowchart of the text content recognition method provided by an exemplary embodiment of the present application, which is described by applying this method to a server. As Figure 2 shown, this method includes:
[0073] Step 201, obtain the target text content.
[0074] In some embodiments, the target text content includes, but is not limited to, medical record documents, clinical trial data, test reports, maintenance logs, customer service records, etc.; if the initial raw data is voice data or video data, optionally, the server uses speech recognition technology to recognize the raw data to obtain the target text content.
[0075] In some embodiments, the target text content includes a quantifier, which is used to represent characters with a certain measurement value. For example, the target text content A is "Xiaoming's heart rate is 90 beats / min", where the quantifier is "90 beats / min", and the target text content B is "Xiaohong's height is 180 cm", where the quantifier is "180 cm". In actual application processes, it is necessary to further study these quantifiers to determine the description object corresponding to the quantifier. The description object includes, but is not limited to, the described entity, the described attribute, the limiting expression, etc., and determine the actual meaning corresponding to the quantifier.
[0076] Optionally, the above-mentioned acquisition method of the target text content includes at least one of the following methods:
[0077] First, the server receives the target text content uploaded by the terminal.
[0078] In some embodiments, the terminal uploads the target text content to be recognized to the server, and the server receives the target text content and determines the quantifier and the corresponding relational word in the target text content. Optionally, when the terminal uploads the target text content to the server, it can also send a recognition requirement for recognizing the target text content to the server. The recognition requirement includes at least one of recognizing the quantifier, recognizing the relational word, and recognizing the matching relationship between the quantifier and the relational word:
[0079] Second, the server receives the text content link sent by the terminal, and obtains the target text content from other servers based on the text content link.
[0080] In some embodiments, when the server receives the text content link sent by the terminal, it obtains the target text content from the server corresponding to the link URL based on the text content link. Optionally, when the terminal sends the text content link to the server, it can also indicate to the server the recognition requirements for the target text content.
[0081] Third, the server receives the multimedia file uploaded by the terminal, and the multimedia file includes at least one of pictures, audio content, and video content.
[0082] Optionally, the server receives the picture uploaded by the terminal, and the picture contains text information. The server uses optical character recognition technology (OCR for short) to convert the text content in the picture into the recognized text content, and determines the recognized text content as the target file content. The server performs corresponding recognition operations on the target file content to determine the quantifiers and the relational words corresponding to the quantifiers in the target file content.
[0083] Optionally, the server receives the audio content / video content uploaded by the terminal, and the audio content / video content can be at least one of call records, video records, etc. The server uses speech recognition technology to convert the audio content / video content into text content. Optionally, when the terminal uploads the audio content / video content, it can upload the subtitle file corresponding to the audio content / video content to the server, and the server directly uses the subtitle file as the target text content to recognize the quantifiers and the relational words corresponding to the quantifiers.
[0084] Fourth, when the execution subject is implemented as a terminal, the terminal can obtain the target text content from the local database; or, the terminal downloads the target text content from the server.
[0085] It should be noted that the above methods for obtaining the target text content are only illustrative examples, and the embodiments of the present application are not limited thereto.
[0086] Step 202, extract quantifiers from the target text content.
[0087] Quantifiers refer to the quantifiers with actual meanings that appear in the unstructured target text content, such as 180 cm and 90 kg. In the embodiments of the present application, the quantifier is a character with unit description, and the unstructured text can be at least one of medical, scientific articles, customer service records, etc.
[0088] Step 203: Based on the context of the quantifier in the target text content, determine a relational word associated with the quantifier from the target text content. The relational word is used to express the meaning of the quantifier in the target text content.
[0089] The relational word is used to express the actual meaning of the quantifier in the target text content (unstructured text) in combination with the context information, including entity relational words, attribute relational words, and restrictive relational words. Among them, the entity relational word refers to the entity described by the quantifier, which can be a person, a thing, etc.; the attribute relational word refers to the attribute described by the quantifier, which can be height, weight, capacity, etc.; the restrictive relational word is an expression that helps the understanding of the quantifier in the context where the quantifier and the described entity appear. Exemplarily, taking the target text content "The average height of the students in class A is 160 cm" as an example, where the quantifier is "160 cm", the entity relational word is "student" / "students in class A", the attribute relational word is "height", and the restrictive relational word is "average". It can be seen that if only the quantifier "160 cm" is known, it is impossible to comprehensively understand the actual meaning described by the quantifier, and the relational word is a restrictive expression of the quantifier. That is to say, the relational word is an important context information for understanding the quantifier.
[0090] The server extracts the quantifier in the target text content, encodes the quantifier and the target text content, obtains the first feature vector corresponding to n characters in the target text content, where n is a positive integer, and performs a relational probability prediction on the first feature vector of the i-th character to obtain the probability value that the i-th character belongs to the relational word, where 1 ≤ i ≤ n.
[0091] In some embodiments, the probability value includes a first probability value and a second probability value. The first probability value is used to represent the probability that the i-th character is the starting character of the relational word, and the second probability value is used to represent the probability that the i-th character is the ending character of the relational word.
[0092] In some embodiments, based on the first probability value corresponding to n characters, determine the first starting character of the relational word from the n characters; based on the second probability value corresponding to n characters, determine the first ending character of the relational word from the n characters. Optionally, take the character with the largest first probability value among the n characters as the first starting character of the relational word, and take the character with the largest second probability value among the n characters as the first ending character of the relational word; or, take the characters with the first probability value greater than a certain preset threshold among the n characters as the first starting characters of the relational word, and take the characters with the second probability greater than a certain preset threshold among the n characters as the first ending characters of the relational word. The preset threshold can be set by the programmer or automatically adjusted based on the relational word probability prediction process. For the above relational word probability prediction formula, please refer to Formula 1 and Formula 2.
[0093] Formula 1:
[0094] Formula 2:
[0095] In Formula 1 and Formula 2, is used to represent the probability value that the i-th character is the first starting character of a relational word, is used to represent the probability value that the i-th character is the first ending character of a relational word, is used to represent the input quantifier, h i is used to represent the vector representation corresponding to the i-th character, and σ is used to represent the sigmoid function, and is used to represent the weight parameter corresponding to the r-th relational word, and is used to represent the bias parameter corresponding to the r-th relational word, 1 ≤ i ≤ n.
[0096] In the embodiments of the present application, r takes the value of 3. When r takes the value of 1, is used to represent the weight value corresponding to the first starting character of the i-th character as an entity relational word, is used to represent the weight value corresponding to the first ending character of the i-th character as an entity relational word, is used to represent the bias value corresponding to the first starting character of the i-th character as an entity relational word, is used to represent the bias value corresponding to the first ending character of the i-th character as an entity relational word.
[0097] When r takes the value of 2, is used to represent the weight value corresponding to the first starting character of the i-th character as an attribute relational word, is used to represent the weight value corresponding to the first ending character of the i-th character as an attribute relational word, is used to represent the bias value corresponding to the first starting character of the i-th character as an attribute relational word, is used to represent the bias value corresponding to the first ending character of the i-th character as an attribute relational word.
[0098] When r takes the value of 3, is used to represent the weight value corresponding to the first starting character of the i-th character as a qualifying relational word, is used to represent the weight value corresponding to the first ending character of the i-th character as a qualifying relational word, is used to represent the bias value corresponding to the first starting character of the i-th character as a qualifying relational word, is used to represent the bias value corresponding to the first ending character of the i-th character as a qualifying relational word.
[0099] In some embodiments, and it can be automatically adjusted during the relative word probability prediction process.
[0100] Exemplarily, the weight parameter is described in combination with a specific scenario. Taking the sentence A "There are about 40 students in classroom 502" as an example, the server extracts the quantifiers "502" and "40" in the sentence A, and performs relative word probability prediction on the quantifiers "502" and "40" respectively. During the prediction process, the weight parameter of the quantifier "502" is less than the weight parameter of the quantifier "40". Finally, the quantifier "40" is studied intensively to provide data support for downstream data aggregation or data prediction.
[0101] Optionally, the characters from the first start character to the first end character are obtained as the relative word corresponding to the quantifier; or, the first start character and the first end character of the relative word corresponding to the quantifier are determined, and the characters corresponding to the first start character and the first end character are marked in the first feature vector. The server performs one-to-one matching on the first start character and the first end character to obtain the coordinate pair corresponding to the relative word, and locates the relative word corresponding to the quantifier from the target text content based on the coordinate pair.
[0102] In some embodiments, the relative word includes at least one of an attribute relative word, an entity relative word, and a restrictive relative word. The determination methods of the three relative words will be further introduced below.
[0103] In response to the relative word including an attribute relative word, the first start character includes an attribute start character, and the first end character includes an attribute end character. Optionally, the characters from the attribute start character to the attribute end character are used as the attribute relative word; or, for the first feature vector obtained by encoding the target text content by the server, the vector corresponding to the character with the largest first probability value is marked as the attribute start character, and the vector corresponding to the character with the second largest probability value is marked as the attribute end character. The server matches the attribute start character and the attribute end character in the first feature vector marked with the attribute relative word to obtain the attribute relative word coordinate pair, and the attribute relative word coordinate pair is used to locate the attribute relative word corresponding to the quantifier from the target text content.
[0104] In response to the relational words including entity relational words, the first starting character includes an entity starting character, and the first ending character includes an entity ending character. Optionally, the characters from the entity starting character to the entity ending character are used as the entity relational word; alternatively, for the first feature vector obtained by the server encoding the target text content, the vector corresponding to the character with the largest first probability value is marked as the entity starting character, and the vector corresponding to the character with the second largest probability value is marked as the entity ending character. The server matches the entity starting character and the entity ending character in the first feature vector marked with the entity relational word to obtain an entity relational word coordinate pair, and this entity relational word coordinate pair is used to locate the entity relational word corresponding to the quantifier in the target text content.
[0105] In response to the relational words including qualification relational words, the first starting character includes a qualification starting character, and the first ending character includes a qualification ending character. Optionally, the characters from the qualification starting character to the qualification ending character are used as the qualification relational word; alternatively, for the first feature vector obtained by the server encoding the target text content, the vector corresponding to the character with the largest first probability value is marked as the qualification starting character, and the vector corresponding to the character with the second largest probability value is marked as the qualification ending character. The server matches the qualification starting character and the qualification ending character in the first feature vector marked with the qualification relational word to obtain a qualification relational word coordinate pair, and this qualification relational word coordinate pair is used to locate the qualification relational word corresponding to the quantifier in the target text content.
[0106] Illustratively, taking the target text content "Zhang Xiaohong's height is 160 cm" as an example, the quantifier "160 cm" in the target text content is extracted, and then the quantifiers related to this quantifier "160 cm" are extracted. When identifying the entity relational word corresponding to this quantifier, the character "Zhang" is used as the entity starting character, "Hong" is used as the entity ending character, and all the characters "Zhang Xiaohong" from the character "Zhang" to the character "Hong" are used as the entity relational word of the quantifier "160 cm"; when identifying the attribute relational word corresponding to this quantifier, the character "Shen" is used as the attribute starting character, the character "Gao" is used as the attribute ending character, and all the characters "Shen Gao" from the character "Shen" to the character "Gao" are used as the attribute relational word of the quantifier "160 cm".
[0107] In the embodiments of the present application, the process of extracting the quantifier and the relational word corresponding to the quantifier can be implemented as a relation extraction model. This relation extraction model includes an entity relational word tagger, an attribute relational word tagger, a qualification relation tagger, etc. Each relation word tagger contains a start tagging sequence and an end tagging sequence. The relation words described by the start tagging sequence and the end tagging sequence are matched to obtain the relational word corresponding to the quantifier.
[0108] Step 204, output the matching relationship between the quantifier and the relational word.
[0109] The server determines the relational word corresponding to the quantifier from the encoded first feature vector, and the relational word includes but is not limited to entity relational words, attribute relational words, and restrictive relational words.
[0110] The server feeds back the relationship between the quantifier and the relational word to the terminal in the form of a table or in the form of a tuple.
[0111] In some embodiments, the server analyzes and processes the matching relationships among the determined multiple relational words, and outputs and feeds back the analysis results to the terminal. Taking the relational words "weight", "average", and "Xiaohong" and the quantifier "50 kg" as an example, the server outputs the matching relationships between the relational words and the quantifier in the form of binary tuples (weight, 50 kg), (average, 50 kg), and (Xiaohong, 50 kg). The server can also judge the relationships among "weight", "average", and "Xiaohong", and output the judgment results in the form of a ternary tuple (Xiaohong, attribute extraction, weight), which is used to indicate that the quantifier describes Xiaohong's weight.
[0112] In some embodiments, the server performs an average calculation on the quantifiers with the same extracted relational words, and uses the average calculation result as a prediction value. For example, for the attribute relational word "weight", the non-restrictive relational word, and the quantifiers of the entity relational word "person", a frequency statistics is performed. When the frequency is greater than a certain preset value, an average calculation is performed on all the quantifiers related to the "weight" of "person" in the statistics, and the average value prediction value is used.
[0113] In summary, the text content recognition method provided by the embodiments of the present application determines the quantifier in the target text content, combines the context content of the quantifier, determines the relational word used to describe or restrict the quantifier from the target text content, and matches the relationship between the extracted relational word and the corresponding quantifier of the quantifier, which greatly improves the information interaction rate, enables the user to have a further in-depth understanding of the quantifier in combination with the relational word, and improves the efficiency of understanding the actual meaning expressed by the quantifier.
[0114] In an alternative embodiment, the method provided by the embodiments of the present application can also be applied to the extraction process of quantifiers. For details, please refer to Figure 3 , Figure 3 is a flowchart of a method for extracting quantifiers provided by another exemplary embodiment of the present application. Taking the application of this method in a server as an example, as Figure 3 shown, the method includes:
[0115] Step 301, encode the target text content.
[0116] The server receives the target text content, encodes the target text content, and obtains second feature vectors corresponding to n characters in the target text content, where n is a positive integer, and labels the n characters in the second feature vectors by means of pointer decoding. In some embodiments, a CRF decoding method or a decoding method of token classification may also be used to extract the quantifiers in the target text content.
[0117] Step 302: Perform a quantifier probability prediction on the second feature vector corresponding to the i-th character to obtain the probability value that the i-th character belongs to a quantifier.
[0118] The server needs to perform a quantifier probability prediction on each character in the second feature vector, determine the characters greater than a certain threshold as quantifiers, or label the character with the largest probability value among the probability values as a quantifier.
[0119] In some embodiments, the probability value includes a third probability value and a fourth probability value. The third probability value is used to represent the probability that the i-th character is the start character of a quantifier, and the fourth probability value is used to represent the probability that the i-th character is the end character of a quantifier.
[0120] Step 303: Extract the quantifiers from the target text content based on the probability values that the n characters belong to quantifiers.
[0121] In some embodiments, based on the third probability values corresponding to the n characters, determine the second start character of the quantifier from the n characters; based on the fourth probability values corresponding to the n characters, determine the second end character of the relational word from the n characters. Optionally, use the character with the largest third probability value among the n characters as the first start character of the quantifier, and use the character with the largest fourth probability value among the n characters as the second end character of the quantifier; or, use the characters with third probability values greater than a certain preset threshold among the n characters as the second start characters of the quantifier, and use the characters with fourth probabilities greater than a certain preset threshold among the n characters as the second end characters of the quantifier. The preset threshold can be set by the programmer or automatically adjusted based on the relational word probability prediction process.
[0122] In some embodiments, the method for extracting quantifiers from the target text content includes the following methods:
[0123] First, perform a quantifier probability prediction on the encoded target text content.
[0124] The server performs pointer encoding on the n characters of the target text content, where n is a positive integer, to obtain the second feature vector corresponding to the target text content, and performs a quantifier probability prediction on the i-th character in the second feature vector, 1 ≤ i ≤ n. The specific quantifier probability prediction process will be described in detail later.
[0125] Second, use natural language processing technology to process the target text content.
[0126] The server uses natural language processing technology to perform sentence segmentation on the target text content. Among them, natural language processing technology includes but is not limited to Conditional Random Field (CRF for short) and sentence segmentation methods based on neural networks. Then, character recognition is performed on the segmented sentences, and the characters that meet the character category are determined as quantifiers.
[0127] In the embodiments of the present application, the quantifiers are mainly extracted by predicting the probability of quantifiers for the encoded target text content. For the specific process, please refer to the following expressions.
[0128] In some embodiments, an input target text content S = {w1, w2,..., w n}, where wi represents the i-th character in the target text content. The server encodes the sentence to obtain a second feature vector R = {r1, r2,..., r n}, where r i is used to represent the encoded vector representation of w i ; calculate the probability distribution for each character in the second feature vector as the second start character to obtain P start = {p1, p2,..., p n}, where P i is used to represent the probability value of w i as the second start character, and the character with the largest probability value is used as the second start character of the quantifier; calculate the probability distribution for each character in the second feature vector as the second end character, P end = {p1, p2,..., p n}, where P i is used to represent the probability value of w i as the second end character, and the character with the largest probability value is used as the second end character of the quantifier; match the second start character and the second end character to obtain the coordinate pair (start, end) of the quantifier, and locate the quantifier from the target text content based on this coordinate pair, where start is used to represent the position of the second start character in the target text content, and end is used to represent the position of the second end character in the target text content.
[0129] In some embodiments, in order to improve the accuracy and recall rate of quantifier recognition, the quantifiers can also be recognized by using multi-fold voting fusion and multi-model union fusion. The following specifically describes multi-fold voting fusion and multi-model union fusion.
[0130] Multi-fold voting fusion means splitting the training set into k parts. Each time, k - 1 parts of the data are used as the training set, and the remaining one part is used as the validation set to complete the training of the model. That is, if there are k parts of data in total, k trainings can be performed, and k models can be obtained by training with different data each time. Each model can be tested on the test set to obtain a test result, and a set of test results is obtained. Count the number of times each quantifier appears in the result set. Retain the quantifiers for which the number of times n that each quantifier appears is greater than a certain threshold to complete the multi-fold voting fusion. Schematically, taking ten-fold voting fusion as an example, there are 100 pieces of data. The first 90 pieces of data are used as the training set (there are 9 quantifier recognition models corresponding to this training set), and the last 10 pieces of data are used as the test set. When there are differences in the obtained quantifiers, the final quantifier can be determined by voting.
[0131] The multi-model union fusion method uses multiple quantifier recognition models to identify the quantifiers in the target text content, and combines the results of multiple quantifier recognition models to determine the final quantifier.
[0132] In summary, the text content recognition method provided in the embodiments of the present application uses the probability prediction method to determine the second start character and the second end character corresponding to the target text content, and determines the characters between the second start character and the second end character as quantifiers. In order to improve the recognition accuracy and recall rate of quantifiers, different decoding methods can also be used to decode to obtain different results, and the results are fused to determine the final quantifier.
[0133] In another exemplary embodiment, the method provided in the embodiments of the present application can also be applied to the classification prediction process of quantifiers. For details, please refer to Figure 4 , Figure 4 which is the flowchart of classifying quantifiers provided in another exemplary embodiment of the present application. Taking the application of this method in a server as an example, as Figure 4 shown, the method includes:
[0134] Step 401, identify the quantifiers from the target text content.
[0135] The server receives the target text content, encodes the target text content, and obtains the second feature vectors corresponding to n characters in the target text content. n is a positive integer, and the n characters in the second feature vectors are labeled by using the pointer decoding method.
[0136] The server performs quantifier probability prediction on each character in the second feature vectors, determines the characters greater than a certain threshold as quantifiers, or labels the character with the largest probability value among the probability values as a quantifier.
[0137] This step is the same as the process from step 301 to step 302, and will not be elaborated here.
[0138] Step 402: Determine the qualifying relational words corresponding to the quantifiers from the target text content.
[0139] For the first feature vector obtained by the server encoding the target text content, the vector corresponding to the character with the largest first probability value is marked as the qualifying start character, and the vector corresponding to the character with the largest second probability value is marked as the qualifying end character. The server matches the qualifying start character and the qualifying end character in the first feature vector marked with the qualifying relational words to obtain a qualifying relational word coordinate pair, which is used to locate the qualifying relational word corresponding to the quantifier from the target text content.
[0140] This step is the same as the process of step 203, and will not be elaborated here.
[0141] Step 403: Classify and predict the quantifiers to obtain a classification result.
[0142] In some embodiments, the target text content includes quantifiers with qualifying expressions, that is, there are corresponding qualifying relational words for the quantifiers. The server needs to classify and predict the quantifiers with qualifying expressions and obtain the classification result corresponding to the quantifiers. The classification result represents the measurement value of the quantifiers, where the classification result includes any one of a quantitative value, a range value, and an approximate value. Exemplarily, taking the text a "greater than 40 kg" as an example, classifying the quantifier "40 kg" obtains that the quantifier belongs to the range value; taking the text b "about 40 kg" as an example, classifying the quantifier "40 kg" obtains that the quantifier belongs to the approximate value; taking the text c "40 kg" as an example, classifying the quantifier "40 kg" obtains that the quantifier belongs to the quantitative value.
[0143] In an exemplary embodiment, the identified quantifiers are encoded to obtain a representation vector c corresponding to the quantifiers i , and the representation vector c i is used as a feature and put into a multi-classifier for classification to obtain the type y corresponding to the quantifiers i , specifically, see formulas 3 to 4 below.
[0144] Formula 3: c i = BERT([x i-n ; …; x n ; …; x i+n )
[0145] Formula 4:
[0146] Among them, in Formulas 3 to 4, xi is used to represent the i-th character in the quantifier, n is used to represent the number of characters in the quantifier, and 1 ≤ i ≤ n. Optionally, the classification process of the quantities involved in the above Formulas 3 to 4 can be trained into a quantifier classification model.
[0147] In summary, the method for identifying text content provided by the embodiments of the present application uses probability prediction to determine the quantifiers in the target text content, encodes the quantifiers, and puts them into a pre-set classifier for classification to obtain the measurement range to which the quantifiers belong, providing certain data support for subsequent research on quantifiers, and at the same time helping to understand the limiting meaning expressed by the quantifiers.
[0148] In an optional embodiment, the method provided by the embodiments of the present application can also be applied to the process of identifying the units of quantifiers. For details, please refer to Figure 5 , Figure 5 which is a flowchart for identifying the units of quantifiers provided by another exemplary embodiment of the present application. Taking the application of this method in a server as an example for illustration, as Figure 5 shown, the method includes:
[0149] Step 501, identify the quantifiers from the target text content.
[0150] In some embodiments, the target text content includes but is not limited to medical record documents, clinical trial data, test reports, maintenance logs, customer service records, etc.; if the initial raw data is voice data or video data, optionally, the server uses speech recognition technology to identify the raw data to obtain the target text content.
[0151] In some embodiments, the target text content includes quantifiers, and the quantifiers are used to represent characters with certain measurement values. For example, for the target text content A "Xiaoming's heart rate is 90 beats / min", the quantifier is "90 beats / min", and for the target text content B "Xiaohong's height is 180 cm", the quantifier is "180 cm". In actual application processes, it is necessary to further study these quantifiers to determine the description objects corresponding to the quantifiers. The description objects include but are not limited to the described entities, described attributes, and limiting expressions, etc., to determine the actual meaning corresponding to the quantifier.
[0152] This step has the same process as Steps 201 to 202, and will not be elaborated here.
[0153] Step 502, determine whether the last character of the quantifier is in the preset unit set.
[0154] The preset unit set is the units that have appeared in the dataset during the training phase to obtain the preset unit set.
[0155] In the embodiments of the present application, the quantifier includes a unit description. When the server identifies the unit in the quantifier, it mainly determines the last character of the quantifier, and the last character is the string after the first non-alphabetic character during forward traversal.
[0156] If the last character is in the preset unit set, step 503 is executed; if the last character is not in the preset unit set, step 504 is executed.
[0157] Step 503: Output the longest unit corresponding to the last character.
[0158] In response to the last character of the quantifier appearing in the preset unit set, output the longest unit in the preset unit set corresponding to the last character, and determine the longest unit as the unit corresponding to the quantifier.
[0159] Step 504: Traverse the quantifier to determine the unit corresponding to the quantifier.
[0160] In response to the last character not appearing in the preset unit set, perform a forward traversal operation on the quantifier, output the character after the first non-alphabetic character during the traversal process, and use the output character as the unit of the quantifier.
[0161] The process of identifying the quantifier unit involved in the embodiments of the present application can be implemented as the following code:
[0162]
[0163]
[0164] The algorithm description of the above code is as follows:
[0165] 1: Perform statistical analysis on the existing data set and record all existing unit sets V.
[0166] 2 - 6: If the end of the input quantity string is in V, select the unit with the longest length as the output.
[0167] 7 - 10: Record the position of the last space at the same time.
[0168] 11 - 14: Perform forward traversal to find the starting position of the possible unit until a non-alphabetic character is encountered or the first digit at the beginning of the quantity is traversed.
[0169] 15 - 17: If an illegal position is traversed, the search fails.
[0170] 18: Return the finally determined unit.
[0171] In summary, for the method provided in the embodiments of the present application, after extracting the quantifiers from the target text content, the units of the quantifiers are identified, and the unit descriptions related to understanding the quantifiers are extracted, which is convenient for determining the research value and importance of the quantifiers in the subsequent quantifier analysis process.
[0172] Combined with the above embodiments, another exemplary embodiment of the present application will be described. For details, please refer to the following Figure 6 , Figure 6 FIG. is a schematic diagram of a method for extracting quantifiers and their relationships based on a deep learning model provided by another exemplary embodiment of the present application. Taking the deep learning model stored in the server as an example for illustration.
[0173] In the embodiments of the present application, the deep learning model includes five models: an encoding model 601, a quantifier extraction model 602, a relationship extraction model 603, a classification model 604, and a unit recognition model 605. The implementation principles of the five modules will be described in detail below.
[0174] The encoding model 601 is used to encode the input target text content to obtain a feature vector corresponding to the target text content. In the subsequent quantifier extraction and relationship word recognition processes, the feature vector encoded by the encoder can be directly used for the quantifier probability prediction and relationship word probability prediction processes.
[0175] The quantifier extraction model 602 is used to identify quantifiers, perform quantifier probability prediction on each character in the feature vector after encoding processing, and determine which character is the starting character corresponding to the quantifier and which character is the ending character corresponding to the quantifier. The judgment rule can be to use the character with a probability value greater than a certain threshold as the starting character / ending character, or to use the character with the largest probability value among the probability values of n characters in the feature vector as the starting character / ending character. And mark the characters corresponding to the starting character and the ending character in the feature vector respectively to obtain a quantifier coordinate pair. When extracting quantifiers subsequently, the specific position of the quantifier can be located from the target text content directly by applying the quantifier coordinate pair. The content of the quantifier probability prediction process mentioned in this step can refer to the content of steps 301 to 303.
[0176] In this embodiment, the quantifiers in the target text content are mainly extracted by using the CRF layer and the PointerNetlayer.
[0177] The relationship extraction model 603 includes a relationship word tagger, which includes a HasQuantity tagger, a HasProperty tagger, and a Qualifies tagger. The HasQuantity tagger is used to tag the entity relationship words corresponding to the quantifiers, the HasProperty tagger is used to tag the attribute relationship words corresponding to the quantifiers, and the Qualifies tagger is used to tag the qualification relationship words corresponding to the quantifiers. Among them, the recognition processes of the entity relationship words, the attribute relationship words, and the qualification relationship words can refer to steps 202 to 204.
[0178] Optionally, the final result obtained by the relationship extraction model 603 can be a relationship word or a matching relationship between a relationship word and a quantifier. Schematically, when the target text content "Xiaohong's height is 156 cm" and the quantifier "156 cm" are input into the relationship extraction model 603, the final result can be "Xiaohong", "height", or the final result can be "(Xiaohong, entity extraction, 156 cm), (Xiaohong, attribute extraction, height), (height, attribute extraction, 156 cm)". This application does not limit the form of the recognition result of this relationship extraction model.
[0179] The classification model 604 is used to judge the category of the quantifier. In some embodiments, the quantifier in the target text content is extracted, and the quantifier is classified by a Softmax classifier, and the classification result corresponding to the quantifier is output. The classification result includes at least one of a quantitative value, a range value, or an approximate value, and the attribute of the quantifier can also be classified. This application does not limit this.
[0180] The unit recognition model 605 is used to recognize the unit in the quantifier. The main unit recognition method is to judge the belonging relationship between the last character of the quantifier and the preset unit set. The belonging relationship includes that the last character is in the preset unit set or the last character is not in the preset unit set. The preset unit set is formed by incorporating the units that have appeared in the training set during the model training process. During the subsequent process of recognizing the quantifier, it is judged whether the last character of the quantifier is in the preset unit set. If it is, the longest unit corresponding to the last character in the preset unit set is output. For example, for the quantifier "20 pieces", the unit "piece" and "piece / min" have appeared in the preset unit set. When outputting the unit, "piece / min" is used as the unit of the quantifier "20 pieces" to avoid losing the characters in the quantifier unit during the process of recognizing the quantifier. If not, traverse the quantifier forward, and use the first character after the first non-alphabetic character as the unit of the quantifier. The specific implementation process can refer to steps 501 to 504.
[0181] In this embodiment, the encoding model 601, the quantifier extraction model 602, the relation extraction model 603, the classification model 604, and the unit recognition model 605 can be applied separately or combined arbitrarily, and the present application does not limit this.
[0182] In summary, a method for identifying quantifiers and their relational words based on deep learning provided by an embodiment of the present application first encodes the target text content, predicts the relational word probability of the encoded feature vector to determine the corresponding quantifier, and based on the quantifier, identifies the entity relational word, the attribute relational word, and the restrictive relational word corresponding to the quantifier from the target text content, and at the same time determines the relationship between the relational word and the quantifier. After identifying the relationship between the quantifier and the relational word corresponding to the quantifier, it is convenient to more deeply understand the actual meaning expressed by the quantifier. The quantifier can also be classified and the unit can be recognized, and the more important quantifier information and the context expression required to understand the quantifier are extracted from the target text content, which improves the information interaction efficiency to a certain extent, enables the user to have a further in-depth understanding of the quantifier in combination with the relational word, and improves the efficiency of understanding the actual meaning expressed by the quantifier.
[0183] Figure 7 is a structural block diagram of a payment device for exchanging resources provided by an exemplary embodiment of the present application, as Figure 6 shown, the device includes: an acquisition module 710, an extraction module 720, a determination module 730, and an output module 740;
[0184] The acquisition module 710 is configured to acquire target text content, and the target text content includes quantifiers;
[0185] The extraction module 720 is configured to extract the quantifiers from the target text content;
[0186] The determination module 730 is configured to determine, from the target text content, a relational word having an associated relationship with the quantifier based on the context content of the quantifier in the target text content, and the relational word is used to express the meaning of the quantifier in the target text content;
[0187] The output module 740 is configured to output the matching relationship between the quantifier and the relational word.
[0188] In an optional embodiment, the determination module 730 is further configured to encode the quantifier and the target text content to obtain a first feature vector corresponding to n characters in the target text content, where n is a positive integer; predict the relational probability of the first feature vector of the i-th character to obtain a probability value that the i-th character belongs to the relational word, 1 ≤ i ≤ n; and determine the relational word from the target text content based on the probability value.
[0189] In an alternative embodiment, the probability value includes a first probability value and a second probability value. The first probability value is used to represent the probability that the i-th character is the starting character of the relational word, and the second probability value is used to represent the probability that the i-th character is the ending character of the relational word;
[0190] The determining module 730 is further configured to determine the first starting character of the relational word from the n characters based on the first probability value corresponding to the n characters; determine the first ending character of the relational word from the n characters based on the second probability value corresponding to the n characters; and obtain the characters from the start of the first starting character to the end of the first ending character as the relational word.
[0191] In an alternative embodiment, the relational word includes at least one of an attribute relational word, an entity relational word, and a qualification relational word;
[0192] In response to the relational word including an attribute relational word, the first starting character includes an attribute starting character, and the first ending character includes an attribute ending character;
[0193] In response to the relational word including an entity relational word, the first starting character includes an entity starting character, and the first ending character includes an entity ending character;
[0194] In response to the relational word including a qualification relational word, the first starting character includes a qualification starting character, and the first ending character includes a qualification ending character.
[0195] In an alternative embodiment, the extraction module 720 is further configured to encode the target text content to obtain a second feature vector corresponding to n characters in the target text content; perform a quantifier probability prediction on the second feature vector corresponding to the i-th character to obtain a probability value that the i-th character belongs to the quantifier; and extract the quantifier from the target text content based on the probability values that the n characters belong to the quantifier.
[0196] In an alternative embodiment, the probability value that the i-th character belongs to the quantifier includes a third probability value and a fourth probability value. The third probability value represents the probability that the i-th character is the starting character of the quantifier, and the fourth probability value is used to represent the probability that the i-th character is the ending character of the quantifier;
[0197] The extraction module 720 is further configured to determine a second start character of the quantifier from the n characters based on the third probability values corresponding to the n characters; determine a second end character of the quantifier from the n characters based on the fourth probability values corresponding to the n characters; and obtain the characters from the start of the second start character to the end of the second end character as the relational word.
[0198] In an alternative embodiment, as Figure 8 shown, the apparatus further includes:
[0199] A classification module 750, configured to perform a classification prediction on the quantifier to obtain a classification result corresponding to the quantifier, where the classification result is used to represent the measurement value of the quantifier, and the classification result includes any one of a quantitative value, a range value, and an approximate value.
[0200] In an alternative embodiment, the output device further includes:
[0201] An output module 740, configured to output the longest unit corresponding to the end character in a preset unit set in response to the end character of the quantifier appearing in the preset unit set, and determine the longest unit as the unit corresponding to the quantifier; or,
[0202] The output module 740 is further configured to perform a forward traversal operation on the quantifier in response to the end character not appearing in the preset unit set, output the character after the first non-alphabetic character in the traversal process, and use the output character as the unit of the quantifier.
[0203] In summary, the text content recognition device provided in the embodiments of the present application determines the quantifier in the target text content, combines the context content of the quantifier, determines the relational word used to describe or limit the quantifier from the target text content, and matches the relationship between the extracted relational word and the quantifier, greatly improving the information interaction rate, enabling the user to have a further in-depth understanding of the quantifier in combination with the relational word, and improving the efficiency of understanding the actual meaning expressed by the quantifier.
[0204] It should be noted that: for the text content recognition device provided in the above embodiments, only the above-mentioned division of each functional module is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the text content recognition device provided in the above embodiments and the embodiments of the text content recognition method belong to the same concept, and the specific implementation process is detailed in the method embodiments, which will not be repeated here.
[0205] Figure 9The structural schematic diagram of a server provided by an exemplary embodiment of the present application is shown. The server can be the Figure 1 server shown. Specifically:
[0206] The server includes a central processing unit (CPU, Central Processing Unit) 901, a system memory 904 including a random access memory (RAM, Random Access Memory) 902 and a read-only memory (ROM, Read Only Memory) 903, and a system bus 905 connecting the system memory 904 and the central processing unit 1201. The server 120 also includes a basic input / output system (I / O system, Input Output System) 906 for transferring information between various components within the computer, and a mass storage device 907 for storing an operating system 913, application programs 914, and other program modules 915.
[0207] The basic input / output system 906 includes a display 908 for displaying information and input devices 909 such as a mouse and a keyboard for user input. Among them, both the display 908 and the input devices 909 are connected to the central processing unit 901 through an input / output controller 910 connected to the system bus 905. The basic input / output system 906 may also include an input / output controller 910 for receiving and processing inputs from multiple other devices such as a keyboard, a mouse, or an electronic stylus. Similarly, the input / output controller 910 also provides output to a display screen, a printer, or other types of output devices.
[0208] The mass storage device 907 is connected to the central processing unit 901 through a mass storage controller (not shown) connected to the system bus 905. The mass storage device 907 and its associated computer-readable medium provide non-volatile storage for the server 120. That is to say, the mass storage device 907 may include a computer-readable medium (not shown) such as a hard disk or a compact disc read-only memory (CD-ROM, Compact Disc Read Only Memory) drive.
[0209] Without loss of generality, computer-readable media can include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes RAM, ROM, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other solid-state storage technologies, CD-ROM, digital versatile disc (DVD) or other optical storage, magnetic tape cartridges, tapes, disk storage or other magnetic storage devices. Of course, those skilled in the art will appreciate that computer storage media is not limited to the above several types. The above-mentioned system memory 904 and mass storage device 907 can be collectively referred to as memory.
[0210] According to various embodiments of the present application, the server can also run on a remote computer on the network connected through a network such as the Internet. That is, the server can be connected to the network 912 through the network interface unit 911 connected to the system bus 905, or in other words, the network interface unit 911 can also be used to connect to other types of networks or remote computer systems (not shown).
[0211] The above-mentioned memory further includes one or more programs, and one or more programs are stored in the memory and configured to be executed by the CPU.
[0212] Embodiments of the present application also provide a computer device, which includes a processor and a memory. At least one instruction, at least one segment of program, code set or instruction set is stored in the memory, and at least one instruction, at least one segment of program, code set or instruction set is loaded and executed by the processor to implement the text content recognition method provided by the above-mentioned method embodiments.
[0213] Embodiments of the present application also provide a computer-readable storage medium, on which at least one instruction, at least one segment of program, code set or instruction set is stored, and at least one instruction, at least one segment of program, code set or instruction set is loaded and executed by the processor to implement the text content recognition method provided by the above-mentioned method embodiments.
[0214] Embodiments of the present application also provide a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to execute the text content recognition method described in any one of the above embodiments.
[0215] Optionally, the computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), solid state drives (SSD), or optical discs, etc. Among them, the random access memory may include resistive random access memory (ReRAM) and dynamic random access memory (DRAM). The serial numbers of the above embodiments of the present application are only for description and do not represent the advantages or disadvantages of the embodiments.
[0216] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above embodiments can be completed by hardware, or can be completed by a program instructing related hardware. The program can be stored in a computer-readable storage medium. The storage medium mentioned above can be read-only memory, a magnetic disk, or an optical disc, etc.
[0217] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for identifying text content, characterized in that, The method includes: Obtaining target text content, where the target text content includes a quantifier; Extracting the quantifier from the target text content; Encoding the quantifier and the target text content to obtain first feature vectors corresponding to n characters in the target text content, where n is a positive integer, and the first feature vector corresponding to the i-th character is used to characterize the i-th character and the quantifier; Performing relationship probability prediction on the first feature vector corresponding to the i-th character to obtain a probability value that the i-th character belongs to a relationship word having an associated relationship with the quantifier, 1 ≤ i ≤ n, and the relationship word is used to express the meaning of the quantifier in the target text content; Determining the relationship word from the target text content based on the probability values that the n characters respectively belong to the relationship word; Outputting the matching relationship between the quantifier and the relationship word.
2. The method according to claim 1, wherein The probability value that the i-th character belongs to the relationship word includes a first probability value and a second probability value, where the first probability value is used to represent the probability that the i-th character is the starting character of the relationship word, and the second probability value is used to represent the probability that the i-th character is the ending character of the relationship word; The determining the relationship word from the target text content based on the probability values respectively corresponding to the n characters includes: Determining a first starting character of the relationship word from the n characters based on the first probability values respectively corresponding to the n characters; Determining a first ending character of the relationship word from the n characters based on the second probability values respectively corresponding to the n characters; Obtaining the characters from the start of the first starting character to the end of the first ending character as the relationship word.
3. The method according to claim 2, wherein The relationship word includes at least one of an attribute relationship word, an entity relationship word, and a qualification relationship word; In response to the relationship word including an attribute relationship word, the first starting character includes an attribute starting character, and the first ending character includes an attribute ending character; In response to the relationship word including an entity relationship word, the first starting character includes an entity starting character, and the first ending character includes an entity ending character; In response to the relationship word including a qualification relationship word, the first starting character includes a qualification starting character, and the first ending character includes a qualification ending character.
4. The method according to any one of claims 1 to 3, characterized in that, The extracting the quantifier from the target text content includes: Encoding the target text content to obtain second feature vectors corresponding to the n characters in the target text content; Performing quantifier probability prediction on the i-th character based on the second feature vector to obtain a probability value that the i-th character belongs to the quantifier; Extracting the quantifier from the target text content based on the probability values that the n characters respectively belong to the quantifier.
5. The method according to claim 4, wherein The probability value that the i-th character belongs to the quantifier includes a third probability value and a fourth probability value, where the third probability value represents the probability that the i-th character is the starting character of the quantifier, and the fourth probability value is used to represent the probability that the i-th character is the ending character of the quantifier; Extracting the quantifier from the target text content based on the probability values that the n characters respectively belong to the quantifier includes: Determining a second starting character of the quantifier from the n characters based on the third probability values respectively corresponding to the n characters; Determining a second ending character of the quantifier from the n characters based on the fourth probability values respectively corresponding to the n characters; Obtaining the characters from the start of the second starting character to the end of the second ending character as the quantifier.
6. The method according to any one of claims 1 to 3, characterized in that After extracting the quantifier from the target text content, it further includes: Performing classification prediction on the quantifier to obtain a classification result corresponding to the quantifier, where the classification result is used to represent the measurement value of the quantifier, and the classification result includes any one of a quantitative value, a range value, and an approximate value.
7. An apparatus for recognizing text content, characterized in that, The device includes: An acquisition module, configured to acquire target text content, where the target text content includes a quantifier; An extraction module, configured to extract the quantifier from the target text content; A determination module, configured to encode the quantifier and the target text content to obtain first feature vectors respectively corresponding to n characters in the target text content, where n is a positive integer, and the first feature vector corresponding to the i-th character is used to characterize the i-th character and the quantifier; performing relationship probability prediction on the first feature vector corresponding to the i-th character to obtain a probability value that the i-th character belongs to a relationship word having an associated relationship with the quantifier, 1 ≤ i ≤ n, where the relationship word is used to express the meaning of the quantifier in the target text content; determining the relationship word from the target text content based on the probability values that the n characters respectively belong to the relationship word; An output module, configured to output a matching relationship between the quantifier and the relationship word.
8. A computer device, characterized in that, The computer device includes a processor and a memory, and at least one program is stored in the memory, and the at least one program is loaded and executed by the processor to implement the text content recognition method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, At least one program is stored in the storage medium, and the at least one program is loaded and executed by a processor to implement the text content recognition method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Electronic medical record processing method and device and equipment
CN110277149A
Core entity labeling method and device and electronic equipment
CN111241832A