Text analysis method and device, electronic equipment and computer readable storage medium
By determining the target separator in text parsing and using preset rules for character segmentation and parsing, credibility screening, the problem of high cost and low efficiency caused by complex parsing rules in the prior art is solved, and simple and fast text parsing is achieved.
Patent Information
- Application Number
- CN202510691968.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-09-02
AI Technical Summary
When processing unstructured text data, the prior art requires writing complicated analytical rules, which leads to increased development and testing costs and low parsing efficiency.
By obtaining the text to be identified, determining the target separator for character segmentation, using preset character parsing rules for parsing, and filtering results based on the parsing credibility, to achieve intelligent text parsing.
Simplifies the text parsing process, reduces development and testing costs, and improves analysis efficiency and accuracy.
Smart Images

Figure CN120579536A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to, but are not limited to, the field of text parsing, and in particular to a text parsing method, apparatus, electronic device, and computer-readable storage medium. Background Art
[0002] With the rapid development of social economy and the continuous advancement of science and technology, life is also moving towards intelligence. Intelligent systems have been increasingly used to uniformly manage related data. For example, in insurance business systems or smart medical systems, data is stored in different forms. Among them, the storage of text-type unstructured data requires the downstream application to use different parsing rules to parse the parameter content in the data when applying text-type unstructured data. In addition, the content format of the data is different, and there are Chinese, numbers, punctuation marks and special characters. This requires downstream developers to use different parsing rules for parsing when parsing extended parameters, resulting in increased development and testing costs. Summary of the Invention
[0003] The following is a summary of the subject matter described in detail herein. This summary is not intended to limit the scope of the claims.
[0004] In order to solve the problems mentioned in the above background technology, the embodiments of the present application provide a text parsing method, device, electronic device and computer-readable storage medium, which make text parsing simpler and faster, and reduce development costs and testing costs.
[0005] In a first aspect, an embodiment of the present application provides a text parsing method, comprising:
[0006] Get the text to be recognized;
[0007] Determining a target delimiter from the text to be recognized;
[0008] Performing character segmentation processing on the text to be recognized according to the target separator to obtain a plurality of segmentation characters;
[0009] Parsing the plurality of segmented characters according to a preset character parsing rule to obtain a plurality of character parsing results, wherein each of the character parsing results carries a parsing reliability;
[0010] When the resolution reliability is greater than a preset reliability threshold, the character resolution result corresponding to the resolution reliability is used as the text resolution result.
[0011] In a second aspect, an embodiment of the present application further provides a text parsing device, the device comprising:
[0012] An acquisition unit, used for acquiring the text to be recognized;
[0013] A recognition unit, configured to determine a target delimiter from the text to be recognized;
[0014] a segmentation unit, configured to perform character segmentation processing on the text to be recognized according to the target separator to obtain a plurality of segmented characters;
[0015] a parsing unit, configured to parse the plurality of segmented characters according to a preset character parsing rule to obtain a plurality of character parsing results, wherein each of the character parsing results carries a parsing reliability;
[0016] The comparison unit is configured to, when the resolution reliability is greater than a preset reliability threshold, use the character resolution result corresponding to the resolution reliability as the text resolution result.
[0017] In a third aspect, an embodiment of the present application further provides an electronic device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the text parsing method as described in the first aspect above is implemented.
[0018] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium storing computer-executable instructions, wherein the computer-executable instructions are used to execute the text parsing method described in the first aspect above.
[0019] According to the text parsing method of the embodiment provided by the present application, there are at least the following beneficial effects: in the process of text parsing, it is first necessary to obtain the text to be recognized, and then determine the target delimiter from the text to be recognized; then, according to the target delimiter, the text to be recognized is subjected to character segmentation processing to obtain multiple segmented characters; then, according to the pre-set character parsing rules, the multiple segmented characters are parsed to obtain multiple character parsing results, wherein each character parsing result carries a parsing credibility; finally, when the parsing credibility is greater than the pre-set credibility threshold, the character parsing result corresponding to the parsing credibility is used as the text parsing result. Through the above technical solution, the target delimiter is first determined from the text to be recognized, and then, according to the target delimiter, the text to be recognized is subjected to character segmentation processing to obtain multiple segmented characters; then, according to the character parsing rules, the multiple segmented characters are parsed to obtain multiple character parsing results, and finally, the character parsing results are subjected to parsing credibility detection processing, thereby realizing intelligent text detection processing, without the need for developers to write complicated recognition rules as in the past, making text parsing simpler and faster, and reducing development costs and testing costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The accompanying drawings are used to provide a further understanding of the technical solution of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the technical solution of the present application and do not constitute a limitation on the technical solution of the present application.
[0021] Figure 1 This is a flowchart of a text parsing method provided by an embodiment of the present application;
[0022] Figure 2 yes Figure 1 A schematic flow chart of a specific implementation of step S200;
[0023] Figure 3 yes Figure 1 A schematic flow chart of a specific implementation of step S300;
[0024] Figure 4 yes Figure 1 A schematic flow chart of a specific implementation of step S400;
[0025] Figure 5 It is executed Figure 1 A schematic flow chart of a specific implementation method after step S400;
[0026] Figure 6 It is executed Figure 5 A schematic diagram of a specific implementation flow chart after step S470;
[0027] Figure 7 It is executed Figure 1 A schematic flow chart of a specific implementation method after step S500;
[0028] Figure 8 is a schematic diagram of a text parsing device provided by an embodiment of the present application;
[0029] Figure 9 This is a schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0030] In order to make the purpose, technical solutions and advantages of this application more clearly understood, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0031] It should be noted that although the device schematics illustrate functional module divisions and the flowcharts illustrate logical sequences, in certain circumstances, the steps shown or described may be performed in a sequence that differs from the module divisions in the device or the sequence in the flowcharts. The terms "first," "second," and so on, used in the specification, claims, and drawings, are used to distinguish similar items and are not necessarily used to describe a specific sequence or precedence.
[0032] It should be noted that, unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0033] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Artificial Intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results.
[0034] AI is a new technical discipline that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence. Artificial intelligence is a branch of computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. Artificial intelligence can simulate the information processes of human consciousness and thinking. It also refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results.
[0035] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0036] Artificial intelligence, or AI, is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to achieve optimal results.
[0037] The servers involved in artificial intelligence technology can be independent servers or cloud servers that provide basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), as well as big data and artificial intelligence platforms.
[0038] The present application provides a text parsing method, device, electronic device and computer-readable storage medium. In the process of text parsing, it is first necessary to obtain the text to be recognized, and then determine the target delimiter from the text to be recognized; then, according to the target delimiter, the text to be recognized is subjected to character segmentation processing to obtain multiple segmentation characters; then, according to a pre-set character parsing rule, the multiple segmentation characters are parsed to obtain multiple character parsing results, wherein each character parsing result carries a parsing credibility; finally, when the parsing credibility is greater than a pre-set credibility threshold, the character parsing result corresponding to the parsing credibility is used as the text parsing result. Through the above technical solution, the target delimiter is first determined from the text to be recognized, and then, according to the target delimiter, the text to be recognized is subjected to character segmentation processing to obtain multiple segmentation characters; then, according to the character parsing rule, the multiple segmentation characters are parsed to obtain multiple character parsing results, and finally, the character parsing result is subjected to parsing credibility detection processing, thereby realizing intelligent text detection processing, and eliminating the need for developers to write complicated recognition rules as in the past, making text parsing simpler and faster, and reducing development costs and testing costs.
[0039] The text parsing method provided in the embodiment of the present application relates to the field of text parsing. The text parsing method provided in the embodiment of the present application can be applied in a terminal, can also be applied in a server side, and can also be software running in a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, or can be configured as a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.
[0040] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected via a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.
[0041] It should be noted that in each specific embodiment of the present application, when it comes to the need to perform relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first, and the collection, use, and processing of such data will comply with relevant laws, regulations, and standards. In addition, when the embodiment of the present application needs to obtain the user's sensitive personal information, the user's separate permission or consent will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the user's separate permission or consent, the necessary user-related data for the normal operation of the embodiment of the present application will be obtained.
[0042] The embodiments of the present application are further described below with reference to the accompanying drawings.
[0043] like Figure 1 As shown, Figure 1 This is a flowchart of a text parsing method provided by an embodiment of the present application, which includes the following steps:
[0044] Step S100: Obtain the text to be recognized.
[0045] The text parsing method provided in the embodiment of the present application first needs to obtain the text to be recognized during the text parsing process. For example, in an insurance business system, when a user needs to create their own account name on the system, they will enter some character information according to their own preferences. For example, the account name they enter can be "aa@123". In order to recognize and process these account names with various complex characters, text parsing is required. In the past, the recognition process required the writing of various complex text recognition rules during the system development and maintenance process, which would reduce the efficiency of system development and extend the development cycle. Alternatively, when a business person uses a paper insurance policy to sign a purchase and sale agreement with a user, after the user completes the policy signing, the business person needs to take an image of the paper policy and upload it for processing. The insurance business system then needs to recognize and process the content filled in on the paper policy and convert it into electronic text. Then, the text parsing method of the embodiment of the present application can be used to recognize and parse the relevant character text. Alternatively, in the field of smart healthcare, when a doctor writes handwritten text on a medical record, prescription, or examination sheet, and later needs to record this text into the smart healthcare system, it is necessary to take a photo of these files and upload them. The system will then perform text recognition processing on the photographed images and convert them into electronic text. The text parsing method of the embodiment of the present application can then be used to perform recognition and parsing processing on the relevant character text.
[0046] It should be noted that during the text parsing process of this application, the user's permission or consent will be obtained before obtaining the text to be identified, and the collection, use and processing of such data will comply with relevant laws, regulations and standards. In addition, when the embodiment of this application needs to obtain the user's sensitive personal information, the user's separate permission or consent will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the user's separate permission or consent, the necessary user-related data for the normal operation of the embodiment of this application will be obtained.
[0047] It is worth noting that the text to be recognized may include letters, numbers, words and special characters; the text parsing method based on the embodiment of the present application can accurately and quickly recognize and process various types of characters in the text to be recognized, thereby greatly improving the efficiency of text recognition.
[0048] Step S200: determining a target separator from the text to be recognized.
[0049] The text parsing method provided by the embodiment of the present application, after obtaining the text to be recognized, can first determine the target delimiter from the text to be recognized, and then perform character segmentation processing on the text to be recognized according to the target delimiter to obtain multiple segmentation characters. For example, the target delimiter in the text to be recognized can include a colon, a space, and an equal sign; when the above target delimiters exist in the text to be recognized, character segmentation processing will be performed according to the target delimiter. For example, a text to be parsed is "My name: Xiao Ming", so the colon in the text to be parsed is the target delimiter, so the characters on both ends of the colon can be split, so that the characters "My name" and "Xiao Ming" can be obtained as two segmentation characters, and the two segmentation characters can be subsequently recognized and parsed separately. Alternatively, a text to be parsed is "Insurance name = accident insurance", so the equal sign in the text to be parsed is the target delimiter, so the characters on both ends of the equal sign can be split, so that the characters "Insurance name" and "Accident insurance" can be obtained as two segmentation characters, and the two segmentation characters can be subsequently recognized and parsed separately. Alternatively, in a smart medical system, a text to be parsed is "patient name: Xiaoqiang". Therefore, the colon in the above text to be parsed is the target separator, so the characters on both ends of the colon can be split and processed, so that the two separated characters "patient name" and "Xiaoqiang" can be obtained. Subsequently, these two separated characters can be identified and parsed separately.
[0050] It is worth noting that in insurance business systems or smart medical systems, the target delimiter can be pre-set by R&D personnel, so that in the subsequent process of determining the target delimiter from the text to be recognized, the matching and determination can be performed according to the delimiter pre-set by the R&D personnel. The entire target delimiter determination process is simple and fast, which improves the efficiency of subsequent text parsing.
[0051] like Figure 2 As shown, determining the target delimiter from the text to be recognized may include the following steps:
[0052] Step S210, extracting a plurality of set delimiters from a preset delimiter database;
[0053] Step S220 , matching the set delimiter with the text to be recognized to obtain the target delimiter.
[0054] For steps S210 to S220, in the process of determining the target delimiter from the text to be recognized, several set delimiters are first extracted from a pre-set delimiter database; subsequently, the set delimiter can be matched with the text to be recognized to obtain the corresponding target delimiter; through the above technical solution, the entire target delimiter determination process can be made simpler and faster.
[0055] It is worth noting that the set delimiters in the delimiter database can be characters pre-set and stored by R&D personnel, and the delimiter database covers as many delimiter characters as possible, and then in the subsequent process of matching the set delimiter with the text to be recognized, the target delimiter can be determined from the text to be recognized, so that the target delimiter determination process can be more accurate and quick. For example, the set delimiters include equal signs, colons, and spaces, etc., and the above-mentioned pre-set delimiters are stored in the delimiter database. Later, in the process of text parsing, delimiters such as equal signs, colons, and spaces can be extracted from the pre-set delimiter database, and then these delimiters can be matched with the text to be recognized, so that the corresponding target delimiter can be selected quickly and accurately.
[0056] Step S300: performing character segmentation processing on the text to be recognized according to the target separator to obtain a plurality of segmented characters.
[0057] The text parsing method provided in the embodiment of the present application can determine the target delimiter from the text to be recognized, and then perform character segmentation processing on the text to be recognized according to the target delimiter to obtain multiple segmentation characters, and then perform parsing processing on the multiple segmentation characters to obtain multiple character parsing results; through the above technical solution, the text parsing process can be made simpler, faster and more accurate. For example, in an insurance business system, when the text to be recognized is "The type of insurance I purchased: commercial insurance", the colon is the target delimiter of the text to be recognized, so when performing character segmentation processing on the text to be recognized according to the colon, the two segmentation characters "The type of insurance I purchased" and "commercial insurance" can be obtained respectively; the two segmentation characters can be further parsed, so that the text parsing process can be made more accurate and faster, and the occurrence of character parsing errors caused by incorrect division of the text to be recognized can be avoided. Alternatively, in a smart medical system, when the text to be recognized is "patient number: patient name: patient gender", the colon is the target separator of the text to be recognized. Therefore, by performing character segmentation processing on the text to be recognized based on the colon, the three segmentation characters of "patient number", "patient name" and "patient gender" can be obtained respectively; subsequently, these three segmentation characters can be further parsed, making the text parsing process more accurate and quick.
[0058] Exemplarily, there can be multiple segmentation characters. For example, in an insurance business system, when the text to be recognized is "The type of insurance I purchased: commercial insurance; My name: Xiao Ming", the colon and semicolon are both target separators. Therefore, in the process of character segmentation of the text to be recognized according to the corresponding target separators, multiple segmentation characters can be obtained. The above-mentioned text to be recognized can be segmented into four segmentation characters: "The type of insurance I purchased", "commercial insurance", "My name" and "Xiao Ming". Subsequently, these four segmentation characters can be parsed to obtain the corresponding character parsing results. The whole process is simple, fast and reliable. Alternatively, in a smart medical system, when the text to be recognized is "patient number: 123; patient name: Xiao Ming; examination item: electrocardiogram", the colon and semicolon are both target separators. Therefore, in the process of character segmentation of the text to be recognized according to the corresponding target separators, multiple segmentation characters can be obtained. The above-mentioned text to be recognized can be segmented into six segmentation characters: "patient number", "123", "patient name", "Xiao Ming", "examination item" and "electrocardiogram". Subsequently, these six segmentation characters can be parsed to obtain the corresponding character parsing results.
[0059] like Figure 3 As shown, character segmentation processing is performed on the text to be recognized according to the target separator to obtain multiple segmented characters, which may include the following steps:
[0060] Step S310, performing initial segmentation processing on the text to be recognized according to the target separator to obtain a plurality of initial segmentation characters;
[0061] Step S320, verifying the integrity of the plurality of initially divided characters to obtain integrity verification information;
[0062] Step S330: When the integrity verification information indicates that the initial segmentation character is a complete character, the initial segmentation character is used as a segmentation character.
[0063] For steps S310 to S330, in the process of performing character segmentation processing on the text to be recognized according to the target delimiter to obtain multiple segmentation characters, firstly, the text to be recognized is initially segmented according to the target delimiter to obtain multiple initial segmentation characters; then the integrity of the multiple initial segmentation characters is verified to obtain integrity verification information; finally, when the integrity verification information indicates that the initial segmentation character is a complete character, the initial segmentation character is used as the segmentation character. Through the above technical solution, after performing initial segmentation processing on the text to be recognized according to the target delimiter to obtain multiple initial segmentation characters, the integrity of the initial segmentation characters obtained by segmentation can be verified; when the initial segmentation characters are verified to be complete characters, the initial segmentation characters can be used as the final segmentation characters, so that subsequent text parsing can be more reasonable and accurate, and the entire text parsing process can be more reliable.
[0064] For example, in an insurance business system, when the text to be recognized is "Insurance buyer: Zhang San", where the colon is the target delimiter, the text to be recognized can be initially divided according to the target delimiter, and two initial division characters "Insurance buyer" and "Zhang San" can be obtained; then the initial division characters obtained by division are verified to obtain integrity verification information; when the integrity verification information indicates that the initial division characters are complete characters, the two initial division characters obtained by division can be used as division characters. For another example, when the text to be recognized is initially divided according to the target delimiter to obtain two initial division characters "Insurance buyer: Zhang" and "San", integrity verification information is obtained when the two initial division characters are subjected to integrity verification, however, the integrity verification information indicates that the initial division characters are not complete characters, therefore the initial division characters obtained by division will not be used as division characters, and the target delimiter will be reused to perform initial division on the text to be recognized. Through the above technical solution, the accuracy of the division of the text to be recognized can be greatly improved. Alternatively, in a smart medical system, when the text to be recognized is "patient name: Xiao Ming", where the colon is the target delimiter, the text to be recognized can be initially divided according to the target delimiter to obtain two initial division characters "patient name" and "Xiao Ming"; the initial division characters obtained by the division are then verified to obtain integrity verification information; when the integrity verification information indicates that the initial division characters are complete characters, the two initial division characters obtained by the division can be used as division characters. For another example, when the text to be recognized is initially divided according to the target delimiter to obtain two initial division characters "patient name: Xiao" and "Ming", integrity verification information is obtained when the integrity verification information is performed on the two initial division characters. However, the integrity verification information indicates that the initial division characters are not complete characters, so the initial division characters obtained by the division will not be used as division characters, and the target delimiter will be reused to perform initial division on the text to be recognized.
[0065] Step S400: parsing the plurality of segmented characters according to a preset character parsing rule to obtain a plurality of character parsing results, wherein each character parsing result carries a parsing reliability.
[0066] The text parsing method provided in the embodiment of the present application can, after obtaining multiple segmented characters, parse the multiple segmented characters according to pre-set character parsing rules to obtain multiple character parsing results, wherein each character parsing result will carry a corresponding parsing credibility. Subsequently, the corresponding character parsing results can be screened and processed according to the parsing credibility corresponding to the character parsing result, thereby further improving the reliability and accuracy of text parsing.
[0067] It is worth noting that the character parsing rules can be parsing rules pre-set by system developers, and the parsing rules can be a pre-trained large language model. According to the pre-trained large language model, each segmented character obtained can be parsed and processed to obtain the corresponding character parsing results. For example, for text, the pre-trained large language model can parse and process the relevant text information; for numbers, the pre-trained large language model can also parse and process the relevant text information; for English letters, the pre-trained large language model can also parse and process the relevant English letters. Through the pre-trained large language model, various types of segmented characters can be well recognized and processed.
[0068] It is worth noting that each character parsing result will carry a parsing credibility; based on the parsing credibility, the accuracy of the character parsing result obtained by the parsing can be judged; when the parsing credibility is greater than a preset credibility threshold, the character parsing result corresponding to the parsing credibility can be used as the text parsing result; when the parsing credibility is not greater than the preset credibility threshold, the character parsing result corresponding to the parsing credibility will not be used as the text parsing result; through the above technical solution, the text parsing process can be more accurate and reliable, which greatly improves the accuracy of text parsing.
[0069] like Figure 4 As shown, parsing multiple segmented characters according to preset character parsing rules to obtain multiple character parsing results may include the following steps:
[0070] Step S410, matching the segmented character with the preset character type information to obtain the target character type corresponding to the segmented character;
[0071] Step S420, selecting a target character parsing rule from character parsing rules according to the target character type;
[0072] Step S430, performing character recognition processing on the corresponding segmented characters according to the target character parsing rule to obtain initial character recognition information;
[0073] Step S440 , performing semantic interpretation processing on the initial character recognition information to obtain a character parsing result.
[0074] In steps S410 to S440, in the process of parsing multiple segmented characters according to a pre-set character parsing rule to obtain multiple character parsing results, the segmented characters are first matched with pre-set character type information to obtain a target character type corresponding to the segmented characters; then, a target character parsing rule is selected from the character parsing rule according to the target character type; then, character recognition processing is performed on the corresponding segmented characters according to the target character parsing rule to obtain initial character recognition information; finally, semantic interpretation processing is performed on the initial character recognition information to obtain a character parsing result. Through the above technical solution, the text parsing process can be made more accurate.
[0075] For example, in an insurance business system, the text to be recognized is "insurance account number: user 123aa". In the process of text segmentation according to the target delimiter, the four segmentation characters "insurance account number", "user", "123" and "aa" can be obtained; subsequently, each segmentation character can be matched with the pre-set character type information to obtain the target character type. For example, "insurance account number" and "user" are text types, "123" is a numeric type, and "aa" is a letter type. Then, the target character parsing rule can be selected from the character parsing rule according to the relevant character type. For example, according to the text type parsing rule, the character initial recognition information can be obtained by performing character recognition processing on "insurance account number" and "user", and then the obtained character initial recognition information can be semantically interpreted to obtain the character parsing result; according to the numeric type parsing rule, the character initial recognition information can be obtained by performing character recognition processing on "123", and then the character initial recognition information can be semantically interpreted to obtain the character parsing result; according to the letter type parsing rule, the character initial recognition information can be obtained by performing character recognition processing on "aa", and then the character initial recognition information can be semantically interpreted to obtain the character parsing result. Alternatively, in a smart medical system, the text to be recognized is "Attending physician: Xiaoqiang; Department: 3". In the process of text segmentation according to the target delimiter, the four segmentation characters "Attending physician", "Xiaoqiang", "Department" and "3" can be obtained. Subsequently, each segmentation character can be matched with the pre-set character type information to obtain the target character type. For example, "Attending physician", "Xiaoqiang" and "Department" are text types, and "3" is a number type. Then, the target character parsing rule can be selected from the character parsing rule according to the relevant character type. For example, according to the text type parsing rule, character recognition processing can be performed on "Attending physician", "Xiaoqiang" and "Department" to obtain initial character recognition information, and then the obtained initial character recognition information can be semantically interpreted to obtain the character parsing result; and according to the number type parsing rule, character recognition processing can be performed on "3" to obtain initial character recognition information, and then the obtained initial character recognition information can be semantically interpreted to obtain the character parsing result.
[0076] Step S500: when the resolution reliability is greater than a preset reliability threshold, the character resolution result corresponding to the resolution reliability is used as the text resolution result.
[0077] The text parsing method provided in the embodiment of the present application parses multiple segmented characters according to pre-set character parsing rules to obtain multiple character parsing results. Since each character parsing result carries a parsing credibility, when the parsing credibility is greater than a pre-set credibility threshold, the character parsing result corresponding to the parsing credibility can be used as the text parsing result, thereby improving the accuracy of text parsing.
[0078] It is worth noting that when the resolution confidence is not greater than a preset confidence threshold, the character resolution result corresponding to the resolution confidence will not be used as the text resolution result, which greatly improves the reliability of text resolution.
[0079] like Figure 5 As shown, after parsing multiple segmented characters according to the preset character parsing rules to obtain multiple character parsing results, the following steps may be included:
[0080] Step S450: If the parsing confidence level is not greater than the confidence threshold, the character parsing result corresponding to the parsing confidence level is deleted, and the segmented characters corresponding to the character parsing result are marked to obtain marked characters.
[0081] Step S460: parsing the marked characters based on the pre-trained character parsing model to obtain a character inference result;
[0082] Step S470, performing credibility verification processing on the character inference result to obtain character credibility;
[0083] Step S480: When the character credibility is greater than the credibility threshold, the character inference result is used as the text parsing result.
[0084] For steps S450 to S480, after parsing multiple segmented characters according to the preset character parsing rules to obtain multiple character parsing results, if the parsing credibility is not greater than the credibility threshold, the character parsing results corresponding to the parsing credibility can be deleted, and the segmented characters corresponding to the character parsing results can be marked to obtain marked characters; then the marked characters are parsed based on the pre-trained character parsing model to obtain character inference results; then the character inference results are credibility verified to obtain character credibility; if the character credibility is greater than the credibility threshold, the character inference results can be used as the text parsing result. Through the above technical solution, the text parsing function can be maximized; only when the initial parsing recognition cannot parse the text normally, the character parsing model will be used to continue to parse the characters that were previously unable to parse normally, and then the credibility of the character parsing will be further verified, which can also greatly improve the accuracy of text parsing.
[0085] It is worth noting that the preset character parsing rules are first used to parse the segmented characters to obtain multiple character parsing results; when the parsing credibility of the character parsing results is not greater than the preset credibility threshold, the character parsing model will be used to parse the previous segmented characters again; given that the functionality of the character parsing model is stronger than the preset character parsing rules, it can well parse the segmented characters that could not be parsed normally before, thereby further improving the feasibility and accuracy of text parsing.
[0086] like Figure 6 As shown, after the character inference result is subjected to credibility verification processing and the character credibility is obtained, the following steps may be included:
[0087] Step S471: If the character credibility is not greater than the credibility threshold, the marked character corresponding to the character inference result is regarded as an abnormal character;
[0088] Step S472: Delete the abnormal characters.
[0089] For steps S471 to S472, after the character speculation results are subjected to credibility verification processing to obtain the character credibility, if the character credibility is not greater than the credibility threshold, the marked characters corresponding to the character speculation results are treated as abnormal characters, and then the abnormal characters are deleted to eliminate the abnormal characters in the text to be recognized, and the user can also be reminded to pay attention to the standardization of the input characters.
[0090] It is worth noting that when the preset character parsing rules and the character parsing model are unable to correctly identify and parse the relevant characters, the corresponding characters will be regarded as abnormal characters, and then the corresponding characters will be eliminated, and the user will be informed that there are abnormal characters in the previously entered characters to be recognized, and pay attention to standardize and correct the relevant characters.
[0091] like Figure 7 As shown, when the resolution reliability is greater than a preset reliability threshold, after taking the character resolution result corresponding to the resolution reliability as the text resolution result, the following steps may be included:
[0092] Step S600, adding tag information to the text parsing result to obtain a text parsing data packet;
[0093] Step S700: transferring the text parsed data packet to a preset memory.
[0094] For steps S600 to S700, when the parsing confidence is greater than the preset confidence threshold, the character parsing result corresponding to the parsing confidence is used as the text parsing result, and then tag information is added to the text parsing result to obtain a text parsing data packet; then the text parsing data packet is transferred to a pre-set memory so that developers can review and process the relevant text parsing process later, so as to facilitate subsequent maintenance of the system.
[0095] It is worth noting that tagging information is used to identify the type, attributes, or purpose of the parsing results. For example, in natural language processing, part-of-speech tags (such as nouns and verbs), entity type tags (such as names and places), and semantic role tags can be added. Tagging information can be attached to the parsing results in the form of metadata. Storage is a preconfigured storage device or storage system used to store text parsing data packets. It can be a local hard drive, a database, or a cloud storage service.
[0096] In addition, if Figure 8 As shown, an embodiment of the present application further provides a text parsing device 10, the device comprising:
[0097] An acquisition unit 100 is used to acquire a text to be recognized;
[0098] The recognition unit 200 is used to determine a target delimiter from the text to be recognized;
[0099] The segmentation unit 300 is configured to perform character segmentation processing on the text to be recognized according to the target separator to obtain a plurality of segmented characters;
[0100] The parsing unit 400 is configured to parse the plurality of segmented characters according to a preset character parsing rule to obtain a plurality of character parsing results, wherein each of the character parsing results carries a parsing reliability;
[0101] The comparison unit 500 is configured to, when the resolution reliability is greater than a preset reliability threshold, use the character resolution result corresponding to the resolution reliability as the text resolution result.
[0102] It should be noted that, in the process of text parsing, it is first necessary to obtain the text to be recognized, and then determine the target delimiter from the text to be recognized; then, according to the target delimiter, the text to be recognized is subjected to character segmentation processing to obtain multiple segmented characters; then, according to the pre-set character parsing rules, the multiple segmented characters are parsed to obtain multiple character parsing results, wherein each character parsing result carries a parsing credibility; finally, when the parsing credibility is greater than the pre-set credibility threshold, the character parsing result corresponding to the parsing credibility is used as the text parsing result. Through the above technical solution, the target delimiter is first determined from the text to be recognized, and then, according to the target delimiter, the text to be recognized is subjected to character segmentation processing to obtain multiple segmented characters; then, according to the character parsing rules, the multiple segmented characters are parsed to obtain multiple character parsing results, and finally, the character parsing results are subjected to parsing credibility detection processing, thereby realizing intelligent text detection processing, and eliminating the need for developers to write complicated recognition rules as in the past, making text parsing simpler and faster, and reducing development costs and testing costs.
[0103] The specific implementation of the text parsing device 10 is substantially the same as the specific embodiment of the above-mentioned text parsing method, and will not be described in detail here.
[0104] In addition, if Figure 9 As shown, an embodiment of the present application further provides an electronic device 700 , which includes: a memory 720 , a processor 710 , and a computer program stored in the memory 720 and executable on the processor 710 .
[0105] The processor 710 and the memory 720 may be connected via a bus or other means.
[0106] The non-transitory software programs and instructions required to implement the text parsing methods of the above embodiments are stored in the memory 720 , and when executed by the processor 710 , the text parsing methods of the above embodiments are performed.
[0107] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.
[0108] In addition, an embodiment of the present application also provides a computer-readable storage medium, which stores computer-executable instructions, and the computer-executable instructions are executed by a processor 710 or a controller, for example, by a processor 710 in the above-mentioned device embodiment, so that the above-mentioned processor 710 can execute the text parsing method in the above-mentioned embodiment.
[0109] The above embodiments may be used in combination, and modules with the same name in different embodiments may be the same or different.
[0110] The foregoing description describes specific embodiments of the present application, and other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in an order different from that described in the embodiments and still achieve the desired results. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0111] The various embodiments in this application are described in a progressive manner. Similar portions between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from the other embodiments. In particular, the device, equipment, and computer-readable storage medium embodiments are generally similar to the method embodiments, so their descriptions are relatively simple. For relevant portions, refer to the descriptions of the method embodiments.
[0112] The apparatus, device, computer-readable storage medium and method provided in the embodiments of the present application correspond to each other. Therefore, the apparatus, device and non-volatile computer storage medium also have similar beneficial technical effects as the corresponding method. Since the beneficial technical effects of the method have been described in detail above, the beneficial technical effects of the corresponding apparatus, device and computer storage medium will not be repeated here.
[0113] In the 1990s, technological improvements could be clearly distinguished as either hardware improvements (for example, improvements to circuit structures like diodes, transistors, and switches) or software improvements (improvements to process flows). However, with the advancement of technology, many process flow improvements today can now be considered direct improvements to hardware circuit structures. Designers almost always create the corresponding hardware circuit structure by programming the improved process flow into the hardware circuit. Therefore, it cannot be said that a process flow improvement cannot be implemented using hardware modules. For example, a programmable logic device (PLD), such as a field programmable gate array (FPGA), is an integrated circuit whose logical function is determined by user programming. Designers can "integrate" a digital system on a PLD by programming it themselves, without having to hire a chip manufacturer to design and manufacture a dedicated integrated circuit chip. Moreover, nowadays, instead of manually fabricating integrated circuit chips, this programming is mostly performed using "logic compiler" software. This is similar to the software compiler used during program development. Before compilation, the original code must also be written in a specific programming language, called a hardware description language (HDL). There is not just one HDL, but many, such as ABEL (Advanced Boolean Express Language), AHDL (Altera Hardware Description Language), Confluence, CUPL (Cornell University Programming Language), HDCal, JHDL (Java Hardware Description Language), Lava, Lola, MyHDL, PALASM, RHDL (Ruby Hardware Description Language), etc. Currently, the most commonly used are VHDL (Very-High-Speed Integrated Circuit Hardware Description Language) and Verilog. Those skilled in the art will also understand that by simply programming the method flow in one of these hardware description languages and then programming it into an integrated circuit, a hardware circuit that implements the logic method flow can be easily obtained.
[0114] The controller can be implemented in any suitable manner. For example, the controller can take the form of a microprocessor or processor and a computer-readable medium storing computer-readable program code (e.g., software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers. Examples of controllers include, but are not limited to, the following microcontrollers: ARC625D, Atmel AT91SAM, Microchip PIC18F26K20, and Silicone Labs C8051F320. The memory controller can also be implemented as part of the control logic of the memory. Those skilled in the art will also know that in addition to implementing the controller in a purely computer-readable program code manner, it is entirely possible to implement the same function in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, such a controller can be considered a hardware component, and the devices included therein for implementing various functions can also be considered as structures within the hardware component. Or even, the devices for implementing various functions can be considered as both software modules implementing the method and structures within the hardware component.
[0115] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0116] For the convenience of description, the above devices are described as being divided into various units according to their functions. Of course, when implementing the embodiments of the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0117] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the embodiments of the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the embodiments of the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0118] This specification is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0119] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0120] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0121] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0122] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0123] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0124] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0125] In the embodiments of the present application, "at least one" refers to one or more, and "more" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent the existence of A alone, the existence of A and B at the same time, and the existence of B alone. Among them, A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b and c can represent: a, b, c, a and b, a and c, b and c or a and b and c, where a, b, c can be single or multiple.
[0126] Embodiments of the present application may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. Embodiments of the present application may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communications network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0127] The various embodiments in this application are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiment is generally similar to the method embodiment, so the description is relatively simple. For relevant parts, refer to the partial description of the method embodiment.
[0128] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application should be included within the scope of the claims of the present application.
Claims
1. A text parsing method, characterized in that: include: Get the text to be recognized; Determining a target delimiter from the text to be recognized; Performing character segmentation processing on the text to be recognized according to the target separator to obtain a plurality of segmentation characters; Parsing the plurality of segmented characters according to a preset character parsing rule to obtain a plurality of character parsing results, wherein each of the character parsing results carries a parsing reliability; When the resolution reliability is greater than a preset reliability threshold, the character resolution result corresponding to the resolution reliability is used as the text resolution result.
2. The text parsing method according to claim 1, wherein: The step of determining a target delimiter from the text to be recognized includes: Extracting a number of set delimiters from a preset delimiter database; The set delimiter is matched with the text to be recognized to obtain the target delimiter.
3. The text parsing method according to claim 1, wherein: The character segmentation process is performed on the text to be recognized according to the target separator to obtain a plurality of segmented characters, including: Performing initial segmentation processing on the text to be recognized according to the target separator to obtain a plurality of initial segmentation characters; Verifying the integrity of the plurality of initially divided characters to obtain integrity verification information; In a case where the integrity verification information indicates that the initial partitioning character is a complete character, the initial partitioning character is used as the segmentation character.
4. The text parsing method according to claim 1, wherein: The parsing of the plurality of segmented characters according to the preset character parsing rules to obtain a plurality of character parsing results includes: Matching the segmentation character with preset character type information to obtain a target character type corresponding to the segmentation character; Selecting a target character parsing rule from the character parsing rules according to the target character type; Performing character recognition processing on the corresponding segmented characters according to the target character parsing rule to obtain initial character recognition information; The character initial recognition information is semantically interpreted to obtain the character parsing result.
5. The text parsing method according to claim 1, wherein: After parsing the plurality of segmented characters according to the preset character parsing rules to obtain a plurality of character parsing results, the method further includes: If the resolution reliability is not greater than the reliability threshold, deleting the character resolution result corresponding to the resolution reliability, and marking the segmented character corresponding to the character resolution result to obtain a marked character; Parsing the marked characters based on a pre-trained character parsing model to obtain a character inference result; Performing credibility verification on the character inference result to obtain character credibility; When the character credibility is greater than the credibility threshold, the character inference result is used as the text parsing result.
6. The text parsing method according to claim 5, characterized in that: After performing credibility verification on the character inference result to obtain character credibility, the method further includes: When the character credibility is not greater than the credibility threshold, treating the marked character corresponding to the character inference result as an abnormal character; The abnormal characters are deleted.
7. The text parsing method according to claim 1, characterized in that: When the resolution reliability is greater than a preset reliability threshold, after taking the character resolution result corresponding to the resolution reliability as the text resolution result, the method further includes: Adding tag information to the text parsing result to obtain a text parsing data packet; The text parsed data packet is transferred to a preset memory.
8. A text analysis device, characterized in that: The device comprises: An acquisition unit, used for acquiring the text to be recognized; A recognition unit, configured to determine a target delimiter from the text to be recognized; a segmentation unit, configured to perform character segmentation processing on the text to be recognized according to the target separator to obtain a plurality of segmented characters; a parsing unit, configured to parse the plurality of segmented characters according to a preset character parsing rule to obtain a plurality of character parsing results, wherein each of the character parsing results carries a parsing reliability; The comparison unit is configured to, when the resolution reliability is greater than a preset reliability threshold, use the character resolution result corresponding to the resolution reliability as the text resolution result.
9. An electronic device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the text parsing method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing computer-executable instructions, characterized in that: The computer-executable instructions are used to execute the text parsing method according to any one of claims 1 to 7.