Loan accounting method and device based on natural language processing, equipment and medium
By segmenting and mapping corporate financial texts based on natural language processing methods and combining them with the principle of loan identity, automatic accounting is achieved, solving the problems of low corporate accounting efficiency and high professional requirements, improving efficiency and saving labor costs.
Patent Information
- Application Number
- CN202210734687.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-27
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2042-06-27
AI Technical Summary
The current corporate accounting process is inefficient and requires high professional qualities from accounting personnel, who need to be familiar with finance and business.
A natural language processing-based method is used to segment the text to be recorded, map it to a pre-trained part-of-speech set, verify the balance of credit and debit based on the credit and debit identity principle, and enter the verified financial information into the database.
It realizes automatic accounting, improves accounting efficiency, reduces the requirements for the professional quality of accounting personnel, and significantly saves labor costs.
Smart Images

Figure CN115034891B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and specifically to a debit and credit accounting method, device, equipment and medium based on natural language processing. Background Art
[0002] The debit-and-credit accounting method is a bookkeeping method that uses the accounting equation as the bookkeeping principle and debit and credit as the bookkeeping symbols to reflect the increase and decrease of economic activities. When a certain economic activity occurs, the three elements of "debit and credit direction, subject, and amount" must be extracted to realize the debit-and-credit accounting. Currently, the debit-and-credit accounting method is widely used within enterprises.
[0003] Currently, enterprises have the following pain points in the bookkeeping process: the entire bookkeeping process is manually operated and inefficient; and bookkeeping is performed by professional accountants who must be familiar with both finance and business, which places high demands on bookkeepers. Summary of the Invention
[0004] In response to the above situation, the embodiments of the present application propose a debit and credit accounting method, device, equipment and medium based on natural language processing, which realizes automatic accounting, not only greatly improves accounting efficiency, but also saves labor costs.
[0005] In a first aspect, an embodiment of the present application provides a debit and credit accounting method based on natural language processing, the method comprising:
[0006] Perform word segmentation on the incoming text to obtain multiple keywords;
[0007] Mapping the obtained multiple keywords to a part-of-speech set to obtain a part-of-speech mapping;
[0008] Based on the principle of loan identity and according to the part-of-speech mapping, verify the loan balance of the text to be recorded;
[0009] The financial information in the to-be-entered text that has passed the verification result will be entered into the database.
[0010] In a second aspect, an embodiment of the present application further provides a debit and credit accounting device based on natural language processing, the device comprising:
[0011] The word segmentation unit is used to segment the text to be entered into the account and obtain multiple keywords;
[0012] A mapping unit, configured to map the obtained multiple keywords to a part-of-speech set to obtain a part-of-speech mapping list;
[0013] a verification unit, configured to verify the debit / credit balance of the to-be-entered text based on the debit / credit identity principle and the part-of-speech mapping list;
[0014] The storage unit is used to store the financial information in the verified documents to be recorded.
[0015] In a third aspect, an embodiment of the present application further provides an electronic device, comprising: a processor; and a memory arranged to store computer-executable instructions, wherein the executable instructions, when executed, enable the processor to perform any of the above methods.
[0016] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, which stores one or more programs. When the one or more programs are executed by an electronic device including multiple applications, the electronic device executes any of the above methods.
[0017] At least one of the above technical solutions adopted in the embodiments of the present application can achieve the following beneficial effects:
[0018] This application is based on natural language processing technology. It splits the text to be recorded and maps the obtained keywords to the part-of-speech set obtained by pre-training to obtain a part-of-speech mapping. Based on the principle of loan identity, the loan balance of the text to be recorded is verified according to the obtained part-of-speech mapping, and the text to be recorded that passes the verification is entered into the warehouse. By flexibly applying natural language processing technology, this application can quickly and accurately identify the key financial information of the text to be recorded and perform fast and accurate accounting. It not only realizes the automatic accounting of loan business and greatly improves the accounting efficiency of the enterprise; it also reduces the professional threshold requirements for bookkeepers and significantly saves labor costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0020] Figure 1 A schematic diagram illustrating a flow chart of a natural language processing-based debit and credit accounting method according to an embodiment of the present application is shown;
[0021] Figure 2 A schematic diagram illustrating a flow chart of a natural language processing-based debit and credit accounting method according to another embodiment of the present application is shown;
[0022] Figure 3 A schematic diagram showing the structure of a debit and credit accounting device based on natural language processing according to an embodiment of the present application is shown;
[0023] Figure 4 This is a structural diagram of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION
[0024] To make the purpose, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the specific embodiments of this application and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0025] The following describes in detail the technical solutions provided by various embodiments of the present application in conjunction with the accompanying drawings.
[0026] The double-entry bookkeeping method uses the accounting equation as its accounting principle and debit and credit as accounting symbols to reflect changes in economic transactions. When a certain economic transaction occurs, we need to extract the three elements of "debit and credit direction, account, and amount" to implement debit and credit accounting. Currently, the double-entry bookkeeping method is widely used within enterprises. The specific accounting process is as follows: identify the debit and credit direction of the current transaction; find the company financial account corresponding to the current transaction; enter the specific amount and verify the debit and credit balance; and store the financial data in the database. Currently, enterprises generally use manual bookkeeping, which is not only inefficient but also requires professional accountants to operate. They must be familiar with finance and understand the business, and the professional qualifications of accountants are high.
[0027] The idea of this application is to address the current situation where most companies have low bookkeeping efficiency and high personnel requirements. This application proposes a fast debit and credit bookkeeping method based on NLP (natural language processing). This method is based on natural language processing and combines big data and artificial intelligence to efficiently process text, thereby realizing automatic bookkeeping. Specifically, by splitting the text to be recorded and mapping the obtained keywords to the part-of-speech set obtained by pre-training, a part-of-speech mapping is obtained. Based on the principle of loan identity, the debit and credit balance of the text to be recorded is verified according to the obtained part-of-speech mapping, and the text to be recorded with a passed verification result is stored in the warehouse, thereby realizing automatic, fast and accurate bookkeeping.
[0028] Figure 1 A flow chart of a natural language processing-based debit and credit accounting method according to an embodiment of the present application is shown. Figure 1 It can be seen that this application includes at least steps S110 to S140:
[0029] Step S110: performing word segmentation processing on the text to be recorded to obtain multiple keywords.
[0030] In this application, the company's bookkeepers do not need to have a professional financial background. They only need to input the company's financial behavior directly into the bookkeeping terminal through text or voice. After the server receives the information, it can quickly identify key financial information through the method of this application, thereby achieving the purpose of fast bookkeeping.
[0031] Financial behavior is usually presented in the form of text to be recorded. It should be noted that the original form of the text to be recorded can be text or voice. If it is text, it can be processed directly. If it is voice, the voice can be recognized as text and then processed. The recognition process can adopt any voice recognition technology in the existing technology, and this application does not limit this.
[0032] Take an enterprise's accounting scenario as an example. In this scenario, the enterprise purchases a piece of production equipment for 1,000 yuan, and the payment is made to the bank inventory. First, the text to be recorded is segmented to obtain multiple keywords. In this scenario, the text to be recorded can be, but is not limited to: The enterprise purchases a piece of production equipment for 1,000 yuan, and the payment is made to the bank inventory. By segmenting the text to be recorded, the multiple keywords obtained can be: enterprise, purchase, production equipment, one, price, 1,000 yuan, payment, bank inventory, payment. In other embodiments, the keywords can be further screened to remove some function words that do not contain actual meaning. This application does not limit this, and selection can be made as needed.
[0033] This application does not limit the word segmentation method. Word segmentation can adopt a dictionary-based rule matching method or a statistics-based machine learning method.
[0034] Dictionary-based word segmentation algorithms are essentially string matching. They use a specific algorithm strategy to match the string to be matched against a sufficiently large dictionary. If a match is found, word segmentation is performed. Different matching strategies are available, including forward maximum matching, reverse maximum matching, bidirectional matching, and full segmentation.
[0035] The statistical word segmentation algorithm is essentially a sequence labeling problem, which labels the words in the sentence according to their position in the word. The main labels are: B (a word at the beginning of a word), E (the last word of a word), M (a word in the middle of a word, there may be multiple), and S (a word represented by a word). For example, "MyBank is the most important product of Ant Financial's Micro Loan Division" is labeled as "BMMESBMMEBMMMESBMEBE", and the corresponding word segmentation result is "MyBank is the most important product of Ant Financial's Micro Loan Division". This type of algorithm is based on machine learning or deep learning, and mainly includes HMM (hidden Markov model), CRF (conditional random field), SVM (support vector machine), and deep learning.
[0036] Step S120: Map the obtained multiple keywords to a part-of-speech set to obtain a part-of-speech mapping.
[0037] The part-of-speech set is pre-trained based on debit and credit bookkeeping scenarios and contains many key words, each representing a part of speech. In some embodiments of this application, these key words include, but are not limited to, debit action, credit action, debit account, credit account, debit amount, and credit amount. Each key word contains a corresponding vocabulary set. Table 1 shows a part-of-speech set according to one embodiment of this application.
[0038] Table 1:
[0039]
[0040] As can be seen from Table 1, in this embodiment, the subject terms include debit behavior, credit behavior, debit account, credit account, debit amount, and credit amount. Each subject term contains a corresponding vocabulary set, and each vocabulary set has similar semantics. Taking debit behavior as an example, the vocabulary set includes words such as "buy", "buy", "purchase", and "procure", all of which represent the act of purchasing.
[0041] The obtained multiple keywords are mapped to the part-of-speech set to obtain the part-of-speech mapping, which can be simply understood as compiling the key financial information in the obtained multiple keywords into a table with the subject words of the part-of-speech set as the table header.
[0042] In other words, not every keyword needs to be filled in, only the key financial information needs to be filled in.
[0043] Table 2 shows a part-of-speech mapping according to an embodiment of the present application. As can be seen from Table 2, the key financial information includes the above-mentioned keywords such as purchase, payment, production equipment, bank inventory, 1,000 yuan, and 1,000 yuan. By matching these key financial information with the subject words of the part-of-speech set, a part-of-speech mapping is formed.
[0044] Table 2:
[0045]
[0046] Step S130: Based on the principle of credit and debit identity and according to the part-of-speech mapping, verify the credit and debit balance of the text to be recorded.
[0047] The borrowing and lending identity principle, i.e., the borrowing and lending identity, can be understood as that there is lending for borrowing and the balance of borrowing and lending. Based on the principle, the balance of borrowing and lending of the to-be-accounted text is verified according to the part-of-speech mapping. Only the financial information involved in the to-be-accounted text that passes the verification is subjected to the accounting operation.
[0048] According to the borrowing and lending identity principle, the verification of the balance of borrowing and lending mainly includes two aspects, one is the behavior verification, and the other is the amount verification. The behavior verification is to verify that there is lending for borrowing, and the amount verification is to verify that the balance of borrowing and lending. In this embodiment, it can be seen that there are behaviors of purchasing and paying, i.e., there is lending for borrowing, and it can also be seen that the amount of the borrower and the amount of the lender are both 1000 yuan, i.e., the balance of borrowing and lending. Therefore, in this embodiment, the verification result of the balance of borrowing and lending of the to-be-accounted text is passed.
[0049] Step S140: performing the warehousing operation on the accounting information in the to-be-accounted text whose verification result is passed.
[0050] Finally, the accounting information involved in the to-be-accounted text whose verification result is passed is subjected to the warehousing operation, and for the to-be-accounted text whose verification result is not passed, a prompt information can be displayed to remind the accounting personnel to modify or check the accounting information involved in the to-be-accounted text.
[0051] As can be seen from the method shown in Figure 1 Based on the natural language processing technology, the to-be-accounted text is split, the obtained keywords are mapped to the part-of-speech set obtained by prior training, the part-of-speech mapping is obtained, the balance of borrowing and lending of the to-be-accounted text is verified based on the borrowing and lending identity principle, and the to-be-accounted text whose verification result is passed is subjected to the warehousing operation. By flexibly using the natural language processing technology, the application can quickly and accurately identify the key financial information of the to-be-accounted text, perform fast and accurate accounting, realize automatic accounting of the borrowing and lending business, greatly improve the accounting efficiency of the enterprise, reduce the professional threshold requirement of the accounting personnel, and significantly save the labor cost.
[0052] In some embodiments of the application, in the above method, the part-of-speech set is obtained according to the following method: obtaining a training sample set, the training sample set including a plurality of borrowing and lending training texts; identifying the training sample set based on a word vector method to obtain a plurality of word set groups, each word set group containing a plurality of words with similar semantics; and attributing each word set group to a preset topic word to obtain the part-of-speech set.
[0053] In some embodiments of the present application, keywords can be set according to the purpose and needs of debit and credit accounting, such as the debit behavior, credit behavior, debit account, credit account, debit amount, credit amount, etc. mentioned above, but this application does not limit the setting of keywords, and they can be modified, added or reduced according to needs.
[0054] Part-of-speech sets can be obtained through machine learning. Specifically, first, a training sample set is collected. The training sample set includes a large number of training texts, and these training texts are all in the field of lending. Then, based on the word vectors in natural language processing, these massive training samples are identified and analyzed to obtain multiple groups of semantically similar vocabulary sets. Each vocabulary set contains multiple words. As mentioned above, purchase, buy, purchase, and purchase are a group of vocabulary sets; payment, spending, expenditure, and consumption are a group of vocabulary sets. These vocabulary sets are then classified under a semantically corresponding preset subject word, and a part-of-speech set as shown in Table 1 can be obtained. It should be noted that Table 1 is only an exemplary description and does not constitute any limitation to this application. The part-of-speech set is not limited to the form of a table. It can be in the form of a document or a regular expression, and this application does not limit this.
[0055] In some embodiments of the present application, in the above method, the obtained multiple keywords are mapped to a part-of-speech set to obtain a part-of-speech mapping, including: determining a target vocabulary set for a keyword; mapping the keyword to a subject word corresponding to the target vocabulary set; and looping the steps of determining a target vocabulary set for a keyword and mapping the keyword to a subject word corresponding to the target vocabulary set to obtain a part-of-speech mapping.
[0056] When assigning part-of-speech sets to multiple keywords, these keywords can be traversed. For a keyword, first determine its corresponding target vocabulary set, then map the keyword to the subject word corresponding to the target vocabulary set, and then assign the next keyword, and execute the loop until all keywords containing key financial information are assigned. In some embodiments of the present application, determining the target vocabulary set of a keyword includes: determining whether a keyword belongs to a certain vocabulary set, and if so, using the vocabulary set as the target vocabulary set for the keyword; if it is determined that a keyword does not belong to all vocabulary sets, then determining the sum of the distances between the keyword and each word in each vocabulary set, and using the vocabulary set with the smallest sum of distances as the target vocabulary set for the keyword.
[0057] That is, for a keyword, first determine whether there is a word in a certain set of vocabulary that is exactly the same as the keyword, if there is, then the set of vocabulary is the target set of vocabulary of the keyword, and the keyword is subsequently attributed to the subject word corresponding to the target set of vocabulary. Taking the above Table 1 as the set of word types, and taking the keyword as purchase as an example, under the subject word debit behavior in Table 1, there are purchase, buy, buy, purchase and purchase, which contain the word “purchase”, that is, the keyword “purchase” belongs to the set of vocabulary “purchase, buy, buy, purchase”, so determine “purchase, buy, buy, purchase” as the target set of vocabulary of “purchase”, and subsequently attribute the keyword “purchase” to the subject word “debit behavior” corresponding to the target set of vocabulary.
[0058] If it is determined that a keyword does not belong to all sets of vocabulary, the idea of clustering can be used to determine the distance sum of the keyword and each word in each set of vocabulary, and the set of vocabulary with the smallest distance sum is determined as the target set of vocabulary of the keyword. Each word can be regarded as a point on a plane, and words with similar semantics will gather together to form a pile (the distance between words is small), so the semantic similarity of a word to the words in each set of vocabulary can be determined by calculating the distance.
[0059] If a word has the smallest distance sum with the words in a set of vocabulary, it means that the semantic similarity of the keyword to the words in the set of vocabulary is the closest, and the set of vocabulary can be determined as the target set of vocabulary of the keyword, and the keyword is subsequently attributed to the subject word corresponding to the target set of vocabulary. Still taking the above Table 1 as the set of word types, and taking the keyword as bought as an example, the distance of the keyword “bought” to the six subject words is L1, L2, L3, L4, L5 and L6, respectively. Determine which one has the smallest value, assume it is L1, then determine the target set of vocabulary of “bought” as “purchase, buy, buy, purchase”, and subsequently attribute the keyword “bought” to the subject word “debit behavior” corresponding to the target set of vocabulary. The distance can be absolute distance, Euclidean distance, etc., which is not limited in the present application.
[0060] In some embodiments of the present application, in the above method, the balance of debit and credit of the text to be accounted for is verified based on the debit and credit identity principle according to the word type mapping, which comprises: generating debit and credit data according to the word type mapping; determining whether the behavior direction of the debit and credit data is opposite and the debit and credit amount is equal according to the debit and credit identity principle; if both are, then the verification result is passed, otherwise, the verification result is failed.
[0061] In some embodiments, in order to verify conveniently, the part-of-speech mapping can also be further separated into debit financial data and credit financial data, so that the data is more clear. Then according to the debit-credit identity principle, the balance is verified. Taking the part-of-speech mapping shown in Table 2 as an example, the debit financial data and the credit financial data obtained after separation are shown in Table 3-a and Table 3-b, respectively.
[0062] Table 3-a:
[0063] Debit behavior Debit Account Debit amount Purchase production equipment 1,000 yuan
[0064] Table 3-b:
[0065] Lender behavior Lender behavior Credit Amount Payment Bank inventory 1,000 yuan
[0066] From Table 3-a and Table 3-b, it can be seen that the debit financial data and the credit financial data are obtained.
[0067] The verification of balance mainly includes two aspects of verification. One is the opposite behavior, i.e. one party credits and the other party debits. The other is the amount of money, i.e. the amount of money of the debit side and the amount of money of the credit side need to be equal. If both of them are satisfied, it is determined that the verification result is passed. If one or both of them are not satisfied, it is determined that the verification result is not passed.
[0068] Figure 2 The flowchart of the natural language processing-based debit-credit accounting method according to another embodiment of the present application is shown, from Figure 2 It can be seen that the present embodiment includes:
[0069] Obtaining the to-be-credited text and the part-of-speech set.
[0070] Carrying out word segmentation on the to-be-credited text to obtain a plurality of keywords.
[0071] Traversing the keywords to determine whether a keyword belongs to a certain set of vocabulary. If it belongs, the set of vocabulary is taken as the target set of vocabulary of the keyword. If it does not belong, the distance sum of the keyword and each word in each set of vocabulary is determined. The set of vocabulary with the smallest distance sum is taken as the target set of vocabulary of the keyword. The keyword is mapped to the subject word corresponding to the target set of vocabulary. Then the next keyword is processed, and the process is repeated for multiple times to obtain the part-of-speech mapping.
[0072] Separating the part-of-speech mapping to obtain the debit financial data and the credit financial data. It is judged whether the behavior direction of the debit financial data and the credit financial data is opposite and whether the debit amount and the credit amount are equal. If yes, the financial information of the to-be-credited text is stored in the database. If not, a prompt information is sent out.
[0073] Figure 3A schematic diagram of a structure of a debit and credit accounting device based on natural language processing according to an embodiment of the present application is shown. Figure 3 It can be seen that the debit and credit accounting device 300 based on natural language processing includes:
[0074] The word segmentation unit 310 is used to perform word segmentation processing on the text to be entered to obtain multiple keywords;
[0075] A mapping unit 320 is configured to map the obtained multiple keywords to a part-of-speech set to obtain a part-of-speech mapping list;
[0076] A verification unit 330 is configured to verify the debit / credit balance of the to-be-entered text based on the debit / credit identity principle and the part-of-speech mapping list;
[0077] The storage unit 340 is used to store the financial information in the verified document to be recorded.
[0078] In some embodiments of the present application, in the above-mentioned device, the part-of-speech set is obtained according to the following method: obtaining a training sample set, the training sample set including multiple loan training texts; based on the word vector method, identifying the training sample set to obtain multiple groups of vocabulary sets, each group of vocabulary sets containing multiple words with similar semantics; assigning each group of vocabulary sets to a preset keyword with corresponding semantics to obtain a part-of-speech set.
[0079] In some embodiments of the present application, in the above-mentioned device, the subject words include: debit behavior, credit behavior, debit account, credit account, debit amount, credit amount.
[0080] In some embodiments of the present application, in the above apparatus, the mapping unit 320 is configured to determine a target vocabulary set for a keyword; map the keyword to a subject word corresponding to the target vocabulary set;
[0081] The steps of determining a target vocabulary set for a keyword and mapping the keyword to a subject word corresponding to the target vocabulary set are performed cyclically to obtain a part-of-speech mapping.
[0082] In some embodiments of the present application, in the above apparatus, the mapping unit 320 is configured to determine whether a keyword belongs to a certain vocabulary set, and if so, use the vocabulary set as the target vocabulary set for the keyword.
[0083] In some embodiments of the present application, in the above-mentioned device, the mapping unit 320 is used to determine the sum of the distances between the keyword and each word in each vocabulary set if it is determined that a keyword does not belong to all groups of vocabulary sets, and to take the vocabulary set with the smallest sum of distances as the target vocabulary set for the keyword.
[0084] In some embodiments of the present application, the verification unit 330 in the above-mentioned device is used to generate debit account data and credit financial data based on the part-of-speech mapping; according to the principle of debit and credit identity, determine whether the behavior directions of the debit account data and the credit financial data are opposite, and whether the debit and credit amounts are equal; if both are true, the verification result is determined to be passed, otherwise, the verification result is determined to be failed.
[0085] It should be noted that the above-mentioned debit and credit accounting device based on natural language processing can implement the above-mentioned debit and credit accounting methods based on natural language processing one by one, and will not be repeated here.
[0086] Figure 4 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present application. Figure 4 At the hardware level, the electronic device includes a processor and, optionally, an internal bus, a network interface, and memory. The memory may include internal memory, such as high-speed random-access memory (RAM), or non-volatile memory, such as at least one disk drive. Of course, the electronic device may also include other hardware required for its services.
[0087] The processor, network interface, and memory can be interconnected via an internal bus, which can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 4 Only one bidirectional arrow is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0088] The memory is used to store programs. Specifically, the program may include program code, which includes computer operating instructions. The memory may include internal memory and non-volatile memory, and provides instructions and data to the processor.
[0089] The processor reads the corresponding computer program from the non-volatile memory into the internal memory and then runs it, forming a debit and credit accounting device based on natural language processing at the logical level. The processor executes the program stored in the memory and is specifically used to perform the aforementioned method.
[0090] The above application Figure 3The methods performed by the natural language processing-based debit and credit accounting device disclosed in the illustrated embodiments can be applied to or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be performed by hardware integrated logic circuits in the processor or by software instructions. The above processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The methods, steps, and logic block diagrams disclosed in the embodiments of this application can be implemented or executed. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method.
[0091] The electronic device may also perform Figure 3 A method for executing a debit and credit accounting device based on natural language processing in Figure 3 The functions of the illustrated embodiment will not be described in detail in the embodiments of the present application.
[0092] The embodiment of the present application also provides a computer-readable storage medium, which stores one or more programs, wherein the one or more programs include instructions, which, when executed by an electronic device including multiple application programs, can enable the electronic device to execute Figure 3 The method executed by the debit and credit accounting device based on natural language processing in the illustrated embodiment is specifically used to execute the aforementioned method.
[0093] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0094] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0095] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0096] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing the instructions executed on the computer or other programmable device for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0097] In a typical configuration, a computing device includes one or more processors (CPUs), input / output interfaces, network interfaces, and memory.
[0098] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0099] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0100] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0101] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Furthermore, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0102] The foregoing is merely an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.
Claims
1. A debit and credit accounting method based on natural language processing, characterized in that: The method comprises: Performing word segmentation on the text to be entered to obtain multiple keywords; wherein the word segmentation includes a dictionary-based rule matching method or a statistics-based machine learning method; Mapping the obtained multiple keywords to a part-of-speech set to obtain a part-of-speech mapping; wherein the part-of-speech set includes multiple subject words, each subject word includes a class of vocabulary sets corresponding to the subject word, and the semantics of each class of vocabulary sets are similar; Based on the principle of loan identity and according to the part-of-speech mapping, verify the loan balance of the text to be recorded; The financial information in the to-be-entered document that has passed the verification result will be stored in the database; The step of mapping the obtained multiple keywords to a part-of-speech set to obtain a part-of-speech mapping includes: Determine a target vocabulary set for keywords; Mapping the keyword to the subject word corresponding to the target vocabulary set; The steps of determining a target vocabulary set for a keyword and mapping the keyword to a subject word corresponding to the target vocabulary set are executed cyclically to obtain a part-of-speech mapping; The step of determining a target vocabulary set for a keyword includes: Determine whether a keyword belongs to a certain vocabulary set, and if so, use the vocabulary set as the target vocabulary set for the keyword; If it is determined that a keyword does not belong to all vocabulary sets, the sum of the distances between the keyword and each word in each vocabulary set is determined, and the vocabulary set with the smallest sum of the distances is used as the target vocabulary set for the keyword.
2. The method according to claim 1, characterized in that The part-of-speech set is obtained according to the following method: Obtaining a training sample set, wherein the training sample set includes a plurality of loan training texts; Based on the word vector method, the training sample set is identified to obtain multiple groups of vocabulary sets, each of which contains multiple words with similar semantics; Each vocabulary set is assigned to a preset keyword with corresponding semantics to obtain a part-of-speech set.
3. The method according to any one of claims 2, characterized in that The subject terms include: debit behavior, credit behavior, debit account, credit account, debit amount, and credit amount.
4. The method according to claim 1, wherein The verification of the debit and credit balance of the to-be-entered text based on the debit and credit identity principle and according to the part-of-speech mapping includes: Generate debit account data and credit financial data according to the part-of-speech mapping; According to the principle of debit-credit identity, determine whether the debit account data and the credit financial data have opposite behavior directions and whether the debit and credit amounts are equal; If both are true, the verification result is determined to be passed; otherwise, the verification result is determined to be failed.
5. A debit and credit accounting device based on natural language processing, characterized in that: The device comprises: A word segmentation unit is used to perform word segmentation processing on the text to be recorded to obtain multiple keywords; wherein the word segmentation processing includes a dictionary-based rule matching method or a statistics-based machine learning method; A mapping unit is used to map the obtained multiple keywords to a part-of-speech set to obtain a part-of-speech mapping list; wherein the part-of-speech set includes multiple subject words, each subject word includes a class of vocabulary sets corresponding to the subject word, and the semantics of each class of vocabulary sets are similar; a verification unit, configured to verify the debit / credit balance of the to-be-entered text based on the debit / credit identity principle and the part-of-speech mapping list; The storage unit is used to store the financial information in the verified documents to be recorded; The mapping unit is used to determine a target vocabulary set for a keyword; Mapping the keyword to the subject word corresponding to the target vocabulary set; The steps of determining a target vocabulary set for a keyword and mapping the keyword to a subject word corresponding to the target vocabulary set are executed cyclically to obtain a part-of-speech mapping; The mapping unit is further configured to determine whether a keyword belongs to a certain vocabulary set, and if so, use the vocabulary set as the target vocabulary set for the keyword; If it is determined that a keyword does not belong to all vocabulary sets, the sum of the distances between the keyword and each word in each vocabulary set is determined, and the vocabulary set with the smallest sum of the distances is used as the target vocabulary set for the keyword.
6. An electronic device comprising: processor; as well as A memory arranged to store computer executable instructions, which when executed cause the processor to perform the method according to any one of claims 1 to 4.
7. A computer-readable storage medium storing one or more programs, which, when executed by an electronic device including a plurality of application programs, causes the electronic device to execute the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Accounting rule-driven scientific research management economic activity interpretation and identification method and device
CN114358699A
Business transaction bookkeeping method for improving bookkeeping accuracy
CN114612092A