Tax guide system and method based on large language model
By applying a tax guidance system based on a large language model in the tax hall, combining digital twin technology and semantic understanding, the problems of low efficiency and insufficient knowledge in the tax processing process are solved, efficient and high-quality tax guidance services are achieved, and labor costs are reduced.
Patent Information
- Application Number
- CN202411943159.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-26
- Publication Date
- 2025-05-06
AI Technical Summary
In the existing lobby tax payment scenario, taxpayers need to manually guide tax due to poor business proficiency, resulting in low efficiency and high cost. The manual tax guidance officer cannot fully cover all tax knowledge, resulting in some difficult questions being unable to be answered in a timely manner.
The tax guidance system based on a large language model is adopted, combining digital twin technology, semantic understanding and emotional recognition to achieve more natural and accurate human-computer interaction, and provide taxpayers with high-quality and real-time tax guidance services.
It improves the efficiency and quality of tax consulting services, reduces the demand for manual tax guidance, reduces labor costs, and provides taxpayers with a better tax payment experience, enhancing the service level of the tax hall.
Smart Images

Figure CN119941419A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of artificial intelligence recognition, large language models and digital virtual humans, and in particular to a tax guidance system and method based on a large language model. Background Art
[0002] In the existing tax payment scenarios in tax halls, taxpayers have poor business proficiency, so they must rely on manual tax guides to answer relevant business questions, which results in low processing efficiency and high labor costs. In actual business handling, since the knowledge structure of manual tax guides cannot fully cover all tax knowledge, some difficult questions may not be answered in a timely manner. With the continuous promotion of the application of digital technologies, various regions have actively embraced new technologies and introduced new products in tax services, continuously innovated tax service models, realized the digitization and intelligence of tax hall-related businesses, and built a complete smart tax ecosystem.
[0003] In the field of tax consulting, traditional manual services have problems such as low efficiency and high cost. Although the existing automated tax guidance system can provide certain consulting services, it often lacks natural language processing capabilities and cannot accurately understand user intentions, resulting in poor user experience.
[0004] With the rapid development of artificial intelligence technology, large language models have made significant progress in the field of natural language processing, promoting the increasing maturity of human-computer interaction technology. Digital twin virtual human technology uses computer graphics, artificial intelligence and sensor technology to create virtual humans that are highly similar to human images and behaviors in the real world. Not only can it realistically reproduce humans visually, but it can also replace real people to achieve interaction and communication. However, in existing technologies, virtual human technology has not yet been efficiently applied to tax consulting services. Summary of the invention
[0005] In view of the shortcomings of the existing technology, the present invention aims to provide a tax guidance system and method based on a large language model, which can achieve more natural and accurate human-computer interaction through digital twin technology, semantic understanding and analysis, emotion recognition and other technologies, and provide high-quality and real-time tax guidance services in the tax field for payers and taxpayers in the process of handling tax-related businesses, thereby improving the efficiency and quality of tax consulting services. It can also replace manual tax guidance to a certain extent and reduce labor costs. Specifically, the present invention provides the following technical solutions:
[0006] In one aspect, the present invention provides a tax guidance system based on a large language model, the system comprising:
[0007] Data collection module, used to collect taxpayer image data and taxpayer question data;
[0008] The data processing module is used to clean, convert, analyze and store the image data and question data, and obtain the voice data set and the image data set; and perform identity recognition based on the image data set;
[0009] Large language model module, used to obtain corresponding answers based on speech datasets and perform taxpayer emotion recognition based on image datasets;
[0010] The digital middle platform module generates a digital twin and controls the expression of the digital twin based on the output of the large language model module.
[0011] A question-answering module, which is used to receive tax questions input by taxpayers and answers output by the large language model module, and send the answers to the output module;
[0012] The output module is used to output the answer based on the answer through the speaker and the twin digital human, and display the facial expression of the twin digital human.
[0013] Preferably, the data processing module performs identity recognition by face recognition based on the image data set: comparing the face feature values obtained by face recognition with the information in the taxpayer information database;
[0014] If the taxpayer has been authenticated by real name, the tax services that can be handled by the taxpayer will be displayed in the interactive interface.
[0015] Preferably, in the data processing module, speech preprocessing is performed on the question data, including noise reduction, pre-emphasis, framing and windowing, so as to obtain a speech data set suitable for acoustic model processing.
[0016] Preferably, the text sequence obtained after the acoustic model processing is corrected and optimized, including part-of-speech tagging, grammatical analysis and semantic understanding.
[0017] Preferably, the digital middle station uses a mesh method to represent the movements and expressions of the twin digital human, and the neural network structure adopts a convolutional neural network.
[0018] Preferably, in the action and expression representation, 51 shape values are used to represent facial actions, and the neural network outputs weights of the 51 shape values, and obtains facial expression and lip shape representation by combining the shape values with the weights.
[0019] Preferably, the digital middle platform is also used to generate body movements of the twin digital human;
[0020] The generation of the body movement is based on the rotation of the joint points, the rotation of the joint points is represented by quaternions, and the result is assigned to the twin digital human skeleton model.
[0021] Preferably, in taxpayer emotion recognition, emotions are judged based on taxpayer image data and voice data, expression data, and text data obtained from taxpayer question data through a deep learning network;
[0022] The deep learning network is a convolutional neural network, a recurrent neural network, or a combination thereof.
[0023] Preferably, the large language model module also performs intent recognition to identify key intents in taxpayer question data for quick response.
[0024] The system also includes a network module, which uses a wired network and wireless Wifi mode to connect the system to the Internet.
[0025] On the other hand, the present invention also provides a tax guidance method based on a large language model, which is applied to the system as described above, and comprises:
[0026] S1. After the system is started and initialized, the taxpayer performs identity authentication and identification;
[0027] S2. The taxpayer raises tax consultation questions by voice. The data collection module collects the sound data and intonation data in the voice, and converts the taxpayer's voice into text through the data processing module. The text and intonation data form a voice data set; the taxpayer's image data passes through the data processing module to obtain an image data set;
[0028] S3, the data processing module sends the speech data set and the image data set to the large language model module and obtains the answer, while performing semantic understanding and emotion recognition;
[0029] S4. The output module broadcasts the answer to the user through the digital twin with appropriate tone and expression based on the output results of the large language model module; if the taxpayer continues to ask questions, return to step S2.
[0030] Compared with the prior art, the present invention has at least the following beneficial effects:
[0031] 1. Integrating multiple advanced technologies, such as artificial intelligence, biometrics, speech recognition (ASR), text-to-speech synthesis (TTS), augmented reality (AR), virtual reality (VR), simulation, large models, cloud computing, etc., the product technology level is advanced and leading.
[0032] 2. AI tax guides can realize intelligent reception; accurately identify taxpayers through face recognition technology; quickly and accurately understand taxpayers' needs through voice or text interaction, answer tax policies and tax handling procedures in real time, and provide detailed business guidance to help taxpayers quickly find self-service tax handling terminals or handling windows; they can also recommend relevant services and policies according to taxpayers' needs. This will reduce taxpayers' waiting time and tax handling links, improve tax handling experience, improve tax handling efficiency and service satisfaction, and also save taxpayers time and energy.
[0033] 3. This solution effectively replaces manual tax guidance, reduces the staff configuration of the tax bureau, saves labor costs, provides better tax services for taxpayers, improves the service level of the tax hall, enhances the public's satisfaction with tax services, innovates the service model of the tax hall, and promotes the digitalization, intelligence and personalization of the tax hall. This is of great significance for building a smart tax ecosystem. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0035] Figure 1 A schematic diagram of the hardware architecture of an embodiment of the present invention;
[0036] Figure 2 A schematic diagram of a software system design according to an embodiment of the present invention;
[0037] Figure 3 is an interactive flow chart of an embodiment of the present invention;
[0038] Figure 4 A schematic diagram of a system framework of an embodiment of the present invention;
[0039] Figure 5 This is an actual product effect diagram of an embodiment of the present invention. DETAILED DESCRIPTION
[0040] The following will be combined with the figures in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0041] The present invention aims to provide a tax guidance system and method based on a large language model, which realizes a more natural and accurate human-computer interaction through digital twin technology, semantic understanding and analysis, emotion recognition and other technologies, and provides high-quality and real-time tax guidance services in the tax field for payers and taxpayers in tax affairs in the process of handling tax-related businesses, thereby improving the efficiency and quality of tax consulting services. To achieve the above purpose, the present invention provides a tax guidance system based on a large language model. In this embodiment, combined with Figure 4 As shown, the system includes:
[0042] 1. Data collection module: used to collect identity information or taxpayer question information data, using hardware collection and software processing. The hardware includes at least a microphone and a binocular camera.
[0043] Furthermore, when collecting data, when collecting tax-related corpus data, it is necessary to ensure that high-quality data is obtained to provide reliable data support for subsequent tasks such as deep computing, target detection, and behavior analysis.
[0044] The data collection module also includes user-free identity authentication. The binocular camera installed in the system can collect taxpayers' facial information in real time, and compare it with the taxpayer information database through feature values to obtain taxpayers' information, thereby providing taxpayers with personalized tax services or problem answers.
[0045] 2. Data processing module: As one of the core components of the system, the data processing module is responsible for cleaning, converting, analyzing and storing the raw data obtained by the data acquisition module in order to provide high-quality data support for subsequent applications and decision-making.
[0046] This module processes the collected face data (such as taxpayers' face data) and sound data to obtain a specific data set. The image data obtained by the binocular camera is converted into useful information through a series of preprocessing, feature extraction, image matching, depth calculation and data analysis steps to provide support for the application.
[0047] 3. Large language model module: used to call the large language model platform, obtain corresponding answers, and realize interactive question and answer;
[0048] In this embodiment, the large language model improves the performance of speech recognition and processing through technologies such as deep neural networks, self-attention mechanisms, end-to-end methods, pre-training and fine-tuning, multi-language and multi-task learning, etc.
[0049] 4. Question and answer module: used to receive tax questions input by users, send them to the large language model module, and receive answers generated by the large language model module;
[0050] In this embodiment, the question-answering module receives tax questions input by users and generates answers through a large language model. In this embodiment, the performance and user experience of the question-answering system are significantly improved by applying methods such as pre-trained language models, natural language processing technology, and multi-round dialogue management.
[0051] 5. Output module: used to output the generated answers to the user, broadcast the answers through the speaker, and realize the changes in facial expressions through the twin digital human.
[0052] In the process of result output, technologies such as image rendering, motion capture, expression capture, speech synthesis, natural language processing, virtual reality and augmented reality, and data visualization are used to provide realistic, vivid and intelligent digital human images and behaviors, and the generated answers are output to users through digital humans.
[0053] In this embodiment, combined with Figure 1 As shown, the hardware of this system uses a dedicated host, an external display, a touch screen, a microphone, a binocular camera, a speaker, and a network module to achieve basic recognition and interaction functions. In addition, the hardware also includes a power module to power the system.
[0054] The host, as the core of the system, completes the initialization and scheduling of other hardware components, and runs the various methods of this embodiment. The microphone is a multi-microphone system composed of a certain number of acoustic sensors arranged according to certain rules, which is used to collect taxpayers' voices. In addition, there are a series of front-end algorithms to complete audio signal processing and recognize corresponding text; the network module adopts a wired network + wireless Wifi mode to connect the device to the Internet; the touch screen receives the contact input signal and displays it through the display. When the graphic button on the screen is touched, the tactile feedback system on the screen can drive various external devices according to the pre-programmed program to realize various operations; the binocular camera can automatically recognize people according to their facial features, and can realize millisecond-level response of portrait detection, key point positioning and face recognition functions in multiple scenes; the speaker is used to read the virtual person's answer.
[0055] The above-mentioned hardware components are divided into five parts according to their functions, namely: data acquisition module, preprocessing module, large language model module, question and answer processing module and result output module. The corresponding methods in the invention are implemented in these modules.
[0056] In this embodiment, the binocular camera used is a visual system that simulates the working principle of human eyes. It obtains images from different perspectives through two cameras, obtains three-dimensional information of the face in real time, and then obtains the corresponding feature values through a specific face recognition algorithm. The invention device then transmits these feature values to the taxpayer information database of the tax bureau through network security, and determines whether it is a taxpayer with real-name authentication based on the comparison results sent back. If it is a non-real-name authenticated taxpayer, in the display interface, the digital twin digital person will actively welcome the taxpayer through voice. At this time, the taxpayer can communicate and communicate with the twin digital person face to face, just like communicating with a manual tax guide; if it is a real-name authenticated taxpayer, in the display interface, in addition to the digital twin digital person actively welcoming the taxpayer through voice, it will also display the tax business that can be handled related to the taxpayer. Therefore, in addition to communicating with the twin digital person, the taxpayer can also handle tax business related to himself.
[0057] When the system communicates with taxpayers, the data collection module will collect taxpayers' expressions in real time through binocular cameras and voices through microphones. The data processing module will then convert the voice into corresponding text and determine the taxpayer's tone of voice. For facial data, the data processing module will obtain the taxpayer's expression. The invented device inputs text based on the processed taxpayer's voice, submits the question to the big model platform through the big model module, and after obtaining the corresponding answer, the twin digital human will take the corresponding expression and tone of voice and play it to the taxpayer through the speaker.
[0058] Taxpayers can consult the digital twins through natural language communication about the latest tax preferential policies, tax laws or tax rates and other general tax issues. They can also learn about common business operation procedures such as the electronic tax bureau, the natural person electronic tax bureau, the social security payment client, and the self-service tax handling machine. They can also experience the process of simulated filling in the declaration form and simulated tax handling. If there are some specific tax handling issues or opinions and suggestions related to corporate taxation, the system provided in this embodiment can also provide help and guidance. If the taxpayer is a real-name authenticated taxpayer, in addition to the above functions, the digital twins can also push the tax-related information and pending business of the related enterprises, and will give reasonable opinions and suggestions. At the same time, they can also guide taxpayers to handle the corresponding tax-related business.
[0059] Combination Figure 2 As shown in the figure, in the data processing module, speech recognition mainly includes speech preprocessing, acoustic model training and post-processing, as follows:
[0060] Speech preprocessing: This step includes noise reduction, pre-emphasis, framing, windowing and other operations, the purpose of which is to convert the original audio signal into a feature sequence suitable for acoustic model processing.
[0061] Acoustic model training: The acoustic model is the core of ASR technology, which converts acoustic features into text sequences. This can be achieved through deep learning-based methods such as deep neural networks (DNN), long short-term memory networks (LSTM), and recurrent neural networks (RNN).
[0062] Post-processing: Correct and optimize the text sequence output by the acoustic model, including operations such as part-of-speech tagging, grammatical analysis, and semantic understanding to improve the accuracy and comprehensibility of ASR.
[0063] In addition, the implementation of speech recognition also includes some key technologies, such as VAD (Voice Activity Detection) technology, which is used to detect the effective speech part in the speech signal; feature extraction, such as MFCC features, which are used to convert sound signals into multi-dimensional vectors; acoustic model (AM), which converts acoustic features into phonemes; and language model (LM), which is used to adjust the recognition results to conform to language logic.
[0064] In terms of the generation of digital twin digital humans, the implementation of this embodiment is as follows:
[0065] Digital human motion and expression generation is provided by the digital middle platform to solve the problem of virtual digital human driving in products and terminals. Traditional motion production and motion capture technology solutions can no longer meet the needs of tens of billions of motion expressions. Through various different input situations, the required motion expressions and lip shapes can be generated intelligently and in real time. This is the multi-template real-time driving solution defined by the digital middle platform.
[0066] The digital middle platform uses artificial intelligence technology and high-quality, large-scale data sets to generate digital human body movement data, facial expression data, and lip shape data through AI, using audio sources, text source inputs, sensor inputs, semantic inputs, script inputs, etc. The digital middle platform provides open solutions to the industry in the form of API / SDK. Whether it is the Metaverse community, the newly added digital human App, or the offline large screen, VR / AR equipment, as long as there are digital humans, it can make driving simpler, more real-time, and lower cost.
[0067] Digital middle station core AI generation capability - action expression generation (Audio to Motion), based on machine learning neural network model. This method does not need to extract specific types of phonemes for data. Thanks to the powerful function fitting ability of neural networks, audio data can be directly mapped to expressions, actions and mouth shapes. For the scheme using neural networks, the most important thing is to choose a reasonable data representation and network structure. This embodiment uses the mesh method to represent actions and expressions, and the network structure uses a convolutional neural network. Convolutional neural networks have proven their powerful ability to extract features in the field of computer vision, and speech features can be processed by convolutional neural networks after being converted into spectra. They are used to process speech spectrum features, which better solves the problems of temporal information dependence and single generation mode. For expressions and mouth shapes, 51 shape values are used to represent facial movements, and the output of the neural network is the weight of these 51 shape values. By combining these shape values, facial expressions and mouth shapes can be reasonably represented. For body movements, the model outputs the rotation of the joint points. This embodiment uses quaternions to represent them. When used, the values output by the model are assigned to the character skeleton in a specific order.
[0068] The analysis and generation of multimodal stylized motion data relies on high-quality large-scale data sets. With the construction of a sophisticated and efficient deep learning network structure and the continuous input of training data, the generative network based on the transformer architecture can easily and efficiently complete the tasks from audio to motion, expression and lip shape.
[0069] In terms of emotion recognition, this embodiment judges a person's emotional state by analyzing voice, expression, text, etc. This embodiment is based on deep learning algorithms, such as convolutional neural networks (CNN) and recurrent neural networks (RNN) and their variants, such as long short-term memory networks (LSTM) and gated recurrent units (GRU). The implementation process of recognition artificial intelligence technology usually includes steps such as data collection, data preprocessing, feature extraction, model training, model evaluation and optimization. Large-scale labeled data is essential for training accurate and reliable recognition models.
[0070] The following describes the interaction method and interaction process of the system provided by the present invention in combination with a specific implementation example. Figure 3 As shown, the system operation and interaction process of this embodiment is as follows:
[0071] 1. The user is authenticated and identified through the system.
[0072] 2. Users ask tax consultation questions by voice. The microphone collects the sound and tone and converts it into text through the data processing module.
[0073] 3. The data processing module sends the converted text of the user's voice to the large language model module and obtains the answer. At the same time, it performs semantic understanding and emotion recognition based on the converted text and the obtained user image recognition results.
[0074] 4. The output module uses the digital twin to broadcast the answer to the user through the speaker with appropriate tone and expression based on the recognition results and the answer obtained.
[0075] 5. The system conducts multiple rounds of questions and answers to actively respond to and feedback user needs.
[0076] 6. The system automatically identifies the intention in the question and provides corresponding tax advice or solutions.
[0077] The system's interactive interface is as follows Figure 5 As shown, the interface mainly includes an interactive digital person, an interactive language and text display, and buttons for users to operate tax affairs.
[0078] A person skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium, and when the program is executed, it can include the processes of the embodiments of the above-mentioned methods. The storage medium can be a disk, an optical disk, a read-only memory (ROM) or a random access memory (RAM), etc.
[0079] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A tax guidance system based on a large language model, characterized in that: The system comprises: Data collection module, used to collect taxpayer image data and taxpayer question data; The data processing module is used to clean, convert, analyze and store the image data and question data, and obtain the voice data set and the image data set; and perform identity recognition based on the image data set; Large language model module, used to obtain corresponding answers based on speech datasets and perform taxpayer emotion recognition based on image datasets; The digital middle platform module generates a digital twin and controls the expression of the digital twin based on the output of the large language model module. A question-answering module, which is used to receive tax questions input by taxpayers and answers output by the large language model module, and send the answers to the output module; The output module is used to output the answer based on the answer through the speaker and the twin digital human, and display the facial expression of the twin digital human.
2. The system according to claim 1, characterized in that The data processing module performs identity recognition by face recognition based on the image data set: comparing the face feature values obtained by face recognition with the information in the taxpayer information database; If the taxpayer has been authenticated by real name, the tax services that can be handled by the taxpayer will be displayed in the interactive interface.
3. The system according to claim 1, characterized in that In the data processing module, speech preprocessing is performed on the question data, including noise reduction, pre-emphasis, framing and windowing, so as to obtain a speech data set suitable for acoustic model processing.
4. The system according to claim 3, characterized in that Correct and optimize the text sequence obtained after acoustic model processing, including part-of-speech tagging, grammatical analysis and semantic understanding.
5. The system according to claim 1, characterized in that The digital middle platform uses a mesh method to represent the movements and expressions of the twin digital human, and the neural network structure adopts a convolutional neural network.
6. The system according to claim 5, characterized in that In the action and expression representation, 51 shape values are used to represent facial actions, and the neural network outputs the weights of the 51 shape values, and obtains the facial expression and lip shape representation by combining the shape values with the weights.
7. The system according to claim 5, characterized in that The digital middle platform is also used to generate body movements of the twin digital human; The generation of the body movement is based on the rotation of the joint points, the rotation of the joint points is represented by quaternions, and the result is assigned to the twin digital human skeleton model.
8. The system according to claim 1, characterized in that In taxpayer emotion recognition, emotions are judged based on taxpayer image data and voice data, expression data, and text data obtained from taxpayer question data through a deep learning network; The deep learning network is a convolutional neural network, a recurrent neural network, or a combination thereof.
9. The system according to claim 1, characterized in that The large language model module also performs intent recognition to identify key intents in taxpayer question data for quick response.
10. A tax guidance method based on a large language model, characterized in that: The method is applied to the system according to any one of claims 1 to 9, and the method comprises: S1. After the system is started and initialized, the taxpayer performs identity authentication and identification; S2. The taxpayer raises tax consultation questions by voice. The data collection module collects the sound data and intonation data in the voice, and converts the taxpayer's voice into text through the data processing module. The text and intonation data form a voice data set; the taxpayer's image data passes through the data processing module to obtain an image data set; S3, the data processing module sends the speech data set and the image data set to the large language model module and obtains the answer, while performing semantic understanding and emotion recognition; S4. The output module broadcasts the answer to the user through the digital twin with appropriate tone and expression based on the output results of the large language model module; if the taxpayer continues to ask questions, return to step S2.