Coal logistics park content generation customer service system based on large model
By introducing large-model-based technology into the intelligent customer service system, the limitations and high maintenance costs of existing systems when dealing with complex user problems are solved, and more efficient and accurate understanding and answering of user problems are achieved, improving user experience and system quality.
Patent Information
- Application Number
- CN202510135406.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-07
- Publication Date
- 2025-05-16
AI Technical Summary
The existing intelligent customer service system shows limitations when dealing with complex and diverse user problems, and requires continuous maintenance and update of the rule database. The maintenance cost is high, and lacks deep semantic understanding capabilities, so it is unable to truly understand user problems.
A large-scale model-based coal logistics park content generation customer service system is used, including a data collection module, a coal logistics industry knowledge base, a customer service question matching estimate module and a question-and-answer generation large model, and quickly understand user questions and generate accurate answers through large-scale model technology.
It significantly improves customer service response speed and quality, can understand complex and diverse questions and provide accurate responses, reduces maintenance costs, and provides personalized services to improve user experience and satisfaction.
Smart Images

Figure CN120012937A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of customer service systems, and in particular relates to a customer service system for coal logistics park content generation based on a large model. Background Art
[0002] At present, as an important link in the coal supply chain, the customer service quality of the coal logistics park directly affects the park's operational efficiency and customer satisfaction, and the efficiency and quality of its customer service system directly affect the overall operational effect of the park. However, traditional customer service systems mostly rely on manual operations, and have problems such as high cost, slow response speed, low processing efficiency, inaccurate information, and inability to provide personalized services. With the rapid development of artificial intelligence technology, especially the widespread application of large model technology in natural language processing and data analysis, it has provided new possibilities for upgrading the customer service system of coal logistics parks.
[0003] The existing technology mainly has the following problems:
[0004] 1. In existing intelligent customer service systems, a rule-based approach is usually used to answer user questions. This approach requires manual writing of a large number of rules and templates to handle various possible user queries. It has limitations when facing complex and diverse user questions, and requires constant maintenance and updating of the rule base, which has a relatively high maintenance cost.
[0005] 2. Existing intelligent customer service systems have limited ability to understand user intent and provide accurate answers. Although some systems use natural language processing technology and machine learning algorithms to improve performance, in actual applications, the questions asked by users are often diverse. Traditional systems lack deep semantic understanding capabilities, resulting in an inability to truly understand user questions and can only provide standard answers based on surface information.
[0006] 3. Existing systems often lack the support of knowledge bases. A knowledge base is a database that stores a large number of question and answer pairs, which can help the system better answer users' questions. However, the knowledge base construction and maintenance process of many existing systems is relatively difficult, resulting in limited quality and practicality of the knowledge base. Summary of the invention
[0007] The purpose of the present invention is to provide a coal logistics park content generation customer service system based on a large model, aiming to solve the problem that in the existing intelligent customer service system, a rule-based method is usually used to answer user questions, requiring manual writing of a large number of rules and templates to handle various possible user queries, showing limitations when facing complex and diverse user questions, requiring continuous maintenance and updating of the rule base, and having relatively high maintenance costs.
[0008] The present invention is implemented as follows: a coal logistics park content generation customer service system based on a large model, the system comprising:
[0009] Data collection module, used to obtain and process the problem data input by the user;
[0010] The coal logistics industry knowledge base is used to store industry knowledge data of the coal logistics industry;
[0011] The customer service question matching degree estimation module is used to calculate the question matching degree index based on the question data and determine the user's question;
[0012] A large model for question-answer generation is used to query the coal logistics industry knowledge base to obtain answers based on user questions or generate answers through a large language model. The large model for question-answer generation includes three parts: front-end, AI service, and back-end.
[0013] The customer service interface module is used to present answers to customers in a variety of ways.
[0014] Preferably, the data collection module includes a customer service problem collection unit and a customer service problem processing unit. The customer service problem collection unit is used to collect text word vector parameters, keyword parameters and character count parameters of questions raised by customers to the intelligent customer service. The customer service problem processing unit is used to clean, deduplicate and annotate the collected data, and perform data preprocessing through word segmentation, removal of stop words and part-of-speech tagging.
[0015] Preferably, the data stored in the coal logistics industry knowledge base includes basic coal knowledge data, coal logistics process data, market dynamics data, technical equipment data, policy and regulation data, and safety management data.
[0016] Preferably, when training a large question-answer generation model, a pre-training strategy is adopted. Pre-training is performed on a large-scale text dataset to learn the general representation and knowledge of the language. Based on the pre-training, fine-tuning is performed on a specific question-answer generation task dataset. During fine-tuning, the learning rate and optimizer are adjusted, and the training dataset is expanded through data enhancement. Regularization strategies are used during training to prevent model overfitting. For question-answer generation tasks involving multimodal information, multimodal fusion is used for representation.
[0017] Preferably, the question matching index is calculated based on the question data, and the process of determining the user question includes calculating text similarity, keyword difference coefficient and sentence length difference coefficient.
[0018] Preferably, the text similarity is characterized by word frequency vector similarity. The word frequency vector is used to represent the text, and the cosine similarity of two texts is calculated. The text similarity TF Similarity is expressed as:
[0019]
[0020] Among them, TF(w i ,d j ) is the word w i In the document d j The word frequency in , i is the word number, j is the document number, and n is the number of words.
[0021] Preferably, the text similarity is characterized by TF-IDF similarity. The TF-IDF vector is used to represent the text, and the cosine similarity of two text vectors is calculated. The formula is TF-IDF Similarity TF-IDF Similarity is expressed as:
[0022]
[0023] Among them, TF(w i ,d j ) is the word w i In the document d j TF-IDF value in, i is the number of words, j is the number of documents, and n is the number of words.
[0024] Preferably, the text similarity is calculated by Word2Vec similarity, which converts the text into the average value of the word vector and calculates the cosine similarity of the two text vectors. The Word2Vec similarity is expressed as:
[0025]
[0026] in, and are the average word vectors of the two texts.
[0027] Preferably, the keyword difference coefficient is characterized by the Jaccard coefficient, the text is converted into a keyword set, and the Jaccard coefficient of two keyword sets is calculated as the keyword difference coefficient, and the formula is:
[0028]
[0029] Among them, S 1 and S 2 They are the keyword sets of the two texts respectively.
[0030] Preferably, the sentence length difference coefficient is characterized by the length difference, and the difference between the lengths of the two text sentences is calculated and converted into the sentence length difference coefficient, and the formula is:
[0031]
[0032] Among them, L 1 and L2 are the sentence lengths of the two texts respectively.
[0033] The present invention provides a coal logistics park content generation customer service system based on a large model, and its beneficial effects include:
[0034] 1. Improve customer service efficiency and quality. Big model technology enables the customer service system to quickly understand user questions and generate accurate answers, significantly improving customer service response speed and reducing user waiting time. Big models have powerful language understanding and generation capabilities, and can accurately understand various complex, vague and even ambiguous questions raised by users, thereby giving more accurate replies. Compared with traditional intelligent customer service, the customer service system based on big models can conduct multiple rounds of dialogues, continuously adjust answers based on user feedback, and better understand user needs.
[0035] 2. Enhance the user experience of the coal logistics park: By analyzing the user's historical interaction data, the system can provide users with personalized service suggestions and improve user satisfaction and loyalty. The system can interact with users in a variety of information forms such as text, voice, and images to improve the richness and naturalness of the interaction. The customer service system based on the large model can provide 24-hour online service to meet users' consulting needs anytime and anywhere.
[0036] 3. Reduce the operating costs of the coal logistics park: The application of the system can replace part of the manual customer service work, thereby reducing the labor costs of the logistics park. By automatically processing a large number of user inquiries, the system can significantly improve the company's operating efficiency and reduce manual processing time. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 A flowchart of a customer service system for content generation in a coal logistics park based on a large model provided by an embodiment of the present invention;
[0038] Figure 2 An architectural diagram of a coal logistics park content generation customer service system based on a large model provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0039] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0040] like Figure 1 and Figure 2 As shown, a coal logistics park content generation customer service system based on a large model is provided in an embodiment of the present invention, and the system includes:
[0041] The data collection module is used to obtain and process the problem data input by the user.
[0042] In this system, the data collection module includes a customer service question collection unit: it is used to collect text word vector parameters, keyword parameters, and character count parameters of questions asked by customers to intelligent customer service; the system first needs to collect a large amount of customer service conversation data related to the coal logistics park, which may come from historical customer service records, user feedback, and FAQ, etc.
[0043] Customer service problem processing unit: cleans, removes duplicates, and labels the collected data to improve data quality. At the same time, data preprocessing is performed through operations such as word segmentation, removal of stop words, and part-of-speech tagging to provide a basis for subsequent feature engineering.
[0044] The customer service question processing unit includes a sentence vector sub-unit, a keyword sub-unit and a sentence length sub-unit. The vector sub-unit transmits the word vector parameters to a preset word vector mathematical model to obtain text similarity, the keyword sub-unit transmits the keyword parameters to the keyword mathematical model to obtain the keyword difference coefficient, and the sentence length sub-unit transmits the character number parameter to the sentence length mathematical model to obtain the sentence length difference coefficient, and transmits the obtained data to the customer service question matching degree estimation module.
[0045] In this embodiment, text similarity is expressed in a variety of ways, including:
[0046] Text similarity is represented by word frequency vector similarity. The word frequency vector is used to represent the text, and the cosine similarity of two texts is calculated. The text similarity TF Similarity is expressed as:
[0047]
[0048] Among them, TF(w i ,d j ) is the word w i In the document d j The word frequency in , i is the word number, j is the document number, and n is the number of words;
[0049] Text similarity is characterized by TF-IDF similarity. TF-IDF vectors are used to represent texts. The cosine similarity of two text vectors is calculated using the formula TF-IDF Similarity. TF-IDF Similarity is expressed as:
[0050]
[0051] Among them, TF(w i ,d j ) is the word w i In the document dj TF-IDF value in, i is the number of words, j is the number of documents, and n is the number of words;
[0052] The text similarity is calculated by Word2Vec similarity, which converts the text into the average value of the word vector and calculates the cosine similarity of the two text vectors. The Word2Vec similarity is expressed as:
[0053]
[0054] in, and are the average word vectors of the two texts.
[0055] The keyword difference coefficient is characterized by the Jaccard coefficient. The text is converted into a keyword set, and the Jaccard coefficient of two keyword sets is calculated as the keyword difference coefficient. The formula is:
[0056]
[0057] Among them, S 1 and S 2 are the keyword sets of the two texts respectively;
[0058] The keyword difference coefficient can also be represented by the edit distance, which is calculated by calculating the edit distance (Levenshtein Distance) of two texts and then converting it into the keyword difference coefficient. The edit distance represents the minimum number of editing operations required to convert one text into another.
[0059] The sentence length difference coefficient is characterized by the length difference. The difference between the lengths of two text sentences is calculated and converted into the sentence length difference coefficient. The formula is:
[0060]
[0061] Among them, L 1 and L 2 are the sentence lengths of the two texts respectively.
[0062] The coal logistics industry knowledge base is used to store industry knowledge data of the coal logistics industry.
[0063] In this embodiment, the coal logistics industry knowledge base is used to quickly and accurately answer users' questions and provide scientific basis and data support for decision-making. Through the interactive communication between users and the system, the knowledge base module can popularize relevant knowledge of the coal logistics industry to users and enhance users' industry awareness. Including but not limited to the following:
[0064] Basic knowledge of coal: including the formation, classification, coalification degree, coal quality indicators and other basic knowledge of coal, providing the customer service system with basic knowledge of the coal industry.
[0065] Coal logistics process: covers the entire process from coal production to consumption, including the business processes and management points of each link such as coal mining, washing and processing, storage, transportation, and sales.
[0066] Policies and Regulations: Includes national policies and regulations, industry standards, local regulations, etc. related to the coal logistics industry to ensure that the customer service system can comply with legal and regulatory requirements when answering questions.
[0067] Market dynamics: including coal market prices, supply and demand, transportation trends and other market dynamic information, providing real-time market data support for the customer service system.
[0068] Technical equipment: Introduces the technical equipment, process flow, technological innovations, etc. commonly used in the coal logistics industry, and provides technical knowledge support for the customer service system.
[0069] Safety management: Emphasize safety management measures, accident prevention, emergency response, etc. in the coal logistics process, and ensure that the customer service system can provide professional guidance when answering safety questions.
[0070] The customer service question matching degree estimation module is used to calculate the question matching degree index based on the question data and determine the user problem.
[0071] In this system, the customer service question matching degree estimation module is used to transmit the text similarity, keyword difference coefficient and sentence length difference coefficient obtained by the customer service question information processing module to the customer service question matching degree mathematical model to obtain the question matching degree index.
[0072] The question-answer generation model is used to query the coal logistics industry knowledge base to obtain answers based on user questions or generate answers through a large language model. The question-answer generation model includes three parts: the front-end, AI service, and the back-end.
[0073] When a user asks a question through the customer service system, the system inputs the user's question into the trained big model. The big model uses its natural language processing capabilities to understand the user's question and generate a corresponding answer. This answer may come from a direct match in the knowledge base, or it may be brand new content generated by the big model based on context and semantic understanding. The system presents the generated answer to the user and interacts with the user further. If the user is not satisfied with the answer or needs further explanation, the system can continue to generate a more detailed or specific answer.
[0074] In this embodiment, the architecture of the large question-answer generation model includes three parts: the front-end, AI service, and the back-end:
[0075] Front desk: The main interface for users to interact with the system, and also the starting point of the entire question-answering system. Users enter questions here, and the system selects different paths to help users find the most appropriate answers based on the complexity of the questions. In order to improve the user experience, the front desk design should be concise and clear, the operation process should be simple and easy to understand, and it should be able to provide timely feedback on the results of the user's questions.
[0076] AI service: The core part of the question-answering system, responsible for understanding and processing questions. It receives questions passed from the front desk and returns corresponding answers based on the data source in the background. AI services include tasks such as question-answer matching, vectorization processing, and large model generation of answers. Through matching algorithms and vectorization processing, AI services can quickly retrieve the content that best matches the user's question to answer. When an exact match cannot be found, the system will call large model services, such as advanced large language models such as Qwen-7B-Chat, to generate more complex and personalized answers.
[0077] Backend: The core of the knowledge base and data processing, mainly responsible for managing and maintaining the data and documents required for system operation. It provides knowledge support for the front desk and AI services to ensure that questions can be fully and accurately answered. Backend administrators can upload local documents within the enterprise as the basic content of the knowledge base. These documents will be automatically parsed and entered into the database to form structured data. In addition, the backend will also perform word segmentation and embedding vectorization on these texts, convert them into vector representations and store them in the vector database for fast semantic retrieval by AI services.
[0078] When training question-answering to generate a large model, the specific methods and strategies are as follows:
[0079] 1. Pre-training strategy: Pre-train on large-scale text datasets to learn general representations and knowledge of language and lay the foundation for subsequent fine-tuning. For example, pre-training on large-scale question-answering datasets enables the model to have the initial ability to generate answers.
[0080] 2. Fine-tuning strategy: Based on pre-training, fine-tune the model on a specific question-answer generation task dataset to make the model better suited to the specific task. During fine-tuning, you can adjust hyperparameters such as learning rate and optimizer to achieve better performance.
[0081] 3. Data enhancement strategy: Expand the training data set through data enhancement technology to improve the robustness and generalization ability of the model. For example, synonym replacement and sentence reorganization can be performed on questions and answers to generate more variants.
[0082] 4. Regularization strategy: Use regularization techniques such as weight decay, dropout, etc. during training to prevent the model from overfitting. For example, use dropout in the encoder and decoder of the model to reduce the risk of overfitting.
[0083] 5. Multimodal fusion strategy: For question-answering tasks involving multimodal information (such as visual question-answering), information from multiple modalities such as text and images can be fused to obtain a more comprehensive representation. For example, visual features and text features are used as input to the model to generate answers related to the image content.
[0084] In addition, some specific training strategies and methods, such as pre-training and fine-tuning, adversarial training, and multi-task learning, are also widely used in the training of large question-answer generation models. By combining these methods and techniques, the performance and training efficiency of the model can be improved.
[0085] Question-answer generation large model training and optimization: Use the processed data to train the large model so that it can understand the relevant knowledge of the coal logistics park and the user's intention. During the training process, the pre-trained model can be used for transfer learning to improve the training effect. At the same time, the distributed training technology is used to speed up the training speed and reduce the training cost. According to the performance of the model during the training process, the learning rate is dynamically adjusted to improve the model convergence speed. In addition, the model performance can be optimized through strategies such as multimodal fusion, personalized services, and continuous learning.
[0086] Customer service interface module: presents the content generated by the large model processing module to customers through voice, text, video and other means to achieve efficient and intelligent customer service.
[0087] The present invention also includes a client (PC & mobile terminal): used to transmit the final answer output by the question matching output module to the user for human-computer interaction. The user enters the question on the client, and the question is then submitted to the question-answering model. The question-answering model generates a customer service answer based on the input user question and returns the customer service answer to the user terminal.
[0088] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. A coal logistics park content generation customer service system based on a large model, characterized in that: The system comprises: Data collection module, used to obtain and process the problem data input by the user; The coal logistics industry knowledge base is used to store industry knowledge data of the coal logistics industry; The customer service question matching degree estimation module is used to calculate the question matching degree index based on the question data and determine the user's question; A large model for question-answer generation is used to query the coal logistics industry knowledge base to obtain answers based on user questions or generate answers through a large language model. The large model for question-answer generation includes three parts: front-end, AI service, and back-end. The customer service interface module is used to present answers to customers in a variety of ways.
2. The coal logistics park content generation customer service system based on a large model according to claim 1 is characterized in that: The data collection module includes a customer service question collection unit and a customer service question processing unit. The customer service question collection unit is used to collect text word vector parameters, keyword parameters and character number parameters of questions raised by customers to intelligent customer service. The customer service question processing unit is used to clean, deduplicate and annotate the collected data, and perform data preprocessing through word segmentation, removal of stop words and part-of-speech tagging.
3. The coal logistics park content generation customer service system based on a large model according to claim 1 is characterized in that: The data stored in the coal logistics industry knowledge base includes coal basic knowledge data, coal logistics process data, market dynamics data, technical equipment data, policy and regulation data, and safety management data.
4. The coal logistics park content generation customer service system based on a large model according to claim 1 is characterized in that: When training a large question-answer generation model, a pre-training strategy is adopted. Pre-training is performed on a large-scale text dataset to learn the general representation and knowledge of the language. Based on the pre-training, fine-tuning is performed on a specific question-answer generation task dataset. During fine-tuning, the learning rate and optimizer are adjusted, and the training dataset is expanded through data enhancement. Regularization strategies are used during training to prevent model overfitting. For question-answer generation tasks involving multimodal information, multimodal fusion is used for representation.
5. The coal logistics park content generation customer service system based on a large model according to claim 1 is characterized in that: The question matching index is calculated based on the question data. The process of determining the user's question includes calculating text similarity, keyword difference coefficient and sentence length difference coefficient.
6. The coal logistics park content generation customer service system based on a large model according to claim 5 is characterized in that: Text similarity is represented by word frequency vector similarity. The word frequency vector is used to represent the text, and the cosine similarity of two texts is calculated. The text similarity TF Similarity is expressed as: Among them, TF(w i ,d j ) is the word w i In the document d j The word frequency in , i is the word number, j is the document number, and n is the number of words.
7. The coal logistics park content generation customer service system based on a large model according to claim 5 is characterized in that: Text similarity is characterized by TF-IDF similarity. TF-IDF vectors are used to represent texts. The cosine similarity of two text vectors is calculated using the formula TF-IDF Similarity. TF-IDF Similarity is expressed as: Among them, TF(w i ,d j ) is the word w i In the document d j TF-IDF value in, i is the number of words, j is the number of documents, and n is the number of words.
8. The coal logistics park content generation customer service system based on a large model according to claim 5 is characterized in that: The text similarity is calculated by Word2Vec similarity, which converts the text into the average value of the word vector and calculates the cosine similarity of the two text vectors. The Word2Vec similarity is expressed as: in, and are the average word vectors of the two texts.
9. The coal logistics park content generation customer service system based on a large model according to claim 5 is characterized in that: The keyword difference coefficient is characterized by the Jaccard coefficient. The text is converted into a keyword set, and the Jaccard coefficient of two keyword sets is calculated as the keyword difference coefficient. The formula is: Among them, S1 and S2 are keyword sets of two texts respectively.
10. The coal logistics park content generation customer service system based on a large model according to claim 5 is characterized in that: The sentence length difference coefficient is characterized by the length difference. The difference between the lengths of two text sentences is calculated and converted into the sentence length difference coefficient. The formula is: Among them, L1 and L2 are the sentence lengths of the two texts respectively.