Customer service reply method and device, equipment and medium

By replacing rich media elements with structured tags in the RAG customer service system and correcting the initial response text based on context, we achieved accurate and complete presentation of rich media elements, solved the problem of URLs occupying too many tokens and generating errors, and improved user experience and system reliability.

CN120632029APending Publication Date: 2025-09-12GUANGZHOU SHANGYUN NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510717723.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

When faced with documents containing rich media elements such as images and hyperlinks, the existing RAG customer service system has the problem of URLs occupying too many tokens and difficulty in accurately presenting rich media elements, affecting user experience and system reliability.

Method used

By replacing rich media elements with structured tags, constructing labeled prompt text, guiding the large language model to generate initial response text, and correcting it based on the contextual relationship of the element tags, it is finally restored to rich media elements, achieving accurate processing of the entire process.

Benefits of technology

It solves the problem of low token utilization, ensures the complete and accurate presentation of multimedia information, and improves user experience and response reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120632029A_ABST
    Figure CN120632029A_ABST
Patent Text Reader

Abstract

The invention relates to a customer service reply method and device, equipment and medium in the technical field of e-commerce, and the method comprises the steps: responding to a customer service consultation request, determining customer service knowledge materials related to a user consultation text carried by the request, and replacing all rich media elements in the customer service knowledge materials with corresponding element tags; constructing a labeled prompt text according to the replaced customer service knowledge material and the user consultation text, guiding a large language model by the labeled prompt text, and determining a target knowledge fragment in the customer service knowledge material and an initial reply text corresponding to the target knowledge fragment; correcting the element tags in the initial reply text according to the element tags in the target knowledge fragment and the context relationship of the element tags to obtain a corrected reply text; and restoring each element tag in the corrected reply text into a corresponding rich media element, and responding to the customer service consultation request by using the obtained target reply text. According to the method, the accuracy of the customer service reply generated based on retrieval enhancement can be ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of e-commerce technology, and in particular to a customer service response method and its corresponding device, computer equipment, and computer-readable storage medium. Background Art

[0002] In today's digital age, Large Language Model (LLM) technology is booming at an unprecedented rate. The question-answering system based on retrieval-augmented generation (RAG) has unique advantages. In the field of intelligent customer service, it provides customers with accurate answers and helps companies improve service efficiency. In knowledge question-answering scenarios, it meets users' query needs for various professional knowledge and becomes a convenient tool for people to obtain information. In terms of document summarization, it can quickly extract the core content of the document, allowing users to quickly grasp key information. Therefore, it is widely used in many fields and plays an important role. The working principle of the RAG model is to first retrieve document fragments related to the user's question, and then use the generative model to produce a response. In this process, the accuracy and richness of the response are effectively improved, providing users with a better service experience.

[0003] However, existing RAG customer service systems present several pressing challenges when dealing with documents containing rich media elements such as images and hyperlinks. For example, images are typically presented as '![image](url)', while hyperlinks are presented as '[alt](url)'. The URLs in these rich media elements often begin with "http". These long URLs occupy a large number of tokens, a limited number of which wastes the LLM window size, reducing the available space for context and limiting the model's ability to accurately understand and fully utilize the document content. Secondly, due to the generative nature and hallucination issues inherent in large language models, the model struggles to accurately and completely represent rich media elements such as images and hyperlinks in the original document when generating responses. This results in errors in these elements in the generated responses, resulting in information that is inconsistent with the original document content and lacks accuracy. This significantly impacts the customer service experience and reduces user confidence in the system's reliability. Furthermore, the search data in the RAG customer service system mostly uses text formats such as Markdown, while the responses generated by the model are usually in plain text. In this case, rich media elements are difficult to automatically embed into the responses and cannot be accurately and completely presented to users, making it difficult to meet users' viewing needs for multimedia information and limiting the system's robustness and reliability in information presentation.

[0004] In view of the shortcomings of traditional technology, the applicant has been engaged in research in related fields for a long time, and has taken a different approach to solve the industry problems in the e-commerce field. Summary of the Invention

[0005] The primary purpose of this application is to solve at least one of the above problems and provide a customer service response method and its corresponding device, computer equipment, and computer program product.

[0006] In order to meet the various objectives of this application, this application adopts the following technical solutions:

[0007] A customer service response method provided to meet one of the purposes of this application includes the following steps:

[0008] Respond to a customer service consultation request, determine customer service knowledge materials related to the user consultation text carried in the request, and replace each rich media element in the customer service knowledge material with a corresponding element tag;

[0009] Constructing a labeled prompt text based on the replaced customer service knowledge material and the user inquiry text, and using the labeled prompt text to guide the large language model to determine the target knowledge segment in the customer service knowledge material and its corresponding initial response text;

[0010] Correcting the element labels in the initial reply text according to the element labels in the target knowledge fragment and their contextual relationships to obtain a corrected reply text;

[0011] Each element tag in the corrected reply text is restored to a corresponding rich media element, and the customer service consultation request is responded to with the obtained target reply text.

[0012] On the other hand, a customer service reply device provided to meet one of the purposes of the present application includes a request response module, a model reasoning module, a reply correction module and a request answer module, wherein the request response module is used to respond to customer service consultation requests, determine the customer service knowledge material related to the user consultation text carried by the request, and replace each rich media element in the customer service knowledge material with a corresponding element label; the model reasoning module is used to construct a labeled prompt text based on the replaced customer service knowledge material and the user consultation text, and guide the large language model with the labeled prompt text to determine the target knowledge segment in the customer service knowledge material and its corresponding initial reply text; the reply correction module is used to correct the element labels in the initial reply text according to the each element label in the target knowledge segment and its contextual relationship to obtain a corrected reply text; the request answer module is used to restore each element label in the corrected reply text to the corresponding rich media element, and reply to the customer service consultation request with the obtained target reply text.

[0013] On the other hand, a computer device provided to meet one of the purposes of the present application includes a central processing unit and a memory, wherein the central processing unit is used to call and run a computer program stored in the memory to execute the steps of the customer service response method described in the present application.

[0014] On the other hand, a computer program product provided to meet another purpose of the present application includes a computer program / instruction, which, when executed by a processor, implements the steps of the method described in any embodiment of the present application.

[0015] The technical solution of this application has many advantages, including but not limited to the following:

[0016] This application first retains the placeholders and necessary semantic information of rich media elements in the prompt text construction stage, and fundamentally solves the technical bottleneck of long texts such as URLs occupying excessive tokens in traditional RAG systems by converting rich media content such as pictures and hyperlinks into structured tags. It not only enables the limited input window of the large language model to accommodate more substantive contextual content, that is, more customer service knowledge related to user consultations, but also establishes a traceable mapping relationship between rich media elements and text content, providing necessary support for subsequent restoration links.

[0017] Secondly, during the response generation phase, given the generative characteristics and "hallucination" problems of large language models, it is difficult to accurately and completely present rich media elements when generating responses. This can easily lead to information errors and inconsistencies with the original document content, seriously impacting the user's consultation experience and the reliability of customer service responses. To address this, the large language model is first guided to extract the target knowledge fragment that answers user inquiries from the customer service knowledge material, and the corresponding initial response text is inferred based on this fragment. Then, based on the contextual relationship of each element label in the target knowledge fragment, the individual element labels in the initial response text are corrected to be consistent with the individual element labels in the target knowledge fragment. This allows us to fix the problem of inconsistencies between labels that may occur during model generation and those in the original document content.

[0018] Furthermore, in the final stage, the tags are restored to achieve seamless integration of rich media elements and generated text, so that the multimedia content in the document can be accurately and completely presented in the final response, completing the reverse conversion process from structured tags to native rich media representation, thus forming a complete technical closed loop.

[0019] In summary, this approach achieves precise, full-process processing of rich media elements in customer service responses. This bidirectional processing of rich media elements not only optimizes the model's processing efficiency and accuracy, but also ensures the complete and accurate presentation of multimedia information. It also addresses the issue of low token utilization and ensures label consistency between responses and substance through a context-aware correction mechanism. More importantly, it enables lossless conversion of rich media content from retrieval of documents containing rich media content to generated responses, meeting users' demand for high-quality multimedia customer service content. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:

[0021] Figure 1 The network architecture of the e-commerce platform exemplified in this application;

[0022] Figure 2 This is a flowchart of a typical embodiment of the customer service response method of this application;

[0023] Figure 3 This is a functional block diagram of the customer service response device for this application;

[0024] Figure 4 This is a schematic diagram of the structure of a computer device used in this application. DETAILED DESCRIPTION

[0025] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and are not to be construed as limiting the present application.

[0026] It will be understood by those skilled in the art that, unless expressly stated otherwise, the singular forms "a", "an", "said" and "the" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present application refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. In addition, "connected" or "coupled" as used herein may include wireless connections or wireless couplings. The term "and / or" used herein includes all or any units and all combinations of one or more associated listed items.

[0027] It will be understood by those skilled in the art that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those skilled in the art to which this application belongs. It should also be understood that terms such as those defined in common dictionaries should be understood to have meanings consistent with their meanings in the context of the prior art and will not be interpreted in an idealized or overly formal sense unless specifically defined as herein.

[0028] like Figure 1 In the network architecture shown, the e-commerce platform 82 is deployed on the Internet to provide corresponding services to its users. Similarly, the devices 80 of the merchant users of the e-commerce platform 82 and the devices 81 of the consumer users are also connected to the Internet to use the services provided by the e-commerce platform.

[0029] The exemplary e-commerce platform 82 provides supply and demand matching of products and / or services to the general public with the help of Internet infrastructure. In the e-commerce platform 82, products and / or services are provided as commodity information. To simplify the description, the concepts of commodity, product, etc. are used in this application to refer to the products and / or services in the e-commerce platform 82, which may specifically be physical products, digital products, tickets, service subscriptions, other offline services, etc.

[0030] In reality, various entities can access the e-commerce platform 82 as users, use the various online services provided by the e-commerce platform 82, and achieve the purpose of participating in the business activities achieved by the e-commerce platform 82. These entities can be natural persons, legal persons, or social organizations. Corresponding to the two types of entities in business activities, merchants and consumers, the e-commerce platform 82 has two corresponding types of users: merchant users and consumer users. In business activities, all entities in the product distribution chain, including manufacturers, sellers, retailers, logistics providers, etc., can use online services on the e-commerce platform 82 as merchant users, while consumers in business activities, including real or potential consumers, can use online services on the e-commerce platform 82 as their corresponding consumer users. In actual business activities, the same entity can act as both a merchant user and a consumer user, and this should be understood flexibly.

[0031] The infrastructure used to deploy the e-commerce platform 82 primarily includes a backend architecture and frontend devices. The backend architecture runs various online services through a service cluster, including platform-facing middleware or frontend services, consumer-facing services, merchant-facing services, etc., to enrich and improve its service functions. The frontend devices primarily encompass the terminal devices used by users to access the e-commerce platform 82 as clients, including but not limited to various mobile terminals, personal computers, point-of-sale devices, etc. For example, a merchant user can use their terminal device 80 to enter product information for their online store, or generate their product information using an interface open to the e-commerce platform. A consumer user can use their terminal device 81 to access the webpage of the online store implemented by the e-commerce platform 82, trigger the shopping process through the shopping button provided on the webpage, and invoke various online services provided by the e-commerce platform 82 during the shopping process, thereby completing the purpose of placing a shopping order.

[0032] In some embodiments, the e-commerce platform 82 may be implemented by a processing facility including a processor and a memory, the processing facility storing a set of instructions that, when executed, cause the e-commerce platform 82 to perform the e-commerce and support functions described herein. The processing facility may be part of a server, client, network infrastructure, mobile computing platform, cloud computing platform, fixed computing platform, or other computing platform, and may provide electronic components of the e-commerce platform 82, merchant devices, payment gateways, application developers, marketing channels, transportation providers, customer devices, point-of-sale devices, and the like.

[0033] The e-commerce platform 82 can be implemented as an online service such as cloud computing, software as a service (SaaS), infrastructure as a service (IaaS), platform as a service (PaaS), desktop as a service (DaaS), hosted software as a service, mobile backend as a service (MBaaS), information technology management as a service (ITMaaS), etc. In some embodiments, the various functional components of the e-commerce platform 82 can be implemented to be suitable for operation on various platforms and operating systems. For example, for an online store, its administrator users can enjoy the same or similar functions regardless of the various embodiments such as iOS, Android, HomonyOS, or web pages.

[0034] The e-commerce platform 82 can implement its corresponding independent website for each merchant to run its corresponding online store, and provide merchants with corresponding business management engine instances for merchants to establish, maintain, and run one or more online stores in one or more independent websites. The business management engine instance can be used for content management, task automation, and data management of one or more online stores, and can configure various specific business processes of the online store through interfaces or built-in components to support the implementation of business activities. The independent website is the infrastructure of the e-commerce platform 82 with cross-border service functions. Merchants can maintain their online stores more centrally and independently based on the independent website. The independent website usually has a domain name and storage space dedicated to the merchant, and different independent websites are relatively independent. The e-commerce platform 82 can provide standardized or personalized technical support for a large number of independent websites, so that merchant users can customize their own business management engine instance and use this business management engine instance to maintain one or more online stores they own.

[0035] The online store can implement backend configuration and maintenance by having the merchant user log in to its business management engine instance as an administrator. With the support of various online services provided by the infrastructure of the e-commerce platform 82, the merchant user can configure various functions in its online store as an administrator, view various data, etc. For example, the merchant user can manage various aspects of its online store, such as viewing the latest activities of the online store, updating the online store product catalog, managing orders, recent visit activities, total order activities, etc.; the merchant user can also view more detailed information about the business and visitors to the merchant's online store by obtaining reports or metrics, such as showing a sales summary of the merchant's overall business, specific sales and participation data of active sales marketing channels, etc.

[0036] The e-commerce platform 82 may provide communication facilities and associated merchant interfaces for providing electronic communications and marketing, such as utilizing electronic message aggregation facilities to collect and analyze communication interactions between merchants, consumers, merchant devices, customer devices, point-of-sale devices, etc., aggregating and analyzing communications, such as for increasing the potential for providing product sales, etc. For example, a consumer may have questions about a product, which may generate a conversation between the consumer and the merchant (or an automated processor-based agent on behalf of the merchant), wherein the communication facility is responsible for the interaction and provides the merchant with analysis on how to increase the probability of sales.

[0037] In some embodiments, applications suitable for installation on terminal devices can be provided to serve the access needs of different users, so that various users can access the e-commerce platform 82 by running applications on the terminal devices, such as the merchant backend module of the online store in the e-commerce platform 82. In the process of implementing business activities through these functions, the e-commerce platform 82 can implement various functions related to supporting business activities as middleware or online services and open corresponding interfaces, and then implant toolkits corresponding to the interface access functions into the application to realize functional expansion and task implementation. The business management engine can include a series of basic functions and expose these functions to online services and / or application calls through APIs. The online services and applications use the corresponding functions by remotely calling the corresponding APIs.

[0038] Supported by the various components of the business management engine instance, the e-commerce platform 82 provides online shopping functionality, enabling merchants to connect with customers in a flexible and transparent manner. Consumers can select items online, create an order, provide a delivery address in the order, and complete payment confirmation for the order. Merchants can then review and complete or cancel the order.

[0039] A customer service reply method of the present application can be programmed as a computer program product and deployed in a client or server for execution. For example, in the exemplary application scenario of the present application, it can be deployed and implemented in the server of an e-commerce customer service platform. The method can be executed by accessing the interface opened after the computer program product is run and performing human-computer interaction with the process of the computer program product through a graphical user interface.

[0040] See also Figure 2 The customer service reply method of the present application, in its typical embodiment, includes the following steps:

[0041] Step S1100: respond to a customer service consultation request, determine customer service knowledge materials related to the user consultation text carried in the request, and replace each rich media element in the customer service knowledge materials with a corresponding element tag;

[0042] The triggering of the customer service consultation request originates from the consultation and customer service operation initiated by the user in the interactive interface of the e-commerce platform (such as the product details page, order page or online customer service portal). The user can be a consumer who browses and / or purchases products on the e-commerce platform, or a store operator who sells products. The customer service consultation request is transmitted to the server via the HTTP / HTTPS protocol, and its data packet structure usually contains header information (such as session ID, user identity identifier) ​​and payload data (user consultation text provided by the user). The method of obtaining the user consultation text includes but is not limited to the following embodiments: keyboard input text received by the touch screen of the mobile terminal; text information converted by the voice recognition module, such as the text generated by the voice consultation input by the user through the microphone after being processed by the ASR (automatic speech recognition) engine; structured text triggered by preset options, such as the consultation content automatically generated after the user selects the "damaged product" option on the return page; extracting the content in the picture provided by the user through OCR (optical character recognition), etc. When a distributed system is adopted, this step can be routed to the message queue after the API gateway receives the request in the system, and the consumer service extracts the text data from RabbitMQ or Kafka.

[0043] Customer service knowledge materials are documents related to user inquiries extracted from the customer service knowledge base. They contain rich media elements, which refer to multimedia content other than pure text, such as images, videos, audio, hyperlinks, tables, mind maps, and so on. These elements typically convey relevant information to users more intuitively. For example, customer service knowledge materials may include a mixed text and image guide, with images showing the steps and text describing the instructions. For example, in a Markdown document, the rich media element for images is represented as '! [image](url)', and the content of hyperlinks is represented as '[alt](url)'.

[0044] First, multiple customer service knowledge documents related to the user consultation text are recalled from a preset customer service knowledge base. In one embodiment, a pre-trained text encoding model is applied to encode semantic vectors for each content block based on the deep semantic information of each pre-split content block from each customer service knowledge document in the customer service knowledge base. A corresponding semantic vector is then encoded based on the deep semantic information of the user consultation text. A vector similarity algorithm is then used to calculate the vector similarity between the semantic vectors of each content block and the semantic vector of the user consultation text. Subsequently, the content blocks are sorted in descending order based on vector similarity. The number of recalled characters is calculated by multiplying the upper limit of the input character count of the large language model by a preset multiplier. Multiple content blocks that are ranked high and that meet the number of recalled characters are selected. A unique identifier for the content block is appended to the header of each content block. Subsequently, the appended content blocks, in the current order, are used to form a document content sequence, serving as customer service knowledge material. The preset ratio is used to control the customer service knowledge materials obtained by recalling, so that after the various rich media elements therein are subsequently replaced with corresponding element tags, there is still sufficient text content to provide for the subsequent guidance of the large language model, and to avoid recalling too many content blocks. The specific value can be set by technical personnel in this field as needed, for example, 5 times.

[0045] The customer service knowledge base is a pre-built structured or semi-structured database that stores any one or more types of customer service knowledge documents, including historical question and answer records, product manuals, and FAQs. For each customer service knowledge document in the customer service knowledge base, the individual content blocks within the document are separated. These blocks are semantically complete portions of the document that can independently answer at least one user inquiry. Those skilled in the art can flexibly implement content block separation based on the disclosure herein.

[0046] The text encoding model is pre-trained to extract the semantics of the input text and represent it as a vector. This ensures that the corresponding vectors of texts with similar semantics are closer in the semantic space. Vector similarity algorithms can be cosine similarity, Jaccard similarity, and other algorithms. Those skilled in the art can choose one to implement as needed.

[0047] Furthermore, each rich media element in the customer service knowledge material is identified. These rich media elements include but are not limited to non-pure text multimedia content such as pictures, videos, audios, hyperlinks, tables, mind maps, etc. For example, pictures are usually represented as! [image](url) in Markdown format documents, and hyperlinks are represented as [alt](url). Obtain a unique element tag pre-generated for each rich media element. The format of the element tag can be [docId_element type N], such as [1...11_picture 1], etc., where docId is the unique identifier of the customer service knowledge document where the rich media element is located, and the element type refers to the multimedia entity corresponding to the rich media element, such as a picture, and N is the serial number of the picture or link in the document. At the same time, establish a mapping table to store each rich media element and its mapped associated element tags for subsequent use in restoring these rich media elements.

[0048] Based on the characteristics of the text expression formats of various rich media elements, corresponding regular expressions can be designed. These regular expressions can be used to identify the corresponding rich media elements. For example, \[.*? \]\((.*?)\) can be used to identify images in Markdown format documents.<audio.*?src="(.*?)".*?> Can be used for audio in this document.

[0049] Step S1200: Constructing a labeled prompt text based on the replaced customer service knowledge material and the user inquiry text, and using the labeled prompt text to guide the large language model to determine the target knowledge segment in the customer service knowledge material and its corresponding initial response text;

[0050] Next, the maximum number of characters in the large language model's input is calculated, minus the total number of characters in the user's inquiry text. This is then subtracted from the total number of characters in the preset prompt template to arrive at the remaining total number of input characters. The top-ranked content from the customer service knowledge material that meets the remaining total number of characters is extracted and embedded into the prompt template along with the user's inquiry text to create the labeled prompt text.

[0051] The prompt template is a fixed-format string. Its design must guide the large language model to retrieve the partial content needed to respond to user inquiries from the customer service knowledge material, organize the language based on this content to generate the corresponding response content, and ultimately output the identifiers and response content based on each part of the content. The content blocks corresponding to each identifier are extracted from the customer service knowledge material to form the target knowledge fragment, and the generated response content is used as the initial response text. For example, the prompt template: "User question: {user inquiry text}\nReference knowledge: {customer service knowledge material}\nPlease generate a response strictly based on the above reference knowledge. Do not make up! Output the response result and the identifiers of each knowledge block referenced in the reference knowledge."

[0052] Step S1300: Correct the element tags in the initial response text according to each element tag and its context relationship in the target knowledge fragment to obtain a corrected response text;

[0053] Correcting the element tags in the initial response text is a key step to ensure the accuracy of multimedia content citation. Therefore, by analyzing the structural relevance between the target knowledge fragment and the initial response text, the reference order and integrity of the element tags are dynamically adjusted. Specifically in implementation, first, the target knowledge fragment and the initial response text are respectively divided into content composition units that are semantically independent, such as sentences and phrases divided by punctuation marks or natural language processing techniques. For each content composition unit in the target knowledge fragment, calculate the reference confidence between this content composition unit and each content composition unit in the initial response text to measure the matching degree of the content unit in the initial response text and the content unit in the target knowledge fragment in multiple dimensions such as semantics and logic. Its determination method can be achieved based on technical means such as text similarity algorithms and semantic understanding models. For example, the BERT model can be used to calculate the semantic similarity of two content units, and the obtained similarity value is used as the reference confidence.

[0054] For each reference confidence, when the reference confidence is greater than or equal to the preset threshold, start the tag synchronization mechanism, and adjust the element tag sequence in the subsequent content after the content composition unit corresponding to the initial response text to be the same as the element tag sequence in the subsequent content after the corresponding content composition unit in the target knowledge fragment. For easy understanding, taking the scenario of product return as an example, if the target knowledge fragment contains an operation guide sequence "[doc123_picture 1][doc123_video 2]", and the initial response text referring to this part of the operation guide is missing this sequence after the corresponding semantic unit, then append this sequence in the target knowledge fragment to the unit in the initial response text. For partial matching situations, for example, the initial response text only contains "[doc123_picture 1]", according to the relative position of the tags in the target knowledge fragment (such as the video tag is after the picture tag), insert the missing "[doc123_video 2]" into the correct position. During the tag correction process, synchronously verify the validity of the tags in the mapping table, and delete the invalid tags not recorded in the mapping table, such as illegal references like "[doc12_picture 1]" inferred by the large language model itself. The preset threshold can be set by those skilled in the art as needed, such as 0.8. In terms of implementation methods, this step can adopt two processing paths: the first is the incremental completion strategy, that is, only supplement the missing tags without changing the existing tag order; the second is the overwrite synchronization strategy, that is, completely replace the corresponding part in the initial response with the complete tag sequence of the target knowledge fragment.

[0055] Step S1400: Restore each element tag in the corrected reply text to the corresponding rich media element, and answer the customer service consultation request with the obtained target reply text.

[0056] Restore the corrected reply text after tag correction and synchronization to a complete customer service reply containing the original multimedia resources, i.e., rich media elements. This depends on the mapping table of rich media elements and element tags established in the early stage. The table stores the corresponding relationships such as [doc123_picture 1] and![image](http: / / cd.example.com / image1.jpg) in the form of mapping pairs. The restoration process needs to perform the following operations:

[0057] First, parse all the element tags in the corrected reply text according to the pre-designed format of the element tags. The format of the element tags follows the standardized naming rule of [docId_element type N]. For example, [1...11_video 3] represents the 3rd video resource in the document with the document ID of 1...11. When parsing, use regular expressions to match the element tags to ensure accurate positioning of all element tags to be replaced.

[0058] Second, perform one-to-one replacement according to the mapping table. For each parsed element tag, retrieve the mapping table to obtain its corresponding rich media element. For example, the tag [456_audio 2] is restored to <audio src="http: cd.example.com audio2.mp3"controls>The audio tag.

[0059] After the replacement is complete, the target response text is transmitted to the requester via HTTP / HTTPS. Transmission is adaptively optimized based on the user's device type: on mobile apps, images and videos use progressive loading; on web pages, CDN acceleration and lazy loading are enabled. The resulting response retains the text representation generated by the large language model while fully restoring key multimedia resources such as instructional diagrams and product demonstration videos, ensuring accurate and complete customer service responses.

[0060] It is not difficult to understand from the above embodiments that compared with the prior art, the present application has many advantages, including at least:

[0061] This application first retains the placeholders and necessary semantic information of rich media elements in the prompt text construction stage, and fundamentally solves the technical bottleneck of long texts such as URLs occupying excessive tokens in traditional RAG systems by converting rich media content such as pictures and hyperlinks into structured tags. It not only enables the limited input window of the large language model to accommodate more substantive contextual content, that is, more customer service knowledge related to user consultations, but also establishes a traceable mapping relationship between rich media elements and text content, providing necessary support for subsequent restoration links.

[0062] Secondly, during the response generation phase, given the generative characteristics and "hallucination" problems of large language models, it is difficult to accurately and completely present rich media elements when generating responses. This can easily lead to information errors and inconsistencies with the original document content, seriously impacting the user's consultation experience and the reliability of customer service responses. To address this, the large language model is first guided to extract the target knowledge fragment that answers user inquiries from the customer service knowledge material, and then infer the corresponding initial response text based on this fragment. Then, based on the contextual relationship of each element label in the target knowledge fragment, the individual element labels in the initial response text are corrected to be consistent with the individual element labels in the target knowledge fragment, so that the problem of inconsistency between labels that may occur during model generation and the labels in the original document content can be fixed.

[0063] Furthermore, in the final stage, the tags are restored to achieve seamless integration of rich media elements and generated text, so that the multimedia content in the document can be accurately and completely presented in the final response, completing the reverse conversion process from structured tags to native rich media representation, thus forming a complete technical closed loop.

[0064] In summary, this approach achieves precise, full-process processing of rich media elements in customer service responses. This bidirectional processing of rich media elements not only optimizes the model's processing efficiency and accuracy, but also ensures the complete and accurate presentation of multimedia information. It also addresses the issue of low token utilization and ensures label consistency between responses and substance through a context-aware correction mechanism. More importantly, it enables lossless conversion of rich media content from retrieval of documents containing rich media content to generated responses, meeting users' demand for high-quality multimedia customer service content.

[0065] In a further embodiment, step S1100, determining customer service knowledge material related to the user consultation text carried in the request, includes the following steps:

[0066] Step S1110: Determine, from a preset customer service knowledge base, a plurality of customer service knowledge documents related to the user consultation text carried in the customer service consultation request;

[0067] In one embodiment, document summaries associated with each customer service knowledge document are first obtained from the customer service knowledge base. A document summary is a high-level summary of the core content of a customer service knowledge document, typically generated by extracting key information from the document using natural language processing techniques. For example, a text summarization algorithm, such as the extractive summarization method based on TF-IDF (Term Frequency-Inverse Document Frequency), can be used to extract keywords and key sentences from the document and combine them into a summary. Alternatively, a generative summarization model based on a neural network, such as the Seq2Seq model, can be used to generate a semantically coherent document summary.

[0068] Next, determine the relevance between each document summary and the user consultation text. The relevance measures the degree of match between the document summary content and the user consultation intention, which can be achieved in a variety of ways. One embodiment is to use a text similarity algorithm, such as a cosine similarity algorithm based on semantic vectors, to vectorize the document summary and the user consultation text separately, and then calculate the cosine value of the angle between the vectors. The closer the cosine value is to 1, the higher the relevance. Another embodiment is to use a deep learning model, such as the BERT (Bidirectional Encoder Representations from Transformers) model, to semantically encode the document summary and the user consultation text, and then determine the relevance by calculating the similarity between the encoded semantic vectors.

[0069] Finally, multiple customer service knowledge documents whose relevance meets preset conditions are determined. The preset conditions can be set based on actual needs. For example, a relevance threshold can be set, and only customer service knowledge documents with relevance above the threshold are determined. Alternatively, a certain number of customer service knowledge documents with the highest relevance are determined.

[0070] Step S1120: Recall document contents that meet a preset character limit from the plurality of customer service knowledge documents to construct customer service knowledge materials.

[0071] Sort each customer service knowledge document in descending order of relevance, calculate the upper limit of the number of input characters of the large language model multiplied by the preset multiplier to obtain the number of recalled characters, filter out multiple customer service knowledge documents that are ranked high and reach the corresponding number of recalled characters, append the unique identifier of the customer service knowledge document to the header of each customer service knowledge document, and then use the customer service knowledge documents appended under the current ranking to form a document content sequence as customer service knowledge material.

[0072] It can be understood that when step S1200 is executed with the customer service knowledge material obtained in this embodiment, the identifiers of each part of the customer service knowledge material, that is, the identifiers of each corresponding customer service knowledge document, can be output from the large language model. Subsequently, the customer service knowledge documents corresponding to each identifier are extracted from the customer service knowledge material to form the target knowledge fragment.

[0073] In this embodiment, firstly, by associating customer service knowledge documents with their historical consultation text sets to form a question-answer mapping relationship, the user's current consultation intention is deeply matched with the historical consultation scenario at the semantic level, effectively solving the semantic deviation problem caused by traditional keyword matching or simple vector retrieval. Secondly, through the dual constraints of question-answer matching ranking and character number limit, it not only guarantees the relevance ranking of recalled documents, but also avoids the risk of excessive context truncation caused by input window limitations of large language models. In particular, in the pre-processing operation of appending the unique identifier of the document, not only is a data foundation established for subsequent label correction and content tracing, but also the precise control of the granularity of knowledge fragments is achieved through the management of identified content blocks, so that the core content most relevant to the user's consultation can be retained first within a limited input window. Relevant content blocks can be recalled first rather than lengthy full-text documents, thereby improving the information density of knowledge materials.

[0074] In a further embodiment, step S1110, determining, from a preset customer service knowledge base, a plurality of customer service knowledge documents related to the user consultation text carried in the customer service consultation request, comprises the following steps:

[0075] Step S1111: Obtain a consulting text set corresponding to each customer service knowledge document in a preset customer service knowledge base;

[0076] The customer service knowledge base is a pre-built structured or semi-structured database that stores customer service knowledge documents such as historical Q&A records, product manuals, and FAQs. Each customer service knowledge document is associated with a set of inquiry texts. The historical inquiry texts in this set are provided by users during past consultations, and the customer service knowledge document contains the knowledge content used to answer these user inquiries.

[0077] For ease of understanding, let's take an example: a customer service knowledge document in the customer service knowledge base regarding "Product Return and Exchange Policy." Its corresponding consultation text set includes historical consultation texts such as "How do I process a product return?" and "What are the conditions for exchanging a product?" These historical consultation texts are collected historically and linked to the customer service knowledge document "Product Return and Exchange Policy." This document contains specific platform or store regulations and / or operational procedures that can be used to respond to these historical consultation texts.

[0078] In actual implementation, there are multiple ways to obtain the consultation text set. In one embodiment, the consultation text set corresponding to each customer service knowledge document is retrieved from the customer service knowledge base through a database query operation. The customer service knowledge base can use a relational database or a non-relational database to store data, wherein the customer service knowledge document and the consultation text set can be associated through a foreign key or other association mechanism. For example, in a relational database, a customer service knowledge document table and a consultation text set table can be set up, and the two can be associated through a foreign key, so as to realize the rapid retrieval of the corresponding consultation text set from the customer service knowledge document table.

[0079] In another embodiment, an indexing mechanism is used to obtain the consultation text set. During the construction of the customer service knowledge base, an index can be created for each customer service knowledge document and its corresponding consultation text set. The index includes the identifier of the customer service knowledge document and the identifier of the consultation text set associated therewith. When the consultation text set needs to be obtained, the corresponding consultation text set can be quickly located through the index, thereby improving retrieval efficiency. For example, an inverted index algorithm can be used to quickly retrieve the consultation text set corresponding to the customer service knowledge document based on the index.

[0080] Step S1112: For each consultation text set, determine the consultation similarity between each historical consultation text in the consultation text set and the user consultation text, and select the largest consultation similarity as the question-answer matching degree between the customer service knowledge document corresponding to the consultation text set and the user consultation text;

[0081] In one embodiment, a single consultation text set is used as an example, and other consultation text sets are similarly applied. A deep learning model, such as the BERT model, is applied to semantically encode historical consultation texts and user consultation texts, and the consultation similarity is determined by calculating the similarity between the encoded semantic vectors. Afterwards, the maximum consultation similarity is screened out. This maximum consultation similarity reflects the degree of similarity between the customer service knowledge document corresponding to the consultation text set and the historical consultation text that best matches the user consultation text, and thus serves as the question-answer matching degree. It is not difficult to understand that the maximum consultation similarity can quantify the confidence that the consultation intent between the historical consultation text and the relatively new user consultation text is consistent. Under this premise, combined with the above, a question-answer relationship is first established between the historical consultation text and the customer service knowledge document, so it can be proved that the customer service knowledge document can answer the user consultation text, which is the maximum consultation similarity. Based on this, the consultation similarity is used as the question-answer matching degree between the customer service knowledge document and the user consultation text.

[0082] Step S1113: Determine that the question-answer matching degree satisfies a plurality of customer service knowledge documents corresponding to preset conditions.

[0083] Preset conditions can be set based on actual needs. For example, a threshold can be set so that only customer service knowledge documents with a question-answer match above the threshold are included. Alternatively, the questions can be sorted from high to low based on the question-answer match, with a certain number of customer service knowledge documents ranked at the top.

[0084] In this embodiment, by pre-establishing a mapping relationship between customer service knowledge documents and a collection of inquiry texts, we deeply bind user inquiry intent to historical question-and-answer scenarios, enabling knowledge retrieval based on real-world user needs. Compared to traditional retrieval strategies that rely solely on document content similarity, this step effectively captures the potential correlation between user inquiries and knowledge documents by mining the semantic features of historical inquiry texts. This is particularly suitable for resolving complex inquiry scenarios where user expressions are ambiguous or ambiguous.

[0085] Furthermore, a deep learning model is used to calculate query similarity and select the maximum value as the question-answer matching degree, achieving precise semantic matching. Furthermore, using the maximum query similarity as the criterion ensures that the recalled knowledge documents contain as many historical queries as possible to fully cover the current user's needs, avoiding the problem of matching a candidate due to the user's personalized expression. This provides high-quality signal input for subsequent knowledge screening.

[0086] In a further embodiment, step S1300, correcting the element tags in the initial reply text according to the element tags in the target knowledge fragment and their contextual relationships to obtain a corrected reply text, includes the following steps:

[0087] Step S1310: Segment the target knowledge fragment and the initial reply text into respective content components, and determine reference confidences between the content components from different texts;

[0088] In one embodiment, the target knowledge fragment and the initial response text are segmented based on sentence boundaries. Sentence boundaries can be identified using sentence segmentation algorithms in natural language processing technology, such as using punctuation marks (such as periods, question marks, exclamation marks, etc.) and line breaks as sentence delimiters. Furthermore, the context-awareness of the language model can be combined to more accurately identify sentence boundaries and avoid segmentation errors caused by the misuse or omission of punctuation marks.

[0089] For example, the target knowledge fragment is "Document ID: 456\nDocument Content: XXX0\n1. Add basic information of logistics:\n[456_Picture 1]\nXXX1\n2. Set freight\n[456_Link 1]\nXXX2". The sentence segmentation algorithm can be used to split it into the following content units: "Document ID: 456"; "Document Content: XXX0"; "1. Add basic information of logistics:"; "[456_Picture 1]"; "XXX1"; "2. Set freight"; "[456_Link 1]"; "XXX2".

[0090] In one embodiment, each content component is first pre-processed to filter out redundant characters, such as useless tags and links. Subsequently, an edit distance score is determined between the pre-processed content component derived from the target knowledge fragment and the pre-processed content component derived from the initial reply text, serving as a reference confidence score between the two content components. An exemplary formula is as follows:

[0091]

[0092] Among them, dist score is the edit distance score, String a is the preprocessed content unit derived from the initial reply text, String b is the preprocessed content unit derived from the target knowledge fragment, editDistance(String a, String b) is the edit distance between the two content units obtained by applying the edit distance algorithm, len(String a) is the total number of characters in the preprocessed content unit derived from the initial reply text, and max(len(String a), 1) is the maximum value between the total number of characters and 1.

[0093] It is not difficult to understand that the smaller the reference confidence, the more relevant the two corresponding content components are. In other words, when the large language model infers the reply content, it refers to the content components derived from the target knowledge fragment, and thus derives the content components derived from the initial reply text. The higher the confidence of this event.

[0094] Step S1320: For each reference confidence, when the reference confidence meets the preset conditions, the element label sequence in the subsequent content after the corresponding content component unit in the initial reply text is adjusted to be consistent with the element label sequence in the subsequent content after the corresponding content component unit in the target knowledge segment.

[0095] When the reference confidence level does not exceed the preset threshold, it is determined that the reference confidence level satisfies the preset condition. The preset threshold can be set as needed by those skilled in the art based on the above disclosure, for example, 0.3. At this point, in a further embodiment, the following steps are performed, including:

[0096] Step S1321: Use the corresponding content component unit in the initial reply text as the first component unit, and use the corresponding content component unit in the target knowledge segment as the second component unit;

[0097] For two content component units corresponding to reference confidence levels that do not exceed a preset threshold, the content component unit derived from the initial reply text is used as the first component unit, and the content component unit derived from the target knowledge fragment is used as the second component unit.

[0098] Step S1322: Determine one by one in lexical order whether the content component units following the second component unit in the target knowledge segment are element tags. If they are element tags, append the content component unit to the element tag sequence. If a content component unit is not determined to be an element tag for the first time, terminate the determination process.

[0099] First, create an empty element label sequence. Then, after the second component unit in the target knowledge fragment, each content component unit arranged in order from first to last is judged one by one to determine whether the content component unit is an element label. In this process, when it is confirmed that the single content component unit currently judged exists in the mapping table, the judgment result at this time is that the content component unit is an element label, and the content component unit is appended to the element label sequence. Then, continue to judge the next content component element after the content component unit; when it is confirmed that the single content component unit currently judged does not exist in the mapping table, the judgment result at this time is that the content component unit is not an element label, and the judgment process ends immediately. It is not difficult to understand that each content component unit in the element label sequence is extracted based on the corresponding contextual relationship.

[0100] Step S1323: When none of the content components in the element tag sequence appears in the subsequent content after the first component unit in the initial reply text, append the element tag sequence after the first component unit;

[0101] This indicates that when the large language model infers the initial reply text, it omits all the content components in the entire element tag sequence of the target knowledge segment. Therefore, the corresponding element tags need to be added back to the initial reply text. Accordingly, insert the element tag sequence immediately after the first component unit to ensure that this part of the content is exactly the same as the target knowledge segment, that is, it contains the same number, order, and specific content of element tags.

[0102] Step S1324: When some content components in the element tag sequence appear in the subsequent content after the first component unit in the initial reply text, insert the remaining content components in the element tag sequence into the subsequent content according to the relative positions of the content components in the element tag sequence;

[0103] This indicates that when the large language model infers the initial reply text, it omits some content components in the entire element tag sequence of the target knowledge segment, that is, the corresponding element tags. Therefore, the corresponding element tags need to be added back to the initial reply text. Accordingly, first, determine each content component that needs to be added back in the element tag sequence. Subsequently, according to the relative positions of each content component that needs to be added back in the element tag sequence, insert each content component that needs to be added back into the corresponding position in the subsequent content immediately following the first component unit. For easy understanding, as an example, the element tag sequence is "[456_Picture 1][456_Picture 2][456_Picture 3]", and only "[456_Picture 1]" appears in the subsequent content after the first component unit in the initial reply text. The content components that need to be added back in the element tag sequence are "[456_Picture 2]" and "[456_Picture 3]". According to the relative positions of each content component in this sequence, that is, the picture 2 tag is after the picture 1 tag, and the picture 3 tag is after the picture 2 tag, insert "[456_Picture 2]" after "[456_Picture 1]", and then insert "[456_Picture 3]" after "[456_Picture 2]" to complete the insertion, and "[456_Picture 1][456_Picture 2][456_Picture 3]" appears after the first component unit.

[0104] Step S1325: Delete the element tags that do not exist in the mapping table from the initial reply text according to the mapping table constructed during the replacement process of the customer service knowledge material. The mapping table includes each rich media element in the customer service knowledge material and its corresponding element tag.

[0105] All element tags in the initial reply text are retrieved, and then, by querying the mapping table, each element tag that is not recorded in the mapping table is found from these element tags, and each element tag is deleted from the initial reply text.

[0106] In this embodiment, by establishing a fine alignment mechanism between content components, the problem of confusion in rich media element references caused by "hallucinations" when large language models generate replies is effectively solved. By segmenting the target knowledge fragment and the initial reply text into semantically independent content components and calculating the reference confidence between cross-text units based on the edit distance algorithm, it is possible to accurately locate the paragraphs in the initial reply text that are semantically related to the target knowledge fragment, thereby providing a reliable association anchor for subsequent label synchronization. Compared with traditional full-text similarity comparisons, this confidence assessment mechanism based on fine-grained text units significantly improves the positioning accuracy of correction operations and avoids the destruction of contextual coherence that may be caused by global adjustments.

[0107] Secondly, through context-aware label sequence reconstruction technology, the order of citation of multimedia elements in the reply text is ensured to be consistent with the original logic of the knowledge source document. When it is detected that the element label in the initial reply text is missing or the order is misplaced, it is dynamically completed and rearranged according to the label sequence after the corresponding unit in the target knowledge fragment. This not only fixes the label omission problem that may occur during model generation, but more importantly, it maintains the inherent progressive relationship of rich media elements in scenarios such as operation instructions and process descriptions. For example, in the product return process, the misalignment of the order of image labels and video labels may cause confusion in user operations. By forcibly synchronizing the label sequence of the target knowledge fragment, the logical rigor of the multimedia content reference is fundamentally guaranteed.

[0108] Furthermore, the introduction of a mapping table-driven tag verification mechanism improves response reliability while strengthening system security. By deleting illegal tags not registered in the mapping table, this eliminates the risk of large language models fabricating false multimedia references and blocks potential security vulnerabilities that could result from malicious injection of illegal links. This two-way verification mechanism not only ensures the legitimacy and traceability of rich media elements in the final response, but also, through the strict mapping relationship between structured tags and the original multimedia resources, establishes a trusted closed loop from knowledge retrieval to response generation.

[0109] In summary, this correction step creatively integrates text alignment, sequence reconstruction, and security verification. While mitigating the effects of large language model hallucinations and maintaining the logical integrity of multimedia content, it also establishes a paradigm for the precise and controllable application of rich media elements in customer service response scenarios. Compared to the lag-prone processing of manual review or rule-based filtering in traditional solutions, this algorithm-driven automated correction process ensures semantic consistency between generated responses and source documents at the multimedia citation level, significantly enhancing the professionalism and user experience of the intelligent customer service system.

[0110] In a further embodiment, step S1320, adjusting the element tag sequence in the subsequent content following the corresponding content component unit in the initial reply text to be consistent with the element tag sequence in the subsequent content following the corresponding content component unit in the target knowledge segment, includes the following steps:

[0111] Step S13201: Using the element tag sequence in the subsequent content of the content component unit in the initial reply text as a first tag sequence, and using the element tag sequence in the subsequent content of the content component unit in the target knowledge segment as a second tag sequence. If the second tag sequence exists and the first tag sequence does not exist, append the second tag sequence after the content component unit in the initial reply text.

[0112] For the target knowledge segment, first, create an empty element label sequence. Then, corresponding to the reference confidence that meets the preset conditions, judge each content component unit in the target knowledge segment according to the word order from first to last, and judge whether the content component unit is an element label. In this process, when it is confirmed that the single content component unit currently judged exists in the mapping table, the judgment result at this time is that the content component unit is an element label, and the content component unit is appended to the element label sequence. Then, continue to judge the next content component element after the content component unit; when it is confirmed that the single content component unit currently judged does not exist in the mapping table, the judgment result at this time is that the content component unit is not an element label, and the judgment process ends immediately. It is not difficult to understand that each content component unit in the element label sequence is extracted based on the corresponding contextual relationship. The element label sequence is used as the first label sequence.

[0113] For the initial response text, first, create another empty sequence of element tags. Subsequently, for the reference confidence that meets the preset conditions, after the corresponding content component unit in the initial response text, each content component unit arranged in the order from first to last is judged one by one to determine whether the content component unit is an element tag. During this process, when it is confirmed that the single content component unit being judged currently exists in the mapping table, the judgment result at this time is that the content component unit is an element tag, and this content component unit is appended to the sequence of element tags. Then, continue to judge the next content component element after this content component unit; when it is confirmed that the single content component unit being judged currently does not exist in the mapping table, the judgment result at this time is that the content component unit is not an element tag, and the judgment process is immediately ended. It is not difficult to understand that each content component unit in the sequence of element tags is obtained by extraction according to the corresponding context relationship. Take this sequence of element tags as the second tag sequence.

[0114] Step S13202: When there exist the second tag sequence and the first tag sequence, replace the first tag sequence with this second tag sequence;

[0115] Directly replace the first tag sequence in the initial response text with the second tag sequence. The core of this operation is to ensure that this sequence must be consistent in the initial response text and the target knowledge fragment. For easy understanding, the sequence of element tags in the target knowledge fragment is "[doc123_Picture 1][doc123_Video 2]", and no matter the sequence of element tags in the initial response text is "[doc123_Picture 1]" or "[doc123_Picture 1][doc123_Video 2]" or "[doc123_Video 3]", it will be replaced with "[doc123_Picture 1][doc123_Video 2]" to ensure that the initial response text contains an accurate and complete sequence of element tags.

[0116] Step S13203: According to the mapping table constructed during the replacement process of the customer service knowledge material, delete the element tags that do not exist in the mapping table from the initial response text. The mapping table includes each rich media element in the customer service knowledge material and its corresponding element tag.

[0117] Retrieve all the element tags in the initial response text. Subsequently, by querying the mapping table, find out each element tag that is not recorded in the mapping table from these element tags, and delete each of these element tags from the initial response text.

[0118] In this embodiment, a dual mechanism of dynamic tag sequence replacement and security verification is used to ensure the integrity and legitimacy of references to rich media elements in response text. First, by comparing the tag sequence differences between the initial response and the target knowledge fragment and implementing forced synchronization, this fundamentally addresses the issues of tag omissions or order deviations caused by generation bias in large language models. This precise and direct replacement based on context sequences can efficiently repair structural errors in the model generation process.

[0119] Secondly, by introducing a mapping table-driven illegal tag filtering mechanism, we improve response security while maintaining the traceability of knowledge references. We remove illegal tags not registered in the mapping table, ensuring that all rich media elements in responses originate from the vetted customer service knowledge base, eliminating false references caused by model "hallucinations." We automatically identify and remove these illegal references, ensuring that users only access officially verified resources certified by the platform, thereby maintaining the reliability and credibility of customer service responses.

[0120] Furthermore, the automated reconstruction and verification of label sequences significantly reduces manual review costs. Traditional solutions rely on manual verification of label consistency between responses and source documents. This step, through algorithm-driven sequence matching and mapping table retrieval, achieves millisecond-level label correction and filtering, ensuring efficient processing while avoiding potential operational delays and omissions caused by manual intervention.

[0121] In summary, this approach not only solves the core pain point of citing rich media elements in intelligent customer service scenarios, but also achieves strict consistency between reply content and knowledge source documents at the multimedia level through automated and structured technical means, while taking into account processing efficiency. It provides key technical support for improving the reliability and large-scale application of intelligent customer service systems.

[0122] In a further embodiment, after step S1400, restoring each element tag in the corrected reply text to a corresponding rich media element and responding to the customer service consultation request with the obtained target reply text, the following steps are included:

[0123] Step S1500: In response to the reply with reading event, determine each to-be-read content in the target reply text to construct a read-with sequence;

[0124] After the customer service reply is generated, it further responds to the reply reading event triggered by the user to realize the reading function. The reply reading event is triggered by the user clicking the "auto play" control on the terminal interface to request the system to play the reply content item by item in a specific order. In the specific implementation, the target reply text is first segmented according to the smallest semantic unit to generate a reading sequence. The word segmentation process uses a combination of rule-based regular expression matching and natural language processing technology to split the target reply text into continuous independent units, namely word units. The types of units include pure word segmentation word units, rich media elements, and single punctuation marks or continuous punctuation marks.

[0125] For example, the target reply text:

[0126] 1. Please refer to the diagram:

[0127] After word segmentation, "[image](url)" forms the sequence ["1"".""Please""refer""to"the"diagram""! [image](url)"], in which the element label and label symbol are retained as independent units.

[0128] When constructing a read sequence, initialize an empty queue and append each word segmentation unit to the end of the queue in sequence according to the original word order of the target reply text.

[0129] Step S1510: For each of the to-be-read contents in the reading sequence, a playback component corresponding to a content attribute of the to-be-read content is called to play the to-be-read content.

[0130] In the front-end interface, a sliding window highlights the content currently being read within the target reply text, displaying a visual colored border. As the current content completes, the window automatically moves to the next content to be read until there is no more content to read, and the window disappears from the interface, guiding the user's attention along the way. On the back-end, each content in the read sequence is sequentially read and its attributes are identified. If the content to be read is non-rich media, such as plain text, a pre-defined text player component is invoked to play it using specific text effects, such as setting the background color of the frame to a highlight color and / or bolding the foreground text to highlight the specific content currently being read. The duration from the start to the end of playback is dynamically calculated based on the length of the text, for example, by allocating a preset number of milliseconds per character and setting minimum and maximum thresholds. After playback is complete, a command is triggered to move the sliding window to the next unit position. If the content to be read is rich media, such as an image or audio, the corresponding media player component is invoked. The playback duration of all rich media elements is determined by the media's own duration, such as the total playback time of a video file, or by user-interactive skipping. For example, for images, the user's image player is invoked to load and display the image content within the interface. For audio, the user's audio player is invoked to play the audio file.

[0131] In the recommended embodiment, an operation console can also be built in the front-end interface to allow users to control the overall and / or single-moment playback progress, including but not limited to rewind, skip, double speed, end, pause, and continue.

[0132] In this embodiment, different playback components can be flexibly called according to the attributes of the content to be read, so as to achieve orderly playback and display of various contents in the target reply text, thereby improving the user's reading and replying experience.

[0133] See also Figure 3 A customer service reply device provided to meet one of the purposes of the present application is a functional embodiment of the customer service reply method of the present application. On the other hand, the device provides a customer service reply device to meet one of the purposes of the present application, including a request response module 1100, a model reasoning module 1200, a reply correction module 1300 and a request answering module 1400, wherein the request response module 1100 is used to respond to the customer service consultation request, determine the customer service knowledge material related to the user consultation text carried by the request, and replace each rich media element in the customer service knowledge material with a corresponding element tag; the model reasoning module 1200 is used to construct a labeled prompt text based on the replaced customer service knowledge material and the user consultation text, and use the labeled prompt text to guide the large language model to determine the target knowledge segment in the customer service knowledge material and its corresponding initial reply text; the reply correction module 1300 is used to correct the element tags in the initial reply text according to the each element tag in the target knowledge segment and its contextual relationship to obtain a corrected reply text; the request answering module 1400 is used to restore each element tag in the corrected reply text to the corresponding rich media element, and respond to the customer service consultation request with the obtained target reply text.

[0134] In a further embodiment, the request response module 1100 includes: a document recall sub-module, used to determine, from a preset customer service knowledge base, multiple customer service knowledge documents related to the user consultation text carried by the customer service consultation request; a material extraction sub-module, used to recall document content that meets a preset character limit from the multiple customer service knowledge documents to construct customer service knowledge material.

[0135] In a further embodiment, the document recall submodule includes: a text set acquisition unit, which is used to obtain the consultation text set corresponding to each customer service knowledge document in a preset customer service knowledge base; a question and answer matching unit, which is used to determine, for each consultation text set, the consultation similarity between each historical consultation text in the consultation text set and the user consultation text, and screen out the maximum consultation similarity as the question and answer matching degree between the customer service knowledge document corresponding to the consultation text set and the user consultation text; a document recall unit, which is used to determine that the question and answer matching degree meets the preset conditions for multiple customer service knowledge documents.

[0136] In a further embodiment, the reply correction module 1300 includes: a confidence determination submodule, which is used to segment the respective content component units in the target knowledge fragment and the initial reply text, and determine the reference confidence between the respective content component units originating from different texts; a label adjustment submodule, which is used to, for each reference confidence, adjust the element label sequence in the subsequent content after the corresponding content component unit in the initial reply text to be consistent with the element label sequence in the subsequent content after the corresponding content component unit in the target knowledge fragment when the reference confidence meets the preset conditions.

[0137] In a further embodiment, the label adjustment submodule includes: a unit marking unit, which is used to take the corresponding content component unit in the initial reply text as the first component unit, and the corresponding content component unit in the target knowledge segment as the second component unit; a sequence construction unit, which is used to judge one by one in the lexical order whether the content component unit after the second component unit in the target knowledge segment is an element tag, and when it is an element tag, the content component unit is appended to the element tag sequence, and when it is judged for the first time that the content component unit is not an element tag, the judgment process is ended; a first sequence appending unit, which is used to append the element tag sequence when the element tag sequence does not appear in the subsequent content after the first component unit in the initial reply text. When any content component unit in the column is inserted, the element tag sequence is appended after the first component unit; a tag insertion unit is used to, when part of the content component units in the element tag sequence appear in the subsequent content after the first component unit in the initial reply text, insert the remaining content component units in the element tag sequence into the subsequent content according to the relative positions between the various content component units in the element tag sequence; a tag deletion unit is used to delete the element tags that do not exist in the mapping table constructed in the process of replacing the customer service knowledge material from the initial reply text, wherein the mapping table includes various rich media elements in the customer service knowledge material and their corresponding element tags.

[0138] In a further embodiment, the label adjustment submodule includes: a second sequence appending unit, which is used to use the element label sequence in the subsequent content of the content component unit in the initial reply text as the first label sequence, and the element label sequence in the subsequent content of the content component unit in the target knowledge fragment as the second label sequence, and when the second label sequence exists and the first label sequence does not exist, append the second label sequence after the content component unit in the initial reply text; a sequence replacement unit, which is used to replace the first label sequence with the second label sequence when the second label sequence and the first label sequence exist; a label deletion unit, which is used to delete the element labels that do not exist in the mapping table constructed during the replacement process of the customer service knowledge material from the initial reply text according to the mapping table. The mapping table includes each rich media element in the customer service knowledge material and its corresponding element labels.

[0139] In a further embodiment, after the request response module 1400, it includes: an event response module, which is used to respond to the reply reading event and determine the various contents to be read in the target reply text to construct a reading sequence; a content playback module, which is used to call the playback component corresponding to the content attribute of each content to be read in the reading sequence to play the content to be read.

[0140] In order to solve the above technical problems, the embodiment of the present application also provides a computer device. Figure 4 As shown, a schematic diagram of the internal structure of a computer device. The computer device includes a processor, a computer-readable storage medium, a memory, and a network interface connected via a system bus. Among them, the computer-readable storage medium of the computer device stores an operating system, a database, and computer-readable instructions, and the database may store a control information sequence, and when the computer-readable instructions are executed by the processor, the processor may implement a customer service reply method. The processor of the computer device is used to provide computing and control capabilities to support the operation of the entire computer device. The memory of the computer device may store computer-readable instructions, and when the computer-readable instructions are executed by the processor, the processor may execute the customer service reply method of the present application. The network interface of the computer device is used to connect and communicate with the terminal. Those skilled in the art will understand that Figure 4 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.

[0141] In this embodiment, the processor is used to execute Figure 3 The memory stores the program code and various data required to execute the specific functions of each module and its submodule in the customer service response device. The network interface is used to transmit data between user terminals or servers. The memory in this embodiment stores the program code and data required to execute all modules / submodules in the customer service response device of this application. The server can call the server's program code and data to execute the functions of all submodules.

[0142] The present application also provides a storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the customer service response method of any embodiment of the present application.

[0143] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments of the present application can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the above-mentioned embodiments of the method. The aforementioned storage medium can be a computer-readable storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).

[0144] In summary, this application can ensure the accuracy of customer service responses generated based on retrieval enhancement and improve the user consultation experience.

[0145] Those skilled in the art will understand that the various operations, methods, steps, measures, and schemes in the processes discussed in this application may be interchanged, changed, combined, or deleted. Furthermore, other steps, measures, and schemes in the various operations, methods, and processes discussed in this application may also be interchanged, changed, rearranged, decomposed, combined, or deleted. Furthermore, the steps, measures, and schemes in the various operations, methods, and processes in the prior art that are open source and disclosed in this application may also be interchanged, changed, rearranged, decomposed, combined, or deleted.

[0146] The above description is only part of the implementation methods of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.< / audio>

Claims

1. A customer service reply method, characterized in that: The steps include: Respond to a customer service consultation request, determine customer service knowledge materials related to the user consultation text carried in the request, and replace each rich media element in the customer service knowledge material with a corresponding element tag; Constructing a labeled prompt text based on the replaced customer service knowledge material and the user inquiry text, and using the labeled prompt text to guide the large language model to determine the target knowledge segment in the customer service knowledge material and its corresponding initial response text; Correcting the element labels in the initial reply text according to the element labels in the target knowledge fragment and their contextual relationships to obtain a corrected reply text; Each element tag in the corrected reply text is restored to a corresponding rich media element, and the customer service consultation request is responded to with the obtained target reply text.

2. The customer service reply method according to claim 1, characterized in that: Determining customer service knowledge materials related to the user inquiry text carried in the request includes the following steps: Determining, from a preset customer service knowledge base, a plurality of customer service knowledge documents related to the user consultation text carried in the customer service consultation request; From the plurality of customer service knowledge documents, document contents that meet a preset character limit are retrieved to construct customer service knowledge materials.

3. The customer service reply method according to claim 2, characterized in that: Determining, from a preset customer service knowledge base, a plurality of customer service knowledge documents related to the user consultation text carried in the customer service consultation request, including the following steps: Obtain the consulting text set corresponding to each customer service knowledge document in the preset customer service knowledge base; For each consultation text set, determine the consultation similarity between each historical consultation text in the consultation text set and the user consultation text, and select the largest consultation similarity as the question-answer matching degree between the customer service knowledge document corresponding to the consultation text set and the user consultation text; Determine that the question-answer matching degree satisfies multiple customer service knowledge documents corresponding to preset conditions.

4. The customer service reply method according to claim 1, characterized in that: Correcting the element labels in the initial reply text according to the element labels and their contextual relationships in the target knowledge fragment to obtain a corrected reply text includes the following steps: Segmenting the target knowledge fragment and the initial response text into respective content components, and determining the reference confidence between the content components from different texts; For each reference confidence, when the reference confidence meets the preset conditions, the element label sequence in the subsequent content after the corresponding content component unit in the initial reply text is adjusted to be consistent with the element label sequence in the subsequent content after the corresponding content component unit in the target knowledge fragment.

5. The customer service reply method according to claim 4, characterized in that: Adjusting the element label sequence in the subsequent content after the corresponding content component unit in the initial response text to be consistent with the element label sequence in the subsequent content after the corresponding content component unit in the target knowledge segment includes the following steps: The corresponding content component unit in the initial response text is used as the first component unit, and the corresponding content component unit in the target knowledge fragment is used as the second component unit; Determine one by one in the target knowledge segment whether the content component units following the second component unit are element tags according to the lexical order. If they are element tags, append the content component unit to the element tag sequence. When it is determined for the first time that the content component unit is not an element tag, the determination process ends. When any content component unit in the element tag sequence does not appear in the subsequent content after the first component unit in the initial reply text, the element tag sequence is appended after the first component unit; When some content components in the element tag sequence appear in the subsequent content after the first component in the initial reply text, the remaining content components in the element tag sequence are correspondingly inserted into the subsequent content according to the relative positions of the content components in the element tag sequence; According to the mapping table constructed during the replacement of the customer service knowledge material, element tags that do not exist in the mapping table are deleted from the initial reply text. The mapping table includes each rich media element in the customer service knowledge material and its corresponding element tags.

6. The customer service reply method according to claim 4, characterized in that: Adjusting the element label sequence in the subsequent content after the corresponding content component unit in the initial response text to be consistent with the element label sequence in the subsequent content after the corresponding content component unit in the target knowledge segment includes the following steps: Using the element tag sequence in the subsequent content of the content component unit in the initial reply text as the first tag sequence, using the element tag sequence in the subsequent content of the content component unit in the target knowledge fragment as the second tag sequence, and when the second tag sequence exists and the first tag sequence does not exist, appending the second tag sequence after the content component unit in the initial reply text; When the second tag sequence and the first tag sequence exist, replacing the first tag sequence with the second tag sequence; According to the mapping table constructed during the replacement of the customer service knowledge material, element tags that do not exist in the mapping table are deleted from the initial reply text. The mapping table includes each rich media element in the customer service knowledge material and its corresponding element tags.

7. The customer service reply method according to claim 1, characterized in that: Restoring each element tag in the corrected reply text to a corresponding rich media element and responding to the customer service consultation request with the obtained target reply text includes the following steps: In response to the reply reading event, each to-be-read content in the target reply text is determined to construct a reading sequence; For each of the to-be-read contents in the reading sequence, a playback component corresponding to a content attribute matching the to-be-read content is called to play the to-be-read content.

8. A customer service response device, characterized in that: include: The request response module is used to respond to customer service consultation requests, determine the customer service knowledge materials related to the user consultation text carried in the request, and replace each rich media element in the customer service knowledge materials with the corresponding element tags; The model inference module is used to construct a labeled prompt text based on the replaced customer service knowledge material and the user inquiry text, and use the labeled prompt text to guide the large language model to determine the target knowledge segment in the customer service knowledge material and its corresponding initial response text; A reply correction module, configured to correct the element labels in the initial reply text according to the element labels and their contextual relationships in the target knowledge fragment, to obtain a corrected reply text; The request response module is used to restore each element tag in the corrected response text to a corresponding rich media element, and respond to the customer service consultation request with the obtained target response text.

9. A computer device comprising a central processing unit and a memory, characterized in that: The central processing unit is configured to call and run a computer program stored in the memory to execute the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that It stores a computer program implemented according to the method described in any one of claims 1 to 7 in the form of computer-readable instructions, and when the computer program is called and executed by a computer, the steps included in the corresponding method are executed.