Knowledge base construction and message processing method, computing device, storage medium and product

By constructing a graph knowledge base and using image features and knowledge data mapping, the problem of intelligent customer service being unable to answer image inquiries was solved, achieving efficient and accurate response processing and reducing the intervention of human customer service.

CN121834003APending Publication Date: 2026-04-10HANGZHOU ALIBABA INT INTERNET IND CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU ALIBABA INT INTERNET IND CO LTD
Filing Date
2025-12-19
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing intelligent customer service systems cannot automatically respond to inquiries containing images, resulting in low response efficiency and increased processing costs for human customer service representatives.

Method used

A graph knowledge base is constructed by acquiring the image features and content of sample consultation images, generating knowledge data using an intelligent recognition model, establishing a mapping relationship between image features and knowledge data, and realizing rapid retrieval of knowledge data and generation of response messages through image feature matching.

Benefits of technology

It improves the efficiency of responding to inquiries containing images, reduces reliance on human customer service, and ensures the accuracy and relevance of response messages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121834003A_ABST
    Figure CN121834003A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a knowledge base construction and message processing method, computing equipment, a storage medium and a product. The method comprises the steps of obtaining a sample consultation image, extracting image features of the sample consultation image, recognizing image content of the sample consultation image by using an intelligent recognition model, generating corresponding knowledge data based on the image content, and forming knowledge entries of an image knowledge base by the image features and the knowledge data of the sample consultation image. The image feature in the knowledge entry is used for generating a response message for the consultation message based on the knowledge data in the knowledge entry under the condition that the image feature in the knowledge entry is matched with the target image in the consultation message. According to the technical scheme provided by the embodiment of the invention, the consultation message containing the target image can be automatically answered based on the constructed graph knowledge base, so that the answering efficiency of the consultation message containing the image is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a knowledge base construction and message processing method, computing device, storage medium and product. Background Technology

[0002] With the development of computer technology, network service systems are providing users with increasingly rich online services. However, users often encounter operational problems while using their client to receive online services. In such cases, users can communicate with the network service system's customer service based on the problems they encounter. For example, they can send an inquiry message to customer service, who will then respond with a reply to resolve the user's problem.

[0003] Currently, online service systems typically prioritize using intelligent customer service systems to communicate with users, replacing human customer service representatives. These intelligent customer service systems usually retrieve the question data corresponding to the inquiry from a question-and-answer knowledge base and use the corresponding response data as the reply message.

[0004] However, this method only applies to text-based inquiry messages. If the inquiry message contains images, it is impossible to find the response data from the question-and-answer knowledge base, and the message must be transferred to human customer service for processing, which greatly affects the response efficiency of the inquiry message. Summary of the Invention

[0005] This application provides a knowledge base construction and message processing method, computing device, storage medium, and product to solve the problem that existing intelligent customer service systems cannot automatically respond to inquiry messages containing images.

[0006] Firstly, this application provides a knowledge base construction method, including: Obtain sample consultation images; Extract the image features of the sample consultation image; The intelligent recognition model is used to identify the image content of the sample consultation image; Based on the image content, generate knowledge data corresponding to the image content; Knowledge entries are constructed from the image features of the sample consultation images and the knowledge data; Based on the knowledge entries, a graph knowledge base is constructed; when the image features in the knowledge entries are used to match the target image in the consultation message, a response message is generated based on the knowledge data in the knowledge entries for the consultation message.

[0007] Secondly, embodiments of this application provide a message processing method, including: Retrieve inquiry messages sent by the user client; If the consultation message includes a target image, extract the target image features of the target image; Based on the target image features, a target knowledge entry matching the target image is searched in a graph knowledge base; the graph knowledge base includes multiple knowledge entries; the knowledge entry includes image features of the sample consultation image and knowledge data; the knowledge data is generated based on the image content of the sample consultation image, and the image content is identified from the sample consultation image using an intelligent recognition model; Based on the knowledge data in the target knowledge item, generate a response message for the inquiry message; The response message is sent to the user terminal.

[0008] Thirdly, this application provides a computing device, including a processing component and a storage component; the storage component stores a computing program; the computer program is used to be called and executed by the processing component to implement the knowledge base construction method of the first aspect above, or to implement the message processing method of the second aspect above.

[0009] Fourthly, this application provides a computer storage medium storing a computer program thereon. When the computer program is executed by a processing component, it implements the knowledge base construction method of the first aspect above, or the message processing method of the second aspect above.

[0010] Fifthly, this application provides a computer program product, including a computer program or instructions, which, when executed by a processing component, implement the knowledge base construction method of the first aspect above, or the message processing method of the second aspect above.

[0011] This embodiment acquires sample consultation images, extracts their image features, and uses an intelligent recognition model to identify the image content they contain. Knowledge data is generated based on the image content, and the image features of the sample consultation image and the generated knowledge data constitute knowledge entries in a graph knowledge base. The image features in these knowledge entries are used to generate a response message based on the knowledge data in the user's consultation message, provided they match the target image. This embodiment specifically constructs a graph knowledge base based on the mapping of image features and knowledge data to respond to consultation messages containing images. Upon receiving a consultation message containing a target image, it can quickly recall the knowledge data associated with the target image through image feature matching and generate a response message based on the recalled knowledge data, eliminating the need for manual customer service processing and improving the response efficiency of consultation messages. Furthermore, this embodiment utilizes an intelligent recognition model to identify the image content of the sample consultation image and generates knowledge data in knowledge entries based on the image content, ensuring a high correlation between the knowledge data and the target image, thereby guaranteeing the accuracy of the final generated response message.

[0012] These or other aspects of this application will become more apparent from the description of the following embodiments. Attached Figure Description

[0013] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 A flowchart of an embodiment of the knowledge base construction method provided in this application is shown; Figure 2 A schematic diagram of the structure of the intelligent recognition model provided in this application is shown; Figure 3 This paper presents a flowchart illustrating a method for constructing a graph knowledge base in a practical application scenario. Figure 4 This paper illustrates a flowchart of a method for constructing a graph knowledge base in another practical application scenario of this application. Figure 5 This paper presents a flowchart illustrating a method for maintaining a graph knowledge base in a practical application scenario. Figure 6 A flowchart of one embodiment of the message processing method provided in this application is shown; Figure 7 This paper presents a flowchart illustrating a method for responding to inquiry messages in a practical application scenario. Figure 8 This paper shows a schematic diagram of the structure of an embodiment of the knowledge base construction apparatus provided in this application; Figure 9A schematic diagram of the structure of one embodiment of the message processing apparatus provided in this application is shown; Figure 10 A schematic diagram of the structure of the computing device provided in this application is shown. Detailed Implementation

[0014] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0015] It should be noted that, in the cases involving user information in the embodiments of this application, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse. In addition, the various models involved in this application (including but not limited to language models or large models) comply with relevant laws and standards.

[0016] It should be noted that the technical solutions in this application are applicable to virtual network environments, and the users described generally refer to "virtual users." Real users can register user accounts on the server through registration to obtain user identities in the network environment. The same user account can log in to the server through different types of client terminals, enabling the server to identify the same user.

[0017] Interactions between the server and the user can be based on user accounts. The data received or sent by the server to the user is also based on the user account; in reality, the user's client, corresponding to the user account, receives or sends data to the server. Furthermore, users can also communicate with each other through their user accounts. Here, "user" can refer to an individual or an organization, such as a company; this application does not impose specific restrictions.

[0018] As disclosed in the background technology, users may encounter problems while the network service system is providing services. In such cases, users can communicate with the network service system's customer service to resolve their questions. To improve response efficiency and reduce response costs, network service systems typically prioritize using automated customer service systems instead of human customer service representatives to communicate with users.

[0019] Currently, intelligent customer service in online service systems typically relies on question-and-answer knowledge bases to determine response messages. For example, a question-and-answer knowledge base contains multiple question-and-answer pairs, each containing a question and its corresponding response. When the online service system receives an inquiry message from a user, it searches the question-and-answer knowledge base for question data that matches the text content of the inquiry message, and uses the corresponding response data as the response message. If no matching question data is found, the inquiry message is transferred to a human customer service representative for processing.

[0020] In real-world customer service scenarios, users often send inquiries via images, meaning the messages sent by the user include images. Due to limitations of the question-and-answer knowledge base, images cannot be directly matched with the question data. As a result, inquiries containing images can only be transferred to human customer service for processing, which greatly affects the efficiency of responding to inquiries. Furthermore, a large number of transfers to human customer service will increase the response cost.

[0021] To address the aforementioned issues, the inventors conceived of constructing a dedicated knowledge base (as shown in the graph knowledge base) for images in consultation messages. This graph knowledge base would then be used to match the corresponding knowledge data (e.g., matching the response data corresponding to the image) to generate a response message. Furthermore, existing knowledge bases typically rely on manual construction and maintenance; therefore, the key to solving this problem lies in how to automatically construct a graph knowledge base for responding to consultation messages containing images.

[0022] To address this, the inventors conducted further research and proposed a solution for constructing a graph knowledge base. The basic idea is to acquire sample consultation images, extract their image features, and use an intelligent recognition model to identify the image content they contain. Knowledge data is then generated based on the image content, and the image features of the sample consultation images and the generated knowledge data constitute the knowledge entries in the graph knowledge base. When the image features in these knowledge entries match the target image in a user's consultation message, a response message is generated based on the knowledge data in these knowledge entries. This embodiment specifically constructs a graph knowledge base based on the mapping of image features and knowledge data to respond to consultation messages containing images. Upon receiving a consultation message containing a target image, the knowledge data associated with the target image can be quickly retrieved through image feature matching, and a response message is generated based on the retrieved knowledge data, eliminating the need for manual customer service processing and improving the response efficiency of consultation messages. Furthermore, this embodiment utilizes an intelligent recognition model to identify the image content of the sample consultation images and generates knowledge data in the knowledge entries based on the image content, ensuring a high correlation between the knowledge data and the target image, thereby guaranteeing the accuracy of the final generated response message.

[0023] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0024] Figure 1 This is a flowchart of an embodiment of a knowledge base construction method provided in this application. The technical solution of this embodiment can be executed by a processing end, which can be a server in an online network service system, or it can be other nodes independent of the server in the online network service system.

[0025] In practical applications, online network service systems typically consist of a user terminal, a client terminal, and a server terminal. The user terminal and client terminal establish connections with the server terminal via a network. The network provides the medium for communication links between the user terminal, client terminal, and server terminal. Networks can include various connection types, such as wired and wireless communication links or fiber optic cables, etc.

[0026] The customer service client can be operated by human customer service personnel to respond to inquiries sent by users. The user client can be operated by consumers to initiate inquiries based on problems arising from their operation of the online network service system.

[0027] The client and client can interact with the server via the network to receive or send messages. For example, the client can sense the input and sending operations of human customer service personnel and send a response message to the server for the inquiry message, which the server then forwards to the client. The client can sense the input operations performed by the consumer and send the corresponding inquiry message to the server. The server can either generate a response message for the inquiry message and send it back to the client, or forward the inquiry message to the client and forward the response message from the client to the client.

[0028] The user or client side can be a browser, an app (application), a web application such as an H5 (HyperText Markup Language 5) application, a mini-program (also known as a lightweight application), or a cloud application. The user or client side can be deployed on electronic devices and depends on the device or certain apps on the device to run. Electronic devices can have displays and support information browsing, such as personal mobile terminals like mobile phones, tablets, personal computers, desktop computers, smart speakers, smartwatches, etc.

[0029] The aforementioned processing or server may include servers that provide various services, such as servers that provide intelligent customer service consultation to users, or servers that build the graph knowledge base on which the response consultation messages depend.

[0030] It should be noted that the processing or server end can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. The server can also be a server in a distributed system, or a server combined with blockchain. The server can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.

[0031] Figure 1 The knowledge base construction method shown may include the following steps: S101, Obtain sample consultation image.

[0032] The sample consultation image can be any image used to characterize operational problems of the network service system. For example, it can be a screenshot of a webpage that allows users to operate the network service system, or it can be a consultation image sent by a user when they encountered an operational problem (i.e., a historical consultation image). In this embodiment, the number of sample consultation images is usually multiple.

[0033] In some embodiments, there are many ways to obtain sample consultation images in this embodiment. For example, one possible method is to take a screenshot of the target webpage provided by the network service system to obtain the sample consultation image. The target webpage can be any webpage provided by the network service system; it can also be a functional webpage in the network service system where operational questions are likely to arise. For example, if the network service system is an e-commerce system, the functional pages could be order placement pages, settings pages, and payment pages, etc. Specifically, since the user terminal of the network service system can be in various forms such as APP, web application, and lightweight application, this embodiment can take screenshots of the target webpage (such as all pages or functional pages) displayed by the network service system for each type of user terminal to obtain the sample consultation image. This embodiment uses the method of taking screenshots of the target webpage of the network service system to obtain the sample consultation image, so that the knowledge entries of the graph knowledge base cover the target webpage of the network service system as much as possible.

[0034] Another possible approach is to determine historical consultation images from the historical session records corresponding to the network service system and use these historical consultation images as sample consultation images. The historical session records can be a complete record of a series of session messages generated during the user's historical interactions with customer service, stored in the network service system. Optionally, this embodiment can select historical session records corresponding to the user's historical interactions with human customer service, so that when subsequently determining the knowledge data corresponding to the historical consultation image, it can be combined with the human customer service's response message to that historical consultation image. Historical consultation images can be images related to the services provided by the network service system contained in the historical session records, such as screenshots of web pages from the network service system sent by the user. Specifically, it can be possible to traverse the historical session records corresponding to the network service system and search for historical consultation images (such as screenshots of web pages from the network service system) sent by the user that are related to the services provided by the network service system, using these as sample consultation images. This embodiment uses historical consultation images from historical session records as sample consultation images to ensure that the constructed graph knowledge base meets the user's actual consultation needs.

[0035] Another possible approach is to combine the two approaches mentioned above. This involves obtaining a sample consultation image by taking screenshots of the target webpage of the network service system, and then further supplementing the sample consultation image by collecting historical consultation images from historical session records. This ensures the comprehensiveness and diversity of the knowledge entries in the constructed graph knowledge base.

[0036] In some embodiments, this embodiment may be triggered at set intervals (e.g., 1 month) or when a webpage update event is detected in the network service system to obtain sample consultation images and perform subsequent operations, continuously improving and supplementing the knowledge entries in the graph knowledge base, so as to achieve regular maintenance and updates of the graph knowledge base and ensure the timeliness of the knowledge entries in the graph knowledge base.

[0037] S102, Extract image features from sample consultation images.

[0038] Here, the so-called image feature can be a feature representation obtained by extracting features from the sample consultation image to characterize the semantic content of the sample consultation image. For example, the image feature can be a multi-dimensional graph vector of the sample consultation image.

[0039] Optionally, this embodiment may employ a feature extraction algorithm to extract features from the sample consultation image to obtain its image features. Alternatively, it may use a convolutional neural network (such as MobileNetV3) to perform convolution and vectorization processing on the sample consultation image to obtain its image features. It should be noted that each sample consultation image has its unique corresponding image features; that is, the image features can uniquely represent the corresponding sample consultation image.

[0040] S103 uses an intelligent recognition model to identify the image content of sample consultation images.

[0041] The so-called image content can be the information obtained after multi-dimensional content analysis of the sample consultation image, used to fully represent the core content of the image. For example, the image content may include, but is not limited to, image description information, text information, and service scenarios of the sample consultation image. The image description information can be a structured description of the subject, text, layout, and other information in the sample consultation image. The text information can be the text information contained in the sample consultation image. The service scenario can be the service category corresponding to the sample consultation image in the online service system. For example, if the online service system is an e-commerce system, the service scenario can include payment, order, registration, etc.

[0042] In this embodiment, the sample consultation image can be input into an intelligent recognition model, which then performs multimodal information parsing on the sample consultation image to obtain the image content of the sample consultation image.

[0043] In one embodiment, this step may involve using an intelligent recognition model to generate image description information for the sample consultation image and extracting the text information contained in the sample consultation image. Specifically, this embodiment may use an intelligent recognition model to identify the image description information of the sample consultation image, such as recognizing the subject, text, and layout of the sample consultation image, and describing the recognition results using natural language to obtain image description information. It may also use OCR (Optical Character Recognition) technology to identify the text information contained in the sample consultation image. This embodiment can directly use the obtained image description information and text information as image content, or it can further perform reasoning and judgment processing on the image description information and text information to determine the service scenario of the sample consultation image, and integrate the model's image description information and text information description content as the image content of the sample consultation image. For example, this embodiment may use the sample consultation image as input data, and then further combine it with one or more conditional information such as content recognition instructions, role settings, content recognition requirements (such as constraints, data output requirements, etc.), thought chains, and example data to generate prompt instructions. The content recognition instructions are used to clearly tell the model what it needs to do, such as "understand the content information contained in the image." Role setting instructs the model to assume specific roles, altering its professionalism, for example, "You are a customer service summary assistant." Content recognition requirements define constraints, such as "You must not fabricate information that does not exist in the image." Data output requirements specify the format or content of the model's output, such as "The description should be concise and not exceed 20 words." The thought chain guides the model through reasoning, including elements like "the reasoning logic for image content parsing." Sample data provides learning samples to help the model understand the operations performed. The generated prompts are input into the intelligent recognition model, which identifies the image description information and text information contained within the sample consultation image. It then performs reasoning and judgment processing on the image description information and the text information to generate the image content of the sample consultation image.

[0044] The various intelligent models involved in this article, including the intelligent recognition model in this embodiment and the intelligent evaluation model, intelligent filtering model, message generation model, etc. that may appear in the corresponding embodiments below, can be language models (LM) or multimodal models (MM) based on artificial intelligence, with the goal of meeting actual needs.

[0045] In one implementation, such as Figure 2As shown, the intelligent recognition model in this embodiment may include, for example, an input layer (Input) 10, an encoder (Encoder) 20, a decoder (Decoder) 30, and an output layer (Output) 40. It may also include a self-attention layer and a feed-forward neural network, etc., and this application does not impose any limitations on this. The input layer 10 is used to receive data input, including text, images, or multimodal data. The encoder 20 is mainly used to convert the input data (usually in sequence form) into a vector representation. This process can incorporate the semantic features of the input data. The decoder 30 is responsible for converting the intermediate representation generated by the encoder 20 into output data (usually in sequence form). The output layer 40 is used to output data, such as layout structure information in JSON format. A self-attention layer is a mechanism that allows a model to focus on other positions in a sequence to better encode information about the current position. A feedforward neural network can perform nonlinear transformations on the output of the self-attention layer to enhance the model's expressive power. The various parts work together to enable the model built on them to perform well in a variety of complex processing tasks, such as natural language processing, computer vision, speech recognition, machine translation, text summarization, and intelligent question answering.

[0046] S104, Based on the image content, generate knowledge data corresponding to the image content.

[0047] The knowledge data corresponding to the image content can be data that supports the solution of related questions and the expansion of domain knowledge for the sample consultation image. For example, the knowledge data may include question data related to the sample consultation image and the corresponding response data; it may also include domain knowledge associated with the sample consultation image.

[0048] There are many ways to generate knowledge data corresponding to the image content in this step, and no limitation is imposed on this method. For example, it can be based on the image content, searching for knowledge data matching the image content from the local knowledge base of the network service system or the network; or it can be based on a knowledge generation model combined with the image content to automatically generate the corresponding knowledge data. If the sample consultation image is a historical consultation image, the knowledge data corresponding to the image content can also be extracted from the historical session messages corresponding to the sample consultation image (such as historical session messages associated with the context of the sample consultation image).

[0049] In some embodiments, if the knowledge data of this embodiment includes question data and response data corresponding to the question data, one implementation method is to search for question data and response data corresponding to the question data that match the image content from a question-and-answer knowledge base based on the image content.

[0050] The question-and-answer knowledge base can be a knowledge base containing multiple question-and-answer knowledge pairs. Each question-and-answer knowledge pair includes a question and its corresponding response. In this embodiment, the question-and-answer knowledge base can be traversed to find at least one question-and-answer knowledge pair that matches the image content. In this case, the question and its corresponding response for each question-and-answer knowledge pair constitute one knowledge data point corresponding to the image content. For example, if the image content is "The image is an after-sales page containing text such as 'refund,' 'return and refund,' and 'price protection,'" then the question and / or response data in the knowledge base are searched for question-and-answer knowledge pairs related to "after-sales," "refund," "return and refund," and "price protection." This yields multiple question data points matching the image content and their corresponding response data. This implementation utilizes an existing question-and-answer knowledge base to determine the standardized question and response data corresponding to the sample consultation image, which not only reduces the difficulty of generating knowledge data but also ensures the accuracy of the knowledge data. It should be noted that, based on image content, the operation of searching for question data and corresponding response data that match the image content from the question-and-answer knowledge base can be implemented by running an intelligent search model or by running a knowledge search script, and there is no limitation on which one is implemented.

[0051] Another approach is to generate question data for sample consultation images based on image content, and then search for corresponding response data from a question-and-answer knowledge base. Specifically, an intelligent question generation model can be used to extract potential question data from the sample consultation images based on the image content, and then the corresponding response data for each extracted question can be found in the question-and-answer knowledge base, thus obtaining the question data corresponding to the image content and its corresponding response data. It should be noted that for question data for which no response data can be found in the question-and-answer knowledge base, the response data can be manually added, or the question data can be discarded; there are no restrictions on this.

[0052] Another implementation method involves retrieving the contextual message corresponding to the sample consultation image from historical consultation records when the sample consultation image is a historical consultation image. Then, using an intelligent generation model, based on the image content and the contextual message, question data corresponding to the image content and response data corresponding to the question data are generated. Specifically, this embodiment can retrieve the contextual message of the sample consultation image from the historical consultation records corresponding to the network service system. This contextual message includes at least a response message for the sample consultation image, such as a response message from a human customer service representative replying to the sample consultation image via a client. Then, using the intelligent generation model, based on the image content and the contextual message, the consultation question (i.e., question data) corresponding to the image content of the sample consultation image and the response data corresponding to the question data are inferred and summarized. For example, response data can be summarized based on the contextual message, and question data corresponding to the response data can be summarized based on the response data and the image content. For example, the intelligent generation model can be used in conjunction with the following prompt to achieve the summary operation of question data and response data. This prompt can be generated based on input data, further combined with operation instructions, role settings, summary requirements (such as constraints, data output requirements, etc.), thought chains, and one or more conditional information from the sample data. The input data can include "image content and contextual messages." Operation instructions explicitly tell the model what to do, such as "based on image content and contextual messages, infer at least one set of consultation questions and their corresponding response data." Role setting can instruct the model to play a specific role, changing its level of expertise, such as "You are an intelligent customer service decision engine with multimodal understanding capabilities." Summary requirements specify constraints and data output requirements; for example, constraints may include "Do not fabricate information that does not exist in the image, do not disclose sensitive field information," and data output requirements may include "No vague descriptions, prohibit the output of personal opinions, etc." The thought process can guide the model through step-by-step reasoning, such as including "Guiding the model to summarize the reasoning logic of question and response data." Sample data can provide learning samples to help the model understand the operations performed.

[0053] This implementation method generates question data and response data by summarizing image content and contextual messages, so as to supplement and improve the knowledge data of sample consultation images, and further improve the knowledge comprehensiveness and accuracy of the constructed graph knowledge base.

[0054] Another possible approach is to combine at least two of the above methods to generate knowledge data corresponding to the image content of the sample consultation image. Optionally, the method for generating the corresponding knowledge data can be determined based on how the sample consultation image was obtained. For example, if the sample consultation image was obtained by taking a screenshot of the target webpage, the knowledge data can be generated using the first two methods; if the sample consultation image was obtained from historical conversation messages, the knowledge data can be generated using the third method. This embodiment uses different approaches to generate the corresponding response data for sample consultation images obtained in different ways, ensuring the accuracy and comprehensiveness of the generated knowledge data.

[0055] S105, knowledge entries are formed from the image features of the sample consultation image and knowledge data.

[0056] For any given sample consultation image, its corresponding image features are unique, but the knowledge data corresponding to its image content may be one or more. For example, for a sample consultation image, multiple knowledge data may be generated, each knowledge data corresponding to a question and a response.

[0057] This embodiment can be used to construct a knowledge entry corresponding to any sample consultation image by combining its image features with each piece of knowledge data. The knowledge entry in this embodiment can be represented in vector form. For example, if the image feature of sample consultation image A is graph vector A, and question 1 and answer 1, as well as question 2 and answer 2, are generated based on the image content of sample consultation image A, then two knowledge entries can be constructed. Knowledge entry 1 can be (graph vector A, question 1, answer 1), and knowledge entry 2 can be (graph vector A, question 2, answer 2). It should be noted that the image features (as shown by the graph vector) in the knowledge entry of this embodiment are used to determine whether the knowledge entry matches the target image in the consultation message during the actual conversation interaction between the intelligent customer service and the user; the knowledge data is used to generate a response message when the image features of any knowledge entry match the target image of the consultation message.

[0058] In some embodiments, a sample consultation image may correspond to multiple knowledge entries. In this case, to further filter out more accurate knowledge entries to generate response data, the content recognition result of the sample consultation image can be added to the knowledge entries. This content recognition result can be the image content obtained by recognizing the sample consultation image using an intelligent recognition model, or it can contain only a portion of the image content, such as service scenarios or text information. In this case, the image features of the sample consultation image, the content recognition result, and the knowledge data constitute a knowledge entry. The content recognition result is used to filter the target knowledge entry corresponding to the response message when the image features of multiple knowledge entries match the target image in the consultation message. For example, if the image feature of sample consultation image A is graph vector A, the content recognition result is image description content A, and the service scenario A, and question 1 and answer 1 are generated based on the image content of sample consultation image A, then a knowledge entry corresponding to sample consultation image A can be (graph vector A, image description content A, service scenario A, question 1, answer 1). Based on this, if the target image contained in the consultation message sent by the user terminal matches the image features of multiple knowledge items, the target knowledge item for generating the response message can be further selected from the multiple knowledge items based on the content recognition results of the multiple knowledge items, so as to improve the accuracy of the response message. The specific selection method will be described in detail in subsequent embodiments.

[0059] S106, Construct a graph knowledge base based on knowledge entries.

[0060] Specifically, when the image features in the knowledge entries are used to match the target image in the consultation message, a response message is generated based on the knowledge data in the knowledge entries for the consultation message. The process of generating a response message for a consultation message based on the constructed graph knowledge base will be described in detail in subsequent embodiments.

[0061] In this embodiment, the knowledge entries formed by each sample consultation image in S105 can be added to the graph knowledge base to realize the construction of the graph knowledge base.

[0062] This embodiment acquires sample consultation images, extracts their image features, and uses an intelligent recognition model to identify the image content they contain. Knowledge data is generated based on the image content, and the image features of the sample consultation image and the generated knowledge data constitute knowledge entries in a graph knowledge base. The image features in these knowledge entries are used to generate a response message based on the knowledge data in the user's consultation message, provided they match the target image. This embodiment specifically constructs a graph knowledge base based on the mapping of image features and knowledge data to respond to consultation messages containing images. Upon receiving a consultation message containing a target image, it can quickly recall the knowledge data associated with the target image through image feature matching and generate a response message based on the recalled knowledge data, eliminating the need for manual customer service processing and improving the response efficiency of consultation messages. Furthermore, this embodiment utilizes an intelligent recognition model to identify the image content of the sample consultation image and generates knowledge data in knowledge entries based on the image content, ensuring a high correlation between the knowledge data and the target image, thereby guaranteeing the accuracy of the final generated response message.

[0063] In practical applications, sample consultation images may contain target areas. These target areas can be prominent features in the image, such as areas highlighted by the user (e.g., areas indicated by a box, arrow, or arrow) or other markers on the page. In this scenario, this embodiment utilizes an intelligent recognition model to identify the image content of the sample consultation image, including: identifying the target areas within the image and recognizing the image content of those target areas. Specifically, the intelligent recognition model first identifies the coordinates of the target areas in the sample consultation image, and then, based on these coordinates, performs image content recognition on the target areas (the specific recognition method has been described in the above embodiment) to obtain the image content of the target areas. This embodiment, for sample consultation images containing target areas, can generate targeted question and response data based on the image content of the target areas, improving the correlation between knowledge data and target areas. This allows for more accurate generation of response messages for subsequent consultation images containing target areas.

[0064] In some embodiments, this embodiment can not only periodically execute operations S101-S106 to expand the knowledge entries in the graph knowledge base, but also identify expired knowledge entries in the graph knowledge base and delete them. Expired knowledge entries may be knowledge entries with low accuracy, or knowledge data corresponding to invalid web pages after service upgrades, etc. This embodiment can periodically identify expired knowledge entries and delete them from the graph knowledge base to achieve continuous updates to the graph knowledge base, realize a self-closing loop of knowledge entries in the graph knowledge base, and continuously improve the quality of knowledge entries in the graph knowledge base.

[0065] Optionally, the method for determining expired knowledge entries in this embodiment includes at least one of the following: Method 1: Define knowledge entries in the graph knowledge base that have not participated in any session message generation operations within a preset time period as expired knowledge entries. Specifically, periodically count knowledge entries in the graph knowledge base that have not matched the target image contained in the consultation message within a preset time period (e.g., within 3 months). These are knowledge entries corresponding to sample consultation images that the user has not consulted; these are considered expired knowledge entries.

[0066] Method 2: Obtain user feedback data for any response message, and if the user feedback data determines that the response message does not meet the satisfaction requirements, designate the knowledge entry that generated the response message as an expired knowledge entry. Specifically, in this embodiment, after sending a response message for an inquiry message containing a target image, a satisfaction feedback prompt can be further sent. This satisfaction feedback prompt is used to prompt the user to reply with feedback data (i.e., user feedback data) regarding whether the currently received response message resolved their problem. For example, the satisfaction feedback prompt could be "Please rate the satisfaction of this reply" or "Did this reply resolve your problem?". In this embodiment, user feedback data on the user's reply to the response message can be obtained. Based on this user feedback data, it can be determined whether the response message meets the satisfaction requirements. If not, it indicates that the quality of the knowledge entry used to generate the response message is not high and needs to be deleted as an expired knowledge entry. If yes, it indicates that the knowledge entry is a high-quality knowledge entry and can be retained. For example, if the user feedback data is a satisfaction rating for the response message, the knowledge entry that generated the response message can be designated as an expired knowledge entry if the satisfaction rating is less than a first preset value. If the user feedback data indicates that the problem is unresolved, the knowledge entry that generated the response message can be considered an expired knowledge entry.

[0067] Method 3: Considering that user feedback data may be subjective and may not accurately reflect the quality of the response message, or that some users may not cooperate in replying to user feedback data, this embodiment can also use an intelligent evaluation model to evaluate the quality of any response message, and if the evaluation result is that the response message does not meet the quality requirements, the knowledge entry that generated the response message will be regarded as an expired knowledge entry.

[0068] Specifically, in this embodiment, any response message, the consultation message of the response message, and user feedback data on the response message can be used as input data. Further, one or more conditions such as quality assessment instructions, role settings, quality assessment requirements (such as constraints, data output requirements, etc.), assessment logic, and sample data are combined to generate prompt instructions. The generated prompt instructions are then input into the intelligent assessment model, which performs a quality assessment on the response data. If the assessment result indicates that the response message does not meet the quality requirements, the knowledge entry that generated the response message is designated as an expired knowledge entry.

[0069] It should be noted that for methods two and three above, the intelligent customer service can perform the operation of determining expired items once after each time it sends a response message to an inquiry message containing a target image, or it can perform the operation of determining expired items once every preset time interval for the intelligent customer service to send a response message to an inquiry message containing a target image within that time interval.

[0070] This embodiment uses one or more of the above three methods in combination to determine expired knowledge entries, which improves the accuracy and timeliness of the determination of expired knowledge entries, thereby improving the quality of knowledge entries in the graph knowledge base.

[0071] Regarding Method 3 above, this embodiment can specifically utilize an intelligent evaluation model to perform quality evaluation on any response message in the following manner: Using the intelligent evaluation model, for any response message, based on the consultation message corresponding to the response message, the response message, and user feedback data for the response message, determine the quality score corresponding to the response message in multiple quality dimensions; and determine the target score of the response message based on the weight values ​​and quality scores corresponding to the multiple quality dimensions.

[0072] In this embodiment, the multiple quality dimensions may include, but are not limited to: problem resolution (i.e., whether the problem is closed-loop and whether the user is satisfied), response timeliness (i.e., whether the conversation rhythm is smooth and whether there is a long wait), service attitude (i.e., whether it is polite and patient), communication clarity (i.e., whether the explanation is clear and easy to understand), emotional empathy (i.e., whether it recognizes and soothes the user's emotions), and process standardization (i.e., whether it follows the standard service process), etc. Specifically, this implementation method may further add requirements for multiple quality dimensions to the above prompt instructions. For example, the prompt instructions at this time can be generated based on input data, further combined with scoring instructions, role settings, scoring requirements (such as the requirements of the above multiple quality dimensions, data output requirements, etc.), thought chain, and one or more conditional information in the example data. The input data may include "the consultation message corresponding to the response message, the response message, and user feedback data for the response message". The scoring instructions are used to clearly tell the model what it needs to do, such as "to score the quality of the response message based on the response message, consultation message, and user feedback data", etc. Role settings can instruct the model to play a specific role to change its professionalism, etc. For example, "You are a quality inspection expert for conversation messages", etc. A thought chain can be used to guide the model through step-by-step reasoning, such as "the reasoning logic of multi-dimensional quality judgment." Sample data can be used to provide learning samples to help the model understand the operations it performs.

[0073] The prompt instruction is input into the intelligent evaluation model to control the model to evaluate the response message from multiple quality dimensions based on the prompt instruction. This yields a quality score for each quality dimension, and then, based on the weight values ​​corresponding to each dimension, a weighted sum of these scores is calculated to obtain a comprehensive score for the response message, i.e., the target score, which serves as the evaluation result. If the target score is less than a second preset value, the response message is determined to have failed to meet quality requirements, and the knowledge entry that generated the response message is designated as an expired knowledge entry. This embodiment comprehensively evaluates the response message from multiple quality dimensions, making the evaluation results closer to the user's actual feedback, thereby ensuring the accuracy of the expired knowledge entries determined based on the evaluation results.

[0074] Figure 3The flowchart illustrates an implementation method for acquiring sample consultation images and constructing a graph knowledge base in a practical application scenario, including the following steps: The processing end performs fully automated screenshot processing on all functional web pages provided by the network service system's website and application to obtain functional web page screenshots (i.e., sample consultation images) (S301). For example, this can be done through screenshot tools or the screenshot function of automated testing tools. Then, an intelligent recognition model is used to analyze the image content of the functional web page screenshots, which may include service scenarios and image description information, recognized text information, etc. (S302), generating a set of mapping relationships (Image 1 (i.e., image identifier), Scenario 1, Image Description 1) (S303). Then, a feature extraction model (such as MobileNetV3) is used to extract the graph vectors of the functional web page screenshots through convolution and vectorization processing (S304), and the mapping relationship is upgraded to (Graph Vector 1, Scenario 1, Image Description 1) (S305). Then, based on the image description information in the mapping relationship, the standard questions and answers corresponding to the image description information are retrieved in the question-and-answer knowledge base (S306). The mapping relationship is then filled with knowledge to form multiple mappings (i.e. multiple knowledge entries), namely (graph vector 1, scene 1, image description 1, question 1, answer 1) ... (graph vector 1, scene 1, image description 1, question n, answer n) (S307). A graph knowledge base is constructed based on the formed multiple mappings (S308).

[0075] To ensure the comprehensiveness and accuracy of knowledge entries in the graph knowledge base, this embodiment... Figure 3 Based on the method shown, further methods can be introduced. Figure 4 The graph knowledge base construction method shown is used to expand the knowledge entries in the graph knowledge base. Specifically, it may include the following steps: Filtering historical session records containing user screenshots in the network service system (S401), where the user screenshot in this embodiment refers to a screenshot of a webpage expressing a user's question about webpage operation. Extracting historical consultation images (i.e., sample consultation images) sent by the user (S402), and using an intelligent recognition model to analyze the image content of the historical consultation images (S403), where the image content may also include: service scenario and image description information, recognized text information, etc. Using an intelligent generation model combined with the contextual messages of the sample consultation images and the image content recognized in S403, summarizing the questions and answers, obtaining a set of mapping relationships (image A, scenario A, content description A, question A, answer A) (S404), then using a feature extraction model to extract the graph vectors of the historical consultation images, updating the mapping relationships (graph vector A, scenario A, content description A, question A, answer A) (S405), and obtaining new knowledge entries to add to the graph knowledge base (S406).

[0076] To ensure the accuracy and timeliness of the graph knowledge base content, this embodiment may also employ... Figure 5The method shown involves periodically performing closed-loop optimization on knowledge entries in the graph knowledge base. Specifically, it includes the following three methods: Method 1: Trigger the above execution periodically (e.g., every 3 days). Figure 3 The operations of S301-S308 shown above, and the above Figure 4 The operations shown in S401-S406 (S501) generate new knowledge entries (S502) and add the newly generated knowledge entries to the graph knowledge base (S503), responding to rapid changes in services by periodically adding new knowledge entries.

[0077] Method 2: Periodically trigger an inspection of each knowledge entry in the graph knowledge base (S504), and count whether the knowledge entry has been recalled within a preset period (e.g., 3 months), i.e. whether it has participated in the session message generation operation (S505). If not, it can be considered that the service of the network service system has changed, and the knowledge entry is marked as an expired knowledge entry (S506), and the expired knowledge entry is deleted from the graph knowledge base (S507). If it has been recalled, the knowledge entry is retained (S508).

[0078] Method 3: Using online inspection of chat messages, monitor received inquiry messages containing screenshots (i.e., target images) (S509). Extract the inquiry message containing the screenshot, its response message, and user feedback data regarding the response message (S510). Call the intelligent evaluation model to analyze the response message across multiple quality dimensions based on the extracted inquiry message, response message, and user feedback data, obtaining quality scores for each quality dimension (S511). Based on the quality scores and weights of multiple quality dimensions, calculate the satisfaction score (i.e., target score) (S512). Determine if the satisfaction score is less than a first preset value (e.g., 60) (S513). If so, mark the knowledge entry that generated the response message as an expired knowledge entry and delete it from the graph knowledge base (S514). If not, retain the knowledge entry that generated the response message (S515). This embodiment uses... Figure 5 The knowledge closed-loop optimization strategy is illustrated. Through mechanisms such as periodic automated webpage screenshotting and analysis, historical conversation message mining, and automated online conversation message inspection, the graph knowledge base is continuously optimized, expired knowledge is deleted, and the knowledge quality of the graph knowledge base is continuously improved, that is, the accuracy and timeliness of the knowledge entries in the graph knowledge base are improved.

[0079] This embodiment employs automated screenshotting of all functional web pages of a website or application. An intelligent recognition model simultaneously performs OCR, target recognition, topic and intent inference, and service scenario classification, producing multi-dimensional semantic structured annotations (service scenario and image content description) to form a searchable semantic layer, significantly reducing manual annotation costs. Furthermore, this solution expands the knowledge entries of the graph knowledge base through weakly supervised knowledge mining of historical session records. Specifically, for historical session messages containing screenshots, combined with image analysis and context, "question-answer" pairs are automatically extracted and quantitatively added to the knowledge base, enabling the knowledge base to grow automatically according to the distribution of actual questions, reducing cold start and bias. In addition, this embodiment introduces a self-healing and closed-loop governance mechanism for service drift, incorporating... Figure 5 The diagram illustrates a three-tiered quality signal-driven knowledge evolution approach. This involves: deleting expired knowledge entries based on participation in response message operations; periodically re-screening and rebuilding the index to adapt to service or webpage upgrades in the network service system; and performing multi-dimensional weighted quality checks on dialogues containing images, with low-scoring samples triggering the removal of associated knowledge entries to prevent errors from becoming entrenched.

[0080] Figure 6 This is a flowchart illustrating an embodiment of a message processing method provided in this application. The technical solution of this embodiment can be executed by a processing end, which can be a server in an online network service system, or it can be another node independent of the server in the online network service system. Specifically, the processing end can execute based on the graph knowledge base constructed in the above embodiment. Figure 6 The message processing method shown may include the following steps: S601, retrieve the inquiry message sent by the user client.

[0081] Specifically, when a network service system provides services to users, users may encounter problems. At this time, users can access the network service system's consultation page through their client. For example, clicking the "Customer Service Consultation" icon triggers access to the consultation page. The client can respond to the user's input on the consultation page, obtain the consultation message entered by the user, and send the consultation message to the processing terminal. The corresponding processing terminal will obtain the consultation message sent by the user.

[0082] S602, when the consultation message includes a target image, extract the target image features of the target image.

[0083] In this embodiment, the target image can be an image representing the content of a user's inquiry. For example, it can be a screenshot of a webpage from a web service system, or an image after editing a webpage screenshot (such as adding a label box).

[0084] Specifically, after receiving an inquiry message, it can be determined whether the inquiry message contains an image (i.e., a target image). If it does, the target image features of the target image can be extracted in accordance with the method for extracting image features of the sample inquiry image described in the above embodiments.

[0085] If the target image is not included, a response message can be generated and sent to the user, following the same processing method as for text messages. For example, the system could search for matching question data from a question-and-answer knowledge base, retrieve the corresponding response data, and generate a response message to send to the user.

[0086] S603, based on the characteristics of the target image, search for target knowledge entries that match the target image in the graph knowledge base.

[0087] The graph knowledge base includes multiple knowledge entries; each knowledge entry includes image features of a sample consultation image and knowledge data; the knowledge data is generated based on the image content of the sample consultation image, which is obtained by recognizing the image from the sample consultation image using an intelligent recognition model. The construction and maintenance process of the graph knowledge base has been described in the above embodiments and will not be repeated here.

[0088] Optionally, this embodiment may involve traversing each knowledge entry in the graph knowledge base, calculating the similarity between the target image features and the image features in each knowledge entry (e.g., calculating the cosine similarity of the features), and selecting one or more knowledge entries with a similarity greater than a third preset value (e.g., 0.9) as the target knowledge entry.

[0089] It should be noted that if there is no target knowledge entry in the graph knowledge base with a similarity greater than the third preset value, the consultation message containing the target image can be forwarded to the customer service terminal so that the human customer service personnel can reply to the consultation message through the customer service terminal. For example, the customer service terminal will display the consultation message and then respond to the input operation of the human customer service personnel to obtain the response message input by the human customer service personnel in response to the consultation message, and feed the response message back to the processing terminal, which will further forward the response message to the user terminal.

[0090] In practical applications, each sample consultation image in the graph knowledge base may contain multiple knowledge data. Based solely on the image feature matching method described above, multiple target knowledge entries with similarity greater than a third preset value may appear. Since target knowledge entries are used to generate response messages, if there are multiple of them, it may affect the accuracy of the response messages. To more accurately filter target knowledge entries, this embodiment may search for at least one candidate knowledge entry that matches the target image in the graph knowledge base based on the target image features; if the number of the at least one candidate knowledge entry is one, the candidate knowledge entry is used as the target knowledge entry; if the number of the at least one candidate knowledge entry is multiple, at least one historical session message corresponding to the consultation message is obtained, and the target knowledge entry is determined from the multiple candidate knowledge entries based on the at least one historical session message.

[0091] Specifically, the process can begin by calculating image feature similarity as described above, searching the graph knowledge base for at least one knowledge entry with a feature similarity greater than a third preset value to the target image as a candidate knowledge entry. Further, it can be determined whether there is only one candidate knowledge entry. If so, this candidate knowledge entry is directly used as the target knowledge entry. If there are multiple candidate knowledge entries, at least one conversation message between the user and the processing end prior to the current consultation message can be obtained as historical conversation messages. Based on these historical conversation messages, the target knowledge entry can be further filtered from the multiple candidate knowledge entries. For example, the user's consultation intent can be clarified by combining at least one historical conversation message and the target image, and then at least one candidate knowledge entry matching the user's consultation intent can be selected as the target knowledge entry. In this embodiment, when image features cannot accurately match knowledge entries, the historical conversation messages of the consultation message are further combined to accurately filter the target knowledge entry, ensuring the accuracy of the generated response message.

[0092] To further improve the comprehensiveness of candidate knowledge item selection, this embodiment may also include obtaining text information contained in the target image; and searching for at least one candidate knowledge item matching the target image from the graph knowledge base based on the target image features and the text information. Specifically, the text information contained in the target image may be extracted using a text extraction algorithm (such as an OCR algorithm) or a text extraction model. Optionally, if the target image contains a target region, the text information contained in the target region may be extracted. One implementation of searching for at least one candidate knowledge item matching the target image from the graph knowledge base based on the target image features and text information may be as follows: First, according to the similarity calculation method described in the above embodiment, search for knowledge items whose image features match the target image features from the graph knowledge base, as a part of the candidate knowledge items. Then, traverse each knowledge item in the graph knowledge base, determine whether the knowledge data is related to the text information (for example, determine whether the subject object of the knowledge data and the text information is the same, or determine whether the knowledge data contains keywords from the text information, etc.), and take the knowledge items corresponding to the knowledge data related to the text information as another part of the candidate knowledge items. This embodiment supports dual-channel indexing of candidate knowledge items through text information (i.e. semantic filtering) and target image features (e.g., vector recall), which balances the recall accuracy and interpretability of candidate knowledge items and improves the controllability of online consultation and response services.

[0093] When the knowledge entries also include the content recognition results of the sample consultation image, another implementation method is to search for at least one candidate knowledge entry in the graph knowledge base whose image features match the target image features; and to search for at least one candidate knowledge entry in the graph knowledge base whose content recognition results match the text information. Specifically, first, following the similarity calculation method described in the above embodiments, knowledge entries whose image features match the target image features are searched in the graph knowledge base as a part of the candidate knowledge entries. Then, each knowledge entry in the graph knowledge base is traversed to determine whether the content recognition results are related to the text information (e.g., whether the image content and the text information contain the same keywords), and the knowledge entries corresponding to the image content related to the text information are taken as another part of the candidate knowledge entries.

[0094] Optionally, if the knowledge entry includes content recognition results and the content recognition results include service scenarios, this embodiment can further filter the candidate knowledge entries after determining the candidate knowledge entries by combining the service scenarios in the content recognition results, such as candidate knowledge entries whose service scenarios are inconsistent with the target image service scenarios, so as to improve the accuracy of candidate knowledge entry determination.

[0095] In this embodiment, to further improve the accuracy of target knowledge item selection, an optional implementation of determining the target knowledge item from multiple candidate knowledge items based on at least one historical session message can be: using an intelligent filtering model, determining the session intent of the consultation message based on at least one historical session message, and determining the correlation between the target image and the multiple candidate knowledge items based on the session intent, at least one historical session message, and the content recognition results corresponding to each of the multiple candidate knowledge items; sorting the multiple candidate knowledge items based on the correlation, and determining the target knowledge item based on the sorting results.

[0096] Specifically, an intelligent filtering model can be used to perform semantic parsing on at least one historical session message to determine the session intent of the currently received session message. Combining the session intent, at least one historical session message, and the content recognition results of multiple candidate knowledge items, the correlation between the target image and the multiple candidate knowledge items can be determined. For example, for any candidate knowledge item, it can be analyzed whether the content recognition result matches the session intent, and whether the service scenario corresponding to the content recognition result is consistent with the service scenario corresponding to the historical session message, to determine the correlation between the target image and the multiple candidate knowledge items. Furthermore, when determining the correlation between the target image and the multiple candidate knowledge items, the relevance between the session intent and the knowledge data of the candidate knowledge items can be further considered. Then, the candidate knowledge items are sorted in descending order of correlation, and one or more of the top-ranked candidate knowledge items are selected as the target knowledge item.

[0097] For example, the intelligent filtering model in this embodiment can be used in conjunction with the following prompt to sort candidate knowledge items. This prompt can be generated based on input data, further combined with operation instructions, role settings, operation requirements (such as constraints, data output requirements, etc.), thought processes, and one or more conditional information from the example data. The input data can include "the target image and its corresponding at least one historical conversation message, and multiple candidate knowledge items." Operation instructions explicitly tell the model what to do, such as "based on the semantic analysis of historical conversation messages and the target image, determine the correlation between multiple candidate knowledge items and the target image, and reorder the multiple candidate knowledge items according to their correlation," etc. Role settings can instruct the model to play a specific role, changing its professionalism, etc. For example, "You are an intelligent customer service decision engine with multimodal understanding capabilities," etc. Operation requirements specify constraints, such as "Do not fabricate information that does not exist in the image, do not disclose sensitive field information," etc. Data output requirements are the requirements for the format or content of the model's output results. The thought chain can be used to guide the model through step-by-step reasoning, such as "the reasoning logic that guides the model to parse the conversational intent, and the reasoning logic that analyzes the correlation between the target image and multiple candidate knowledge items." Sample data can be used to provide learning samples to help the model understand the operations performed.

[0098] S604, Generate a response message for the inquiry message based on the knowledge data in the target knowledge item.

[0099] Optionally, if the knowledge data in the target knowledge entry consists of question data and its corresponding response data, the included response data can be directly used as the response message to the inquiry message. Alternatively, a message generation model can be used to further process the response data in the target knowledge entry (such as adjusting language style, increasing emotional empathy, etc.) to generate a response message to the inquiry message that conforms to natural speech description. If the knowledge data in the target knowledge entry is domain knowledge, the domain knowledge can be directly used as the response message to the inquiry message, or a message generation model can be used to further process the domain knowledge to generate a response message to the inquiry message that conforms to natural speech description.

[0100] In some embodiments, if there are multiple target knowledge items, the relevance of the knowledge data of the multiple target knowledge items (such as the relevance of corresponding question data) can be determined. If the relevance is greater than a fourth preset value, the knowledge data of the multiple target knowledge items can be integrated to generate a response message containing multiple knowledge data. If the relevance is less than or equal to the fourth preset value, the knowledge data of the multiple target knowledge items can be combined to generate multiple response messages for the inquiry message.

[0101] S605 sends a response message to the user terminal.

[0102] In this embodiment, the processing end can send one or more response messages generated by S704 to the user end through the intelligent customer service account.

[0103] This embodiment pre-constructs a graph knowledge base based on the mapping of image features and knowledge data. Upon receiving a query message containing a target image, it can quickly recall target knowledge entries that match the target image features from the graph knowledge base, achieving efficient knowledge retrieval. Furthermore, this embodiment can also generate response messages based on the recalled knowledge data to quickly reply to the user's query message.

[0104] In one embodiment, the consultation message may include not only the target image but also a consultation question described in text form, i.e., a text message; and the knowledge data includes question data and corresponding response data. In this case, for the target image in the consultation message, the method described in the above embodiment can be used to search for a target knowledge entry matching the target image in the graph knowledge base; for the text message, the response data corresponding to the text message can be searched in the question-and-answer knowledge base. For example, question data with a similarity greater than a fifth preset value to the text message can be searched in the question-and-answer knowledge base, and its corresponding response data can be obtained as the response data corresponding to the text message. Based on the response data contained in the target knowledge entry and the response data corresponding to the text message, a response message for the consultation message is generated. Specifically, one implementation method is to integrate the response data determined based on the target image and the response data corresponding to the text message (e.g., merging according to a certain strategy) to generate a response message for the consultation message. Another implementation method is to further filter the two response data and select the higher-quality response data to generate a response message.

[0105] In one embodiment, another possible way to generate a response message for a consultation message based on the response data contained in the target knowledge entry and the response data corresponding to the text message may include: determining whether the text message is associated with the target image; if so, determining target response data from the response data contained in the target knowledge entry and the response data corresponding to the text message based on at least one historical session message corresponding to the consultation message, and generating a response message for the consultation message based on the target response data; if not, generating a response message for the consultation message based on the response data corresponding to the text message and using the response data contained in the target knowledge entry as parameter data.

[0106] Specifically, determining whether a text message is relevant to a target image can be done by judging whether the main subject or service scenario of the text information in the text message and the target image are consistent. If they are consistent, they are considered relevant; otherwise, they are not. If they are relevant, the user's inquiry intent can be analyzed based on at least one historical conversation message corresponding to the inquiry message. Then, target response data that matches the user's inquiry intent can be determined from the response data contained in the target knowledge item and the response data corresponding to the text message. Then, according to the method of generating response messages based on response data described in the above embodiments, a response message for the inquiry message can be generated based on the target response data. For example, the target response data can be adjusted using a message generation model to output a response message that conforms to natural language description.

[0107] If they are unrelated, the response data corresponding to the text message can be used as the primary basis for generating the response message, while the response data contained in the target knowledge item can be used as reference data (i.e., secondary basis). The response data determined by both methods can be integrated to generate a response message specifically for the inquiry message. For example, a message generation model can be used to generate a response message that conforms to natural language description, primarily reflecting the response data corresponding to the text message, and briefly covering the response data corresponding to the target knowledge item, based on the response data corresponding to the text message and using the response data contained in the target knowledge item as reference data. Alternatively, both methods of determining the response data can be used as the content of the response message, but corresponding attention prompts can be added for different response data. These attention prompts can be used to guide the user on the order of attention for different response data. For example, assuming the response data corresponding to the text message is data 1 and the response data contained in the target knowledge item is data 2, the generated response message could be, "Please first check if data 1 can answer your question. If not, please check data 2."

[0108] This embodiment addresses consultation messages that simultaneously contain both target images and text messages. It not only analyzes the response data corresponding to both the target image and the text message, but also generates response messages using different methods based on the relevance between the target image and the text message, thereby improving the accuracy of the response messages.

[0109] In a real-world scenario of responding to user inquiries, this embodiment can be based on... Figure 7The method shown generates a response message for an inquiry message, specifically including the following steps: receiving an inquiry message containing a target image (S701), performing convolutional vectorization on the target image using a feature extraction model to obtain the target image's graph vector (i.e., target image features) (S702), traversing each knowledge entry in the knowledge base, calculating the cosine similarity between the graph vector in each knowledge entry and the graph vector of the target image (S703), filtering candidate knowledge entries with a similarity greater than a third preset value (e.g., 0.9) (S704), determining whether a candidate knowledge entry has been selected (S705), and if not, forwarding the inquiry message to the customer service terminal for a human customer service representative to reply with a response message, i.e., transferring the message to manual processing (S706). If candidate knowledge items are selected, the system further determines whether there are multiple candidate knowledge items (S707). If not, a response message for the inquiry message is generated directly based on the candidate knowledge item and sent to the user (S708). If there are multiple candidate knowledge items, at least one historical session message corresponding to the inquiry message is obtained as the current session context (S709). Using an intelligent filtering model, the multiple candidate knowledge items are sorted based on the current session context (S710), and one or more of the top-ranked candidate knowledge items are selected as the most suitable target knowledge item to generate the response message (S711). This embodiment enables rapid response to user inquiries via images, reduces user waiting time, and improves user satisfaction and user experience.

[0110] The detailed implementation methods and beneficial effects of each step in this embodiment have been described in detail in the foregoing embodiments, and will not be elaborated here.

[0111] It should be noted that some processes described in the above embodiments and accompanying drawings include multiple operations that appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order they appear in this document, or they may be executed in parallel. The operation numbers, such as S102, S103, etc., are merely used to distinguish different operations and do not represent any execution order. Furthermore, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should also be noted that the descriptions such as "first" and "second" in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.

[0112] Figure 8 This invention illustrates a structural schematic diagram of an embodiment of the knowledge base construction apparatus provided in this application. The apparatus includes: Image acquisition module 801 is used to acquire sample consultation images; The first feature extraction module 802 is used to extract image features from the sample consultation image; Image recognition module 803 is used to recognize the image content of the sample consultation image using an intelligent recognition model; The knowledge generation module 804 is used to generate knowledge data corresponding to the image content based on the image content; Entry construction module 805 is used to construct knowledge entries from the image features of the sample consultation image and the knowledge data; The knowledge base construction module 806 is used to construct a graph knowledge base based on the knowledge entries; when the image features in the knowledge entries match the target image in the consultation message, a response message for the consultation message is generated based on the knowledge data in the knowledge entries.

[0113] In some embodiments, the image acquisition module 801 is specifically used to take a screenshot of the target webpage provided by the network service system to obtain a sample consultation image; or, to determine a historical consultation image from the historical session records corresponding to the network service system and use the historical consultation image as a sample consultation image.

[0114] In some embodiments, the knowledge data includes question data and corresponding response data; the knowledge generation module 804 is specifically used to search for question data matching the image content and corresponding response data from a question-and-answer knowledge base based on the image content; or, based on the image content, generate question data for the sample consultation image and search for corresponding response data from the question-and-answer knowledge base; or, if the sample consultation image is a historical consultation image, obtain the context message corresponding to the sample consultation image from the historical session record, and use an intelligent generation model to generate question data corresponding to the image content and corresponding response data based on the image content and the context message.

[0115] In some embodiments, the image recognition module 803 is specifically used to identify a target region in the sample consultation image using an intelligent recognition model, and to identify the image content of the target region.

[0116] In some embodiments, the image recognition module 803 is further configured to generate image description information of the sample consultation image using an intelligent recognition model, and to extract text information contained in the sample consultation image.

[0117] In some embodiments, the entry construction module 805 is further specifically used to construct knowledge entries from the image features of the sample consultation image, the content recognition result, and the knowledge data; the content recognition result is used to filter the target knowledge entries corresponding to the response message when the image features in multiple knowledge entries all match the target image in the consultation message.

[0118] In some embodiments, the apparatus further includes an update module, configured to determine expired knowledge entries in the graph knowledge base and delete the expired knowledge entries from the graph knowledge base; the determination of expired knowledge entries includes at least one of the following methods: identifying knowledge entries in the graph knowledge base that have not participated in any session message generation operation within a preset time period as expired knowledge entries; obtaining user feedback data for any response message, and if the user feedback data determines that the response message does not meet satisfaction requirements, identifying the knowledge entry that generated the response message as an expired knowledge entry; using an intelligent evaluation model to perform quality evaluation on any response message, and if the evaluation result indicates that the response message does not meet quality requirements, identifying the knowledge entry that generated the response message as an expired knowledge entry.

[0119] In some embodiments, when the update module uses an intelligent evaluation model to perform quality evaluation on any response message, it specifically uses the intelligent evaluation model to determine the quality score of the response message in multiple quality dimensions based on the consultation message corresponding to the response message, the response message, and user feedback data for the response message; and determines the target score of the response message according to the weight values ​​and quality scores corresponding to the multiple quality dimensions.

[0120] Figure 8 The knowledge base construction device can execute Figure 1 The implementation principle and technical effects of the knowledge base construction method described in the illustrated embodiments will not be repeated here. The specific methods by which each module and unit of the knowledge base construction apparatus in the above embodiments perform operations have been described in detail in the embodiments related to this method, and will not be elaborated upon here.

[0121] Figure 9 This application provides a schematic diagram of the structure of a message processing apparatus according to one embodiment. The apparatus includes: The message acquisition module 901 is used to acquire inquiry messages sent by the user client; The second feature extraction module 902 is used to extract target image features of the target image when the consultation message includes a target image. The knowledge search module 903 is used to search for target knowledge entries that match the target image from a graph knowledge base based on the target image features; the graph knowledge base includes multiple knowledge entries; the knowledge entries include image features of the sample consultation image and knowledge data; the knowledge data is generated based on the image content of the sample consultation image, and the image content is obtained by recognizing the sample consultation image using an intelligent recognition model; The message generation module 904 is used to generate a response message for the consultation message based on the knowledge data in the target knowledge item; The message sending module 905 sends the response message to the user terminal.

[0122] In some embodiments, the knowledge lookup module 903 is specifically configured to: search for at least one candidate knowledge entry that matches the target image in a graph knowledge base based on the target image features; if the number of the at least one candidate knowledge entry is one, use the candidate knowledge entry as the target knowledge entry; if the number of the at least one candidate knowledge entry is multiple, obtain at least one historical session message corresponding to the consultation message, and determine the target knowledge entry from the multiple candidate knowledge entries based on the at least one historical session message.

[0123] In some embodiments, the apparatus further includes: a text acquisition module, configured to acquire text information contained in the target image; the knowledge search module 903 is specifically configured to search for at least one candidate knowledge entry that matches the target image from a graph knowledge base based on the features of the target image and the text information.

[0124] In some embodiments, the knowledge entry further includes the content recognition result of the sample consultation image; the knowledge search module 903 is also specifically used to search for at least one candidate knowledge entry from the graph knowledge base whose image features match the target image features; and to search for at least one candidate knowledge entry from the graph knowledge base whose content recognition result matches the text information.

[0125] In some embodiments, the knowledge lookup module 903 is further configured to utilize an intelligent filtering model to determine the conversation intent of the consultation message based on the at least one historical conversation message, and to determine the correlation between the target image and the multiple candidate knowledge entries based on the conversation intent, the at least one historical conversation message, and the content recognition results corresponding to each of the multiple candidate knowledge entries; to sort the multiple candidate knowledge entries based on the correlation, and to determine the target knowledge entry based on the sorting results.

[0126] In some embodiments, the consultation message further includes a text message; the knowledge data includes question data and response data corresponding to the question data; the device further includes: a response data lookup module, used to look up the response data corresponding to the text message from a question-and-answer knowledge base; the message generation module 904 is specifically used to generate a response message for the consultation message based on the response data contained in the target knowledge entry and the response data corresponding to the text message.

[0127] In some embodiments, the message generation module 904 is specifically used to determine whether the text message is associated with the target image; if so, based on at least one historical session message corresponding to the consultation message, target response data is determined from the response data contained in the target knowledge entry and the response data corresponding to the text message, and a response message for the consultation message is generated based on the target response data; if not, a response message for the consultation message is generated based on the response data corresponding to the text message and using the response data contained in the target knowledge entry as reference data.

[0128] Figure 9 The message processing device can perform Figure 6 The implementation principle and technical effects of the message processing method described in the illustrated embodiments will not be repeated here. The specific methods by which each module and unit of the message processing device in the above embodiments performs its operations have been described in detail in the embodiments related to this method, and will not be elaborated upon here.

[0129] Figure 10 This is a schematic diagram of the structure of one embodiment of a computing device provided in this application. Figure 10 As shown, in practice, the computing device may include a storage component 1001 and a processing component 1002.

[0130] Storage component 1001 is used to store computer programs and can be configured to store various other data to support operation on a computing device. Examples of this data include instructions for any application or method used to operate on the computing device, data structures, contact data, phone book data, messages, pictures, videos, etc.

[0131] Processing component 1002, coupled to storage component 1001, is used to execute computer programs in storage component 1001 for implementing, etc. Figure 1 The knowledge base construction method shown, or its implementation as follows Figure 6 The message processing method shown.

[0132] Furthermore, such as Figure 10As shown, the computing device may also include other components such as a communication component 1003, a display component 1004, a power supply component 1005, and an audio component 1006. Figure 10 The diagram only shows some components and does not mean that the device includes only these components. Figure 10 The components shown. Additionally... Figure 10 The components within the dashed box are optional, not mandatory, and their specific requirements depend on the product form of the computing device. The computing device in this embodiment can be a terminal device such as a desktop computer, laptop computer, smartphone, or IoT (Internet of Things) device, or a server-side device such as a conventional server, cloud server, or server array. If the computing device in this embodiment is implemented as a terminal device such as a desktop computer, laptop computer, or smartphone, it may include... Figure 10 The components within the dashed box; if the computing device in this embodiment is implemented as a conventional server, cloud server, or server array, etc., then it may not include... Figure 10 The component within the dashed box.

[0133] The processing component described above includes one or more processors to execute computer instructions to complete all or part of the steps in the method described above. Alternatively, the processing component may be implemented as one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the method described above.

[0134] The aforementioned storage components can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0135] The aforementioned communication component is configured to facilitate wired or wireless communication between the device housing the communication component and other devices. The device housing the communication component can access wireless networks based on communication standards, such as mobile communication networks, or combinations thereof. In one exemplary embodiment, the communication component receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel.

[0136] The aforementioned display components may include a screen, which may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also the duration and pressure associated with the touch or swipe operation.

[0137] The aforementioned power supply components provide power to various components within the device in which they reside. These power supply components may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power to the device in which they reside.

[0138] The aforementioned audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC) configured to receive external audio signals when the device containing the audio component is in an operating mode, such as call mode, recording mode, or voice recognition mode. The received audio signals can be further stored in memory or transmitted via a communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals.

[0139] Accordingly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to implement the steps in the above-described method embodiments. The computer-readable storage medium includes volatile or non-volatile components, or a combination thereof, and can be removable or non-removable. Examples of computer-readable storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random-access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), flash memory or other memory technologies, CD-ROM, Digital Video Disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium. Accordingly, this application also provides a computer program product, which includes a computer program or instructions that, when executed by a processor, cause the processor to implement the steps in the above method embodiments. It should be understood that each step or combination of steps in the above method flow can be implemented by the computer program or instructions. Furthermore, these computer programs or instructions can be applied to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device, enabling the processor of the general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to function as an apparatus for implementing the corresponding functions in the above method embodiments.

[0140] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0141] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0142] Finally, it should be noted that the above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A method for constructing a knowledge base, characterized in that, include: Obtain sample consultation images; Extract the image features of the sample consultation image; The intelligent recognition model is used to identify the image content of the sample consultation image; Based on the image content, generate knowledge data corresponding to the image content; Knowledge entries are constructed from the image features of the sample consultation images and the knowledge data; Based on the aforementioned knowledge entries, a graph knowledge base is constructed; When the image features in the knowledge entry are used to match the target image in the consultation message, a response message is generated based on the knowledge data in the knowledge entry for the consultation message.

2. The method according to claim 1, characterized in that, The obtained sample consultation images include: Take screenshots of the target webpages provided by the network service system to obtain sample consultation images; or, Historical consultation images are determined from the historical session records corresponding to the network service system, and these historical consultation images are used as sample consultation images.

3. The method according to claim 2, characterized in that, The knowledge data includes question data and corresponding response data. Generating knowledge data corresponding to the image content based on the image content includes: Based on the image content, search the question data that matches the image content and the corresponding answer data from the question-and-answer knowledge base; or, Based on the image content, generate question data for the sample consultation image, and search for the corresponding response data from the question-and-answer knowledge base; or, When the sample consultation image is a historical consultation image, the context message corresponding to the sample consultation image is obtained from the historical session record, and the intelligent generation model is used to generate question data corresponding to the image content and response data corresponding to the question data based on the image content and the context message.

4. The method according to claim 3, characterized in that, The image content identified using the intelligent recognition model for the sample consultation image includes: The target region in the sample consultation image is identified using an intelligent recognition model, and the image content of the target region is also identified.

5. The method according to any one of claims 1-4, characterized in that, The image content identified using the intelligent recognition model for the sample consultation image includes: The intelligent recognition model is used to generate image description information of the sample consultation image and extract the text information contained in the sample consultation image.

6. The method according to any one of claims 1-4, characterized in that, The knowledge entries, composed of the image features of the sample consultation images and the knowledge data, include: The image features of the sample consultation image, the content recognition result, and the knowledge data constitute a knowledge entry; the content recognition result is used to filter the target knowledge entry corresponding to the response message when the image features of multiple knowledge entries match the target image in the consultation message.

7. The method according to any one of claims 1-4, characterized in that, Also includes: Identify expired knowledge entries in the graph knowledge base and delete them from the graph knowledge base; The methods for determining expired knowledge entries include at least one of the following: In the graph knowledge base, knowledge entries that have not participated in any session message generation operation within a preset time period are considered expired knowledge entries. Obtain user feedback data for any response message, and if the user feedback data determines that the response message does not meet the satisfaction requirements, then the knowledge entry for generating the response message will be designated as an expired knowledge entry. Using an intelligent evaluation model, the quality of any response message is evaluated, and if the evaluation result indicates that the response message does not meet the quality requirements, the knowledge entry that generated the response message is designated as an expired knowledge entry.

8. The method according to claim 7, characterized in that, The process of using an intelligent evaluation model to perform quality evaluation on any response message includes: Using an intelligent evaluation model, for any response message, based on the consultation message corresponding to the response message, the response message, and user feedback data for the response message, a quality score corresponding to the response message in multiple quality dimensions is determined. The target score of the response message is determined based on the weight values ​​and quality scores corresponding to the multiple quality dimensions.

9. A message processing method, characterized in that, include: Retrieve inquiry messages sent by the user client; If the consultation message includes a target image, extract the target image features of the target image; Based on the target image features, a target knowledge entry matching the target image is searched in a graph knowledge base; the graph knowledge base includes multiple knowledge entries; the knowledge entry includes the image features of the sample consultation image and knowledge data; The knowledge data is generated based on the image content of the sample consultation image, and the image content is obtained by recognizing the sample consultation image using an intelligent recognition model; Based on the knowledge data in the target knowledge item, generate a response message for the inquiry message; The response message is sent to the user terminal.

10. The method according to claim 9, characterized in that, The step of searching for target knowledge entries that match the target image from the graph knowledge base based on the target image features includes: Based on the features of the target image, at least one candidate knowledge entry that matches the target image is searched from the graph knowledge base; If the number of the at least one candidate knowledge entry is one, the candidate knowledge entry shall be used as the target knowledge entry; If there are multiple candidate knowledge entries, obtain at least one historical session message corresponding to the consultation message, and determine the target knowledge entry from the multiple candidate knowledge entries based on the at least one historical session message.

11. The method according to claim 10, characterized in that, The method further includes: Obtain the text information contained in the target image; The step of searching for at least one candidate knowledge entry matching the target image from the graph knowledge base based on the target image features includes: Based on the target image features and the text information, at least one candidate knowledge entry matching the target image is searched from the graph knowledge base.

12. The method according to claim 11, characterized in that, The knowledge entries also include the content recognition results of the sample consultation images; The step of searching for at least one candidate knowledge entry matching the target image from the graph knowledge base based on the target image features and the text information includes: Search the graph knowledge base for at least one candidate knowledge entry whose image features match the target image features; In addition, at least one candidate knowledge entry is searched from the graph knowledge base that matches the content recognition result with the text information.

13. The method according to any one of claims 10-12, characterized in that, The step of determining the target knowledge entry from the plurality of candidate knowledge entries based on the at least one historical session message includes: Using an intelligent filtering model, based on the at least one historical conversation message, the conversation intent of the consultation message is determined, and based on the conversation intent, the at least one historical conversation message, and the content recognition results corresponding to each of the multiple candidate knowledge items, the correlation between the target image and the multiple candidate knowledge items is determined. Based on the relevance, the multiple candidate knowledge entries are sorted, and the target knowledge entry is determined according to the sorting result.

14. The method according to any one of claims 9-12, characterized in that, The consultation message also includes a text message; the knowledge data includes question data and corresponding response data. The method further includes: Search the question-and-answer knowledge base for the response data corresponding to the text message; Based on the knowledge data in the target knowledge entry, generating a response message for the inquiry message includes: Based on the response data contained in the target knowledge entry and the response data corresponding to the text message, a response message is generated for the consultation message.

15. The method according to claim 14, characterized in that, The step of generating a response message for the inquiry message based on the response data contained in the target knowledge item and the response data corresponding to the text message includes: Determine whether the text message is associated with the target image; If so, based on at least one historical session message corresponding to the consultation message, target response data is determined from the response data contained in the target knowledge entry and the response data corresponding to the text message, and a response message for the consultation message is generated based on the target response data; If not, a response message is generated based on the response data corresponding to the text message, and using the response data contained in the target knowledge entry as reference data.

16. A computing device, characterized in that, This includes processing components and storage components; The storage component stores a computer program; the computer program is invoked and executed by the processing component to implement the knowledge base construction method as described in any one of claims 1-8, or to implement the message processing method as described in any one of claims 9-15.

17. A computer-readable storage medium, characterized in that, It stores a computer program, which, when executed by a processing component, implements the knowledge base construction method as described in any one of claims 1-8, or the message processing method as described in any one of claims 9-15.

18. A computer program product, characterized in that, It includes a computer program or instructions that, when executed by a processing component, implement the knowledge base construction method as described in any one of claims 1-8, or the message processing method as described in any one of claims 9-15.