Product appearance information intelligent processing method and system based on dialogue context
Through multimodal data processing and the use of intent recognition, entity extraction and feature fusion technologies, the problem of low retrieval accuracy caused by the diversity of user-uploaded pictures and text descriptions is solved, and efficient appearance retrieval result generation is achieved.
Patent Information
- Application Number
- CN202511241098.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-02
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-09-02
AI Technical Summary
The existing product appearance retrieval and compliance inspection system has low query accuracy due to the diversity of images and text descriptions uploaded by users. It cannot accurately present real related results and cannot meet user needs.
By processing multimodal data (text and images), using intent recognition, entity extraction, redundant information processing, feature extraction and fusion, we generate effective features that represent appearance retrieval intent, and achieve accurate understanding and processing of user appearance retrieval intent.
It improves the accuracy and efficiency of AI dialogue agents in appearance retrieval tasks and provides accurate retrieval results.
Smart Images

Figure CN120744147A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the technical field of data processing systems or methods specifically suitable for administrative, commercial, financial, management, supervisory or forecasting purposes, and in particular relates to a method and system for intelligently processing product appearance information based on conversation context. Background Art
[0002] Merchant users of e-commerce platforms can use the related services of third-party similar product query and compliance testing systems for products to be listed before the products are listed or during the development stage, so as to promptly find out whether the products have similar appearance, trademarks, etc.
[0003] At present, relevant product appearance retrieval and compliance detection systems are already able to search and present results for similar appearance patents based on product images and / or text descriptions uploaded by users. However, in actual usage scenarios, due to the diversity and uncontrollability of the forms of expression of product images uploaded by users and text descriptions actually entered, such as the mixing of non-relevant images, the current relevant systems do not use query algorithms for recognition and differentiation, resulting in low query accuracy. The real related result mask cannot be accurately presented to the user end in a large amount of noisy results, and cannot meet user usage needs. Summary of the Invention
[0004] This application provides a method and system for intelligent processing of product appearance information based on conversation context. This application utilizes multimodal data (text and images) and achieves accurate understanding and processing of user appearance retrieval intentions through operations such as intent recognition, entity extraction, redundant information processing, feature extraction and fusion. It effectively extracts key information from complex data, optimizes data quality, generates effective features that represent appearance retrieval intentions, and then accurately executes appearance retrieval operations, ultimately providing users with accurate retrieval results, thereby improving the accuracy and efficiency of AI conversational agents in appearance retrieval tasks.
[0005] In the first aspect, the present application provides a method for intelligent processing of product appearance information based on conversation context, which is applied to an AI conversation agent, and the method includes: obtaining first conversation data of a first conversation event between a user and the AI conversation agent, the first conversation data including a text data set and a picture data set, the picture data set including a first picture uploaded by a picture editing component of the AI conversation agent and / or a second picture generated by the AI conversation agent based on the conversation data in the first text data set, or the picture data set is empty; performing intent recognition, entity extraction, redundant information decoupling, feature extraction, and feature fusion on the text data set and the picture data set to obtain an appearance retrieval feature that represents the appearance retrieval intent; performing an appearance retrieval operation based on the appearance retrieval feature to obtain an appearance retrieval result; and displaying the appearance retrieval result.
[0006] In the second aspect, the present application provides a product appearance information intelligent processing system based on dialogue context, including a terminal device and a server, wherein the terminal device is used to execute the steps performed by the terminal device as described in any method in the first aspect; and the server is used to execute the steps performed by the server as described in any method in the first aspect.
[0007] As can be seen, in an embodiment of the present application, first conversation data of a first conversation event between a user and an AI conversational agent is obtained, the first conversation data comprising a text dataset and an image dataset, the image dataset comprising a first image uploaded by the AI conversational agent's image editing component and / or a second image generated by the AI conversational agent based on the conversation data in the first text dataset, or the image dataset being empty; intent recognition, entity extraction, redundant information decoupling, feature extraction, and feature fusion are performed on the text dataset and the image dataset to obtain an appearance retrieval feature representing the appearance retrieval intent; an appearance retrieval operation is performed based on the appearance retrieval feature to obtain an appearance retrieval result; and the appearance retrieval result is displayed. As can be seen, the present application utilizes multimodal data (text and images) and, through operations such as intent recognition, entity extraction, redundant information processing, feature extraction, and fusion, achieves accurate understanding and processing of the user's appearance retrieval intent. This effectively extracts key information from complex data, optimizes data quality, generates effective features representing the appearance retrieval intent, accurately executes the appearance retrieval operation, and ultimately provides the user with accurate retrieval results, thereby improving the accuracy and efficiency of the AI conversational agent in appearance retrieval tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0009] Figure 1 A schematic diagram of the results of a product appearance information intelligent processing system provided in an embodiment of the present application; Figure 2 A flowchart of a method for intelligently processing product appearance information based on conversation context provided in an embodiment of the present application; Figure 3 This is the appearance detection dialogue page of the AI dialogue agent provided in the embodiment of the present application; Figure 4 A schematic diagram of a picture editing component provided in an embodiment of the present application; Figure 5A schematic diagram of the structure of a server provided in an embodiment of the present application. DETAILED DESCRIPTION
[0010] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0011] The terms "first," "second," and the like in the specification and claims of this application and the accompanying drawings are used to distinguish between different objects, not to describe a particular order. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to the process, method, product, or apparatus.
[0012] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0013] In the embodiments of this application, "and / or" describes the relationship between associated objects and indicates that three relationships can exist. For example, "A and / or B" can represent the following three situations: A exists alone; A and B exist simultaneously; and B exists alone. A and B can be singular or plural.
[0014] In the embodiments of the present application, the symbol " / " can indicate that the preceding and following objects are in an "or" relationship. In addition, the symbol " / " can also represent a division sign, that is, performing a division operation. For example, A / B can mean A divided by B.
[0015] In the embodiments of the present application, "at least one item" or similar expressions refers to any combination of these items, including any combination of single items or plural items, and refers to one or more, and multiple refers to two or more. For example, at least one item (item) of a, b, or c can represent the following seven situations: a, b, c, a and b, a and c, b and c, a, b, and c. Among them, each of a, b, and c can be an element or a set containing one or more elements.
[0016] In the embodiments of this application, "equal to" can be used in conjunction with "greater than" and is applicable to the technical solution adopted when "greater than" is used, and can also be used in conjunction with "less than" and is applicable to the technical solution adopted when "less than" is used. When "equal to" is used in conjunction with "greater than", it should not be used in conjunction with "less than"; when "equal to" is used in conjunction with "less than", it should not be used in conjunction with "greater than".
[0017] This application provides a method and system for intelligent processing of product appearance information based on conversation context. This application utilizes multimodal data (text and images) and achieves accurate understanding and processing of user appearance retrieval intentions through operations such as intent recognition, entity extraction, redundant information processing, feature extraction and fusion. It effectively extracts key information from complex data, optimizes data quality, generates effective features that represent appearance retrieval intentions, and then accurately executes appearance retrieval operations, ultimately providing users with accurate retrieval results, thereby improving the accuracy and efficiency of AI conversational agents in appearance retrieval tasks.
[0018] The following specific embodiments describe in detail the technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0019] See also Figure 1 and Figure 2 , Figure 1 This is a schematic diagram of the results of a product appearance information intelligent processing system provided by an embodiment of the present application. Figure 2 A flowchart of a method for intelligently processing product appearance information based on conversation context provided in an embodiment of the present application.
[0020] The product appearance information intelligent processing system 1 includes a server 10 and a terminal device 20 , and the server 10 and the terminal device 20 are communicatively connected.
[0021] The server 10 may specifically include a server that is applied to one side of the network platform and is responsible for data processing in the background, which can realize functions such as data transmission and data processing. It can be a physical server or a server cluster or distributed system composed of multiple physical servers. In this embodiment, the number of servers is not specifically limited. Alternatively, it can also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0022] Specifically, the terminal device 20 may include a front-end device applied to the user side and capable of performing functions such as data collection and data transmission in a signal-free scenario. The device may be a user equipment (UE) such as a mobile phone, smart phone, laptop computer, digital broadcast receiver, personal digital assistant (PDA), tablet computer (PAD), handheld device, vehicle-mounted device, wearable device, computing device or other processing device connected to a wireless modem, mobile station (MS), mobile terminal, etc. Alternatively, the terminal device 20 may also be a software application that can run on the above electronic devices.
[0023] The terminal device 20 is Figure 2 The execution body of the method for intelligently processing product appearance information based on the dialogue context shown in the figure includes the following steps S201 to S204: Step S201: Obtain first dialogue data of a first dialogue event between a user and the AI dialogue agent.
[0024] Among them, the first dialogue data includes a text dataset and a picture dataset, the picture dataset includes a first picture uploaded by the picture editing component of the AI dialogue agent and / or a second picture generated by the AI dialogue agent based on the dialogue data in the first text dataset, or the picture dataset is empty.
[0025] Among them, the user enters text information data through multiple rounds of dialogue with the AI dialogue agent, uploads the first picture through the picture editing component of the AI agent, or generates a second picture based on the dialogue data, or there is no picture, and the picture data set is empty; and obtains the text data set and picture data set based on the entered text information data and / or picture data.
[0026] For user-entered text data, the AI conversational agent automatically generates targeted clarification questions based on user interaction. When the system detects ambiguous or incomplete user input, it helps users clarify their needs. Spelling and semantic error correction mechanisms are introduced to handle user input errors and improve system robustness. During the query process, user input is anonymized to adhere to data privacy and security standards.
[0027] Step S202 , performing intent recognition, entity extraction, redundant information decoupling, feature extraction, and feature fusion on the text dataset and the image dataset to obtain appearance retrieval features representing appearance retrieval intent.
[0028] Intent recognition for text datasets involves using natural language processing techniques, such as deep learning-based pre-trained models like BERT and GPT variants, to comprehensively analyze the vocabulary, grammatical structure, and semantic information in the text. The model learns from large amounts of text data annotated with intent. When presented with new text data, it predicts the intent category it belongs to, such as whether it's a query for descriptive text about a particular product's appearance or reviews comparing the appearance of different products, thereby identifying intent trends related to appearance retrieval.
[0029] Intent recognition for image datasets involves using computer vision techniques, such as convolutional neural network (CNN) architecture models. CNNs can capture visual features at different levels of an image to determine the underlying intent. For example, a clothing image can be used to showcase a style or focus on the details of a single item. This allows the identification of image intent related to appearance-based search.
[0030] After this step, the text and image each have preliminary intent classification results, providing direction for subsequent processing.
[0031] In some embodiments, the intent classification results include but are not limited to: graphic interface appearance retrieval category, physical product appearance retrieval category, text trademark retrieval category, graphic trademark retrieval category, mixed graphic and text trademark retrieval category, and e-commerce platform policy retrieval category.
[0032] The Graphical Interface Appearance search category focuses on similarities in the appearance of graphical interfaces for various software, electronic devices, and other devices. For example, this category searches for graphical layout, icon style, and graphical control settings within a graphical interface based on color matching, icon shape and arrangement, and the distribution of interactive areas within the interface. The Physical Product Appearance search category primarily searches for similarities in the physical appearance of physical products. The Text Trademark search category primarily searches for similarities in trademarks composed of text. The Graphic Trademark search category primarily searches for similarities in trademarks composed of graphic elements. The Graphic and Text Mixed Trademark search category is applicable to trademarks composed of both text and graphics, taking into account the combination of text and image, the characteristics of the text and graphics, and their interrelationships. The E-commerce Platform Policy search category searches for platform policies regarding product image display specifications, trademark use restrictions, and standards for determining appearance infringement before listing products on the e-commerce platform. This ensures that product display and trademark use comply with platform rules and avoids penalties for violations.
[0033] Entity extraction for text datasets involves identifying the text for appearance-based search intent and then using a named entity recognition (NER) algorithm to extract key entities. For example, for the query "I'm looking for images of a red flip phone," entities closely related to appearance attributes, such as "red" and "flip phone," would be extracted, simplifying the text into its core descriptive elements.
[0034] Entity extraction from image datasets refers to the use of image segmentation and target detection techniques to identify object entities in images. For example, it can detect the entity of a sofa in a home scene image, annotate its outline, position and other information, and sort out visual entities directly related to the appearance. These extracted entities will serve as materials for more precise processing in subsequent steps.
[0035] In specific implementation, in some embodiments, intent recognition and entity extraction for text datasets and image datasets include the following steps: (1) Multilingual support: Introducing multilingual word segmentation and word vector technology to support cross-language intent recognition and entity extraction. (2) Entity recognition optimization: Using named entity recognition (NER) and the ontology library of the patent field, accurately extract key entities related to design and improve the accuracy of retrieval. (3) Model training: Using pre-trained language models (such as BERT, GPT) for fine-tuning, combined with the corpus of the patent field, to improve the accuracy of intent recognition. (4) Intent classification and similarity calculation: Using a hierarchical intent classification system and using measurement methods such as cosine similarity to improve the accuracy of intent matching. (5) Dynamic intent update and multi-round dialogue tracking: Through the context-aware mechanism and dialogue history, the user intent model is updated in real time to ensure effective tracking of multi-round conversations.
[0036] Among them, the decoupling of redundant information for text datasets means: for the text after entity extraction, removing redundant content such as modifiers and background descriptions that are irrelevant to the appearance retrieval intent, so that the text information can focus more on the key elements of appearance.
[0037] Decoupling redundant information from image datasets means removing elements in the image background that interfere with appearance judgment by using background removal algorithms or image cropping techniques. For example, if a product image is surrounded by complex promotional labels and irrelevant ornaments, the algorithm can be used to highlight the main body of the product and eliminate residual visual information, making the subsequent fused information purer.
[0038] After this step, the text and image data are further "purified", which is conducive to feature fusion.
[0039] Feature extraction refers to obtaining various characteristic information from text or images that can represent their content. For example, for image data, pre-trained visual models (such as ResNet and ViT) are used to extract image features, automatically extracting color, texture, and shape features. For text data, pre-trained language models (such as BERT and GPT) are used to extract text features. These methods, such as counting word frequencies, calculating TF-IDF values, or using word embeddings to obtain semantic vectors, leverage deep learning models to capture contextual features, providing critical support for tasks such as text classification and sentiment analysis.
[0040] Unlike entity extraction, entity extraction focuses on isolating specific objects from images and text. These objects are entities that can be clearly named and located, such as identifying the product itself and its components in a product image. Feature extraction abstracts the entire image or text content and focuses on how to use digital features to describe the visual characteristics of the image, rather than focusing on the specific names of the entities.
[0041] Specifically, the feature fusion of entity objects and feature objects is mapped to a unified vector space through cross-modal comparative learning, supporting similarity calculation of multimodal features. This method can effectively handle the differences between data of different modalities, improving the consistency of feature representation and the accuracy of similarity calculation.
[0042] Step S203: performing an appearance search operation according to the appearance search feature to obtain an appearance search result.
[0043] The appearance retrieval operation implements efficient multimodal retrieval logic, fusing text features describing appearance with image features representing appearance to generate appearance-sensing retrieval features. This retrieval feature is then used to calculate similarity with candidate data in the database, using algorithms such as Euclidean distance and cosine similarity. An adaptive sorting algorithm is also introduced to optimize the ranking of candidate results based on user preferences and feedback. For example, it can prioritize appearance data that is most similar to the text and image data entered by the user, sorted from high to low similarity.
[0044] During each round of conversation, users can provide positive or negative feedback, such as whether the retrieved most similar appearance data is accurate. The system leverages this feedback and employs online learning algorithms to adjust the ranking model and retrieval strategy in real time. Through long-term interactions, the system builds a user preference model and provides personalized retrieval results.
[0045] Step S204: display the appearance search result.
[0046] As can be seen, in an embodiment of the present application, first conversation data of a first conversation event between a user and an AI conversational agent is obtained, the first conversation data comprising a text dataset and an image dataset, the image dataset comprising a first image uploaded by the AI conversational agent's image editing component and / or a second image generated by the AI conversational agent based on the conversation data in the first text dataset, or the image dataset being empty; intent recognition, entity extraction, redundant information decoupling, feature extraction, and feature fusion are performed on the text dataset and the image dataset to obtain an appearance retrieval feature representing the appearance retrieval intent; an appearance retrieval operation is performed based on the appearance retrieval feature to obtain an appearance retrieval result; and the appearance retrieval result is displayed. As can be seen, the present application utilizes multimodal data (text and images) and, through operations such as intent recognition, entity extraction, redundant information processing, feature extraction, and fusion, achieves accurate understanding and processing of the user's appearance retrieval intent. This effectively extracts key information from complex data, optimizes data quality, generates effective features representing the appearance retrieval intent, accurately executes the appearance retrieval operation, and ultimately provides the user with accurate retrieval results, thereby improving the accuracy and efficiency of the AI conversational agent in appearance retrieval tasks.
[0047] In some embodiments, the image dataset is empty; obtaining the first conversation data of the first conversation event between the user and the AI conversation agent includes: obtaining multiple rounds of text conversation data of the first conversation event between the user and the AI conversation agent; and performing associated data screening on the multiple rounds of text conversation data according to the conversation topic consistency constraint to obtain the first conversation data.
[0048] Among them, conversation topic keywords are extracted from the text conversation data of each round in the multiple rounds of text conversation data to obtain multiple conversation topic keyword groups. Each conversation topic keyword group contains at least one topic keyword, which may specifically include explicit conversation topic keywords and implicit conversation topic keywords. Explicit conversation topic keywords are directly extracted keywords, and implicit conversation topic keywords are keywords obtained based on semantic understanding and expansion.
[0049] In some embodiments, the text dataset and the image dataset are subjected to intent recognition, entity extraction, redundant information decoupling, feature extraction, and feature fusion to obtain appearance retrieval features representing appearance retrieval intent, including: performing intent recognition on the first conversation data to obtain an intent recognition category as a graphical interface appearance retrieval category; performing entity extraction on the first conversation data according to the graphical interface appearance retrieval category to obtain text entities; performing redundant information decoupling processing on the first conversation data according to the graphical interface appearance retrieval category to obtain processed first conversation data; performing text encoding operation on the processed first conversation data to obtain text features; and fusing the text entities and the text features to obtain appearance retrieval features representing appearance retrieval intent.
[0050] Among them, when intent recognition is performed based on the first conversation data and it is determined that the intent recognition category is a graphical interface appearance retrieval category, entity extraction is performed on the first conversation data to obtain text entities. Exemplarily, the text entities include but are not limited to the color matching, icon shape and arrangement, interface interaction area distribution and other features of the graphical interface appearance; redundant information decoupling is performed on the first conversation data, and text encoding operations are performed on the processed first conversation data to obtain text features; the text entities and text features are fused to obtain appearance retrieval features.
[0051] It can be seen that in this embodiment, after the graphical interface appearance retrieval category is determined through intent recognition, entity extraction is performed to obtain key features, such as color matching, icon shape and arrangement, interface interaction area distribution, etc., to accurately anchor the graphical interface appearance elements that users are concerned about. This can avoid interference from other irrelevant information and focus the retrieval on core needs. Redundant information decoupling processing further purifies the data, removes noise, and ensures the accuracy of subsequent text encoding and feature fusion. After such processing, the appearance retrieval features obtained are highly consistent with user needs, the retrieval results are more accurate, and false detections and missed detections are reduced.
[0052] In some embodiments, the image dataset includes a second image generated by the AI dialogue agent based on the dialogue data in the first text dataset; and obtaining the first dialogue data of the first dialogue event between the user and the AI dialogue agent includes the following steps A1 to A5: Step A1: Obtain multiple rounds of text conversation data of the first conversation event between the user and the AI conversation agent.
[0053] Step A2: performing image correlation analysis and differentiation processing on the multiple rounds of text conversation data to obtain a non-correlated first conversation data set and a correlated second conversation data set.
[0054] Step A3: extract the image data from the second conversation data set and perform deduplication processing to obtain the second image.
[0055] Step A4: extract text data from the second conversation data set and merge it with the first conversation data set to obtain a third conversation data set.
[0056] Step A5: screening the third conversation data set for related data according to the conversation topic consistency constraint to obtain the first conversation data.
[0057] The first conversation data set consists of conversation data unrelated to the image, while the second conversation data set consists of conversation data related to the image. A second image is generated based on the image-related conversation data in the second conversation data set. The text data in the second conversation data set is filtered out and merged with the first conversation data set to generate a third conversation data set. The third conversation data set is then filtered for related data to obtain at least one first conversation data set with the same topic.
[0058] It can be seen that in this embodiment, by performing a series of processes such as image correlation analysis, classification, extraction, deduplication, merging, and screening based on topic consistency on multiple rounds of text conversation data, effective separation and integration of image-related and irrelevant information in the conversation data is achieved, data redundancy is removed, and data quality is improved, so that the final conversation data is more consistent in terms of topic. This provides a high-quality, clearly structured, and thematically clear data foundation for subsequent in-depth analysis, model training, and intelligent applications (such as intelligent customer service) based on conversation data, enhances the availability and effectiveness of the data in actual application scenarios, helps to improve the accuracy of the relevant systems' understanding of user intent, and optimizes the service experience and system performance.
[0059] In some embodiments, performing intent recognition, entity extraction, redundant information decoupling, feature extraction, and feature fusion on the text dataset and the image dataset to obtain appearance retrieval features representing appearance retrieval intent includes: Performing intent recognition on the first conversation data and the second image to obtain an intent recognition category as a graphical interface appearance retrieval category; performing a first entity extraction on the first conversation data according to the graphical interface appearance retrieval category to obtain a descriptive entity; and performing a second entity extraction on the second image according to the graphical interface appearance retrieval category to obtain a visual entity; Performing first redundant information decoupling and first feature extraction on the first conversation data according to the graphical interface appearance retrieval category to obtain text features; performing second redundant information decoupling and second feature extraction on the second image according to the graphical interface appearance retrieval category to obtain image features; Feature fusion is performed on the description entity, the visual entity, the text feature, and the image feature to obtain the appearance retrieval feature.
[0060] As can be seen, in this embodiment, when the conversation data includes text data and image data, and the search category is identified as a graphical interface appearance search category based on the conversation data and image data, descriptive entities are obtained based on the first conversation data, visual entities are obtained based on the second image, text features are obtained from the first conversation data, and image features are obtained from the second image. Finally, feature fusion is performed based on the descriptive entities, visual entities, text features, and entity features to obtain appearance search features. This can accurately capture the user's intention for graphical interface appearance search, extract accurate and comprehensive entity information from the conversation data and images, remove redundant interference, obtain high-quality text and image features, and effectively fuse these features. This generates a fused feature that fully represents the appearance of the display interface, providing an accurate and rich information foundation for subsequent graphical interface appearance search.
[0061] In some embodiments, the AI dialogue agent is expressed in the terminal device as an appearance detection dialogue page, and obtaining the first dialogue data of the first dialogue event between the user and the AI dialogue agent includes: detecting a first selection operation for a graphic user interface option in the main dialogue interaction interface; in response to the first selection operation, displaying the appearance detection dialogue page, the appearance detection dialogue page including a text dialogue component and the picture editing component; obtaining the text data set through the text dialogue component, and obtaining the picture data set through the picture editing component.
[0062] For example, see Figure 3 , Figure 3 The appearance detection dialogue page of the AI dialogue agent provided in the embodiment of this application is as follows Figure 3As shown, the appearance detection dialogue page 3 includes a text dialogue component 31 and an image editing component 32. The text dialogue component 31 includes a user input area 311 and a dialogue display area 312. Users can enter text information in the user input area 311 and then click the send button to interact with the AI dialogue agent. Users can also directly add text files to the user input area 311, enter processing instructions, and then click the send button to interact with the AI dialogue agent. Among them, users upload pictures through the image editing component 32.
[0063] In some embodiments, the image dataset includes a first image uploaded by the image editing component of the AI dialogue agent; and the image dataset includes a display interface image of a graphical interface product and a physical product appearance image of a physical product carrying the graphical interface product; the text dataset includes the name and purpose of the graphical interface product, and descriptive information of the display interface image; performing intent recognition, entity extraction, redundant information decoupling, feature extraction, and feature fusion on the text dataset and the image dataset to obtain an appearance retrieval feature representing the appearance retrieval intent includes: Performing intent recognition on the text dataset and the image dataset to obtain an intent recognition category that is a graphical interface appearance retrieval category; performing a first entity extraction on the text dataset based on the graphical interface appearance retrieval category to obtain a descriptive entity; and performing a second entity extraction on the image dataset based on the graphical interface appearance retrieval category to obtain a visual entity; performing an associated image convergence operation on the image dataset based on the graphical interface appearance retrieval category and the description information of the display interface diagram to obtain a picture dataset after removing the physical product appearance diagram; determining global text features based on the name and the purpose; and determining local text features of the display interface diagram based on the description information; performing an encoding operation on the display interface diagram to obtain image features of the display interface diagram; and performing feature fusion on the descriptive entity, the visual entity, the local text features, the global text features and the image features of the display interface diagram to obtain the appearance retrieval feature.
[0064] In some embodiments, performing an associated image convergence operation on the image dataset based on the graphical interface appearance retrieval category and the description information of the display interface image to obtain an image dataset after removing the physical product appearance image includes: performing object recognition on the display interface images in the image dataset to obtain one or more display objects of each display interface image, where a single display object refers to a graphic composed of one or more graphic elements and having independent visual connotations; performing a first software interface element association analysis on the display objects to obtain a first software interface element association of each display object; performing a second software interface element association analysis on each display object based on the description information of the display interface image to obtain a second software interface element association of each display object; fusing the first software interface element association and the second software interface element association to obtain a software interface element fusion association of each display object; determining a software interface association of the display interface image based on the one or more software interface element fusion associations of the one or more display objects; and updating the image dataset based on a comparison result of the software interface association and the association corresponding to the graphical interface appearance retrieval category to obtain the image dataset after removing the physical product appearance image.
[0065] In some embodiments, the first software interface element association analysis of the display object is performed to obtain the first software interface element association of each display object, including: performing graphic semantic analysis on each display object to obtain one or more graphic semantic keywords that are adapted to each display object; using the one or more graphic semantic keywords as query identifiers, querying a preset first mapping relationship set to obtain one or more existence attributes that are adapted to the graphic semantic keywords, the existence attributes including virtual existence attributes and physical existence attributes, the first mapping relationship combination including a correspondence between graphic semantic keywords and existence attributes; determining the first software interface element association of each display object according to the proportion of the number of virtual existence attributes in the one or more existence attributes.
[0066] In some embodiments, the second software interface element association analysis is performed on each display object based on the descriptive information of the display interface diagram to obtain the second software interface element association of each display object, including: performing text semantic analysis on the descriptive information of the display interface diagram to obtain one or more text semantic keywords adapted for each display object; using the one or more text semantic keywords as query identifiers, querying a preset second mapping relationship set to obtain one or more existence attributes of the adapted text semantic keywords, the existence attributes including virtual existence attributes and physical existence attributes, the second mapping relationship combination including a correspondence between text semantic keywords and existence attributes; determining the second software interface element association of each display object based on the proportion of the number of virtual existence attributes in the one or more existence attributes.
[0067] Among them, see Figure 4 , Figure 4 Schematic diagram of the picture editing component provided in the embodiment of the present application. For the picture editing component 32 in the appearance detection dialogue page 3, after the user clicks the picture editing component 32, the following is displayed: Figure 4 The picture editing component display page 4 shown includes a main view editing component 41 and a general view editing component 42. The main view editing component 41 is used to enter the main view, and the general view editing component 42 is used to enter the multiple state change diagrams.
[0068] In some embodiments, the display interface diagram includes a main view and multiple state change diagrams, the description information includes first description information of the main view and the first state change diagram, and second description information of the second state change diagram; determining the local text features of the display interface diagram based on the description information includes: determining the first local text features of the main view based on the first description information and the image information of the main view; determining the second local text features of the first state change diagram based on the first description information and the image information of the first state change diagram; determining the third local text features of the second state change diagram based on the second description information and the image information of the second state change diagram.
[0069] In some embodiments, determining the first local text feature of the main view based on the first descriptive information and the image information of the main view includes: extracting the first usage connotation feature of the first display object in the image information of the main view based on the first descriptive information to obtain the first usage connotation feature; determining the second local text feature of the first state change diagram based on the first descriptive information and the image information of the first state change diagram includes: extracting the second usage connotation feature of the second display object in the image information of the first state change diagram based on the first descriptive information to obtain the second usage connotation feature, wherein the first display object includes the second display object; determining the third local text feature of the second state change diagram based on the second descriptive information and the image information of the second state change diagram includes: extracting the third usage connotation feature of the third display object in the image information of the second state change diagram based on the second descriptive information to obtain the third usage connotation feature, and the first display object is different from the third display object.
[0070] With the above Figure 2 For details on the embodiment shown, please refer to Figure 5 , Figure 5 A schematic diagram of the structure of a server provided in an embodiment of the present application is shown in FIG. Figure 5 As shown, the server 10 includes a processor 51, a memory 53, a communication interface 52, and one or more programs 531. The one or more programs 531 are stored in the memory 53 and are configured to be executed by the processor 51. The above programs include methods for executing the methods described in the above embodiments.
[0071] An embodiment of the present application also provides a computer storage medium, wherein the computer storage medium stores a computer program for electronic data exchange, and the computer program enables a computer to execute part or all of the steps of any method described in the above method embodiments, and the above computer includes an electronic device.
[0072] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.
[0073] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0074] In the several embodiments provided in this application, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical or other forms.
[0075] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0076] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0077] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, or the part that contributes to the existing technology, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a memory and includes a number of instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the above-mentioned methods in each embodiment of the present application. The aforementioned memory includes: U disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), mobile hard disk, magnetic disk or optical disk, etc. Various media that can store program code.
[0078] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments can be completed by a program instructing related hardware. The program can be stored in a computer-readable memory, which may include a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0079] The above is a detailed introduction to the embodiments of the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method and core idea of the present application. At the same time, for those skilled in the art, according to the idea of the present application, there may be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
Claims
1. A method for intelligently processing product appearance information based on conversation context, characterized in that: Applied to an AI dialogue agent, the method includes: Obtaining first conversation data of a first conversation event between a user and the AI conversation agent, where the first conversation data includes a text dataset and an image dataset, where the image dataset includes a first image uploaded by an image editing component of the AI conversation agent and / or a second image generated by the AI conversation agent based on the conversation data in the first text dataset, or the image dataset is empty; Performing intent recognition, entity extraction, redundant information decoupling, feature extraction, and feature fusion on the text dataset and the image dataset to obtain appearance retrieval features representing appearance retrieval intent; Performing an appearance retrieval operation according to the appearance retrieval feature to obtain an appearance retrieval result; The appearance search result is displayed.
2. The method according to claim 1, characterized in that The image dataset is empty; and obtaining first dialogue data of a first dialogue event between the user and the AI dialogue agent includes: Obtaining multiple rounds of text conversation data of a first conversation event between the user and the AI conversation agent; The associated data of the multiple rounds of text conversation data are screened according to a conversation topic consistency constraint condition to obtain the first conversation data.
3. The method according to claim 1, characterized in that The image dataset includes a second image generated by the AI dialogue agent based on the dialogue data in the first text dataset; and obtaining first dialogue data of a first dialogue event between the user and the AI dialogue agent includes: Obtaining multiple rounds of text conversation data of a first conversation event between the user and the AI conversation agent; Performing image correlation analysis and differentiation processing on the multiple rounds of text conversation data to obtain a non-correlated first conversation data set and a correlated second conversation data set; Extracting image data from the second conversation data set and performing deduplication processing to obtain the second image; extracting text data from the second conversation data set and merging it with the first conversation data set to obtain a third conversation data set; The third conversation data set is screened for associated data according to the conversation topic consistency constraint to obtain the first conversation data.
4. The method according to claim 3, characterized in that The performing of intent recognition, entity extraction, redundant information decoupling, feature extraction, and feature fusion on the text dataset and the image dataset to obtain appearance retrieval features representing appearance retrieval intent includes: Performing intent recognition on the first conversation data and the second image to obtain an intent recognition category as a graphical interface appearance retrieval category; performing a first entity extraction on the first conversation data according to the graphical interface appearance retrieval category to obtain a descriptive entity; and performing a second entity extraction on the second image according to the graphical interface appearance retrieval category to obtain a visual entity; Performing first redundant information decoupling and first feature extraction on the first conversation data according to the graphical interface appearance retrieval category to obtain text features; performing second redundant information decoupling and second feature extraction on the second image according to the graphical interface appearance retrieval category to obtain image features; Feature fusion is performed on the description entity, the visual entity, the text feature, and the image feature to obtain the appearance retrieval feature.
5. The method according to claim 1, wherein The AI dialogue agent is presented in the terminal device as an appearance detection dialogue page, and the obtaining of first dialogue data of a first dialogue event between the user and the AI dialogue agent includes: detecting a first selection operation on a graphical user interface option in the primary dialog interaction interface; In response to the first selection operation, displaying an appearance detection dialogue page, the appearance detection dialogue page including a text dialogue component and the image editing component; The text data set is obtained through the text dialogue component, and the image data set is obtained through the image editing component.
6. The method according to claim 5, characterized in that The picture dataset includes a first picture uploaded by the picture editing component of the AI dialogue agent; and The image dataset includes a display interface image of a graphical interface product and an appearance image of a physical product carrying the graphical interface product; the text dataset includes the name and purpose of the graphical interface product, and description information of the display interface image; The performing of intent recognition, entity extraction, redundant information decoupling, feature extraction, and feature fusion on the text dataset and the image dataset to obtain appearance retrieval features representing appearance retrieval intent includes: Performing intent recognition on the text dataset and the image dataset to obtain an intent recognition category as a graphical interface appearance retrieval category; Performing a first entity extraction on the text dataset according to the graphical interface appearance retrieval category to obtain a descriptive entity; and performing a second entity extraction on the image dataset according to the graphical interface appearance retrieval category to obtain a visual entity; performing an associated image convergence operation on the image dataset according to the graphical interface appearance retrieval category and the description information of the display interface image, to obtain an image dataset after removing the physical product appearance image; Determining global text features based on the name and the purpose; and determining local text features of the display interface image based on the description information; Performing an encoding operation on the display interface image to obtain image features of the display interface image; The descriptive entity, the visual entity, the local text feature, the global text feature and the image feature of the display interface image are subjected to feature fusion to obtain the appearance retrieval feature.
7. The method according to claim 6, characterized in that The step of performing an associated image convergence operation on the image dataset according to the graphical interface appearance retrieval category and the description information of the display interface image to obtain an image dataset after excluding the physical product appearance image includes: Performing object recognition on the display interface images in the image dataset to obtain one or more display objects of each display interface image, where a single display object refers to a graphic composed of one or more graphic elements and having independent visual connotations; Performing a first software interface element relevance analysis on the display objects to obtain a first software interface element relevance of each display object; Performing a second software interface element relevance analysis on each display object according to the description information of the display interface diagram to obtain a second software interface element relevance of each display object; performing fusion processing on the first software interface element correlation degree and the second software interface element correlation degree to obtain the software interface element fusion correlation degree of each display object; Determining the software interface relevance of the display interface diagram according to the fusion relevance of one or more software interface elements of the one or more display objects; The image dataset is updated according to a comparison result of the software interface association degree and the association degree corresponding to the graphical interface appearance retrieval category, to obtain an image dataset after excluding the physical product appearance image.
8. The method according to claim 7, characterized in that The performing first software interface element relevance analysis on the display objects to obtain the first software interface element relevance of each display object includes: Performing graphic semantic analysis on each display object to obtain one or more graphic semantic keywords adapted to each display object; Using the one or more graphic semantic keywords as query identifiers, querying a preset first mapping relationship set to obtain one or more existence attributes adapted to the graphic semantic keywords, the existence attributes including virtual existence attributes and physical existence attributes, the first mapping relationship combination including a correspondence between the graphic semantic keywords and the existence attributes; The first software interface element association degree of each display object is determined according to the quantity ratio of the virtual presence attribute in the one or more presence attributes.
9. The method according to claim 8, characterized in that The performing the second software interface element relevance analysis on each display object according to the description information of the display interface diagram to obtain the second software interface element relevance of each display object includes: Performing text semantic analysis on the description information of the display interface image to obtain one or more text semantic keywords adapted to each display object; Using the one or more text semantic keywords as query identifiers, querying a preset second mapping relationship set to obtain one or more existence attributes of the adapted text semantic keywords, the existence attributes including virtual existence attributes and physical existence attributes, the second mapping relationship combination including a correspondence between the text semantic keywords and the existence attributes; The second software interface element association degree of each display object is determined according to the quantity ratio of the virtual presence attribute in the one or more presence attributes.
10. A product appearance information intelligent processing system based on conversation context, characterized in that: Including terminal equipment and servers, among which, The terminal device is configured to execute the steps executed by the terminal device in the method according to any one of claims 1 to 9; The server is configured to execute the steps executed by the server in the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Multi-robot dialogue method and system for software-as-a-service platform
CN117573834A
System and Method for Extremely Efficient Image and Pattern Recognition and Artificial Intelligence Platform
US20220121884A1