Method and system for intelligent processing of product appearance information based on dialogue context

By using multimodal data processing and employing intent recognition, entity extraction, and feature fusion techniques, the problem of low retrieval accuracy caused by the diversity of user-uploaded images and text descriptions was solved, achieving efficient appearance-based retrieval results.

CN120744147BActive Publication Date: 2025-11-07深圳市睿观信息科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511241098.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-02
Publication Date
2025-11-07
Estimated Expiration
2045-09-02

AI Technical Summary

Technical Problem

Existing product appearance retrieval and compliance testing systems suffer from low query accuracy due to the diverse forms of images and text descriptions uploaded by users. They cannot accurately present true and relevant results and thus fail to meet user needs.

Method used

By processing multimodal data (text and images), and utilizing intent recognition, entity extraction, redundant information processing, feature extraction and fusion, effective features representing appearance retrieval intent are generated, enabling accurate understanding and processing of users' appearance retrieval intent.

Benefits of technology

It improves the accuracy and efficiency of AI dialogue agents in appearance retrieval tasks, and provides accurate search results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744147B_ABST
    Figure CN120744147B_ABST
Patent Text Reader

Abstract

The application provides a product appearance information intelligent processing method and system based on dialogue context, belongs to the technical field of data processing systems or methods specially applicable to administrative, commercial, financial, management, supervision or prediction purposes, utilizes multi-modal data (text and pictures), realizes accurate understanding and processing of user appearance retrieval intention through operations such as intention recognition, entity extraction, redundant information processing, feature extraction and fusion, effectively extracts key information from complex data, optimizes data quality, generates effective features representing appearance retrieval intention, and then accurately performs appearance retrieval operation, finally provides accurate retrieval results for users, and improves the accuracy and efficiency of AI dialogue intelligent agents in appearance retrieval tasks.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of data processing systems or methods specially adapted for administrative, commercial, financial, managerial, supervisory or predictive purposes, and particularly relates to a product appearance information intelligent processing method and system based on dialogue context. BACKGROUND

[0002] The merchant user of the e-commerce platform can use the related services of the third-party similar product query and compliance detection system for the product to be listed before the product is listed or in the development stage, so as to timely learn whether the product has similar appearance, trademark, etc.

[0003] At present, the related product appearance retrieval and compliance detection system can search and present the results of similar appearance patents according to the product pictures and / or text descriptions uploaded by the user, but in the actual use scene, due to the diversity and uncontrollability of the performance forms of the product pictures uploaded by the user and the text descriptions actually entered, for example, mixed non-related pictures, the query algorithm of the related system does not make identification and distinction, which leads to low query accuracy, the real associated results cannot be accurately presented to the user end in a large number of noise results, and the user use demand cannot be met. SUMMARY

[0004] The present application provides a product appearance information intelligent processing method and system based on dialogue context, which utilizes multi-modal data (text and picture), realizes accurate understanding and processing of user appearance retrieval intention through intention recognition, entity extraction, redundant information processing, feature extraction and fusion, effectively extracts key information from complex data, optimizes data quality, generates effective features representing appearance retrieval intention, and then accurately performs appearance retrieval operation, finally provides accurate retrieval results for users, and improves the accuracy and efficiency of AI dialogue intelligent agent in appearance retrieval task.

[0005] In a first aspect, the present application provides a product appearance information intelligent processing method based on dialogue context, applied to an AI dialogue intelligent agent, the method comprising: obtaining first dialogue data of a first dialogue event between a user and the AI dialogue intelligent agent, the first dialogue data comprising a text data set and a picture data set, the picture data set comprising a first picture uploaded by a picture editing component of the AI dialogue intelligent agent and / or a second picture generated by the AI dialogue intelligent agent according to dialogue data in the first text data set, or the picture data set is empty; performing intention recognition, entity extraction, redundant information decoupling, feature extraction and feature fusion on the text data set and the picture data set to obtain appearance retrieval features representing appearance retrieval intention; performing appearance retrieval operation according to the appearance retrieval features to obtain appearance retrieval results; and displaying the appearance retrieval results.

[0006] In a second aspect, the present application provides a product appearance information intelligent processing system based on dialogue context, comprising a terminal device and a server, wherein the terminal device is configured to perform the steps performed by the terminal device in any one of the methods of the first aspect; and the server is configured to perform the steps performed by the server in any one of the methods of the first aspect.

[0007] It can be seen that, in the embodiments of the present application, by obtaining first dialogue data of a first dialogue event of a user and an AI dialogue agent, the first dialogue data includes a text data set and a picture data set, the picture data set includes a first picture uploaded by a picture editing component of the AI dialogue agent and / or a second picture generated by the AI dialogue agent according to dialogue data in the first text data set, or the picture data set is empty; the text data set and the picture data set are subjected to intent recognition, entity extraction, redundant information decoupling, feature extraction, and feature fusion to obtain appearance retrieval features representing an appearance retrieval intent; an appearance retrieval operation is performed according to the appearance retrieval features to obtain an appearance retrieval result; and the appearance retrieval result is displayed. It can be seen that, by using multi-modal data (text and picture), the present application realizes accurate understanding and processing of the appearance retrieval intent of the user through operations such as intent recognition, entity extraction, redundant information processing, feature extraction, and feature fusion, effectively extracts key information from complex data, optimizes data quality, generates effective features representing the appearance retrieval intent, and then accurately performs the appearance retrieval operation, so as to finally provide the user with accurate retrieval results, thereby improving the accuracy and efficiency of the AI dialogue agent in the appearance retrieval task. BRIEF DESCRIPTION OF DRAWINGS

[0008] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort.

[0009] Figure 1 A result schematic diagram of a product appearance information intelligent processing system provided by an embodiment of the present application;

[0010] Figure 2 A flow schematic diagram of a product appearance information intelligent processing method based on dialogue context provided by an embodiment of the present application;

[0011] Figure 3 An appearance detection dialogue page of an AI dialogue agent provided by an embodiment of the present application;

[0012] Figure 4 A schematic diagram of a picture editing component provided by an embodiment of the present application;

[0013] Figure 5 A structural schematic diagram of a server provided by an embodiment of the present application. DETAILED DESCRIPTION

[0014] In order to enable persons skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by persons skilled in the art without creative work fall within the scope of protection of the present application.

[0015] The terms “first”, “second”, and the like in the specification of the present application, the claims and the above drawings are used to distinguish different objects, and are not used to describe a specific sequence. In addition, the terms “include” and “have” and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but can optionally include other steps or units not listed or can optionally include other steps or units inherent to the process, method, product or device.

[0016] In this document, the term “embodiment” means that the specific features, structures or characteristics described in connection with the embodiment can be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor is it mutually exclusive or alternative to other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0017] In the embodiments of the present application, “and / or” describes the association relationship of the associated objects, which means that there can be three kinds of relationships. For example, A and / or B can represent the following three cases: A exists alone; A and B exist simultaneously; B exists alone. Wherein, A and B can be singular or plural.

[0018] In the embodiments of the present application, the symbol “ / ” can represent that the associated objects before and after it are in an “or” relationship. In addition, the symbol “ / ” can also represent the division sign, that is, to perform division operation. For example, A / B can represent A divided by B.

[0019] In the embodiments of the present application, “at least one” or similar expressions refer to any combination of these items, including any combination of single or multiple items, refer to one or more, and multiple refers to two or more. For example, at least one of a, b or c can represent the following seven cases: a, b, c, a and b, a and c, b and c, a, b and c. Among them, each of a, b and c can be an element or a set containing one or more elements.

[0020] In the embodiments of the present application, “equal to” can be used with greater than, which is applicable to the technical solutions adopted when greater than, or can be used with less than, which is applicable to the technical solutions adopted when less than. When equal to is used with greater than, it is not used with less than; when equal to is used with less than, it is not used with greater than.

[0021] The present application provides a product appearance information intelligent processing method and system based on dialogue context, which utilizes multi-modal data (text and picture), realizes accurate understanding and processing of user appearance retrieval intention through intention recognition, entity extraction, redundant information processing, feature extraction and fusion, effectively extracts key information from complex data, optimizes data quality, generates effective features representing appearance retrieval intention, and then accurately performs appearance retrieval operation, finally provides accurate retrieval results for users, and improves the accuracy and efficiency of AI dialogue intelligent agent in appearance retrieval task.

[0022] The technical solutions of the present application and how the technical solutions of the present application solve the above technical problems will be described in detail in the specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of the present application will be described below with reference to the accompanying drawings.

[0023] Please refer to Figure 1 and Figure 2 , Figure 1 The result schematic diagram of a product appearance information intelligent processing system provided by the embodiments of the present application is shown in Figure 2 The flowchart of a product appearance information intelligent processing method based on dialogue context provided by the embodiments of the present application is shown in

[0024] The product appearance information intelligent processing system 1 includes a server 10 and a terminal device 20, and the server 10 and the terminal device 20 are in communication connection.

[0025] Specifically, server 10 may include a backend server responsible for data processing, applied on the network platform side, capable of data transmission, data processing, and other functions. It can be a physical server, a server cluster composed of multiple physical servers, or a distributed system. In this embodiment, the number of servers is not specifically limited. Alternatively, it may be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.

[0026] Specifically, terminal device 20 may include a front-end device applied to the user side, capable of data acquisition and data transmission in signal-free scenarios. This device may be a user equipment (UE) such as a mobile phone, smartphone, laptop, digital broadcast receiver, personal digital assistant (PDA), or tablet computer (PAD), a handheld device, in-vehicle device, wearable device, computing device, or other processing device connected to a wireless modem, a mobile station (MS), or a mobile terminal. Alternatively, terminal device 20 may also be a software application capable of running on the aforementioned electronic devices.

[0027] Terminal device 20 Figure 2 The execution entity of the intelligent processing method for product appearance information based on dialogue context shown in the figure includes the following steps S201-S204:

[0028] Step S201: Obtain the first dialogue data of the first dialogue event between the user and the AI ​​dialogue agent.

[0029] The first dialogue data includes a text dataset and an image dataset. The image dataset includes a first image uploaded by the image editing component of the AI ​​dialogue agent and / or a second image generated by the AI ​​dialogue agent based on the dialogue data in the first text dataset, or the image dataset is empty.

[0030] In this process, users input text information data through multi-round dialogue with the AI ​​dialogue agent, upload the first image through the AI ​​agent's image editing component, or generate the second image based on the dialogue data, or have no image, or the image dataset is empty; and obtain the text dataset and image dataset based on the input text information data and / or image data.

[0031] Among them, for the text data entered by the user, the AI dialogue agent automatically generates targeted clarification questions from the user dialogue interaction dimension when the system detects that the user input is ambiguous or incomplete, helping the user to clarify the demand. Introduce spelling correction and semantic correction mechanism, fault tolerance processing of user's misinput, improve the robustness of the system. In the query process, the input of the user is anonymized, and the data privacy and security standards are followed.

[0032] Step S202, the text data set and the picture data set are subjected to intent recognition, entity extraction, redundant information decoupling, feature extraction, and feature fusion to obtain appearance retrieval features representing appearance retrieval intent.

[0033] Among them, the intent recognition of the text data set refers to: using natural language processing technology, such as pre-trained models based on deep learning, such as BERT, GPT variants, etc., to comprehensively analyze the vocabulary, syntax structure, and semantic information in the text. The model will learn a large amount of text data with labeled intent graphs. When facing new text data, it predicts the intent category it belongs to, such as query descriptive text of certain product appearance, or compare different product appearance reviews, and identifies the intent tendency related to appearance retrieval.

[0034] The intent recognition of the picture data set refers to: using computer vision technology, such as convolutional neural network (CNN) architecture model. CNN can capture different levels of visual features of pictures, and then judge the potential intent behind the pictures, such as a clothing picture, which is used to show the dressing style, or focus on the appearance details of the clothing single product, so as to distinguish the picture intent related to appearance retrieval.

[0035] After this step, the text and pictures have preliminary intent classification results respectively, providing direction for subsequent processing.

[0036] In some embodiments, the intent classification result includes but is not limited to: graphical interface appearance retrieval category, entity product shape retrieval category, text trademark retrieval category, graphic trademark retrieval category, mixed text and graphic trademark retrieval category, and e-commerce platform policy retrieval category.

[0037] Among them, the graphical interface appearance retrieval category focuses on the similarity search of the appearance of various software, electronic device and other graphical interfaces. For example, the graphical layout, icon style, graphical control setting and other aspects in the graphical interface. Through this retrieval category, the similar graphical interface appearance is retrieved according to the characteristics such as color collocation, icon shape and arrangement, interface interaction area distribution and the like. The entity product shape retrieval category is mainly aimed at the similarity search of the external shape of the entity product. The text trademark retrieval category is mainly aimed at the similarity search of the trademark composed of text. The graphical trademark retrieval category is mainly aimed at the similarity search of the trademark composed of graphical elements. The graphic-text mixed trademark retrieval category is suitable for the similarity search of the trademark composed of text and graphics, and the graphic-text combination mode, text and graphic characteristics and their mutual relationship are comprehensively considered during the retrieval. The e-commerce platform policy retrieval category refers to the policies of the platform regarding the commodity picture display specification, trademark use restriction, appearance infringement judgment standard and the like before the commodity of the merchant on the e-commerce platform, so as to ensure that the commodity display, trademark use and other behaviors conform to the platform rules and avoid being punished for violation.

[0038] Among them, the entity extraction of the text data set refers to: on the basis of identifying the text of the appearance retrieval intention, the key entity is extracted through the named entity recognition (NER) algorithm. For example, for “I want to find a picture of a red flip phone”, the entities such as “red” and “flip phone” closely related to the appearance attribute are extracted, and the text is simplified into the core appearance description element.

[0039] The entity extraction of the picture data set refers to: the object entity in the picture is identified by using image segmentation and target detection technology, for example, the sofa entity is detected from a home scene picture, and the contour, position and other information of the sofa entity are labeled, and the visual entity directly related to the appearance is sorted out. These extracted entities will be used as the material for more accurate processing in the subsequent steps.

[0040] In specific implementation, in some embodiments, the intent recognition and entity extraction of the text data set and the picture data set include the following steps: (1) multi-language support: introduce multi-language segmentation and word vector technology to support cross-language intent recognition and entity extraction. (2) entity recognition optimization: use named entity recognition (NER) and ontology library in the patent field to accurately extract key entities related to design and improve the accuracy of retrieval. (3) model training: use pre-trained language models (such as BERT and GPT) for fine-tuning, combine with patent field corpus, and improve the accuracy of intent recognition. (4) intent classification and similarity calculation: use a hierarchical intent classification system and use cosine similarity and other measurement methods to improve the accuracy of intent matching. (5) dynamic intent update and multi-round dialogue tracking: update the user intent model in real time through the context perception mechanism and dialogue history to ensure effective tracking of multi-round conversations.

[0041] Among them, the decoupling of redundant information for the text data set refers to removing the redundant content such as adjectives, background descriptions and other redundant content irrelevant to the appearance retrieval intention from the text after entity extraction, so that the text information is more focused on the appearance key elements.

[0042] The decoupling of redundant information for the picture data set refers to removing elements that interfere with appearance judgment in the picture background, using background removal algorithm or image cropping technology, such as a product picture with complex promotional labels and irrelevant ornaments around the product, highlighting the product main body through algorithm, eliminating residual visual information, so that the subsequent fused information is more pure.

[0043] After this step, the text and picture data are further "purified", which is beneficial to feature fusion.

[0044] Among them, feature extraction refers to obtaining various feature information that can represent the content of text or picture, for example, for picture data, using pre-trained visual models (such as ResNet, ViT) to extract image features, automatically extracting color features, texture features, shape features, etc. For text data, use pre-trained language models (such as BERT, GPT) to extract text features. Like counting word frequency, calculating TF-IDF value, or using word embedding to get semantic vector, using deep learning model to capture context features, so as to provide key support for text classification, sentiment analysis and other tasks.

[0045] Different from entity extraction, entity extraction focuses on separating specific objects or objects from pictures and texts. These objects are entities that can be named and located, such as identifying the product itself and the product components in product pictures. Feature extraction is an abstract feature extraction of the entire picture or text content, which does not focus on the specific entity object name, but focuses on how to describe the visual characteristics of the image with digital features.

[0046] Among them, the feature fusion of entity objects and feature objects maps entity objects and feature objects to a unified vector space through cross-modal contrast learning, supporting similarity calculation of multi-modal features. This method can effectively handle the differences between different modal data, improve the consistency of feature representation and the accuracy of similarity calculation.

[0047] Step S203, performing appearance retrieval operation according to the appearance retrieval feature to obtain appearance retrieval result.

[0048] The appearance retrieval operation implements efficient multi-modal retrieval logic, fuses text features describing the appearance and picture features representing the appearance, obtains appearance retrieval features, and calculates the similarity with candidate data in the database according to the appearance retrieval features, such as calculating the similarity by using an Euclidean distance algorithm or a cosine similarity algorithm. An adaptive ranking algorithm is introduced, and the candidate results are optimized and ranked in combination with user preferences and feedback. For example, the appearance data most similar to the text data and the picture data entered by the user can be preferentially displayed in a ranking mode from high to low similarity.

[0049] The user can provide positive or negative feedback in each round of dialogue, such as whether the most similar appearance data retrieved is accurate. The system uses these feedbacks, adopts an online learning algorithm, and adjusts the ranking model and the retrieval strategy in real time. Through long-term interaction, the system establishes a user preference model and provides personalized retrieval results for the user.

[0050] In step S204, the appearance retrieval result is displayed.

[0051] It can be seen that, in the embodiment of the present application, the first dialogue data of the first dialogue event of the user and the AI dialogue agent is obtained, the first dialogue data includes a text data set and a picture data set, the picture data set includes a first picture uploaded by a picture editing component of the AI dialogue agent and / or a second picture generated by the AI dialogue agent according to dialogue data in the first text data set, or the picture data set is empty; the text data set and the picture data set are subjected to intent recognition, entity extraction, redundant information decoupling, feature extraction, and feature fusion to obtain appearance retrieval features representing an appearance retrieval intent; an appearance retrieval operation is performed according to the appearance retrieval features to obtain an appearance retrieval result; and the appearance retrieval result is displayed. It can be seen that, by using multi-modal data (text and picture), the present application realizes accurate understanding and processing of the appearance retrieval intent of the user through operations such as intent recognition, entity extraction, redundant information processing, feature extraction, and feature fusion, effectively extracts key information from complex data, optimizes data quality, generates effective features representing the appearance retrieval intent, and then accurately performs the appearance retrieval operation to finally provide accurate retrieval results for the user, thereby improving the accuracy and efficiency of the AI dialogue agent in the appearance retrieval task.

[0052] In some embodiments, the picture data set is empty; the first dialogue data of the first dialogue event of the user and the AI dialogue agent is obtained by: obtaining multi-round text dialogue data of the first dialogue event of the user and the AI dialogue agent; and performing associated data filtering on the multi-round text dialogue data according to a dialogue theme consistency constraint condition to obtain the first dialogue data.

[0053] Wherein, the dialogue topic keyword extraction is performed on the text dialogue data of each round of the multi-round text dialogue data to obtain a plurality of dialogue topic keyword groups, each dialogue topic keyword group containing at least one topic keyword, which can specifically contain an explicit dialogue topic keyword and an implicit dialogue topic keyword. The explicit dialogue topic keyword is a directly extracted keyword, and the implicit dialogue topic keyword is a keyword obtained based on semantic understanding and expansion.

[0054] In some embodiments, the intent recognition, entity extraction, redundant information decoupling, feature extraction, and feature fusion on the text data set and the picture data set to obtain appearance retrieval features representing appearance retrieval intent include: performing intent recognition on the first dialogue data to obtain an intent recognition category of a graphical interface appearance retrieval category; performing entity extraction on the first dialogue data according to the graphical interface appearance retrieval category to obtain text entities; performing redundant information decoupling processing on the first dialogue data according to the graphical interface appearance retrieval category to obtain processed first dialogue data; performing text encoding operation on the processed first dialogue data to obtain text features; and performing fusion processing on the text entities and the text features to obtain appearance retrieval features representing appearance retrieval intent.

[0055] Wherein, when the intent recognition category is determined to be the graphical interface appearance retrieval category according to the intent recognition on the first dialogue data, entity extraction is performed on the first dialogue data to obtain text entities. Exemplarily, the text entities include but are not limited to color matching of graphical interface appearance, icon shape and arrangement, interface interaction region distribution, and the like. Redundant information decoupling processing is performed on the first dialogue data, and text encoding operation is performed on the processed first dialogue data to obtain text features. Fusion processing is performed on the text entities and the text features to obtain appearance retrieval features.

[0056] As can be seen, in this embodiment, after the intent recognition determines the graphical interface appearance retrieval category, entity extraction is performed to obtain key features such as color matching, icon shape and arrangement, interface interaction region distribution, and the like, which accurately anchor the graphical interface appearance elements that the user is concerned about. This can avoid interference of other irrelevant information and focus the retrieval on the core demand. Redundant information decoupling processing further purifies the data, removes noise, and ensures the accuracy of subsequent text encoding and feature fusion. After such processing, the obtained appearance retrieval features are highly consistent with the user demand, the retrieval result is more accurate, and the false positives and false negatives are reduced.

[0057] In some embodiments, the picture data set includes a second picture generated by the AI dialogue agent according to the dialogue data in the first text data set; and the first dialogue data of the first dialogue event between the user and the AI dialogue agent is obtained through the following steps A1-A5:

[0058] Step A1, obtaining multi-round text conversation data of a first conversation event of a user with the AI conversation agent.

[0059] Step A2, performing picture relevance analysis and classification processing on the multi-round text conversation data to obtain a first conversation data set of irrelevant data and a second conversation data set of relevant data.

[0060] Step A3, extracting picture data in the second conversation data set and performing deduplication processing to obtain the second picture.

[0061] Step A4, extracting text data in the second conversation data set and merging with the first conversation data set to obtain a third conversation data set.

[0062] Step A5, performing associated data filtering on the third conversation data set according to a conversation theme consistency constraint condition to obtain the first conversation data.

[0063] Among them, the first conversation data set is conversation data irrelevant to pictures, and the second conversation data set is conversation data related to pictures. The second picture is generated according to the conversation data related to the picture in the second conversation data set, and the text data in the second conversation data set is filtered and merged with the first conversation data set to obtain a third conversation data set. The third conversation data set is filtered to obtain at least one first conversation data with the same theme.

[0064] As can be seen, in this embodiment, through a series of processes such as picture relevance analysis, classification, extraction, deduplication, merging, and filtering based on theme consistency on multi-round text conversation data, effective separation and integration of picture-related and picture-irrelevant information in the conversation data are realized, data redundancy is removed, data quality is improved, and the conversation data obtained finally is more consistent in theme, providing a high-quality, clear-structured, and theme-specific data basis for subsequent in-depth analysis, model training, and intelligent applications (such as intelligent customer service, etc.) based on conversation data, enhancing the usability and effectiveness of data in actual application scenarios, and helping to improve the accuracy of user intent understanding of related systems, optimizing service experience and system performance.

[0065] In some embodiments, the intent recognition, entity extraction, redundant information decoupling, feature extraction, and feature fusion on the text data set and the picture data set to obtain appearance retrieval features representing appearance retrieval intent include:

[0066] Performing intent recognition on the first conversation data and the second picture to obtain an intent recognition category of a graphical interface appearance retrieval category;

[0067] perform first entity extraction on the first dialogue data according to the graphic interface appearance retrieval category to obtain a description entity, and perform second entity extraction on the second picture according to the graphic interface appearance retrieval category to obtain a visual entity;

[0068] perform first redundancy information decoupling and first feature extraction on the first dialogue data according to the graphic interface appearance retrieval category to obtain text features;

[0069] perform second redundancy information decoupling and second feature extraction on the second picture according to the graphic interface appearance retrieval category to obtain image features;

[0070] perform feature fusion on the description entity, the visual entity, the text features, and the image features to obtain the appearance retrieval features.

[0071] It can be seen that in the embodiment, when the dialogue data includes text data and picture data, and the retrieval category is identified as the graphic interface appearance retrieval category according to the dialogue data and the picture data, the description entity is obtained according to the first dialogue data, the visual entity is obtained according to the second picture, the text features are obtained through the first dialogue data, the image features are obtained through the second picture, and finally the appearance retrieval features are obtained through feature fusion of the description entity, the visual entity, the text features, and the entity features. The intention of the user for graphic interface appearance retrieval can be accurately captured, accurate and comprehensive entity information can be extracted from the dialogue data and the picture, redundant interference can be removed, high-quality text and image features can be obtained, and these features can be effectively fused. Thus, the fusion features that can fully represent the display interface appearance are generated, and accurate and rich information basis is provided for subsequent graphic interface appearance retrieval.

[0072] In some embodiments, the AI dialogue agent is in the form of an appearance detection dialogue page in the terminal device, the first dialogue data of the first dialogue event between the user and the AI dialogue agent is obtained by detecting a first selection operation on an image user interface option in a main dialogue interaction interface, displaying an appearance detection dialogue page in response to the first selection operation, the appearance detection dialogue page including a text dialogue component and a picture editing component, obtaining the text data set through the text dialogue component, and obtaining the picture data set through the picture editing component.

[0073] For example, see Figure 3 , Figure 3 The appearance detection dialogue page of the AI dialogue agent provided in the embodiments of the present application is as follows Figure 3As shown, the appearance detection dialogue page 3 includes a text dialogue component 31 and a picture editing component 32, wherein the text dialogue component 31 includes a user input area 311 and a dialogue display area 312, the user can input text information in the user input area 311, and then click the send button to interact with the AI dialogue agent; the user can also directly add a text file in the user input area 311 and input a processing instruction, and then click the send button to interact with the AI dialogue agent. Among them, the user uploads pictures through the picture editing component 32.

[0074] In some embodiments, the picture dataset includes a first picture uploaded through the picture editing component of the AI dialogue agent; and the picture dataset includes a display interface diagram of a graphical interface product and an entity product appearance diagram of an entity product carrying the graphical interface product; the text dataset includes a name, a use, and description information of the display interface diagram of the graphical interface product; the text dataset and the picture dataset are subjected to intent recognition, entity extraction, redundant information decoupling, feature extraction, and feature fusion to obtain appearance retrieval features representing appearance retrieval intent, including:

[0075] The text dataset and the picture dataset are subjected to intent recognition to obtain an intent recognition category of a graphical interface appearance retrieval category; the text dataset is subjected to first entity extraction according to the graphical interface appearance retrieval category to obtain description entities; and the picture dataset is subjected to second entity extraction according to the graphical interface appearance retrieval category to obtain visual entities; the picture dataset is subjected to an associated picture convergence operation according to the graphical interface appearance retrieval category and the description information of the display interface diagram to obtain a picture dataset after excluding the entity product appearance diagram; global text features are determined according to the name and the use; and local text features of the display interface diagram are determined according to the description information; the display interface diagram is subjected to an encoding operation to obtain image features of the display interface diagram; the description entities, the visual entities, the local text features, the global text features, and the image features of the display interface diagram are subjected to feature fusion to obtain the appearance retrieval features.

[0076] In some embodiments, the picture data set is subjected to an associated picture convergence operation according to the graphical interface appearance retrieval category and the description information of the display interface graph, to obtain a picture data set after the entity product appearance graph is removed, which comprises: performing object recognition on the display interface graph in the picture data set to obtain one or more display objects of each display interface graph, wherein a single display object refers to a graph composed of one or more graphical elements and having independent visual connotation; performing first software interface element correlation degree analysis on the display objects to obtain a first software interface element correlation degree of each display object; performing second software interface element correlation degree analysis on each display object according to the description information of the display interface graph to obtain a second software interface element correlation degree of each display object; performing fusion processing on the first software interface element correlation degree and the second software interface element correlation degree to obtain a software interface element fusion correlation degree of each display object; determining a software interface correlation degree of the display interface graph according to one or more software interface element fusion correlation degrees of the one or more display objects; and updating the picture data set according to a comparison result of the software interface correlation degree and the correlation degree corresponding to the graphical interface appearance retrieval category, to obtain a picture data set after the entity product appearance graph is removed.

[0077] In some embodiments, the first software interface element correlation degree of each display object is obtained by performing first software interface element correlation degree analysis on the display objects, which comprises: performing graphical semantic analysis on each display object to obtain one or more graphical semantic keywords adapted to each display object; querying a preset first mapping relationship set with the one or more graphical semantic keywords as query identifiers to obtain one or more existence attributes adapted to the graphical semantic keywords, wherein the existence attributes comprise virtual existence attributes and entity existence attributes, and the first mapping relationship set comprises a corresponding relationship between graphical semantic keywords and existence attributes; and determining the first software interface element correlation degree of each display object according to a proportion of the number of virtual existence attributes in the one or more existence attributes.

[0078] In some embodiments, the second software interface element correlation degree analysis of each display object according to the description information of the display interface graph comprises: performing text semantic analysis on the description information of the display interface graph to obtain one or more text semantic keywords adapted to each display object; taking the one or more text semantic keywords as a query identifier, querying a preset second mapping relationship set to obtain one or more existence attributes adapted to the text semantic keywords, the existence attributes comprising virtual existence attributes and entity existence attributes, and the second mapping relationship set comprising a corresponding relationship between a text semantic keyword and an existence attribute; and determining the second software interface element correlation degree of each display object according to a proportion of the number of virtual existence attributes in the one or more existence attributes.

[0079] In some embodiments, the display interface graph comprises a main view and a plurality of state change graphs, the description information comprises first description information of the main view and a first state change graph, and second description information of a second state change graph; and the local text features of the display interface graph are determined according to the description information, comprising: determining a first local text feature of the main view according to the first description information and image information of the main view; determining a second local text feature of the first state change graph according to the first description information and image information of the first state change graph; and determining a third local text feature of the second state change graph according to the second description information and image information of the second state change graph. Figure 4 Figure 4 In some embodiments, the display interface graph comprises a main view and a plurality of state change graphs, the description information comprises first description information of the main view and a first state change graph, and second description information of a second state change graph; and the local text features of the display interface graph are determined according to the description information, comprising: determining a first local text feature of the main view according to the first description information and image information of the main view; determining a second local text feature of the first state change graph according to the first description information and image information of the first state change graph; and determining a third local text feature of the second state change graph according to the second description information and image information of the second state change graph. Figure 4

[0080] In some embodiments, the display interface graph comprises a main view and a plurality of state change graphs, the description information comprises first description information of the main view and a first state change graph, and second description information of a second state change graph; and the local text features of the display interface graph are determined according to the description information, comprising: determining a first local text feature of the main view according to the first description information and image information of the main view; determining a second local text feature of the first state change graph according to the first description information and image information of the first state change graph; and determining a third local text feature of the second state change graph according to the second description information and image information of the second state change graph.

[0081] ​​In some embodiments, determining the first local text features of the main view based on the first description information and the image information of the main view includes: extracting a first purpose connotation feature from a first display object in the image information of the main view based on the first description information to obtain a first purpose connotation feature; determining the second local text features of the first state change map based on the first description information and the image information of the first state change map includes: extracting a second purpose connotation feature from a second display object in the image information of the first state change map based on the first description information to obtain a second purpose connotation feature, wherein the first display object includes the second display object; determining the third local text features of the second state change map based on the second description information and the image information of the second state change map includes: extracting a third purpose connotation feature from a third display object in the image information of the second state change map based on the second description information to obtain a third purpose connotation feature, wherein the first display object is different from the third display object.

[0082] With the above Figure 2 The embodiments shown are consistent; please refer to [link / reference]. Figure 5 , Figure 5 This application provides a schematic diagram of the structure of a server, as shown in the embodiment of the present application. Figure 5 As shown, the server 10 includes a processor 51, a memory 53, a communication interface 52, and one or more programs 531, which are stored in the memory 53 and configured to be executed by the processor 51. The programs include methods for performing the methods described in the above embodiments.

[0083] This application also provides a computer storage medium storing a computer program for electronic data interchange, which causes a computer to perform some or all of the steps of any of the methods described in the above method embodiments, wherein the computer includes an electronic device.

[0084] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0085] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0086] In several embodiments provided in the present application, it should be understood that the disclosed apparatus can be implemented in other manners. For example, the division of the apparatus embodiments described above is merely illustrative, and the division of the units can be different, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electrical, mechanical or other forms.

[0087] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, i.e., can be located in one place or distributed on multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.

[0088] In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0089] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a computer readable memory. Based on this understanding, the technical solutions of the present application essentially or the part of the prior art that contributes to the technical solutions or all or part of the technical solutions can be embodied in the form of a software product, which is stored in a memory and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the above-mentioned method of each embodiment of the present application. The aforementioned memory includes: a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.

[0090] Those of ordinary skill in the art can understand that all or part of the steps of the various methods of the above embodiments can be completed by a program instructing relevant hardware, and the program can be stored in a computer readable memory, which can include a flash disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0091] The above has introduced the embodiments of the present application in detail, and the principles and implementation manners of the present application are described by applying specific examples; the above embodiment description is only used for helping to understand the method of the present application and its core idea; meanwhile, for the general technical personnel in the art, according to the idea of the present application, the specific implementation manner and application range will have changes; and according to the above, the content of the specification should not be understood as the limitation of the present application.

Claims

1. A product appearance information intelligent processing method based on dialogue context, characterized in that, The method applied to an AI conversation agent comprises: obtaining first conversation data of a first conversation event of a user with the AI conversation agent, the first conversation data comprising a text data set and a picture data set, the picture data set comprising a first picture uploaded by a picture editing component of the AI conversation agent and / or a second picture generated by the AI conversation agent according to conversation data in the first text data set, or the picture data set being empty; performing intent recognition, entity extraction, redundant information decoupling, feature extraction and feature fusion on the text data set and the picture data set to obtain appearance retrieval features representing an appearance retrieval intent; performing an appearance retrieval operation according to the appearance retrieval features to obtain an appearance retrieval result; displaying the appearance retrieval result; the picture data set being empty; the obtaining of the first conversation data of the first conversation event of the user with the AI conversation agent comprises: obtaining multi-round text conversation data of the first conversation event of the user with the AI conversation agent; and performing associated data screening on the multi-round text conversation data according to a conversation theme consistency constraint condition to obtain the first conversation data; the picture data set comprising the second picture generated by the AI conversation agent according to the conversation data in the first text data set; the obtaining of the first conversation data of the first conversation event of the user with the AI conversation agent comprises: obtaining multi-round text conversation data of the first conversation event of the user with the AI conversation agent; performing picture correlation analysis and differentiation processing on the multi-round text conversation data to obtain a first conversation data set that is irrelevant and a second conversation data set that is relevant; extracting picture data in the second conversation data set and performing de-duplication processing to obtain the second picture; extracting text data in the second conversation data set and merging the text data with the first conversation data set to obtain a third conversation data set; and performing associated data screening on the third conversation data set according to a conversation theme consistency constraint condition to obtain the first conversation data; the AI conversation agent is in the form of an appearance detection conversation page in a terminal device; the obtaining of the first conversation data of the first conversation event of the user with the AI conversation agent comprises: detecting a first selection operation on an image user interface option in a main conversation interaction interface; and in response to the first selection operation, displaying an appearance detection conversation page, the appearance detection conversation page comprising a text conversation component and the picture editing component; the text data set is obtained through the text conversation component, and the picture data set is obtained through the picture editing component.

2. The method of claim 1, wherein, the performing of intent recognition, entity extraction, redundant information decoupling, feature extraction and feature fusion on the text data set and the picture data set to obtain appearance retrieval features representing an appearance retrieval intent comprises: performing intent recognition on the first conversation data and the second picture to obtain an intent recognition category that is a graphical interface appearance retrieval category; According to the graphic interface appearance retrieval category, first entity extraction is performed on the first dialogue data to obtain a description entity, and second entity extraction is performed on the second picture according to the graphic interface appearance retrieval category to obtain a visual entity; According to the graphic interface appearance retrieval category, first redundant information decoupling and first feature extraction are performed on the first dialogue data to obtain text features; According to the graphic interface appearance retrieval category, second redundant information decoupling and second feature extraction are performed on the second picture to obtain image features; The description entity, the visual entity, the text features, and the image features are fused to obtain the appearance retrieval features.

3. The method of claim 1, wherein, The picture dataset includes a first picture uploaded by the picture editing component of the AI dialogue intelligent agent; and The picture dataset includes a display interface diagram of a graphic interface product and an entity product appearance diagram of an entity product carrying the graphic interface product; and the text dataset includes a name, a use, and description information of the display interface diagram of the graphic interface product; The text dataset and the picture dataset are subjected to intent recognition, entity extraction, redundant information decoupling, feature extraction, and feature fusion to obtain appearance retrieval features representing an appearance retrieval intent, including: The text dataset and the picture dataset are subjected to intent recognition to obtain an intent recognition category as a graphic interface appearance retrieval category; According to the graphic interface appearance retrieval category, first entity extraction is performed on the first dialogue data to obtain a description entity, and second entity extraction is performed on the second picture according to the graphic interface appearance retrieval category to obtain a visual entity; According to the graphic interface appearance retrieval category and the description information of the display interface diagram, an associated picture convergence operation is performed on the picture dataset to obtain a picture dataset after the entity product appearance diagram is removed; According to the name and the use, global text features are determined, and according to the description information, local text features of the display interface diagram are determined; The display interface diagram is subjected to an encoding operation to obtain image features of the display interface diagram; The description entity, the visual entity, the local text features, the global text features, and the image features of the display interface diagram are fused to obtain the appearance retrieval features.

4. The method of claim 3, wherein, According to the graphic interface appearance retrieval category and the description information of the display interface diagram, an associated picture convergence operation is performed on the picture dataset to obtain a picture dataset after the entity product appearance diagram is removed, including: Object recognition is performed on the display interface diagram in the picture dataset to obtain one or more display objects of each display interface diagram, and a single display object refers to a graphic composed of one or more graphic elements and having independent visual connotation; First software interface element correlation degree analysis is performed on the display objects to obtain first software interface element correlation degrees of each display object; Second software interface element correlation degree analysis is performed on each display object according to the description information of the display interface diagram to obtain second software interface element correlation degrees of the display object; and Second software interface element correlation degree analysis is performed on each display object according to the description information of the display interface diagram to obtain second software interface element correlation degrees of the display object. fusing the first software interface element correlation degree and the second software interface element correlation degree to obtain a software interface element fused correlation degree of each display object; determining a software interface correlation degree of the display interface graph according to one or more software interface element fused correlation degrees of the one or more display objects; updating the picture data set according to a comparison result of the software interface correlation degree and a correlation degree corresponding to the graphical interface appearance retrieval category, to obtain a picture data set after eliminating the entity product appearance graph.

5. The method of claim 4, wherein, The first software interface element correlation degree analysis on the display objects comprises: performing graphical semantic analysis on each display object to obtain one or more graphical semantic keywords adapted to each display object; taking the one or more graphical semantic keywords as a query identifier to query a preset first mapping relationship set to obtain one or more existence attributes adapted to the graphical semantic keywords, the existence attributes comprising virtual existence attributes and entity existence attributes, and the first mapping relationship set comprising a corresponding relationship between graphical semantic keywords and existence attributes; determining the first software interface element correlation degree of each display object according to a proportion of the number of virtual existence attributes in the one or more existence attributes.

6. The method of claim 5, wherein, The second software interface element correlation degree analysis on each display object according to the description information of the display interface graph comprises: performing text semantic analysis on the description information of the display interface graph to obtain one or more text semantic keywords adapted to each display object; taking the one or more text semantic keywords as a query identifier to query a preset second mapping relationship set to obtain one or more existence attributes adapted to the text semantic keywords, the existence attributes comprising virtual existence attributes and entity existence attributes, and the second mapping relationship set comprising a corresponding relationship between text semantic keywords and existence attributes; determining the second software interface element correlation degree of each display object according to a proportion of the number of virtual existence attributes in the one or more existence attributes.

7. A product appearance information intelligent processing system based on dialogue context, characterized by, The terminal device and the server are included, wherein, the terminal device is configured to perform steps as described in the terminal device in any one of claims 1-6; the server is configured to perform steps in any one of claims 1-6.

Citation Information

Patent Citations

  • Multi-robot dialogue method and system for software-as-a-service platform

    CN117573834A

  • System and Method for Extremely Efficient Image and Pattern Recognition and Artificial Intelligence Platform

    US20220121884A1