Expression picture matching method and device, medium and equipment

By performing feature extraction and correlation analysis on the client's query request information, matching expression pictures are selected, which solves the problem of insufficient diversity of expression pictures in the prior art, and achieves a more efficient and accurate expression image distribution service.

CN120020911APending Publication Date: 2025-05-20TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311539381.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-17
Publication Date
2025-05-20

AI Technical Summary

Technical Problem

In the existing emoticon image distribution services, the emoticon image categories are single and lack of diversity, which cannot effectively match user needs, making it difficult for users to obtain the most suitable expression pictures in a short period of time.

Method used

By obtaining the client's query request information, performing feature extraction processing, and obtaining the request feature information, including request text feature data and request semantic feature data. Then, the correlation degree between the text label data of the candidate expression picture and the request text feature data, and the correlation degree between the semantic vector data of the candidate expression picture and the request semantic feature data are determined, so as to filter out the target candidate expression picture matching the query request information.

Benefits of technology

It improves the efficiency and accuracy of the emoticon image distribution service, enhances the user experience, improves the diversity and coverage of recall results, and ensures the matching and matching with user needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120020911A_ABST
    Figure CN120020911A_ABST
Patent Text Reader

Abstract

The invention discloses an emoticon matching method and device, a medium and equipment, and relates to the technical field of the Internet, and the method comprises the steps: obtaining query request information which is sent by a client and aims at an emoticon; performing feature extraction processing based on the query request information to obtain request text feature data and request semantic feature data; determining a first association degree between the text label data of each candidate expression picture and the request text feature data; determining a second association degree between the semantic vector data of each candidate expression picture and the request semantic feature data; determining at least one target candidate expression picture from the plurality of candidate expression pictures according to the first association degree and the second association degree; and obtaining a to-be-distributed picture sequence matched with the query request information according to the at least one target candidate expression picture. According to the method, the matching degree of the to-be-distributed emoticon picture and the query request information can be improved, and the efficiency and quality of emoticon picture distribution service are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of Internet technologies, and particularly to an expression picture matching method, apparatus, medium, and device. Background Art

[0002] Artificial Intelligence (AI) is a comprehensive technology in computer science. By studying the design principles and implementation methods of various intelligent machines, machines are enabled to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject, covering a wide range of fields, such as several major directions including natural language processing, machine learning, and deep learning. With the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0003] In the distribution service of expression pictures, the use of artificial intelligence technology to provide or recommend rich expression pictures for users has promoted the interaction between users. However, in related technologies, the types of expression pictures provided or recommended are single, lacking diversity, and cannot well match the needs of users, resulting in the problem that users cannot obtain the most satisfactory expression pictures in a short time. Summary of the Invention

[0004] To improve the efficiency and quality of the expression picture distribution service, this application provides an expression picture matching method, apparatus, medium, and device. The technical solutions are as follows:

[0005] In a first aspect, this application provides an expression picture matching method, which includes:

[0006] Obtain query request information for an expression picture sent by a client;

[0007] Based on the query request information, perform feature extraction processing to obtain request feature information, where the request feature information includes request text feature data and request semantic feature data;

[0008] Determine a first correlation degree between the text label data of each candidate expression picture among multiple candidate expression pictures and the request text feature data; the text label data is content feature data described in text obtained by performing multi-dimensional analysis on the corresponding candidate expression picture;

[0009] Determine a second correlation degree between the semantic vector data of each candidate expression picture and the request semantic feature data; the semantic vector data is content feature data described in a semantic vector obtained by performing multi-dimensional analysis on the corresponding candidate expression picture;

[0010] According to the first correlation degree and the second correlation degree, determine at least one target candidate expression picture from the multiple candidate expression pictures;

[0011] Based on the at least one target candidate expression picture, obtain a sequence of pictures to be distributed that matches the query request information, where the sequence of pictures to be distributed is part or all of the at least one target candidate expression picture.

[0012] Optionally, the method further includes:

[0013] Perform extraction processing on text elements of each candidate expression picture to obtain the original text data displayed on each candidate expression picture;

[0014] Perform text analysis processing on the original text data displayed on each candidate expression picture to obtain the text analysis result of each candidate expression picture, where the text analysis result includes semantic analysis result, entity recognition result, and sentiment analysis result;

[0015] Perform label extraction processing on the text analysis result of each candidate expression picture to obtain the first label data of each candidate expression picture;

[0016] Perform extraction processing on image features of each candidate expression picture to obtain the first image feature data of each candidate expression picture;

[0017] Perform image classification processing based on the first image feature data of each candidate expression picture to obtain the second label data of each candidate expression picture;

[0018] The text label data of each candidate expression picture includes the first label data of each candidate expression picture and the second label data of each candidate expression picture.

[0019] Optionally, the method further includes:

[0020] Determine the text description information of each candidate expression picture, where the text description information includes the original text data of each candidate expression picture and the text label data of each candidate expression picture;

[0021] Input the text description information of each candidate expression picture into a contrastive learning model to perform text feature extraction processing to obtain the text feature data of each candidate expression picture;

[0022] Map the text feature data of each candidate expression picture to the corresponding semantic vector space of the contrastive learning model to obtain the text semantic vector data of each candidate expression picture;

[0023] Input each of the candidate expression images into the contrastive learning model for image feature extraction processing to obtain the second image feature data of each candidate expression image;

[0024] Map the second image feature data of each candidate expression image to the semantic vector space to obtain the image semantic vector data of each candidate expression image;

[0025] The semantic vector data of each candidate expression image includes the text semantic vector data of each candidate expression image and the image semantic vector data of each candidate expression image.

[0026] Optionally, determining the first correlation degree between the text label data of each candidate expression image among the multiple candidate expression images and the request text feature data includes:

[0027] Obtain the label index table of the multiple candidate expression images, where the label index table contains the text label data of each candidate expression image;

[0028] Based on the text matching algorithm, determine the first correlation degree between the text label data of each candidate expression image and the request text feature data; the first correlation degree indicates the text matching degree between each candidate expression image and the query request information.

[0029] Optionally, determining the second correlation degree between the semantic vector data of each candidate expression image and the request semantic feature data includes:

[0030] Obtain the semantic vector data of each candidate expression image; the semantic vector data of each candidate expression image is determined based on the contrastive learning model, each candidate expression image, and the text description information of each candidate expression image; the text description information of each candidate expression image includes the text label data of each candidate expression image and the original text data shown on each candidate expression image;

[0031] Based on the vector space distance algorithm, determine the second correlation degree between the semantic vector data of each candidate expression image and the request semantic feature data; the second correlation degree indicates the semantic vector similarity between each candidate expression image and the query request information.

[0032] Optionally, the method further includes:

[0033] Determine the request object information associated with the client in the query request information;

[0034] Determine the collaborative relationship between the historical request object in the historical distribution information and the multiple candidate expression pictures, and the collaborative relationship is represented by a graph model;

[0035] Based on the request object information and the collaborative relationship, determine at least one first candidate expression picture from the multiple candidate expression pictures;

[0036] Update the at least one target candidate expression picture according to the at least one first candidate expression picture.

[0037] Optionally, obtaining the picture sequence to be distributed that matches the query request information according to the at least one target candidate expression picture includes:

[0038] Sort the at least one target candidate expression picture according to the first correlation degree corresponding to each target candidate expression picture and the second correlation degree corresponding to each target candidate expression picture in the at least one target candidate expression picture to obtain a first picture sequence;

[0039] Determine the request object information associated with the client in the query request information;

[0040] Determine the background information corresponding to the query request information, where the background information includes interaction context information, service type information, and picture quality requirement information;

[0041] Filter the first picture sequence according to the picture quality requirement information to obtain a second picture sequence;

[0042] Determine the first matching degree between each target candidate expression picture in the second picture sequence and the request object information;

[0043] Determine the second matching degree between each target candidate expression picture in the second picture sequence and the interaction context information;

[0044] Determine the third matching degree between each target candidate expression picture in the second picture sequence and the service type information;

[0045] Re - sort the second picture sequence according to the first matching degree, the second matching degree, and the third matching degree to obtain the picture sequence to be distributed.

[0046] Optionally, the method further includes:

[0047] In the case where the number of the at least one target candidate expression picture is lower than a preset number threshold, determine the key text information of the query request information;

[0048] Obtain multiple basic expression pictures corresponding to the multiple candidate expression pictures; none of the multiple basic expression pictures contains text;

[0049] Obtain the index vector data of each basic expression picture among the multiple basic expression pictures, where the index vector data is text label data and semantic vector data in vector form;

[0050] Determine at least one target basic expression picture from the multiple basic expression pictures according to the similarity between the index vector data of each basic expression picture and the requested semantic feature data;

[0051] Synthesize the key text information and the at least one target basic expression picture to obtain at least one second candidate expression picture;

[0052] Update the at least one target candidate expression picture according to the at least one second candidate expression picture.

[0053] In a second aspect, the present application provides an expression picture matching device, and the device includes:

[0054] An acquisition module, configured to acquire query request information for an expression picture sent by a client;

[0055] A feature extraction module, configured to perform feature extraction processing based on the query request information to obtain request feature information, where the request feature information includes request text feature data and request semantic feature data;

[0056] A text correlation determination module, configured to determine a first correlation between the text label data of each candidate expression picture among the multiple candidate expression pictures and the request text feature data; the text label data is content feature data described in text obtained by performing multi-dimensional analysis on the corresponding candidate expression picture;

[0057] A vector correlation determination module, configured to determine a second correlation between the semantic vector data of each candidate expression picture and the request semantic feature data; the semantic vector data is content feature data described in a semantic vector obtained by performing multi-dimensional analysis on the corresponding candidate expression picture;

[0058] A recall module, configured to determine at least one target candidate expression picture from the multiple candidate expression pictures according to the first correlation and the second correlation;

[0059] A sequence generation module, configured to obtain a picture sequence to be distributed that matches the query request information according to the at least one target candidate expression picture, where the picture sequence to be distributed is part or all of the at least one target candidate expression picture arranged in a preset order.

[0060] In a third aspect, the present application provides a computer-readable storage medium, in which at least one instruction or at least one program segment is stored, and the at least one instruction or at least one program segment is loaded and executed by a processor to implement an expression picture matching method as described in the first aspect.

[0061] In a fourth aspect, the present application provides a computer device, which includes a processor and a memory. At least one instruction or at least one program segment is stored in the memory, and the at least one instruction or at least one program segment is loaded and executed by the processor to implement an expression picture matching method as described in the first aspect.

[0062] In a fifth aspect, the present application provides a computer program product, which includes computer instructions, and when the computer instructions are executed by a processor, an expression picture matching method as described in the first aspect is implemented.

[0063] The expression picture matching method, device, medium and device provided by the present application have the following technical effects:

[0064] In the solution provided by the present application, based on the query request information for the expression picture sent by the client, feature extraction processing is performed to obtain request feature information including request text feature data and request semantic feature data. Feature extraction and characterization are performed on the query request information from the text dimension and the semantic vector dimension, so that the demand characteristics of the user for the expression picture can be mined more comprehensively and accurately. In the solution provided by the present application, the content of multiple candidate expression pictures is also understood and analyzed in more detail from multiple dimensions, and text label data and semantic vector data of each candidate expression picture are generated. Furthermore, according to the first correlation degree between the text label data of each candidate expression picture and the request text feature data, and the second correlation degree between the semantic vector data of each candidate expression picture and the request semantic feature data, at least one target candidate expression picture that matches the query request information can be screened out more efficiently, accurately and richly and diversely. The distribution sequence obtained according to at least one target candidate expression picture can meet the needs of the expression picture distribution service.

[0065] In the solution provided by the present application, the diversity, coverage and matching degree with user needs of the recall results can be greatly improved, the efficiency and accuracy of the expression picture distribution can be effectively improved, and the user experience can be optimized.

[0066] The additional aspects and advantages of the present application will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present application. Description of the Drawings

[0067] To more clearly illustrate the technical solutions and advantages in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0068] Figure 1 It is a schematic diagram of the implementation environment of an expression picture matching method provided by an embodiment of the present application;

[0069] Figure 2 It is a schematic flowchart of an expression picture matching method provided by an embodiment of the present application;

[0070] Figure 3 It is a schematic flowchart of a graphic and text synthesis provided by an embodiment of the present application;

[0071] Figure 4 It is a schematic flowchart of a [specific content not filled] provided by an embodiment of the present application;

[0072] Figure 5 It is a schematic flowchart of a [specific content not filled] provided by an embodiment of the present application;

[0073] Figure 6 It is a schematic diagram of the architecture of an expression picture distribution system provided by an embodiment of the present application;

[0074] Figure 7 It is a schematic diagram of an expression picture matching device provided by an embodiment of the present application;

[0075] Figure 8 It is a schematic diagram of the hardware structure of a device for implementing an expression picture matching method provided by an embodiment of the present application. Detailed implementation manners

[0076] Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making. Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics.

[0077] The solutions provided in the embodiments of this application involve technologies such as Machine Learning (ML), Deep Learning (DL), and Computer Vision (CV) in artificial intelligence.

[0078] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. Pre-trained models are the latest development results of deep learning, integrating the above technologies.

[0079] Deep Learning (DL) is a major research direction in the field of Machine Learning (ML). It is introduced into machine learning to make it closer to the original goal - artificial intelligence. Deep learning is to learn the internal laws and representation levels of sample data, and the information obtained in these learning processes is very helpful for the interpretation of data such as text, images, and sounds. Its ultimate goal is to enable machines to have the ability of analysis and learning like humans, and be able to recognize data such as text, images, and sounds. Deep learning is a complex machine learning algorithm, and the effects achieved in speech and image recognition far exceed those of previous related technologies. Deep learning has achieved many results in search technology, data mining, machine learning, machine translation, natural language processing, multimedia learning, speech, recommendation and personalization technology, and other related fields. Deep learning enables machines to imitate human activities such as audiovisual and thinking, solves many complex pattern recognition problems, and makes great progress in artificial intelligence-related technologies.

[0080] Computer Vision Technology (CV) Computer vision is a science that studies how to make machines "see". Further, it refers to using cameras and computers to replace human eyes for machine vision such as target recognition, tracking, and measurement, and further performing graphics processing to make the computer-processed images more suitable for human eyes to observe or be transmitted to instruments for detection. As a scientific discipline, computer vision studies related theories and technologies, and attempts to build artificial intelligence systems that can obtain information from images or multi-dimensional data. Large model technology has brought important changes to the development of computer vision technology. Pretrained models in the visual field such as swin-transformer, ViT, V-MOE, and MAE can be quickly and widely applied to downstream specific tasks after fine-tuning. Computer vision technology usually includes image processing, image recognition, image semantic understanding, image retrieval, OCR, video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D technology, virtual reality, augmented reality, simultaneous localization and mapping, etc. technologies, and also includes common biometric recognition technologies such as face recognition and fingerprint recognition.

[0081] The solution provided in the embodiments of this application can be deployed in the cloud, which also involves cloud technology, etc.

[0082] Cloud technology: It refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or a local area network to achieve data computing, storage, processing, and sharing. It can also be understood as the general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. It can form a resource pool, be used as needed, and is flexible and convenient. The background services of the technical network system require a large amount of computing and storage resources, such as video websites, picture websites, and more portal websites. With the high development and application of the Internet industry, in the future, each item may have its own identification mark and needs to be transmitted to the background system for logical processing. Data at different levels will be processed separately, and various industry data requires the support of a powerful system. Therefore, cloud technology needs to be supported by cloud computing. Cloud computing is a computing model that distributes computing tasks on a resource pool composed of a large number of computing devices, enabling various application systems to obtain computing power, storage space, and information services as needed. The network that provides resources is called the "cloud". The resources in the "cloud" seem to users to be infinitely expandable, can be obtained at any time, used as needed, expanded at any time, and paid according to usage. As a basic capability provider of cloud computing, a cloud computing resource pool platform will be established, abbreviated as the cloud platform, generally referred to as Infrastructure as a Service (IaaS). Various types of virtual resources are deployed in the resource pool for external customers to select and use. The cloud computing resource pool mainly includes: computing devices (which can be virtual machines, including operating systems), storage devices, and network devices.

[0083] To improve the efficiency and quality of the emoticon picture distribution service, the embodiments of the present application provide an emoticon picture matching method, device, medium, and device. The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions from beginning to end.

[0084] It should be noted that in the description of the present application, the claims and the above-mentioned drawings, terms such as "first" and "second" are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or server that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products or devices.

[0085] It can be understood that in the specific implementation of the present application, when related data such as object information and historical distribution information are involved, when the above embodiments of the present application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of the related data need to comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0086] Please refer to Figure 1 , which is a schematic diagram of the implementation environment of an expression picture matching method provided by an embodiment of the present application. As Figure 1 shown, the implementation environment may at least include a client 01 and a server 02.

[0087] Specifically, the client 01 may include devices such as smart phones, desktop computers, tablet computers, laptop computers, vehicle-mounted terminals, digital assistants, smart wearable devices, and voice interaction devices, and may also include software running on the devices. For example, web pages provided by some service providers to users, or applications provided by these service providers to users. Specifically, the client 01 may be used to obtain input information of the user for the expression picture, generate a query request information, and send the query request information to the server 02. The client 01 may also be used to receive the picture sequence to be distributed returned by the server 02 and display it.

[0088] Specifically, the server 02 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The server 02 can include a network communication unit, a processor, a memory, and so on. The terminal and the server can be directly or indirectly connected through wired or wireless communication methods, and this application does not limit this. Specifically, the server 02 can be used to perform feature extraction processing based on the query request information to obtain request feature information, where the request feature information includes request text feature data and request semantic feature data; determine the first correlation degree between the text label data of each candidate expression picture among the multiple candidate expression pictures and the request text feature data, where the text label data is content feature data described in text obtained by performing multi-dimensional analysis on the corresponding candidate expression picture; determine the second correlation degree between the semantic vector data of each candidate expression picture and the request semantic feature data, where the semantic vector data is content feature data described by a semantic vector obtained by performing multi-dimensional analysis on the corresponding candidate expression picture; determine at least one target candidate expression picture from the multiple candidate expression pictures according to the first correlation degree and the second correlation degree; and obtain a picture sequence to be distributed that matches the query request information according to the at least one target candidate expression picture, where the picture sequence to be distributed is part or all of the at least one target candidate expression picture.

[0089] The embodiments of this application can also be implemented in combination with cloud technology. Cloud technology (Cloud technology) refers to a hosting technology that unifies a series of resources such as hardware, software, and networks within a wide area network or a local area network to achieve data calculation, storage, processing, and sharing, and can also be understood as the general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. Cloud technology requires cloud computing as a support. Cloud computing is a computing model that distributes computing tasks on a resource pool composed of a large number of computing devices, enabling various application systems to obtain computing power, storage space, and information services as needed. The network that provides resources is called the "cloud". Specifically, the server 02 and the database are located in the cloud, and the server 02 can be a physical machine or a virtualized machine.

[0090] The following introduces a method for matching expression pictures provided by this application. Figure 2It is a flowchart of a method for matching expression pictures provided by an embodiment of the present application. The present application provides the method operation steps as described in the embodiment or flowchart, but based on routine or non-creative labor, it may include more or fewer operation steps. The step sequence listed in the embodiment is only one way among the execution sequences of numerous steps and does not represent the only execution sequence. When the actual system or server product executes, it can be executed in the order of the method shown in the embodiment or the accompanying drawings or executed in parallel (for example, in an environment of parallel processors or multi-threaded processing). Please refer to Figure 2 , a method for matching expression pictures provided by an embodiment of the present application may include the following steps:

[0091] S210: Obtain the query request information for the expression picture sent by the client.

[0092] A method for matching an expression picture provided by an embodiment of the present application can be applied to the distribution service of expression pictures. The distribution service of expression pictures may include, but is not limited to, application scenarios such as searching for expression pictures, recommending expression pictures, and associating expression pictures. In the distribution service of expression pictures, the server responds to the query request information sent by the user based on the client, filters out the matching expression picture sequence, and pushes the expression picture sequence to the client for the user to use.

[0093] In the embodiment of the present application, the query request information may include, but is not limited to, the user's input information, current interaction information, request object information associated with the user and the client, interaction context information, service type information, etc. Among them, the user's input information can be of text type or picture type, meeting the requirements of applications such as searching for pictures by text and searching for pictures by picture; among them, the current interaction information can indicate the current interaction content of the user in interaction activities such as chatting, to meet the application requirements of actively recommending expression pictures.

[0094] S220: Based on the query request information, perform feature extraction processing to obtain request feature information, and the request feature information includes request text feature data and request semantic feature data.

[0095] In an embodiment of the present application, the feature extraction process based on the query request information includes, but is not limited to, text understanding processing, picture understanding processing, semantic vector representation processing, etc. The obtained request text feature data is request feature data of text type, indicating attribute information such as entity name, action name, emotional state, geographical location, etc. in the query request information; the request semantic feature data is request feature data of vector type, representing attribute information such as entity name, action name, emotional state, geographical location, etc. in the query request information in the form of a vector. It can be understood that the request text feature data can indicate the content features to be displayed, the request semantic feature data can indicate implicit vector features, and the request text feature data and the request semantic feature data more comprehensively and accurately indicate the demand features for the expression pictures.

[0096] In an embodiment of the present application, step S220 may be implemented as:

[0097] S221: When the input information in the query request information is of text type, perform text understanding processing on the input information to obtain request text feature data; the request text feature data includes feature data indicating the emotional state.

[0098] Specifically, the text understanding processing includes, but is not limited to, processing such as parsing, rewriting, and intent recognition of the query request information. Among them, the parsing processing may include preprocessing such as case conversion and traditional / simplified Chinese conversion, word segmentation processing, entity recognition, determination of word weights, etc. The rewriting processing may include processing such as error correction, expansion, and normalization of the input information. The intent recognition processing may combine the historical expression picture distribution information of the same user to recognize the user's intent or the intensity of the intent, etc.

[0099] In an embodiment of the present application, the request text feature data obtained by performing text understanding processing on the input information is request feature data of text type, which can indicate attribute information such as entity name, action name, emotional state, geographical location, etc. of the request. Considering that the expression pictures distributed in the embodiments of the present application are also called emoji packs, the emotional state included in the input information will be emphasized in text understanding.

[0100] Exemplarily, in an application for searching for expression pictures through input text, common input information can be in the following forms of expression: (1) entity names, such as popular names of people, character names, names of movies and TV dramas, etc., as well as aliases and abbreviations of entities; (2) relatively colloquial descriptions of emotions and actions, such as "happy", "hug", "want to sleep", etc.; (3) combinations of entities and actions, such as "come on, little friend", "XX said you're right"; (4) popular memes, phrases, such as "awesome", "professional team", "who am I and where am I"; (5) grammatically complete sentences, such as "Hello, nice to meet you". It is possible to perform text understanding on the input information based on the above several forms of expression, extract entity names, action names therein, identify the emotions or emotional states, etc. as request feature data, and represent them in the form of text characters.

[0101] S223: Perform vector representation processing on the request text feature data to obtain request semantic feature data.

[0102] The vector representation processing of text semantics is also to represent the text with an embedding vector. It can be understood that converting the request feature data in text form into request feature data in vector form is also to convert the request feature data into the semantic vector space, and use higher-order, deep and implicit feature vectors to indicate the demand features for expression pictures.

[0103] In a feasible implementation manner, the vector representation processing of text semantics can be implemented based on the BERT (Bidirectional Encoder Representation from Transformers, a bidirectional language representation model) model.

[0104] In a feasible implementation manner, when the input information in the query request information is of text type, the vector representation processing of text semantics can be directly performed on the input information to obtain more original request semantic feature data.

[0105] In the above embodiments, extracting and representing the input information of text type from the text dimension and the semantic vector dimension can more comprehensively and accurately mine the demand features of users for expression pictures.

[0106] In another embodiment of the present application, step S220 can also be implemented as:

[0107] S222: When the input information in the query request information is of picture type, perform picture analysis processing on the input information to obtain text description information corresponding to the input information.

[0108] It can be understood that in the emoji picture distribution service, similar emoji pictures can be provided for the user according to the picture input by the user. Then, when the input information is of the picture type, the picture analysis and processing performed on the input information may include, but are not limited to, OCR recognition (Optical Character Recognition), text understanding for recognizing text, image detection and classification, etc. Among them, OCR recognition refers to recognizing the original text data in the input picture and taking the original text data as a type of text description information; text understanding for recognizing text can perform entity recognition, sentiment analysis, etc. on the recognized text based on natural language processing technology; image detection and classification processing can be based on computer vision technology, and detect or recognize the entity categories, action categories, and expression categories that appear in the input picture through the extracted image pixel point feature data, and generate a type of text description information.

[0109] S224: Determine the request text feature data based on the text description information.

[0110] S226: Perform vector representation processing on the request text feature data for text semantics to obtain the request semantic feature data.

[0111] In a feasible implementation manner, steps S224 to S226 are similar to the foregoing steps S221 to S223, and will not be elaborated here.

[0112] In another feasible implementation manner, when the input information is of the picture type, in addition to obtaining the request semantic feature data corresponding to the request text feature data, the feature vector representing the image content extracted during the picture understanding process can also be used as the request semantic feature data.

[0113] In another feasible implementation manner, the input picture and its text description information are input into the aligned CLIP (Contrastive Language-Image Pre-Training) model to obtain the image semantic vector data corresponding to the input picture and the text semantic vector data corresponding to the text description information, and the image semantic vector data and the text semantic vector data are used as the request semantic feature data.

[0114] In the above embodiments, feature extraction and representation are performed on the input information of the picture type from the text dimension and the semantic vector dimension, which can more comprehensively and accurately mine the demand characteristics of the user for emoji pictures.

[0115] In another embodiment of the present application, in addition to the input information, text analysis or graph analysis processing can also be performed on the text and graph information in the foregoing current interaction information, and the specific process is the same and will not be elaborated here.

[0116] S230: Determine a first correlation degree between the text label data of each candidate expression picture among the multiple candidate expression pictures and the request text feature data.

[0117] In the embodiments of the present application, in order to meet the requirements of the expression picture distribution service, a plurality of candidate expression pictures are pre-stored in a database, and the text label data and semantic vector data of each candidate expression picture are pre-generated to meet the requirements of the recall process. The text label data is content feature data described in text obtained by performing multi-dimensional analysis on the corresponding candidate expression picture.

[0118] In the embodiments of the present application, from the perspective of text, by examining the first correlation degree between the text label data of each candidate expression picture among the multiple candidate expression pictures and the request text feature data, a part of the candidate expression pictures can be recalled. The first correlation degree between the text label data of each of this part of the candidate expression pictures and the request text feature data meets a preset first matching condition, and the first matching condition can indicate the threshold that the first correlation degree corresponding to the recalled candidate expression picture needs to meet.

[0119] Specifically, step S230 can be implemented as:

[0120] S231: Obtain a label index table of the multiple candidate expression pictures.

[0121] Feasibly, the label index table of the multiple candidate expression pictures can be stored in an inverted index system, and the label index table can include the text label data of each candidate expression picture.

[0122] S232: Based on a text matching algorithm, determine a first correlation degree between the text label data of each candidate expression picture and the request text feature data; the first correlation degree indicates the text matching degree between each candidate expression picture and the query request information.

[0123] The text matching algorithm refers to an algorithm that determines whether two texts are similar or matched by calculating the similarity or distance between texts under certain semantic and syntactic rules. Commonly used calculation methods for text similarity or matching degree include algorithms such as TF-IDF (Term Frequency-Inverse Document Frequency) algorithm or BM25 (Best Matching, 25 is the number of iterations) algorithm, etc.

[0124] In the above embodiments, based on the idea of text recall, the first correlation degree between the text label data of each candidate expression picture and the request text feature data is determined, so as to quickly and accurately screen out a part of candidate expression pictures that match the request text feature data.

[0125] S240: Determine the second correlation degree between the semantic vector data of each candidate expression picture and the request semantic feature data.

[0126] In the embodiments of the present application, from the perspective of semantic vectors, by examining the second correlation degree between the semantic vector data of each candidate expression picture among multiple candidate expression pictures and the request semantic feature data, a part of candidate expression pictures can be recalled. The second correlation degree between the semantic vector data of each of these candidate expression pictures and the semantic vector data meets the preset second matching condition. Among them, the semantic vector data is content feature data described by semantic vectors obtained by performing multi-dimensional analysis on the corresponding candidate expression pictures, and the second matching condition can indicate the threshold that the second correlation degree corresponding to the recalled candidate expression pictures needs to meet.

[0127] Specifically, step S240 can be implemented as:

[0128] S231: Obtain the semantic vector data of each candidate expression picture; the semantic vector data of each candidate expression picture is determined based on a contrastive learning model, each candidate expression picture, and the text description information of each candidate expression picture; the text description information of each candidate expression picture includes the text label data of each candidate expression picture and the original text data shown on each candidate expression picture.

[0129] Among them, the contrastive learning model can be an aligned CLIP (Contrastive Language-Image Pre-Training) model. Inputting each candidate expression picture and its text description information into the aligned CLIP model can obtain the corresponding semantic vector data. Among them, the text description information can include the text label data of the corresponding candidate expression picture and the original text data shown on the corresponding candidate expression picture, that is, inputting a set of matching candidate expression pictures and text description information into the CLIP model, and encoding them respectively to obtain image semantic vector data and text semantic vector data. The image semantic vector data and the text semantic vector data constitute the semantic vector data. The text description information and semantic vector data corresponding to the candidate expression pictures generated by the contrastive learning model can be stored in the expression picture content vector system to serve the subsequent recall processing based on semantic vectors.

[0130] S232: Based on the vector space distance algorithm, determine the second correlation degree between the semantic vector data of each candidate expression picture and the requested semantic feature data; the second correlation degree indicates the semantic vector similarity between each candidate expression picture and the query request information.

[0131] The vector space distance algorithm is used to measure the distance between two vectors. Common vector space distances can include Euclidean distance, cosine distance, Manhattan distance, etc.

[0132] It can be understood that the query request information for expression pictures usually contains many phrases related to daily emotions and actions. If only text character-level recall is performed during recall, it is very easy to result in insufficient recall results or low matching degrees. In the embodiments of this application, through the recall method based on semantic vectors, the recall results can be enriched and the matching degree between the recall results and the query request information can be improved. Exemplarily, there may be few candidate expression pictures matched by the text "beaming with joy", but through the similarity in the dimension of semantic vectors, related texts such as "smiling from ear to ear" and "delighted" can be associated, and then more matching candidate expression pictures can be recalled according to the related texts.

[0133] Feasibly, based on vector index libraries such as Faiss (Facebook AI Similarity Search, an approximate nearest neighbor search library) and ElasticSearch (a search data analysis engine), the nearest neighbor search of the requested semantic feature data and the semantic vector data can be realized, so as to achieve the millisecond-level recall effect under the scale of tens of millions.

[0134] In the above embodiments, based on the idea of vector recall, determine the second correlation degree between the semantic vector data of each candidate expression picture and the requested semantic feature data, so as to quickly and accurately screen out a part of the candidate expression pictures that match the requested semantic feature data, enrich the recall results, and improve the recall matching degree.

[0135] In an application scenario provided by an embodiment of the present application, before providing a distribution service for expression pictures, multiple candidate expressions can be processed through picture understanding, text understanding, embedding representation, etc., to establish a unified portrait database of expression pictures. The portrait database stores text label data and semantic vector data for each candidate expression picture. These information can provide basic feature inputs for subsequent distribution services of expression pictures, and ultimately improve the accuracy and conversion click-through rate of the distribution process. Among them, the text label data can indicate content features including but not limited to entities, copywriting, actions, emotions, categories, styles, etc. in the candidate expression pictures. The generation of labels generally combines algorithms and manual annotation, and the annotation based on algorithms is mainly implemented from multiple dimensions such as expression picture understanding and expression text understanding. The semantic vector data can be obtained by performing semantic vector representation on the text label data and the original text data obtained in the previous embodiment, or can be obtained by independently performing multi-dimensional analysis such as expression picture understanding and expression text understanding on the candidate expression pictures.

[0136] In an embodiment of the present application, specifically, before providing a distribution service for expression pictures, the method may further include:

[0137] S310: Perform extraction processing on text elements of each candidate expression picture to obtain the original text data displayed on each candidate expression picture.

[0138] Feasibly, text elements can be extracted and recognized based on OCR to obtain the original text data displayed on the candidate expression pictures. Generally, the original text data displayed on the candidate expression pictures is an oral phrase or short sentence. Introducing oral sentences emphatically in the process of multi-dimensional analysis of candidate expression pictures can improve the matching accuracy for expression pictures.

[0139] S320: Perform text analysis processing on the original text data displayed on each candidate expression picture to obtain the text analysis result of each candidate expression picture. The text analysis result includes semantic analysis result, entity recognition result, and sentiment analysis result.

[0140] In the process of text analysis processing, the text included in the candidate expression pictures can be processed through word segmentation and part-of-speech tagging, syntactic analysis, semantic analysis, entity analysis, relation extraction, sentiment analysis, etc. based on NLP (Natural Language Processing) technology to obtain text analysis results including but not limited to semantic analysis results, entity recognition results, and sentiment analysis results.

[0141] S330: Perform label extraction processing on the text analysis result of each candidate expression picture to obtain the first label data of each candidate expression picture.

[0142] Based on the text analysis results of each candidate expression picture, extract the results that can be used as label data, and the obtained first label data can indicate the content characteristics of the corresponding candidate expression picture from the perspectives of entities, actions, emotions, people, etc. included.

[0143] S340: Extract the image features of each candidate expression picture to obtain the first image feature data of each candidate expression picture.

[0144] S350: Perform image classification processing based on the first image feature data of each candidate expression picture to obtain the second label data of each candidate expression picture.

[0145] Exemplarily, a multi-level category label specifically constructed for expression pictures and a ViT (Vision Transformer) model can be used to construct an expression picture classification model. Furthermore, classification can be performed based on the basic image features extracted by the ViT model to obtain the second label data.

[0146] In the embodiments of the present application, the text label data of each candidate expression picture includes the first label data of each candidate expression picture and the second label data of each candidate expression picture.

[0147] Furthermore, a label index table can be constructed based on the text label data of each candidate expression picture. The label index table represents the mapping relationship between the identification data of the candidate expression picture and the label data of the candidate expression picture, and the label index table is stored in an inverted index system for subsequent text-based recall processing.

[0148] In the above embodiments, the content features of each candidate expression picture are understood, analyzed, and described in multiple dimensions, and the obtained text label data can effectively improve the accuracy and recall efficiency of recalling candidate expression pictures based on text.

[0149] In an embodiment of the present application, specifically, before providing the distribution service of expression pictures, the method may further include:

[0150] S410: Determine the text description information of each candidate expression picture, where the text description information includes the original text data of each candidate expression picture and the text label data of each candidate expression picture.

[0151] In an embodiment of the present application, semantic vector data can be obtained by performing semantic vector representation on the text label data and the original text data obtained in the previous embodiment. This method can effectively improve the processing efficiency of multi-dimensional analysis of candidate expression pictures, and at the same time ensure the consistency of the content features indicated by the text label data and the semantic vector data of the same candidate expression picture.

[0152] In another embodiment of the present application, multi-dimensional analysis such as expression picture understanding and expression text understanding can also be independently performed on the candidate expression pictures to obtain semantic vector data.

[0153] S420: Input the text description information of each candidate expression picture into the contrastive learning model, perform text feature extraction processing, and obtain the text feature data of each candidate expression picture.

[0154] S430: Map the text feature data of each candidate expression picture to the corresponding semantic vector space of the contrastive learning model to obtain the text semantic vector data of each candidate expression picture.

[0155] In one embodiment of the present application, the contrastive learning model can adopt the CLIP (Contrastive Language-Image Pre-Training) model. The CLIP model includes a text encoder and an image encoder. The text description information can be input into the text encoder to perform extraction and encoding processing of text features, so as to map the text feature data into the semantic vector space (also called the hidden space of the CLIP model) to obtain the corresponding text semantic vector data.

[0156] S440: Input each candidate expression picture into the contrastive learning model, perform image feature extraction processing, and obtain the second image feature data of each candidate expression picture.

[0157] S450: Map the second image feature data of each candidate expression picture to the semantic vector space to obtain the image semantic vector data of each candidate expression picture.

[0158] Specifically, the candidate expression picture can be input into the image encoder to perform extraction and encoding processing of image features, so as to map the second image feature data into the same semantic vector space (also called the hidden space of the CLIP model) to obtain the corresponding image semantic vector data.

[0159] It can be understood that in the multi-dimensional analysis of candidate expression pictures, the CLIP model plays the role of semantic vector representation. Since a set of input text-image data is the candidate expression picture and the text description information of the candidate expression picture, that is, it is default that a set of input text-image data is matched, there is no need to calculate the similarity between the text semantic vector data and the image semantic vector data of the same candidate expression picture. In another feasible implementation, the similarity between the text semantic vector data and the image semantic vector data of the same candidate expression picture can also be calculated. Furthermore, candidate expression pictures with unmatched text and images can be screened out according to the similarity, and the candidate expression pictures with unmatched text and images can be screened out or the text description information can be determined again.

[0160] In an embodiment of the present application, the semantic vector data of each candidate expression picture includes the text semantic vector data of each candidate expression picture and the image semantic vector data of each candidate expression picture.

[0161] It can be understood that the CLIP model itself actually learns the spatial mapping relationship between text and matching pictures, and the similarity between the semantic vector data corresponding to the matching text and pictures meets the preset matching conditions. In addition to representing the semantic vector of the candidate expression picture, the CLIP model can also be applied to the recall process based on the semantic vector. That is, in step S220 and step S240, the request text feature information corresponding to the request feature information can be input into the CLIP model to obtain the corresponding request semantic feature data. Furthermore, the similarity between the semantic vector data of each candidate expression picture and the request semantic feature data is calculated, and proximity search is performed according to the similarity, so as to recall some target candidate expression pictures.

[0162] In an embodiment of the present application, the used CLIP model can be obtained by further training and alignment based on the general CLIP model. In the process of further training and alignment, multiple groups of matching text-image data can be constructed in the following ways: One is to collect a validation data set from the expression picture database. The original text information in each candidate expression picture in the validation data set can be extracted and recognized through OCR, and the text description information can be formed by combining the label data of each candidate expression picture, so as to construct a data set of candidate expression pictures and text description information; The other is to construct a data set between the historical query request information and the historical distributed expression pictures through the historical distribution records of the expression pictures. By performing the contrastive learning method on the above data, the CLIP model can be further upgraded to be more suitable for the semantic vector representation of expression pictures or the matching of expression pictures.

[0163] In the above embodiments, by using the text description information of the candidate expression pictures and the corresponding generated text semantic vector data and image semantic vector data for the candidate expression pictures, the processing efficiency of multi-dimensional analysis of the candidate expression pictures can be effectively improved, and at the same time, the consistency of the content features indicated by the text label data and the semantic vector data of the same candidate expression picture can be ensured; at the same time, the use of a symmetric learning model can effectively improve the efficiency of semantic vector representation processing; the obtained semantic vector data of the candidate expression pictures can be stored in the expression picture vector system to serve subsequent recall processing based on semantic vectors, thereby improving the recall accuracy and richness of the target candidate expression pictures from the semantic dimension and reducing semantic ambiguity and misunderstanding.

[0164] S250: Determine at least one target candidate expression picture from multiple candidate expression pictures according to the first correlation degree and the second correlation degree.

[0165] In the embodiments of the present application, multi-channel recall according to the first correlation degree and the second correlation degree can screen out a rich and matching number of target candidate expression pictures. The first correlation degree between the text label data of the target candidate expression picture and the request text feature data meets a preset first matching condition, and the second correlation degree between the semantic vector data of the target candidate expression picture and the semantic vector data meets a preset second matching condition, where the first matching condition or the second matching condition can indicate a threshold of the correlation degree or a ranking threshold sorted according to the correlation degree.

[0166] In an embodiment of the present application, the method may further include:

[0167] S251: Determine the request object information associated with the client in the query request information.

[0168] The request object information may include, but is not limited to, the account information corresponding to the request account, the preference information corresponding to the request account, the type information of the client, etc.

[0169] S252: Determine the cooperation relationship between the historical request object in the historical distribution information and multiple candidate expression pictures. The cooperation relationship is represented by a graph model.

[0170] In the cooperation relationship represented by the graph model, any historical request object and any candidate expression picture correspond to a node in the graph model. According to the historical interaction information, the jump relationship between the nodes can be determined, and then a directed graph model is generated. The cooperation relationship can also indicate the social propagation situation of the candidate expression pictures.

[0171] S253: Based on the request object information and the cooperation relationship, determine at least one first candidate expression picture from multiple candidate expression pictures.

[0172] Feasibly, based on the DeepWalk algorithm (random walk algorithm) or the GraphSAGE algorithm (Graph Sampling and Aggregation, a node embedding learning algorithm for graph neural networks), a feature vector representation of each node is generated. Then, an approximate node can be matched according to the feature vector representation corresponding to the request object information, and a first candidate expression picture can be determined according to the directed path involved in the approximate node.

[0173] S254: Update at least one target candidate expression picture according to at least one first candidate expression picture.

[0174] In the above embodiments, according to the collaborative relationship between the historical request object reflected by the historical distribution information and multiple candidate expression pictures, the recall result can also be enriched by the recall method using the graph model.

[0175] In the above embodiments, through the multi-channel recall method, the generalization ability of recall can be improved, the recall result can be greatly enriched, and the problem of insufficient recall result quantity caused by the colloquial words in the query request information can be avoided. In addition, targeted filtering may be required for different business scenarios, such as filtering according to the source, type, size, etc. of the candidate expression pictures. The specific filtering conditions can be set according to the strategy of the business scenario and are not specifically limited here.

[0176] In another embodiment of the present application, as Figure 3 shown, the method may further include:

[0177] S510: Determine the key text information of the query request information when the number of at least one target candidate expression picture is lower than the preset quantity threshold.

[0178] Feasibly, the query request information can be segmented to obtain multiple segments, and according to the word frequency or inverse document frequency corresponding to each segment in the inverted index system, the word weight of each segment is determined. Then, important segments are screened out according to the word weight as the key text information.

[0179] Feasibly, the important segments can also be expanded into sentences with complete meaning expressions, and the sentence is used as the key text information.

[0180] S520: Obtain multiple basic expression pictures corresponding to multiple candidate expression pictures; none of the multiple basic expression pictures contains text.

[0181] In the embodiments of the present application, the text of multiple pre-stored candidate expression pictures can also be erased to obtain the corresponding multiple basic expression pictures.

[0182] S530: Obtain the index vector data of each basic expression picture among multiple basic expression pictures, where the index vector data is text label data and semantic vector data in vector form.

[0183] S540: Determine at least one target basic expression picture from multiple basic expression pictures according to the similarity between the index vector data of each basic expression picture and the requested semantic feature data.

[0184] In the embodiments of the present application, both multiple basic expression pictures and multiple candidate expression pictures have corresponding text label data and semantic vector data. During the synthesis process, in order to balance the synthesis efficiency and the graphic-text matching degree, calculate the vector similarity between the requested semantic feature data and the index vector data of each basic expression picture, and then determine at least one target basic expression picture from multiple basic expression pictures according to the vector similarity.

[0185] S550: Synthesize the key text information and at least one target basic expression picture to obtain at least one second candidate expression picture.

[0186] Feasibly, the key text information and at least one target basic expression picture can be synthesized based on a synthesis strategy or rule. Exemplarily, the synthesis strategy or rule can be: (a) When the number of characters in the key text information does not exceed 6 text characters, display the key text information in one line. When the number of characters in the key text information exceeds 7 text characters, display the key text information in line breaks, and the displayed font is regular script; (b) Font position: centered at the bottom; (c) Font color: When the picture has a dark background, the font color is white. When the picture has a light background, the font color is white.

[0187] In addition, on the server side, the synthesis relationship between the key text information and at least one target basic expression picture can be constructed first, and then synthesized in real time when the client interacts with a certain second candidate expression picture, which can ensure the user experience while improving the efficiency of graphic-text synthesis processing.

[0188] S315: Update at least one target candidate expression picture according to at least one second candidate expression picture.

[0189] In the above embodiments, when the number of at least one target candidate expression pictures recalled is lower than the preset number threshold, through the synthesis process of the key text information and the target basic expression picture, the recall result can be enriched, providing more choices for users and improving the user experience.

[0190] S260: Obtain a picture sequence to be distributed that matches the query request information according to at least one target candidate expression picture, where the picture sequence to be distributed is part or all of the at least one target candidate expression picture.

[0191] In the embodiments of the present application, the sorting process performed on at least one target candidate expression picture recalled through multiple channels can effectively improve the interaction potential of the entire picture sequence to be distributed, that is, it helps to improve the conversion rate of the entire picture sequence to be distributed and promote the interaction activity.

[0192] Specifically, the sorting process may include processes such as rough sorting, fine sorting, and re - sorting. Among them, rough sorting and fine sorting can be implemented based on an interaction index prediction model. Through the interaction index values corresponding to each target candidate expression picture predicted by the interaction index prediction model, such as conversion rate, click - through rate, etc., at least one target candidate expression picture is sorted. There are differences in the model complexity, the number of interaction indexes, etc. between the interaction index prediction model used for rough sorting and the interaction index prediction model used for fine sorting. Generally, rough sorting is a rough screening of at least one target candidate expression picture, and fine sorting will reduce the number of at least one target candidate expression picture to a preset number of pictures. And re - sorting is adjusted according to the personalized needs of users and the business operation needs.

[0193] In an embodiment of the present application, step S260 may be implemented as:

[0194] S261: Sort at least one target candidate expression picture according to the first correlation degree corresponding to each target candidate expression picture and the second correlation degree corresponding to each target candidate expression picture in at least one target candidate expression picture, and obtain a first picture sequence.

[0195] Specifically, the first correlation degree or the second correlation degree corresponding to each target candidate expression picture can be normalized to the same interval, and the normalized first correlation degree or second correlation degree is used as a sorting factor for the corresponding target candidate expression picture in the sorting process to perform a preliminary sorting process on at least one target candidate expression picture, and obtain a first picture sequence.

[0196] S262: Determine the request object information associated with the client in the query request information.

[0197] The request object information may include, but is not limited to, the account information corresponding to the request account, the preference information corresponding to the request account, the type information of the client, etc.

[0198] S263: Determine the background information corresponding to the query request information, where the background information includes interaction context information, business type information, and picture quality requirement information.

[0199] The interaction context information may indicate the interaction context information of the user when the query request information is sent; the service type information may indicate the specific application scenario of the emoji picture distribution service and the operation metrics to be considered in this application scenario, etc.; the picture quality requirement information may indicate the pre-set picture resolution conditions, picture sizes, etc.

[0200] S264: Screen the first picture sequence according to the picture quality requirement information to obtain a second picture sequence.

[0201] In the case where the user has pre-set the requirements for the quality of emoji pictures, determine whether each target candidate emoji picture in the first picture sequence meets the user's requirements for the quality of emoji pictures, and delete the target candidate emoji pictures that do not meet the requirements from the first picture sequence to obtain a second picture sequence.

[0202] S265: Determine the first matching degree between each target candidate emoji picture in the second picture sequence and the request object information.

[0203] S266: Determine the second matching degree between each target candidate emoji picture in the second picture sequence and the interaction context information.

[0204] S267: Determine the third matching degree between each target candidate emoji picture in the second picture sequence and the service type information.

[0205] The first matching degree, the second matching degree, and the third matching degree respectively represent the matching fitness degrees between the corresponding target candidate emoji pictures and the request object, the interaction context, and the service type. The matching fitness degrees with the request object, the interaction context, and the service type can be used as sorting factors for personalized re-ranking to better fit the user's preferences, adapt to the interaction context and service type, etc.

[0206] S268: Re-rank the second picture sequence according to the first matching degree, the second matching degree, and the third matching degree to obtain a picture sequence to be distributed.

[0207] In the above embodiments, by combining the request object information and the background information, screening and re-ranking at least one target candidate emoji picture can make the obtained picture sequence to be distributed better fit the user's preferences, adapt to the interaction context and service type, improve the user experience, and also meet the operation requirements.

[0208] Figure 4 It is a schematic diagram of two application scenarios of the emoji picture distribution service provided by the embodiments of the present application. From left to right, they are emoji search and emoji recommendation. In the application scenario of emoji search, the user can enter text information in the text box, and the server can match a picture sequence to be distributed based on the query request information containing the text information and display it on the interface of the client, such asFigure 4 as shown in Emoji Picture 1, Emoji Picture 2, etc. in []. In the application scenario of emoji recommendation, in response to a click instruction executed by the user on the plus (+) control, the client generates query request information and sends it to the server. The server matches a suitable sequence of pictures to be distributed based on a series of emoji pictures collected by the user indicated by the query request information, such as Emoji Picture 11, Emoji Picture 12, etc., and displays them on the interface of the client for the user to select. In addition, in the application scenario of emoji association, query request information is generated based on the text input by the user in the chat input box. The server matches a sequence of pictures to be distributed in response to the query request information, and displays the target candidate emoji pictures in the sequence of pictures to be distributed in the input display window for the user to select.

[0209] Figure 5 is a schematic diagram of personalized settings for an emoji search application provided by an embodiment of the present application. As Figure 5 shown, when searching for emoji pictures on the client, in addition to being able to input text, such as "so envious", the number of target candidate emoji pictures in the sequence of pictures to be distributed displayed can be set. Further, the size and resolution of the target candidate emoji pictures can also be set. A certain number or a certain size and resolution of emoji pictures matching the input text will be displayed in the emoji search application interface. For example Figure 5 the envious Emoji Picture 1, Envious Emoji Picture 2, etc. in [].

[0210] Figure 6 is a schematic diagram of the architecture of an emoji picture distribution system provided by an embodiment of the present application. As Figure 6 shown, before providing the distribution service of emoji pictures, multiple candidate emojis in the emoji central library data can be processed, such as picture understanding, text understanding, background image erasure, and embedding characterization (i.e., Emb understanding in the figure), etc., to establish a unified emoji portrait database. The emoji understanding data of each candidate emoji picture, including emoji classification, semantic tags, vectors, etc., is stored in the emoji portrait database. In addition, the emoji understanding data of the basic emoji pictures corresponding to the candidate emoji pictures is also included in the emoji portrait database. Extract the text label data of each candidate emoji picture from the emoji understanding data in the emoji portrait database, and generate a label index table to serve the inverted index system; extract the semantic vector data of each candidate emoji picture from the emoji understanding data in the emoji portrait database to serve the emoji content vector system.

[0211] In the application stage, query request information (Query) is generated according to the input of the end user. The input can be text or an emoticon picture. After performing Query understanding processing on the query request information, request text feature data and request semantic feature data are obtained. The inverted index system performs text recall based on the request text feature data and the text label data of each candidate emoticon picture. The emoticon content vector system performs vector recall based on the request semantic feature data and the semantic vector data of each candidate emoticon picture. In addition, recall based on a graph model can also be performed to obtain at least one target candidate emoticon picture. The emoticon sorting system performs rough sorting and fine sorting on at least one target candidate emoticon picture. Further, re-sorting can also be performed according to the personalized preferences of the end user and the operation requirements to obtain a to-be-distributed sequence, and the target candidate emoticon pictures in the to-be-distributed sequence are displayed on the terminal in the sequence order for the user to select. In the case where the number of recall results is insufficient, an emoticon picture and text synthesis service can also be used to synthesize the key text information obtained by Query understanding with at least one target basic emoticon picture that is matched, and return it to the emoticon sorting system.

[0212] As can be seen from the above embodiments, in an emoticon picture matching method provided in the present application, based on the query request information for emoticon pictures sent by the client, feature extraction processing is performed to obtain request feature information including request text feature data and request semantic feature data. Feature extraction and characterization of the query request information from the text dimension and the semantic vector dimension can more comprehensively and accurately mine the demand characteristics of the user for emoticon pictures. In the solution provided in the present application, the content of multiple candidate emoticon pictures is also more carefully understood and analyzed from multiple dimensions, and the text label data and semantic vector data of each candidate emoticon picture are generated. Furthermore, according to the first correlation degree between the text label data of each candidate emoticon picture and the request text feature data, and the second correlation degree between the semantic vector data of each candidate emoticon picture and the request semantic feature data, at least one target candidate emoticon picture that matches the query request information can be more efficiently, accurately, and diversely screened out. The to-be-distributed sequence obtained according to at least one target candidate emoticon picture can meet the needs of the emoticon picture distribution service. In the solution provided in the present application, the diversity, coverage, and matching degree with the user's needs of the recall results can be greatly improved, the efficiency and accuracy of emoticon picture distribution can be effectively improved, and the user's usage experience can be optimized. When applied to an interaction scenario, an efficient, accurate, and rich emoticon picture distribution solution can also promote the interaction between users, enhance the fun and diversity of the interaction, and improve the activity of the interaction platform.

[0213] The embodiment of the present application also provides an emoticon picture matching device 700, as Figure 7 shown, the device may include:

[0214] An acquisition module 710, configured to acquire query request information for an expression picture sent by a client;

[0215] A feature extraction module 720, configured to perform feature extraction processing based on the query request information to obtain request feature information, where the request feature information includes request text feature data and request semantic feature data;

[0216] A text correlation determination module 730, configured to determine a first correlation between text label data of each candidate expression picture among a plurality of candidate expression pictures and the request text feature data; the text label data is content feature data described in text obtained by performing multi-dimensional analysis on the corresponding candidate expression picture;

[0217] A vector correlation determination module 740, configured to determine a second correlation between semantic vector data of each candidate expression picture and the request semantic feature data; the semantic vector data is content feature data described in a semantic vector obtained by performing multi-dimensional analysis on the corresponding candidate expression picture;

[0218] A recall module 750, configured to determine at least one target candidate expression picture from the plurality of candidate expression pictures according to the first correlation and the second correlation;

[0219] A sequence generation module 760, configured to obtain a picture sequence to be distributed that matches the query request information according to the at least one target candidate expression picture, where the picture sequence to be distributed is part or all of the at least one target candidate expression picture arranged in a preset order.

[0220] In an embodiment of the present application, the apparatus 700 may further include:

[0221] A text element extraction unit, configured to perform text element extraction processing on each candidate expression picture to obtain original text data displayed on each candidate expression picture;

[0222] A text analysis unit, configured to perform text analysis processing on the original text data displayed on each candidate expression picture to obtain a text analysis result of each candidate expression picture, where the text analysis result includes a semantic analysis result, an entity recognition result, and an emotion analysis result;

[0223] A label extraction unit, configured to perform label extraction processing on the text analysis result of each candidate expression picture to obtain first label data of each candidate expression picture;

[0224] The first image feature extraction unit is configured to perform image feature extraction processing on each of the candidate expression pictures to obtain first image feature data of each of the candidate expression pictures;

[0225] The image classification unit is configured to perform image classification processing according to the first image feature data of each of the candidate expression pictures to obtain second label data of each of the candidate expression pictures;

[0226] The text label data of each of the candidate expression pictures includes the first label data of each of the candidate expression pictures and the second label data of each of the candidate expression pictures.

[0227] In an embodiment of the present application, the apparatus 700 may further include:

[0228] The text description information determination unit is configured to determine the text description information of each of the candidate expression pictures, where the text description information includes the original text data of each of the candidate expression pictures and the text label data of each of the candidate expression pictures;

[0229] The text feature extraction unit is configured to input the text description information of each of the candidate expression pictures into a contrastive learning model to perform text feature extraction processing to obtain text feature data of each of the candidate expression pictures;

[0230] The text semantic vector mapping unit is configured to map the text feature data of each of the candidate expression pictures to the semantic vector space corresponding to the contrastive learning model to obtain text semantic vector data of each of the candidate expression pictures;

[0231] The second image feature extraction unit is configured to input each of the candidate expression pictures into the contrastive learning model to perform image feature extraction processing to obtain second image feature data of each of the candidate expression pictures;

[0232] The image semantic vector mapping unit is configured to map the second image feature data of each of the candidate expression pictures to the semantic vector space to obtain image semantic vector data of each of the candidate expression pictures;

[0233] The semantic vector data of each of the candidate expression pictures includes the text semantic vector data of each of the candidate expression pictures and the image semantic vector data of each of the candidate expression pictures.

[0234] In an embodiment of the present application, the text correlation determination module 730 may include:

[0235] A label index table acquisition unit for acquiring a label index table of the multiple candidate expression pictures, where the label index table contains text label data of each candidate expression picture;

[0236] A first correlation degree determination unit for determining a first correlation degree between the text label data of each candidate expression picture and the request text feature data based on a text matching algorithm; the first correlation degree indicates the text matching degree between each candidate expression picture and the query request information.

[0237] In an embodiment of the present application, the vector correlation degree determination module 740 may include:

[0238] A semantic vector data acquisition unit for acquiring semantic vector data of each candidate expression picture; the semantic vector data of each candidate expression picture is determined based on a contrastive learning model, each candidate expression picture, and the text description information of each candidate expression picture; the text description information of each candidate expression picture includes the text label data of each candidate expression picture and the original text data displayed on each candidate expression picture;

[0239] A second correlation degree determination unit for determining a second correlation degree between the semantic vector data of each candidate expression picture and the request semantic feature data based on a vector space distance algorithm; the second correlation degree indicates the semantic vector similarity between each candidate expression picture and the query request information.

[0240] In an embodiment of the present application, the device 700 may further include:

[0241] A first request object information determination unit for determining request object information associated with the client in the query request information;

[0242] A collaboration relationship determination unit for determining a collaboration relationship between the historical request object in the historical distribution information and the multiple candidate expression pictures, where the collaboration relationship is represented by a graph model;

[0243] A graph recall unit for determining at least one first candidate expression picture from the multiple candidate expression pictures based on the request object information and the collaboration relationship;

[0244] A first update unit for updating the at least one target candidate expression picture according to the at least one first candidate expression picture.

[0245] In an embodiment of the present application, the sequence generation module 760 may include:

[0246] The first sorting unit is configured to sort the at least one target candidate expression picture according to the first correlation degree corresponding to each target candidate expression picture and the second correlation degree corresponding to each target candidate expression picture among the at least one target candidate expression pictures, so as to obtain a first picture sequence.

[0247] The request object information second determination unit is configured to determine the request object information associated with the client in the query request information.

[0248] The background information determination unit is configured to determine the background information corresponding to the query request information, where the background information includes interaction context information, service type information, and picture quality requirement information.

[0249] The screening unit is configured to screen the first picture sequence according to the picture quality requirement information to obtain a second picture sequence.

[0250] The first matching degree determination unit is configured to determine the first matching degree between each target candidate expression picture in the second picture sequence and the request object information.

[0251] The second matching degree determination unit is configured to determine the second matching degree between each target candidate expression picture in the second picture sequence and the interaction context information.

[0252] The third matching degree determination unit is configured to determine the third matching degree between each target candidate expression picture in the second picture sequence and the service type information.

[0253] The second sorting unit is configured to re-sort the second picture sequence according to the first matching degree, the second matching degree, and the third matching degree to obtain the picture sequence to be distributed.

[0254] In an embodiment of the present application, the apparatus 700 may further include:

[0255] The key text extraction unit is configured to determine the key text information of the query request information when the number of the at least one target candidate expression picture is lower than a preset number threshold.

[0256] The basic expression picture acquisition unit is configured to acquire a plurality of basic expression pictures corresponding to the plurality of candidate expression pictures; none of the plurality of basic expression pictures contains text.

[0257] The index vector data acquisition unit is configured to acquire the index vector data of each basic expression picture among the plurality of basic expression pictures, where the index vector data is vector-form text label data and semantic vector data.

[0258] A similarity recall unit, configured to determine at least one target basic expression picture from the multiple basic expression pictures according to the similarity between the index vector data of each basic expression picture and the requested semantic feature data;

[0259] A synthesis unit, configured to synthesize the key text information and the at least one target basic expression picture to obtain at least one second candidate expression picture;

[0260] A second update unit, configured to update the at least one target candidate expression picture according to the at least one second candidate expression picture.

[0261] It should be noted that, for the device provided in the above embodiments, when implementing its functions, only the division of the above function modules is used for illustration. In actual applications, the above functions can be allocated to different function modules according to needs, that is, the internal structure of the device is divided into different function modules to complete all or part of the functions described above. In addition, the device provided in the above embodiments and the method embodiments belong to the same concept, and the specific implementation process is detailed in the method embodiments and will not be repeated here.

[0262] An embodiment of the present application provides a computer device, which includes a processor and a memory. At least one instruction or at least one program segment is stored in the memory, and the at least one instruction or the at least one program segment is loaded and executed by the processor to implement an expression picture matching method as provided in the above method embodiment.

[0263] Figure 8 The figure shows a schematic hardware structure diagram of a device for implementing an expression picture matching method provided in an embodiment of the present application. The device may participate in forming or include the device or system provided in the embodiment of the present application. As Figure 8 shown, the device 10 may include one or more (shown as 1002a, 1002b,..., 1002n in the figure) processors 1002 (the processor 1002 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 1004 for storing data, and a transmission device 1006 for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 8 the structure shown is only schematic and does not limit the structure of the above electronic device. For example, the device 10 may further include more or fewer components than Figure 8 shown in the figure, or have a different configuration from Figure 8 shown in the figure.

[0264] It should be noted that one or more of the above-mentioned processors 1002 and / or other data processing circuits can generally be referred to as "data processing circuits" herein. The data processing circuit can be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit can be a single independent processing module, or be incorporated in whole or in part into any one of other elements in the device 10 (or mobile device). As involved in the embodiments of the present application, the data processing circuit is a kind of processor control (such as the selection of a variable resistance terminal path connected to an interface).

[0265] The memory 1004 can be used to store software programs and modules of application software, such as the program instructions / data storage devices corresponding to the methods described in the embodiments of the present application. The processor 1002 executes various functional applications and data processing by running the software programs and modules stored in the memory 1004, that is, to implement the above-mentioned method for matching expression pictures. The memory 1004 can include high-speed random access memory, and can also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 1004 can further include a memory remotely set relative to the processor 1002, and these remote memories can be connected to the device 10 through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0266] The transmission device 1006 is used to receive or send data via a network. Specific examples of the above-mentioned network can include the wireless network provided by the communication provider of the device 10. In one instance, the transmission device 1006 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 1006 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0267] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables the user to interact with the user interface of the device 10 (or mobile device).

[0268] The embodiments of the present application also provide a computer-readable storage medium, which can be set in a server to store at least one instruction or at least one segment of program related to implementing a method for matching expression pictures in the method embodiments. The at least one instruction or the at least one segment of program is loaded and executed by the processor to implement the method for matching expression pictures provided by the above-mentioned method embodiments.

[0269] Optionally, in this embodiment, the above storage medium may be located in at least one of multiple network servers in a computer network. Optionally, in this embodiment, the above storage medium may include, but is not limited to: various media that can store program codes such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs.

[0270] The embodiment of the present invention also provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and these computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads these computer instructions from the computer-readable storage medium, and the processor executes these computer instructions, so that the computer device executes an expression picture matching method provided in the above various optional embodiments.

[0271] It should be noted that: the above sequence of the embodiments of the present application is only for description and does not represent the superiority or inferiority of the embodiments. And the above specific embodiments of the present application have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be executed in a different order from that in the embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0272] Each embodiment in this application is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the device, equipment, and storage medium, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.

[0273] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a magnetic disk, or an optical disc, etc.

[0274] The above are only the preferred embodiments of the present application and are not used to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A facial expression picture matching method, characterized in that: Applied to a server, the method comprises: Get the query request information for emoticon images sent by the client; Based on the query request information, feature extraction processing is performed to obtain request feature information, wherein the request feature information includes request text feature data and request semantic feature data; Determine a first degree of association between text label data of each candidate emoticon picture among a plurality of candidate emoticon pictures and the requested text feature data; the text label data is content feature data described in text and obtained by performing multi-dimensional analysis on the corresponding candidate emoticon pictures; Determine a second degree of association between the semantic vector data of each candidate expression picture and the requested semantic feature data; the semantic vector data is content feature data described by a semantic vector obtained by performing a multi-dimensional analysis on the corresponding candidate expression picture; Determining at least one target candidate expression picture from the plurality of candidate expression pictures according to the first association degree and the second association degree; According to the at least one target candidate expression picture, a sequence of pictures to be distributed that matches the query request information is obtained, and the sequence of pictures to be distributed is part or all of the at least one target candidate expression picture.

2. The method according to claim 1, characterized in that The method further comprises: Extracting text elements from each candidate expression picture to obtain original text data displayed on each candidate expression picture; Performing text analysis on the original text data displayed on each candidate expression picture to obtain a text analysis result of each candidate expression picture, wherein the text analysis result includes a semantic analysis result, an entity recognition result, and a sentiment analysis result; Performing label extraction processing on the text analysis results of each candidate expression picture to obtain first label data of each candidate expression picture; Performing image feature extraction processing on each candidate expression picture to obtain first image feature data of each candidate expression picture; Performing image classification processing according to the first image feature data of each candidate expression picture to obtain second label data of each candidate expression picture; The text label data of each candidate expression picture includes first label data of each candidate expression picture and second label data of each candidate expression picture.

3. The method according to claim 2, characterized in that The method further comprises: Determine text description information of each candidate emoticon picture, wherein the text description information includes original text data of each candidate emoticon picture and text label data of each candidate emoticon picture; Inputting the text description information of each candidate expression picture into the contrastive learning model, performing text feature extraction processing, and obtaining text feature data of each candidate expression picture; Mapping the text feature data of each candidate expression picture to the semantic vector space corresponding to the contrastive learning model to obtain the text semantic vector data of each candidate expression picture; Input each candidate expression picture into the contrastive learning model, perform image feature extraction processing, and obtain second image feature data of each candidate expression picture; Mapping the second image feature data of each candidate expression picture to the semantic vector space to obtain image semantic vector data of each candidate expression picture; The semantic vector data of each candidate expression picture includes text semantic vector data of each candidate expression picture and image semantic vector data of each candidate expression picture.

4. The method according to claim 1, characterized in that: The determining of a first degree of association between the text label data of each candidate emoticon picture among a plurality of candidate emoticon pictures and the requested text feature data comprises: Obtain a label index table of the plurality of candidate emoticon images, wherein the label index table includes text label data of each candidate emoticon image; Based on a text matching algorithm, a first degree of association between the text label data of each candidate emoticon image and the request text feature data is determined; the first degree of association indicates a text matching degree between each candidate emoticon image and the query request information.

5. The method according to claim 1, characterized in that: The determining of the second degree of association between the semantic vector data of each candidate expression picture and the requested semantic feature data comprises: Acquire semantic vector data of each candidate expression picture; the semantic vector data of each candidate expression picture is determined based on a contrastive learning model, each candidate expression picture and text description information of each candidate expression picture; the text description information of each candidate expression picture includes text label data of each candidate expression picture and original text data displayed on each candidate expression picture; Based on a vector space distance algorithm, a second degree of association between the semantic vector data of each candidate expression picture and the requested semantic feature data is determined; the second degree of association indicates the semantic vector similarity between each candidate expression picture and the query request information.

6. The method according to claim 1, characterized in that The method further comprises: Determining request object information associated with the client in the query request information; Determine a collaborative relationship between a historical request object in the historical distribution information and the plurality of candidate emoticon images, wherein the collaborative relationship is represented by a graph model; Based on the request object information and the collaborative relationship, determining at least one first candidate emoticon picture from the plurality of candidate emoticon pictures; According to the at least one first candidate expression picture, the at least one target candidate expression picture is updated.

7. The method according to claim 1, characterized in that The step of obtaining a sequence of pictures to be distributed that matches the query request information according to the at least one target candidate expression picture includes: According to the first correlation degree corresponding to each target candidate expression picture in the at least one target candidate expression picture and the second correlation degree corresponding to each target candidate expression picture, the at least one target candidate expression picture is sorted to obtain a first picture sequence; Determining request object information associated with the client in the query request information; Determine background information corresponding to the query request information, the background information including interaction context information, service type information and picture quality requirement information; Screening the first picture sequence according to the picture quality requirement information to obtain a second picture sequence; Determine a first matching degree between each target candidate expression picture in the second picture sequence and the requested object information; Determining a second matching degree between each target candidate expression picture in the second picture sequence and the interaction context information; Determining a third matching degree between each target candidate expression picture in the second picture sequence and the business type information; The second picture sequence is reordered according to the first matching degree, the second matching degree, and the third matching degree to obtain the picture sequence to be distributed.

8. The method according to claim 1, characterized in that The method further comprises: When the number of the at least one target candidate emoticon image is lower than a preset number threshold, determining key text information of the query request information; Acquire multiple basic expression pictures corresponding to the multiple candidate expression pictures; none of the multiple basic expression pictures contain text; Obtaining index vector data of each basic expression picture in the plurality of basic expression pictures, wherein the index vector data is text label data and semantic vector data in vector form; Determining at least one target basic expression picture from the multiple basic expression pictures according to the similarity between the index vector data of each basic expression picture and the requested semantic feature data; Synthesizing the key text information and the at least one target basic expression picture to obtain at least one second candidate expression picture; According to the at least one second candidate expression picture, the at least one target candidate expression picture is updated.

9. An expression picture matching device, characterized in that: The device comprises: An acquisition module is used to obtain query request information for emoticon images sent by the client; A feature extraction module, used to perform feature extraction processing based on the query request information to obtain request feature information, wherein the request feature information includes request text feature data and request semantic feature data; A text relevance determination module, used to determine a first relevance between text label data of each candidate emoticon picture among a plurality of candidate emoticon pictures and the requested text feature data; the text label data is content feature data described in text and obtained by performing multi-dimensional analysis on the corresponding candidate emoticon pictures; A vector association degree determination module is used to determine a second association degree between the semantic vector data of each candidate expression picture and the requested semantic feature data; the semantic vector data is content feature data described by a semantic vector obtained by performing a multi-dimensional analysis on the corresponding candidate expression picture; A recall module, configured to determine at least one target candidate expression picture from the plurality of candidate expression pictures according to the first association degree and the second association degree; A sequence generation module is used to obtain a sequence of pictures to be distributed that matches the query request information based on the at least one target candidate expression picture, wherein the sequence of pictures to be distributed is part or all of the at least one target candidate expression picture arranged in a preset order.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement an expression picture matching method as described in any one of claims 1 to 8.

11. A computer device, characterized in that: The computer device includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or at least one program is loaded and executed by the processor to implement an expression picture matching method as described in any one of claims 1 to 8.