Processing Method, Device and Server for Emoji Pictures

By extracting and processing keyframes of dynamic emoticon package pictures, combining text and image semantics processing rules, the problem of difficult to recognize dynamic emoticon package pictures in the existing technology is solved, and accurate recognition and interactive experience are improved.

CN114880512BActive Publication Date: 2025-05-27CHINA CONSTRUCTION BANK
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210445438.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-26
Publication Date
2025-05-27
Estimated Expiration
2042-04-26

AI Technical Summary

Technical Problem

The prior art is difficult to accurately and automatically identify the semantic content represented by dynamic emoticon pictures, resulting in the inability to accurately reply to users and affect the interactive experience.

Method used

By obtaining the target emoticon image, extracting keyframe images, combining preset text semantic processing rules and image semantic processing rules, the image group is processed to obtain text semantic content and image semantic content, and then determining the target semantic content.

Benefits of technology

It realizes accurate semantic recognition of dynamic emoticon pack pictures, accurately determines semantic content, and improves user interaction experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114880512B_ABST
    Figure CN114880512B_ABST
Patent Text Reader

Abstract

This specification provides a method, apparatus, and server for processing meme pictures. The method relates to the field of artificial intelligence technology. Based on this method, in the intelligent customer service scenario, when a target user uses a target meme picture such as a dynamic meme picture in the conversation interface, the server can first extract multiple key frame pictures from the target picture sequence corresponding to the target meme picture according to a preset extraction rule to obtain a target picture group; then process the target picture group according to a preset semantic processing rule, obtain and determine the target semantic content represented by the target meme picture based on the text semantic content and / or image semantic content of the target meme picture. Thus, it can automatically and accurately identify and determine the target semantic content represented by the target meme picture used by the target user, and accurately determine the matching target reply content according to the target semantic content for timely reply, improving the user's interaction experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification belongs to the field of artificial intelligence technology, and particularly relates to a method, device, and server for processing emoji pictures. Background Art

[0002] In the customer service reply scenario, processing devices such as servers need to identify the true semantic content of the user based on the content data input by the user in the dialogue interface for automatic reply.

[0003] However, in addition to text data, users sometimes use emoji pictures in the dialogue interface. Based on existing methods, it is often difficult to accurately and automatically identify the semantic content represented by the emoji pictures. Especially when the emoji pictures used by the user are dynamic emoji pictures, the recognition error is relatively large when performing semantic recognition processing based on existing methods. This will further lead to inaccurate automatic reply to the user and affect the user's interaction experience.

[0004] In response to the above problems, no effective solution has been proposed yet. Summary of the Invention

[0005] The method, device, and server for processing emoji pictures provided in this specification can automatically and accurately identify and determine the target semantic content represented by the target emoji picture used by the target user in the dialogue interface, and accurately determine the matching target reply content for reply according to the target semantic content, improving the user's interaction experience.

[0006] This specification provides a method for processing emoji pictures, including:

[0007] Obtain a target emoji picture; wherein, the target emoji picture includes the emoji picture used by the target user in the dialogue interface; the target emoji picture includes a dynamic emoji picture;

[0008] Extract a plurality of key frame pictures from the target picture sequence corresponding to the target emoji picture according to a preset extraction rule to obtain a target picture group;

[0009] Process the target picture group according to a preset semantic processing rule to obtain the text semantic content and / or image semantic content of the target emoji picture; wherein, the preset semantic processing rule at least includes a preset text semantic processing rule and a preset image semantic processing rule;

[0010] Determine the target semantic content represented by the target emoji picture according to the text semantic content and / or image semantic content of the target emoji picture.

[0011] In one embodiment, according to a preset extraction rule, multiple key-frame images are extracted from the image sequence corresponding to the target emoji image to obtain a target image group, including:

[0012] Parse the target emoji image to obtain the target image sequence corresponding to the target emoji image;

[0013] According to the preset extraction rule, extract a first number of key-frame images from the first image sequence in the target image sequence; extract a second number of key-frame images from the second image sequence in the target image sequence; wherein, the images in the first image sequence are displayed earlier in the target image sequence than the images in the second image sequence; the value of the first number is less than the value of the second number;

[0014] Arrange the multiple key-frame images according to the display order to obtain a target image group.

[0015] In one embodiment, according to a preset semantic processing rule, process the target image group to obtain the text semantic content of the target emoji image, including:

[0016] According to the preset text semantic processing rule, detect whether there are text characters in the key-frame images in the target image group;

[0017] In the case of determining that there are text characters in the key-frame images in the target image group, determine the key-frame images with text characters as the first type of key-frame images;

[0018] Call a preset OCR recognition model to process the first type of key-frame images to obtain the text semantic content of the first type of key-frame images;

[0019] According to the display order of the first type of key-frame images in the target image sequence, combine the text semantic content of multiple first type of key-frame images to obtain the text semantic content of the target emoji image.

[0020] In one embodiment, call a preset OCR recognition model to process the first type of key-frame images to obtain the text semantic content of the first type of key-frame images, including:

[0021] Call a preset OCR recognition model to process the first type of key-frame images to extract the text characters in the first type of key-frame images;

[0022] Call a preset character classification model to process the text characters in the first type of key-frame images to determine the character type of the text characters; wherein, the character type includes printed characters and non-printed characters;

[0023] According to the character type of the text characters, use the matching semantic recognition method to perform semantic recognition on the text characters in the first type of key frame pictures, so as to obtain the text semantic content of the first type of key frame pictures.

[0024] In one embodiment, when it is determined that the character type of the text characters is printed characters, according to the character type of the text characters, use the matching semantic recognition method to perform semantic recognition on the text characters in the first type of key frame pictures, so as to obtain the text semantic content of the first type of key frame pictures, including:

[0025] According to the preset text semantic processing rules, use the preset character template to perform text feature matching with the text characters in the first type of key frame pictures to obtain the corresponding text matching result;

[0026] According to the text matching result, determine the text semantic content of the first type of key frame pictures.

[0027] In one embodiment, when it is determined that the character type of the text characters is non-printed characters, according to the character type of the text characters, use the matching semantic recognition method to perform semantic recognition on the text characters in the first type of key frame pictures, so as to obtain the text semantic content of the first type of key frame pictures, including:

[0028] According to the preset text semantic processing rules, use the preset text character semantic recognition model to process the text characters in the first type of key frame pictures to obtain the initial semantic recognition result; wherein, the preset text character semantic recognition model is trained using the character sample data containing non-printed text characters;

[0029] Obtain the context correlation data of the target meme picture in the target user's conversation interface;

[0030] Adjust the initial semantic recognition result according to the context correlation data to obtain the text semantic content of the first type of key frame pictures.

[0031] In one embodiment, according to the preset semantic processing rules, process the target picture group to obtain the image semantic content of the target meme picture, including:

[0032] According to the preset image semantic processing rules, extract the image features of the key frame pictures;

[0033] Match the image features of the key frame pictures with the preset expression database to obtain the corresponding image matching result; wherein, the preset expression database stores multiple preset meme pictures and the semantic labels corresponding to the preset meme pictures;

[0034] Determine the semantic label of the preset emoji picture that matches the image features of the key frame picture according to the image matching result, as the image semantic content of the key frame picture;

[0035] According to the display order of the key frame pictures in the target picture sequence, combine the image semantic contents of multiple key frame pictures to obtain the image semantic content of the target emoji picture.

[0036] In one embodiment, after matching the image features of the key frame picture with the preset emoji database, the method further includes:

[0037] In the case where it is determined that there is no preset emoji picture in the preset emoji database that matches the image features of the key frame picture according to the image matching result, obtain the context association data of the target emoji picture in the dialogue interface of the target user according to the preset image semantic processing rules;

[0038] Combine the image features of the key frame picture and the context association data of the target emoji picture to obtain a combined feature;

[0039] Process the combined feature using a preset comprehensive semantic recognition model to obtain a corresponding comprehensive recognition result; wherein, the preset comprehensive semantic recognition model is trained using comprehensive sample data including image features and context association data;

[0040] Determine the image semantic content of the key frame picture according to the comprehensive recognition result.

[0041] In one embodiment, determining the target semantic content represented by the target emoji picture according to the text semantic content and / or image semantic content of the target emoji picture includes:

[0042] Detect whether the reliability of the text semantic content of the target emoji picture meets the preset reliability requirements;

[0043] In the case where it is determined that the reliability of the text semantic content of the target emoji picture meets the preset reliability requirements, determine the text semantic content of the target emoji picture as the target semantic content represented by the target emoji picture.

[0044] In one embodiment, after detecting whether the reliability of the text semantic content of the target emoji picture meets the preset reliability requirements, the method further includes:

[0045] In the case where it is determined that the reliability of the text semantic content of the target emoji picture does not meet the preset reliability requirements, count the proportion of the number of key frame pictures in the target picture group whose image features match the preset emoji database;

[0046] Detect whether the reliability of the image semantic content of the target emoji picture meets the preset reliability requirements according to the proportion of the number of key-frame pictures in the target picture group whose image features match the preset emoji database;

[0047] When it is determined that the reliability of the image semantic content of the target emoji picture meets the preset reliability requirements, determine the image semantic content of the target emoji picture as the target semantic content represented by the target emoji picture.

[0048] In one embodiment, after detecting whether the reliability of the image semantic content of the target emoji picture meets the preset reliability requirements, the method further includes:

[0049] When it is determined that the reliability of the image semantic content of the target emoji picture does not meet the preset reliability requirements, determine the target semantic content represented by the target emoji picture by combining the text semantic content and the image semantic content of the target emoji picture.

[0050] In one embodiment, after determining the target semantic content represented by the target emoji picture, the method further includes:

[0051] Determine a matching target reply text according to the target semantic content;

[0052] In the dialogue interface of the target user, reply the target reply text to the target user.

[0053] In one embodiment, the method further includes:

[0054] Obtain the position information of the target emoji picture in the dialogue interface of the target user;

[0055] According to the position information, when it is determined that the target emoji picture is the starting conversation initiated by the target user in the dialogue interface, determine the preset greeting text as the matching target reply text.

[0056] This specification also provides a processing device for emoji pictures, including:

[0057] An acquisition module, configured to acquire a target emoji picture; wherein, the target emoji picture includes the emoji picture used by the target user in the dialogue interface; the target emoji picture includes a dynamic emoji picture;

[0058] An extraction module, configured to extract a plurality of key-frame pictures from the target picture sequence corresponding to the target emoji picture according to a preset extraction rule to obtain a target picture group;

[0059] A processing module, configured to process a target picture group according to preset semantic processing rules to obtain the text semantic content and / or image semantic content of the target meme picture; wherein, the preset semantic processing rules at least include a preset text semantic processing rule and a preset image semantic processing rule;

[0060] A determination module, configured to determine the target semantic content represented by the target meme picture according to the text semantic content and / or image semantic content of the target meme picture.

[0061] This specification also provides a server, including a processor and a memory for storing instructions executable by the processor. When the processor executes the instructions, the steps of the processing method of the meme picture are implemented.

[0062] This specification also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the processing method of the meme picture are implemented.

[0063] Based on the processing method, device and server of the meme picture provided in this specification, in the intelligent customer service scenario, when a target user uses a target meme picture such as a dynamic meme picture in a dialogue interface, the server can first extract, according to preset extraction rules, multiple key frame pictures with a high probability of containing important semantic content from the target picture sequence corresponding to the target meme picture with a large data volume to obtain a target picture group with a small data volume; then process the target picture group according to preset semantic processing rules that at least include a preset text semantic processing rule and a preset image semantic processing rule to obtain the text semantic content and / or image semantic content of the target meme picture; and then can determine the target semantic content represented by the target meme picture according to the text semantic content and / or image semantic content of the target meme picture. Thus, it can automatically and accurately identify and determine the target semantic content represented by the target user's target meme picture, and accurately determine the matching target reply content according to the target semantic content for timely reply, improving the user's interaction experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] To more clearly illustrate the embodiments of this specification, the following will briefly introduce the drawings required for the embodiments. The drawings described below are only some embodiments recorded in this specification. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained according to these drawings.

[0065] Figure 1 A flowchart of the processing method of the meme picture provided by an embodiment of this specification;

[0066] Figure 2It is a schematic diagram of an embodiment of applying the method for processing meme pictures provided in the embodiments of this specification in a scenario example;

[0067] Figure 3 It is a schematic diagram of an embodiment of applying the method for processing meme pictures provided in the embodiments of this specification in a scenario example;

[0068] Figure 4 It is a schematic diagram of an embodiment of applying the method for processing meme pictures provided in the embodiments of this specification in a scenario example;

[0069] Figure 5 It is a schematic diagram of an embodiment of applying the method for processing meme pictures provided in the embodiments of this specification in a scenario example;

[0070] Figure 6 It is a schematic diagram of an embodiment of applying the method for processing meme pictures provided in the embodiments of this specification in a scenario example;

[0071] Figure 7 It is a schematic diagram of an embodiment of applying the method for processing meme pictures provided in the embodiments of this specification in a scenario example;

[0072] Figure 8 It is a schematic diagram of the structural composition of a server provided in an embodiment of this specification;

[0073] Figure 9 It is a schematic diagram of the structural composition of a device for processing meme pictures provided in an embodiment of this specification. Detailed implementation manners

[0074] In order to enable those skilled in the art to better understand the technical solutions in this specification, the following will clearly and completely describe the technical solutions in the embodiments of this specification with reference to the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of this specification, rather than all the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this specification.

[0075] Refer to Figure 1 As shown, the embodiments of this specification provide a method for processing meme pictures. Specifically, when implementing this method, it may include the following content:

[0076] S101: Obtain a target meme picture; wherein, the target meme picture includes the meme pictures used by the target user in the conversation interface; the target meme picture includes dynamic meme pictures;

[0077] S102: Extract multiple key frame images from the target image sequence corresponding to the target emoticon image according to a preset extraction rule to obtain a target image group;

[0078] S103: Process the target image group according to a preset semantic processing rule to obtain the text semantic content and / or image semantic content of the target emoticon image; wherein, the preset semantic processing rule at least includes a preset text semantic processing rule and a preset image semantic processing rule;

[0079] S104: Determine the target semantic content represented by the target emoticon image according to the text semantic content and / or image semantic content of the target emoticon image.

[0080] In some embodiments, the above method for processing emoticon images can be specifically applied to an intelligent customer service scenario.

[0081] Specifically, when a target user needs to consult a service provider (e.g., the network service platform of XX product) about relevant business issues (e.g., after-sales issues of XX product, consultation on pre-sales preferential activities, etc.), the target user can first establish a customer service consultation dialogue interface with the customer service server of the service provider through a terminal device. Then, the target user can describe their problems by inputting dialogue content data including text data (e.g., "Are there any preferential activities for XX product?") and emoticon images in the dialogue interface. For details, please refer to Figure 2 as shown.

[0082] Among them, the above emoticon image can be specifically understood as a kind of image data that is mostly applied in social software and expresses the emotions and / or semantics of users through static images or dynamic images, etc.

[0083] The terminal device can obtain the above dialogue content data through the dialogue interface and feedback the above dialogue content data to the customer service server. Refer to Figure 3 as shown. The customer service server can perform semantic recognition processing on the text data and emoticon images in the content data respectively to obtain the final semantic content corresponding to the dialogue content data. Among them, when the customer service server performs recognition processing on the emoticon images in the dialogue content data, it can apply the method for processing emoticon images provided in this specification to accurately recognize the semantic content of the emoticon images and reduce recognition errors.

[0084] Furthermore, the customer service server can determine the target problem of the target user according to the final semantic content; then find the target reply problem that matches the target problem of the target user; and finally reply the above target reply text to the target user through the dialogue interface to automatically and efficiently reply to the business problems of the target user.

[0085] In this embodiment, referring to Figure 3 as shown, the above-mentioned customer service server may specifically include a background server applied to one side of a network service platform (XX product network service platform) that can implement functions such as data transmission and data processing. Specifically, the customer service server may be, for example, an electronic device with data operation, storage functions, and network interaction functions. Or, the customer service server may also be a software program running in this electronic device that provides support for data processing, storage, and network interaction. In this embodiment, the number of servers included in the customer service server is not specifically limited. The customer service server may specifically be one server, or several servers, or a server cluster formed by several servers.

[0086] In this embodiment, the terminal device may specifically include a front end applied to the user side that can implement functions such as data collection and data transmission. Specifically, the terminal device may be, for example, electronic devices such as desktop computers, tablet computers, laptop computers, and smart phones. Or, the terminal device may also be a software application that can run in the above-mentioned electronic devices. For example, it may be a certain APP running on a smart phone.

[0087] In some embodiments, the terminal device may send the conversation content data input by the target user in the conversation interface to the customer service server. After receiving the above-mentioned conversation content data, the customer service server may first detect whether there are emoji pictures in the conversation content data. Considering that the recognition and processing of emoji pictures are more difficult than the recognition and processing of conventional text, when the customer service server detects that there are emoji pictures in the conversation content data, it may extract the emoji pictures from the conversation content data as the target emoji pictures to be recognized and processed currently. Correspondingly, the customer service server may obtain the target emoji pictures.

[0088] In some embodiments, the above-mentioned target emoji pictures may specifically include: static emoji pictures and dynamic emoji pictures. Among them, the above-mentioned static emoji pictures usually only contain one frame of picture, and this frame of picture is presented to the user in a static manner to express relevant emotions or / and semantics. The above-mentioned dynamic emoji pictures usually contain a picture sequence composed of multiple frames of pictures, and the multiple frames of pictures in this picture sequence are presented to the user in a dynamic manner to express relevant emotions or / and semantics.

[0089] In this embodiment, the recognition and processing of dynamic emoji pictures are mainly taken as an example for specific description. The recognition and processing of static emoji pictures can refer to the following embodiments of the recognition and processing of dynamic emoji pictures. For this, this specification will not elaborate.

[0090] In some embodiments, the above-mentioned preset extraction rules can be specifically determined after learning a large number of sample dynamic emoji pictures. By learning and sorting out a large number of sample dynamic emoji pictures, and combining with human language habits, the following layout characteristics are found: In a dynamic emoji, the pictures containing relatively important semantic content often tend to be concentrated in the pictures with a relatively late display order in the picture sequence. Based on the above layout characteristics, further statistical analysis is carried out on the picture sequences of a large number of sample dynamic emoji pictures to obtain the preset extraction rules. Among them, the pictures in the picture sequence corresponding to the dynamic emoji that have a relatively high probability of containing relatively important semantic content can be recorded as key frame pictures.

[0091] In some embodiments, referring to Figure 4 As shown, according to the above-mentioned preset extraction rules, multiple key frame pictures are extracted from the picture sequence corresponding to the target emoji picture to obtain a target picture group. Specifically, in implementation, it may include the following content:

[0092] S102-1: Analyze the target emoji picture to obtain the target picture sequence corresponding to the target emoji picture;

[0093] S102-2: According to the preset extraction rules, extract the first number of key frame pictures from the first picture sequence in the target picture sequence; extract the second number of key frame pictures from the second picture sequence in the target picture sequence; wherein, the pictures in the first picture sequence are displayed earlier in the target picture sequence than the pictures in the second picture sequence; the value of the first number is less than the value of the second number.

[0094] S102-3: Arrange the multiple key frame pictures according to the display order to obtain a target picture group.

[0095] In some embodiments, the above-mentioned first number, first picture sequence, second number, and second picture sequence can be specifically determined according to the preset extraction rules.

[0096] Specifically, the above-mentioned first picture sequence can be the picture sequence of the first 30% of the pictures with relatively early display order in the target picture sequence corresponding to the target emoji picture. The above-mentioned second picture sequence can be the picture sequence of the last 30% of the pictures with relatively late display order in the target picture sequence corresponding to the target emoji picture. The value of the first number is greater than the second number. The specific values of the first number and the second number are determined according to the total number of pictures included in the target picture sequence.

[0097] For example, if the target picture sequence contains 100 pictures, the first picture sequence can be the first 30 pictures with a relatively early display order among these 100 pictures, and the second picture sequence can be the last 30 pictures with a relatively late display order among these 100 pictures. The above-mentioned first quantity can specifically be 5 pictures, and the above-mentioned second quantity can specifically be 10 pictures.

[0098] In specific implementation, according to the preset extraction rule, a first quantity of key frame pictures can be randomly extracted from the first picture sequence in the target picture sequence; at the same time, according to the preset extraction rule, a second quantity of key frame pictures can be extracted from the second picture sequence in the target picture sequence, so that pictures with a relatively small quantity but a relatively high probability of containing relatively important semantic content can be screened out from the target picture sequence with a large number of frames as key frame pictures to participate in subsequent recognition processing. In this way, on the one hand, the data processing volume can be reduced and the overall data processing efficiency can be improved; on the other hand, the interference of redundant pictures on the final recognition result can also be reduced, and the accuracy of subsequent recognition processing can be improved.

[0099] In some embodiments, the above-mentioned preset semantic processing rules at least include a preset text semantic processing rule and a preset image semantic processing rule. Among them, the above-mentioned preset text semantic processing rule is a processing rule for recognizing the text characters in the picture to extract the semantic content contained in the text characters. The above-mentioned preset image semantic processing rule is a processing rule for recognizing the image content of non-text characters in the picture to extract the semantic content contained in the image content.

[0100] Based on the above embodiments, by using the preset semantic processing rules that at least include the preset text semantic processing rule and the preset image semantic processing rule, semantic recognition processing can be performed on the picture from two different dimensions of text characters and the image content of non-text characters, so that the semantic content in the picture can be recognized more accurately and comprehensively.

[0101] In some embodiments, the above-mentioned processing of the target picture group according to the preset semantic processing rule to obtain the text semantic content of the target meme picture may specifically include the following contents when implemented:

[0102] S1: According to the preset text semantic processing rule, detect whether there are text characters in the key frame pictures in the target picture group;

[0103] S2: In the case of determining that there are text characters in the key frame pictures in the target picture group, determine the key frame pictures with text characters as the first type of key frame pictures;

[0104] S3: Process the first type of key-frame pictures by invoking a preset OCR (Optical Character Recognition) model to obtain the text semantic content of the first type of key-frame pictures;

[0105] S4: Combine the text semantic contents of multiple first type of key-frame pictures according to the display order of the first type of key-frame pictures in the target picture sequence to obtain the text semantic content of the target meme picture.

[0106] Among them, the above-mentioned preset OCR (Optical Character Recognition) model can be specifically understood as a pre-trained model that can recognize and extract text characters from pictures.

[0107] Through the above embodiments, key-frame pictures without text characters can be filtered out from the key-frame pictures first, and only key-frame pictures with text characters are retained to participate in the subsequent recognition and processing of text semantic content, thereby effectively reducing the data processing volume and improving the overall processing efficiency.

[0108] In some embodiments, refer to Figure 5 As shown, when specifically implementing the process of processing the first type of key-frame pictures by invoking a preset OCR model to obtain the text semantic content of the first type of key-frame pictures, the following contents may be included:

[0109] S1: Invoke a preset OCR model to process the first type of key-frame pictures to extract the text characters in the first type of key-frame pictures;

[0110] S2: Invoke a preset character classification model to process the text characters in the first type of key-frame pictures to determine the character type of the text characters; among them, the character type includes printed characters and non-printed characters;

[0111] S3: According to the character type of the text characters, use a matching semantic recognition method to perform semantic recognition on the text characters in the first type of key-frame pictures to obtain the text semantic content of the first type of key-frame pictures.

[0112] In some embodiments, the above-mentioned preset character classification model can be specifically understood as a pre-trained neural network model that can recognize whether the input text characters are printed characters or non-printed characters.

[0113] Before specific implementation, an initial first classification model can be constructed; use the labeled character samples marked with printed characters and non-printed characters as training data to train the initial first classification model to obtain a preset character classification model that meets the requirements.

[0114] In some embodiments, refer to Figure 5As shown, when it is determined that the character type of the text character is printed text, according to the character type of the text character, the semantic recognition method that matches is used to perform semantic recognition on the text character in the first type of key frame picture, so as to obtain the text semantic content of the first type of key frame picture. Specifically in implementation, it may include the following content:

[0115] S1: According to the preset text semantic processing rule, use the preset character template to perform text feature matching with the text character in the first type of key frame picture to obtain the corresponding text matching result;

[0116] S2: According to the text matching result, determine the text semantic content of the first type of key frame picture.

[0117] Among them, the preset character template stores standard printed text characters and the text semantic information corresponding to each printed text character. Correspondingly, the above text matching result can specifically be the semantic information corresponding to the standard printed text character that matches the text character in the first type of key frame picture.

[0118] Based on the above embodiment, for the text characters of the printed text type that are better recognizable and more standard in the first type of key frame picture, the preset character template can be used to quickly determine the contained text semantic content.

[0119] In some embodiments, refer to Figure 5 As shown, when it is determined that the character type of the text character is non-printed text, according to the character type of the text character, the semantic recognition method that matches is used to perform semantic recognition on the text character in the first type of key frame picture, so as to obtain the text semantic content of the first type of key frame picture. Specifically in implementation, it may include the following content:

[0120] S1: According to the preset text semantic processing rule, use the preset text character semantic recognition model to process the text character in the first type of key frame picture to obtain the initial semantic recognition result; among them, the preset text character semantic recognition model is trained using the character sample data containing non-printed text characters;

[0121] S2: Obtain the context-related data of the target emoji picture in the dialogue interface of the target user;

[0122] S3: Adjust the initial semantic recognition result according to the context-related data to obtain the text semantic content of the first type of key frame picture.

[0123] Among them, the above preset text character semantic recognition model can specifically be understood as a pre-trained neural network model that can perform semantic recognition on the text character in the picture.

[0124] The above context-related data can be specifically understood as text data adjacent to the target expression picture. Refer to Figure 2 as shown. For example, in a conversation interface, the text data "Is there still a discount for XX products?" entered by the user before entering the target expression pack picture can be used as the context-related data of the target expression pack picture.

[0125] Before specific implementation, a preset text character semantic recognition model can be trained in the following way: construct an initial first prediction model; obtain character samples; where the character samples at least include non-printing type character samples; label the corresponding text semantic information on the character samples to obtain the labeled sample data; use the labeled sample data to train the initial first prediction model to obtain the preset text character semantic recognition model.

[0126] Based on the above embodiments, for the difficult-to-identify and non-standard non-printing type text characters in the first type of key frame pictures, the preset text character semantic recognition model can be used to accurately determine the included text semantic content.

[0127] In some embodiments, after obtaining the text semantic content of the first type of key frame pictures, when the method is specifically implemented, it may further include: obtaining the text semantic content of the adjacent pictures of the current first type of key frame pictures in the target picture group; correcting the text semantic content of the current first type of key frame pictures according to the text semantic content of the adjacent pictures. Thus, the text semantic content of the first type of key frame pictures with higher accuracy and smaller error can be obtained.

[0128] In some embodiments, refer to Figure 6 as shown. The above-mentioned processing of the target picture group according to the preset semantic processing rules to obtain the image semantic content of the target expression pack picture may specifically include the following contents when implemented:

[0129] S1: Extract the image features of the key frame pictures according to the preset image semantic processing rules;

[0130] S2: Match the image features of the key frame pictures with the preset expression database to obtain the corresponding image matching result; where the preset expression database stores multiple preset expression pack pictures and the semantic labels corresponding to the preset expression pack pictures;

[0131] S3: According to the image matching result, determine the semantic label of the preset expression pack picture that matches the image features of the key frame pictures as the image semantic content of the key frame pictures;

[0132] S4: According to the display order of the key-frame pictures in the target picture sequence, combine the image semantic contents of multiple key-frame pictures to obtain the image semantic content of the target emoji picture.

[0133] Among them, a preset emoji database stores multiple preset emoji pictures and semantic labels corresponding to the preset emoji pictures.

[0134] The above-mentioned preset emoji database can be obtained by clustering a large number of sample emoji pictures in advance to get multiple preset emoji pictures that can represent a certain type of graphic semantic information; then, according to the represented image semantic information, set the semantic labels corresponding to the preset emoji pictures to construct the preset emoji database.

[0135] The above-mentioned preset emoji database can also retain the historical emoji pictures processed during application; and at every preset time interval, use the historical emoji pictures retained in the previous preset time interval to update the preset emoji database, so as to gradually obtain a relatively rich, comprehensive and relatively better-effect preset emoji database.

[0136] Based on the above embodiments, for the emoji pictures that relatively often appear in the key-frame pictures, the corresponding image semantic content can be determined relatively quickly and accurately by using the preset emoji database.

[0137] In some embodiments, as shown in Figure 6 After matching the image features of the key-frame pictures with the preset emoji database, when the method is specifically implemented, the following contents can also be included:

[0138] S1: When it is determined according to the image matching result that there is no preset emoji picture in the preset emoji database that matches the image features of the key-frame pictures, obtain the context-related data of the target emoji picture in the target user's dialogue interface according to the preset image semantic processing rules;

[0139] S2: Combine the image features of the key-frame pictures and the context-related data of the target emoji picture to obtain a combined feature;

[0140] S3: Process the combined feature by using a preset comprehensive semantic recognition model to obtain a corresponding comprehensive recognition result; among them, the preset comprehensive semantic recognition model is trained by using comprehensive sample data including image features and context-related data;

[0141] S4: Determine the image semantic content of the key-frame pictures according to the comprehensive recognition result.

[0142] Among them, the above-mentioned preset comprehensive semantic recognition model can be specifically understood as a neural network model that has been pre-trained to recognize and determine the image semantic content of a picture in the context of context association based on the combined features of the image features of the spliced picture with the input and the context association data.

[0143] Before specific implementation, the preset comprehensive semantic recognition model can be trained in the following way: construct an initial second prediction model; obtain sample dialogue data; among them, the sample dialogue data at least includes sample emoji pictures; extract sample image features from the sample emoji pictures; at the same time, obtain the context association data of the sample emoji pictures from the sample dialogue data; splice the sample image features of the sample emoji pictures with the context association data of the sample emoji pictures to obtain sample data; label the image semantic content represented by the sample emoji pictures in the sample data to obtain the labeled sample data; use the labeled sample data to train the initial second prediction model to obtain the preset comprehensive semantic recognition model.

[0144] Based on the above embodiments, for the emoji pictures that relatively rarely appear in the key frame pictures, the corresponding image semantic content can be accurately determined by using the preset comprehensive semantic recognition model.

[0145] In some embodiments, after obtaining the image semantic content of the key frame picture, when the method is specifically implemented, it may further include: obtaining the image semantic content of the adjacent pictures of the current key frame picture in the target picture group; correcting the image semantic content of the current key frame picture according to the image semantic content of the adjacent pictures. Thus, the image semantic content of the key frame picture with higher accuracy and smaller error can be obtained.

[0146] In some embodiments, the above-mentioned determining the target semantic content represented by the target emoji picture according to the text semantic content and / or image semantic content of the target emoji picture may specifically include:

[0147] S1: Detect whether the reliability of the text semantic content of the target emoji picture meets the preset reliability requirements;

[0148] S2: In the case where it is determined that the reliability of the text semantic content of the target emoji picture meets the preset reliability requirements, determine the text semantic content of the target emoji picture as the target semantic content represented by the target emoji picture.

[0149] In some embodiments, when specifically implementing whether the reliability of the text semantic content of the detected target meme picture meets the preset reliability requirements, it may include: counting the proportion of the first type of quantity of the first type of key-frame pictures in the target picture group; counting the proportion of the second type of quantity of the first type of key-frame pictures with the character type of printed characters in the target picture group; obtaining the confidence parameter output by the preset text character semantic recognition model when processing the first type of key-frame pictures with the character type of non-printed characters; determining the reliability of the text semantic content of the target meme picture according to the proportion of the first type of quantity, the proportion of the second type of quantity, and the confidence parameter; comparing the reliability with the preset reliability threshold to determine whether the reliability of the text semantic content of the target meme picture meets the preset reliability requirements.

[0150] In some embodiments, when specifically implementing, the target picture group may be processed according to the preset semantic processing rules first to obtain the text semantic content of the target meme picture; then it is detected whether the reliability of the text semantic content of the target meme picture meets the preset reliability requirements. When it is determined whether the reliability of the text semantic content of the target meme picture meets the preset reliability requirements, the target picture group may not be processed according to the preset semantic processing rules anymore; instead, the text semantic content of the target meme picture is directly determined as the target semantic content represented by the target meme picture. Thus, the target semantic content represented by the target meme picture can be determined relatively quickly.

[0151] In some embodiments, after detecting whether the reliability of the image semantic content of the target meme picture meets the preset reliability requirements, when specifically implementing the method, the following content may further be included: when it is determined that the reliability of the image semantic content of the target meme picture does not meet the preset reliability requirements, the target semantic content represented by the target meme picture is determined by combining and using the text semantic content and the image semantic content of the target meme picture.

[0152] In some embodiments, when specifically implementing determining the target semantic content represented by the target meme picture by combining and using the text semantic content and the image semantic content of the target meme picture, it may include: splicing the text semantic content and the image semantic content of the target meme picture to obtain the spliced semantic content; calling a semantic prediction model to process the spliced semantic content to determine the target semantic content.

[0153] In some embodiments, when specifically implementing determining the target semantic content represented by the target meme picture by combining and using the text semantic content and the image semantic content of the target meme picture, it may further include: jointly using the text semantic content and the image semantic content of the target meme picture for mutual verification to determine the target semantic content.

[0154] Based on the above embodiments, by combining the text semantic content and the image semantic content of the target emoji picture, the target semantic content represented by the target emoji picture can be determined more accurately, reducing the recognition error.

[0155] In some embodiments, after determining the target semantic content represented by the target emoji picture, when the method is specifically implemented, the following content may further be included:

[0156] S1: Determine a target reply text that matches according to the target semantic content;

[0157] S2: In the dialogue interface of the target user, reply the target user with the target reply text.

[0158] Specifically, the customer service server can prepare a preset reply text library in advance; among them, the preset reply text library can store preset reply texts corresponding to multiple common preset questions.

[0159] When specifically implemented, the customer service server can obtain the complete and comprehensive semantic content of the dialogue content data for the target user according to the target semantic content recognized based on the target emoji picture in the dialogue content data and the semantic content recognized based on the text data in the dialogue content data; then, according to the complete and comprehensive semantic content, accurately determine the corresponding target question; furthermore, the preset reply text library can be retrieved according to the target question to find the preset reply text corresponding to the preset question that matches the target question as the target reply text.

[0160] Specifically, as shown in Figure 7 The customer service server can return the determined target reply text to the terminal device. The terminal device receives the target reply text. For example, the text data "There is currently a promotion of 50 off for every 100 spent that suits you. The link...". Further, the terminal device can promptly display the target reply text to the target user in the dialogue interface, so as to achieve an automatic reply to the target user.

[0161] In some embodiments, when the method is specifically implemented, it may further include: obtaining the position information of the target emoji picture in the dialogue interface of the target user; according to the position information, when determining that the target emoji picture is the starting dialogue initiated by the target user in the dialogue interface, determining the preset greeting text as the target reply text that matches.

[0162] This is because, after learning and organizing a large number of user conversation records in advance, it is found that: many users' meme pictures sent at the beginning in the conversation interface are often just for greeting and have no actual semantics. Therefore, in addition to identifying and processing meme pictures in terms of text semantics dimension and image semantics dimension, it is also possible to introduce an analysis and recognition of meme pictures based on the conversation position dimension. In this way, for meme pictures used in some special conversation situations, the corresponding semantic content can be more quickly identified and determined.

[0163] In some embodiments, when it is impossible to determine the target semantic content represented by the target meme picture, the current conversation interface where the target meme picture is located can also be screenshot and the screenshot can be forwarded to the customer service staff for manual reply, so as to provide an accurate reply to the user in a timely manner and enable the user to obtain a better interaction experience.

[0164] As can be seen from the above, based on the meme picture processing method provided in the embodiments of this specification, in the intelligent customer service scenario, when a target user uses a target meme picture such as a dynamic meme picture in the conversation interface, multiple key frame pictures can be extracted from the target picture sequence corresponding to the target meme picture according to the preset extraction rules to obtain a target picture group; then, according to the preset semantic processing rules that at least include the preset text semantic processing rules and the preset image semantic processing rules, the target picture group is processed to obtain the text semantic content and / or the image semantic content of the target meme picture; then, according to the text semantic content and / or the image semantic content of the target meme picture, the target semantic content represented by the target meme picture is determined. Thus, it is possible to automatically and accurately identify and determine the target semantic content represented by the target user's target meme picture, and accurately determine the matching target reply content according to the target semantic content for reply, improving the user's interaction experience.

[0165] The embodiments of this specification also provide a server, including a processor and a memory for storing instructions executable by the processor. When specifically implemented, the processor can execute the following steps according to the instructions: obtain a target meme picture; wherein, the target meme picture includes the meme picture used by the target user in the conversation interface; the target meme picture includes a dynamic meme picture; extract multiple key frame pictures from the target picture sequence corresponding to the target meme picture according to the preset extraction rules to obtain a target picture group; process the target picture group according to the preset semantic processing rules to obtain the text semantic content and / or the image semantic content of the target meme picture; wherein, the preset semantic processing rules at least include the preset text semantic processing rules and the preset image semantic processing rules; determine the target semantic content represented by the target meme picture according to the text semantic content and / or the image semantic content of the target meme picture.

[0166] To be able to more accurately complete the above instructions, refer to Figure 8 As shown, the embodiment of this specification also provides another specific server. Among them, the server includes a network communication port 801, a processor 802, and a memory 803. The above structures are connected by internal cables so that each structure can perform specific data interactions.

[0167] Among them, the network communication port 801 can specifically be used to obtain target emoji pictures; among them, the target emoji pictures include the emoji pictures used by the target user in the conversation interface; the target emoji pictures include dynamic emoji pictures.

[0168] The processor 802 can specifically be used to extract multiple key frame pictures from the target picture sequence corresponding to the target emoji pictures according to a preset extraction rule to obtain a target picture group; process the target picture group according to a preset semantic processing rule to obtain the text semantic content and / or image semantic content of the target emoji pictures; among them, the preset semantic processing rule at least includes a preset text semantic processing rule and a preset image semantic processing rule; determine the target semantic content represented by the target emoji pictures according to the text semantic content and / or image semantic content of the target emoji pictures.

[0169] The memory 803 can specifically be used to store corresponding instruction programs.

[0170] In this embodiment, the network communication port 801 can be bound to different communication protocols, so as to send or receive different data virtual ports. For example, the network communication port can be a port responsible for web data communication, can also be a port responsible for FTP data communication, and can also be a port responsible for mail data communication. In addition, the network communication port can also be a physical communication interface or communication chip. For example, it can be a wireless mobile network communication chip, such as GSM, CDMA, etc.; it can also be a Wifi chip; it can also be a Bluetooth chip.

[0171] In this embodiment, the processor 802 can be implemented in any appropriate manner. For example, the processor can take the form of, for example, a microprocessor or a processor and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, application specific integrated circuit (ASIC), programmable logic controller, and embedded microcontroller, etc. This specification does not make a limitation.

[0172] In this embodiment, the memory 803 may include multiple levels. In a digital system, anything that can store binary data can be a memory. In an integrated circuit, a circuit with a storage function that has no physical form is also called a memory, such as RAM, FIFO, etc. In a system, a storage device with a physical form is also called a memory, such as a memory module, a TF card, etc.

[0173] An embodiment of this specification also provides a terminal device, including a processor and a memory for storing processor-executable instructions. When specifically implemented, the processor may execute the following steps according to the instructions: obtain a target meme picture; where the target meme picture includes the meme pictures used by the target user in the conversation interface; the target meme picture includes dynamic meme pictures; extract a plurality of key frame pictures from the target picture sequence corresponding to the target meme picture according to a preset extraction rule to obtain a target picture group; process the target picture group according to a preset semantic processing rule to obtain the text semantic content and / or image semantic content of the target meme picture; where the preset semantic processing rule at least includes a preset text semantic processing rule and a preset image semantic processing rule; determine the target semantic content represented by the target meme picture according to the text semantic content and / or image semantic content of the target meme picture.

[0174] An embodiment of this specification also provides a computer storage medium based on the above data display method. When the computer program instructions stored in the computer storage medium are executed, the following is realized: obtain a target meme picture; where the target meme picture includes the meme pictures used by the target user in the conversation interface; the target meme picture includes dynamic meme pictures; extract a plurality of key frame pictures from the target picture sequence corresponding to the target meme picture according to a preset extraction rule to obtain a target picture group; process the target picture group according to a preset semantic processing rule to obtain the text semantic content and / or image semantic content of the target meme picture; where the preset semantic processing rule at least includes a preset text semantic processing rule and a preset image semantic processing rule; determine the target semantic content represented by the target meme picture according to the text semantic content and / or image semantic content of the target meme picture.

[0175] In this embodiment, the storage medium includes, but is not limited to, a Random Access Memory (RAM), a Read-Only Memory (ROM), a Cache, a Hard Disk Drive (HDD), or a Memory Card. The memory can be used to store computer program instructions. The network communication unit can be set according to the standards specified by the communication protocol and is an interface for network connection communication.

[0176] In this embodiment, the functions and effects specifically implemented by the program instructions stored in the computer storage medium can be explained by comparison with other embodiments and will not be elaborated here.

[0177] This specification also provides a computer program product, including a computer program, which when executed by a processor, implements the following steps: receiving a data query request initiated by a target user through a terminal device; wherein, the data query request carries at least a data table identifier of a target data table to be queried and a user identifier of the target user; responding to the data query request, querying a blockchain according to the data table identifier of the target data table to obtain the target data table; wherein, the target data table at least includes a binding relationship set; wherein, the binding relationship set stores the binding relationship between the content data of a cell in the target data table and the user identifier; performing data hiding processing on the target data table according to the binding relationship set and the user identifier of the target user to obtain a target display table for the target user; sending the target display table to the terminal device; wherein, the terminal device is used to display the target display table to the target user.

[0178] This specification also provides another computer program product, including a computer program, which when executed by a processor, implements the following steps: obtaining a target emoji picture; wherein, the target emoji picture includes the emoji picture used by the target user in the conversation interface; the target emoji picture includes a dynamic emoji picture; extracting a plurality of key frame pictures from the target picture sequence corresponding to the target emoji picture according to a preset extraction rule to obtain a target picture group; processing the target picture group according to a preset semantic processing rule to obtain the text semantic content and / or image semantic content of the target emoji picture; wherein, the preset semantic processing rule at least includes a preset text semantic processing rule and a preset image semantic processing rule; determining the target semantic content represented by the target emoji picture according to the text semantic content and / or image semantic content of the target emoji picture.

[0179] Refer to Figure 9As shown, at the software level, the embodiments of this specification also provide a data display device, which may specifically include the following structural modules:

[0180] An acquisition module 901, which may specifically be used to acquire target meme pictures; wherein, the target meme pictures include the meme pictures used by the target user in the conversation interface; the target meme pictures include dynamic meme pictures;

[0181] An extraction module 902, which may specifically be used to extract a plurality of key-frame pictures from the target picture sequence corresponding to the target meme picture according to a preset extraction rule, to obtain a target picture group;

[0182] A processing module 903, which may specifically be used to process the target picture group according to a preset semantic processing rule, to obtain the text semantic content and / or image semantic content of the target meme picture; wherein, the preset semantic processing rule at least includes a preset text semantic processing rule and a preset image semantic processing rule;

[0183] A determination module 904, which may specifically be used to determine the target semantic content represented by the target meme picture according to the text semantic content and / or image semantic content of the target meme picture.

[0184] In some embodiments, when the above extraction module 902 is specifically implemented, it may extract a plurality of key-frame pictures from the picture sequence corresponding to the target meme picture according to the preset extraction rule in the following manner to obtain a target picture group: Analyze the target meme picture to obtain the target picture sequence corresponding to the target meme picture; Extract a first number of key-frame pictures from the first picture sequence in the target picture sequence according to the preset extraction rule; Extract a second number of key-frame pictures from the second picture sequence in the target picture sequence; wherein, the pictures in the first picture sequence are displayed earlier in the target picture sequence than the pictures in the second picture sequence; The value of the first number is less than the value of the second number; Arrange the plurality of key-frame pictures according to the display order to obtain a target picture group.

[0185] In some embodiments, when the above-mentioned processing module 903 is specifically implemented, it may process the target picture group according to the following manner based on the preset semantic processing rules to obtain the text semantic content of the target emoji picture: Detect whether there are text characters in the key frame pictures of the target picture group according to the preset text semantic processing rules; in the case of determining that there are text characters in the key frame pictures of the target picture group, determine the key frame pictures with text characters as the first type of key frame pictures; process the first type of key frame pictures by calling the preset OCR recognition model to obtain the text semantic content of the first type of key frame pictures; combine the text semantic content of multiple first type of key frame pictures according to the display order of the first type of key frame pictures in the target picture sequence to obtain the text semantic content of the target emoji picture.

[0186] In some embodiments, when the above-mentioned processing module 903 is specifically implemented, it may process the first type of key frame pictures by calling the preset OCR recognition model according to the following manner to obtain the text semantic content of the first type of key frame pictures, including: Call the preset OCR recognition model to process the first type of key frame pictures to extract the text characters in the first type of key frame pictures; call the preset character classification model to process the text characters in the first type of key frame pictures to determine the character type of the text characters; wherein, the character type includes printed characters and non-printed characters; according to the character type of the text characters, use the matching semantic recognition method to perform semantic recognition on the text characters in the first type of key frame pictures to obtain the text semantic content of the first type of key frame pictures.

[0187] In some embodiments, when the above-mentioned processing module 903 is specifically implemented, in the case of determining that the character type of the text characters is printed characters, it may also be used to perform text feature matching between the preset character template and the text characters in the first type of key frame pictures according to the preset text semantic processing rules to obtain the corresponding text matching result; according to the text matching result, determine the text semantic content of the first type of key frame pictures.

[0188] In some embodiments, when the above-mentioned processing module 903 is specifically implemented, in the case of determining that the character type of the text characters is non-printed characters, it may also be used to process the text characters in the first type of key frame pictures by using the preset text character semantic recognition model according to the preset text semantic processing rules to obtain the initial semantic recognition result; wherein, the preset text character semantic recognition model is trained by using the character sample data containing non-printed text characters; obtain the context association data of the target emoji picture in the target user's dialogue interface; adjust the initial semantic recognition result according to the context association data to obtain the text semantic content of the first type of key frame pictures.

[0189] In some embodiments, when the above-mentioned processing module 903 is specifically implemented, it may process the target picture group according to the preset semantic processing rules in the following manner to obtain the image semantic content of the target meme picture: extract the image features of the key-frame pictures according to the preset image semantic processing rules; match the image features of the key-frame pictures with the preset expression database to obtain the corresponding image matching result; wherein, the preset expression database stores a plurality of preset meme pictures and the semantic labels corresponding to the preset meme pictures; determine the semantic label of the preset meme picture that matches the image features of the key-frame pictures according to the image matching result as the image semantic content of the key-frame picture; combine the image semantic contents of multiple key-frame pictures according to the display order of the key-frame pictures in the target picture sequence to obtain the image semantic content of the target meme picture.

[0190] In some embodiments, when the above-mentioned processing module 903 is specifically implemented, after matching the image features of the key-frame pictures with the preset expression database, it can also be used to obtain the context-related data of the target meme picture in the target user's conversation interface according to the preset image semantic processing rules when it is determined that there is no preset meme picture in the preset expression database that matches the image features of the key-frame pictures; combine the image features of the key-frame pictures and the context-related data of the target meme picture to obtain the combined features; process the combined features with the preset comprehensive semantic recognition model to obtain the corresponding comprehensive recognition result; wherein, the preset comprehensive semantic recognition model is trained with comprehensive sample data including image features and context-related data; determine the image semantic content of the key-frame picture according to the comprehensive recognition result.

[0191] In some embodiments, when the above-mentioned determination module 904 is specifically implemented, it may determine the target semantic content represented by the target meme picture according to the text semantic content and / or image semantic content of the target meme picture in the following manner: detect whether the reliability of the text semantic content of the target meme picture meets the preset reliability requirements; when it is determined that the reliability of the text semantic content of the target meme picture meets the preset reliability requirements, determine the text semantic content of the target meme picture as the target semantic content represented by the target meme picture.

[0192] In some embodiments, when the above-mentioned determination module 904 is specifically implemented, after detecting whether the reliability of the text semantic content of the target emoji picture meets the preset reliability requirements, it can also be used to count the proportion of the number of key frame pictures in the target picture group whose image features match the preset emoji database when it is determined that the reliability of the text semantic content of the target emoji picture does not meet the preset reliability requirements; according to the proportion of the number of key frame pictures in the target picture group whose image features match the preset emoji database, detect whether the reliability of the image semantic content of the target emoji picture meets the preset reliability requirements; when it is determined that the reliability of the image semantic content of the target emoji picture meets the preset reliability requirements, determine the image semantic content of the target emoji picture as the target semantic content represented by the target emoji picture.

[0193] In some embodiments, when the above-mentioned determination module 904 is specifically implemented, after detecting whether the reliability of the image semantic content of the target emoji picture meets the preset reliability requirements, it can also be used to determine the target semantic content represented by the target emoji picture by combining the text semantic content and the image semantic content of the target emoji picture when it is determined that the reliability of the image semantic content of the target emoji picture does not meet the preset reliability requirements.

[0194] In some embodiments, after the above-mentioned device determines the target semantic content represented by the target emoji picture, it can also be used to determine a matching target reply text according to the target semantic content; and reply the target reply text to the target user in the dialogue interface of the target user.

[0195] In some embodiments, when the above-mentioned device is specifically implemented, it can also be used to obtain the position information of the target emoji picture in the dialogue interface of the target user; according to the position information, when it is determined that the target emoji picture is the starting conversation initiated by the target user in the dialogue interface, determine the preset greeting text as the matching target reply text.

[0196] It should be noted that the units, devices, modules, etc. described in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. For the convenience of description, when describing the above devices, they are divided into various modules according to functions and described separately. Of course, when implementing this specification, the functions of each module can be implemented in the same or multiple software and / or hardware, or the modules implementing the same function can be realized by a combination of multiple sub-modules or sub-units, etc. The device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be in electrical, mechanical or other forms.

[0197] As can be seen from the above, based on the processing device for emoji pictures provided in the embodiments of this specification, in the intelligent customer service scenario, when a target user uses a target emoji picture such as a dynamic emoji picture in the dialogue interface, multiple key frame pictures can be extracted from the target picture sequence corresponding to the target emoji picture according to a preset extraction rule to obtain a target picture group; then, according to a preset semantic processing rule that at least includes a preset text semantic processing rule and a preset image semantic processing rule, the target picture group is processed to obtain the text semantic content and / or image semantic content of the target emoji picture; then, according to the text semantic content and / or image semantic content of the target emoji picture, the target semantic content represented by the target emoji picture is determined. Thus, it can automatically and accurately identify and determine the target semantic content represented by the target user's target emoji picture, and accurately determine a matching target reply content for reply according to the target semantic content, improving the user's interaction experience.

[0198] Although the present specification provides method operation steps as described in the embodiments or flowcharts, more or fewer operation steps may be included based on conventional or non-creative means. The order of steps listed in the embodiments is only one way among the execution orders of numerous steps and does not represent the only execution order. When the actual device or client product is executed, it may be executed in the order of the method shown in the embodiments or the drawings or executed in parallel (for example, in a parallel processor or multi-threaded processing environment, or even in a distributed data processing environment). The terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, product or device comprising a series of elements not only includes those elements but also includes other elements not expressly listed, or also includes elements inherent to such process, method, product or device. Without further limitation, there is no exclusion of additional identical or equivalent elements in the process, method, product or device comprising the said elements. The terms such as "first", "second" are used to denote names and do not denote any particular order.

[0199] As is also known to those skilled in the art, in addition to implementing the controller in the form of pure computer-readable program code, the method steps can be logically programmed to enable the controller to implement the same functions in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, embedded microcontrollers, etc. Therefore, such a controller can be regarded as a hardware component, and the devices included therein for implementing various functions can also be regarded as the structures within the hardware component. Or even, the devices for implementing various functions can be regarded as either software modules for implementing the method or the structures within the hardware component.

[0200] The present specification can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, classes, etc. that perform specific tasks or implement specific abstract data types. The present specification can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0201] As can be seen from the description of the above embodiments, those skilled in the art can clearly understand that this specification can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of this specification can essentially be embodied in the form of a software product, which can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, mobile terminal, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments of this specification.

[0202] The various embodiments in this specification are described in a progressive manner. For the same or similar parts among the various embodiments, reference can be made to each other. The key point of each embodiment is to illustrate the differences from other embodiments. This specification can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on.

[0203] Although this specification is depicted through embodiments, those of ordinary skill in the art know that this specification has many variations and changes without departing from the spirit of this specification. It is hoped that the appended claims will cover these variations and changes without departing from the spirit of this specification.

Claims

1. A method for processing meme pictures, characterized in that, it includes: Obtain a target meme picture; wherein, the target meme picture includes the meme pictures used by the target user in the conversation interface in the intelligent customer service scenario; the target meme picture includes dynamic meme pictures; According to a preset extraction rule, extract multiple key frame pictures from the target picture sequence corresponding to the target meme picture to obtain a target picture group; including: parsing the target meme picture to obtain the target picture sequence of the target meme picture; according to the preset extraction rule, extract the first number of key frame pictures from the first picture sequence in the target picture sequence; extract the second number of key frame pictures from the second picture sequence in the target picture sequence; the pictures in the first picture sequence are displayed earlier in the target picture sequence than the pictures in the second picture sequence; the value of the first number is less than the value of the second number; arrange the multiple key frame pictures according to the display order to obtain the target picture group; the preset extraction rule is learned from the sample dynamic meme pictures; According to a preset semantic processing rule, process the target picture group to obtain the text semantic content and / or image semantic content of the target meme picture; wherein, the preset semantic processing rule at least includes a preset text semantic processing rule and a preset image semantic processing rule; According to the text semantic content and / or image semantic content of the target meme picture, determine the target semantic content represented by the target meme picture; including: when the reliability of the text semantic content of the target meme picture does not meet the preset reliability requirement, and the reliability of the image semantic content does not meet the preset reliability requirement, splice the text semantic content and the image semantic content of the target meme picture to obtain the spliced semantic content; call a semantic prediction model to process the spliced semantic content to determine the target semantic content.

2. The method according to claim 1, characterized in that, According to a preset semantic processing rule, process the target picture group to obtain the text semantic content of the target meme picture, including: According to a preset text semantic processing rule, detect whether there are text characters in the key frame pictures in the target picture group; In the case of determining that there are text characters in the key frame pictures in the target picture group, determine the key frame pictures with text characters as the first type of key frame pictures; Call a preset OCR recognition model to process the first type of key frame pictures to obtain the text semantic content of the first type of key frame pictures; According to the display order of the first type of key frame pictures in the target picture sequence, combine the text semantic content of the multiple first type of key frame pictures to obtain the text semantic content of the target meme picture.

3. The method according to claim 2, characterized in that, Call a preset OCR recognition model to process the first type of key frame pictures to obtain the text semantic content of the first type of key frame pictures, including: Call a preset OCR recognition model to process the first type of key frame pictures to extract the text characters in the first type of key frame pictures; Call a preset character classification model to process the text characters in the first type of key-frame images to determine the character types of the text characters; wherein, the character types include printed characters and non-printed characters; According to the character types of the text characters, use a matching semantic recognition method to perform semantic recognition on the text characters in the first type of key-frame images to obtain the text semantic content of the first type of key-frame images.

4. The method according to claim 3, wherein, When it is determined that the character type of the text character is a printed character, according to the character type of the text character, use a matching semantic recognition method to perform semantic recognition on the text characters in the first type of key-frame images to obtain the text semantic content of the first type of key-frame images, including: According to the preset text semantic processing rules, use a preset character template to perform text feature matching with the text characters in the first type of key-frame images to obtain the corresponding text matching result; According to the text matching result, determine the text semantic content of the first type of key-frame images.

5. The method according to claim 3, wherein, When it is determined that the character type of the text character is a non-printed character, according to the character type of the text character, use a matching semantic recognition method to perform semantic recognition on the text characters in the first type of key-frame images to obtain the text semantic content of the first type of key-frame images, including: According to the preset text semantic processing rules, use a preset text character semantic recognition model to process the text characters in the first type of key-frame images to obtain an initial semantic recognition result; wherein, the preset text character semantic recognition model is trained using character sample data containing non-printed text characters; Obtain the context-related data of the target meme image in the target user's dialogue interface; Adjust the initial semantic recognition result according to the context-related data to obtain the text semantic content of the first type of key-frame images.

6. The method according to claim 1, wherein, According to the preset semantic processing rules, process the target image group to obtain the image semantic content of the target meme image, including: According to the preset image semantic processing rules, extract the image features of the key-frame images; Match the image features of the key-frame images with a preset expression database to obtain the corresponding image matching result; wherein, the preset expression database stores multiple preset meme images and semantic labels corresponding to the preset meme images; According to the image matching result, determine the semantic label of the preset meme image that matches the image features of the key-frame image as the image semantic content of the key-frame image; According to the display order of the key-frame images in the target image sequence, combine the image semantic contents of multiple key-frame images to obtain the image semantic content of the target meme image.

7. The method according to claim 6, wherein, After matching the image features of the key-frame images with the preset expression database, the method further includes: In the case where it is determined, according to the image matching result, that there is no preset meme picture in the preset expression database that matches the image features of the key-frame picture, obtain the context-related data of the target meme picture in the dialogue interface of the target user according to the preset image semantic processing rules; Combine the image features of the key-frame picture and the context-related data of the target meme picture to obtain combined features; Process the combined features using a preset comprehensive semantic recognition model to obtain a corresponding comprehensive recognition result; wherein, the preset comprehensive semantic recognition model is trained using comprehensive sample data including image features and context-related data; Determine the image semantic content of the key-frame picture according to the comprehensive recognition result.

8. The method according to claim 1, wherein, determining the target semantic content represented by the target meme picture according to the text semantic content and / or the image semantic content of the target meme picture includes: detecting whether the reliability of the text semantic content of the target meme picture meets the preset reliability requirements; in the case where it is determined that the reliability of the text semantic content of the target meme picture meets the preset reliability requirements, determining the text semantic content of the target meme picture as the target semantic content represented by the target meme picture.

9. The method according to claim 8, wherein, after detecting whether the reliability of the text semantic content of the target meme picture meets the preset reliability requirements, the method further includes: in the case where it is determined that the reliability of the text semantic content of the target meme picture does not meet the preset reliability requirements, counting the proportion of the number of key-frame pictures in the target picture group whose image features match the preset expression database; detecting whether the reliability of the image semantic content of the target meme picture meets the preset reliability requirements according to the proportion of the number of key-frame pictures in the target picture group whose image features match the preset expression database; in the case where it is determined that the reliability of the image semantic content of the target meme picture meets the preset reliability requirements, determining the image semantic content of the target meme picture as the target semantic content represented by the target meme picture.

10. The method according to claim 9, wherein, after detecting whether the reliability of the image semantic content of the target meme picture meets the preset reliability requirements, the method further includes: in the case where it is determined that the reliability of the image semantic content of the target meme picture does not meet the preset reliability requirements, determining the target semantic content represented by the target meme picture by combining and using the text semantic content and the image semantic content of the target meme picture.

11. The method according to claim 1, wherein, after determining the target semantic content represented by the target meme picture, the method further includes: determining a matching target reply text according to the target semantic content; in the dialogue interface of the target user, reply the target reply text to the target user.

12. The method according to claim 11, wherein, the method further includes: Obtain the position information of the target emoji picture in the conversation interface of the target user; According to the position information, when it is determined that the target emoji picture is the starting conversation initiated by the target user in the conversation interface, determine the preset greeting text as the matching target reply text.

13. An emoji picture processing device, Characterized in that, Comprising: An acquisition module, configured to acquire a target emoji picture; wherein, the target emoji picture includes an emoji picture used by a target user in a conversation interface in an intelligent customer service scenario; the target emoji picture includes a dynamic emoji picture; An extraction module, configured to extract a plurality of key frame pictures from the target picture sequence corresponding to the target emoji picture according to a preset extraction rule, to obtain a target picture group; specifically, the extraction module is configured to: parse the target emoji picture to obtain the target picture sequence of the target emoji picture; extract a first number of key frame pictures from the first picture sequence in the target picture sequence according to the preset extraction rule; extract a second number of key frame pictures from the second picture sequence in the target picture sequence; the pictures in the first picture sequence are displayed earlier in the target picture sequence than the pictures in the second picture sequence; the value of the first number is less than the value of the second number; arrange the plurality of key frame pictures according to the display order to obtain a target picture group; the preset extraction rule is learned from a sample dynamic emoji picture; A processing module, configured to process the target picture group according to a preset semantic processing rule to obtain the text semantic content and / or image semantic content of the target emoji picture; wherein, the preset semantic processing rule at least includes a preset text semantic processing rule and a preset image semantic processing rule; A determination module, configured to determine the target semantic content represented by the target emoji picture according to the text semantic content and / or image semantic content of the target emoji picture; specifically, the determination module is configured to: when the reliability of the text semantic content of the target emoji picture does not meet the preset reliability requirement, and the reliability of the image semantic content does not meet the preset reliability requirement, splice the text semantic content and the image semantic content of the target emoji picture to obtain the spliced semantic content; call a semantic prediction model to process the spliced semantic content to determine the target semantic content.

14. A server, Characterized in that, Comprising a processor and a memory for storing processor-executable instructions, and when the processor executes the instructions, the steps of the method according to any one of claims 1 to 12 are implemented.

15. A computer program product, Characterized in that, Including a computer program, and when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 12 are implemented.

Citation Information

Patent Citations

  • An electronic document generation method and equipment

    CN109710907A

  • Multi-modal emotion analysis method for emoji package of social platform

    CN112651448A

  • Automated Video-To-Text System

    US20070273696A1