Message display method and device, electronic equipment and storage medium
By analyzing the association relationship between messages and extracting content positioning information, key content is automatically highlighted in the target message, solving the problem of low message viewing efficiency and achieving the effect of quickly identifying key content.
Patent Information
- Application Number
- CN202410029008.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-04
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, the message recipient needs to view the instant messaging message one by one when viewing the message to understand the key content, resulting in low efficiency in viewing the message.
By analyzing the association relationship between multiple messages of the message sender, the content positioning information is extracted, the key content is determined from the context message associated with the target message, and the key content is automatically highlighted when viewing the target message.
Improve message viewing efficiency, allowing message recipients to quickly identify and view key content that message senders want to emphasize.
Smart Images

Figure CN120255992A_ABST
Abstract
Description
Background Art
[0002] With the rapid development of Internet technology, electronic devices are increasingly widely used, and people's requirements for the performance of electronic devices are also getting higher and higher.
[0003] Currently, an object can communicate with other objects through an electronic device. Taking the chat scenario as an example, usually, the message sender can edit instant messaging messages, such as text messages, voice messages, picture messages, video messages, etc., in the chat interface of the chat application program in the electronic device. Then, after the instant messaging message is edited, the message sender can trigger the electronic device to send the instant messaging message it edited, and the message recipient can receive and view the instant messaging message sent by the message sender through the electronic device.
[0004] Currently, the message display function of electronic devices is relatively single, and it can only display each candidate historical message sent by the message sender one by one. The message recipient can also only view these instant messaging messages one by one. Therefore, when the message recipient wants to understand the key content in the messages sent by the message sender, they can only view them one by one by themselves, resulting in low message viewing efficiency.
[0005] Therefore, a message display strategy is needed to enable the message recipient to easily understand the key content in the messages sent by the message sender and improve the message viewing efficiency. Summary of the Invention
[0006] Embodiments of the present application provide a message display method, device, electronic device, and storage medium to improve the message viewing efficiency.
[0007] A message display method provided by an embodiment of the present application includes:
[0008] Display the chat interface;
[0009] Display the target message sent by the message sender in the chat interface;
[0010] In response to a viewing operation on the target message, highlight the key content in the target message; wherein, the key content is determined from the target message based on content location information; the content location information is information with a prompting effect extracted from at least one context message associated with the target message.
[0011] It should be emphasized that in the specific implementation of the present application, it involves chat data between objects, such as the target message and context message listed above. When the above embodiments of the present application are applied to specific products or technologies, the permission or consent of the object needs to be obtained, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.
[0012] Another message display method provided by an embodiment of the present application includes:
[0013] Obtain at least one context message associated with the target message;
[0014] Extract content location information of the key content for the target message mentioned in the at least one context message; wherein, the content location information is information with a prompting effect;
[0015] After receiving a viewing request for the target message, feed back the content location information to the client, so that the client, in response to a viewing operation for the target message, highlights the key content in the target message based on the content location information.
[0016] A message display device provided by an embodiment of the present application includes:
[0017] A first display unit for displaying a chat interface; and displaying a target message sent by a message sender in the chat interface;
[0018] A second display unit for highlighting key content in the target message in response to a viewing operation for the target message; wherein, the key content is determined from the target message based on content location information; the content location information is extracted from at least one context message associated with the target message and is information with a prompting effect.
[0019] Optionally, the target message is an audio message or a video message, and the second display unit is specifically configured to:
[0020] Use each message frame corresponding to the specified time represented by the content location information as a target frame, or use each message frame containing the key content represented by the content location information as a target frame; jump to the position corresponding to the start frame of each target frame, automatically play the target message starting from the position corresponding to the start frame, and highlight the key content in the target message.
[0021] Optionally, the second display unit is specifically configured to:
[0022] Highlight the key content in the target message; or circle out the key content in the target message;
[0023] Wherein, the key content is at least one of a target paragraph, a target keyword, and a target entity.
[0024] Optionally, if the key content is a target entity, the second display unit is specifically configured to:
[0025] Circle out the target entity in the target message according to the outline of the target entity.
[0026] Optionally, when the target message is an audio message, the second display unit is specifically configured to:
[0027] Convert the audio message into a text message, and highlight the key content in the converted text message.
[0028] Optionally, the device further includes:
[0029] An association unit, configured to present an association control at a first relevant position of the context message if it is determined that there is an association between a context message and the target message;
[0030] In response to a message association operation triggered by the association control corresponding to the context message, associate the target message with the context message, and present the target message at a second relevant position of the context message.
[0031] Optionally, if the association control is a text control describing the message type of the target message, the association unit is specifically configured to:
[0032] In response to a message association operation triggered by the association control corresponding to the context message, present a message selection interface corresponding to the message type; wherein, the message selection interface is used to present each candidate historical message that conforms to the message type and is published by the message sender in the current chat window within a set historical period;
[0033] In response to a selection operation of selecting the target message from the candidate historical messages, associate the context message with the target message.
[0034] Optionally, if the association control is a text control describing the message type of the target message, the association unit is specifically configured to:
[0035] In response to a message association operation triggered by the association control corresponding to the context message, present corresponding message selection controls at a third relevant position of each candidate historical message that conforms to the message type in the current chat window;
[0036] In response to a selection operation of selecting the target message from the current chat window based on each message selection control, associate the context message with the target message.
[0037] Optionally, the association unit is further configured to:
[0038] In response to a viewing operation on the target message triggered at the second relevant position, present the target message and highlight the key content in the target message.
[0039] Optionally, the display distance between the target message and the context message in the current chat window exceeds a preset distance, or the number of messages between the target message and the context message exceeds a preset number.
[0040] Optionally, the device further includes an analysis unit for determining the association between the context message and the target message in at least one of the following ways:
[0041] The context message contains content existing in the target message;
[0042] The context message contains at least one of time positioning information and location positioning information.
[0043] Another message display device provided by an embodiment of the present application includes:
[0044] A message acquisition unit for acquiring at least one context message associated with the target message;
[0045] An information extraction unit for extracting content positioning information of the key content for the target message mentioned in the at least one context message; wherein the content positioning information is information with a prompting effect;
[0046] A feedback unit for, after receiving a viewing request for the target message, feeding back the content positioning information to the client, so that the client, in response to a viewing operation on the target message, highlights the key content in the target message based on the content positioning information.
[0047] Optionally, the information extraction unit is specifically configured to:
[0048] If the content positioning information includes location positioning information, match the at least one context message based on a preset first regular expression, and extract the location positioning information in the at least one context message according to the matching result;
[0049] If the content positioning information includes time positioning information, match the at least one context message based on a preset second regular expression, and extract the time positioning information in the at least one context message according to the matching result;
[0050] If the content location information includes entity description information, based on a pre-trained deep learning model, extract the entity description information in the at least one context message; wherein, the deep learning model includes an entity classification layer, and the deep learning model is fine-tuned based on a historical instant messaging message dataset with pre-annotated entity categories.
[0051] Optionally, the target message is a non-long text picture message, the content location information includes entity description information, and the key content is the target entity;
[0052] Then the feedback unit is specifically configured to:
[0053] Feedback the content location information to the client, so that the client obtains a pre-constructed first entity dictionary; wherein, the first entity dictionary includes a plurality of first key-value pairs, the key in each first key-value pair is an entity category, and the value in each first key-value pair is an array of expression words associated with the entity category;
[0054] Identify the target entity category corresponding to the entity in the non-long text picture message;
[0055] Match the entity description information included in the content location information with the array of expression words associated with the target entity category in the first entity dictionary;
[0056] If the match is successful, use the corresponding entity as the target entity and highlight the target entity in the non-long text picture message.
[0057] Optionally, the device further includes a first construction unit for constructing the first entity dictionary in the following manner:
[0058] Extract the expression words related to the entity from the historical instant messaging messages;
[0059] According to the extracted expression words, construct an array of expression words associated with each entity category;
[0060] Based on the constructed array of expression words, store the index relationship between the entity category and the array of expression words in dictionary form to construct the first entity dictionary.
[0061] Optionally, the target message is a video message, the content location information includes time location information and entity description information, and the key content is the target entity;
[0062] Then the feedback unit is specifically configured to:
[0063] Feed the content location information back to the client, so that the client automatically jumps to the target frame position based on the time location information and starts playing the video message automatically from the target frame position; and
[0064] Identify the entities in the video frame corresponding to the target frame;
[0065] If there is a target entity in the identified entities that matches the entity description information included in the content location information, highlight the target entity in the target frame of the video message.
[0066] Optionally, if the target message is a video message, the device further includes an association unit for determining the context message associated with the target message in the following manner:
[0067] Obtain a second entity dictionary pre-constructed for the video message; wherein, the second entity dictionary includes at least one second key-value pair, the key in each second key-value pair is a video timestamp, and the value in each second key-value pair is the entity category of the entity in the video frame corresponding to the video timestamp.
[0068] For each context message, match the entity description information in the context message with the second entity dictionary;
[0069] If the match is successful, determine that there is an association between the video message and the context message.
[0070] Optionally, the device further includes a second construction unit for constructing the second entity dictionary in the following manner:
[0071] Process the video message frame by frame to identify the entities included in each video frame of the video message;
[0072] Based on the recognition result, when it is determined that the entities in the video frames included in the video message change, insert a new second key-value pair into the second entity dictionary, and the video timestamp represented by the key in the new key-value pair is the video timestamp corresponding to the video frame when the change occurs this time.
[0073] An electronic device provided by an embodiment of the present application includes a processor and a memory. Among them, the memory stores a computer program, and when the computer program is executed by the processor, the processor is caused to execute the steps of any of the above message display methods.
[0074] An embodiment of the present application provides a computer-readable storage medium, which includes a computer program. When the computer program runs on an electronic device, the computer program is used to cause the electronic device to execute the steps of any of the above message display methods.
[0075] An embodiment of the present application provides a computer program product, the computer program product includes a computer program, and the computer program is stored in a computer-readable storage medium; when a processor of an electronic device reads the computer program from the computer-readable storage medium, the processor executes the computer program, so that the electronic device executes the steps of any one of the above message display methods.
[0076] The beneficial effects of the present application are as follows:
[0077] An embodiment of the present application provides a message display method, device, electronic device and storage medium. In the present application, for multiple messages sent by a message sender, it is possible to determine whether there is an association between the multiple messages. In the case where the multiple messages are associated, it is possible to determine the context messages associated with the target message, extract the content location information with a prompting effect mentioned in these associated context messages, and then, when viewing the target message, based on the extracted content location information, automatically highlight in the target message the key content determined based on the content location information in at least one context message.
[0078] Since the present application can automatically analyze the association relationship between the messages sent by the same message sender, and then, based on this association relationship, extract the content location information from the context messages associated with the target message, so as to automatically mark the key content in the target message, that is, the content that the message sender wants the message recipient to focus on. Furthermore, the message recipient can focus on viewing the key content according to the highlighted mark, effectively improving the viewing efficiency of the message.
[0079] Other features and advantages of the present application will be described in the subsequent specification, and some of them will become obvious from the specification, or be understood by implementing the present application. The objectives and other advantages of the present application can be achieved and obtained through the structures specifically pointed out in the written specification, claims, and drawings. Brief Description of the Drawings
[0080] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0081] Figure 1 is an optional schematic diagram of an application scenario in an embodiment of the present application;
[0082] Figure 2 is an implementation flowchart of a message display method in an embodiment of the present application;
[0083] Figure 3Schematic diagram of automatically highlighting key content in a long text picture message in an embodiment of the present application;
[0084] Figure 4 Schematic diagram of automatically highlighting key content in a document message in an embodiment of the present application;
[0085] Figure 5 Schematic diagram of automatically highlighting key content in a link message in an embodiment of the present application;
[0086] Figure 6 Schematic diagram of automatically highlighting key content in a normal text message in an embodiment of the present application;
[0087] Figure 7 Schematic diagram of the first method of selecting key content in a picture message in an embodiment of the present application;
[0088] Figure 8 Schematic diagram of the second method of selecting key content in a picture message in an embodiment of the present application;
[0089] Figure 9 Schematic diagram of the first method of selecting key content in a video frame in an embodiment of the present application;
[0090] Figure 10 Schematic diagram of the second method of selecting key content in a video message in an embodiment of the present application;
[0091] Figure 11 Schematic diagram of highlighting key content in a voice message in an embodiment of the present application;
[0092] Figure 12 Schematic diagram of message association in an embodiment of the present application;
[0093] Figure 13 Schematic diagram of a message selection interface in an embodiment of the present application;
[0094] Figure 14 Schematic diagram of an associated picture in an embodiment of the present application;
[0095] Figure 15 As shown, it is a schematic diagram of a method for viewing associated messages in an embodiment of the present application;
[0096] Figure 16 Schematic diagram of a message selection control in an embodiment of the present application;
[0097] Figure 17 Schematic diagram of an associated text message in an embodiment of the present application;
[0098] Figure 18As shown, it is a schematic diagram of another way to view associated messages in an embodiment of the present application;
[0099] Figure 19 It is a flowchart of the implementation of another message display method in an embodiment of the present application;
[0100] Figure 20 It is an interaction logic diagram between a terminal device and a server in an embodiment of the present application;
[0101] Figure 21 It is a schematic diagram of the composition structure of a message display device in an embodiment of the present application;
[0102] Figure 22 It is a schematic diagram of the composition structure of another message display device in an embodiment of the present application;
[0103] Figure 23 It is a schematic diagram of a hardware composition structure of an electronic device applying an embodiment of the present application;
[0104] Figure 24 It is a schematic diagram of a hardware composition structure of another electronic device applying an embodiment of the present application. Specific implementation manners
[0105] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, rather than all, of the embodiments of the technical solutions of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments recorded in this application document without creative efforts shall fall within the scope of protection of the technical solutions of the present application.
[0106] Some concepts involved in the embodiments of the present application are introduced below.
[0107] Object: Refers to the person using the terminal device, such as a user.
[0108] Target message and context message: They are a set of relative concepts. In a chat scenario (whether it is a group chat or a private chat), the message sender can send multiple instant messaging messages to the message recipient continuously or discontinuously. These messages may have no association with each other, or there may be a certain association.
[0109] Specifically, the target message can be any one of multiple messages sent from the same message sender to the same message recipient. Other messages sent by the message sender to the message recipient before or after sending the target message can all be referred to as context messages of the target message. Among these context messages, those associated with the target message can be referred to as context messages associated with the target message.
[0110] Content location information: Specifically, it refers to a location hint proposed by one message for the key content in another message among two associated messages, which has a hinting effect. For example, if the two associated messages are message X1 and message X2, where the message sender hopes that the message recipient focuses on a certain content (i.e., the key content) in message X1 and hints at this key content in message X2, then the relevant hint in message X2 is content location information.
[0111] Specifically, according to the focus hinted by the content location information, whether it is a hint for the location of the content or a hint for the characteristics of the content itself, etc., the content location information in the embodiments of the present application can be divided into two major categories: location-based and descriptive information. Among them, the location-based content location information can be further divided into two categories: time location information and location location information; the descriptive information can specifically be entity description information. Specifically, the time location information refers to information describing the time when the key content appears in the target message; the location location information refers to information describing the location where the key content appears in the target message; the entity description information refers to information describing an entity, specifically, it can be the name of the entity, the characteristics of the entity (such as shape characteristics, dressing characteristics, gender characteristics), etc.
[0112] Multimodal message: It refers to messages of multiple modalities, including: text, picture, video, audio, etc. In the embodiments of the present application, the target message can be a multimodal message, such as a multi-text content message, a picture message, a video message, an audio message, etc.
[0113] Long-text picture message and non-long-text picture message: According to the amount of text content included in the picture message, the picture message can be divided into a long-text picture message and a non-long-text picture message. Among them, the long-text picture message specifically refers to a picture message whose content is mainly text content, and the non-long-text picture message specifically refers to a picture message whose content is mainly image content.
[0114] In the embodiments of the present application, the picture message is divided as follows: If the number of words contained in the picture exceeds the first threshold, the picture message is considered a long-text picture message; otherwise, if the number of words contained in the picture does not exceed the first threshold, the picture message is considered a non-long-text picture message. Or, if the proportion of the words contained in the picture in the entire content of the image exceeds the second threshold, the picture message is considered a long-text picture message; otherwise, if the proportion of the words contained in the picture in the entire content of the image does not exceed the second threshold, the picture message is considered a non-long-text picture message.
[0115] The embodiments of the present application relate to artificial intelligence (AI) and machine learning technologies, and are designed based on machine learning (ML) and speech technology in artificial intelligence.
[0116] Artificial intelligence is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, and is a theory, method, technology, and application system that can perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.
[0117] Artificial intelligence technology is a comprehensive discipline that involves a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-trained model technology, operation / interaction systems, mechatronics, etc. Among them, the pre-trained model, also known as the large model or the basic model, can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0118] The message display method in the present application involves some processing of messages, such as identifying and classifying entities in video messages, picture messages, etc., and converting speech messages into text, which will involve machine learning, natural language processing (NLP), speech technology, etc.
[0119] Machine Learning (ML) is an interdisciplinary field that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rote learning. Pre-trained models are the latest development results of deep learning, integrating the above technologies.
[0120] NLP is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers in natural language. Natural language processing involves natural language, that is, the language people use in daily life, and is closely related to linguistic research; at the same time, it involves computer science and mathematics. The pre-trained model, an important technology for model training in the field of artificial intelligence, has evolved from the large language model in the field of NLP. After fine-tuning, the large language model can be widely applied to downstream tasks. Natural language processing technologies usually include text processing, semantic understanding, machine translation, robot question answering, knowledge graph, and other technologies.
[0121] As in the embodiments of the present application, machine learning technology and NLP technology can be used to train a deep learning model for entity recognition and classification. Furthermore, the trained deep learning model can be applied to extract entity description information from messages, detect entities and entity categories in messages, and so on.
[0122] The key technologies of Speech Technology include Automatic Speech Recognition (ASR), Text To Speech (TTS), and voiceprint recognition technology. Enabling computers to listen, see, speak, and feel is the future development direction of human-computer interaction, and voice has become one of the most promising human-computer interaction methods in the future. The large model technology has brought a revolution to the development of speech technology. Pre-trained models such as WavLM and UniSpeech that follow the Transformer architecture have strong generalization and versatility and can excellently complete speech processing tasks in various directions.
[0123] As in the embodiments of the present application, the automatic speech recognition technology can be used to recognize audio messages and automatically convert the audio messages into text messages so as to highlight key content in the text messages and so on.
[0124] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in multiple fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, driverless, autonomous driving, drones, digital twins, virtual humans, robots, artificial intelligence generated content (AIGC), conversational interaction, smart healthcare, smart customer service, game AI, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.
[0125] The following briefly introduces the design concept of the embodiments of this application:
[0126] With the rapid development of Internet technology, electronic devices are more and more widely used, and people's requirements for the performance of electronic devices are also getting higher and higher.
[0127] Currently, an object can communicate with other objects through an electronic device. Taking the chat scenario as an example, usually, the message sender can edit an instant messaging message, such as a text message, a voice message, a picture message, a video message, etc., in the chat interface of a chat application program (such as instant messaging software) in the electronic device. Then, after the instant messaging message is edited, the message sender can trigger the electronic device to send the edited instant messaging message, and the message recipient can receive and view the instant messaging message sent by the message sender through the electronic device.
[0128] Specifically, instant messaging (IM) software is an application program that transmits instant messages between objects, which can facilitate the communication between objects. For example, objects usually communicate between one client and one client or between one client and multiple clients.
[0129] In related technologies, in order to facilitate the communication between objects and quickly reply to the questions raised by the other party, some current instant messaging software supports recommending reply content (such as map location, photos, etc.) according to the chat content. In this way, one party can quickly reply to the other party according to the recommended reply content by automatically clicking the reply.
[0130] For example, according to the question "Where are you", multiple picture replies are recommended, and the object can select one picture from the multiple recommended replies to directly reply; or for example, according to the question "Where are you", multiple location methods are recommended, and the object can select one location method from the multiple recommended replies to directly send the location for reply, and so on.
[0131] However, the above method only improves the message reply efficiency of the object to a certain extent. When the message sender sends multiple instant messaging messages, when the message receiver wants to know the key content in the messages sent by the message sender, they can only view them one by one, and cannot timely understand the key content in these messages, resulting in a low message viewing efficiency.
[0132] In view of this, the embodiments of the present application propose a message display method, device, electronic device and storage medium. In the present application, for multiple messages sent by the message sender, it is possible to determine whether there is an association between the multiple messages. When there is an association between the multiple messages, the context messages associated with the target message can be determined, and the content location information with a prompting effect mentioned in these associated context messages can be extracted. Furthermore, when viewing the target message, based on the extracted content location information, the key content determined based on the content location information in at least one context message can be automatically highlighted in the target message.
[0133] Since the present application can automatically analyze the association relationship between the messages sent by the same message sender, and then, based on this association relationship, extract the content location information from the context messages associated with the target message, so as to automatically mark the key content in the target message based on this, that is, the content that the message sender wants the message receiver to focus on. Furthermore, the message receiver can focus on viewing the key content according to the highlighted mark, effectively improving the message viewing efficiency.
[0134] The following describes the preferred embodiments of the present application with reference to the accompanying drawings of the specification. It should be understood that the preferred embodiments described herein are only used to illustrate and explain the present application, and are not used to limit the present application. And without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.
[0135] As Figure 1 shown, it is a schematic diagram of the application scenario of the embodiments of the present application. The application scenario diagram includes two terminal devices 110 and a server 120.
[0136] In the embodiments of the present application, the terminal device 110 includes, but is not limited to, devices such as mobile phones, tablet computers, laptop computers, desktop computers, e - book readers, intelligent voice interaction devices, intelligent home appliances, vehicle terminals, etc.; a client related to message display can be installed on the terminal device. The client can be software (such as a browser, instant messaging software, short - video software, etc.), or a web page, a small program, etc. The server 120 is the background server corresponding to the software, web page, small program, etc., or a server specifically used for message display. The present application does not make specific limitations. The server 120 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.
[0137] It should be noted that the message display method in each embodiment of the present application can be executed by an electronic device, and the electronic device can be the terminal device 110 or the server 120. That is, the method can be executed independently by the terminal device 110 or the server 120, or jointly executed by the terminal device 110 and the server 120. For example, when jointly executed by the terminal device 110 and the server 120, an instant messaging software can be installed on the terminal device 110. Communication can be carried out between the message sender A and the message receiver B based on the instant messaging software installed on their respective terminal devices 110. For example, the message sender A sends multiple messages to the message receiver B, including a target message and at least one context message associated with the target message; the message receiver B can view the chat interface between the two through the instant messaging software on the corresponding terminal device 110, and at least the target message can be displayed in the chat interface. Furthermore, when the message receiver B wants to view the specific content of the target message, the corresponding instant messaging software will send a viewing request for the target message to the server 120 through the terminal device 110. The server 120 can feedback the content location information extracted based on the context message associated with the target message to the instant messaging software logged in by the message receiver B through the terminal device 110, and then highlight the key content in the target message based on the content location information.
[0138] In an alternative embodiment, the terminal device 110 and the server 120 can communicate through a communication network.
[0139] In an alternative embodiment, the communication network is a wired network or a wireless network.
[0140] It should be noted that Figure 1The above is only an example, and actually the number of terminal devices and servers is not limited, and no specific limitation is made in the embodiments of the present application.
[0141] In the embodiments of the present application, when the number of servers is multiple, the multiple servers can form a blockchain, and the servers are nodes on the blockchain; as the message display method disclosed in the embodiments of the present application, the message data involved therein can be saved on the blockchain, for example, target messages, context messages associated with the target messages, key content, content location information, etc.
[0142] In addition, the embodiments of the present application can be applied to various scenarios, including but not limited to scenarios such as cloud technology, artificial intelligence, intelligent transportation, and assisted driving.
[0143] It should be emphasized that in the specific implementation manners of the present application, for relevant data such as communication messages in the chat scenario, when the above embodiments of the present application are applied to specific products or technologies, user permission or consent is required, and the collection, use, and processing of relevant data need to comply with relevant laws, regulations, and standards of relevant countries and regions.
[0144] Next, in combination with the above-described application scenarios, the message display method provided by the exemplary embodiments of the present application will be described with reference to the accompanying drawings. It should be noted that the above application scenarios are only shown for the convenience of understanding the spirit and principle of the present application, and the embodiments of the present application are not limited in this regard.
[0145] Refer to Figure 2 As shown, it is a flowchart of the implementation of a message display method provided by the embodiments of the present application. Taking the client as the execution subject as an example, the specific implementation process of the method is as follows S21 to S23:
[0146] S21: The client displays a chat interface.
[0147] S22: The client displays the target message sent by the message sender in the chat interface.
[0148] In the embodiments of the present application, the chat interface can specifically be a private chat interface between two objects or a group chat interface between multiple objects, and no specific limitation is made herein. The messages in the embodiments of the present application can be of any message type, such as text, pictures, videos, audios, etc.
[0149] In the private chat scenario between two objects, there is a message sender and a message receiver. The message sender can send multiple instant messaging messages to the message receiver continuously or discontinuously, and these messages may have no association or may have a certain association.
[0150] In the above private chat scenario, specifically, the target message can be any one of multiple messages sent by the same message sender to the same message recipient. Other messages sent by the message sender to the message recipient before or after sending the target message can all be referred to as the context messages of the target message. Among these context messages, those associated with the target message can be referred to as the context messages associated with the target message. Specifically, the method of analyzing whether two messages are associated can be found in the following embodiments and will not be elaborated here.
[0151] Specifically, in addition to displaying the target message, the chat interface can also display one or more context messages associated with the target message.
[0152] Generally, the target message refers to the type of message that contains the key content (i.e., the crucial content) that the message sender hopes the message recipient will view; the target message can be a multi-modal message, such as a multi-text content message, a picture message, a video message, an audio message, etc.
[0153] In the embodiments of the present application, a multi-text content message refers to a message that contains a lot of text content, such as a document, an external link, a long text picture, etc. Among them, the external link in this article specifically refers to a link containing text content such as an article; a long text picture refers to a picture that contains a lot of text content.
[0154] Specifically, the "a lot of text content" in the embodiments of the present application can specifically be that the number of words contained in the message exceeds a preset threshold (which can be flexibly set according to experience); or the proportion of the words contained in the message exceeds a preset threshold (which can be flexibly set according to experience), etc.; specifically, there can also be other determination methods, which will not be elaborated one by one here.
[0155] In the embodiments of the present application, for a picture message, according to the amount of text content contained in the picture message, the picture message can be divided into the following two types:
[0156] For example, if the number of words contained in the picture exceeds the first threshold, it is considered that the picture message is a long text picture message; conversely, if the number of words contained in the picture does not exceed the first threshold, it is considered that the picture message is a non-long text picture message.
[0157] Among them, the above first threshold can be flexibly set according to actual needs, such as 500, 1000, etc.
[0158] Or, if the proportion of the words contained in the picture in the entire content of the image exceeds the second threshold, it is considered that the picture message is a long text picture message; conversely, if the proportion of the words contained in the picture in the entire content of the image does not exceed the second threshold, it is considered that the picture message is a non-long text picture message.
[0159] Among them, the above-mentioned second threshold can also be flexibly set according to actual needs, such as 50%, 80%, etc.
[0160] It should be noted that the above two methods of dividing long text picture messages and non-long text picture messages are only simple examples. Any method of dividing long text picture messages and non-long text picture messages according to the amount of text content contained in the picture message is applicable to the embodiments of the present application, and will not be elaborated one by one here.
[0161] Similar to the above private chat scenario, in the group chat scenario among multiple objects, there is at least one message sender and multiple message receivers. For the same message sender, the message sender can send multiple instant messaging messages to each message receiver in the group continuously or discontinuously. These messages may have no association with each other, or there may be a certain association.
[0162] In the group chat scenario of the embodiments of the present application, each message receiver can view the target message in multiple messages sent by the same message sender or different message senders through the message display method in the embodiments of the present application, and highlight the key content in the target message, provided that there is at least one context message associated with the target message among these multiple messages, and content positioning information for positioning the key content in the target message can be extracted from these context messages.
[0163] It should be emphasized that in the group chat scenario, the target message and the context message associated with the target message can come from the same message sender or different message senders. Specifically, when coming from different message senders, it can be that some messages come from the same message sender, some messages come from different message senders, or all come from different message senders, etc., which are not specifically limited in this article. For example, the target message and a context message associated with the target message are sent by administrator A, and then administrator B supplements and sends another context message associated with the target message, and so on.
[0164] For the sake of simple illustration below, it is taken as an example that the target message and its associated context message come from the same message sender (of course, if the target message and its associated context message come from different message senders, the specific processing methods are the same or basically similar, and will not be repeated here):
[0165] In the case where the target message is generally a multimodal message, if the same message sender sends multiple messages to the same message receiver, then the multimodal messages among these multiple messages can be used as the target message, and the messages sent within a preset time before or after the target message, or several messages before or after the target message, among these multiple messages, can be used as the context message of the target message.
[0166] Optionally, the association between the context message of the target message and the target message can be determined by at least one of the following methods:
[0167] Method 1: The context message contains the content existing in the target message.
[0168] Specifically, if two messages sent by the same message sender to the same message recipient contain the same content, it can be considered that there is an association between these two messages. Specifically, if there are words, image descriptions, etc. existing in the target message in the context message, it can be considered that the context message contains the content existing in the target message.
[0169] For example, if a message is "She also mentioned Person A", and the keyword "Person A" appears in the picture message (i.e., the target message) above this message, it can be determined that "She also mentioned Person A" is associated with the picture message above it; another example, if a message is "Do you see the one wearing black-framed glasses?", and there is a person wearing black-framed glasses in the video message above this message, it can be determined that "Do you see the one wearing black-framed glasses?" is associated with the video message above it; and so on.
[0170] It should be noted that in this method, among two associated messages, generally the one with more content is the "target message", and the one with less content is the "context message associated with the target message".
[0171] Method 2: The context message contains at least one of the time positioning information and the location positioning information.
[0172] Among them, the time positioning information and the location positioning information belong to the positioning-type content positioning information in this article. Therefore, if the context message contains "positioning-type" descriptions, it can be considered that there is an association between this context message and the target message.
[0173] For example, there are "positioning-type" words related to location in the context message, such as "the first paragraph", "the first row", "xx part", etc.
[0174] Another example, there are "positioning-type" words related to time in the context message, such as "the 4th second", "the last minute", "xx moment", "xx time period", etc.
[0175] It should be noted that in this method, among two associated messages, generally the one containing the above time positioning information or location positioning information is the "context message", and before or after this context message, the one related to the time positioning or location positioning is the "target message associated with this context message".
[0176] It should be emphasized here that the acquisition of data such as the above-mentioned target messages and context messages is with the permission or consent of the user, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions.
[0177] In the embodiments of the present application, the association relationship between messages can be determined in advance in the above manner. In this way, when the message recipient wants to view the target message, the context message associated with the target message can be automatically retrieved to locate the key content in the target message, facilitating the message recipient to quickly understand the key content emphasized by the message sender and improving the message reading efficiency.
[0178] It should be noted that the above-listed several ways to analyze whether there is a relationship between two messages are just simple examples. In addition, other ways are also applicable to the embodiments of the present application and will not be elaborated one by one here.
[0179] In addition, it should be noted that the publication times of two associated messages are generally within a certain preset time range, such as one minute, one hour, one day, etc. Specifically, it can be flexibly set according to actual needs, the chat frequency between the message sender and the message recipient, etc. And, there is a certain publication order between these two associated messages. Generally, the target message is in the front and the context message is in the back; of course, it can also be that the context message is in the front and the target message is in the back, and they belong to the context order.
[0180] S23: The client, in response to the viewing operation for the target message, highlights the key content in the target message; wherein, the key content is determined from the target message based on the content location information; the content location information is information with a prompting effect extracted from at least one context message associated with the target message.
[0181] Specifically, the prompting effect in this article specifically refers to: prompting the key content in the target message, which can specifically be prompting the position where the content appears, or prompting the characteristics of the content itself, etc.
[0182] According to the above content, the content location information in the embodiments of the present application can be divided into two major categories: positioning type and descriptive type information. Among them, the positioning type of content location information can be further divided into two categories: time positioning information and location positioning information; the descriptive type information can specifically be entity description information. Specifically, the time positioning information refers to information describing the time when the key content appears in the target message; the location positioning information refers to information describing the position where the key content appears in the target message; the entity description information refers to information describing an entity, which can specifically be the name of the entity, the characteristics of the entity (such as shape characteristics, dressing characteristics, gender characteristics), etc.
[0183] In the embodiments of the present application, there are many ways to highlight key content in a message, such as highlighting, magnifying, bolding, underlining, circling, marking with different colors, etc.
[0184] Optionally, one implementation manner of S23 is as follows:
[0185] Highlight the key content in the target message; or, circle the key content in the target message;
[0186] Wherein, the key content may be at least one of a target paragraph, a target keyword, and a target entity.
[0187] Wherein, the target paragraph may specifically be any one or a combination of multiple paragraphs, a paragraph fragment, multiple paragraph fragments, etc. Among them, the paragraph fragment may specifically be a part of a paragraph, such as one sentence, multiple sentences, etc. in a paragraph.
[0188] The target keyword may specifically be a person's name, a place name, a proper noun, etc.; the target entity may specifically be a living thing (such as a person, an animal, a plant, etc.), a non-living thing (such as a building, a table, a chair, etc.).
[0189] The following will respectively give examples to illustrate the methods of highlighting key content in different types of target messages (multimodal messages):
[0190] First, an example will be given to illustrate the method of highlighting key content in a multi-text content message according to the context message:
[0191] Specifically, when the target message is a multi-text content message such as a document, a long text picture, or an external link, the key content is also text content, such as a target paragraph or a target keyword.
[0192] An optional implementation manner is that when the target message is a multi-text content message, the key content can be highlighted in the multi-text content message, where the key content is at least one of a target paragraph and a target keyword.
[0193] Specifically, for a multi-text content message (including a long text picture, a document, an external link, etc.), the system intelligently locates the key content in the multi-text content message according to the associated context message sent by the message sender A, and automatically highlights the key key content when the message receiver B clicks on the multi-text content message.
[0194] Taking the multi-text content message as a long text picture message as an example, as Figure 3 shown, it is a schematic diagram of automatically highlighting key content in a long text picture message in the embodiments of the present application.
[0195] Specifically, Figure 3What is shown represents the chat window that message recipient B can view, which is the chat interface with message sender A. In this chat window, there is a message sent by message recipient B: "…", and there are also 4 consecutive messages sent by message sender A. One of them is a long text with a picture message, and the other three are text messages, such as "The first paragraph is written very well", "She also mentioned character A", "…".
[0196] Through the above-listed relevance analysis method, it can be known that message S302: "The first paragraph is written very well" contains the location positioning information "the first paragraph", and message S303: "She also mentioned character A" contains the same entity word "character A" as the long text with a picture message. Moreover, considering the publishing time, publishing order, etc. of these messages, it can be known that the long text with a picture message S301 is related to message S302: "The first paragraph is written very well" and message S303: "She also mentioned character A".
[0197] Specifically, the content positioning information included in these two related context messages includes paragraph location positioning information, such as "the first paragraph", and entity description information, such as "character A". Based on this content positioning information, when message recipient B clicks to view the long text with a picture message, the picture shown on the right side can be presented Figure 3 and the first paragraph in the picture, such as S304, and the entity word "character A" that appears in the picture, such as S305, will be highlighted.
[0198] Taking the multi-text content message as a document message as an example, as Figure 4 shown, it is a schematic diagram of automatically highlighting key content in a document message in an embodiment of the present application.
[0199] Specifically, Figure 4 What is shown represents the chat window that message recipient B can view, which is the chat interface with message sender A. In this chat window, there is a message sent by message recipient B: "…", and there are also 4 consecutive messages sent by message sender A. One of them is a document message, which is a PDF document related to a visa application form, and the other three are text messages, such as "Pay attention to the passport information section", "…", "…".
[0200] Through the above-listed relevance analysis method, it can be known that message S402: "Pay attention to the passport information section" contains the location positioning information "the passport information section", and considering the publishing time, publishing order, etc. of these messages, it can be known that the document message S401 is related to message S402: "Pay attention to the passport information section".
[0201] Specifically, the content positioning information contained in this associated context message includes the location positioning information of the paragraph, such as "passport information section". Based on this content positioning information, when the message recipient B clicks to view the document message, the Figure 4 The visa application form shown on the right in the figure and highlight the passport information section of the document, such as S403.
[0202] For example, take a multi-text content message as an external link message, such as Figure 5 As shown, it is a schematic diagram of automatically highlighting key content in an external link message in an embodiment of the present application.
[0203] Specifically, Figure 5 What is shown is the chat window between the message receiver B and the message sender A, in which a message "..." is displayed on the chat interface sent by the message receiver B, and 4 messages sent consecutively by the message sender A are displayed, one of which is a link message, which is a link to a post, and the other three are text messages, such as "The little * book is too funny", "Look at the part about avoiding thunderbolts", and "...".
[0204] Through the correlation analysis method listed above, it can be known that message S502: "The little * book is too funny" contains the entity description information "little * book", and message S503: "Look at the part about avoiding thunderbolts" contains the location positioning information "the part about avoiding thunderbolts", and combined with the release time, release order, etc. of these messages, it can be known that the linked message S501 is associated with message S502: "The little * book is too funny" and message S503: "Look at the part about avoiding thunderbolts".
[0205] Specifically, the content positioning information contained in the two related context messages includes the location positioning information of the paragraph, such as "the paragraph about avoiding thunderbolts" and the entity description information "small book". Based on this content positioning information, when the message recipient B clicks to view the link message, it can be presented Figure 5 The main text shown on the right side of the link is highlighted, and the entity word "小*书" in the link is highlighted, such as S504, S505, S506, and the part about "避雷贴那段", such as S507.
[0206] In addition to the above-mentioned types of multi-text content messages, the target message may also be a common text message containing a lot of text content.
[0207] like Figure 6 As shown, it is a schematic diagram of automatically highlighting key content in a common text message in an embodiment of the present application.
[0208] Specifically, Figure 6What is shown represents the chat window that message recipient B can view, which is the chat interface with message sender A. In this chat window, three consecutive messages sent by message sender A are presented, all of which are ordinary text messages, such as "Here is an event notice", "Event notice: To welcome the arrival of the xx anniversary and further promote the xx spirit, xx specially holds a series of xx activities. The details are as follows: Tonight (June 25th) at 7:30, an outdoor movie screening will be held at xx location. The movie to be screened is 《x movie》. Welcome everyone to actively participate~", "Pay attention to the time and don't be late", and there is also a text message "..." sent by message recipient B.
[0209] Through the above-listed correlation analysis method, it can be known that message S603: "Pay attention to the time and don't be late" contains the entity word (which is a kind of entity description information) "time", and combined with the release time, release order, etc. of these messages, it can be known that text message S603 is related to message S602: "Here is an event notice", "Event notice: To welcome the arrival of the xx anniversary and further promote the xx spirit, xx specially holds a series of xx activities. The details are as follows: Tonight (June 25th) at 7:30, an outdoor movie screening will be held at xx location. The movie to be screened is 《x movie》. Welcome everyone to actively participate~".
[0210] Specifically, the content location information included in this associated context message has the entity description information "time". Based on this content location information, when message recipient B clicks to view this text message, etc., the interface shown on the right side can be presented Figure 6 as shown on the right side, and the time information "Tonight (June 25th) at 7:30" in this text message is highlighted, such as S604.
[0211] In addition, for ordinary text messages, it is not necessary to click to view. When message recipient B opens the chat window with message sender A, the interface shown on the right side is automatically presented Figure 6 as shown on the right side; or when message recipient B browses to this text message S602, the time information in this text message is automatically highlighted, etc.
[0212] It should be noted that the above is an example with the target message being a multi-text content message. In addition, the target message can also be a picture message, a video message, an audio message, etc.
[0213] An optional implementation method is that when the target message is a picture message, according to the amount of text content included in the picture message, the picture message can be divided into a long-text picture message and a non-long-text picture message. The specific division method can refer to the above embodiments, and the repeated parts will not be elaborated.
[0214] Specifically, for long text picture messages (which belong to multi-text content messages), the specific implementation can refer to the above embodiments. For example, Figure 3 as shown, the repeated parts will not be elaborated again.
[0215] For non-long text picture messages, if the picture message does not contain text, the key content it contains should be the target entity; if the picture message contains a small amount of text, the key content it contains can be the target entity or the target keyword.
[0216] Taking the key content as the target entity as an example, an optional implementation is to circle the key content in the picture message.
[0217] In the embodiments of the present application, regular graphics or irregular graphics can be used to circle the target entity. Specifically, a part of the target entity or the entire target entity can be circled.
[0218] Taking the example of using a circle to circle, as Figure 7 shown, it is the first schematic diagram of circling the key content in the picture message in the embodiments of the present application.
[0219] Specifically, Figure 7 shown represents the chat window that the message recipient B can view, and in the chat interface displayed in this chat window, there are two messages sent by the message recipient B, namely "……", "……", and there are also 4 consecutive messages sent by the message sender A, one of which is a picture message, which is a class graduation photo as shown in message S701, and the other three are text messages, such as message S702: "Unexpectedly, I saw our class graduation photo", message S703: "I am the sixth from the right in the second row", "……".
[0220] Through the above-listed correlation analysis method, it can be known that message S703: "I am the sixth from the right in the second row" contains the location positioning information "the sixth from the right in the second row", and combined with the release time, release order, etc. of these messages, it can be known that there is a correlation between the picture message S701 and message S703: "I am the sixth from the right in the second row".
[0221] Specifically, the content positioning information contained in this associated context message has the location positioning information of the target entity, such as "the sixth from the right in the second row". Based on this content positioning information, when the message recipient B clicks to view this picture message, the picture shown on the right side of Figure 7 will be presented, and the target entity (i.e., the message sender A in this class photo) will be circled, as shown in S704. Among them, S704 takes the example of circling through a circle.
[0222] It should be noted that Figure 7The circled selection method shown is just a simple example. In addition, other regular graphics can be used for circled selection, such as rectangles. Of course, irregular graphics can also be used for circled selection.
[0223] Optionally, an implementation manner of circled selection according to an irregular graphic is: circle the target entity in the target message according to the contour of the target entity.
[0224] Such as Figure 8 shown, which is the second schematic diagram of circled selection of key content in a picture message in the embodiments of the present application.
[0225] Specifically, Figure 8 shown represents the chat window that the message recipient B can view, which shows two messages sent by the message recipient B, namely "..." and "...", in the chat interface displayed in this chat window. There are also 4 consecutive messages sent by the message sender A, one of which is a picture message, such as the message S801. This picture is a screenshot of a TV drama, and the other three are text messages, such as the message S802: "Sister, look quickly", the message S803: "The heroine of this drama is so beautiful", and "...".
[0226] Through the above-listed correlation analysis method, it can be known that the message S803: "The heroine of this drama is so beautiful" contains the entity description information "heroine", and combined with the release time, release order, etc. of these messages, it can be known that the picture message S801 is related to the message S803: "The heroine of this drama is so beautiful".
[0227] Specifically, the content location information included in this associated context message has the entity description information of the target entity (specifically, the feature description vocabulary), such as "heroine". Based on this content location information, when the message recipient B clicks to view this picture message, the picture shown on the right side of Figure 8 can be presented, and the target entity (that is, the heroine in this TV drama screenshot) is circled, such as S804. Among them, S804 is circled according to the contour of the heroine as an example.
[0228] It should be noted that Figure 8 the circled selection method shown is just a simple example. In addition, other irregular graphics can be used for circled selection, and no specific limitation is made herein.
[0229] In the above implementation manner, the key content in the target message can be automatically marked with emphasis through the circled selection method. Based on this idea, the present application can ensure that when the object views the message, all key information will be highlighted, saving a lot of time and effort for the object. The object does not need to view each message one by one, and can find the most critical information at a glance.
[0230] It should be noted that the above takes the target message as a picture message as an example. In addition, the target message can also be a video message, an audio message, etc.
[0231] When the target message is an audio message or a video message, an optional implementation manner of S23 is as follows:
[0232] S231: Use each message frame corresponding to the specified time represented by the content location information as the target frame, or use each message frame containing the key content represented by the content location information as the target frame.
[0233] Among them, when the content location information includes time location information, the specified time represented by the content location information can be determined based on the time location information. Specifically, the specified time can be a moment or a time period. Therefore, the target frame determined based on the time location can be one frame or multiple frames.
[0234] In addition, in addition to the method of positioning according to time location, it can also be positioned according to the content itself. For example, if the key content represented by the content location information is an entity, then each message frame containing the entity can be used as the target frame. Since the message frames containing the entity may be one frame or multiple frames, the target frame determined based on the content itself can be one frame or multiple frames.
[0235] S232: Jump to the position corresponding to the starting frame among each target frame, automatically play the target message starting from the position corresponding to the starting frame, and highlight the key content in the target message.
[0236] Among them, the starting frame among each target frame refers to the message frame with the smallest corresponding timestamp among each target frame.
[0237] Specifically, if there are multiple target frames and they are consecutive message frames, when automatically playing the target message starting from the position corresponding to the starting frame, only these multiple consecutive message frames can be played.
[0238] If there are multiple target frames and they are non - consecutive message frames, when automatically playing the target message starting from the position corresponding to the starting frame, only these multiple consecutive message frames can be played, or the starting frame and all subsequent message frames can be played starting from the starting frame.
[0239] If there is only one target frame and this target frame is the starting frame, when automatically playing the target message starting from the position corresponding to the starting frame, the target frame and all subsequent message frames can be played starting from this target frame.
[0240] In addition, when highlighting key content in the target message, the key content can be highlighted only in the starting frame of the playback, or in each message frame containing the key content during the playback.
[0241] It should be noted that the above-listed ways of playing the target message and highlighting the key content are only simple examples. In addition, other ways of "automatically playing the target message starting from the position corresponding to the starting frame and highlighting the key content in the target message" are also applicable to the embodiments of this application, and will not be elaborated one by one here.
[0242] Taking a video message as an example, the message display method in the embodiments of this application will be further described below:
[0243] In the embodiments of this application, the video message can be a short video, a long video, etc., and no specific limitation is made herein.
[0244] Optionally, for a video message, it can be understood as being composed of many frames of video pictures, and each frame of video picture can be understood as a picture. Therefore, the way of highlighting key content in the video picture can refer to the way of highlighting key content in the picture message listed above. Specifically, the video picture can also be divided in the way of the amount of text content included as listed above, or it can not be divided; and when highlighting the target entity in the video picture, regular graphics or irregular graphics can be used, etc. For details, refer to the above embodiments, and the repeated parts will not be elaborated.
[0245] For example, for a video message, the system intelligently locates the target object and time node in the video according to the context message with a prompting function sent by message sender A, and automatically circles out the target object and automatically marks the time node when message recipient B clicks on the message.
[0246] As Figure 9 shown, it is a schematic diagram of the first way to circle key content in the video picture in the embodiments of this application.
[0247] Specifically, Figure 9 shown represents the chat window that message recipient B can view, which shows two messages sent by message recipient B, namely "……", "……", in the chat interface displayed in the chat window. There are also 4 consecutive messages sent by message sender A, one of which is a video message, which is a short video, and the other three are text messages, such as "At the 4th second", "Do you see the one wearing black-framed glasses?", "……".
[0248] From the above-listed relevance analysis method, it can be seen that message S902: "At the 4th second" contains the time positioning information "the 4th second"; message S903: "Do you see the one wearing black-framed glasses?" contains the characteristic vocabulary description "wearing black-framed glasses" of the entity. And by combining the sending time, sending order, etc. of these messages, it can be known that the picture message S901 is related to message S902: "At the 4th second" and message S903: "Do you see the one wearing black-framed glasses?".
[0249] Specifically, the content positioning information contained in these two related context messages includes the time information of the target entity's appearance, such as "the 4th second", and the descriptive vocabulary information of the target entity, such as "wearing black-framed glasses". Based on this content positioning information, when the message recipient B clicks to view this video message, it can automatically jump to the 4th second, and at this time, the target entity (the person wearing black-framed glasses) is circled in the video picture, such as S904. Among them, S904 takes the example of circling with a regular graphic of a circle.
[0250] In addition, it should be noted that this time point of "the 4th second" can also be highlighted in the video progress bar, such as enlarged and highlighted, as shown in S905.
[0251] As Figure 10 shown, it is the second schematic diagram of selecting key content in a video message in the embodiment of the present application.
[0252] Specifically, Figure 10 shown represents the chat window that the message recipient B can view, which is the chat interface with the message sender A. In this chat interface, two messages sent by the message recipient B, namely "..." and "...", are presented, and three consecutive messages sent by the message sender A are also presented. One of them is a video message, which is a short video, and the other two are text messages, such as "At the 4th and 6th seconds, do you see the one wearing black-framed glasses?" and "...".
[0253] From the above-listed relevance analysis method, it can be seen that message S1002: "At the 4th and 6th seconds, do you see the one wearing black-framed glasses?" contains both the time positioning information "the 4th and 6th seconds" and the characteristic description vocabulary "wearing black-framed glasses" of the entity. And by combining the sending time, sending order, etc. of these messages, it can be known that the video message S1001 is related to message S1002: "At the 4th and 6th seconds, do you see the one wearing black-framed glasses?".
[0254] Specifically, the content location information included in this associated context message includes the time information of the target entity's appearance, such as "the 4th second", "the 6th second", and the descriptive vocabulary information of the target entity, such as "wearing black-framed glasses". Based on this content location information, when the message recipient B clicks to view this video message, it can automatically jump to the 4th second, and at this time, the target entity (the person wearing black-framed glasses) is circled in the video picture, such as S1003. Among them, S1003 is an example of circling according to the outline of this person.
[0255] In addition, it should be noted that these two time points, "the 4th second" and "the 6th second", can also be highlighted in the video progress bar, such as enlarged and highlighted, as shown in S1004 and S1005.
[0256] It should be noted that for picture messages or video messages, the above is an example of highlighting one target entity; of course, if multiple target entities are mentioned in the context message associated with the target message, then these multiple target entities need to be highlighted separately in the target message. Specifically, the highlighting method for each target entity can refer to the above embodiments and will not be repeated here.
[0257] In the above embodiment, in addition to automatically highlighting the key content in the target message, it can also automatically jump to the part containing the key content. The object not only does not need to view each message one by one, but also this automatic jump function enables the object to view the key information in the first time, further saving the time and energy of the object to browse the target message.
[0258] Next, taking the audio message as an example, the message display method in the embodiments of the present application will be further described:
[0259] In the embodiments of the present application, the audio message can be a voice message, an audio file (such as a recording, a song), etc., and will not be specifically limited herein.
[0260] Optionally, when the target message is an audio message, the audio message can be converted into a text message, and the key content is highlighted in the converted text message.
[0261] Specifically, for a voice message, similar to the above-listed video message, it can automatically jump to the determined start frame to start playing, and can automatically convert the voice message into text (or can also be converted into text after being triggered by the user), and then highlight the key content (such as the target keyword) in the converted text message.
[0262] As Figure 11 shown, it is a schematic diagram of highlighting key content in a voice message in the embodiments of the present application.
[0263] Specifically,Figure 11 What is shown is the chat window that the message recipient B can view, which is the chat interface with two messages sent by the message recipient B, namely "..." and "...", and three consecutive messages sent by the message sender A. One of them is a voice message with a duration of 59 seconds, and the other two are text messages, such as "A brief introduction to Course x starts at the 20th second. You can learn about it first." and "...".
[0264] Through the above-listed correlation analysis method, it can be known that the message S1102: "A brief introduction to Course x starts at the 20th second. You can learn about it first." contains both the time positioning information "the 20th second" and the entity description information "Course x". Combining the publishing time, publishing order, etc. of these messages, it can be known that there is a correlation between the voice message S1101 and the message S1102: "A brief introduction to Course x starts at the 20th second. You can learn about it first."
[0265] Specifically, the content positioning information contained in this associated context message includes the time information of the target entity's appearance, such as "the 20th second", and the descriptive vocabulary information of the target entity, such as "Course x". Based on this content positioning information, when the message recipient B clicks to view this voice message, it can automatically jump to the 20th second, and automatically convert this voice message into a text message. In this text message, the target entity (Course x) and the text after the 20th second, such as S1104, are highlighted. In addition, the part after the 20th second can also be marked in the voice bar of the voice message, as shown in S1103.
[0266] It should be noted that the above is an example of a voice message. In addition, other types of audio messages follow the same principle, and can all be automatically jumped to and automatically converted to text for marking, etc., which will not be elaborated one by one here.
[0267] In the above embodiment, in addition to automatically jumping to the part containing the key content, the audio message can also be automatically converted into a text message so as to mark the key content in the text message, enabling the object to quickly hear the key information while directly browsing the converted text content, providing the object with multiple ways to quickly understand the key information. The object not only does not need to view each message one by one, but on the basis of seeing the key information in the first time, it also enriches the object's way of viewing information, saves the object's time and energy for browsing the target message, and improves the object's message reading experience.
[0268] In addition, it should be noted that when sensitive information is detected in the target message (such as involving some sensitive topics, such as a certain leader, a certain system, a certain policy, etc.), the above functions will not be triggered, that is, the key content in the target message will no longer be automatically highlighted according to the context message. In this case, the ordinary message presentation method can be adopted.
[0269] For example, for a non-long text picture message, if the picture is not a sensitive picture, the system will intelligently locate the target object in the picture according to the context message with a prompting function sent by the message sender A, and automatically circle the target object when the message recipient B clicks on the message. If the picture is a sensitive picture, when the message recipient B clicks on the message, the target object will not be circled out, and only the picture message will be simply presented.
[0270] It should be noted that the above briefly lists the methods for highlighting the key content in the target messages of several different message types. The present application also takes into account that if the target message is far from the context message with a prompting function (which can be abbreviated as the reference message), when the system detects that a certain reference message (i.e., the context message) sent by the same message sender to the same message recipient is related to other messages (i.e., the target message), an association control will be displayed after the reference message, and the message recipient can manually associate the reference message with other messages based on the association control.
[0271] Specifically, having a prompting function means what is listed above, that is, the context message contains some positioning information, such as time positioning information, location positioning information, etc.; or, the context message contains some key information, which specifically refers to some emphasized and concerned information, such as "She also mentioned Person A" and "Do you see the one wearing black-framed glasses" listed above; and so on.
[0272] An optional implementation method is that if it is determined that there is an association between a context message and a target message, the association control can be presented at the first relevant position of the context message; furthermore, the object can manually trigger an association operation based on the association control, and the client responds to the message association operation triggered by the association control corresponding to the context message, associates the target message with the context message, and presents the target message at the second relevant position of the context message.
[0273] In the above embodiments, the distance between messages may not be considered. If the system detects an association between two messages, these two messages can be automatically associated, or an association control can be presented at the first relevant position of a reference message (i.e., context message) with a prompting function for an object to manually associate them. The distance between messages can also be considered. That is, when it is determined that the distance between two messages is far, these two messages can be automatically associated, or when it is determined that the distance between two messages is far, an association control can be presented at the first relevant position of a reference message with a prompting function for an object to manually associate them.
[0274] Among them, the first relevant position of the context message can specifically be a position near the message, such as to the right of the message, above the message, below the message, etc.
[0275] Similarly, the second relevant position of the context message can also specifically be a position near the message, such as to the right of the message, above the message, below the message, etc.
[0276] It should be noted that the first relevant position and the second relevant position can be the same or different. In the embodiments of the present application, after the message recipient manually associates messages, the display of the association control can be cancelled at the first relevant position of the context message. That is, after successful association, the association control can no longer be continuously displayed at the first relevant position of the context message.
[0277] Optionally, the following method can be used to analyze whether the distance between the target message and its context message is far:
[0278] If the display distance between the target message and the context message in the current chat window exceeds a preset distance (such as one screen), it can be considered that the distance between these two messages is far; or, if the number of messages between the target message and the context message exceeds a preset number (such as 10), it can be considered that the distance between these two messages is far.
[0279] It should be noted that the above-listed preset distance, preset number, etc. are only simple examples. In addition, they can also be set to other distances or numbers, which can be specifically set flexibly according to experience or actual needs and are not specifically limited here.
[0280] In short, the embodiments of the present application are undoubtedly a great help for use scenarios that require searching for key information in multi-modal content in different formats and for quickly viewing key information in a short time, effectively improving the viewing efficiency of the object to view key information in various modal contents and greatly optimizing the object experience.
[0281] The following is an example to illustrate the association between the above-mentioned messages with reference to the drawings:
[0282] Such as Figure 12As shown in the figure, it is a schematic diagram of message association in an embodiment of the present application. Specifically, if the target message is far from the reference message, when the system detects that a certain reference message is associated with a certain target message, an association control is displayed after the reference message. The style of the association control is not specifically limited in this article. For example, it can be a button of "can be associated with xx message", where "xx" represents the type of the message. For example Figure 12 In the reference message S1201 "The first paragraph is written so well", it contains location positioning information and is a context message with a prompt function. Moreover, before this message, the message sender A also sent a multi-modal message, which is a long text picture message containing text content in multiple paragraphs. Through analysis, it can be known that this long text picture message is the target message associated with the reference message. Therefore, an association control can be presented next to the reference message. As shown in S1202, when the message recipient B clicks this button, the target message can be manually associated, and a highlight mark appears on the target message after association.
[0283] Optionally, the association control is a text control describing the message type of the target message. As shown in S1202, on this basis, when associating the target message with the context message, an optional implementation method is as follows:
[0284] The client presents a message selection interface corresponding to the message type in response to a message association operation triggered by the association control corresponding to the context message. In the embodiment of the present application, the message selection interface is used to present each candidate historical message that meets the message type and is published by the message sender who sent the context message in the current chat window within a set historical period; the message recipient can select the target message to be associated from the message selection interface.
[0285] Subsequently, the client associates the context message with the target message in response to a selection operation of selecting the target message from each candidate historical message.
[0286] For example Figure 13 As shown in the figure, it is a schematic diagram of a message selection interface in an embodiment of the present application. As Figure 13 shown, the association control is a text control describing the message type of the target message, and the target message belongs to the picture type, as shown in S1301.
[0287] After the message recipient B clicks the association control, a message selection interface related to pictures can be presented, such as Figure 13 the interface shown on the right in the figure. Among them, the 11 pictures from Picture 1 to Picture 11 are historical picture messages published by the message sender A during the chat process between the message sender A and the message recipient B within a set historical period (such as today, within three days, within seven days, etc.).
[0288] Specifically, the message recipient B can select the target message to be associated from these historical picture messages. For example, Figure 13 as shown, select "Picture 5", and then click the "OK" button. Then, the client can, in response to the selection operation of selecting the target message from each candidate historical message, associate the context message with the target message.
[0289] Furthermore, the associated target message can be presented at the second relevant position of the context message.
[0290] For example, Figure 14 as shown, it is a schematic diagram of associating pictures in an embodiment of the present application. In Figure 14 after the message recipient B selects "Picture 5" and confirms, the picture 5 can be automatically associated with "The writing in the first paragraph is excellent", and the associated picture is presented above "The writing in the first paragraph is excellent", as shown in S1401.
[0291] It should be noted that the above takes the association between "The writing in the first paragraph is excellent" and "Picture 5" as an example. In fact, there may be multiple context messages associated with "Picture 5". In this case, for each context message, the above method can be used to associate it with "Picture 5"; or, there may be multiple target messages associated with "The writing in the first paragraph is excellent". In this case, the message recipient can select multiple pictures in the message selection interface to associate with this context message, etc., which will not be elaborated here one by one.
[0292] In addition, it should be noted that the above takes the association of picture messages as an example. In addition, when associating text messages, video messages, voice messages, etc., a similar method can be used, which will not be elaborated here one by one.
[0293] In the embodiment of the present application, through the above method, two messages with an association relationship can be associated. In this case, since the associated target message is presented at the second relevant position of the context message, when the message recipient browses this context message, the target message at the second relevant position can also be clicked to view. An optional implementation method is as follows:
[0294] The client presents the target message in response to the viewing operation triggered at the second relevant position for the target message, and highlights the key content in the target message.
[0295] In an embodiment of the present application, if the message recipient clicks to view the target message at the second relevant position, the recipient can first jump to the target message. If the message recipient further clicks on the target message (this step can also be ignored, and the recipient can directly jump and view it), the target message will be presented, and the key content will be highlighted in the target message. Alternatively, if the message recipient clicks to view the target message at the second relevant position, the recipient may not jump but directly present the target message, and the key content will be highlighted in the target message; and so on.
[0296] In summary, regardless of which implementation method is used, it supports the object to view the target message, that is, the key content in the target message, from the associated message.
[0297] As Figure 15 shown, it is a schematic diagram of a method for viewing associated messages in an embodiment of the present application. As Figure 15 shown, after the message recipient B clicks to view the associated picture, the interface shown on the right can be presented, presenting the details of the picture and marking the first paragraph in the picture. Figure 15
[0298] Optionally, the associated control is a text control describing the message type of the target message. As shown in S1202, on this basis, when associating the target message with the context message, another optional implementation method is as follows:
[0299] The client responds to a message association operation triggered by an associated control corresponding to the context message, and presents corresponding message selection controls at the third relevant positions of each candidate historical message; then, the message recipient can select the target message to be associated based on the message selection controls.
[0300] In an embodiment of the present application, each of the candidate historical messages is an instant messaging message that conforms to the message type and is sent by the message sender who sent the context message in the current chat window during a set historical period.
[0301] Subsequently, the client responds to a selection operation of selecting the target message from the current chat window based on each message selection control, and associates the context message with the target message.
[0302] Figure 16 As Figure 16 shown, it is a schematic diagram of a message selection control in an embodiment of the present application. As shown, the associated control is a text control describing the message type of the target message, and the target message belongs to the text type, as shown in S1601.
[0303] Figure 16 After the message recipient B clicks the associated control, the message selection controls corresponding to the text messages sent by the message sender A in the current chat window during a set historical period can be presented, as Figure 16S1602, S1603, S1604, S1605, etc. in the interface shown on the right side.
[0304] On this basis, the message recipient B can slide to view more candidate historical messages, such as Figure 17 shown in the left interface of Figure 17 which is a schematic diagram of an associated text message in an embodiment of the present application. Assume that Figure 17 on the basis shown, the message recipient B browses from these candidate historical messages and finds that the message S1701 contains many paragraphs and may be related to the context message "The first paragraph is written very well" in the previous text, and can then select this message for association.
[0305] After association, the interface shown on the left side of Figure 18 will be presented. Among them, Figure 18 is a schematic diagram of another way to view associated messages in an embodiment of the present application. Figure 18 S1801 in
[0306] indicates that the associated target message is presented at the second relevant position of the context message. Figure 18 On this basis, after the message recipient B clicks to view the associated text message, the interface shown on the right side of
[0307] will be presented, jumping to this text message and marking the first paragraph in this text message, as shown in S1802.
[0308] It should be noted that the above mainly gives an example of the message display method in the embodiment of the present application from the product performance on the client side. Next, the message display method in the embodiment of the present application will be further described mainly from the technical implementation details on the server side:
[0309] Refer to Figure 19 shown, which is a flowchart of the implementation of another message display method provided by the embodiment of the present application. Taking the server as the execution entity as an example, the specific implementation process of this method is as follows S191 - S193:
[0310] S191: The server obtains at least one context message associated with the target message;
[0311] Among them, the target message and at least one context message may come from the same message sender or different message senders. For specific details, refer to the above embodiments, and repeated parts will not be elaborated here.
[0312] Specifically, when the message sender A sends multiple messages to the message receiver B, the server can analyze whether there is an association between these messages based on the methods listed above to determine one or more context messages associated with the target message, which will not be repeated here.
[0313] S192: The server extracts the content location information of the key content for the target message mentioned in at least one context message; where the content location information is information with a prompting effect.
[0314] S193: After the server receives a viewing request for the target message, it feeds back the content location information to the client, so that the client, in response to the viewing operation for the target message, highlights the key content in the target message based on the content location information.
[0315] Specifically, after determining the association relationship between multiple messages sent by the same message sender, the server can extract the content location information of the key content for the target message mentioned in these context messages, so that after feeding it back to the client, when the client views the target message, it can locate and highlight the key content in the target message according to these content location information.
[0316] Considering that the messages in the embodiments of the present application can be of any message type (also called modality), such as text, picture, video, audio, etc., for different types of messages, the forms of the content they contain are different, and the corresponding content location information may also be different.
[0317] In the embodiments of the present application, the content location information can be divided into two major categories: location-based and descriptive information. Among them, the location-based content location information can be further divided into two categories: time location information and position location information; the descriptive information can specifically be entity description information.
[0318] Specifically, the time location information is information that describes the time when the key content appears in the target message; the position location information is information that describes the position where the key content appears in the target message; the entity description information is information that describes an entity, specifically, it can be entity words (such as person names, place names, proper nouns), features possessed by the entity (such as shape features, dressing features, gender features), etc.
[0319] Generally speaking, for the extraction tasks of the above two major categories of content location information, in the embodiments of the present application, a method combining a matching rule-based method and a deep learning-based method can be used to solve this task.
[0320] The extraction methods of different content location information are briefly described as follows:
[0321] (1) If the content location information includes location positioning information, at least one context message associated with the target message is matched based on a preset first regular expression, and the location positioning information in this at least one context message is extracted according to the matching result.
[0322] Taking the paragraph location information as an example of the location positioning information, generally, in the context message with a prompting function sent by the message sender A, the paragraph location information can be described as such as "the first page", "the third paragraph", "the fourth line", "the seventh chapter", "the last paragraph", "the first three lines", "the penultimate paragraph", "the first two paragraphs", etc. Such information can be matched by the first regular expression listed as follows:
[0323] (?:last|beginning|previous|penultimate|the)(?:one|two|three|four|five|six|seven|eight|nine|ten|[0-9]+)(?:page|paragraph|line|chapter)
[0324] When the matching is successful, the corresponding paragraph location information can be recorded for subsequent steps of highlighting key content.
[0325] It should be noted that the first regular expression listed above is only a simple example. In addition, other similar regular expressions can be used for matching, which will not be elaborated here one by one.
[0326] In the above implementation, the location positioning information in the message can be quickly extracted through the regular expression, laying a foundation for subsequent marking of key content in the target message (such as the target paragraph, entity at the target location, etc.).
[0327] (2) If the content location information includes time positioning information, at least one context message associated with the target message is matched based on a preset second regular expression, and the time positioning information in this at least one context message is extracted according to the matching result.
[0328] Generally, for target messages of types such as videos and audios, the context messages associated with them may contain time positioning information.
[0329] Specifically, the method of extracting the time location information sent by the message sender is similar to the method of extracting the paragraph location information listed above. In a chat dialogue scenario, the timestamp information of video, audio, etc. can be expressed in the form of "4th second", "1 minute and 30 seconds", "2:05", etc. This form can be matched with a regular expression (i.e., the second regular expression in this article), and the text expression is converted into a timestamp in seconds. When the message recipient clicks on the video or audio message, the timestamp information can be used to jump to the corresponding position.
[0330] In the above implementation, the time location information in the message can be quickly extracted through regular expressions, laying a foundation for subsequent marking of key content in the target message and automatically jumping to play the target message.
[0331] (3) If the content location information includes entity description information, then based on a pre-trained deep learning model, extract the entity description information from at least one context message associated with the target message; wherein the deep learning model includes an entity classification layer, and the deep learning model is fine-tuned based on a historical instant messaging message dataset with pre-labeled entity categories.
[0332] Specifically, the context message with a prompt function sent by the message sender A may contain entity description information such as "A person". In the embodiment of the present application, the task of identifying such entity description information such as names of people, places, proper nouns, etc. in the context message is a named entity recognition (NER) task.
[0333] In the embodiment of the present application, this task can be handled by using a deep learning model. Specifically, the deep learning model includes but is not limited to: Recurrent Neural Network (RNN), Long Short-Term Memory (LSTM), Transformer. These models can process a large amount of unlabeled or labeled data and automatically learn the representation and features of named entities.
[0334] In an embodiment of the present application, an end-to-end and unsupervised pre-trained model BERT can be used as the basis of the deep learning model, and an entity classification layer is added on this basis, and fine-tuned using an annotated NER dataset.
[0335] The deep learning model obtained based on this fine-tuning can be used to match entity words such as "A person", "passport information", "small book", "lightning avoidance post" in the previous text (that is, a type of entity description information in this article).
[0336] When the matching is successful, the corresponding entity description information can be recorded for the subsequent step of highlighting key content.
[0337] In the above embodiment, by identifying entities, entity categories, etc. in the message and extracting the entity description information in the message, it lays a foundation for subsequently marking key information related to entities in the target message (such as target keywords, target entities, etc.).
[0338] It should be noted that the above simply lists the extraction methods of different content location information. After extracting the content location information, the content location information can be fed back to the client so that the client can locate the key content based on this and highlight the key content.
[0339] Among them, the key content can be at least one of a target paragraph, a target keyword, and a target entity. Different key contents can be highlighted in different ways.
[0340] The following separately lists the highlighting processes of key content in several types of target messages:
[0341] First, taking the multi-text content message as an example, the following implementation method can be adopted when highlighting the key content in the multi-text content message:
[0342] When the message recipient B opens the multi-text content message (including long text pictures, documents, external links, etc.), the content location information obtained through the above steps (such as one or more of the paragraph position information and the entity description information) can be used to highlight the multi-text content message.
[0343] Considering that the multi-text content message can be further divided into several types, the following separately describes the specific technical solutions corresponding to several types case by case:
[0344] (1) Long text picture message.
[0345] Specifically, for the long text picture message, first, the text in the long text picture needs to be recognized through the Optical Character Recognition (OCR) technology, and the layout structure and paragraph boundaries in the long text picture are recognized by analyzing the spacing, indentation, and alignment between text lines, etc.
[0346] For example, the Tesseract OCR engine can be used to perform OCR recognition on the long text picture. Tesseract is an open-source OCR engine that can recognize image files in various formats and convert them into text. Of course, other OCR engines can also be used, which are not specifically limited in this article.
[0347] After parsing the long text image message based on the above method, if the content location information extracted for the long text image message includes paragraph location information, the corresponding paragraphs can be highlighted in the following way:
[0348] Specifically, after the layout structure and paragraph boundaries are recognized by the OCR engine, according to the extracted paragraph location information, calculate the image coordinates of the target paragraph to be highlighted in the long text image; furthermore, according to the image coordinates, use the Open Source Computer Vision Library (OpenCV) to process the original long text image, and color the position to be highlighted with a transparent color to highlight the target paragraph.
[0349] If the content location information extracted for the long text image message includes entity description information, the corresponding keywords can be highlighted in the following way:
[0350] Specifically, after the text is recognized by the OCR engine, match the extracted entity description information with the recognized text to find the target keywords to be highlighted in the text. Furthermore, after finding the corresponding keywords, calculate the image coordinates of the keywords in the long text image according to the position, size, rotation information, etc. of the keywords in the text; furthermore, use the OpenCV computer vision library to process the original long text image, and color the position to be highlighted with a transparent color to highlight the target keywords.
[0351] (2) Document message.
[0352] Specifically, for a document message, if the content location information extracted for the document message includes paragraph location information, the corresponding paragraphs can be highlighted in the following way:
[0353] Taking a PDF document as an example, a PDF document consists of many basic elements called "objects". However, in a PDF document, there is no dedicated paragraph object to represent paragraphs. Paragraphs are represented by the visual grouping and arrangement of text content. A PDF document uses a content stream object to describe elements such as text, graphics, and images on a page. The content stream consists of a series of graphics operations, which are organized in the drawing order.
[0354] Therefore, in the embodiments of the present application, in order to highlight the target paragraph in the PDF document, first, it is necessary to parse the content stream of the PDF document, extract the text position (Td) and text matrix (Tm) operation instructions in the content stream, and distinguish paragraphs according to the vertical spacing. Among them, the content stream sets the vertical position of the text through Td and Tm operations. If the vertical distance between two text blocks is greater than a certain threshold (for example, 1.5 times the line height), they are regarded as different paragraphs. After distinguishing the paragraphs, according to the extracted paragraph position information, find the target paragraph to be highlighted. For the target paragraph, create a new Annot annotation object, add a reference to the annotation object in the Annots entry of the corresponding content stream, and then insert it into the PDF document to achieve the effect of highlighting the target paragraph.
[0355] If the content location information extracted for the document message includes entity description information, the corresponding keywords can be highlighted in the following manner:
[0356] First, parse the text of the PDF document, perform string matching on the extracted entity description information and the parsed text to find the target keywords to be highlighted in the text. After finding the corresponding keywords, calculate their rectangular positions in the content stream. Combine the calculated positions, create a new Annot annotation object, add a reference to the annotation object in the Annots entry of the corresponding content stream, and then insert it into the PDF document to achieve the effect of highlighting the target keywords.
[0357] (3) External link message.
[0358] The external link web page is presented by a tree-structured HTML document. Therefore, it is necessary to parse the HTML document of the external link. Specifically, when parsing HTML, first traverse its tags to obtain the tree structure of the tags.
[0359] Specifically, for an external link, if the content location information extracted for the external link includes paragraph position information, the corresponding paragraphs can be highlighted in the following manner:
[0360] Generally, web articles will appear in 、 <section> 、 <article>Among these three tags. These three tags are usually used <h1>、< / h1> <h2>、< / h2> <h3>、< / h3> <h4>、< / h4> <h5>Labels are used to represent the title, and 、 , Use tags such as to distinguish the body paragraphs. Therefore, the paragraph information can be well distinguished according to the opening and closing of the tags. Among them, the opening of the tag refers to "<", and the closing of the tag refers to ">".
[0361] Due to the diversity of tags, different websites may have different usage methods. The embodiments of the present application are completed using a machine learning model trained by an annotated dataset (such as Common Crawl). Among them, Common Crawl is a large publicly available web crawl dataset containing billions of web pages. Using this dataset, models for tasks such as HTML parsing, information extraction, and content extraction can be trained and evaluated.
[0362] After the paragraph information in the web page is parsed by the machine learning model, according to the extracted paragraph position information, the target paragraph to be highlighted is found, and a specific CSS style is added to its tag, so as to achieve the effect of highlighting the target paragraph.
[0363] For external links, if the content location information extracted for the external link includes entity description information, the corresponding keywords can be highlighted in the following way:
[0364] First, parse the text in the web page hypertext markup language (HTML) file, and perform string matching between the extracted entity description information and the parsed text to find the target keywords to be highlighted in the text. After finding the corresponding keywords, use html for the keywords Wrap the label and add specific Cascading Style Sheets (CSS) styles to achieve the effect of highlighting the target keywords.
[0365] The above briefly lists several ways to highlight key content in multi-text content messages. In addition, the multi-text content message can also be other text messages. Moreover, other highlighting methods can be adopted, which will not be elaborated here one by one.
[0366] It should be noted that the above takes the target message as a multi-text content message as an example. In addition, the target message can also be a picture message, a video message, an audio message, etc.
[0367] Taking a non-long text picture message as an example below, the following implementation method can be adopted to highlight the key content in the non-long text picture message:
[0368] For non-long text picture messages (also known as regular picture messages), the key content in such messages is generally a target entity, such as a target object, a target person, etc.
[0369] An optional implementation method is that the target message is a non-long text picture message, the content location information includes entity description information, and the key content is the target entity; then S193 can be implemented according to the following process, including the following steps S19311 to S19341:
[0370] S19311: Feed the content location information back to the client so that the client can obtain a pre-constructed first entity dictionary; where the first entity dictionary contains multiple first key-value pairs, the key in each first key-value pair is an entity category, and the value in each first key-value pair is an array of expression words associated with the entity category.
[0371] It should be noted that the first entity dictionary in the embodiments of the present application can be constructed by the client or the server in the following ways:
[0372] First, extract the expression words related to the entity from the historical instant messaging messages; then, according to the extracted expression words, construct an array of expression words associated with each entity category; finally, based on the constructed array of expression words, store the index relationship between the entity category and the array of expression words in dictionary form to construct the first entity dictionary, which contains multiple first key-value pairs, the key in each first key-value pair is an entity category, and the value in each first key-value pair is an array of expression words associated with the entity category.
[0373] Among them, the historical instant messaging messages can be messages sent or received by any object in a chat scenario. When constructing the array of expression words based on the extracted entity expression words, specifically, the expression words describing entities of the same category are constructed into a group, and then storing the index relationship between the entity category and the array of expression words can be used for subsequent steps.
[0374] In a chat scenario, there can be many text expressions for each entity category by the message sender or the message receiver. For example, the possible expression words for the entity "bicycle" may be "bicycle", "bike", "shared bike", "yellow bike", "pedal bike", etc., and these expression words should all be matched to the "bicycle" entity. To achieve this matching, in the embodiments of the present application, the output of the SSD algorithm is organized in the form of a dictionary, where the keys of the dictionary are category labels (usually strings or integers), and the values are an array of expression words. For example:
[0375] {
[0376] "MALE": ["man", "male", "guy", "little brother", "handsome guy"],
[0377] "FEMALE": ["woman", "female", "gal", "little sister", "beauty"],
[0378] "BIKE": ["bicycle", "bike", "shared bike", "yellow bike", "pedal bike"],
[0379] "CAR": ["car", "SUV", "sedan", "private car", "sports car", "A car brand", "B car brand", "C car brand"],
[0380] }
[0381] The above first entity dictionary contains 4 first key-value pairs, listing the arrays of expression words corresponding to 4 entity categories respectively. It should be noted that the first entity dictionary here is just a simple example. In the actual application process, the constructed first entity dictionary should contain more first key-value pairs, which will not be elaborated here one by one.
[0382] It should be noted that if the above first entity dictionary is constructed by the server, the client can obtain it from the server side. In addition, as time goes by, the historical instant messaging messages will continue to increase, so the first entity dictionary can also be dynamically updated and improved based on this. In this case, the client can pull the latest constructed first entity dictionary from the server side.
[0383] After constructing the first entity dictionary, the entities recognized in the target message can be subsequently matched with this dictionary to determine whether the target message contains the target entities that need to be highlighted. The specific process is as follows:
[0384] S19321: Identify the target entity category corresponding to the entity included in the non-long text picture message.
[0385] Specifically, if the non-long text picture message contains only one entity, one target entity category can be determined; if the non-long text picture message contains multiple entities, multiple target entity categories can be determined.
[0386] For each target entity category, the following operations are respectively performed:
[0387] S19331: Match the entity description information included in the content location information with the array of expression words associated with the target entity category in the first entity dictionary.
[0388] S19341: If the match is successful, the corresponding entity is taken as the target entity, and the target entity is highlighted in the non-long text picture message.
[0389] Specifically, for these historical instant messaging messages of the picture type, taking the picture containing the target object as an example, it is necessary to identify the target object in the picture. Specifically, a target detection algorithm can be used for identification, such as using the Single Shot MultiBox Detector (SSD) algorithm for identification.
[0390] SSD is a fast and accurate target detection algorithm that can detect multiple entities in a single inference for a picture; SSD uses feature maps of multiple scales to predict bounding boxes and categories, and is applicable to entities of different sizes.
[0391] Since the embodiments of the present application are mainly applied to the chat scenario, a relatively lightweight dataset can be used to train a model based on the SSD algorithm, and then entity recognition is performed on the picture based on this model. For example, the MobileNet-SSD model pre-trained with the Common Objects in Context (COCO) dataset (a commonly used target detection dataset) can identify 80 different entity categories, such as people, bicycles, cars, dogs, etc.
[0392] When the expression word (i.e., entity description information) of the entity included in the context message associated with the target message hits any string in the array of expression words corresponding to the target entity category identified based on the target message, it is considered a match. At this time, a circled selection graphic is drawn based on the rectangular border (which can also be a contour, etc.) of the identified entity (denoted as the target entity) to achieve the effect of highlighting the target entity.
[0393] In addition, when the content location information proposed for the target message includes location positioning information, such as "the sixth from the right in the second row" listed above, entity recognition can be performed on the target message of non-long text picture type based on this content location information, and the target entity that conforms to this location positioning information can be located. Furthermore, a selection graphic can be drawn based on the rectangular border (or contour, etc.) of the recognized target entity to achieve the effect of highlighting the target entity.
[0394] It should be noted that the above takes the target message as a non-long text picture message as an example. In addition, the target message can also be a video message, an audio message, etc.
[0395] The following briefly describes the method of highlighting key content in video messages and audio messages:
[0396] For video messages, the content location information can include at least one of time positioning information and entity description information, and the key content is the target entity. Assuming that the content location information includes both time positioning information and entity description information at the same time, the following steps S19312 - S19332 can be implemented:
[0397] S19312: Feed the content location information back to the client, so that the client automatically jumps to the target frame position based on the time positioning information and starts playing the video message automatically from the target frame position;
[0398] S19322: Identify the entities in the video frame corresponding to the target frame;
[0399] S19332: If there is a target entity in the recognized entities that matches the entity description information included in the content location information, highlight the target entity in the target frame of the video message.
[0400] Specifically, for video messages, the video time information (i.e., time positioning information) sent by the message sender can be extracted. In the context of a chat conversation, the time stamp information of the video can be expressed in forms such as "the 4th second", "1 minute and 30 seconds", "2:05", etc. This form can be matched using a regular expression, and the text expression can be converted into a video time stamp in seconds. When the message recipient clicks on the video message, the time stamp information can be used to jump to the specified position of the video, that is, the starting frame position in the target frame mentioned above.
[0401] In addition, keyword information (i.e., entity description information) sent by the message sender can be extracted. After the keyword information is extracted, specific entities in the video need to be identified. In order to quickly locate and highlight the entities mentioned by the chatters in the chat scenario, after jumping to the specified position of the video using the timestamp information in the previous step, similar to the implementation manner related to the non-long text picture message above, the SSD object detection algorithm can be used to identify the images of each target frame, and the detected entities are matched with the keyword information. If the match is successful, then based on the entity grid border output by the SSD algorithm, drawing is performed to achieve the effect of circling the target entity.
[0402] In addition, it should be noted that image recognition can also be performed on the message frames after the target frame. If the target entity is recognized, it can also be circled in the above manner.
[0403] For audio messages, if the audio message is a voice message, then first convert the voice message into a text message, and then the subsequent processing process is similar to that of the above multi-text content message part. That is, if the content location information extracted from the voice message includes time location information, then based on the time specified by the time location information, the corresponding converted text part can be determined, and then a color with transparency is applied to the position of the text part that needs to be highlighted to highlight the key content. If the content location information extracted from the voice message includes entity description information, then the extracted entity description information is matched with the converted text to find the target keyword that needs to be highlighted in the text. After determining the corresponding timestamp in the converted text, a color with transparency is applied to the text part that needs to be highlighted corresponding to the timestamp to highlight the target keyword.
[0404] The corresponding paragraphs are highlighted in the following way, which will not be repeated here; for audio message files of pure music type, the audio track needs to be extracted and the content of the audio track needs to be recognized to convert the audio message into a text message, and then the subsequent processing process is the same as that of the above voice message, which will not be repeated here.
[0405] In summary, the ways to highlight key content in several different types of target messages are simply listed, but the above are only simple examples, and other methods can also be adopted, which will not be elaborated one by one here.
[0406] In the above embodiments, by using OCR technology, the text content in pictures can be quickly recognized. By using NER technology, the key entity information in the chat content can be more accurately recognized. By using SSD technology, the key entities in images and videos can be recognized. Then, this method matches the extracted key entity information with the historical chat records of the object. This means that no matter which part of the content in the chat the object mentions, whether it is text, picture, file, web page or video, this method can quickly locate the key information mentioned in the chat content. This not only improves the efficiency of information search, but also greatly optimizes the object experience.
[0407] In addition, it should be noted that in step S193, the server can also independently determine the key content that needs to be highlighted in the target message based on the content positioning information. Furthermore, the relevant information of the determined key content is fed back to the client for highlighting by the client.
[0408] Specifically, whether it is the server or the terminal device (client) that determines the key content, the specific determination method of the key content is basically the same. For details, please refer to the above embodiments and will not be repeated here.
[0409] The above briefly lists several methods for highlighting key content in target messages of different message types. The present application also considers that if the target message is far from the context message with a prompt function (which can be simply referred to as the reference message), when the system detects that a certain reference message (i.e., the context message) sent by the same message sender to the same message recipient is related to other messages (i.e., the target message), an association control is displayed after the reference message, and the message recipient can manually associate the reference message with other messages based on the association control.
[0410] In the embodiments of the present application, in order to implement associated messages, it is necessary to preprocess the message when receiving a multimodal message (i.e., the target message), and then match the context message with the multimodal target message.
[0411] Specifically, the target message can be divided into multiple types. The following briefly illustrates for different message types respectively:
[0412] For the target message sent by the message sender, if the target message is a multi-text content message such as a long text picture, document, external link, etc., a series of keywords can be extracted based on the extraction of picture, document, and web page text, combined with NER technology, and an index from keywords to content can be established. For example, "D Bookstore", "Rock Music", "Winter in S City", "Ju* Road", "A Person" etc. can be extracted.
[0413] For the context message of the target message sent by the message sender, the NER technology can be combined to extract the entity description information therein. For the specific implementation manners, please refer to the above embodiments, and the repeated parts will not be elaborated. After the entity description information is extracted, it is combined with the above index and the keywords extracted from the above multi-text content message for matching. If the matching is successful, it is determined to be associated.
[0414] Furthermore, when the message receiver receives the message, it can perform a jump according to the associated information.
[0415] Similar to the multi-text content message listed above, if the target message is an audio message, first, based on the implementation manners listed above, the audio message can be converted into a text message. Then, similar to the processing manner of the multi-text content message above, the NER technology can be combined to extract a series of keywords in the converted text message and establish an index from the keywords to the content.
[0416] For the context message of the target message sent by the message sender, the NER technology can be combined to extract the entity description information therein. For the specific implementation manners, please refer to the above embodiments, and the repeated parts will not be elaborated. After the entity description information is extracted, it is combined with the above index and the keywords extracted from the above converted text message for matching. If the matching is successful, it is determined to be associated.
[0417] Furthermore, when the message receiver receives the message, it can perform a jump according to the associated information.
[0418] If the target message is a video message, an index of time stamps and entity categories needs to be established, and based on this index, it is analyzed whether the messages are associated. An optional implementation manner is to determine the context message associated with the target message through the following method:
[0419] First, obtain a second entity dictionary pre-constructed for the video message; wherein, the second entity dictionary contains at least one second key-value pair, the key in each second key-value pair is a video time stamp, and the value in each second key-value pair is the entity category of the entity in the video frame corresponding to the video time stamp.
[0420] It should be noted that the second entity dictionary in the embodiments of the present application can be constructed by the client or the server through the following method:
[0421] First, perform frame-by-frame processing on the video message to identify the entities included in each video frame of the video message; then, based on the recognition result, when the entities in each video frame included in the video message change, insert a new second key-value pair into the second entity dictionary, and the video time stamp represented by the key in the new key-value pair is the video time stamp corresponding to the video frame when the change occurs this time.
[0422] Specifically, the SSD algorithm can be used to process each frame of the video, identify the entities in each frame of the video image, and then store the index relationship in the form of a dictionary. The key of the dictionary is the timestamp, and the value of the dictionary is the entity label in the SSD algorithm corresponding to the entity recognized in the video image of this timestamp (the label represents the category of the entity). Only when the entities in the image change, new key-value pairs are inserted. For example:
[0423] {
[0424] "00:01":["MALE","CAR"],
[0425] "00:05":["MALE","DOG"],
[0426] "00:10":["MALE","FEMALE"],
[0427] "00:20":["MALE","FEMALE","WINE"]
[0428] }
[0429] Among them, the above dictionary indicates that in this video message, men and cars are recognized in the video images during the time period from 00:01 to 00:04; men and dogs are recognized in the video images during the time period from 00:05 to 00:09; men and women are recognized in the video images during the time period from 00:10 to 00:19; men, women, and wine are recognized in the video images during the time period from 00:20 to the end of the video.
[0430] Furthermore, for each context message, the entity description information in the context message is matched with the second entity dictionary; if the match is successful, it is determined that there is an association between the video message and the context message.
[0431] Specifically, for the context message of the target message sent by the message sender, the entity description information can be extracted by combining the NER technology. For the specific implementation method, refer to the above embodiments, and the repeated parts will not be elaborated. After the entity description information is extracted, it is matched with the above dictionary. When the entities recognized in the video appear in the context message, it is determined that there is an association.
[0432] Furthermore, when the message receiver receives the message, it can perform a jump according to the association information.
[0433] Similar to the video messages listed above, a video can be understood as a combination of multiple frames of images. If the target message is a picture message (here it refers to a long text picture message), the SSD algorithm can be directly used to identify the entities in the picture message and determine the corresponding entity categories, and then an index of the content and the entity categories is established.
[0434] Furthermore, for each context message, entity description information can be extracted by combining with NER technology. For the specific implementation manners, please refer to the above embodiments, and the repeated parts will not be elaborated here. After the entity description information is extracted, it is matched with the above index. When the entity recognized in the picture appears in the context message, it is determined that there is an association.
[0435] Furthermore, when the message receiver receives a message, it can jump according to the association information.
[0436] It should be noted that the above uses the target messages of different message types as examples to simply illustrate the association between the context message and the target message. In addition, the association between the target messages of other message types and the context message also applies to the embodiments of the present application, which will not be elaborated here one by one.
[0437] In the above implementation manner, by automatically associating messages with an association relationship and supporting the jump of associated messages, it can be ensured that when the object is far away from the target message and the context message with reference prompts, it can also quickly view the target message through the above method; and in this way, the key content in the target message will also be highlighted, saving a lot of time and energy for the object. The object does not need to view each message one by one, and can find the most critical information at a glance. This not only improves the efficiency of information search, but also greatly optimizes the object experience.
[0438] Similar to the client side above, it should be emphasized that the relevant technical solutions on the server side involve communication messages between objects in a chat scenario, etc. The collection, use, and processing of these data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and have obtained the permission or consent of the object, which will not be repeated here.
[0439] Refer to Figure 20 As shown in Figure 20 , it is a schematic diagram of the interaction logic between a terminal device and a server in an embodiment of the present application. Among them, Figure 20 A logical example is listed in which when the message sender A sends multiple messages to the message receiver B, the key content in the target message to be viewed is highlighted according to the available reference context messages. Specifically, the message sender A can edit an instant messaging message in the chat interface of a chat application in the terminal device, and then after the instant messaging message is edited, the message sender A can trigger the terminal device to send the edited instant messaging message, and the message receiver B can receive and view the instant messaging message sent by the message sender A through the corresponding terminal device.
[0440] Specifically, the instant messaging message sent by the message sender A to the message receiver B needs to be forwarded by the server. In the related art, the server only performs simple forwarding, while in the embodiment of the present application, the server will analyze whether there is an association between these messages. In the case of an association, the server can execute Figure 20 the steps in Figure 20 , extract the content location information in these context messages, and feedback it to the terminal device of the message receiver B. Furthermore, when the message receiver B clicks to view the target message, based on this content location information, the key content in the target message can be highlighted, as shown in Figure 20 Figure 20 . Of course, the specific implementation logic of this process can be referred to the above embodiment, and the repeated parts will not be elaborated.
[0441] It should be noted that Figure 20 the scenarios listed in Figure 20 are only simple examples. In addition, other scenarios are also applicable to the embodiments of the present application, which will not be elaborated one by one here.
[0442] Since the present application can automatically analyze the association relationship between the messages sent by the same message sender, and then, based on this association relationship, extract the content location information from the context messages associated with the target message, so as to automatically mark the key content in the target message based on this, that is, the content that the message sender wants the message receiver to focus on. Furthermore, the message receiver can focus on viewing the key content according to the highlighted mark, effectively improving the viewing efficiency of the message.
[0443] Based on the same inventive concept, the embodiment of the present application also provides a message display device. As shown in Figure 21 Figure 21 , it is a schematic structural diagram of the message display device 2100, which may include:
[0444] The first display unit 2101 is used to display the chat interface; display the target message sent by the message sender in the chat interface;
[0445] The second display unit 2102 is configured to highlight key content in the target message in response to a viewing operation on the target message; wherein, the key content is determined from the target message based on content positioning information; the content positioning information is extracted from at least one context message associated with the target message and is information with a prompting effect.
[0446] Optionally, the target message is an audio message or a video message, and the second display unit 2102 is specifically configured to:
[0447] Use each message frame corresponding to the specified time represented by the content positioning information as the target frame, or use each message frame including the key content represented by the content positioning information as the target frame;
[0448] Jump to the position corresponding to the starting frame among each target frame, automatically play the target message starting from the position corresponding to the starting frame, and highlight the key content in the target message.
[0449] Optionally, the second display unit 2102 is specifically configured to:
[0450] Highlight the key content in the target message; or, circle out the key content in the target message;
[0451] Wherein, the key content is at least one of a target paragraph, a target keyword, and a target entity.
[0452] Optionally, if the key content is a target entity, the second display unit 2102 is specifically configured to:
[0453] Circle out the target entity in the target message according to the outline of the target entity.
[0454] Optionally, when the target message is an audio message, the second display unit 2102 is specifically configured to:
[0455] Convert the audio message into a text message, and highlight the key content in the converted text message.
[0456] Optionally, the apparatus further includes:
[0457] An association unit 2103, configured to present an association control at a first relevant position of the context message if it is determined that there is an association between a context message and the target message;
[0458] In response to a message association operation triggered by the association control corresponding to the context message, associate the target message with the context message, and present the target message at a second relevant position of the context message.
[0459] Optionally, if the association control is a text control describing the message type of the target message, the association unit 2103 is specifically configured to:
[0460] In response to a message association operation triggered by an association control corresponding to a context message, a message selection interface corresponding to the message type is presented; wherein, the message selection interface is used to present each candidate historical message that conforms to the message type and is published by the message sender in the current chat window within a set historical period;
[0461] In response to a selection operation of selecting a target message from each candidate historical message, the context message is associated with the target message.
[0462] Optionally, if the association control is a text control that describes the message type of the target message, the association unit 2103 is specifically configured to:
[0463] In response to a message association operation triggered by an association control corresponding to a context message, a corresponding message selection control is presented at a third relevant position of each candidate historical message that conforms to the message type in the current chat window;
[0464] In response to a selection operation of selecting a target message from the current chat window based on each message selection control, the context message is associated with the target message.
[0465] Optionally, the association unit 2103 is further configured to:
[0466] In response to a viewing operation on the target message triggered at a second relevant position, the target message is presented, and the key content is highlighted in the target message.
[0467] Optionally, the display distance between the target message and the context message in the current chat window exceeds a preset distance, or the number of messages between the target message and the context message exceeds a preset number.
[0468] Optionally, the apparatus further includes an analysis unit 2104, configured to determine that there is an association between the context message and the target message in at least one of the following ways:
[0469] The context message contains content existing in the target message;
[0470] The context message contains at least one of time positioning information and location positioning information.
[0471] Based on the same inventive concept, an embodiment of the present application further provides another message display apparatus. As Figure 22 shown, it is a schematic structural diagram of the message display apparatus 2200, and may include:
[0472] A message acquisition unit 2201, configured to acquire at least one context message associated with the target message;
[0473] An information extraction unit 2202, configured to extract content location information of key content for a target message mentioned in at least one context message; wherein, the content location information is information with a prompting effect;
[0474] A feedback unit 2203, configured to, after receiving a viewing request for a target message, feedback the content location information to a client, so that the client, in response to a viewing operation on the target message, highlights the key content in the target message based on the content location information.
[0475] Optionally, the information extraction unit 2202 is specifically configured to:
[0476] If the content location information includes location location information, match at least one context message based on a preset first regular expression, and extract the location location information in the at least one context message according to the matching result;
[0477] If the content location information includes time location information, match at least one context message based on a preset second regular expression, and extract the time location information in the at least one context message according to the matching result;
[0478] If the content location information includes entity description information, extract the entity description information in at least one context message based on a pre-trained deep learning model; wherein, the deep learning model includes an entity classification layer, and the deep learning model is fine-tuned based on a historical instant messaging message dataset with pre-labeled entity categories.
[0479] Optionally, the target message is a non-long text picture message, the content location information includes entity description information, and the key content is a target entity;
[0480] Then the feedback unit 2203 is specifically configured to:
[0481] Feedback the content location information to the client, so that the client obtains a pre-constructed first entity dictionary; wherein, the first entity dictionary includes multiple first key-value pairs, the key in each first key-value pair is an entity category, and the value in each first key-value pair is an array of expression words associated with the entity category;
[0482] Identify the target entity category corresponding to the entity in the non-long text picture message;
[0483] Match the entity description information included in the content location information with the array of expression words associated with the target entity category in the first entity dictionary;
[0484] If the match is successful, use the corresponding entity as the target entity and highlight the target entity in the non-long text picture message.
[0485] Optionally, the apparatus further includes a first construction unit 2204, configured to construct a first entity dictionary in the following manner:
[0486] Extract the entity-related expression words from the historical instant messaging messages;
[0487] According to the extracted expression words, construct an expression word array associated with each entity category;
[0488] Based on the constructed expression word array, store the index relationship between the entity category and the expression word array in the form of a dictionary, and construct the first entity dictionary.
[0489] Optionally, the target message is a video message, the content location information includes time location information and entity description information, and the key content is the target entity;
[0490] Then, the feedback unit 2203 is specifically configured to:
[0491] Feedback the content location information to the client, so that the client automatically jumps to the target frame position based on the time location information, and starts playing the video message automatically from the target frame position; and
[0492] Identify the entities in the video picture corresponding to the target frame;
[0493] If a target entity matching the entity description information included in the content location information exists in the identified entities, highlight the target entity in the target frame of the video message.
[0494] Optionally, if the target message is a video message, the apparatus further includes an association unit 2205, configured to determine the context message associated with the target message in the following manner:
[0495] Obtain a second entity dictionary pre-constructed for the video message; wherein, the second entity dictionary includes at least one second key-value pair, the key in each second key-value pair is a video timestamp, and the value in each second key-value pair is the entity category of the entity in the video picture corresponding to the video timestamp;
[0496] For each context message, match the entity description information in the context message with the second entity dictionary;
[0497] If the match is successful, determine that there is an association between the video message and the context message.
[0498] Optionally, the apparatus further includes a second construction unit 2206, configured to construct a second entity dictionary in the following manner:
[0499] Process the video message frame by frame, and identify the entities included in each frame of the video picture in the video message;
[0500] Based on the recognition result, when it is determined that an entity in each video frame included in the video message has changed, a new second key-value pair is inserted into the second entity dictionary, and the video timestamp represented by the key in the new key-value pair is the video timestamp corresponding to the video frame at the time of this change.
[0501] In this application, for multiple messages sent by a message sender, it is possible to determine whether there is an association between the multiple messages. In the case where there is an association between the multiple messages, the context messages associated with the target message can be determined, and the content location information with a prompting effect mentioned in these associated context messages can be extracted. Furthermore, when viewing the target message, based on the extracted content location information, the key content determined based on the content location information in at least one context message can be automatically highlighted in the target message.
[0502] Since this application can automatically analyze the association relationship between messages sent by the same message sender, furthermore, based on this association relationship, content location information can be extracted from the context messages associated with the target message, so as to automatically mark the key content in the target message, that is, the content that the message sender wants the message recipient to focus on. Furthermore, the message recipient can focus on viewing the key content according to the highlighted mark, effectively improving the message viewing efficiency.
[0503] For the convenience of description, the above parts are divided into various modules (or units) according to functions and described separately. Of course, when implementing this application, the functions of the various modules (or units) can be implemented in the same or multiple software or hardware.
[0504] In the embodiments of this application, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.
[0505] After introducing the message display method and device according to the exemplary embodiments of this application, next, an electronic device according to another exemplary embodiment of this application is introduced.
[0506] Those skilled in the art can understand that various aspects of this application can be implemented as a system, a method, or a program product. Therefore, various aspects of this application can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or an implementation combining hardware and software aspects, which can be collectively referred to as "circuit", "module", or "system" here.
[0507] Based on the same inventive concept as the above method embodiments, an electronic device is also provided in the embodiments of the present application. In one embodiment, the electronic device may be a server, such as Figure 1 the server 120 shown. In this embodiment, the structure of the electronic device may be as Figure 23 shown, including a memory 2301, a communication module 2303, and one or more processors 2302.
[0508] The memory 2301 is used to store computer programs executed by the processor 2302. The memory 2301 may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system and programs required to run the instant messaging function, etc.; the data storage area may store various instant messaging information and operation instruction sets, etc.
[0509] The memory 2301 may be a volatile memory, such as a random-access memory (RAM); the memory 2301 may also be a non-volatile memory, such as a read-only memory, a flash memory, a hard disk drive (HDD), or a solid-state drive (SSD); or the memory 2301 is any other medium that can be used to carry or store a desired computer program in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 2301 may be a combination of the above memories.
[0510] The processor 2302 may include one or more central processing units (CPUs) or be a digital processing unit, etc. The processor 2302 is used to implement the above message display method when calling the computer program stored in the memory 2301.
[0511] The communication module 2303 is used to communicate with terminal devices and other servers.
[0512] In the embodiments of the present application, the specific connection medium between the above memory 2301, communication module 2303, and processor 2302 is not limited. In the embodiments of the present application Figure 23 it is shown that the memory 2301 and the processor 2302 are connected through a bus 2304. The bus 2304 is described in thick lines in Figure 23 The connection manners between other components are only for illustrative purposes and are not to be taken as limiting. The bus 2304 may be divided into an address bus, a data bus, a control bus, etc. For ease of description, Figure 23 is described only by a thick line, but does not describe having only one bus or one type of bus.
[0513] The computer storage medium is stored in the memory 2301, and computer-executable instructions are stored in the computer storage medium. The computer-executable instructions are used to implement the message display method of the embodiments of the present application. The processor 2302 is used to execute the above message display method, as Figure 19 shown.
[0514] In another embodiment, the electronic device may also be other electronic devices, such as Figure 1 the terminal device 110 shown. In this embodiment, the structure of the electronic device may be as Figure 24 shown, including: a communication component 2410, a memory 2420, a display unit 2430, a camera 2440, a sensor 2450, an audio circuit 2460, a Bluetooth module 2470, a processor 2480 and other components.
[0515] The communication component 2410 is used to communicate with the server. In some embodiments, it may include a Wireless Fidelity (WiFi) module. The WiFi module belongs to short-range wireless transmission technology. The electronic device can help the object send and receive information through the WiFi module.
[0516] The memory 2420 can be used to store software programs and data. The processor 2480 executes various functions and data processing of the terminal device 110 by running the software programs or data stored in the memory 2420. The memory 2420 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices. The memory 2420 stores an operating system that enables the terminal device 110 to run. In the present application, the memory 2420 can store the operating system and various application programs, and can also store the computer program for executing the message display method of the embodiments of the present application.
[0517] The display unit 2430 can also be used to display the information input by the object or the information provided to the object, as well as the graphical user interface (GUI) of various menus of the terminal device 110. Specifically, the display unit 2430 may include a display screen 2432 provided on the front of the terminal device 110. Among them, the display screen 2432 can be configured in the form of a liquid crystal display, a light-emitting diode, etc. The display unit 2430 can be used to display the operation interface of the client in the embodiments of the present application, etc.
[0518] The display unit 2430 can also be used to receive input digital or character information, and generate signal inputs related to the object settings and function controls of the terminal device 110. Specifically, the display unit 2430 may include a touch screen 2431 disposed on the front of the terminal device 110, which can collect touch operations of an object on or near it, such as clicking a button, dragging a scroll bar, etc.
[0519] Among them, the touch screen 2431 can cover the display screen 2432, or the touch screen 2431 and the display screen 2432 can be integrated to implement the input and output functions of the terminal device 110. After integration, it can be simply referred to as a touch display screen. In this application, the display unit 2430 can display application programs and corresponding operation steps.
[0520] The camera 2440 can be used to capture static images, and an object can publish the images captured by the camera 2440 through an application. The camera 2440 can be one or more. An object generates an optical image through a lens and projects it onto a photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the processor 2480 to convert it into a digital image signal.
[0521] The terminal device may further include at least one sensor 2450, such as an acceleration sensor 2451, a distance sensor 2452, a fingerprint sensor 2453, and a temperature sensor 2454. The terminal device may also be configured with other sensors such as a gyroscope, a barometer, a hygrometer, a thermometer, an infrared sensor, a light sensor, and a motion sensor.
[0522] The audio circuit 2460, the speaker 2461, and the microphone 2462 can provide an audio interface between the object and the terminal device 110. The audio circuit 2460 can transmit the electrical signal converted from the received audio data to the speaker 2461, and the speaker 2461 converts it into a sound signal for output. The terminal device 110 may also be configured with a volume button for adjusting the volume of the sound signal. On the other hand, the microphone 2462 converts the collected sound signal into an electrical signal, which is received by the audio circuit 2460 and then converted into audio data. The audio data is then output to the communication component 2410 to be sent to, for example, another terminal device 110, or the audio data is output to the memory 2420 for further processing.
[0523] The Bluetooth module 2470 is used to interact with other Bluetooth devices having Bluetooth modules through the Bluetooth protocol. For example, the terminal device can establish a Bluetooth connection with a wearable electronic device (such as a smart watch) that also has a Bluetooth module through the Bluetooth module 2470, so as to perform data interaction.
[0524] The processor 2480 is the control center of the terminal device, connecting various parts of the entire terminal using various interfaces and lines. By running or executing software programs stored in the memory 2420, and calling data stored in the memory 2420, it executes various functions of the terminal device and processes data. In some embodiments, the processor 2480 may include one or more processing units; the processor 2480 may also integrate an application processor and a baseband processor. Among them, the application processor mainly processes the operating system, object interfaces, and application programs, etc., and the baseband processor mainly processes wireless communication. It can be understood that the above baseband processor may not be integrated into the processor 2480. In this application, the processor 2480 can run the operating system, application programs, object interface display, and touch response, as well as the message display method of the embodiments of this application. In addition, the processor 2480 is coupled to the display unit 2430.
[0525] In some possible implementation manners, various aspects of the message display method provided in this application can also be implemented in the form of a program product, which includes a computer program. When the program product runs on an electronic device, the computer program is used to cause the electronic device to execute the steps in the message display method according to various exemplary embodiments of this application described above in this specification. For example, the electronic device can execute steps such as Figure 2 or Figure 19 shown in.
[0526] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0527] The program product of an embodiment of the present application may be a portable compact disc read-only memory (CD-ROM) and include a computer program, and may be run on an electronic device. However, the program product of the present application is not limited thereto. In this document, a readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with a command execution system, apparatus, or device.
[0528] A readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a readable computer program is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. A readable signal medium may also be any readable medium other than a readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with a command execution system, apparatus, or device.
[0529] The computer program contained on a readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0530] The computer program for performing the operations of the present application may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The computer program may be executed entirely on the target electronic device, partially on the target electronic device, executed as a stand-alone software package, partially on the target electronic device and partially on a remote electronic device, or entirely on a remote electronic device or server. In the case of a remote electronic device, the remote electronic device may be connected to the target electronic device through any type of network including a local area network (LAN) or a wide area network (WAN), or may be connected to an external electronic device (e.g., connected through the Internet using an Internet service provider).
[0531] It should be noted that although several units or subunits of the apparatus are mentioned in the foregoing detailed description, such a division is merely exemplary and not mandatory. In fact, according to the embodiments of the present application, the features and functions of two or more of the above-mentioned units may be embodied in one unit. Conversely, the features and functions of one unit described above may be further divided and embodied by multiple units.
[0532] In addition, although the operations of the method of the present application are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all the shown operations must be performed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step for execution, and / or one step may be decomposed into multiple steps for execution.
[0533] Those skilled in the art should understand that the embodiments of the present application may be provided as a method, a system, or a computer program product. Therefore, the present application may take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable computer programs.
[0534] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0535] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0536] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are performed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0537] Although the preferred embodiments of the present application have been described, additional changes and modifications can be made to these embodiments by those skilled in the art once they learn of the basic creative concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications that fall within the scope of the present application.
[0538] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations. < / h5> < / article> < / section>
Claims
1. A message display method, characterized in that, The method includes: Display a chat interface; Display a target message sent by a message sender in the chat interface; In response to a viewing operation on the target message, highlight key content in the target message; wherein, the key content is determined from the target message based on content location information; the content location information is extracted from at least one context message associated with the target message and is information with a prompting effect.
2. The method according to claim 1, wherein If the target message is an audio message or a video message, then highlighting key content in the target message includes: Regarding each message frame corresponding to the specified time represented by the content location information as a target frame, or regarding each message frame containing the key content represented by the content location information as a target frame; jump to the position corresponding to the start frame of each target frame, automatically play the target message starting from the position corresponding to the start frame, and highlight the key content in the target message.
3. The method according to claim 1, characterized in that Highlighting key content in the target message includes: Highlight the key content in the target message; or, circle the key content in the target message; Wherein, the key content is at least one of a target paragraph, a target keyword, and a target entity.
4. The method according to claim 3, characterized in that, If the key content is a target entity, then circling the key content in the target message includes: Circle the target entity in the target message according to the outline of the target entity.
5. The method according to claim 1, wherein When the target message is an audio message, highlighting key content in the target message includes: Convert the audio message into a text message, and highlight the key content in the converted text message.
6. The method according to any one of claims 1 to 5, characterized in that, The method further includes: If it is determined that there is an association between a context message and the target message, present an association control at a first relevant position of the context message; In response to a message association operation triggered by the association control corresponding to the context message, associate the target message with the context message, and present the target message at a second relevant position of the context message.
7. The method according to claim 6, characterized in that, If the association control is a text control describing the message type of the target message, then in response to a message association operation triggered by the association control corresponding to the context message, associating the target message with the context message includes: In response to a message association operation triggered by the association control corresponding to the context message, present a message selection interface corresponding to the message type; wherein, the message selection interface is used to present each candidate historical message that meets the message type and is published by the message sender in the current chat window within a set historical period; In response to a selection operation of selecting the target message from each candidate historical message, associate the context message with the target message.
8. The method according to claim 6, characterized in that, If the association control is a text control describing the message type of the target message, then in response to a message association operation triggered by the association control corresponding to the context message, associating the target message with the context message includes: In response to a message association operation triggered by an associated control corresponding to the context message, at the third relevant position of each candidate historical message, a corresponding message selection control is presented; each of the candidate historical messages is an instant messaging message that meets the message type and is published by the message sender in the current chat window during a set historical period. In response to a selection operation of selecting the target message from the current chat window based on each message selection control, the context message is associated with the target message.
9. The method according to claim 6, characterized in that, The method further includes: In response to a viewing operation on the target message triggered at the second relevant position, the target message is presented, and the key content is highlighted in the target message.
10. The method according to claim 6, wherein The display distance between the target message and the context message in the current chat window exceeds a preset distance, or the number of messages between the target message and the context message exceeds a preset number.
11. The method according to any one of claims 1 to 5, characterized in that, The association between the context message and the target message is determined by at least one of the following methods: The context message contains content existing in the target message. The context message contains at least one of time positioning information and location positioning information.
12. A message display method, characterized in that, The method includes: Obtain at least one context message associated with the target message. Extract the content positioning information of the key content for the target message mentioned in the at least one context message; wherein, the content positioning information is information with a prompting effect. After receiving a viewing request for the target message, the content positioning information is fed back to the client, so that the client, in response to a viewing operation on the target message, highlights the key content in the target message based on the content positioning information.
13. The method according to claim 12, wherein The extracting the content positioning information of the key content for the target message mentioned in the at least one context message includes: If the content positioning information includes location positioning information, match the at least one context message based on a preset first regular expression, and extract the location positioning information in the at least one context message according to the matching result. If the content positioning information includes time positioning information, match the at least one context message based on a preset second regular expression, and extract the time positioning information in the at least one context message according to the matching result. If the content positioning information includes entity description information, extract the entity description information in the at least one context message based on a pre-trained deep learning model; wherein, the deep learning model includes an entity classification layer, and the deep learning model is fine-tuned based on a historical instant messaging message dataset with pre-labeled entity categories.
14. The method according to claim 12, wherein The target message is a non-long text picture message, the content positioning information contains entity description information, and the key content is a target entity. Then the feeding back the content positioning information to the client so that the client highlights the key content in the target message based on the content positioning information includes: Feedback the content location information to the client so that the client can obtain a pre-constructed first entity dictionary; wherein, the first entity dictionary contains multiple first key-value pairs, the key in each first key-value pair is an entity category, and the value in each first key-value pair is an array of expression words associated with the entity category; Identify the target entity category corresponding to the entity in the non-long text picture message; Match the entity description information included in the content location information with the array of expression words associated with the target entity category in the first entity dictionary; If the match is successful, use the corresponding entity as the target entity and highlight the target entity in the non-long text picture message.
15. The method according to claim 14, wherein Construct the first entity dictionary in the following way: Extract the expression words related to the entity from the historical instant messaging messages; Construct an array of expression words associated with each entity category according to the extracted expression words; Based on the constructed array of expression words, store the index relationship between the entity category and the array of expression words in dictionary form to construct the first entity dictionary.
16. The method according to claim 12, wherein The target message is a video message, the content location information includes time location information and entity description information, and the key content is the target entity; Then the feedback of the content location information to the client so that the client highlights the key content in the target message based on the content location information includes: Feedback the content location information to the client so that the client automatically jumps to the target frame position based on the time location information and starts playing the video message automatically from the target frame position; and Identify the entity in the video frame corresponding to the target frame; If there is a target entity in the recognized entities that matches the entity description information included in the content location information, highlight the target entity in the target frame of the video message.
17. The method according to any one of claims 12 to 16, characterized in that If the target message is a video message, determine the context message associated with the target message in the following way: Obtain a pre-constructed second entity dictionary for the video message; wherein, the second entity dictionary contains at least one second key-value pair, the key in each second key-value pair is a video timestamp, and the value in each second key-value pair is the entity category of the entity in the video frame corresponding to the video timestamp; For each context message, match the entity description information in the context message with the second entity dictionary; If the match is successful, determine that the video message is associated with the context message.
18. The method according to claim 17, wherein Construct the second entity dictionary in the following way: Process the video message frame by frame to identify the entity included in each video frame of the video message; Based on the recognition result, when the entity in each video frame included in the video message changes, insert a new second key-value pair into the second entity dictionary, and the video timestamp represented by the key in the new key-value pair is the video timestamp corresponding to the video frame when the change occurs this time.
19. A message display device, characterized in that, Include: A first display unit for displaying a chat interface; Display the target message sent by the message sender in the chat interface; A second display unit is used to highlight key content in the target message in response to a viewing operation on the target message; wherein the key content is determined from the target message based on content positioning information; and the content positioning information is information extracted from at least one context message associated with the target message and has a prompting effect.
20. A message display device, characterized in that, include: A message acquisition unit, used to acquire at least one context message associated with the target message; An information extraction unit, configured to extract content location information for key content of the target message mentioned in the at least one context message; wherein the content location information is information having a prompting function; The feedback unit is used to feed back the content location information to the client after receiving a viewing request for the target message, so that the client responds to the viewing operation for the target message and highlights the key content in the target message based on the content location information.
21. An electronic device, characterized in that, The method comprises a processor and a memory, wherein the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of any one of the methods of claims 1 to 18.
22. A computer-readable storage medium, characterized in that, It comprises a computer program. When the computer program is run on an electronic device, the computer program is used to enable the electronic device to execute the steps of any one of the methods of claims 1 to 18.
23. A computer program product, characterized in that, It includes a computer program, which is stored in a computer-readable storage medium; when a processor of an electronic device reads the computer program from the computer-readable storage medium, the processor executes the computer program, so that the electronic device executes the steps of any method described in claims 1 to 18.