Processing method, terminal device, and storage medium
Patent Information
- Application Number
- CN202610815481.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-05
- Publication Date
- 2026-08-18
AI Technical Summary
[0015] This application, through the aforementioned technical solution, can understand chat information of any modality within the chat content of a target chat window. After obtaining chat information of any modality to be replied to, it can intelligently generate or retrieve a corresponding chat reply based on the understood chat context semantic information. This reduces or avoids the tedious operation of manually entering chat replies after the user has manually understood the chat context semantic information. The technical solution of this application enables chat assistance functionality, solving the problem of how to intelligently assist in replying to chat information of any modality in a chat window, thereby improving the user experience.
Smart Images

Figure CN122601623A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of terminal technology, specifically to a processing method, terminal device, and storage medium. Background Technology
[0002] With the rapid development of mobile internet and smart terminal technologies, social communication applications have become a primary means of daily communication for users. Current chat content in social communication applications exhibits a multimodal trend, encompassing traditional text modalities as well as other non-text modalities. When using social communication applications, users need to understand the various modalities of chat information in the current chat window before manually entering replies.
[0003] In conceiving and implementing this application, the inventors discovered at least the following problem: how to achieve intelligent reply assistance for chat information of any modality in a chat window is a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0004] To address the aforementioned technical problems, this application provides a processing method, terminal device, and storage medium that can intelligently assist in replying to chat messages of any modality in a chat window, making the user's chat reply operation simpler and more convenient.
[0005] In a first aspect, this application provides a processing method applicable to a terminal device, comprising the steps of: obtaining at least one chat message from the chat content of a target chat window, wherein the modality of all chat messages in the chat content includes at least one modality; and outputting at least one chat reply based on the at least one chat message.
[0006] In one embodiment, the processing method of this application further includes at least one of the following: A chat reply must include at least part of a chat message. Chat replies are used to respond to at least one chat message.
[0007] In one implementation, outputting at least one chat reply based on at least one chat message includes: outputting at least one chat reply corresponding to at least one chat message based on at least one of the following: chat content, user profile, and chat context.
[0008] In one implementation, obtaining at least one chat message from the chat content of a target chat window includes: obtaining the chat content of the target chat window within a preset time period, wherein the modality of all chat messages in the chat content includes at least one of text modality, image modality, voice modality, and video modality; and obtaining at least one chat message from the chat content.
[0009] In one embodiment, the processing method of this application further includes: outputting multiple chat replies in the candidate area of the target chat window based on at least one chat message; in response to a selection operation on the multiple chat replies, displaying at least one selected chat reply in the input box of the target chat window, wherein the position of the input box is different from that of the candidate area; or, in response to a selection operation on the multiple chat replies, sending at least one selected chat reply.
[0010] In one implementation, outputting at least one chat reply based on at least one chat message includes: outputting a corresponding clarification prompt for semantically ambiguous content in at least one chat message; and outputting at least one chat reply for replying with explicit semantic content in response to a confirmation operation on the clarification prompt.
[0011] In one implementation, outputting at least one chat reply based on at least one chat message includes: obtaining model input information corresponding to at least one chat message; and outputting at least one chat reply based on the model input information and the AI model.
[0012] In one embodiment, the processing method of this application further includes at least one of the following: For semantically ambiguous content detected from at least one chat message, output the corresponding clarification prompt; In response to the clarification confirmation action for the clarification prompt, correct the corresponding chat message; In response to a clarification confirmation action for a clarification prompt, update the model input information.
[0013] Secondly, this application also provides a terminal device, including: a memory and a processor, wherein the memory stores a processing program or instructions, and the processing program or instructions, when executed by the processor, implement the steps of the method described above.
[0014] Secondly, this application also provides a storage medium storing a computer program or instructions, which, when executed by a terminal device, implement the steps of the processing method described above.
[0015] This application, through the aforementioned technical solution, can understand chat information of any modality within the chat content of a target chat window. After obtaining chat information of any modality to be replied to, it can intelligently generate or retrieve a corresponding chat reply based on the understood chat context semantic information. This reduces or avoids the tedious operation of manually entering chat replies after the user has manually understood the chat context semantic information. The technical solution of this application enables chat assistance functionality, solving the problem of how to intelligently assist in replying to chat information of any modality in a chat window, thereby improving the user experience. Attached Figure Description
[0016] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application. To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, those skilled in the art can obtain other drawings based on these drawings without any creative effort.
[0017] Figure 1 This is a schematic diagram of the hardware structure of a mobile terminal provided in an embodiment of this application.
[0018] Figure 2 This is a communication network system architecture diagram provided for an embodiment of this application.
[0019] Figure 3 This is a flowchart illustrating a processing method provided in an embodiment of this application.
[0020] Figure 4 This is a flowchart illustrating another processing method provided in an embodiment of this application.
[0021] The realization of the objectives, functional features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. The accompanying drawings have illustrated specific embodiments of this application, which will be described in more detail below. These drawings and textual descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concepts of this application to those skilled in the art through reference to specific embodiments. Detailed Implementation
[0022] It should be understood that although the terms "first," "second," "third," etc., may be used herein to describe various information, these terms are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. The terms "or," "and / or," and "including at least one of the following," as used in this application, can be interpreted as inclusive or meaning any one or any combination thereof. For example, "including at least one of the following: A, B, C" means "any one of the following: A; B; C; A and B; A and C; B and C; A and B and C," and similarly, "A, B, or C" or "A, B, and / or C" means "any one of the following: A; B; C; A and B; A and C; B and C; A and B and C." Exceptions to this definition only occur when combinations of elements, functions, steps, or operations are inherently mutually exclusive in some manner.
[0023] It should be understood that although the steps in the flowcharts of this application's embodiments are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least a portion of the sub-steps or stages of other steps.
[0024] It should be noted that step designations such as S11 and S12 are used in this document for the purpose of more clearly and concisely describing the corresponding content, and do not constitute a substantial limitation on the order. In specific implementation, those skilled in the art may execute S12 first and then S11, etc., but these should all be within the protection scope of this application.
[0025] It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.
[0026] In the following description, the use of suffixes such as "module," "part," or "unit" to denote elements is solely for the purpose of illustrative purposes and has no specific meaning in itself. Therefore, "module," "part," or "unit" may be used interchangeably.
[0027] Terminal devices can be implemented in various forms. For example, the terminal devices described in this application may include terminal devices such as mobile phones, tablets, laptops, handheld computers, personal digital assistants (PDAs), portable media players (PMPs), navigation devices, wearable devices, smart bracelets, pedometers, and fixed terminals such as digital TVs and desktop computers.
[0028] The following description will use a mobile terminal as an example. Those skilled in the art will understand that, apart from elements specifically designed for mobile purposes, the construction according to the embodiments of this application can also be applied to fixed-type terminals.
[0029] Please see Figure 1This is a schematic diagram of the hardware structure of a mobile terminal implementing various embodiments of this application. The mobile terminal 100 may include: an RF (Radio Frequency) unit 101, a WiFi module 102, an audio output unit 103, an A / V (Audio / Video) input unit 104, a sensor 105, a display unit 106, a user input unit 107, an interface unit 108, a memory 109, a processor 110, and a power supply 111, etc. Those skilled in the art will understand that... Figure 1 The mobile terminal structure shown does not constitute a limitation on the mobile terminal. The mobile terminal may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0030] The following is combined with Figure 1 A detailed introduction to each component of the mobile terminal: The radio frequency unit 101 can be used for receiving and transmitting signals during information transmission or calls. Specifically, it receives downlink information from the base station and processes it with the processor 110; additionally, it transmits uplink data to the base station. Typically, the radio frequency unit 101 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low-noise amplifier, and a duplexer. Furthermore, the radio frequency unit 101 can also communicate wirelessly with networks and other devices. The aforementioned wireless communications may use any communication standard or protocol, including but not limited to GSM (Global System of Mobile communication), GPRS (General Packet Radio Service), CDMA2000 (Code Division Multiple Access 2000), WCDMA (Wideband Code Division Multiple Access), TD-SCDMA (Time Division-Synchronous Code Division Multiple Access), FDD-LTE (Frequency Division Duplexing-Long Term Evolution), TDD-LTE (Time Division Duplexing-Long Term Evolution), 5G, and 6G.
[0031] WiFi is a short-range wireless transmission technology. Mobile terminals, through the WiFi module 102, can help users send and receive emails, browse web pages, and access streaming media, providing users with wireless broadband internet access. Although Figure 1 WiFi module 102 is shown, but it is understood that it is not a necessary component of a mobile terminal and can be omitted as needed without changing the nature of the invention.
[0032] The audio output unit 103 can convert audio data received by the radio frequency unit 101 or the WiFi module 102 or stored in the memory 109 into audio signals and output them as sound when the mobile terminal 100 is in call signal receiving mode, call mode, recording mode, voice recognition mode, broadcast receiving mode, etc. Furthermore, the audio output unit 103 can also provide audio output related to specific functions performed by the mobile terminal 100 (e.g., call signal receiving sound, message receiving sound, etc.). The audio output unit 103 may include a speaker, a buzzer, etc.
[0033] The A / V input unit 104 is used to receive audio or video signals. The A / V input unit 104 may include a graphics processing unit (GPU) 1041 and a microphone 1042. The GPU 1041 processes image data of still images or videos acquired by an image capture device (such as a camera) in video capture mode or image capture mode. The processed image frames can be displayed on the display unit 106. The image frames processed by the GPU 1041 can be stored in the memory 109 (or other storage media) or transmitted via the radio frequency unit 101 or the WiFi module 102. The microphone 1042 can receive sound (audio data) in operating modes such as telephone call mode, recording mode, and voice recognition mode, and can process such sound into audio data. The processed audio (voice) data can be converted into a format that can be transmitted to a mobile communication base station via the radio frequency unit 101 in telephone call mode. The microphone 1042 can implement various types of noise cancellation (or suppression) algorithms to eliminate (or suppress) noise or interference generated during the reception and transmission of audio signals.
[0034] The mobile terminal 100 also includes at least one sensor 105, such as a light sensor, a motion sensor, and other sensors. Optionally, the light sensor includes an ambient light sensor and a proximity sensor. Optionally, the ambient light sensor can adjust the brightness of the display panel 1061 according to the ambient light level, and the proximity sensor can turn off the display panel 1061 and / or backlight when the mobile terminal 100 is moved to the ear. As a type of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in various directions (generally three axes), and can detect the magnitude and direction of gravity when stationary. It can be used for applications that recognize the phone's posture (such as landscape / portrait switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc. Other sensors that may be configured in the phone, such as fingerprint sensors, pressure sensors, iris sensors, molecular sensors, gyroscopes, barometers, hygrometers, thermometers, and infrared sensors, will not be described in detail here.
[0035] The display unit 106 is used to display information input by the user or information provided to the user. The display unit 106 may include a display panel 1061, which may be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), or the like.
[0036] User input unit 107 can be used to receive input numerical or character information, and generate key signal inputs related to user settings and function control of the mobile terminal. Optionally, user input unit 107 may include touch panel 1071 and other input devices 1072. Touch panel 1071, also known as touch screen, can collect touch operations on or near the user (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near touch panel 1071), and drive corresponding connection devices according to a pre-set program. Touch panel 1071 may include two parts: a touch detection device and a touch controller. Optionally, the touch detection device detects the user's touch position and the signal generated by the touch operation, and transmits the signal to the touch controller; the touch controller receives touch information from the touch detection device, converts it into touch point coordinates, sends it to processor 110, and can receive and execute commands sent by processor 110. In addition, touch panel 1071 can be implemented using various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch panel 1071, the user input unit 107 may also include other input devices 1072. Optionally, other input devices 1072 may include, but are not limited to, one or more of the following: physical keyboard, function keys (such as volume control buttons, power buttons, etc.), trackball, mouse, joystick, etc., without being specifically limited here.
[0037] Optionally, the touch panel 1071 may cover the display panel 1061. When the touch panel 1071 detects a touch operation on or near it, it transmits the information to the processor 110 to determine the type of touch event. Subsequently, the processor 110 provides corresponding visual output on the display panel 1061 based on the type of touch event. Although in Figure 1 In this embodiment, the touch panel 1071 and the display panel 1061 are two independent components to realize the input and output functions of the mobile terminal. However, in some embodiments, the touch panel 1071 and the display panel 1061 can be integrated to realize the input and output functions of the mobile terminal. The specific implementation is not limited here.
[0038] Interface unit 108 serves as an interface through which at least one external device can connect to mobile terminal 100. For example, the external device may include a wired or wireless headset port, an external power supply (or battery charger) port, a wired or wireless data port, a memory card port, a port for connecting a device with an identification module, an audio input / output (I / O) port, a video I / O port, a headphone port, and so on. Interface unit 108 may be used to receive input (e.g., data, power, etc.) from the external device and transmit the received input to one or more elements within mobile terminal 100, or it may be used to transmit data between mobile terminal 100 and the external device.
[0039] The memory 109 can be used to store software programs and various data. The memory 109 may primarily include a program storage area and a data storage area. Optionally, the program storage area may store the operating system, applications required for at least one function (such as sound playback, image playback, etc.), etc.; the data storage area may store data created based on the use of the mobile phone (such as audio data, phonebook, etc.). Furthermore, the memory 109 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.
[0040] The processor 110 is the control center of the mobile terminal. It connects various parts of the mobile terminal via various interfaces and lines. By running or executing software programs and / or modules stored in the memory 109, and by calling data stored in the memory 109, it performs various functions and processes data of the mobile terminal, thereby providing overall monitoring of the mobile terminal. The processor 110 may include one or more processing units; preferably, the processor 110 may integrate an application processor and a modem processor. Optionally, the application processor mainly handles the operating system, user interface, and applications, while the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 110.
[0041] The mobile terminal 100 may also include a power supply 111 (such as a battery) that supplies power to various components. Preferably, the power supply 111 can be logically connected to the processor 110 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system.
[0042] although Figure 1 As not shown, the mobile terminal 100 may also include a Bluetooth module, etc., which will not be described in detail here.
[0043] To facilitate understanding of the embodiments of this application, the communication network system on which the mobile terminal of this application is based is described below.
[0044] Please see Figure 2 , Figure 2 This application provides a communication network system architecture diagram. The communication network system is an LTE system based on the universal mobile communication technology. The LTE system includes a UE (User Equipment) 201, an E-UTRAN (Evolved UMTS Terrestrial Radio Access Network) 202, an EPC (Evolved Packet Core) 203, and the operator's IP services 204, which are connected in sequence.
[0045] Optionally, UE201 can be the aforementioned terminal 100, which will not be described in detail here.
[0046] E-UTRAN202 includes eNodeB2021 and other eNodeB2022s. Optionally, eNodeB2021 can connect to other eNodeB2022s via backhaul (e.g., X2 interface). eNodeB2021 connects to EPC203 and can provide UE201 with access to EPC203.
[0047] EPC203 may include an MME (Mobility Management Entity) 2031, an HSS (Home Subscriber Server) 2032, other MMEs 2033, an SGW (Serving Gateway) 2034, a PGW (Packet Data Network Gateway) 2035, and a PCRF (Policy and Charging Rules Function) 2036, etc. Optionally, MME2031 is the control node that handles signaling between UE201 and EPC203, providing bearer and connection management. HSS2032 is used to provide registers to manage functions such as the Home Location Register (not shown in the figure) and stores user-specific information such as service characteristics and data rates. All user data can be sent through SGW2034. PGW2035 can provide UE 201 IP address allocation and other functions. PCRF2036 is the policy and charging control decision point for service data flow and IP bearer resources. It selects and provides available policy and charging control decisions for the policy and charging enforcement function unit (not shown in the figure).
[0048] IP services 204 may include the Internet, intranet, IMS (IP Multimedia Subsystem), or other IP services.
[0049] Although the above description uses the LTE system as an example, those skilled in the art should know that this application is not only applicable to the LTE system, but also to other wireless communication systems, such as GSM, CDMA2000, WCDMA, TD-SCDMA, 5G and future new network systems (such as 6G), etc., without limitation.
[0050] Based on the above-described mobile terminal hardware structure and communication network system, various embodiments of this application are proposed.
[0051] First Embodiment Reference Figure 3 , Figure 3 An exemplary flowchart of a processing method is shown. The processing method of this embodiment can be applied to terminal devices (e.g., mobile phones, computers, tablets, etc.), including: S11: Obtain at least one chat message from the chat content of the target chat window, wherein the modality of all chat messages in the chat content includes at least one modality; S12: Output at least one chat reply based on at least one chat message.
[0052] Optionally, the target chat window may represent the interface of a social communication application to which a chat reply is to be processed and / or specified by the user or system.
[0053] Optionally, the target chat window includes, but is not limited to, at least one of the following: a chat window selected through a window selection operation, a chat window displayed the day before yesterday, or a chat window of a selected specific social communication application.
[0054] Optionally, chat content can represent a set of chat messages organized along a time dimension, encompassing one or more modalities. Chat content can also carry semantic information related to the chat context.
[0055] Optionally, the modality of chat messages can characterize the information form type presented by chat messages at the perception and expression levels.
[0056] Optionally, the modalities of all chat messages in the chat content include, but are not limited to, one or more modalities such as text modal, image modal, voice modal, and video modal, and also include emoticon modal.
[0057] Optionally, the text module can represent a semantic expression form that uses character symbols (such as text symbols, number symbols, etc.) as information carriers and relies on visual perception.
[0058] Optionally, the image module can represent a time-aware form of information expression that uses a static two-pixel matrix as the information carrier.
[0059] Optionally, the speech module can represent the form of information transmission that uses time-series sound wave signals as information carriers and relies on auditory perception.
[0060] Optionally, the video module can represent an information transmission form that uses temporally correlated audio-visual composite data or video data as the information carrier and relies on visual perception or visual and auditory fusion perception.
[0061] Optionally, an emoji modality can represent an expression of emotional meaning based on visual perception, using one or more regular or irregular symbols as information carriers.
[0062] Optionally, chat messages can represent semantic units that can be arranged chronologically in the target chat window and carried in a specific modality.
[0063] Optionally, the one or more chat messages obtained can be chat messages to be replied to according to preset rules or chat messages that can be used as a reference for chat dialogue.
[0064] Optionally, the preset rules can characterize the selection and determination rules used to filter and determine chat information to be processed from the chat content.
[0065] Optionally, the preset rules include, but are not limited to, at least one of the following: Chat messages marked as unread are identified as messages awaiting a reply. The most recently received chat message in the chat content is identified as the chat message to be replied to; The chat information selected through the selection operation will be identified as chat information to be replied to or as chat information to be used as a reference in the chat conversation.
[0066] Optionally, chat replies have a clear dialogue response function, the core purpose of which is to respond to the intent, questions or emotions carried by one or more chat messages.
[0067] Optionally, the modality of the chat reply can be consistent with or differ from the modality of some or all of the acquired chat information; thus, the technical solution of this embodiment can both generate chat replies of the same modality (such as replying to text with text) and generate chat replies across modalities (such as replying to text with emoticons), so as to realize the multimodal chat reply function.
[0068] Optionally, the processing method of this embodiment also includes, but is not limited to, at least one of the following: A chat reply must include at least part of a chat message. Chat replies are used to respond to at least one chat message.
[0069] Optionally, the chat content of the target chat window can be obtained directly through a specific API call of the social communication application to which the target chat window belongs, and / or it can be obtained by collecting and processing the chat information presented in the target chat window through system-level functions of the terminal device (such as system-level screen recording, screenshot, audio recording, etc.).
[0070] Optionally, the technical solution of this embodiment can obtain the chat content of the target chat window through the system-level function of the terminal device, and then implement the processing method of this embodiment through the system-level function. In this way, the technical solution of this embodiment can realize the system-level chat assistance function, so that the chat reply corresponding to the chat information to be replied to can be intelligently generated or obtained in various social communication applications of the terminal device. In other words, the system-level chat assistance function implemented in this embodiment can support cross-application use.
[0071] Optionally, the technical solution of this embodiment can generate chat replies by selecting chat replies corresponding to at least one chat message based on the matching method of a rule engine, and / or by intelligently generating chat replies corresponding to at least one chat message based on the content generation capability of an AI model.
[0072] Optionally, step S12: outputting at least one chat reply based on at least one chat message, including: obtaining model input information corresponding to at least one chat message; and outputting at least one chat reply based on the model input information and the AI model.
[0073] Optionally, the model input information can represent the normalized data form that drives the AI model to perform the response generation task.
[0074] Optionally, the technical solution of this embodiment can leverage the powerful semantic understanding, knowledge reasoning, and content generation capabilities of AI models (such as pre-trained large language models, multimodal large models, etc.) to generate natural, fluent, semantically coherent, and socially appropriate response content, breaking through the upper limit of the expressive capabilities of traditional template matching or rule engines.
[0075] For example, step S12: output at least one chat reply based on at least one chat message. The specific implementation method is as follows: perform multi-module information understanding and semantic extraction on one or more chat messages to form a structured semantic vector (equivalent to the aforementioned model input information); generate reply content based on the chat context semantic information and semantic vector provided by the chat content through a generative model (equivalent to the aforementioned AI model) to output at least one chat reply.
[0076] Optionally, output at least one chat reply based on at least one chat message, including: outputting a corresponding clarification prompt for semantically ambiguous content in at least one chat message; and outputting at least one chat reply for replying to explicit semantic content in response to a confirmation operation on the clarification prompt.
[0077] Optionally, before outputting a corresponding clarification prompt for semantically ambiguous content in at least one chat message, the following steps may be included: detecting whether the chat message contains referential keywords; when referential keywords are detected in the chat message, performing a completion determination based on the chat content; and when the determination is that the content cannot be incomplete, determining that there is semantically ambiguous content in the chat message.
[0078] Optionally, referential keywords can represent words and / or phrases used to refer to or identify entities such as people, things, places, and times, for example, this, here, over here, that, there, over there, like that, this person, that person, this person, that person, he, you, etc.
[0079] For example, if the semantic content of a specific chat message is "I want the big one", and it is determined that there is an unclear meaning based on the chat content or the context of the specific chat message, then the specific chat message is deemed to have semantic ambiguity. In this case, a clarification prompt is output: "What exactly is the big one you are referring to?" After the user manually or verbally inputs the specific clarification content "I want the big watermelon" (i.e., the confirmation operation for the clarification prompt), the specific chat message is corrected to the corresponding clear semantic content, thereby allowing at least one chat reply corresponding to the corrected specific chat message to be output.
[0080] Optionally, the processing method of this embodiment further includes at least one of the following: For semantically ambiguous content detected from at least one chat message, output the corresponding clarification prompt; In response to the clarification confirmation action for the clarification prompt, correct the corresponding chat message; In response to a clarification confirmation action for a clarification prompt, update the model input information.
[0081] Optionally, the updated model input information can be fed into the AI model so that the AI model outputs at least one chat reply.
[0082] Optionally, the technical solution of this embodiment can clarify and correct the semantically ambiguous content in the chat information to obtain chat information and / or model input information with clear semantic content, thereby enabling intelligent generation of chat replies that are close to the user's actual information reply needs, and thus improving the user experience.
[0083] Optionally, the processing method provided in this embodiment may further include: constructing a knowledge base for the AI model based on the type tags and chat content of historical chat windows.
[0084] Optionally, step S12: Output at least one chat reply based on at least one chat message, including: obtaining the target type tag corresponding to the target chat window; obtaining model input information based on at least one chat message and the target type tag; calling reference chat content corresponding to the target type tag in the knowledge base through the AI model, and performing reply generation processing through the AI model based on the model input information and the reference chat content to obtain at least one chat reply corresponding to at least one chat message.
[0085] Optionally, the type tags include social relationship type tags and chat scenario type tags.
[0086] Understandably, chat content tends to be repetitive for the same type of social relationship or chat scenario. For example, if a user's job responsibilities include responding to legal inquiries, and the target chat window's social relationship type tag is "legal inquiries group," then the topics of inquiry and the user's responses within this group tend to be repetitive. Therefore, the AI model can access historical chat content related to the legal inquiries group as a reference for generating candidate chat responses. This allows the AI model to generate responses that better match the user's needs and preferred replying style.
[0087] The technical solution of this embodiment can enable AI models to learn and inherit the distribution of high-frequency topics, standard response paradigms, and professional terminology systems under specific categories by calling knowledge bases driven by type tags.
[0088] The processing method of this embodiment can be applied to terminal devices, including step S11: obtaining at least one chat message from the chat content of the target chat window, wherein the modality of all chat messages in the chat content includes at least one modality; S12: outputting at least one chat reply based on the at least one chat message. Thus, the technical solution of this embodiment can understand chat messages of any modality in the chat content of the target chat window. After obtaining any chat message to be replied to in any modality, it can intelligently generate or obtain a chat reply corresponding to the chat message to be replied to based on the understood chat context semantic information, thereby reducing or avoiding the tedious operation of manually inputting chat replies after the user has manually understood the chat context semantic information. Through the technical solution of this embodiment, chat assistance functions can be realized, solving the problem of how to achieve intelligent reply assistance for chat messages of any modality in the chat window, thereby improving the user experience.
[0089] Second Embodiment Based on the technical concept of the first embodiment of this application, this embodiment provides another processing method, including: S21: Obtain at least one chat message from the chat content of the target chat window, wherein the modality of all chat messages in the chat content includes at least one modality; S22: Based on at least one of the following three factors—chat content, user profile, and chat context—output at least one chat reply corresponding to at least one chat message.
[0090] Optionally, the modalities of all chat messages in the chat content include, but are not limited to, one or more modalities such as text modal, image modal, voice modal, video modal, and emoticon modal.
[0091] Optionally, the chat content of the target chat window can be obtained directly through a specific API call or backend database of the social communication application to which the target chat window belongs, and / or it can be obtained by collecting and processing the chat information presented in the target chat window through system-level functions of the terminal device (such as system-level screen recording, screenshot, audio recording, etc.).
[0092] Optionally, the technical solution of this embodiment can obtain the chat content of the target chat window through the system-level function of the terminal device, and then implement the processing method of this embodiment through the system-level function. In this way, the technical solution of this embodiment can realize the system-level chat assistance function, so that no matter what social communication application the user is using on the terminal device, chat information can be obtained through a unified system-level collection mechanism, realizing the cross-application chat assistance function of "one-time deployment, global application", significantly improving the applicability and deployment flexibility of the technical solution.
[0093] Optionally, the chat content of the target chat window is obtained by collecting chat information presented in the target chat window through the system-level functions of the terminal device and then processing it, including but not limited to at least one of the following: Capture a screenshot of the target chat window using the screenshot function, and / or add a corresponding timestamp; Record the video data stream displayed in the target chat window using the screen recording function, and / or add corresponding timestamps; The audio recording function records the first audio data in the current environment and / or adds the corresponding timestamp; The audio recording function records the second audio data played by the player on the terminal device, and / or adds the corresponding timestamp; The system performs alignment and fusion processing on at least one of the following four data sources: a window screenshot with added timestamps, a video data stream, first audio data, and second audio data, to obtain chat content that can carry semantic information about the chat context.
[0094] Optionally, the technical solution of this embodiment can simultaneously capture dual-channel information of visual presentation and auditory output through a system-level acquisition method: window screenshots and video data streams record the visual content of the screen, the first audio data collects environmental speech (such as the user's voice), and the second audio data captures the audio output of the device. These four elements work together to completely restore the multimodal presentation state of the chat window, avoiding information loss or semantic breaks caused by a single acquisition method, and providing a high-fidelity raw data foundation for subsequent intelligent understanding. Furthermore, by adding timestamps to various types of acquired data, alignment processing can be performed in the time domain, thereby accurately restoring the chronological order and / or concurrent relationships of chat information. This ensures that the semantic content corresponding to the acquired data maintains logical coherence during fusion and reconstruction, enabling the generated structured chat context semantic information to accurately reflect the dialogue evolution process, thereby improving the contextual adaptability of the generated response.
[0095] Optionally, the alignment and blending process includes, but is not limited to, at least one of the following: Convert timestamped audio data (such as first audio data, second audio data) into text content; Perform semantic parsing on the screenshot of the window with added timestamps to obtain the parsed content; Perform semantic parsing on the timestamped video data stream to obtain the video parsing content; Align at least one of the text content, screenshot analysis content, and video analysis content in the temporal and semantic domains to obtain the alignment result. The alignment results are processed by a fusion model to extract features and reconstruct the data, generating or outputting structured chat context semantic information as the chat content.
[0096] Optionally, the window screenshot includes at least one of the following: text-modal chat messages, emoji-modal chat messages, image-modal chat messages, and other visual chat messages.
[0097] Optionally, the technical solution of this embodiment can transform heterogeneous modal data into a unified semantic representation (such as text content, screenshot parsing content, video parsing content, etc.), and then perform feature extraction and fusion reconstruction through a fusion model. This can ensure that the semantic content corresponding to the collected data remains logically coherent during fusion reconstruction, so that the generated structured chat context semantic information accurately reflects the dialogue evolution process, thereby improving the context adaptability of the generated response.
[0098] Optionally, before performing semantic parsing on the window screenshots, multiple collected window screenshots can be stitched together as needed to obtain a stitched window screenshot, so that the stitched window screenshot can reflect the historical dialogue information as a whole.
[0099] Optionally, semantic parsing is performed on the window screenshot, including at least one of the following: The text content in a screenshot is analyzed using OCR recognition technology to obtain the screenshot text parsing content; Image semantic analysis is used to identify and parse emojis in window screenshots to obtain the emoji parsing content of the screenshots; Image semantic analysis is used to identify and parse images in window screenshots to obtain the content of the screenshot images; To obtain the screenshot parsing content, at least one of the three—screenshot text parsing content, screenshot emoticon parsing content, and screenshot image parsing content—must be rearranged according to the arrangement of the contents in the window screenshot.
[0100] Optionally, the audio content can be converted into text content using this speech-to-text module (such as an ASR system).
[0101] Optionally, semantic parsing of the video data stream to obtain the video parsing content can be achieved using a mature video parsing model.
[0102] Optionally, the technical solution of this embodiment can obtain the chat content of the target chat window through the system-level function of the terminal device, and then implement the processing method of this embodiment through the system-level function. In this way, the technical solution of this embodiment can realize the system-level chat assistance function, so that the chat reply corresponding to the chat information to be replied to can be intelligently generated or obtained in various social communication applications of the terminal device. In other words, the system-level chat assistance function implemented in this embodiment can support cross-application use.
[0103] Optionally, a user profile can represent personalized information or a set of personalized information that describes at least one of a user's social attributes, expressive characteristics, and identity positioning.
[0104] Optionally, the user profile includes, but is not limited to, at least one of the following: social relationships with contacts in the target chat window, user conversation habits, and user identity location information.
[0105] Optionally, the social relationship with the contacts in the target chat window can characterize and define the social connection attributes between the user and the contacts in the target chat window. Social relationships can encompass the level of intimacy, relationship type, and historical interaction patterns between the user and specific contacts. Its main function is to provide a basis for social distance calibration in response generation, determine the address system, tone scale, topic boundaries, and etiquette norms, so that the output response conforms to the communication expectations in the context of a specific social relationship.
[0106] Optionally, user dialogue habit information can represent feature information or a set of feature information that characterizes a user's stable language output pattern, such as catchphrases, language organization, sentence structure preferences, rhetorical style, and vocabulary selection tendencies—personalized expression identifiers. The main function of user dialogue habit information is to provide a style transfer benchmark for response generation, enabling machine-generated content to simulate the user's unique language fingerprint, thereby enhancing the authenticity, sense of identity, and consistency of the response.
[0107] Optionally, user identity positioning information can represent the characteristic information or set of characteristic information of the positioned user's identity and status in a specific chat context. User identity positioning information, for example, covers the user's role responsibilities, power level, obligation boundaries, and image requirements in scenarios such as the workplace, family, and community. Its main function is to provide a basis for anchoring the stance in response generation, ensuring that the response content conforms to the user's social identity expectations, maintaining role consistency, and avoiding the risk of identity conflict.
[0108] Optionally, the chat context can represent dynamic contextual information describing at least one of the following: the user's current spatiotemporal environment, the context of recent activities, and the chat mode.
[0109] Optionally, the chat scenarios may include at least one of the following: the current environment, recent activity patterns, device status information, and chat modes (such as work chat mode or personal chat mode).
[0110] Optionally, recent activity patterns can represent statistical contextual information about the sequence of user behaviors and their derived characteristics within a specific time window. The main function of recent activity patterns is to provide user behavior context for response generation, support the prediction of users' current interests and potential needs, and optimize response timing and content adaptation.
[0111] Optionally, device status information can characterize real-time hardware context information describing the current operating status, resource conditions, and functional configuration of the terminal device, such as foreground running applications, background service processes, power level, storage capacity, computing load, network type, signal strength, etc.
[0112] Optionally, chat modes can represent categorized contextual information that defines the functional context and behavioral norms to which a user's current social interaction belongs, such as work chat mode, personal chat mode, and business negotiation mode. The main function of chat modes is to provide a social domain framework for response generation, enabling accurate switching and adherence to norms in contexts such as formal / casual, task / emotional, and public / private.
[0113] Optionally, the current environmental scenario can characterize real-time contextual information describing the user's current spatiotemporal physical state and external natural conditions, such as current time, date, holidays, time period characteristics, geographical location, location type, movement status, weather conditions, temperature, lighting, and noise level. The main function of the current environmental scenario is to provide spatiotemporal background constraints for response generation, influencing the urgency of the response, the level of detail, the choice of greeting, and the expected usability.
[0114] Optionally, obtaining at least one chat message from the chat content of the target chat window may include: obtaining the chat content of the target chat window within a preset time period; and obtaining at least one chat message from the chat content.
[0115] Optionally, the technical solution of this embodiment limits a preset time period and extracts a portion of the chat content from the complete chat content within a specific time window. This effectively filters out redundant historical information, focuses on the relevant context of the current conversation, avoids noise interference and waste of computing resources caused by the full historical chat content, improves the targeting of information acquisition and processing efficiency, and ensures that subsequent replies are generated based on the context that is closest to the current interaction needs.
[0116] Optionally, step S22: before outputting at least one chat reply corresponding to at least one chat message based on at least one of the three factors—chat content, user profile, and chat context—includes: outputting a corresponding clarification prompt for semantically ambiguous content in at least one chat message; and correcting the corresponding chat message in response to a confirmation operation for the clarification prompt.
[0117] Optionally, before outputting a corresponding clarification prompt for semantically ambiguous content in at least one chat message, the process may include: detecting whether the chat messages in the chat content include referential keywords, and / or detecting whether at least one chat message to be replied to from the chat content includes referential keywords; when referential keywords are detected in the chat messages, performing completion processing on the chat messages including referential keywords based on the contextual semantic information provided by the chat content to obtain completion processing results; filtering chat messages that still include referential keywords in the completion processing results, and determining that there is semantically ambiguous content in the chat messages that still include referential keywords. Thus, the technical solution of this embodiment can intelligently and automatically complete the chat messages containing semantically ambiguous content based on the chat context semantic information provided by the chat content when it is initially determined that there is semantically ambiguous content in the chat messages. If the aforementioned completion process still fails to eliminate at least part of the semantically ambiguous content in the chat messages, it can further correct the semantically ambiguous content in at least part of the chat messages by outputting clarification prompts and receiving confirmation operations for the clarification prompts (such as manual completion operations by the user). This ensures that the chat messages in the chat content and / or the chat messages to be replied to have clear semantic content, so that subsequent chat replies that are more in line with the user's expectations can be obtained or generated for chat messages with clear semantic content, thereby improving the user experience; and / or, ensure that all chat messages in the corrected chat content have clear semantic content, thereby enabling the corrected chat content to provide more accurate chat context semantic information.
[0118] Optionally, step S22: outputting at least one chat reply corresponding to at least one chat message based on at least one of the three factors: chat content, user profile, and chat context, may include: obtaining model input information corresponding to at least one chat message based on at least one of the three factors: chat content, user profile, and chat context; and outputting at least one chat reply based on the model input information and the AI model.
[0119] Thus, the technical solution of this embodiment can comprehensively consider at least one of the three factors—chat content, user profile, and chat context—to construct model input information that meets the current scenario, thereby enabling the AI model to generate chat responses that better meet user expectations. Specifically, the chat content provides semantic information about the chat context, the user profile injects personalized expression styles, and the chat context provides dynamic contextual constraints. The collaboration of these three factors forms a three-dimensional decision input, allowing the AI model's chat response generation to simultaneously meet the triple standards of semantic relevance, style consistency, and contextual adaptation. Furthermore, through information such as social relationships, conversational habits, and identity positioning provided by the user profile, the AI model can simulate the language fingerprint and social identity of a specific user. This allows the generated chat responses to move beyond generic expressions and become personalized outputs tailored to the specific user's unique catchphrases, sentence structure preferences, and role stance, enhancing the realism and user identification of the chat responses and achieving an expression simulation effect as if the specific user were speaking themselves. Furthermore, chat scenarios can provide a dynamic constraint framework for chat reply generation, enabling AI models to automatically perceive the user's current spatiotemporal state and interaction conditions, and adjust the urgency, detail, modality selection, and formality of the reply in real time, thereby achieving contextualized intelligent responses that are appropriate to the time and place.
[0120] Optionally, the processing method provided in this embodiment may further include: outputting multiple chat replies in the candidate area of the target chat window based on at least one chat message; in response to a selection operation on the multiple chat replies, displaying at least one selected chat reply in the input box of the target chat window; or, in response to a selection operation on the multiple chat replies, sending at least one selected chat reply.
[0121] Optionally, the input box and the candidate area can be in different or the same position. Thus, when the positions are the same, a compact overlay can be formed, reducing the distance the eye moves; when the positions are different, regional functions can be separated, avoiding the candidate responses from obscuring or interfering with the input content.
[0122] Optionally, the technical solution of this embodiment outputs multiple chat replies in the candidate area, leaving the final decision-making power to the user, forming a collaborative interaction mode of machine generation, user filtering, and confirmation of sending; in addition, the technical solution of this embodiment also provides a dual-path response mechanism of input box display and direct sending, wherein the response mechanism of input box display supports secondary editing and polishing by the user; the response of direct sending can realize one-click quick confirmation in a timely manner, which is convenient for dialogue communication that pursues efficiency.
[0123] Optionally, the user's selection actions can reflect the user's personalized preferences (such as style preference, detail preference, and emotional tone). In this way, selection behavior data can be continuously collected to optimize the model parameters or knowledge base of the AI model, forming a positive feedback loop of generation, selection, learning, and fine-tuning, and driving continuous improvement in the quality of chat response generation by the AI model.
[0124] In some implementations, chat reply generation technology mainly includes two types: one is input method AI-assisted technology, and the other is large model-assisted polishing.
[0125] The main working principle of AI-assisted input method technology is as follows: after a user enters part of the text in a chat or text input box, they click the "AI-assisted writing" or "optimized expression" button to trigger the function. The input method application calls its built-in Natural Language Processing (NLP) algorithm to generate subsequent sentences or optimized text based on the currently entered text. Therefore, AI-assisted input method technology relies on the input method engine and the target application interface, and can only read the text that the user has pre-entered in the input box and then generate the corresponding result. The shortcomings of AI-assisted input method technology are: 1. Lack of contextual understanding: Relying on the input method engine and application interface, it cannot obtain the complete chat context, and can only process the partial text in the current input box, without obtaining or parsing the complete chat context; 2. Insufficient personalization: The generated content is mostly templated or universal sentence patterns, lacking personalized customization based on the user's identity, relationship, and habits; 3. Poor situational adaptability: It is difficult to dynamically optimize the generation strategy based on the user's historical usage patterns or real-time environment (time, location, device status, etc.); 4. Dependence on user input triggering: Generation or optimization can only be executed after the user manually enters certain content, and it cannot actively respond to the rhythm of the conversation. The main working method of large-model-assisted polishing is as follows: users copy and paste chat or writing content from the original application to a standalone AI application or AI webpage, send a polishing or generation request to the large language model (LLM), the LLM processes the content and displays the polished content, and then the user manually copies the displayed polished content and pastes it back to the original application. Therefore, large-model-assisted polishing relies on users actively inputting or pasting content and cannot directly access the text content and context of the original application. The drawbacks of large-model-assisted polishing are: 1. Cumbersome steps: It requires multiple manual copying and pasting of content, which cannot be completed instantly in the original chat interface; 2. Scene fragmentation: The model's operating environment is independent of the chat application, and it cannot seamlessly integrate with social scenarios; 3. Poor scene adaptability: The generated results are more suitable for writing or copywriting, rather than customized for short texts in social conversations. The common shortcomings of AI-assisted input method technology and large-scale model-assisted editing are: 1. Insufficient context understanding and acquisition capabilities: It cannot obtain complete context information, including historical messages, images, and emoticons, in real time across applications and conversations, resulting in a disconnect between the reply and the actual dialogue environment; 2. Inability to proactively respond to the rhythm of the conversation: It generally relies on users manually inputting or copying content as triggering conditions, lacking the ability to automatically perceive the context and generate replies during real-time conversations; 3. Dependence on target application interfaces: It usually requires the open APIs or internal interfaces of the target social communication application to extract information such as text and images. This not only requires the target application to provide interfaces, but is also subject to the limitations of interface permissions and call frequency. Once the interface is not open or the policy changes, the function will not function properly.
[0126] The technical solution of this embodiment can overcome some of the aforementioned implementation defects. It can acquire screenshots, video streams, audio data, and other data of the target chat window through system-level functions. Then, it can perform multimodal data processing according to time sequence to obtain complete chat content carrying the semantic information of the chat context of the target chat window. This allows it to capture chat content in the chat window of any social communication application on the terminal device without requiring the social communication application to provide interface support, ensuring the universality of chat assistance functions and usability across application scenarios. Furthermore, the technical solution of this embodiment can extract at least one chat message to be replied to in real time according to the rhythm of the conversation based on the real-time acquired chat content, and automatically generate the corresponding chat reply. Thus, the technical solution of this embodiment can achieve chat content acquisition and intelligent chat reply generation at the operating system level, while maintaining high privacy, low latency, and cross-application adaptability.
[0127] The technical solution of this embodiment can collect text, images, emoticons, and other visual information presented in the chat window of the current foreground social communication application in real time through system-level functions (such as screenshot mechanisms). Combined with externally received and played or local voice input (such as real-time user voice commands or voice dialogue records) collected by system-level functions, and without relying on the target application interface or backend database, it utilizes a multimodal fusion algorithm to extract features and semantically associate multiple sources of data to achieve accurate context reconstruction and obtain chat content. Therefore, the technical solution of this embodiment can achieve cross-application chat content acquisition. To further improve generation accuracy, the technical solution of this embodiment can also combine the constructed user profile and chat scenario to form a personalized generation condition vector, driving a local or edge-secure AI model to complete the generation of chat replies. In addition, it can collect user selection data on generated chat replies and perform feedback processing to continuously optimize the user profile, context-aware model, and AI model through a data closed-loop mechanism, achieving continuous improvement of the reply strategy.
[0128] Based on the technical concept of the foregoing embodiments, the foregoing embodiments will be illustrated by specific scenario examples below: In one example, a processing method includes: 1.1 Detect whether the current context is a social chat scenario; 1.2 In a social chat scenario, display the "Help me say something" function button on the screen; 1.3 After the function button is triggered, voice data and screenshots of the chat window are collected through system-level functions; 1.4 Transcribe the speech data into text content, perform scrolling stitching and / or OCR / image recognition on the captured window screenshots to obtain the screenshot parsing content; 1.5 Perform multimodal fusion and context reconstruction (time and semantic alignment) on the text content and screenshot parsing content to obtain chat content; 1.6 Based on the chat content, the obtained user profile, and the chat context, obtain the generation condition vector (i.e., the aforementioned model input information). 1.7 Based on the generated condition vector and AI model, output multiple candidate chat replies corresponding to the target chat information in the chat content; 1.8 In response to the user's selection of candidate chat replies, display the selected candidate chat reply value input box for sending information; or, in response to the user's error correction operation on the candidate chat reply, record the error correction and editing data, and provide optimization feedback for the AI model.
[0129] In another example: When users handle after-sales matters in a product consultation application, the application scenarios of the chat assistance function built based on the technical solution of this embodiment include: When the scene recognition module detects that the user is currently in the chat window of the product consultation application, a "Help Me Say It" floating button is displayed at the top of the screen as the entry point for enabling the chat reply assistance function. Users can click this button at any time during communication with customers to receive intelligent reply suggestions; If a customer receives goods and finds the packaging damaged, becomes agitated, and uses strong language to express their frustration, and the user wants to calm the customer while trying to salvage the transaction but is unsure of the best approach, the user can click the "Help Me Say It" button in the chat window to request suggested responses. After the button is triggered, the following actions can be taken: 1. System-level screenshot capture of the chat window; 2. Combine the screenshots of the windows above and below the current conversation in the chat window to form a complete chat flow screenshot; 3. Use Optical Character Recognition (OCR) technology to extract the text parsing content from the chat log screenshots, and perform image recognition on the images contained in the chat log screenshots (such as photos of damaged packaging) to extract the image parsing content; 4. Based on the screenshot text parsing content, screenshot image parsing content, and the order of chat bubbles, the text is segmented by the speaker to obtain screenshot parsing content that distinguishes between the customer's and user's statements; 5. When users send voice messages during chat, the system-level voice acquisition module acquires audio data with user authorization, and calls the local speech-to-text (ASR) module to convert the audio data into text content and mark the emotional tendency (such as "anger" or "anxiety"). 6. The screenshot parsing content and text content are timestamped and the context is reconstructed to obtain chat content that ensures semantic consistency; 7. Based on the context analysis module, the chat scenario is identified as "handling return issues", the communication style is formal business, the other party's identity is a customer, and the other party's emotion analysis result is dissatisfaction; 8. Obtain locally stored user profiles. User profiles indicate that the user is a self-operated online store owner of handmade products, has high after-sales requirements, and prefers a concise and direct communication style in business communications. 9. Construct a multi-dimensional conditional vector based on chat content, chat context, and user profile; 10. Input the multi-dimensional conditional vector into a local or secure edge model to generate multiple candidate chat responses, for example: Candidate response 1: Prioritize comforting the customer and building trust; Candidate response 2: Guide negotiation and try to salvage the order through compensation or coupons.
[0130] 11. The system detects that the user has selected candidate reply 2, and after editing the compensation amount, it sends it to the customer.
[0131] 12. Record user selection preferences, error correction magnitude and direction to obtain feedback information, and continuously optimize candidate response style and content based on historical usage frequency and feedback information to achieve data closed-loop learning and optimization; 13. Save all learning records locally to create a personalized communication style model for each user.
[0132] In another example, one processing method includes: 3.1 The system detects that the chat content corresponding to the chat window contains fruit images, and calls the image semantic analysis module to identify the object category (e.g., apple, watermelon) and extract the shape features (e.g., size) of the fruit to obtain the image parsing content; 3.2 The system-level voice acquisition module is called to obtain the user's real-time voice input "I want the big one", and automatic speech recognition (ASR) is performed locally to convert the voice into text content. The text content is appended with a timestamp and emotion tag (such as neutral or relaxed tone). 3.3 By using a multimodal fusion engine, the text content and image parsing content are aligned in the temporal and semantic domains to form a complete conversation semantic graph (i.e., chat content). 3.4 When it is detected that the "big one" in the complete conversation semantic graph has a non-unique reference (it could be an apple or a watermelon), and when it is found through context matching of the complete conversation semantic graph that both have the feature of "big", candidate clarification questions are generated based on the context, such as: "Do you mean the big apple or the big watermelon?" 3.5 Perform a confirmation operation on the clarification query to determine that "the big one" refers to "watermelon", and then correct the complete conversation semantic graph; 3.6 Based on the corrected conversation semantic graph and AI model, output and display multiple candidate chat responses corresponding to the reply "I want the big one"; 3.7 In response to the selection operation of the displayed candidate chat replies, inject the selected candidate chat reply into the input box, and wait for the user to confirm or send it; 3.8 Record the relationship between the context corresponding to the conversation semantic graph and the selected candidate chat replies, and display the user profile and context matching rules locally to improve the accuracy of generating chat replies in subsequent scenarios.
[0133] This application also provides a terminal device, including a memory and a processor. The memory stores a processing program or instructions, which, when executed by the processor, implement the steps of the processing method in any of the above embodiments.
[0134] This application also provides a storage medium storing a processing program or instructions, which, when executed by a processor, implement the steps of the processing method in any of the above embodiments.
[0135] In the embodiments of the terminal device and storage medium provided in this application, all the technical features of any of the above-described processing method embodiments may be included. The extended and explanatory content of the specification is basically the same as that of the embodiments of the above methods, and will not be repeated here.
[0136] This application also provides a computer program product, which includes computer program code. When the computer program code is run on a computer, it causes the computer to perform the methods described in the various possible implementations above.
[0137] This application also provides a chip, including a memory and a processor. The memory is used to store a computer program, and the processor is used to call and run the computer program from the memory, so that a device with the chip installed performs the methods described in the various possible implementations above.
[0138] It is understood that the above scenarios are merely examples and do not constitute a limitation on the application scenarios of the technical solutions provided in the embodiments of this application. The technical solutions of this application can also be applied to other scenarios. For example, as those skilled in the art will know, with the evolution of system architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.
[0139] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0140] The steps in the method of this application embodiment can be adjusted, combined, or deleted according to actual needs.
[0141] The units in the device of this application embodiment can be merged, divided, and deleted according to actual needs.
[0142] In this application, the same or similar terms, concepts, technical solutions and / or application scenario descriptions are generally described in detail only when they appear for the first time. When they appear again, they are generally not repeated for the sake of brevity. When understanding the technical solutions and other contents of this application, the same or similar terms, concepts, technical solutions and / or application scenario descriptions that are not described in detail later can be referred to their previous relevant detailed descriptions.
[0143] In this application, the descriptions of the various embodiments have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0144] The technical features of the present application can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of the present application.
[0145] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as described above, and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, controlled terminal, or network device, etc.) to execute the methods of each embodiment of this application.
[0146] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a storage medium or transmitted from one storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, storage disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).
[0147] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A processing method characterized by, include: Obtain at least one chat message from the chat content of the target chat window, wherein the modality of all chat messages in the chat content includes at least one modality; Output at least one chat reply based on at least one chat message.
2. The processing method as described in claim 1, characterized in that, The method further includes at least one of the following: The chat reply includes a portion of the content of the at least one chat message; The chat reply is used to respond to the at least one chat message.
3. The processing method as described in claim 1 or 2, characterized in that, The step of outputting at least one chat reply based on the at least one chat message includes: Based on at least one of the chat content, user profile, and chat context, output at least one chat reply corresponding to at least one chat message.
4. The processing method as described in claim 1 or 2, characterized in that, The step of obtaining at least one chat message from the chat content of the target chat window includes: The chat content of the target chat window within a preset time period is obtained, wherein the modality of all chat information in the chat content includes at least one of text modality, image modality, voice modality, and video modality; Obtain at least one chat message from the chat content.
5. The processing method as described in claim 1 or 2, characterized in that, The method further includes: Based on at least one chat message, multiple chat replies are output in the candidate area of the target chat window; In response to a selection operation on the multiple chat replies, the input box of the target chat window displays at least one selected chat reply from the multiple chat replies, and the position of the input box is different from that of the candidate area; or, in response to a selection operation on the multiple chat replies, at least one selected chat reply from the multiple chat replies is sent.
6. The processing method as described in claim 1 or 2, characterized in that, The step of outputting at least one chat reply based on the at least one chat message includes: For any semantically ambiguous content in the at least one chat message, output a corresponding clarification prompt; In response to a confirmation action for the clarification prompt, at least one chat reply with explicit semantic content is output.
7. The processing method as described in claim 1 or 2, characterized in that, The step of outputting at least one chat reply based on the at least one chat message includes: Obtain model input information corresponding to the at least one chat message; Based on the input information of the model and the AI model, output at least one chat reply.
8. The processing method as described in claim 6, characterized in that, The method further includes at least one of the following: For semantically ambiguous content detected from at least one chat message, output the corresponding clarification prompt; In response to the clarification confirmation operation for the clarification prompt, the corresponding chat information is corrected; In response to the clarification confirmation operation for the clarification prompt, the model input information is updated.
9. A terminal device, characterized in that, include: A memory and a processor, the memory storing a processing program or instructions which, when executed by the processor, implement the processing method as described in any one of claims 1 to 8.
10. A storage medium, characterized in that, The storage medium stores a computer program or instructions, which, when executed by a terminal device, implement the processing method as described in any one of claims 1 to 8.