Chat information sending method, electronic equipment, readable storage medium and program product
By launching an auxiliary chat service in the game, collecting audio data for semantic recognition and content expansion, and automatically sending chat messages, the inconvenience and inefficiency of sending game chat messages are solved, thus improving the gaming experience.
Patent Information
- Application Number
- CN202410690655.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-30
- Publication Date
- 2025-12-02
AI Technical Summary
The inconvenience and inefficiency of sending chat messages for gamers negatively impacts their gaming experience.
When the auxiliary chat service function is detected to be activated, audio data is collected for semantic recognition, language environment information is identified and content is expanded, the target player's name is automatically identified and extended semantic text is sent.
It improves the convenience and efficiency of chat message interaction, enhances the gaming experience, reduces operation steps, and improves the real-time nature of communication.
Smart Images

Figure CN121041702A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technology, and in particular to a method for sending chat messages, an electronic device, a readable storage medium, and a program product. Background Technology
[0002] With the widespread adoption of mobile devices, more and more users are using smartphones and other smart devices for gaming. This has led to a diversified gaming market, with a plethora of game products of various types. Against this backdrop, users' demands for game quality and experience are constantly increasing; they are paying more attention to details and smooth controls. Therefore, improving the user experience from the smallest details and optimizing game controls to provide a better user experience has become particularly important and necessary. Only in this way can user needs be truly met, user satisfaction enhanced, and a company able to stand out in a highly competitive market.
[0003] In related technologies, the main way for game players to interact during gameplay is by sending text or voice messages. This requires finding the interaction target (teammate's name) or the chat room entrance on the currently running game interface, and then going through multiple steps such as clicking the game voice dialog box, typing on the keyboard, or clicking the voice button to successfully send the message. These steps are cumbersome and the messages are not sent in a timely manner, which greatly affects the user's experience during gameplay.
[0004] In other words, the relevant technologies have the following shortcomings in chat message sending scenarios:
[0005] 1. Inconvenient operation: When you only want to send a message to a specific player, you need to find the corresponding player and switch to the chat room interface. The keypad has few keys, and for some complex words or phrases, players may need to click multiple times to input them.
[0006] 2. Low efficiency of chat message interaction: During gameplay, players need to first click to find the dialogue box of the person they want to interact with, and then input text or voice using the numeric keypad. This method is not only inefficient, but also easily distracts players and affects the gaming experience. Summary of the Invention
[0007] The main purpose of this application is to provide a method for sending chat messages, an electronic device, a readable storage medium, and a program product, aiming to solve the technical problem of how to improve the convenience and efficiency of players' chat message interaction.
[0008] To achieve the above objectives, this application provides a method for sending chat messages, including:
[0009] When the activation of the auxiliary chat service function is detected, audio data is collected, semantic recognition is performed on the audio data, and semantic text corresponding to the audio data is obtained. The audio data includes voice information of the target player's name corresponding to the player to be chatted with.
[0010] The semantic text is identified by its linguistic environment information, and the semantic text is expanded based on the identified linguistic environment information to obtain expanded semantic text. The linguistic environment information includes linguistic intent and / or vocal emotion, and the content expansion includes at least one of the following: conversion of linguistic description, expansion of linguistic expression, and addition of emoticons.
[0011] Based on the current game screen, the target player's name is identified, and the extended semantic text is automatically sent to the instant chat window corresponding to the player to be chatted with.
[0012] In addition, to achieve the above objectives, this application also provides an electronic device, which includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the chat message sending method described above.
[0013] In addition, to achieve the above objectives, this application also provides a readable storage medium, which is a computer-readable storage medium, on which a program implementing a chat message sending method is stored, and the program implementing the chat message sending method is executed by a processor to implement the steps of the chat message sending method as described above.
[0014] In addition, to achieve the above objectives, this application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the chat message sending method described above.
[0015] This application provides a method for sending chat messages. In this embodiment, when the activation of the auxiliary chat service function is detected, audio data is collected, and semantic recognition is performed on the audio data to obtain the semantic text corresponding to the audio data. The audio data includes the voice information of the target player's name corresponding to the player to be chatted with. Then, the linguistic environment information of the semantic text is identified, and the semantic text is expanded according to the identified linguistic environment information to obtain expanded semantic text. The linguistic environment information includes language intent and / or voice emotion. The forms of content expansion include at least one of the following: conversion of the description method of the language text, expansion of the expression of the language text, and addition of emoticons. Based on the current game screen, the target player's name is identified, and the extended semantic text is automatically sent to the instant chat window corresponding to the player to be chatted with. This allows the application to enable the auxiliary chat service function based on different game content and game status, either through screen recognition or user-initiated activation. Once activated, the user's speech is automatically recorded, and the user's voice is converted into text in real time. During text recognition, intent recognition is performed on the user's speech. If specific game-related content is identified, game strategies and predictions can be automatically provided. For example, if the user is joking, their words can be converted into humorous statements or poems. If the user wants to express a certain emotion, emoticons can be automatically provided. Based on the extended semantic text obtained in the previous step, the auxiliary function service automatically matches and sends it to the teammate who needs to chat. This embodiment of the application automatically sends game interaction messages quickly and efficiently, enhancing the user's gaming experience and effectively improving the convenience and efficiency of players' chat information interaction.
[0016] It is worth mentioning that this application provides an AI (Artificial Intelligence) assisted chat method. By using AI to recognize and understand the intentions of game users, it not only provides users with game strategies and communication content, but also intelligently identifies the chat location on the game interface. Through system auxiliary functions, it automatically triggers the sending of chat content, greatly assisting users in both chat content and operation during the game, and enhancing the user's gaming experience. Attached Figure Description
[0017] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.
[0019] Figure 1 This is a flowchart illustrating an embodiment of the chat message sending method of this application.
[0020] Figure 2 This is a flowchart illustrating Embodiment 2 of the chat message sending method of this application;
[0021] Figure 3 This is a flowchart illustrating Embodiment 3 of the chat message sending method of this application;
[0022] Figure 4 This is a diagram of an AI recognition framework from a specific embodiment of this application.
[0023] Figure 5 This is a conceptual model diagram of an AI-powered automatic chat human-computer interaction interface according to a specific embodiment of this application.
[0024] Figure 6 This is a schematic diagram of the AI recognition and automatic sending process in a specific embodiment of this application.
[0025] Figure 7 This is a schematic diagram of a game screen text recognition process in a specific embodiment of this application.
[0026] Figure 8 This is a schematic diagram of a user recording text recognition process in a specific embodiment of this application.
[0027] Figure 9 This is a schematic diagram of the device structure of the hardware operating environment involved in the chat message sending method in this application embodiment.
[0028] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0029] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0030] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0031] In related technologies, the main way for game players to interact during gameplay is by sending text or voice messages. This requires finding the interaction target (teammate's name) or the chat room entrance on the currently running game interface, and then going through multiple steps such as clicking the game voice dialog box, typing on the keyboard, or clicking the voice button to successfully send the message. These steps are cumbersome and the messages are not sent in a timely manner, which greatly affects the user's experience during gameplay.
[0032] In other words, the relevant technologies have the following shortcomings in chat message sending scenarios:
[0033] 1. Inconvenient operation: When you only want to send a message to a specific player, you need to find the corresponding player and switch to the chat room interface. The keypad has few keys, and for some complex words or phrases, players may need to click multiple times to input them.
[0034] 2. Low efficiency of chat message interaction: During gameplay, players need to first click to find the dialogue box of the person they want to interact with, and then input text or voice using the numeric keypad. This method is not only inefficient, but also easily distracts players and affects the gaming experience.
[0035] The main solution of this application is as follows: upon detecting the activation of the auxiliary chat service function, audio data is collected, semantic recognition is performed on the audio data to obtain semantic text corresponding to the audio data, wherein the audio data includes voice information of the target player's name corresponding to the player to be chatted; the linguistic environment information of the semantic text is identified, and the semantic text is expanded according to the identified linguistic environment information to obtain expanded semantic text, wherein the linguistic environment information includes linguistic intent and / or voice emotion, and the form of content expansion includes at least one of the following: conversion of linguistic description, expansion of linguistic expression, and addition of emoticons; based on the current game screen, the target player's name is identified, and the expanded semantic text is automatically sent to the instant chat window corresponding to the player to be chatted.
[0036] This application first collects audio data upon detecting the activation of the assisted chat service function, performs semantic recognition on the audio data to obtain the semantic text corresponding to the audio data, wherein the audio data includes voice information of the target player's name corresponding to the player to be chatted with, then identifies the linguistic environment information of the semantic text, and expands the semantic text according to the identified linguistic environment information to obtain expanded semantic text, wherein the linguistic environment information includes language intent and / or voice emotion, and the form of content expansion includes at least one of the following: transformation of the description method of the language text, expansion of the expression of the language text, and addition of emoticons. Based on the current game screen, the system identifies the target player's name and automatically sends the extended semantic text to the corresponding instant chat window of the player to be chatted with. This allows the application to enable auxiliary chat services based on different game content and states, either through screen recognition or user-initiated activation. Once activated, the system automatically records the user's speech and converts it into text in real time. During text recognition, the system performs intent recognition on the user's speech. If specific game-related content is identified, game strategies and predictions can be automatically provided. For example, if the system recognizes the user as joking, the speech can be converted into humorous statements or poems. If the user wants to express a certain emotion, emoticons can be automatically provided. The extended semantic text obtained from the previous step is automatically matched and sent to the teammate who needs to chat through the auxiliary service. This application automatically sends game interaction messages quickly and efficiently, enhancing the user's gaming experience and effectively improving the convenience and efficiency of players' chat information interaction.
[0037] It should be noted that the executing subject of this application is a terminal device, which may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Tablets), PMPs (Portable Media Players), and in-vehicle terminals (such as in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers, or any electronic device capable of performing the above functions. This application does not specifically limit the scope of the application. The following description uses a terminal device as the executing subject to illustrate the various embodiments of this application.
[0038] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0039] Please refer to Figure 1 , Figure 1This is a flowchart illustrating the first embodiment of the chat message sending method of this application.
[0040] This application proposes a chat message sending method according to a first embodiment. The system of the terminal includes a first logical screen for running a first application. The chat message sending method includes steps S100 to S300:
[0041] Step S100: When the activation of the auxiliary chat service function is detected, audio data is collected, semantic recognition is performed on the audio data, and semantic text corresponding to the audio data is obtained. The audio data includes voice information of the target player's name corresponding to the player to be chatted with.
[0042] It should be noted that, in this embodiment, the auxiliary chat service function is a service function that can help players communicate more conveniently and efficiently in game applications. This auxiliary chat service function integrates advanced speech recognition, natural language processing and intelligent analysis technologies, providing players with a more intelligent and personalized chat assistance method, aiming to enhance the social interaction experience in the game, reduce communication barriers, and improve the interaction efficiency and fun between players.
[0043] It is worth mentioning that the auxiliary chat service function in this embodiment has two activation mechanisms. First, it is automatically triggered by the terminal device's system based on game screen recognition within the game application: the system can intelligently monitor specific situations or player behavior during the game process and automatically activate the auxiliary chat service function. For example, in team cooperation, trading areas, or PVP (Player versus Player) scenarios, the system automatically intervenes when it senses that players may have a communication need. Second, it can be actively triggered by the user through virtual buttons or voice commands, allowing the user to manually activate the auxiliary chat service when needed, giving the user flexible and autonomous control, and increasing the flexibility and convenience of using the auxiliary chat service.
[0044] In a game scenario, once the auxiliary chat service function is detected to be activated, the terminal device automatically turns on the microphone and collects the audio data input by the user with the user's permission. At the same time, it performs semantic recognition on the collected audio data to obtain the semantic text corresponding to the audio data.
[0045] It should be noted that the audio data contains voice information of the target player's name corresponding to the player to be chatted with. The player to be chatted with refers to the game player that the user wants to initiate a chat with or is trying to communicate through the auxiliary chat service. The target player's name refers to the game name of the player to be chatted with in the game application or the nickname that the user has set for the player to be chatted with.
[0046] It's easy to understand that the audio data contains voice information of the target player's name corresponding to the player to be chatted with. This is to ensure that the auxiliary chat service function can correctly identify the person the user wants to chat with, i.e., the player to be chatted with, and better understand the context and intent of the conversation. This ensures that the information is correctly conveyed to the designated player, and can achieve precise interaction even in the complex social context of multiplayer online games, enhancing the targeting and efficiency of communication.
[0047] Furthermore, in this embodiment, the step of semantic recognition of audio data to obtain the corresponding semantic text is achieved by converting the audio data into text form using speech recognition technology. This step involves analyzing the sound waveform, extracting features, and using machine learning models (such as deep neural networks) to match the sound with the most likely text sequence, thereby converting and parsing the original audio data into text information with clear semantics, i.e., semantic text, providing a foundation for subsequent operations such as content expansion and message sending.
[0048] In this embodiment, when the auxiliary chat service function is activated, audio data is collected, and semantic recognition is performed on the collected audio data to obtain the semantic text corresponding to the audio data. This allows users to quickly convey their intentions through voice without manually inputting information during intense gameplay, reducing operation steps, enhancing the continuity and immersive experience of the game, and making in-game communication faster and more timely, thus improving in-game communication efficiency.
[0049] For example, in one feasible implementation, the step of performing semantic recognition on the audio data to obtain the semantic text corresponding to the audio data may include steps S110 to S130:
[0050] Step S110: Perform a second preprocessing on the audio data to obtain preprocessed audio data.
[0051] It should be noted that in this embodiment, the second preprocessing includes at least one of denoising, sampling rate adjustment, and quantization. Denoising aims to eliminate background noise in the audio data, such as fan noise, traffic noise, or other non-speech signals, ensuring the clarity of the main speech content. This is generally achieved through digital signal processing techniques, such as noise suppression algorithms or spectral subtraction, to separate a clean speech signal. Sampling rate adjustment is necessary because different semantic recognition systems require specific sampling rates for the input audio. When the sampling rate of the original audio data does not meet the requirements, it needs to be adjusted to a suitable value using resampling techniques, such as the common 44.1kHz, 48kHz, or 16kHz, to match the optimal performance standards of the semantic recognition engine. Quantization, by converting continuous analog signals into discrete digital signals, reduces the storage space required for audio data processing and transmission, thus improving the processing efficiency. Those skilled in the art will understand that appropriate quantization can effectively control the amount of data without significantly sacrificing speech quality.
[0052] This embodiment can effectively improve the quality and format of the audio signal by performing a second preprocessing on the raw audio data to be acquired, making it more suitable for processing and analysis, thereby improving the robustness and efficiency of the entire semantic recognition process.
[0053] Step S120: Extract audio features from the second preprocessed audio data using Mel frequency cepstral coefficients to obtain target audio features;
[0054] As those skilled in the art will know, Mel Frequency Cepstral Coefficients (MFCCs) are a feature extraction technique widely used in speech recognition and audio processing. MFCCs are inspired by the human ear's perception of different frequencies of sound, and they can effectively capture the spectral features of human speech that are closely related to the content of the speech.
[0055] It is easy to understand that the target audio features in this embodiment are extracted from the second preprocessed audio data using Mel-frequency cepstral coefficients. This embodiment extracts the target audio features from the second preprocessed audio data using Mel-frequency cepstral coefficients, essentially converting the preprocessed audio signal (i.e., the second preprocessed audio data) into a set of numerical feature vectors (i.e., the target audio features). These feature vectors can effectively characterize the dynamic spectral changes in the speech signal, highly summarizing the spectral information in the original audio data, especially information related to speech content and speaker characteristics. They can filter out irrelevant noise and redundant information, retaining only the parts closely related to speech content and semantics, thereby improving the accuracy and reliability of speech recognition. Simultaneously, the process of extracting audio features converts the complex audio signal into a format that machine learning models can understand and process, making the data results more concise and consistent, thus accelerating the efficiency and accuracy of semantic recognition from machine learning models.
[0056] Step S130: Input the target audio features into a pre-trained speech recognition deep neural network model to identify the semantic text corresponding to the audio data.
[0057] As those skilled in the art will recognize, a neural network is a computational model inspired by the structure and function of biological nervous systems, designed to make predictions or decisions by learning the inherent patterns in data. It consists of a large number of artificial neurons (or nodes), interconnected by connection weights to form a complex network structure. Neurons receive input signals, transform these signals through weighted summation and nonlinear activation functions, and pass the results to neurons in the next layer. This process is repeated until a final output is produced. A typical neural network includes at least one input layer (receiving raw data), one output layer (providing the final result), and possibly one or more hidden layers (processing intermediate information). The core capability of neural networks lies in their ability to automatically learn feature representations from training data, thereby solving various complex problems, including classification, regression, clustering, and generation.
[0058] A neural network model refers to a specifically defined neural network structure and its operating mechanism, including but not limited to the number of layers (input, hidden, and output layers), the number of nodes in each layer, the connection methods between nodes (fully connected, convolutional connections, etc.), the activation functions used (such as Sigmoid, ReLU, etc.), the loss function (such as cross-entropy loss), and the optimization algorithm (such as gradient descent). The model also includes the parameters that need to be learned during training, namely the connection weights and biases. A neural network model is a concrete implementation of neural network theory; by configuring different parameters and structures, it can be optimized and applied to different types of tasks.
[0059] In short, a neural network is a concept, while a neural network model is an instantiation of this concept in a practical application, with a defined structure and configuration, capable of performing specific computational tasks.
[0060] Accordingly, Deep Neural Networks (DNNs) are a special type of neural network, characterized by having multiple hidden layers, typically at least two. These hidden layers enable the network to learn more complex and abstract feature representations of the data. Deep neural networks emerged because, in certain tasks, shallow networks (networks with only one or a few hidden layers) cannot effectively capture enough information to solve the problem. As the number of layers increases, deep neural networks can learn different levels of features in the data hierarchically, from low-level edge detection to high-level object recognition. This capability has led to significant success for DNNs in fields such as image recognition, speech recognition, and natural language processing. A deep neural network model is an instance of this concept in practical applications.
[0061] It should be noted that the speech recognition deep neural network model in this embodiment is a deep neural network model specifically designed for speech recognition. It has the function of converting the target audio features extracted through Mel frequency cepstral coefficients into text output. In other words, after the target audio features extracted through Mel frequency cepstral coefficients are input into the speech recognition deep neural network model, the speech recognition deep neural network model can output the semantic text corresponding to the original audio data.
[0062] It is easy to understand that this speech recognition deep neural network model is pre-trained. The training process of this speech recognition deep neural network model has been studied to some extent by those skilled in the art, and this embodiment will not elaborate on it further.
[0063] This application embodiment inputs the target audio features into a pre-trained speech recognition deep neural network model to identify the semantic text corresponding to the audio data. By utilizing an advanced deep neural network model, it performs high-precision and efficient speech recognition, automatically converting the user's voice messages into text. This achieves efficient speech-to-text conversion, allowing players to communicate in-game using natural language without manual input. It enables players to maintain their game state while communicating effectively, providing strong real-time performance and greatly improving the smoothness and convenience of game interaction, thus enhancing the user experience.
[0064] Step S200: Identify the linguistic environment information of the semantic text, and expand the content of the semantic text according to the identified linguistic environment information to obtain expanded semantic text. The linguistic environment information includes linguistic intent and / or vocal emotion. The content expansion includes at least one of the following: conversion of the descriptive method of the language text, expansion of the expression of the language text, and addition of emoticons.
[0065] It should be noted that, in this embodiment, linguistic context information refers to the comprehensive contextual features related to the semantic text to be analyzed, which mainly includes two core aspects:
[0066] Linguistic intention refers to the purpose or action that the speaker intends to achieve through speech. For example, is the speaker asking for information, issuing instructions, expressing emotions (such as happiness or dissatisfaction), stating facts, or engaging in a debate? Understanding linguistic intention helps the system accurately determine how to appropriately respond to or further process this information.
[0067] Voice emotion refers to the emotional tone or state of mind conveyed in speech. This depends not only on the content of the speech but also on nonverbal cues such as tone, speed, and intensity. By analyzing voice signals, the system can identify possible emotional states of the speaker, such as happiness, sadness, anger, and surprise. This is crucial for adding appropriate emotional expressions (such as selecting corresponding emoticons) to chat content.
[0068] Furthermore, it should be noted that in this embodiment, content expansion refers to adapting or enriching the original semantic text based on the identified linguistic environment information. This includes, but is not limited to, rewriting the expression (i.e., transforming the way the language is described), adding descriptive details (i.e., expanding the expression of the language), and adding emotional elements (such as emoticons, i.e., adding emoticons) to better adapt to the communication context and convey the intention.
[0069] According to the analyzed language intention and speech emotion, this embodiment determines which content expansion form to adopt, such as whether a more formal expression is needed, adding adjectives to enrich the description or配合适当的表情符号来增强信息的情感色彩, for example, in a scenario where multiple players are online gaming and attacking dungeon monsters, the original semantic text is "救救我", when it is recognized that the speech intention corresponding to this semantic text is urgent, the system can automatically expand the original semantic text "救救我" to "救救我!!!" and an emoji representing "urgent".
[0070] It is worth mentioning that when identifying the language environment information and performing content expansion according to the language environment information, considering that the cultural backgrounds of different regions are different, the language environment information that can be analyzed and recognized for the same semantic text is different in different cultural backgrounds, and when performing corresponding content expansion, the form and content to be expanded are also different. For example, in the Western cultural background, multiple consecutive exclamation marks can be used to expand "urgent" to "急!!!" to express a tense and urgent mood, while in the East Asian cultural background, in addition to direct text emphasis, it is generally more inclined to use more implicit but equally expressive words or short sentences for emergency situations. For example, in the Chinese cultural background, the description method of the language text of "急" can be converted and expanded into the idiom "十万火急" unique to the Chinese cultural background.
[0071] This embodiment can more accurately convey the true intentions and emotions of players, reduce misunderstandings, and accelerate information circulation by automatically identifying the language environment information and adaptively expanding the content according to the language environment information. In addition, users do not need to spend extra energy to describe in detail or find appropriate emojis (especially during intense game operations), and the system can automatically complete these tasks, making the chat content more vivid and interesting while allowing players to focus more on the game itself. This not only enhances the interactivity between players and the fun of the game, but also helps to build a more friendly and active game community and strengthen the connection and sense of belonging among players.
[0072] Exemplarily, in a feasible implementation manner, the steps of identifying the language environment information of the semantic text and performing content expansion on the semantic text according to the identified language environment information may include:
[0073] Step S210, using a generative artificial intelligence AIGC tool to identify the language environment information of the semantic text, and using the AIGC tool to perform content expansion on the semantic text according to the identified language environment information.
[0074] As those skilled in the art will recognize, an Artificial Intelligence Generated Content (AIGC) tool is a tool that uses artificial intelligence technology to create content in various forms such as text, images, audio, and video. It can automatically generate creative content in various forms such as text, images, and audio based on given conditions or inputs.
[0075] In this embodiment, after the semantic text obtained from semantic recognition in step S100 is input into the AIGC tool, the AIGC tool first uses natural speech processing technology to parse the input semantic text, extracting keywords, phrase results, sentiment tendencies, etc., and then performs contextual understanding based on the extracted information to analyze the intent behind the semantic text. At the same time, it judges the speaker's emotional state through word selection, sentence structure, etc., thereby realizing intent recognition and emotion detection, determining the speaker's linguistic intent and vocal emotion, that is, recognizing the linguistic environment information of the semantic text, and then adding relevant details or supplementary information according to the recognized linguistic environment information to make the content more complete and rich. At the same time, it adjusts the tone and format of the output text to adapt to the recognized context (such as formal, humorous, comforting). In addition, this embodiment can also generate more personalized extended semantic text according to user habits or historical communication records.
[0076] Step S300: Based on the current game screen, identify the name of the target player and automatically send the extended semantic text to the instant chat window corresponding to the player to be chatted with.
[0077] It's easy to understand that the current game screen refers to the real-time game status view displayed on the user's terminal device screen while playing the game. The real-time chat window refers to the area where chat history is displayed when players chat with each other in the game.
[0078] In this embodiment, the process of identifying the target player's name based on the current game screen can be achieved by taking a screenshot of the current game screen, then performing text recognition and text extraction on the current game screen to obtain all text information under the current game screen, and then combining the semantic text corresponding to the audio data to identify the target player's name, determine the player to whom the extended semantic text needs to be sent, and then automatically send the extended semantic text to the instant chat window corresponding to the player to be chatted.
[0079] It's easy to understand that since the audio data contains the voice information of the target player's name corresponding to the player being chatted with, the semantic text obtained by semantic recognition of the audio data contains the target player's name. In this embodiment, the target player's name can be identified by performing word segmentation on two pieces of text (the semantic text corresponding to the audio data and all text information under the current game screen), converting the text into individual words or characters, comparing the word segmentation results of the two pieces of text, finding the same fields, and thus identifying the target player's name.
[0080] It is worth mentioning that before performing text recognition and extraction on the current game screen, an object detection algorithm can be used to identify areas on the current game screen that may contain the target player's name, thereby performing text recognition and extraction on those areas and effectively improving the recognition efficiency of the target player's name.
[0081] This embodiment first collects audio data when the auxiliary chat service function is detected to be activated, performs semantic recognition on the audio data to obtain the semantic text corresponding to the audio data. The audio data includes the voice information of the target player's name corresponding to the player to be chatted with. Then, the linguistic environment information of the semantic text is identified. Based on the identified linguistic environment information, the semantic text is expanded to obtain expanded semantic text. The linguistic environment information includes language intent and / or voice emotion. The forms of content expansion include at least one of the following: conversion of the description method of the language text, expansion of the expression of the language text, and addition of emoticons. Based on the current game screen, the target player's name is identified, and the extended semantic text is automatically sent to the instant chat window corresponding to the player to be chatted with. This embodiment allows for the activation of the auxiliary chat service function based on different game content and states, either through screen recognition or user-initiated activation. Once activated, the user's speech is automatically recorded, and then converted into text in real time. During text recognition, intent is identified in the user's speech. If specific game-related content is identified, game strategies and predictions can be automatically provided. For example, if the user is joking, their words can be converted into humorous statements or poems; if the user wants to express a certain emotion, emoticons can be automatically provided. The extended semantic text obtained from the previous step is automatically matched and sent to the teammate who needs to chat through the auxiliary function service. This embodiment, by automatically sending game interaction messages, is fast and efficient, enhancing the user's gaming experience and effectively improving the convenience and efficiency of players' chat information interaction.
[0082] Based on the first embodiment described above, a chat message sending method according to a second embodiment of this application is proposed.
[0083] In the second embodiment of this application, the same or similar content as in the above embodiments can be referred to the above description, and will not be repeated hereafter.
[0084] Please refer to Figure 2 , Figure 2 This is a flowchart illustrating the second embodiment of the chat message sending method of this application.
[0085] In this embodiment, the step of identifying the target player's name based on the current game screen may include steps S310 to S330:
[0086] Step S310: Perform a first preprocessing on the current game screen to obtain the current game screen after the first preprocessing;
[0087] It should be noted that, in this embodiment, the first preprocessing includes image reading, image resizing, and image grayscale conversion. Image reading refers to acquiring the current game screen through methods such as screenshotting; image resizing refers to adjusting the image resolution using algorithms such as interpolation, scaling the image to a more suitable size for processing while maintaining a certain aspect ratio, to meet the requirements of subsequent image processing algorithms and reduce computational resource consumption; image grayscale conversion refers to converting each pixel of the current game screen into a single grayscale value, simplifying image information, removing color information, and retaining only grayscale values, thereby reducing computational complexity and accelerating subsequent processing.
[0088] This embodiment performs a first preprocessing on the current game screen to reduce the amount of data while optimizing image quality and structure, resulting in a preprocessed current game screen. This improves the efficiency of subsequent processing steps and the accuracy of the final target player name recognition.
[0089] Step S320: Extract image features from the first preprocessed current game screen to obtain target image features;
[0090] As those skilled in the art will know, image features refer to the quantitative representations extracted from an image that describe its basic information such as content, structure, texture, color, and shape. These features are the foundation of computer vision tasks such as image analysis, recognition, and classification. Image feature extraction involves applying appropriate algorithms or models to extract the desired features from an image.
[0091] It is easy to understand that in this embodiment, the target image features are the image features extracted from the current game screen after the first preprocessing, which are required to identify the text information in the current game screen.
[0092] It is worth mentioning that, in order to more accurately identify the text information in the current game screen, the target image features to be extracted in this embodiment may include edge features, texture features, local invariant features, directional features, connected component features, and structural features.
[0093] Step S330: Based on the target image features, identify the text information in the current game screen, and based on the text information in the current game screen, identify the target player's name.
[0094] In this embodiment, there are multiple ways to identify text information in the current game screen based on target image features.
[0095] For example, in a first feasible implementation, the step of identifying text information in the current game screen based on the target image features may include step S331:
[0096] Step S331: Based on a preset temporal classification algorithm, the target image features are decoded to output the text information in the current game screen.
[0097] As those skilled in the art will recognize, time series classification algorithms are a class of machine learning algorithms specifically designed to process time series data, with the aim of identifying and predicting category labels in time series. Time series data refers to a series of data points arranged in chronological order, which typically reflect a phenomenon or measurement that changes over time, such as stock prices, weather data, and electrocardiogram signals.
[0098] In this embodiment, since the game screen may be constantly changing during the process from user voice input to voice-to-text recognition and then to game screen text recognition, causing changes in the text information on the game screen, this embodiment starts by capturing the current game screen from the moment the user inputs voice and the terminal device collects audio data through the microphone. Then, it uses a temporal classification algorithm to decode the target image features of the captured game screen during this period, thereby accurately obtaining the text information in the current game screen.
[0099] For example, in a first feasible implementation, the step of identifying text information in the current game screen based on the target image features may include step S332:
[0100] Step S332: Input the target image features into a pre-trained image recognition convolutional neural network model to recognize the text information in the current game screen.
[0101] As those skilled in the art will recognize, Convolutional Neural Networks (CNNs) are an important branch of deep neural networks. Their unique feature lies in the introduction of convolutional layers, specifically designed to process data with a grid structure, such as images and videos. Compared to traditional deep neural networks, CNNs use a set of learnable filters that slide across the input data through convolutional layers to detect local features. This mechanism not only significantly reduces the number of parameters but also preserves spatial information, enabling the network to remain robust to transformations such as translation, scaling, and rotation. Furthermore, the use of pooling layers further reduces the dimensionality of the data, extracts key features, and increases the model's generalization ability. A convolutional neural network model is an instance of this concept in practical applications.
[0102] In this embodiment, the image recognition convolutional neural network model is a convolutional neural network model specifically designed for recognizing text information in images. It has the ability to convert features extracted from images into text information output. That is, after the target image features are input into the image recognition convolutional neural network model, the model, through a series of carefully designed convolutional layers, pooling layers, and other advanced structures, not only identifies visual patterns in the image, but also accurately parses the text content composed of these patterns, thereby outputting text information that perfectly matches the text displayed on the current game screen (i.e., the text information in the current game screen). This achieves an effective conversion from complex image data to intuitive text descriptions, providing strong technical support for understanding and interacting with game content.
[0103] Based on the first embodiment described above, a chat message sending method according to a third embodiment of this application is proposed.
[0104] In the third embodiment of this application, the same or similar content as the above embodiments can be referred to the above description, and will not be repeated hereafter.
[0105] Please refer to Figure 3 , Figure 3 This is a flowchart illustrating the third embodiment of the chat message sending method of this application.
[0106] In this embodiment, the step of automatically sending the extended semantic text to the instant chat window corresponding to the player to be chatted may include steps S340 to S350:
[0107] Step S340: Based on the target player's name, locate and enter the chat interface corresponding to the player to be chatted, wherein the chat interface includes a chat message input box and a chat message send button;
[0108] It should be noted that, in this embodiment, the chat interface refers to the entire visible area in which the user interacts with other players (including players to be chatted with) in the game, including but not limited to the chat history display area (i.e., the instant chat window), chat information input box, emoticon selection, attachment upload option, and chat information send button.
[0109] This embodiment locates and accesses the chat interface of the player to be chatted with based on their name. All operations are completed automatically by the terminal device. Users no longer need to manually browse player lists or search for specific players' chat entries within complex game interfaces, reducing operational steps and making in-game communication smoother and more efficient, enhancing the game's interactivity and immersion. In tense game matches or cooperative tasks, this embodiment can quickly locate the target player's chat interface, instantly transmitting key information and preventing users from delaying tactical deployments or cooperation opportunities due to searching for players, thus improving team collaboration efficiency. Furthermore, convenient chat interface access allows for more frequent and natural communication between players, helping to form closer social connections, promoting in-game community building, and increasing player engagement.
[0110] It is worth mentioning that when locating the chat interface of the player to be chatted based on the target player's name, the system can automatically search for the player to be chatted corresponding to the target player's name from the game's chat system or the user's friend list in the game, and automatically jump to the chat interface of the player to be chatted. Alternatively, it can obtain the coordinates of the target player's name in the current game screen, and then simulate touching the coordinates to generate and execute the instruction to locate and enter the chat interface of the player to be chatted.
[0111] For example, in one possible implementation, step S340 may include step S341:
[0112] Step S341: Determine the first coordinate position of the target player's name in the current game screen, simulate the first control command corresponding to the first coordinate position using an automatic simulated touch tool, and locate and enter the chat interface corresponding to the player to be chatted based on the first control command.
[0113] It should be noted that the automatic simulated touch tool in this embodiment is a software tool used to simulate the touch input behavior of human users so as to operate the touch screen interface of electronic devices without actual physical contact. The automatic simulated touch tool can automatically execute touch operations such as clicking, swiping, and long pressing on the screen according to preset instructions or scripts.
[0114] In this embodiment, the first coordinate position refers to the coordinate position of the target player's name in the current game screen, and the first control command is used to locate and enter the chat interface corresponding to the player to be chatted with. It is easy to understand that in situations requiring actual physical contact, the user needs to manually touch the first coordinate position to trigger the first control command for locating and entering the chat interface corresponding to the player to be chatted with. However, in this embodiment, the user does not need to operate manually; the automatic touch simulation tool can simulate the process of the user manually touching the first coordinate position, automatically generating and triggering the first control command for locating and entering the chat interface corresponding to the player to be chatted with.
[0115] It is worth mentioning that the automatic touch simulation tool can also generate a touch command for the first coordinate position based on the first coordinate position, simulating the process of the user manually touching the first coordinate position, so as to trigger the first control command used to locate and enter the chat interface corresponding to the player to be chatted.
[0116] After generating the first control command, the terminal device executes the first control command and locates and enters the chat interface corresponding to the player to be chatted with, so as to facilitate the input and sending of subsequent extended semantic text.
[0117] Step S350: Automatically trigger the input of the extended semantic text into the chat information input box, and automatically touch the chat information send button to automatically send the extended semantic text to the instant chat window corresponding to the player to be chatted.
[0118] In this embodiment, after locating and entering the chat interface corresponding to the player to be chatted, the terminal device automatically triggers the input of extended semantic text into the chat information input box of the chat interface, and automatically touches the chat information send button in the chat interface to send the extended semantic text entered into the chat information input box to the instant chat window corresponding to the player to be chatted. This eliminates the need for users to manually input and send information, simplifies the operation process of communication in the game, improves the efficiency and convenience of in-game communication, and avoids typos and miscommunications that may occur during manual operation.
[0119] For example, in one feasible implementation, the step of automatically touching the chat message sending key may include step S351:
[0120] Step S351: Determine the second coordinate position of the chat message sending key in the chat interface, and use an automatic touch simulation tool to simulate the second control command corresponding to the second coordinate position, and automatically touch the chat message sending key according to the second control command.
[0121] It should be noted that, in this embodiment, the second coordinate position refers to the coordinate position of the chat message send button in the chat interface, and the second control command is used to send the chat message entered in the chat message input box to the instant chat window.
[0122] It is easy to understand that when actual physical contact is required, the user needs to manually touch the second coordinate position (that is, touch the chat message send button) to trigger the second control command for sending the chat message entered in the chat message input box to the instant chat window. However, in this embodiment, the user does not need to operate manually. The automatic touch simulation tool can simulate the process of the user manually touching the second coordinate position, automatically generate and trigger the second control command for sending the chat message entered in the chat message input box to the instant chat window.
[0123] In this embodiment, after the extended semantic text is input into the chat message input box, the terminal device automatically generates and triggers a second control command to send the chat message entered in the chat message input box to the instant chat window based on the second coordinate position of the chat message send button in the chat interface. This sends the extended semantic text entered in the chat message input box to the instant chat window corresponding to the player to be chatted with, so that the user does not need to manually input and send information, thereby simplifying the operation process of communication in the game and improving the efficiency and convenience of in-game communication.
[0124] To facilitate understanding of the technical concept or principle of the chat message sending method of this application as described in the above embodiments, a specific embodiment is provided below:
[0125] like Figure 4As shown, this specific embodiment adds a game AI (Artificial Intelligence) chat service framework to the system framework of the terminal device. This game AI chat service framework includes a "RecordMonitorService" module, an "AI Recognition" module, an "AutoService" data processing module, and an "Accessibility" module. The "RecordMonitorService" module is responsible for monitoring the game application. When it detects that a user has opened a chat input box or triggered AI chat (i.e., detected the activation of the accessibility chat service), it automatically activates the microphone through the "RecordMonitorService" module to record audio (i.e., collect audio data). During the recording process, the "AI Recognition" module processes the user's input speech in real time (including semantic recognition and content expansion), outputting the text content corresponding to the speech recognition. Simultaneously, the "RecordMonitorService" module captures a screenshot of the current game screen, and the "AI Recognition" module processes the screenshot image in real time, outputting the text content corresponding to the image recognition. Then, the "AutoService" data processing module performs word segmentation and matching to determine the target player's name. The system determines the first coordinate position corresponding to the target player's name from the current game screen. Then, the "Accessibility" module automatically opens the chat interface corresponding to the player to be chatted based on the first coordinate position. Next, the "RecordMonitorService" module obtains the image of the chat interface and hands it over to the "AI Recognition" module for processing. The "AutoService Data Processing" module determines the coordinate position of the chat message input box and the coordinate position of the chat message send button in the chat interface. Finally, the "Accessibility" module automatically inputs the expanded text content (i.e., expanded semantic text) into the chat message input box of the chat interface based on the coordinate position of the chat message input box. At the same time, based on the coordinate position of the chat message send button, it automatically sends the chat message in the chat message input box to the instant chat window corresponding to the player to be chatted, which is the chat history display area.
[0126] like Figure 5As shown, after the assisted chat service function is enabled, in this specific embodiment, the AI chat recognition interface model deployed on the terminal product (i.e., a generative artificial intelligence tool integrating speech recognition and image recognition functions) automatically acquires the voice input of the game player through the microphone (i.e., collects audio data), performs speech recognition on the voice input, converts the speech into text (i.e., performs semantic recognition on the audio data to obtain the semantic text corresponding to the audio data), and performs intent recognition (i.e., recognizes the linguistic environment information of the semantic text) and content expansion (i.e., expands the semantic text according to the recognized linguistic environment information to obtain expanded semantic text). Then, it performs image recognition on the current game screen to identify the names of the game teammates in the game interface. The system first sets the target player's name and its clickable coordinates (i.e., the first coordinate position), as well as the coordinates of the chat input box. Then, it outputs the text or voice content to be sent (i.e., outputs extended semantic text), and sends the teammate's name coordinates and the coordinates of the chat box in the game interface (i.e., outputs the first coordinate position and the coordinates of the chat input box). The auxiliary function performs a simulated click operation (actually, the automatic simulated touch tool integrated in the auxiliary chat service performs the simulated click operation), automatically triggering chat input and sending (i.e., automatically triggering the input of extended semantic text into the chat input box and automatically touching the chat send button to automatically send the extended semantic text to the instant chat window corresponding to the player to be chatted).
[0127] like Figure 6As shown, due to the different positions of the game chat box and teammate names on the game screen in different games, this specific embodiment activates the AI intelligent chat service (i.e., the auxiliary chat service function) through screen recognition or user-initiated activation based on different game content and game status. After the AI intelligent chat service is activated (i.e., the auxiliary chat service function is detected), the recording service is started, automatically recording the user's speech (i.e., collecting audio data), and audio features are extracted and decoded using a speech recognition deep neural network, converting the user's speech into text in real time (i.e., performing semantic recognition on the audio data to obtain the corresponding semantic text). During the text recognition process, AIGC performs intent recognition on the user's speech (i.e., recognizing the linguistic environment information of the semantic text). If specific game-related content is recognized, game strategies and predictions can be automatically provided; if the user is joking, their words can be converted into humorous statements or poems; if the user wants to express a certain emotion, emoticons can be automatically provided (i.e., expanding the semantic text based on the recognized linguistic environment information to obtain expanded semantic text). Then, the game screen image is automatically captured, and image features are extracted using convolutional neural network image recognition to obtain the text in the game screen (i.e., the target image features are input into a pre-trained image recognition convolutional neural network model to recognize the text information in the current game screen) and its corresponding coordinates, as well as the coordinate position information of the game chat box (including the chat message input box and the chat message send button). Next, the text content obtained by AIGC expansion is compared with the text content obtained by image recognition. A word segmentation matching algorithm is used to match the two texts based on the context, thereby obtaining the same text and determining the name of the game teammate to be sent (i.e., the target player's name). Then, based on the game teammate's name and the image recognition result, the position coordinate information corresponding to the game teammate's name (i.e., the first coordinate position) is obtained. After obtaining the location coordinates of the game teammate's name and the coordinates of the game chat box, the system service simulates a click, automatically adds the text content obtained by AIGC to the chat information input box (i.e. automatically triggers the input of extended semantic text into the chat information input box), and calls the simulated click interface (i.e. automatically simulates touch tools) to automatically send the text content obtained by AIGC to the designated teammate (i.e. automatically touches the chat information send button to automatically send the extended semantic text to the instant chat window of the player to be chatted).
[0128] like Figure 7As shown, the specific process of text recognition in the game screen in this embodiment includes: first, image cropping (i.e., first preprocessing) of the game screen obtained by screenshot; then, feature extraction using the optimized nonlinear activation function; and finally, prediction using the model optimized by the early stopping method (i.e., image recognition convolutional neural network model) to obtain the text information in the current game screen.
[0129] like Figure 8 As shown, the specific process of user speech-to-text recognition in this embodiment includes: first, acquiring the user-input recording data (i.e., acquiring audio data); then, preprocessing the recording data (i.e., second preprocessing), including removal operations, frame segmentation, windowing, etc.; then, extracting audio features through Mel-frequency cepstral coefficients; and finally, using the optimized trained model (i.e., speech recognition deep neural network model) for decoding to obtain the recognized text (i.e., the semantic text corresponding to the audio data).
[0130] It should be noted that the specific embodiments listed are only for the purpose of assisting in understanding this application and do not constitute a limitation on the chat message sending method of this application. Any simple modifications based on this technical concept are all within the protection scope of this application.
[0131] In addition, please refer to Figure 9 , Figure 9 This is a schematic diagram of the device structure of the hardware operating environment involved in the chat message sending method in this application embodiment.
[0132] This application also provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the chat message sending method in the above embodiments.
[0133] The following is for reference. Figure 9 The diagram illustrates a structural schematic of an electronic device suitable for implementing embodiments of this application. The electronic devices in these embodiments may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 9 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of this application.
[0134] like Figure 9 As shown, the electronic device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the electronic device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows electronic devices to communicate wirelessly or wiredly with other devices to exchange data. While electronic devices with various systems are shown in the figures, it should be understood that implementation or possession of all the systems shown is not required. More or fewer systems may be implemented alternatively.
[0135] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0136] The electronic device provided in this application, employing the chat message sending method in the above embodiments, can solve the technical problem of how to improve the convenience and efficiency of players' chat message interaction. Compared with the prior art, the beneficial effects of the electronic device provided in this application are the same as those of the chat message sending method provided in the above embodiments, and other technical features of the electronic device are the same as those disclosed in the previous embodiment method, and will not be repeated here.
[0137] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0138] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0139] In addition, this application also provides a readable storage medium, which is a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, which are used to execute the chat message sending method in the above embodiments.
[0140] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0141] The aforementioned computer-readable storage medium may be included in an electronic device or may exist independently without being assembled into an electronic device.
[0142] The aforementioned computer-readable storage medium carries one or more programs. When the aforementioned one or more programs are executed by an electronic device, the electronic device causes the following actions: upon detecting the activation of the auxiliary chat service function, to collect audio data, perform semantic recognition on the audio data to obtain semantic text corresponding to the audio data, wherein the audio data includes voice information of the target player's name corresponding to the player to be chatted with; to recognize the linguistic context information of the semantic text, and to expand the semantic text based on the recognized linguistic context information to obtain expanded semantic text, wherein the linguistic context information includes linguistic intent and / or vocal emotion, and the form of content expansion includes at least one of the following: conversion of the descriptive method of the linguistic text, expansion of the expression of the linguistic text, and addition of emoticons; and, based on the current game screen, to recognize the target player's name and automatically send the expanded semantic text to the instant chat window corresponding to the player to be chatted with.
[0143] Computer program code for performing the operations of this application can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0144] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0145] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0146] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described chat message sending method, thereby solving the technical problem of how to improve the convenience and efficiency of players' chat message interaction. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the chat message sending method provided in the above embodiments, and will not be repeated here.
[0147] The above description is only a part of the embodiments of this application and does not limit the patent scope of this application. All equivalent structural transformations made under the technical concept of this application and using the contents of the specification and drawings of this application, or direct / indirect applications in other related technical fields, are included in the patent protection scope of this application.
Claims
1. A method for sending chat messages, characterized in that, include: When the activation of the auxiliary chat service function is detected, audio data is collected, semantic recognition is performed on the audio data, and semantic text corresponding to the audio data is obtained. The audio data includes voice information of the target player's name corresponding to the player to be chatted with. The semantic text is identified by its linguistic environment information, and the semantic text is expanded based on the identified linguistic environment information to obtain expanded semantic text. The linguistic environment information includes linguistic intent and / or vocal emotion, and the content expansion includes at least one of the following: conversion of the descriptive method of the language text, expansion of the expression of the language text, and addition of emoticons. Based on the current game screen, the target player's name is identified, and the extended semantic text is automatically sent to the instant chat window corresponding to the player to be chatted with.
2. The chat message sending method as described in claim 1, characterized in that, Based on the current game screen, the identification of the target player's name includes: The current game screen is subjected to a first preprocessing step to obtain the current game screen after the first preprocessing step; Image features are extracted from the preprocessed current game screen to obtain the target image features; Based on the target image features, text information in the current game screen is identified, and based on the text information in the current game screen, the target player's name is identified.
3. The chat message sending method as described in claim 2, characterized in that, Based on the target image features, text information in the current game screen is identified, including: Based on a preset temporal classification algorithm, the features of the target image are decoded to output the text information in the current game screen; or, The target image features are input into a pre-trained image recognition convolutional neural network model to identify the text information in the current game screen.
4. The chat message sending method as described in claim 1, characterized in that, Automatically sending the extended semantic text to the instant chat window corresponding to the player to be chatted includes: Based on the target player's name, locate and enter the chat interface corresponding to the player to be chatted with, wherein the chat interface includes a chat message input box and a chat message send button; The system automatically triggers the input of the extended semantic text into the chat message input box and automatically touches the chat message send button to automatically send the extended semantic text to the instant chat window corresponding to the player to be chatted with.
5. The chat message sending method as described in claim 4, characterized in that, Based on the target player's name, locate and enter the chat interface corresponding to the player to be chatted with, including: The first coordinate position of the target player's name in the current game screen is determined. Then, the first control command corresponding to the first coordinate position is simulated by an automatic touch simulation tool. Based on the first control command, the player is located and enters the chat interface corresponding to the player to be chatted.
6. The chat message sending method as described in claim 4, characterized in that, The steps of automatically pressing the chat message send button include: The second coordinate position of the chat message sending key in the chat interface is determined, and the second control command corresponding to the second coordinate position is simulated by an automatic touch simulation tool. The chat message sending key is automatically touched according to the second control command.
7. The chat message sending method according to any one of claims 1 to 6, characterized in that, The steps of identifying the linguistic context information of the semantic text and expanding the content of the semantic text based on the identified linguistic context information include: The generative artificial intelligence (AIGC) tool is used to identify the linguistic context information of the semantic text, and the AIGC tool is used to expand the content of the semantic text based on the identified linguistic context information.
8. The chat message sending method as described in any one of claims 1 to 6, characterized in that, Perform semantic recognition on the audio data to obtain the semantic text corresponding to the audio data, including: The audio data is subjected to a second preprocessing to obtain the second preprocessed audio data; The target audio features are obtained by extracting audio features from the second preprocessed audio data using Mel frequency cepstral coefficients. The target audio features are input into a pre-trained speech recognition deep neural network model to identify the semantic text corresponding to the audio data.
9. An electronic device, characterized in that, The electronic device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the chat message sending method as described in any one of claims 1 to 8.
10. A readable storage medium, characterized in that, The readable storage medium is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the chat message sending method as described in any one of claims 1 to 8.
11. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the chat message sending method as described in any one of claims 1 to 8.