Information pushing method, storage medium, electronic equipment and computer program product
By analyzing call audio data in real time, determining semantic information, and filtering relevant information from data sources, the problem of not being able to provide relevant information in real time during a call is solved, thus improving communication efficiency and interactivity.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZTE CORP
- Filing Date
- 2025-06-18
- Publication Date
- 2026-04-17
AI Technical Summary
During a call, users cannot obtain relevant information related to the chat topic in real time, resulting in low communication efficiency and poor interactivity.
By acquiring audio data, semantic information is determined using speech recognition and natural language processing technologies. Target information associated with the semantic information is then filtered from multiple data sources and pushed to the target device in real time.
It enables seamless access to information highly relevant to the discussion topic during a call, improving communication efficiency and interactivity while avoiding call interruptions or increased operational complexity.
Smart Images

Figure CN121887856A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communications, and more specifically, to an information push method, storage medium, electronic device, and computer program product. Background Technology
[0002] In today's digital society, instant messaging has become an indispensable part of people's daily lives and work. Basic calls, as the foundation of communication, carry a wealth of information exchange needs. With the development of technology, users are no longer satisfied with transmitting information solely through voice, but rather desire seamless access to real-time trending news and relevant data during calls to deepen the conversation, improve communication efficiency, and enrich the call experience.
[0003] However, users who want to access real-time, relevant information during a call must go through a series of cumbersome steps. For example, if both parties are discussing a technological trend and want to find the latest developments, they must interrupt the call, manually open a news app or use a search engine, search for keywords, filter suitable information, and then return to the call screen to share it with the other party. This process not only disrupts the continuity of the call and increases the complexity of user operations, but also reduces the interactivity and real-time nature of the call, affecting the quality and efficiency of communication.
[0004] There is currently no effective solution to the problem that related technologies cannot provide relevant information in real time based on the conversation topics of the two parties during a call.
[0005] Therefore, it is necessary to improve the relevant technology to overcome the aforementioned defects. Summary of the Invention
[0006] This application provides an information push method, storage medium, electronic device, and computer program product to at least solve the problem in related technologies that it is impossible to provide relevant related information in real time based on the chat topics of the two parties during a call.
[0007] According to one embodiment of this application, an information push method is provided, comprising: acquiring audio data of a target object during a communication process, and determining first semantic information of the audio data; determining target information associated with the first semantic information, and pushing the target information to a target device, wherein the target device is a device that has a binding relationship with at least one of the target objects.
[0008] According to another embodiment of this application, an information push device is provided, comprising: a determining module, configured to acquire audio data of a target object during a communication process and determine first semantic information of the audio data; and a pushing module, configured to determine target information associated with the first semantic information and push the target information to a target device, wherein the target device is a device that has a binding relationship with at least one target object.
[0009] According to yet another embodiment of this application, a computer-readable storage medium is also provided, wherein a computer program is stored therein, and the computer program is configured to perform the steps in any of the above method embodiments when it is run.
[0010] According to yet another embodiment of this application, an electronic device is also provided, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.
[0011] According to yet another embodiment of this application, a computer program product is also provided, including a computer program that, when executed by a processor, implements the steps in any of the above method embodiments.
[0012] Through the above embodiments of this application, audio data of the target object during communication is obtained, and first semantic information of the audio data is determined; target information associated with the first semantic information is determined, and the target information is pushed to the target device, wherein the target device is a device that has a binding relationship with at least one of the target objects. That is, the embodiments of this application obtain and analyze the audio data of both parties in real time, determine their semantic information, and intelligently push target information that is highly related to the discussion topic on this basis, thereby achieving the goal of seamlessly obtaining related information during communication. Therefore, it can solve the problem in related technologies that it is impossible to provide relevant related information in real time according to the chat topics of the two parties during a call. Attached Figure Description
[0013] Figure 1 This is a hardware structure block diagram of a computer terminal for an information push method according to an embodiment of this application;
[0014] Figure 2 This is a flowchart (a) of an information push method according to an embodiment of this application;
[0015] Figure 3 This is a flowchart (II) of an information push method according to an embodiment of this application;
[0016] Figure 4 This is a flowchart (III) of the information push method according to an embodiment of this application;
[0017] Figure 5 This is a structural block diagram (a) of an information push device according to an embodiment of this application;
[0018] Figure 6 This is an architecture diagram of the information acquisition module according to an embodiment of this application;
[0019] Figure 7 This is a flowchart of model training according to an embodiment of this application;
[0020] Figure 8 This is a flowchart illustrating the knowledge graph construction and updating process according to embodiments of this application;
[0021] Figure 9 This is a flowchart (four) of the information push method according to an embodiment of this application;
[0022] Figure 10 This is a structural block diagram (II) of an information push device according to an embodiment of this application. Detailed Implementation
[0023] The embodiments of this application will be described in detail below with reference to the accompanying drawings and examples.
[0024] It should be noted that the terms "first," "second," etc., in the specification, claims, and drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.
[0025] The methods and embodiments provided in this application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Taking running on a computer terminal as an example, Figure 1 This is a hardware structure block diagram of the computer terminal used in the embodiments of the method of this application. For example... Figure 1 As shown, a computer terminal may include one or more ( Figure 1 Only one is shown in the diagram. A processor 102 (which may include, but is not limited to, a microprocessor MPU or a programmable logic device FPGA, etc.) and a memory 104 for storing data are also shown. The computer terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the computer terminal described above. For example, the computer terminal may also include components that are more complex than those described above. Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.
[0026] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to the information push method in this embodiment. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, thus implementing the above-described method. The memory 104 may include high-speed random access memory and non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to a computer terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0027] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by a communication provider for the computer terminal. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module used for wireless communication with the Internet.
[0028] This embodiment provides a method for pushing information on a computer terminal. Figure 2 This is a flowchart (a) of an information push method according to an embodiment of this application, such as... Figure 2 As shown, the process includes the following steps:
[0029] Step S202: Obtain audio data of the target object during the communication process, and determine the first semantic information of the audio data;
[0030] Audio data between target parties (such as the call initiator and receiver) is captured using the microphone of a mobile terminal. Then, speech recognition technology converts the audio into text, and natural language processing (NLP) algorithms are used to analyze the text content to determine the first semantic information of the audio data.
[0031] It should be noted that the target objects mentioned above can be two or more users.
[0032] Step S204: Determine the target information associated with the first semantic information and push the target information to the target device, wherein the target device is a device that has a binding relationship with at least one target object.
[0033] The target information can include news, article summaries, encyclopedic knowledge, etc., related to the discussion topic. Through intelligent integration with multiple data sources, such as news websites, social media platforms, and authoritative encyclopedic databases, the target information is automatically filtered and obtained based on the primary semantic information.
[0034] It should be noted that the corresponding data source is determined based on the communication type of the target audience. For example, when the communication type is public (such as non-work-related communication), the data source is a public data source, such as news websites, social media platforms, and authoritative encyclopedia databases; when the communication type is private (such as work-related communication), the data source is a non-public data source, such as internal company databases and knowledge bases.
[0035] For example, in the case of work-related communication, audio data is converted into text in real time using speech recognition technology. An enterprise-level semantic understanding model is then used to analyze the converted text and identify primary semantic information closely related to the company's business, such as project names, customer names, and product codes. Based on the primary semantic information, target information associated with the primary semantic information is filtered from the company's information database. The target information could be project progress reports, customer relationship management details, market analysis reports, etc.
[0036] It should be noted that the target device with the information push function enabled is determined, and the target information is pushed to the target device. The number of target devices can be one or more.
[0037] Optionally, if target information associated with the first semantic information is determined, the target information is pushed to the target device during the communication process.
[0038] That is, based on this information, the most relevant target information is determined, and this information is then pushed to the user in real time without interrupting the call or reducing the communication experience.
[0039] This application embodiment acquires and analyzes the audio data of both parties in a call in real time to determine its semantic information, and then intelligently pushes target information that is highly relevant to the discussion topic. This achieves the goal of seamlessly acquiring related information during the communication process, and thus solves the problem in related technologies that it is impossible to provide relevant related information in real time based on the chat topics of the two parties during a call.
[0040] Optionally, such as Figure 3 As shown, step S202 above can be implemented in the following way:
[0041] Step S301: Obtain reference keywords from the first data source, wherein the frequency of the reference keywords appearing in the first data source is greater than a preset frequency;
[0042] Step S302: Clean the reference keywords and determine the topic classification of the cleaned reference keywords;
[0043] Step S303: Establish the preset keyword library based on the cleaned reference keywords and the topic classification of the cleaned reference keywords;
[0044] It should be noted that in steps S301-S303 above, during the call, keywords need to be identified and matched to push related information. To ensure the quality and relevance of keywords, this application embodiment uses a strategy to obtain reference keywords from a first data source. The "first data source" can be a social platform (such as Weibo, WeChat), news website, professional forum, etc., which contain a large number of topics discussed by users.
[0045] High-frequency words are extracted from the primary data source, i.e., keywords that appear more frequently than a preset frequency in the data source. The preset frequency should be set based on data analysis and experience to filter out truly representative and popular words, while avoiding interference from overly common filler words or irrelevant words.
[0046] Even high-frequency words may contain irrelevant or misleading keywords, such as certain slang, advertising slogans, or colloquial expressions. These words are not suitable for matching related information. Therefore, the obtained reference keywords need to be cleaned to remove words that do not conform to the preset criteria, such as overly colloquial, advertising-related words, or filler words irrelevant to the topic. The cleaning process can include various natural language processing techniques such as regular expression matching, word frequency statistical analysis, and sentiment analysis.
[0047] Natural language processing techniques, such as TF-IDF, TextRank, or deep learning-based topic classification models, are used to classify the cleaned reference keywords into topics. By analyzing the context and co-occurrence relationships of words, the topic domain to which the words belong can be determined, thus establishing a keyword library containing topic classification information. Finally, the cleaned and topic-classified reference keywords are integrated into the library to form a pre-defined keyword library. The establishment and updating of the keyword library is dynamic and can be adjusted periodically or in real time according to data changes to ensure the timeliness and relevance of the keywords.
[0048] Step S304: Obtain the text information of the audio data and determine the first keyword in the text information;
[0049] Optionally, determining the first keyword in the text information includes: determining the weight value of each word in the text information based on a first algorithm, and determining candidate keywords among multiple words based on the weight value of each word;
[0050] A word graph model is constructed based on the candidate keywords, wherein the nodes of the word graph model are the candidate keywords, and the edges of the word graph model are the correlation between pairs of candidate keywords;
[0051] The weight of a node is determined based on the correlation degree corresponding to its outgoing edges and the correlation degree corresponding to its incoming edges, and the first keyword is determined from the candidate keywords based on the weight of the node.
[0052] Once call logs are converted to text, they contain a large number of words. Some of these words are closely linked to relevant information, while others are merely filler words in everyday conversation and lack informational value. Therefore, it is necessary to differentiate the contribution of each word to determine which words are more likely to indicate the current hot topics of discussion.
[0053] The importance of each word to the conversational text is calculated using algorithms such as TF-IDF (Term Frequency-Inverse Document Frequency) or deep learning methods. TF-IDF is a statistical method used to evaluate the importance of a word to a specific document in a collection of documents or a corpus. Term frequency (TF) reflects the number of times a word appears in a document, while inverse document frequency (IDF) measures the general importance of the word in the entire corpus; the rarer the word, the higher its IDF value. By considering both factors, each word is assigned a weight value; the higher the value, the more important the word is in the current conversational topic.
[0054] Subsequently, a word graph model was constructed based on these keywords, where each keyword is treated as a node, and the correlation between nodes is determined by the semantic similarity or co-occurrence frequency of the keywords, forming connecting edges. The weight of each node in the word graph model was analyzed; the weight calculation combined the correlation between the node's outgoing and incoming edges. The outgoing edge correlation reflects the connection between the keyword and other keywords, while the incoming edge correlation reflects the central position of the keyword in the call text. By calculating the node weights, the keywords closely related to the call topic in the word graph model, i.e., the primary keywords, were identified.
[0055] Imagine a phone call is taking place about "new environmentally friendly technologies," during which terms such as "solar energy," "battery recycling," and "carbon dioxide emissions" are mentioned.
[0056] Step 1: Calculate word weights;
[0057] Using the TF-IDF algorithm, the system assigned weights to these words. For example, "solar energy" appeared 10 times in the call text, but 500 times in the entire corpus; "battery recycling" appeared 5 times, but 100 times in the corpus; and "carbon dioxide emissions" appeared 3 times, but only 30 times in the corpus.
[0058] Based on the above information, the system calculated that "solar energy" has a low weight, "battery recycling" has a medium weight, and "carbon dioxide emissions" has the highest weight.
[0059] Step 2: Determine candidate keywords;
[0060] The threshold is set at the 75th percentile of all word weights. After calculating the weights of all words, the system determines this threshold to filter candidate keywords.
[0061] Based on the threshold screening, only "carbon dioxide emissions" exceeded the set standard and was therefore selected as a candidate keyword.
[0062] Step 3: Construct a word graph model and obtain related information;
[0063] Based on the identified candidate keyword "carbon dioxide emissions", the system constructs a word graph model to further analyze its semantic relationships and related information requirements in the call text.
[0064] Based on the word graph model, the system automatically obtains the latest related information on "carbon dioxide emissions" from relevant data sources (such as news APIs, professional forums, microblogs, etc.), such as the latest environmental protection policies, technological breakthroughs, industry trends, etc.
[0065] Step 4: Push related information.
[0066] Finally, the filtered relevant information is pushed to both parties' call interfaces in real time in the form of "information bubbles". Users can directly obtain summaries and links related to the current topic of "carbon dioxide emissions" through simple clicks or swipes, without interrupting the call or manually searching for information.
[0067] Step S305: Determine the second keyword based on the matching degree between the first keyword and the words in the preset keyword library;
[0068] Based on the first keyword, a matching process is performed with a pre-defined keyword library to determine a second keyword with high similarity. The keyword library contains domain vocabulary and topic classifications. By calculating the similarity (e.g., cosine similarity, Euclidean distance) between the first keyword and other keywords, a second keyword with a similarity greater than a pre-defined threshold is determined. The second keyword is assigned a specific topic classification within the pre-defined keyword library, and this topic classification is then used to determine the first topic classification of the current call audio data.
[0069] Optionally, step S305 can be implemented by matching the first keyword with words in the keyword database to be blocked, so as to determine the words to be blocked in the first keyword;
[0070] The words to be blocked in the first keyword are deleted, and the second keyword is determined based on the matching degree between the deleted first keyword and the words in the preset keyword library, so as to determine the second keyword whose similarity to the deleted first keyword is greater than the preset similarity.
[0071] In this embodiment of the application, after determining the first keyword, it is matched with a pre-set keyword library to be blocked to identify sensitive words that may be included. The keyword library to be blocked contains words that the user does not wish to mention or receive during the call, such as politically sensitive words, words related to adult content, and words related to violence.
[0072] Deleting the words to be blocked matched in the first keyword prevents the push of sensitive information and protects the privacy and data security of both parties in the call.
[0073] Through the above embodiments, by matching and processing keywords with the target word library, sensitive words in the call can be effectively blocked, while accurately locating the current discussion topic category, thereby pushing relevant information that is highly related to the topic and contains no sensitive content to the user.
[0074] Step S306: Determine the first topic category of the audio data based on the second keyword, wherein the first semantic information includes: the first keyword and the first topic category.
[0075] Optionally, determining the first topic category of the audio data based on the second keyword includes: determining the first topic category of the audio data based on the topic category corresponding to the second keyword.
[0076] Optionally, such as Figure 4 As shown, the "determining the target information associated with the first semantic information" in step S204 above includes:
[0077] Step S401: Train the target model;
[0078] Optionally, the target model is trained through the following steps: preprocessing historical data source information to obtain processed historical data source information; dividing the processed historical data source information into a training set and a validation set; training steps: training the target model according to the training set, and validating the trained target model according to the validation set to determine the performance information of the trained target model; repeating the training steps until the performance information meets the preset performance information, and obtaining the trained target model;
[0079] The collected historical data sources are cleaned and standardized to remove irrelevant noise data, such as advertisements and duplicate content. Preprocessing operations such as text segmentation and stemming are also performed to ensure the quality and consistency of the input data to the model. The preprocessed historical data sources are then divided into training and validation sets; for example, 80% of the data is used as the training set for model training, and the remaining 20% is used as the validation set to evaluate model performance.
[0080] First, the initial model is trained using a training set, enabling it to identify keywords and topic classifications related to text. Next, the model's prediction accuracy and stability are tested using a validation set, and performance data such as precision, recall, and F1 score are collected. If the model's performance does not meet the preset performance standards (i.e., the metrics are below a set threshold), model parameters such as learning rate and regularization coefficient are readjusted, and training and validation are repeated until the model's performance meets the preset conditions.
[0081] Optionally, after the initial training of the model is completed, new audio data and text information are continuously collected from the real-time call scenario. This new data undergoes the same preprocessing as historical data, and the model is fine-tuned using an online learning algorithm to update the model parameters and knowledge graph to adapt to the dynamic changes in the call content. This includes: preprocessing the incremental data source information to obtain processed incremental data source information; inputting the processed incremental data source information into the trained target model to obtain the second semantic information of the processed incremental data source information; and optimizing the parameters of the trained target model based on the backpropagation algorithm when the loss value between the second semantic information and the preset semantic information is greater than the preset loss.
[0082] The new data source information undergoes preprocessing, such as data cleaning, standardization, and feature extraction, to ensure consistent data format for easy model processing. Next, this processed incremental data is input into the target model to obtain its second semantic information. After obtaining the second semantic information, the difference between it and the preset semantic information expected to be output by the target model is compared, i.e., the loss value is calculated. If the loss value is greater than a preset threshold, it indicates that the model's understanding of the new information is flawed, or that the new information differs significantly from the original data distribution. In this case, the model parameters need to be optimized to improve the model's adaptability and accuracy to the new data.
[0083] The backpropagation algorithm allows for the adjustment of model parameters based on the calculated loss value, thereby optimizing model performance. Backpropagation is a commonly used neural network training method that adjusts network weights layer by layer based on the error between the output and the expected output, ultimately reducing the overall loss.
[0084] Through the above embodiments, the parameters of the target model are dynamically optimized by the backpropagation algorithm, which effectively improves the model's ability to process new data and its accuracy.
[0085] Step S402: Determine the first information associated with the first semantic information in the second data source;
[0086] Step S403: Determine the first feature information of the first information based on the trained target model, and determine the second feature information of the first information based on the knowledge graph;
[0087] Step S404: Determine the target information from the first information based on the first feature information and / or the second feature information;
[0088] It should be noted that in steps S402-S404 above, after the chat keywords and topics of the two parties in the call are analyzed through speech recognition and natural language processing technology, relevant articles, news clips or discussion posts will be quickly located and captured in the second data source (such as news websites, social media, professional forums) based on the first keyword and topic, using keyword search and topic matching algorithms.
[0089] To further refine the primary information, a trained target model, such as a deep learning model or a natural language processing model, is used to perform in-depth analysis of the primary information. The target model performs deep semantic understanding of the primary information and extracts key features, such as sentiment, timeliness, and credibility rating, which are the primary feature information of the primary information.
[0090] In addition to model analysis, it is also necessary to consider the position and value of information in a broader context and knowledge system. Therefore, this application embodiment also uses knowledge graphs (such as Wikipedia's knowledge graph or a terminology graph of a professional field) to determine the entity relationships, historical background, and precise meanings of domain terms involved in the first information, i.e., the second characteristic information of the first information, thereby enhancing the depth and breadth of information analysis.
[0091] After collecting and analyzing the first and second feature information of the first information, the target information fragments are selected from the first information set based on the comprehensive score of the feature information, such as the timeliness, credibility, and matching degree with user interests.
[0092] For example, User A and User B are discussing new trends in the development of "electric vehicles".
[0093] 1. Keyword and Topic Identification:
[0094] The system monitors and converts call audio into text in real time, identifies the keywords "electric vehicle" and "new energy," and recognizes the call topic as "new energy vehicle."
[0095] 2. Filter the first information from the second data source:
[0096] Based on keywords and themes, articles and discussion posts about "electric vehicles" and "new energy vehicles" were crawled from pre-set secondary data sources (such as professional websites and social media in the new energy industry) to form a collection of primary information.
[0097] 3. Deep Feature Analysis:
[0098] Using a trained target model, we perform deep semantic understanding on the first piece of information to evaluate its sentiment, novelty, and reliability.
[0099] 4. Knowledge graph query:
[0100] By querying the knowledge graph, we identified the entity relationships involved in "electric vehicles," such as Tesla, BYD, and battery technology, as well as the historical background of the development of "new energy vehicles" in the automotive industry and future predictions.
[0101] 5. Comprehensive analysis and target information determination:
[0102] By integrating the first and second feature information, the most relevant and recommended information to users A and B was finally determined, such as a news summary about "Tesla's latest electric vehicle range improvement" and a professional analysis about "BYD's latest breakthrough in electric vehicle battery technology".
[0103] 6. Target information push:
[0104] Finally, the selected target information is displayed in real time as a voice bubble on the call interface of users A and B. Users can click or long-press to view the news summary or jump to the original link.
[0105] Through the above embodiments, the process of determining and pushing the most relevant and deeply analyzed related information to the call interface through deep feature analysis and knowledge graph query ensures the relevance of the information and also improves the timeliness and cognitive value of the information.
[0106] Step S405: Determine the sorting of the target information based on the first feature information and / or the second feature information of the target information; push the target information to the target device according to the sorting of the target information;
[0107] The primary characteristic information includes indicators such as relevance and timeliness of the target information. The relevance of the target information is assessed based on the primary semantic information determined during the call, such as keywords and topic classifications. Furthermore, considering the characteristic that the value of related information diminishes over time, the publication time of the target information is also monitored and recorded to ensure that only the latest and most valuable content is pushed to the system.
[0108] The second set of characteristics includes factors such as the emotional tone of the information, the authority of the information source, and the richness of the information (e.g., the inclusion of multimedia elements). These characteristics are comprehensively evaluated by analyzing the text content, author reputation, and media format of the target information, thereby further refining the information delivery strategy.
[0109] The first and / or second feature information are used as input to the sorting algorithm. Through comparison, weighting, and other strategies, the priority of the target information is determined. Finally, following the determined target information sorting, the information is pushed to the terminal devices of both parties in the call in the optimal order. The user interface can be designed as a list, with related information arranged sequentially for easy browsing and selection by the user.
[0110] Step S406: Retrain the target model using incremental data.
[0111] Optionally, the target model is retrained using incremental data, including: preprocessing the incremental data source information to obtain processed incremental data source information; inputting the processed incremental data source information into the trained target model to obtain second semantic information of the processed incremental data source information; and optimizing the parameters of the trained target model based on the backpropagation algorithm when the loss value between the second semantic information and the preset semantic information is greater than the preset loss.
[0112] Optionally, after the initial training of the model is completed, new audio data and text information are continuously collected from real-time call scenarios. This new data is preprocessed in the same way as historical data, and the model is fine-tuned using online learning algorithms to update model parameters and knowledge graphs to adapt to the dynamic changes in call content.
[0113] The new data source information undergoes preprocessing, such as data cleaning, standardization, and feature extraction, to ensure consistent data format for easy model processing. Next, this processed incremental data is input into the target model to obtain its second semantic information. After obtaining the second semantic information, the difference between it and the preset semantic information expected to be output by the target model is compared, i.e., the loss value is calculated. If the loss value is greater than a preset threshold, it indicates that the model's understanding of the new information is flawed, or that the new information differs significantly from the original data distribution. In this case, the model parameters need to be optimized to improve the model's adaptability and accuracy to the new data.
[0114] The backpropagation algorithm allows for the adjustment of model parameters based on the calculated loss value, thereby optimizing model performance. Backpropagation is a commonly used neural network training method that adjusts network weights layer by layer based on the error between the output and the expected output, ultimately reducing the overall loss.
[0115] Through the above embodiments, the parameters of the target model are dynamically optimized by the backpropagation algorithm, which effectively improves the model's ability to process new data and its accuracy.
[0116] Optionally, after obtaining the second semantic information of the processed incremental data source information, the method further includes: determining the sensitivity of each parameter in the trained target model to the processed historical data source information; constructing a loss function based on the sensitivity; and calculating the loss value between the second semantic information and the preset semantic information based on the loss function.
[0117] As time progresses, information needs to be processed from new data sources, such as trending news and emerging trends on social media. Therefore, AI agents collect and preprocess these new data sources, transforming them into incremental data that the model can understand. During model updates, it is crucial to avoid forgetting existing knowledge. To this end, it is necessary to understand which parameters in the model are most sensitive to existing data—that is, which parameter adjustments are most likely to affect the prediction accuracy of historical data.
[0118] Therefore, in this embodiment of the application, the Elastic Weight Consolidation (EWC) technique is used. It determines which parameters should be more conservative when updating the model by calculating the contribution of parameters to the prediction results of historical data, i.e., the sensitivity of the parameters, so as to avoid large deviations in the understanding of historical data.
[0119] Optionally, the aforementioned sensitivity can be understood as the Fisher information matrix. After the initial training, the importance of each parameter is estimated by calculating the expected value of the second derivative of the target model output with respect to the parameter θ (i.e., the Fisher information matrix F) on the processed historical data source information. The Fisher information matrix reflects the sensitivity of the model parameters to the old dataset.
[0120] A composite loss function was then constructed based on parameter sensitivity. This function not only considers the prediction error of incremental data (new data) but also includes a penalty term for the decrease in prediction accuracy for historical data. Through this dual consideration, it ensures that model updates focus on adaptability to new data without sacrificing accuracy for historical data. Using the constructed loss function, the difference between the second semantic information and the preset semantic information is calculated; this difference is represented by the loss value. The smaller the loss value, the more accurate the model's understanding of the new information after the update, and the smaller the impact on historical data, and vice versa.
[0121] Optionally, define the loss function L for the new data. new (θ), and based on this, an EWC memory regularization term is introduced to form a joint loss function: Where, θ 0 These are parameters obtained during the learning phase of old data (i.e., processed historical data source information), λ is the regularization factor used to balance the importance of learning new and old data, and F is the Fisher information matrix.
[0122] Optionally, before determining the first feature information of the first information according to the target model, the process includes: identifying the first entity information in the historical data source information according to an entity recognition algorithm, and identifying the second entity information in the data source information according to a preset knowledge graph entity; identifying the relationship information between pairs of entities according to a syntactic analysis algorithm and a relational reasoning algorithm, wherein the entity is the entity indicated by the first entity information or the second entity information; and constructing the knowledge graph according to the first entity information, the second entity information, and the relationship information between the pairs of entities.
[0123] Entity recognition algorithms are used to identify mentioned entities, such as "Tesla" and "Apple," from historical data sources (such as text data from past call records). These entity information forms the basic elements of the knowledge graph.
[0124] Based on a pre-defined knowledge graph, entities in the data source information are further identified and linked or expanded with existing entities. The pre-defined knowledge graph contains a wide range of entity information, such as historical events, well-known figures, and brands, which helps to improve the graph's coverage and the relationships between entities.
[0125] Syntactic analysis algorithms can be used to parse out relationships between entities, such as the association between "Tesla" and "electric vehicles" or the founder relationship between "Apple" and "Steve Jobs". Furthermore, relational reasoning algorithms can not only identify directly mentioned relationships but also infer implicit connections between entities; for example, inferring from news reports that "Tesla's" stock price may fluctuate due to changes in the "electric vehicle" market.
[0126] Finally, the identified first entity information, second entity information, and relationship information between entities are integrated to construct or update a knowledge graph. In the graph, entities are nodes, and relationships are edges, forming a network structure that reflects the connections between entities.
[0127] Optionally, acquiring the target's audio data during the communication process includes: determining whether the target has activated the push function; and if the target has activated the push function, acquiring the target's audio data during the communication process.
[0128] To protect user privacy and improve user experience, the call-based related information push service can only be activated with user authorization. In this embodiment, a user interface is designed that allows all parties in a call to activate or deactivate the related information push function via voice commands, touch operations, or settings options. Activation of the function requires explicit user instructions, such as the keyword "activate voice bubble" or clicking a button on the interface.
[0129] After a user activates the push notification function, audio data is captured in real time through the microphones on both devices during the call, and then the audio data is converted into text information.
[0130] To better understand the process of the above information push method, the implementation flow of the above information push method will be described below in conjunction with optional embodiments, but it is not intended to limit the technical solution of the embodiments of this application.
[0131] This embodiment provides a method for pushing information. Figure 5 This is a structural block diagram (I) of an information push device according to an embodiment of this application, as shown below. Figure 5 As shown, it includes the following modules:
[0132] I. Terminal;
[0133] II. Speech Recognition Server:
[0134] The speech recognition server is used to acquire keywords from the conversation between the two parties. Upon recognizing a word that matches a pre-set keyword database, the speech recognition server quickly responds, generates a specific business control command, and sends the command to the command processing module to activate the system's related information push function.
[0135] III. User Command Processing Module:
[0136] The user command processing module handles commands to enable or disable the associated information push function. The associated information push function can be activated via the phone's keypad or directly via voice.
[0137] IV. Keyword Analysis Module:
[0138] The keyword analysis module parses call content to identify high-frequency topic keywords, thereby improving the accuracy and timeliness of related information pushes. Through Natural Language Processing (NLP) technology, combined with a lightweight database of commonly used chat topic keywords, it acquires and categorizes topics that appear in daily conversations, such as technology, entertainment, and lifestyle. Furthermore, this module integrates an anti-interference word database and a sensitive word database, eliminating invalid or inappropriate words to ensure that keyword extraction is both efficient and secure.
[0139] While everyday conversations cover a wide range of topics, they often revolve around a few core areas. This makes it possible to build a lightweight keyword library for frequently used chat topics. For example, under the "Technology" category, the system would further subdivide it into several subcategories such as smartphones, artificial intelligence, computer hardware, and technology news. Each category contains a series of highly relevant keywords, thus constructing a comprehensive keyword library.
[0140] The keyword database is established in the following ways:
[0141] 1) First, common chat topics will be meticulously categorized and organized to construct a multi-level, structured classification system. This system is not limited to primary categories, such as technology, entertainment, lifestyle, sports, health, education, and finance, but will be further subdivided into secondary categories and sub-topics to increase the depth and breadth of the keyword database. For example, under the primary category of "technology," subcategories will be made into "smartphones," "artificial intelligence," "computer hardware," and "technology news," among others.
[0142] 2) Collect high-frequency keywords and related phrases for each topic to ensure the richness and representativeness of the keyword database. On the one hand, refer to hot word lists released by authoritative institutions, industry analysis reports, and hot topic reviews from mainstream news media. These resources can provide keywords that are highly representative and attract significant attention in specific fields or time periods. On the other hand, utilize advanced web crawling technology to scrape massive amounts of user-generated content from highly interactive websites such as social media platforms, online forums, and Q&A communities. Through in-depth mining and analysis of this text data, extract the core keywords and typical expressions for each chat topic.
[0143] In addition, the annotation team can manually filter and annotate some data in detail to supplement and improve the completeness of the keyword database.
[0144] 3) Keywords that are not highly relevant to the topic, are too obscure or difficult to understand, or are used very infrequently will be removed to avoid wasting resources on irrelevant information. Next, based on key indicators such as keyword frequency of appearance in different data sources, search popularity, and dissemination scope, combined with evaluations and subjective judgments from domain experts, a weight value will be assigned to each keyword. The weight value directly reflects the importance and representativeness of the keyword in the corresponding topic.
[0145] For example, within the "Movies" subcategory of the "Entertainment" topic, keywords such as "box office," "director," and "actors" will be given higher weight because they typically enjoy high attention and discussion in movie-related chat content. Conversely, relatively detailed and non-core terms such as "cinema seat number" will be given lower weight.
[0146] In addition, after establishing the keyword database, it is also necessary to ensure that the keyword database is lightweight, specifically:
[0147] S11: Based on the original design intent of the system and the actual needs of the target user group, an upper limit was set for the number of keywords in the keyword library. For example, the range is controlled between 500 and 1000 keywords. By limiting the number of keywords, we ensured the conciseness of the keyword library, making it more focused on high-frequency and representative words, and improving the efficiency and accuracy of keyword matching.
[0148] S12: Regularly conduct statistical analysis of the usage frequency of words in the keyword database. Keywords that have not been used for a long time or whose usage frequency is below a set threshold will be removed from the database or merged. Simultaneously, monitor new developments in social networks, news media, and other fields, and evaluate newly emerging trending keywords. If these keywords demonstrate high potential and relevance, outdated or less important keywords will be replaced in a timely manner to ensure that the keyword database always reflects current hot topics and trends.
[0149] S13: By identifying and merging synonyms or near-synonyms. For example, unify emotional words such as "happy," "joyful," and "joyful" into "happy," and flexibly expand upon them according to the specific context in practical applications.
[0150] After establishing a keyword database, semantic information in the communication data is determined based on the keyword database, including:
[0151] S21: Employ multiple keyword extraction techniques to ensure that the keywords extracted from the preprocessed call text are both representative and accurately reflect the semantic information of the topic. Specific implementation strategies include:
[0152] The TF-IDF algorithm calculates the importance weight of each word by multiplying its term frequency (TF) and inverse document frequency (IDF) in the text. Based on this, a batch of candidate keywords with higher weights is selected.
[0153] TextRank algorithm: A graph-based ranking algorithm that evaluates the semantic connections and importance between words. By constructing a co-occurrence network of words, it identifies and strengthens keywords that play a central role in topic discussions, thereby extracting the most representative and semantically relevant keyword combinations.
[0154] Semantic resource expansion: By combining authoritative semantic dictionaries such as HowNet and using semantic similarity calculation methods, words that do not appear directly in the keyword database but are semantically closely related to the current topic are identified. These words are then converted into semantically relevant keywords.
[0155] S22: During the matching process between keywords and a pre-defined chat topic keyword database, a strategy combining exact matching and fuzzy matching is employed. Specific implementation strategies include:
[0156] Exact match: By direct comparison, locate keywords that match the vocabulary in the keyword library highly, and obtain their corresponding topic classification and weight information.
[0157] Fuzzy matching: Considering the diversity of languages and the variations in vocabulary, this application introduces a fuzzy matching mechanism to identify keywords that are highly similar to words in the keyword database in terms of shape, pronunciation, or semantics (the similarity threshold is set above 80%). By associating these similar keywords with the same topic category, the recall rate of topic identification is improved, ensuring that the core of the topic can be accurately captured even in contexts using synonyms or regional dialects.
[0158] Anti-interference and sensitive word filtering: To prevent casual conversation words (such as "hehe" or "haha") from interfering with topic identification, extracted keywords are compared with a preset anti-interference word library. Keywords matching the anti-interference word library will be filtered by the system to ensure the accuracy and focus of topic identification. Simultaneously, the system will also compare with a sensitive word library, automatically blocking politically sensitive words, violence, adult content, and other inappropriate information.
[0159] V. Information Acquisition Module:
[0160] The information acquisition module is used to collect and analyze key parts of information highly relevant to the call content in real time. Through a collaborative working model between the Agent and the LLM (Large Language Model), the timeliness and accuracy of the information are ensured. For example... Figure 6 As shown, the information acquisition module includes: data source layer, data acquisition layer, LLM analysis layer, and hotspot identification layer.
[0161] 1. Data source layer:
[0162] The data source layer is the origin of related information, including mainstream social media platforms (such as Weibo and Xiaohongshu), news media websites (such as Xinhua News Agency and Toutiao), professional forums and communities (such as Zhihu and GitHub), and authoritative encyclopedic materials such as Wikipedia. Through technologies such as social media APIs, RSS subscriptions, and web crawlers, data is provided to the system, forming the foundation of related information. It should be noted that the definition of the data source layer is not limited to public news hotspots. Depending on the specific use case, such as internal enterprise communication needs, the data source layer can also be extended to company information databases and professional databases to achieve more customized and professional information delivery.
[0163] 2. Data Acquisition Layer:
[0164] The agent interacts with the aforementioned data sources to execute efficient data collection tasks. The collected data first undergoes a series of preprocessing steps, including data cleaning, standardization, and feature extraction, to ensure the consistency and analyzability of the information.
[0165] Data preprocessing mainly consists of the following three steps:
[0166] Data cleaning: This involves removing non-text elements such as HTML tags, special characters, and emojis from the text. Regular expressions are also used to clean and replace noisy information (such as advertising links and redundant punctuation). For duplicate social media posts or news information, text hash comparison or fuzzy matching algorithms are employed for data deduplication.
[0167] Data standardization: The collected text is uniformly converted into lowercase form, and word form restoration and stemming techniques are used to normalize anagrams or cognate words to their basic forms.
[0168] Feature extraction: The TF-IDF algorithm is used to calculate the relative importance of each word in the text, constructing a feature vector for the text. Furthermore, word embedding models such as Word2Vec and GloVe are used to map words to a vector space with semantic information, thereby obtaining semantic similarity and contextual relationships between words. For time-series data, time features such as hour, date, and day of the week are extracted and combined with popularity metrics such as likes, comments, and shares for comprehensive analysis.
[0169] 3. LLM Analysis Layer:
[0170] The LLM analysis layer leverages the powerful learning and understanding capabilities of large-scale language models to perform deep semantic analysis on preprocessed data, such as identifying keywords and themes related to the information, and analyzing sentiment. The incremental learning-based AI agent design allows the model to be quickly adjusted and updated based on only a small amount of new data after initial training. This online learning capability ensures that the system can still respond promptly and accurately in environments with rapidly changing related information, without the need for frequent large-scale training from scratch, significantly saving computational resources and time costs.
[0171] This application also provides a model training method in an LLM analysis layer, such as... Figure 7 As shown, the model training method includes the following steps:
[0172] Step S701: Select model architecture;
[0173] The Transformer architecture was chosen as the core model, primarily based on its self-attention mechanism, which can capture long-distance dependencies and rich semantic information in text. Within the PyTorch framework, a basic Transformer model was constructed using torch.nn.TransformerEncoder and torch.nn.TransformerDecoder. For the characteristics of related text information, hyperparameters such as the embedding layer dimension, the number of encoder layers, and the number of decoder layers were fine-tuned.
[0174] Step S702: Obtain initial training data (equivalent to historical data source information in the above embodiment);
[0175] Representative samples are selected from historical trending data to serve as the initial training dataset for the model. This data should cover different types of trending topics, text styles, and time spans to ensure that the model can learn a wide range of knowledge and feature patterns.
[0176] Step S703: Divide the training set and the validation set;
[0177] The initial training data is divided into a training set and a validation set, with the training dataset divided into the training set and the validation set in a ratio of 80% / 20%.
[0178] Step S704: Define the loss function and optimization algorithm;
[0179] Apply the Adam optimizer and an appropriate loss function (use cross-entropy loss for classification tasks and mean squared error loss for generation tasks);
[0180] Step S705: Model training;
[0181] The model is trained iteratively on the training set.
[0182] Step S706: Verify model performance;
[0183] Regularly evaluate model performance on the validation set, monitoring key metrics such as accuracy, recall, F1 score, and loss. Adjust parameters such as learning rate and regularization based on validation results to prevent overfitting. When validation metrics stabilize after several epochs with minimal fluctuations, it indicates that the model has completed training and reached a good equilibrium.
[0184] Step S707: Incremental learning AI Agent implementation;
[0185] Step S7071: Online learning algorithm: Ensure the model can process and learn new data points or mini-batch data in real time. After data preprocessing, perform forward propagation calculations to obtain prediction results. Update model parameters using the backpropagation algorithm and optimizer.
[0186] Step S7072: Strategies to prevent catastrophic forgetting: Introduce memory regularization terms (such as EWC) to balance the learning of new and old data.
[0187] Step S7073: Incremental Learning: Establish an incremental learning process manager responsible for data flow scheduling, model parameter update coordination, and knowledge graph storage and updates. When new hot data arrives, preprocess it first, then update the model parameters, periodically store the model status, monitor performance indicators, and automatically trigger retraining or fine-tuning when necessary.
[0188] Optionally, this application also provides a method for constructing and maintaining a knowledge graph, such as... Figure 8 As shown, it includes:
[0189] Step S801: In the case of knowledge graph updates, identify and link entities in the acquired information, including:
[0190] Step S8011: Use the NER function of SpaCy or NLTK to identify entities such as people, places, organizations, and events from hotspot text.
[0191] Step S8012: Link the new entity with existing entities in the knowledge graph, for example, by using clustering algorithms or similarity calculation methods (such as string similarity, semantic similarity) for matching.
[0192] Step S8013: Add the identified entities to the knowledge graph and assign them unique identifiers and attribute information;
[0193] Step S802: Extracting relationships between entities, including:
[0194] Step S8021: Identify the types of relationships between entities (such as publishing, participating, influencing, etc.) through dependency parsing and pattern matching, and create or update the corresponding edges in the knowledge graph;
[0195] Step S8022: Update the edges and attributes in the knowledge graph: Add relational attributes (time, location, etc.).
[0196] Step S8023: Relationship Reasoning: Using algorithms such as rule-based reasoning and path-based reasoning, potential new relationships are derived based on existing relationships and added to the graph.
[0197] Step S803: Optimize map storage, including:
[0198] Step S8031: Select a graph database (such as Neo4j) or RDF triple format to store the knowledge graph;
[0199] Step S8032: Index optimization and query optimization.
[0200] Step S804: Store the updated knowledge graph.
[0201] 4. Information Recognition Layer:
[0202] To accurately and promptly identify current trending topics and conduct in-depth analysis of their popularity levels, sentiment trends, etc., such as... Figure 9 As shown, the specific steps are as follows:
[0203] Step S901: Build a real-time inference engine;
[0204] Step S9011: Preprocessing new related information text: For newly collected related information text, preprocessing operations are first performed, such as noise removal, word segmentation and standardization. Then feature extraction is performed, and techniques such as TF-IDF, Word2Vec or GloVe are used to convert the text into a form that the model can understand.
[0205] Step S9012: Input incremental learning model: Input the preprocessed text features into the incremental learning-based model to obtain key indicators such as the popularity prediction and sentiment of the text topic.
[0206] Step S9013: Query relevant entities and relationships in the knowledge graph: Use the updated knowledge graph to query entities and relationships related to the text. Through algorithms such as rule-based reasoning and Bayesian network reasoning, delve deeper into the event background, scope of influence, and possible development trends to obtain deeper semantic information.
[0207] Step S9014: Generate a comprehensive analysis report: Combine the model prediction results with the reasoning information from the knowledge graph to generate a comprehensive analysis report on the related information. The report includes not only topic popularity and sentiment analysis, but also background interpretation of the event and potential development predictions.
[0208] Step S902: Adaptive decision-making algorithm;
[0209] Step S9021: Determine Information Display Priority: Based on the real-time analysis results of related information, the algorithm determines the priority of information display. For events with high popularity and timeliness, priority is given to pushing them to user groups who are interested in or active in these topics, ensuring the timeliness and relevance of the information.
[0210] Step S9022: Determine push strategy adjustment: The decision algorithm will dynamically adjust based on user feedback behavior (such as bubble click rate, browsing duration, likes and interactions) and the dynamic changes of hot topics (popularity trend, sentiment in comments, etc.).
[0211] Step S9023: Dynamically adjust the push strategy: By continuously learning user preferences and the evolution of trending topics, the decision-making algorithm can intelligently optimize the push content to meet users' personalized needs and improve overall experience satisfaction.
[0212] Step S903: Push related information.
[0213] VI. Call Media Stream RTP Processing Module;
[0214] The call media stream RTP processing module is responsible for maintaining the stability of the call media stream under the RTP protocol, ensuring high-quality transmission of audio and video calls. It inserts relevant information bubbles at appropriate times during the call without negatively impacting call quality, achieving a harmonious coexistence of information push and call experience.
[0215] VII. Dynamic speech bubble display module;
[0216] The dynamic voice bubble display module presents relevant keywords on the call interface. Users can click to view a summary, and long-press to jump to a 24-hour timeline of trending topics related to the topic, including news and encyclopedic information. The module also supports links to relevant trending news, allowing users to continue exploring the topic after the call ends, meeting their information needs in different scenarios.
[0217] 8. Intelligent anti-interference module;
[0218] The intelligent anti-interference module is used for at least one of the following:
[0219] 1. A dynamic algorithm is employed to adjust the number of speech bubbles based on the speaking speed of both parties in the call (limited to a maximum of 3 per minute). This ensures that the bubbles do not excessively interfere with the conversation, maintaining topic coherence while providing appropriate supplementary information. Users typically discuss a topic with relevance within a specific timeframe, rather than changing the topic with each sentence. The LLM analysis layer's reasoning and decision-making steps prioritize relevant information (popularity, user preference) within a given timeframe and push it to the user. In practical applications, frequent pop-ups and changes in the user interface can also negatively impact the user experience, thus requiring a limit on the number of speech bubbles pushed per minute.
[0220] 2. Establish an anti-interference keyword database: Analyze call data using NLP technology, and identify and archive words that frequently appear in casual conversation but have low information content, such as "hehe", "haha", and "nothing much", through algorithms such as TF-IDF and TextRank, to form an anti-interference keyword database.
[0221] 3. Establish a sensitive word database: Collect relevant words for sensitive topics such as politics, adult content, violence, and crime, and build a sensitive word database to ensure the appropriateness and legality of pushed information.
[0222] 4. User-defined blocking: Allows users to adjust the blocked word library according to their personal preferences, and protects users' customized information through encrypted storage and transmission, eliminating the risk of sensitive information leakage.
[0223] Through the collaborative work of the above-mentioned functional components, the embodiments of this application can push relevant information in real time according to the chat topics of the two parties during a basic call, thereby improving the interactivity and practicality of the call.
[0224] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0225] This embodiment also provides an information push device, which is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0226] Figure 10 This is a structural block diagram of an information push device according to an embodiment of this application, such as... Figure 10 As shown, the device includes:
[0227] The determining module 1002 is used to acquire audio data of the target object during the communication process and determine the first semantic information of the audio data;
[0228] The push module 1004 is used to determine target information associated with the first semantic information and push the target information to a target device, wherein the target device is a device that has a binding relationship with at least one target object.
[0229] The above-described device acquires audio data of the target object during communication and determines the first semantic information of the audio data; it then determines target information associated with the first semantic information and pushes the target information to the target device, wherein the target device is a device that has a binding relationship with at least one of the target objects. That is, this embodiment of the application acquires and analyzes the audio data of both parties in real time, determines their semantic information, and intelligently pushes target information that is highly related to the discussion topic, thereby achieving the goal of seamlessly acquiring related information during communication. Therefore, it can solve the problem in related technologies that it is impossible to provide relevant related information in real time based on the chat topics of the two parties during a call.
[0230] In an exemplary embodiment, a determining module is configured to determine first semantic information of the audio data, including: acquiring text information of the audio data and determining a first keyword in the text information; determining a second keyword based on the matching degree between the first keyword and words in a preset keyword library; and determining a first topic category of the audio data based on the second keyword, wherein the first semantic information includes: the first keyword and the first topic category.
[0231] In an exemplary embodiment, a determining module is configured to determine a first keyword in the text information, the method comprising: determining a weight value for each word in the text information based on a first algorithm, and determining candidate keywords among multiple words based on the weight value of each word; constructing a word graph model based on the candidate keywords, wherein the nodes of the word graph model are the candidate keywords, and the edges of the word graph model are the correlation degrees between pairs of candidate keywords; determining the weight of a node based on the correlation degrees corresponding to the outgoing edges and the incoming edges of the node, and determining the first keyword among the candidate keywords based on the weight of the node.
[0232] In an exemplary embodiment, the determining module matches the first keyword with words in a preset keyword library, including: matching the first keyword with words in a keyword library to be blocked to determine words to be blocked in the first keyword; deleting words to be blocked in the first keyword; and determining a second keyword based on the matching degree between the deleted first keyword and words in the preset keyword library.
[0233] In an exemplary embodiment, a determining module is configured to obtain reference keywords from a first data source, wherein the frequency of the reference keywords appearing in the first data source is greater than a preset frequency; clean the reference keywords and determine the topic classification of the cleaned reference keywords; and establish the preset keyword library based on the cleaned reference keywords and the topic classification of the cleaned reference keywords.
[0234] In one exemplary embodiment, the push module is configured to determine first information associated with the first semantic information in a second data source; determine first feature information of the first information according to a trained target model, and determine second feature information of the first information according to a knowledge graph; and determine the target information in the first information based on the first feature information and / or the second feature information.
[0235] In one exemplary embodiment, the push module is configured to determine the order of the target information based on first feature information and / or second feature information of the target information; and push the target information to the target device according to the order of the target information.
[0236] In an exemplary embodiment, the push module is used to preprocess historical data source information to obtain processed historical data source information; divide the processed historical data source information into a training set and a validation set; training steps: train the target model according to the training set, and validate the trained target model according to the validation set to determine the performance information of the trained target model; repeat the training steps until the performance information meets the preset performance information, and obtain the trained target model.
[0237] In an exemplary embodiment, the push module is configured to preprocess the incremental data source information to obtain processed incremental data source information; input the processed incremental data source information into a trained target model to obtain second semantic information of the processed incremental data source information; and optimize the parameters of the trained target model based on a backpropagation algorithm when the loss value between the second semantic information and preset semantic information is greater than a preset loss.
[0238] In an exemplary embodiment, the push module is configured to determine the sensitivity of each parameter in the trained target model to processed historical data source information; construct a loss function based on the sensitivity; and calculate the loss value between the second semantic information and preset semantic information based on the loss function.
[0239] In an exemplary embodiment, the push module is configured to identify first entity information in the historical data source information according to an entity recognition algorithm, and identify second entity information in the data source information according to a preset knowledge graph entity; identify relationship information between pairs of entities according to a syntactic analysis algorithm and a relational reasoning algorithm, wherein the entity is the entity indicated by the first entity information or the second entity information; and construct the knowledge graph according to the first entity information, the second entity information and the relationship information between the pairs of entities.
[0240] In one exemplary embodiment, a determining module is configured to determine whether the target object has activated the push function; if the target object has activated the push function, the module acquires audio data of the target object during the communication process.
[0241] In one exemplary embodiment, the push module is configured to push the target information to a target device during communication if target information associated with the first semantic information is determined.
[0242] It should be noted that the above modules can be implemented by software or hardware. For the latter, they can be implemented in the following ways, but are not limited to: all the above modules are located in the same processor; or, the above modules are located in different processors in any combination.
[0243] Embodiments of this application also provide a computer-readable storage medium storing a computer program, wherein the computer program is configured to execute the steps in any of the above method embodiments when run.
[0244] In one exemplary embodiment, the aforementioned computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.
[0245] Embodiments of this application also provide an electronic device including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.
[0246] In one exemplary embodiment, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.
[0247] Embodiments of this application also provide a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above method embodiments.
[0248] Embodiments of this application also provide another computer program product, including a non-volatile computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in any of the above method embodiments.
[0249] Embodiments of this application also provide a computer program that includes computer instructions stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the steps in any of the above method embodiments.
[0250] Specific examples in this embodiment can be found in the examples described in the above embodiments and exemplary implementations, and will not be repeated here.
[0251] Obviously, those skilled in the art should understand that the modules or steps of this application described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. They can be implemented using computer-executable program code, and thus can be stored in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those presented here, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, this application is not limited to any particular combination of hardware and software.
[0252] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the principles of this application should be included within the protection scope of this application.
Claims
1. A method for pushing information, characterized in that, include: Acquire audio data of the target object during the communication process, and determine the first semantic information of the audio data; Determine target information associated with the first semantic information and push the target information to a target device, wherein the target device is a device that has a binding relationship with at least one target object.
2. The method according to claim 1, characterized in that, Determining the first semantic information of the audio data includes: Obtain the text information of the audio data and determine the first keyword in the text information; The second keyword is determined based on the matching degree between the first keyword and the words in the preset keyword library; The first topic category of the audio data is determined based on the second keyword, wherein the first semantic information includes the first keyword and the first topic category.
3. The method according to claim 2, characterized in that, Determining the first keyword in the text information, the process includes: The weight value of each word in the text information is determined based on the first algorithm, and candidate keywords are determined from multiple words based on the weight value of each word. A word graph model is constructed based on the candidate keywords, wherein the nodes of the word graph model are the candidate keywords, and the edges of the word graph model are the correlation between pairs of candidate keywords; The weight of a node is determined based on the correlation degree corresponding to its outgoing edges and the correlation degree corresponding to its incoming edges, and the first keyword is determined from the candidate keywords based on the weight of the node.
4. The method according to claim 2, characterized in that, Based on the matching degree between the first keyword and words in the preset keyword library, a second keyword is determined, including: The first keyword is matched with words in the keyword database to be blocked in order to determine the words to be blocked in the first keyword. The words to be blocked in the first keyword are deleted, and the second keyword is determined based on the matching degree between the deleted first keyword and the words in the preset keyword library.
5. The method according to claim 2, characterized in that, Before determining the second keyword based on the matching degree between the first keyword and words in the preset keyword library, the method further includes: Obtain reference keywords from a first data source, wherein the frequency of the reference keywords appearing in the first data source is greater than a preset frequency; The reference keywords are cleaned, and the topic classification of the cleaned reference keywords is determined. The preset keyword library is established based on the cleaned reference keywords and the topic classification of the cleaned reference keywords.
6. The method according to claim 1, characterized in that, Determining the target information associated with the first semantic information includes: Determine the first information associated with the first semantic information in the second data source; The first feature information of the first information is determined based on the trained target model, and the second feature information of the first information is determined based on the knowledge graph. The target information is determined from the first information based on the first feature information and / or the second feature information.
7. The method according to claim 6, characterized in that, Pushing the target information to the target device includes: The sorting of the target information is determined based on the first feature information and / or the second feature information of the target information; The target information is pushed to the target device according to the sorting of the target information.
8. The method according to claim 6, characterized in that, Before determining the first feature information of the first information based on the trained target model, the method further includes: The historical data source information is preprocessed to obtain the processed historical data source information; The processed historical data source information is divided into a training set and a validation set; Training steps: Train the target model using the training set, and validate the trained target model using the validation set to determine the performance information of the trained target model; The training steps are repeated until the performance information matches the preset performance information, and the trained target model is obtained.
9. The method according to claim 6, characterized in that, After determining the first feature information of the first information based on the trained target model, the method further includes: The incremental data source information is preprocessed to obtain the processed incremental data source information; The processed incremental data source information is input into the trained target model to obtain the second semantic information of the processed incremental data source information; If the loss value between the second semantic information and the preset semantic information is greater than the preset loss, the parameters of the trained target model are optimized based on the backpropagation algorithm.
10. The method according to claim 9, characterized in that, After obtaining the second semantic information of the processed incremental data source information, the method further includes: Determine the sensitivity of each parameter in the trained target model to the processed historical data source information; Construct a loss function based on the sensitivity; The loss value between the second semantic information and the preset semantic information is calculated based on the loss function.
11. The method according to claim 8, characterized in that, Before determining the first feature information of the first information based on the target model, the following steps are included: The system identifies the first entity information in the historical data source information based on the entity recognition algorithm, and identifies the second entity information in the data source information based on the preset knowledge graph entity recognition algorithm. The relationship information between pairs of entities is identified based on syntactic analysis algorithms and relational reasoning algorithms, wherein the entity is the entity indicated by the first entity information or the second entity information; The knowledge graph is constructed based on the first entity information, the second entity information, and the relationship information between the pairs of entities.
12. The method according to claim 1, characterized in that, Obtain audio data of the target audience during the communication process, including: Determine whether the target object has enabled the push function; If the target object activates the push function, acquire the target object's audio data during the communication process.
13. The method according to any one of claims 1 to 12, characterized in that, Pushing the target information to the target device includes: If target information associated with the first semantic information is determined, the target information is pushed to the target device during the communication process.
14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein the computer program, when executed by a processor, implements the steps of the method described in any one of claims 1 to 13.
15. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method described in any one of claims 1 to 13.
16. A computer program product, characterized in that, Includes a computer program, which, when executed by a processor, implements the steps of the method described in any one of claims 1 to 13.