Text pushing method and device, electronic equipment and storage medium

By extracting text keywords and setting classification criteria, business card category clusters are formed, and the matching degree is calculated. This solves the problem of inaccurate information push in existing technologies, realizes personalized and intelligent text push, and improves user satisfaction and information reception rate.

CN121350341APending Publication Date: 2026-01-16SHENZHEN VALUE ONLINE INFORMATION POLYTRON TECH INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511321270.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-15
Publication Date
2026-01-16

AI Technical Summary

Technical Problem

Existing text push methods rely on static user profiles and simple keyword matching, which fail to fully uncover users' real needs and interests, resulting in low accuracy of information push and low user satisfaction.

Method used

By extracting keywords from the text to be pushed, setting classification criteria, forming business card category clusters, and calculating the matching degree between the text and the clusters, user interests can be accurately located to achieve personalized information push.

Benefits of technology

It improved the accuracy of information delivery and user satisfaction, avoided information overload, and increased information reception rate and processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121350341A_ABST
    Figure CN121350341A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of text pushing, and discloses a text pushing method and device, electronic equipment and a storage medium, and the method comprises the steps: extracting a plurality of keywords in a to-be-pushed text, deeply understanding a core theme of content and a real demand of a user, and pushing the to-be-pushed text to the to-be-pushed text; the traditional limitation that only static user portraits and simple keyword matching are relied on is broken through, subsequent classification benchmark setting and business card classification enable user groups to be reasonably and dynamically divided, strong pertinence of information pushing is ensured, interference of irrelevant information is avoided, and finally, by calculating the matching degree of a to-be-pushed text and a business card classification cluster, the information pushing efficiency is improved. The interests of the user can be accurately positioned, personalized and intelligent information pushing is realized, and the user can receive information conforming to the preference of the user. The method has the beneficial effects that the receiving rate and the processing efficiency of the information are improved, the problem of user information overload is effectively solved, and the information pushing precision and the user satisfaction degree are greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of text push technology, and in particular to a text push method, apparatus, electronic device, and storage medium. Background Technology

[0002] With the rapid development of information technology, users are bombarded with a massive amount of text messages every day. In this context, effectively delivering relevant information to the right users has become a pressing issue. In practice, different users have significantly different preferences and needs for information. Simple information pushes can easily lead to information overload, ultimately causing users to filter and ignore the messages. Traditional information push methods often rely on static user profiles and simple keyword matching, failing to fully uncover users' true needs and interests, thus affecting both the accuracy of information pushes and user satisfaction. Summary of the Invention

[0003] Therefore, it is necessary to propose a text push method, device, electronic device and storage medium to address the existing text push problems.

[0004] A text push method, the method comprising:

[0005] Obtain multiple texts to be pushed, and extract multiple text keywords from the multiple texts to be pushed;

[0006] A classification criterion is set based on multiple text keywords;

[0007] Based on the aforementioned classification criteria, each preset business card is classified to obtain multiple business card classification clusters;

[0008] Calculate the matching degree between each of the texts to be pushed and the business card category cluster;

[0009] Based on the matching degree, at least one matching business card category cluster is selected for each text to be pushed;

[0010] Based on the selection results, the text to be pushed will be sent to the users of the corresponding business card category cluster.

[0011] Further, the step of calculating the matching degree between each of the texts to be pushed and the business card category cluster includes:

[0012] Obtain the text keywords and the corresponding number of keywords for each of the texts to be pushed;

[0013] Set corresponding dimension thresholds for each text keyword;

[0014] Detect the number of target business cards in each business card category cluster that is greater than or equal to the stated dimension threshold;

[0015] According to the formula Calculate the matching degree between each of the texts to be pushed and each business card category cluster; where A represents the business card category cluster, B represents the text to be pushed, cos(A,B) represents the matching degree, and s i t represents the number of target business cards corresponding to the i-th text keyword in the business card category cluster. i This represents the number of keywords corresponding to the i-th keyword in the text to be pushed, where n represents the total number of text keywords. w i This represents the weight corresponding to the i-th text keyword.

[0016] Furthermore, the formula Before calculating the matching degree between each of the texts to be pushed and each business card category cluster, the method further includes:

[0017] The text keywords are divided into multiple levels according to preset rules;

[0018] Set the weight of the lowest-ranking text keyword to w. c ;

[0019] According to the formula Assign weights to the remaining levels of the text keywords, where w c w represents the minimum weight. x R represents the weight of the xth level. x This represents the preset parameters for level x, where n is the default parameter. x This represents the number of text keywords corresponding to level x. Level x+1 is lower than level x. c represents the number of levels set.

[0020] Furthermore, the step of extracting multiple text keywords from the multiple texts to be pushed includes:

[0021] The multiple texts to be pushed are segmented into words to obtain several words of the multiple texts to be pushed;

[0022] The aforementioned word segments are each converted into corresponding word vectors;

[0023] Based on a pre-defined word database, extract the target word vector from the word vectors;

[0024] Obtain the preceding and following word vectors of each target word vector in the text to be pushed and concatenate them to obtain the comprehensive word vector of the target word vector;

[0025] The comprehensive word vector is input into a preset keyword judgment model to determine whether the target word vector is the text keyword; wherein, the keyword judgment model is generated by training a deep neural network model using the comprehensive word vector as input and the corresponding result of whether it is a text keyword as output.

[0026] Furthermore, the step of setting classification criteria based on multiple text keywords includes:

[0027] The multiple text keywords are converted into text vectors according to a preset vector transformation relationship;

[0028] Multiple classification functions are set using a preset linear classifier. Among them, b t =b t-1 +m t And b1 = m1, m t b represents the relevant constant corresponding to the text vector. t The bias is represented by t, which is a positive integer, w represents the preset weight vector, and f t (x) represents the t-th classification function, x represents the text vector, and W is a preset parameter;

[0029] Each of the aforementioned classification functions is used as a classification criterion.

[0030] Furthermore, before the step of classifying each preset business card based on the classification criterion to obtain multiple business card classification clusters, the method further includes:

[0031] The paper business cards are scanned by a preset scanning device to obtain the first business card information of each paper business card;

[0032] The unique identifier in the first business card information is compared with the unique identifier in the preset information database to obtain the corresponding second business card information in the preset information database;

[0033] By combining the information from the first business card and the information from the second business card, and performing deduplication, the preset business card is obtained.

[0034] Further, the step of selecting at least one matching business card category cluster for each text to be pushed based on the matching degree includes:

[0035] Determine whether the matching degree between each of the texts to be pushed and the business card category cluster is greater than the matching degree threshold;

[0036] The texts to be pushed with a matching degree greater than the matching degree threshold are assigned to the business card category cluster in the first round;

[0037] Determine whether there are any business card category clusters with target text to be pushed but no matching text after the first round of allocation;

[0038] If there are no matching business card category clusters for the target text to be pushed, then a preset number of business card category clusters are selected for a second round of allocation based on the matching degree between the target text to be pushed and each business card category cluster.

[0039] A text push device, the device comprising:

[0040] The text keyword extraction module is used to obtain multiple texts to be pushed and extract multiple text keywords from the multiple texts to be pushed;

[0041] The classification criterion setting module is used to set a classification criterion based on multiple text keywords;

[0042] The business card classification cluster acquisition module is used to classify each preset business card based on the classification criteria to obtain multiple business card classification clusters;

[0043] The matching degree calculation module is used to calculate the matching degree between each of the texts to be pushed and the business card category cluster;

[0044] Based on the matching degree, at least one matching business card category cluster is selected for each text to be pushed;

[0045] The text push module is used to push the text to be pushed to users of the corresponding business card category cluster based on the selection results.

[0046] An electronic device includes a memory and a processor, the memory storing a computer program that, when executed by the processor, causes the processor to perform the following steps:

[0047] Obtain multiple texts to be pushed, and extract multiple text keywords from the multiple texts to be pushed;

[0048] A classification criterion is set based on multiple text keywords;

[0049] Based on the aforementioned classification criteria, each preset business card is classified to obtain multiple business card classification clusters;

[0050] Calculate the matching degree between each of the texts to be pushed and the business card category cluster;

[0051] Based on the matching degree, at least one matching business card category cluster is selected for each text to be pushed;

[0052] Based on the selection results, the text to be pushed will be sent to the users of the corresponding business card category cluster.

[0053] A computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to perform the following steps:

[0054] Obtain multiple texts to be pushed, and extract multiple text keywords from the multiple texts to be pushed;

[0055] A classification criterion is set based on multiple text keywords;

[0056] Based on the aforementioned classification criteria, each preset business card is classified to obtain multiple business card classification clusters;

[0057] Calculate the matching degree between each of the texts to be pushed and the business card category cluster;

[0058] Based on the matching degree, at least one matching business card category cluster is selected for each text to be pushed;

[0059] Based on the selection results, the text to be pushed will be sent to the users of the corresponding business card category cluster.

[0060] The beneficial effects of this invention are as follows: By extracting multiple keywords from the text to be pushed, a deep understanding of the core theme of the content and the user's real needs is achieved. This breaks through the limitations of traditional methods that rely solely on static user profiles and simple keyword matching. Subsequent classification benchmarks and business card categorization allow user groups to be rationally and dynamically divided, ensuring highly targeted information pushes and avoiding interference from irrelevant information. Finally, by calculating the matching degree between the text to be pushed and the business card categorization clusters, the user's interests can be accurately located, achieving personalized and intelligent information pushes. This enables users to receive information that matches their preferences, thereby improving information reception rate and processing efficiency, effectively solving the problem of user information overload, and greatly improving the accuracy of information pushes and user satisfaction. Attached Figure Description

[0061] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0062] in:

[0063] Figure 1 This is a diagram illustrating the application environment of a text push method in one embodiment;

[0064] Figure 2 Here is a flowchart of a text push method in one embodiment;

[0065] Figure 3This is a structural block diagram of a text push device in one embodiment;

[0066] Figure 4 This is a structural block diagram of an electronic device in one embodiment. Detailed Implementation

[0067] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0068] Figure 1 This is a diagram illustrating a text push application environment in one embodiment. (Refer to...) Figure 1 This text push method is applied to a text push system. The text push system includes a terminal 110 and a server 120. The terminal 110 and server 120 are connected via a network. The terminal 110 can be a desktop terminal or a mobile terminal, and the mobile terminal can be at least one of a mobile phone, tablet, or laptop. The server 120 can be a standalone server or a server cluster consisting of multiple servers. The terminal 110 is used to acquire multiple texts to be pushed, and the server 120 is used to push the texts to be pushed to users in the corresponding business card category clusters.

[0069] like Figure 2 As shown, in one embodiment, a text push method is provided. This method can be applied to both terminals and servers; this embodiment uses terminal application as an example. The text push method specifically includes the following steps:

[0070] S1: Obtain multiple texts to be pushed, and extract multiple text keywords from the multiple texts to be pushed;

[0071] S2: Set classification criteria based on multiple text keywords;

[0072] S3: Classify each preset business card based on the classification criteria to obtain multiple business card classification clusters;

[0073] S4: Calculate the matching degree between each of the texts to be pushed and the business card category cluster;

[0074] S5: Select at least one matching business card category cluster for each text to be pushed, based on the matching degree;

[0075] S6: Based on the selection results, push the text to be pushed to the users of the corresponding business card category cluster.

[0076] As described in step S1 above, multiple texts to be pushed are acquired, and multiple text keywords are extracted from these texts. A series of texts to be pushed are collected; these texts can come from user-created content, including but not limited to social media posts, news articles, blog content, or messages sent by users. After acquiring these texts, natural language processing (NLP) techniques can be used for keyword extraction, such as term frequency-inverse document frequency (TF-IDF), named entity recognition (NER), or deep learning-based models. Through these techniques, core words or phrases in the text can be identified, and these keywords can accurately reflect the theme and important content of the text.

[0077] As described in step S2 above, classification benchmarks are established based on multiple text keywords. These benchmarks are typically set based on predefined topic tags, keyword groups, or other features. A classification framework can be constructed to organize the keywords, potentially using hierarchical or non-hierarchical classification methods. Through these benchmarks, the system can create a representative classification model for classifying the text to be pushed in subsequent steps.

[0078] As described in step S3 above, each preset business card is classified based on the classification criteria, resulting in multiple business card category clusters. According to the set classification criteria, the preset business cards are categorized. By referring to the previously set classification criteria, the system classifies the business cards according to their characteristics and attributes, forming multiple business card category clusters, where each cluster includes multiple preset business cards. These category clusters help the system identify different user groups, such as professional groups, interest groups, or social circles. By effectively classifying business cards, the system can more accurately match push content to relevant user groups, improving push effectiveness and user satisfaction.

[0079] As described in step S4 above, the matching degree between each text to be pushed and the business card category cluster is calculated. The matching degree between each text to be pushed and the already classified business card category cluster is calculated. The calculation method can be cosine similarity, Euclidean distance, or a deep similarity measurement method based on a machine learning model. In this way, the system can quantify the correlation, i.e., the matching degree, between the text to be pushed and different business card category clusters.

[0080] As described in step S5 above, at least one matching business card category cluster is selected for each text to be pushed, based on the matching degree. According to the obtained matching degree, the system will select at least one business card category cluster with the highest matching degree for each text to be pushed. The selection method can be to set some thresholds to ensure that only category clusters with matching degrees higher than a certain standard are selected. This ensures that the content pushed to users is more targeted, increasing their willingness to read or interact. In practical applications, multiple matching category clusters can also be selected for each text to be pushed, thus providing the possibility of multiple pushes for the same text, increasing exposure and user interaction opportunities.

[0081] As described in step S6 above, the text to be pushed is sent to users in the corresponding business card category cluster based on the selection results. The specific push method can be varied, including in-application notifications, emails, and SMS messages. Furthermore, the system can optimize the push strategy based on real-time user feedback and interaction behavior, thereby improving user stickiness and engagement.

[0082] In one embodiment, step S4, which calculates the matching degree between each of the texts to be pushed and the business card category cluster, includes:

[0083] S401: Obtain the text keywords and the corresponding number of keywords for each of the texts to be pushed;

[0084] S402: Set the corresponding dimension threshold for each text keyword;

[0085] S403: Detect the number of target business cards in each business card category cluster that is greater than or equal to the dimension threshold;

[0086] S404: According to the formula Calculate the matching degree between each of the texts to be pushed and each business card category cluster; where A represents the business card category cluster, B represents the text to be pushed, cos(A,B) represents the matching degree, and s i t represents the number of target business cards corresponding to the i-th text keyword in the business card category cluster. i This represents the number of keywords corresponding to the i-th keyword in the text to be pushed, where n represents the total number of text keywords. w i This represents the weight corresponding to the i-th text keyword.

[0087] As described in step S401 above, the text keywords and their corresponding counts for each text to be pushed are obtained. This requires collecting and processing the texts to be pushed, extracting the text keywords, and calculating the frequency of each keyword for subsequent matching calculations. Text keywords represent the core theme or important information in the text, and the number of keywords provides basic data for understanding the depth and breadth of the text content. This can be achieved through Natural Language Processing (NLP) techniques, such as word frequency statistics and TF-IDF calculation. The collected keywords and their corresponding counts provide basic data for subsequent classification and matching calculations, ensuring that the system can accurately identify the theme of the text.

[0088] As described in step S402 above, a corresponding dimension threshold is set for each text keyword. In this sub-step, the system sets a dimension threshold for each keyword based on the keyword characteristics of each text to be pushed. This dimension threshold is usually a standard set based on actual application needs and may be affected by various factors, such as keyword relevance, importance, and user needs. Setting an appropriate threshold can help the system filter out key keywords and avoid keywords that appear less frequently but have little impact on the overall content of the text. Keywords that reach or exceed the dimension threshold will be considered keywords with a high matching weight, while those that do not may be ignored. By defining dimension thresholds, the system can better focus on important information and topics, thereby improving the effectiveness and accuracy of matching degree calculation. In a specific embodiment, a correspondence table between text keywords and corresponding dimension thresholds can be pre-set. Then, based on the text keywords, the corresponding dimension threshold can be obtained by querying the correspondence table. It should be noted that the dimension threshold includes multiple dimensions and the threshold of that dimension. That is, different keywords correspond to at least one dimension and its threshold. If multiple keywords correspond to the same dimension and have multiple thresholds, the smaller threshold can be taken as the final threshold, or the threshold with the most text keywords can be selected as the final threshold.

[0089] As described in step S403 above, the number of target business cards in each business card category cluster that is greater than or equal to the dimension threshold is detected. Each business card category cluster is analyzed, and the number of target business cards within it that meet the set dimension threshold is detected. Target business cards refer to business cards related to the keywords in the text to be pushed. Only when the relevant keywords in the business cards reach or exceed the dimension threshold are these business cards considered valid matches. If the number of business cards in a certain business card category cluster is greater than or equal to the dimension threshold, the system marks it as a potential target category cluster, filtering out the business card category clusters most relevant to the text to be pushed, thereby ensuring that the pushed content is more relevant to the user.

[0090] As described in step S404 above, the matching degree between each text to be pushed and each business card category cluster is calculated according to the formula. The formula for calculating the matching degree between each text to be pushed and each business card category cluster takes into account various factors, such as the number of occurrences of keywords in the text, the number of target business cards, and the weight of the keywords. The closer the calculated matching degree is to 1, the more similar the text to be pushed is to that business card category cluster; the closer the calculated matching degree is to -1, the less similar the text to be pushed is to that business card category cluster. As for the formula... It uses a multiplication of quantity by weight to improve the weight calculation of each text keyword, and the formula... Using only simple weights is to reduce the influence of text keywords in the text to be pushed, and to prevent them from affecting the final calculation result due to too many text keywords in the text to be pushed.

[0091] In one embodiment, the formula Before step S404, which calculates the matching degree between each of the texts to be pushed and each business card category cluster, the method further includes:

[0092] S4041: Divide the text keywords into multiple levels according to preset rules;

[0093] S4042: Set the weight of the lowest-ranking text keyword to w. c ;

[0094] S4043: According to the formula Assign weights to the remaining levels of the text keywords, where w c w represents the minimum weight. x R represents the weight of the xth level. x This represents the preset parameters for level x, where n is the default parameter. x This represents the number of text keywords corresponding to level x. Level x+1 is lower than level x. c represents the number of levels set.

[0095] As described in step S4041 above, the text keywords are divided into multiple levels according to preset rules. This classification of text keywords based on predefined rules aims to categorize keywords into several levels according to factors such as importance, relevance, and frequency of occurrence. For example, a level system can be set up, including high, medium, and low levels, or more levels can be set, depending on actual needs. Keywords of different levels will play different roles in the matching degree calculation. High-level keywords refer to key topic words, which are crucial for text comprehension, while low-level keywords may contain supplementary information. Through this level classification, the system can give greater weight to more important keywords in subsequent matching degree calculations, thus more accurately reflecting the core content of the text and optimizing the relevance and effectiveness of the push notifications.

[0096] As described in step S4042 above, the weight of the lowest-level text keyword is set. In this sub-step, the system assigns a base weight to the text keywords classified as the lowest level. This base weight is usually set to a low value, such as 0.1 or 0.2, with the specific value depending on the system's design considerations and application scenarios. This weight setting is to clarify the status and influence of the lowest-level keywords in the overall matching degree calculation. The lowest-level keywords are usually less related to the main theme of the text and play a supplementary role; therefore, their weight is set relatively low, which helps guide the system to prioritize more important keywords during matching. This decision not only clarifies the keyword weight system structure but also lays the foundation for subsequent weight calculation and matching degree analysis. This hierarchical weight-based design aims to achieve more intelligent text push, enabling the system to make more accurate judgments when processing complex text content through different levels of weight settings.

[0097] As described in step S4043 above, the weights of the remaining levels of text keywords are set according to the formula. Weights are assigned to the remaining levels of text keywords based on the established ranking system. Generally, higher-level keywords have larger weights to reflect their importance and relevance within the text content. Only a minimum weight needs to be set for the lowest-level text keywords, and then the weights of the remaining levels are set sequentially according to the formula. It should be understood that R... x The values ​​can vary with the level, or they can all be the same parameter. The target weights set should satisfy R. x The weight is set to >0, thus assigning a weight to each text keyword. It's important to note that the weight should not be set too high to avoid inaccurate similarity calculations. In this way, the system can effectively reflect the core content of the text, thereby improving the relevance of content recommendations and user experience.

[0098] In one embodiment, step S1, which involves extracting multiple text keywords from the multiple texts to be pushed, includes:

[0099] S101: Perform word segmentation on the multiple texts to be pushed to obtain several word segments of the multiple texts to be pushed;

[0100] S102: Convert the aforementioned word segments into corresponding word vectors;

[0101] S103: Extract the target word vector from the word vector according to the preset word database;

[0102] S104: Obtain the word vectors before and after each target word vector in the text to be pushed and concatenate them to obtain the comprehensive word vector of the target word vector;

[0103] S105: Input the comprehensive word vector into a preset keyword judgment model to determine whether the target word vector is the text keyword; wherein, the keyword judgment model is generated by training a deep neural network model using the comprehensive word vector as input and the corresponding result of whether it is a text keyword as output.

[0104] As described in step S101 above, the multiple texts to be pushed are segmented into several words to obtain several word segments of the multiple texts to be pushed. The purpose of word segmentation is to divide the continuous character sequence of text into independent words or phrases for subsequent analysis and processing. In Chinese text processing, word segmentation is often more complex than in languages ​​such as English because Chinese does not have clear word boundaries. Therefore, it is usually necessary to use word segmentation tools (such as jieba, thulac, etc.) or algorithms (such as maximum matching, hidden Markov models, etc.). After this processing, the texts to be pushed will be converted into several word segments, which are the basis for subsequent keyword extraction and data analysis.

[0105] As described in step S102 above, the segmented words are converted into corresponding word vectors. The segmentation results are then converted into corresponding word vectors. Word vectors are a way to represent words as multi-dimensional real-valued vectors, allowing computers to better understand the relationships between word meanings. Typically, this process uses word embedding techniques such as Word2Vec, GloVe, and FastText. These models learn the position of each word in the semantic space based on contextual information from a large corpus, thereby generating its vector representation. Through word vectors, the similarity between words can be measured by the distance and direction of the vectors. For example, words in the same semantic domain will have relatively close word vectors. This step is crucial for subsequent keyword extraction because by constructing word vectors, the system can utilize the computer's mathematical capabilities to perform a deeper analysis of language, capturing the potential relationships and semantic information between words.

[0106] As described in step S103 above, target word vectors are extracted from the word vectors based on a pre-established word database. A pre-built and maintained word database is used to identify and extract target word vectors with specific meanings within the constructed word vectors. The word database may contain domain-specific terms, popular words, industry keywords, etc. The system identifies valuable target word vectors by matching word vectors with entries in the database. This process typically involves calculating the similarity between vectors, which can be done using distance metrics (such as cosine similarity) to determine whether the target word vector is similar to or matches entries in the database. Potential keywords are filtered from a large number of word segments to ensure that subsequent analysis focuses on more important components. An effective word database not only improves the accuracy of keyword extraction but also reduces the interference of irrelevant words on the results, improving the efficiency of information extraction.

[0107] As described in step S104 above, the word vectors preceding and following each target word vector in the text to be pushed are obtained and concatenated to obtain a comprehensive word vector of the target word vector. The contextual information of the extracted target word vectors in the text to be pushed is integrated. By obtaining the word vectors preceding and following the target word vector and concatenating them together, a comprehensive word vector is formed. This concatenation operation is to fully utilize contextual information and enhance the semantic representation of the target word. Context usually has a significant impact on the meaning of the target word; therefore, extracting and concatenating the word vectors preceding and following the target word vector from the context allows the comprehensive word vector to present more comprehensive semantic features. For example, if the target word is "bank," the words preceding and following it might be "something" and "transfer," and the comprehensive word vector will simultaneously reflect the meaning of "bank" in these contexts. This processing not only improves the quality of keyword extraction but also provides richer semantic information for subsequent keyword judgment.

[0108] As described in step S105 above, the comprehensive word vector is input into a preset keyword judgment model to determine whether the target word vector is a keyword in the text. The judgment model is typically generated by training a deep neural network (such as CNN, RNN, Transformer, etc.). During training, the comprehensive word vector is used as input, and the result of whether it is a keyword is output. This model can learn from a large amount of training data, extract potential features, and establish complex nonlinear mapping relationships to achieve accurate judgment. During prediction, the model outputs whether the target word vector is a valid keyword based on the features of the input comprehensive word vector. Specifically, the keyword judgment model can be obtained by training a pre-constructed first neural network model based on a preset word set. Each word in the preset word set includes a word vector and its corresponding word label. When training the pre-constructed neural network model, the word vector in each word data is used as the input to the neural network model, and the word label corresponding to the word vector in each word data is used as the output of the neural network model. Through training, the neural network model can learn the correspondence between all possible word vectors and word labels. The trained neural network model is then used as the keyword judgment model.

[0109] In one embodiment, step S2, which sets a classification criterion based on a plurality of text keywords, includes:

[0110] S201: Convert the multiple text keywords into text vectors according to the preset vector transformation relationship;

[0111] S202: Multiple classification functions are set using a preset linear classifier. Among them, b t =b t-1 +m t And b1 = m1, m t b represents the relevant constant corresponding to the text vector. t The bias is represented by t, which is a positive integer, w represents the preset weight vector, and f t (x) represents the t-th classification function, x represents the text vector, and W is a preset parameter;

[0112] S203: Use the classification functions described in each clause as the classification criterion.

[0113] As described in step S201 above, multiple text keywords are converted into text vectors according to a preset vector transformation relationship. The previously extracted text keywords are converted into text vectors based on the preset vector transformation relationship. This process utilizes a pre-trained word embedding model to map text keywords into a high-dimensional space, generating corresponding vector representations. The preset vector transformation relationship may be based on different word embedding algorithms (such as Word2Vec, GloVe, or FastText). These algorithms, through training on large amounts of corpus, capture the semantic and contextual relationships between words, thereby forming vectors with certain distributional characteristics. After conversion, each text keyword becomes numerically comparable, providing basic data for subsequent classification and processing. Text vectors typically have a high dimensionality and can contain rich semantic information; therefore, the representational power of vectors is crucial in subsequent analysis. Through this conversion, the system can not only understand the meaning of a single word but also process the complex semantics of phrases or multiple keywords, laying the foundation for setting the final classification benchmark.

[0114] As described in step S202 above, multiple classification functions are set using a preset linear classifier. This process involves constructing a mathematical model, typically using a linear decision boundary to classify text vectors. Each classification function can be represented by a linear equation. By continuously adjusting the weight vector and bias, the classifier learns how to classify different text vectors into different categories. The preset linear classifier is usually trained based on historical data and can effectively distinguish the data to be classified. By setting multiple classification functions, the system can achieve more refined classification, which is crucial for subsequent text matching and processing. Each classification function represents a classification criterion, laying the foundation for subsequent text classification and enabling the system to flexibly classify according to different needs.

[0115] As described in step S203 above, each of the classification functions is used as a classification benchmark. These benchmarks are then used to classify or match new text. Each classification function essentially defines a classification boundary, through which the system can determine which category a new text (input as a text vector) should be classified into. When the vector of the new text interacts with the classification benchmark and the corresponding importance metric is calculated, the system can quickly and accurately classify the text. Based on these classification functions, subsequent operations such as text analysis and keyword push can be performed more easily. For example, in information push applications, the classification benchmark guides the system to determine which text content should be pushed to the corresponding users, ensuring the relevance and effectiveness of the information.

[0116] In one embodiment, before step S3 of classifying each preset business card based on the classification criterion to obtain multiple business card classification clusters, the method further includes:

[0117] S211: Scan paper business cards using a preset scanning device to obtain the first business card information for each paper business card;

[0118] S212: Compare the unique identifier information in the first business card information with the unique identifier information in the preset information database to obtain the corresponding second business card information in the preset information database;

[0119] S213: Combine the information from the first business card and the information from the second business card, and perform deduplication to obtain the preset business card.

[0120] As described in step S211 above, paper business cards are scanned using a preset scanning device to obtain the first business card information for each card. The scanning device can be a high-performance OCR (Optical Character Recognition) scanner, capable of recognizing text, graphics, and other information on the paper business cards, such as names, phone numbers, email addresses, and company names. During the scanning process, the device captures all information on the business cards and parses it into machine-readable text format.

[0121] As described in step S212 above, the unique identifier information in the first business card information is compared with the unique identifier information in a preset information database to obtain the corresponding second business card information in the preset information database. The unique identifier information extracted from the "first business card information," such as name, email address, and phone number, is compared with the unique identifier information stored in the preset information database. The purpose of this comparison is to find the "second business card information" corresponding to the first business card information. Effective comparison of unique identifier information can utilize hash algorithms, similarity matching, or other similarity detection methods to ensure the identification and retrieval of the corresponding record already existing in the information database. Through this comparison, the system can reduce duplicate data, avoid information redundancy, and ensure the uniqueness and accuracy of each record.

[0122] As described in step S213 above, the first business card information and the second business card information are combined and deduplicated to obtain the preset business card. The previously obtained first business card information and the second business card information obtained through comparison are combined and deduplicated. The combination process includes analyzing the information content of both, identifying and merging identical or duplicate data fields, such as combining phone numbers from paper business cards with existing phone numbers in the database, and selecting appropriate information to retain. Deduplication not only eliminates inconsistencies in information but also ensures the accuracy and consistency of the final data. This typically involves using conditional judgments and data merging algorithms to remove duplicates, supplement missing fields, and standardize formats. During this process, the system may also introduce quality control standards, such as information priority, to determine which parts of information from different sources should be retained. Finally, the deduplicated and combined business card information forms the "preset business card," which not only provides clean and integrated information for subsequent classification processes but also provides users with a consistent digital business card solution, improving the efficiency of information management.

[0123] In one embodiment, step S5, which selects at least one matching business card category cluster for each text to be pushed based on the matching degree, includes:

[0124] S501: Determine whether the matching degree between each of the texts to be pushed and the business card category cluster is greater than the matching degree threshold;

[0125] S502: Perform a first round of allocation between the texts to be pushed that have a matching degree greater than the matching degree threshold and the business card category cluster;

[0126] S503: Determine whether there are any business card category clusters with target text to be pushed that do not match after the first round of allocation;

[0127] S504: If there are no matching business card category clusters for the target text to be pushed, then a preset number of business card category clusters are selected for a second round of allocation based on the matching degree between the target text to be pushed and each business card category cluster.

[0128] As described in step S501 above, it is determined whether the matching degree between each text to be pushed and the business card category cluster is greater than a matching degree threshold. It is necessary to determine the matching degree between each text to be pushed and the business card category cluster. The matching degree threshold is a key parameter, usually set based on historical data, average values, or business needs. It determines which texts to be pushed have a significant relationship with the business card category cluster; for example, it might be set to 0.6. By comparing the matching degrees, the system can identify which texts to be pushed have strong relevance and which may be less relevant or irrelevant information. This process ensures the accuracy of pushed content, helps avoid the transmission of invalid information, and improves user experience and satisfaction. If the matching degree is lower than the threshold, the system considers these texts to be pushed as not matching the existing business card category cluster and therefore does not assign them. This determination ensures that subsequent steps are more efficient, focusing only on texts with high matching degrees.

[0129] As described in step S502 above, the texts to be pushed with a matching degree greater than the matching degree threshold are allocated in the first round to the business card category clusters. This allocation process essentially pushes the filtered, highly matching texts to the relevant user groups. In this way, the system can effectively deliver valuable information to the users most likely to be interested. The matching degree used in the first round of allocation reflects the similarity between the text content and the user's business card information, promptly meeting the user's needs. In this round of allocation, the system may use specific algorithms and strategies, such as prioritizing the push of the text with the highest matching degree, to ensure the effectiveness and relevance of the pushed content. Through this process, the texts to be pushed can be connected with the correct users, promoting higher user engagement and enthusiasm.

[0130] As described in step S503 above, it is determined whether there are any target texts to be pushed that do not match any business card category clusters after the first round of allocation. An assessment is made to determine whether there are still some target texts to be pushed that fail to match any business card category clusters after the first round of allocation. The key to this determination is to confirm whether the first round of allocation has covered all pending texts to be pushed. If it is found that there are target texts to be pushed that do not match any business card category clusters, this may mean that although the matching degree of the text to be pushed is relatively high, it still does not reach the preset matching degree threshold, or the relevance between the text to be pushed and the existing business card category clusters is insufficient. Evaluating unmatched texts to be pushed is an important step in optimizing the push process, because these texts to be pushed may contain potentially important information and business opportunities.

[0131] As described in step S504 above, if there are no matching business card category clusters for the target text to be pushed, a preset number of business card category clusters are selected for a second round of allocation based on the matching degree between the target text to be pushed and each business card category cluster. If such texts to be pushed are found, the system will re-evaluate based on the specific matching degree between these target texts to be pushed and each business card category cluster to select a preset number of business card category clusters for a second round of allocation. This more flexible strategy, by including business card category clusters with lower matching degrees, enriches the breadth of matching. This approach ensures that each target text to be pushed has more matching opportunities, thereby improving the comprehensiveness and effectiveness of information delivery. Through this strategy, the system not only avoids the creation of information silos but also increases the likelihood of users encountering new information, especially in the context of rapid market changes and diversified user preferences, further meeting the needs of different users and improving the overall user experience.

[0132] Reference Figure 3 The present invention also provides a text push device, the device comprising:

[0133] The text keyword extraction module 902 is used to acquire multiple texts to be pushed and extract multiple text keywords from the multiple texts to be pushed;

[0134] The classification criterion setting module 904 is used to set a classification criterion based on multiple text keywords;

[0135] The business card classification cluster acquisition module 906 is used to classify each preset business card based on the classification criteria to obtain multiple business card classification clusters;

[0136] The matching degree calculation module 908 is used to calculate the matching degree between each of the texts to be pushed and the business card category cluster;

[0137] The business card category cluster selection module 910 is used to select at least one matching business card category cluster for each text to be pushed based on the matching degree.

[0138] The text push module 912 is used to push the text to be pushed to the user of the corresponding business card category cluster according to the selection result.

[0139] In one embodiment, the matching degree calculation module 908 includes:

[0140] The keyword count acquisition submodule is used to acquire the text keywords and corresponding keyword counts for each of the texts to be pushed.

[0141] The dimension threshold setting submodule is used to set the corresponding dimension threshold for each text keyword.

[0142] The target business card quantity detection submodule is used to detect the number of target business cards in each business card category cluster that is greater than or equal to the dimension threshold.

[0143] The matching degree calculation submodule is used to calculate the matching degree according to the formula. Calculate the matching degree between each of the texts to be pushed and each business card category cluster; where A represents the business card category cluster, B represents the text to be pushed, cos(A,B) represents the matching degree, and s i t represents the number of target business cards corresponding to the i-th text keyword in the business card category cluster. i This represents the number of keywords corresponding to the i-th keyword in the text to be pushed, where n represents the total number of text keywords.

[0144] w i This represents the weight corresponding to the i-th text keyword.

[0145] In one embodiment, the matching degree calculation module 908 includes:

[0146] The level division submodule is used to divide the text keywords into multiple levels according to preset rules;

[0147] The first weight setting submodule is used to set the weight of the lowest-ranking text keyword to w. c ;

[0148] The second weight setting submodule is used to set weights according to the formula. Assign weights to the remaining levels of the text keywords, where w c w represents the minimum weight. x R represents the weight of the xth level. x This represents the preset parameters for level x, where n is the default parameter. x This represents the number of text keywords corresponding to level x. Level x+1 is lower than level x. c represents the number of levels set.

[0149] In one embodiment, the text keyword extraction module 902 includes:

[0150] The word segmentation processing submodule is used to perform word segmentation processing on multiple texts to be pushed, and obtain several words of the multiple texts to be pushed;

[0151] The word vector conversion submodule is used to convert the several word segments into corresponding word vectors;

[0152] The target word vector extraction submodule is used to extract the target word vector from the word vectors based on a preset word database;

[0153] The concatenation submodule is used to obtain the word vectors before and after each target word vector in the text to be pushed and concatenate them to obtain the comprehensive word vector of the target word vectors;

[0154] The input submodule is used to input the comprehensive word vector into a preset keyword judgment model to determine whether the target word vector is the text keyword; wherein, the keyword judgment model is generated by training a deep neural network model using the comprehensive word vector as input and the corresponding result of whether it is a text keyword as output.

[0155] In one embodiment, the classification criterion setting module 904 includes:

[0156] The conversion submodule is used to convert multiple text keywords into text vectors according to a preset vector conversion relationship;

[0157] The classification function setting module is used to set multiple classification functions using a preset linear classifier. Among them, b t =b t-1 +m t And b1 = m1, m t b represents the relevant constant corresponding to the text vector. t The bias is represented by t, which is a positive integer, w represents the preset weight vector, and f t (x) represents the t-th classification function, x represents the text vector, and W is a preset parameter;

[0158] The classification benchmark acquisition submodule is used to take each of the classification functions as the classification benchmark.

[0159] In one embodiment, the text push device further includes:

[0160] The first business card information acquisition module is used to scan paper business cards through a preset scanning device to obtain the first business card information of each paper business card;

[0161] The second business card information acquisition module is used to compare the unique identifier information in the first business card information with the unique identifier information in the preset information database to obtain the corresponding second business card information in the preset information database.

[0162] The preset business card acquisition module is used to combine the information of the first business card and the information of the second business card, and perform deduplication to obtain the preset business card.

[0163] In one embodiment, the business card category selection module 910 includes:

[0164] The first judgment submodule is used to determine whether the matching degree between each of the texts to be pushed and the business card category cluster is greater than the matching degree threshold.

[0165] The first round of allocation submodule is used to perform a first round of allocation between the text to be pushed with a matching degree greater than the matching degree threshold and the business card category cluster;

[0166] The second judgment submodule is used to determine whether there are any business card category clusters with target text to be pushed that have not been matched after the first round of allocation;

[0167] The second-round allocation submodule is used to select a preset number of business card category clusters for the second round of allocation if there are no matching business card category clusters for the target text to be pushed.

[0168] Figure 4 An internal structural diagram of an electronic device in one embodiment is shown. This electronic device can specifically be a terminal or a server, and more specifically, a computer device. Figure 4 As shown, the electronic device includes a processor, a memory, and a network interface connected via a system bus. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and may also store a computer program. When executed by the processor, this computer program enables the processor to implement a text push method. The internal memory may also store a computer program, which, when executed by the processor, enables the processor to implement the text push method. Those skilled in the art will understand that... Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the electronic device to which the present application is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0169] In one embodiment, an electronic device is provided, including a memory and a processor, the memory storing a computer program that, when executed by the processor, causes the processor to perform the following steps:

[0170] Obtain multiple texts to be pushed, and extract multiple text keywords from the multiple texts to be pushed;

[0171] A classification criterion is set based on multiple text keywords;

[0172] Based on the aforementioned classification criteria, each preset business card is classified to obtain multiple business card classification clusters;

[0173] Calculate the matching degree between each of the texts to be pushed and the business card category cluster;

[0174] Based on the matching degree, at least one matching business card category cluster is selected for each text to be pushed;

[0175] Based on the selection results, the text to be pushed will be sent to the users of the corresponding business card category cluster.

[0176] By extracting multiple keywords from the text to be pushed, the system deeply understands the core theme of the content and the user's real needs, breaking through the limitations of traditional methods that rely solely on static user profiles and simple keyword matching. Subsequent categorization benchmarks and business card classification allow for the rational and dynamic segmentation of user groups, ensuring highly targeted information pushes and avoiding interference from irrelevant information. Finally, by calculating the matching degree between the text to be pushed and the business card category clusters, the system can accurately pinpoint user interests, achieving personalized and intelligent information pushes. Ultimately, users are increasingly able to receive information that matches their preferences, thereby improving information reception and processing efficiency, effectively solving the problem of user information overload, and greatly enhancing the accuracy of information pushes and user satisfaction.

[0177] In one embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, causes the processor to perform the following steps:

[0178] Obtain multiple texts to be pushed, and extract multiple text keywords from the multiple texts to be pushed;

[0179] A classification criterion is set based on multiple text keywords;

[0180] Based on the aforementioned classification criteria, each preset business card is classified to obtain multiple business card classification clusters;

[0181] Calculate the matching degree between each of the texts to be pushed and the business card category cluster;

[0182] Based on the matching degree, at least one matching business card category cluster is selected for each text to be pushed;

[0183] Based on the selection results, the text to be pushed will be sent to the users of the corresponding business card category cluster.

[0184] By extracting multiple keywords from the text to be pushed, the system deeply understands the core theme of the content and the user's real needs, breaking through the limitations of traditional methods that rely solely on static user profiles and simple keyword matching. Subsequent categorization benchmarks and business card classification allow for the rational and dynamic segmentation of user groups, ensuring highly targeted information pushes and avoiding interference from irrelevant information. Finally, by calculating the matching degree between the text to be pushed and the business card category clusters, the system can accurately pinpoint user interests, achieving personalized and intelligent information pushes. Ultimately, users are increasingly able to receive information that matches their preferences, thereby improving information reception and processing efficiency, effectively solving the problem of user information overload, and greatly enhancing the accuracy of information pushes and user satisfaction.

[0185] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.

[0186] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0187] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.

Claims

1. A text push method, characterized by, The method comprises: obtaining a plurality of to-be-pushed texts, and extracting a plurality of text keywords in the plurality of to-be-pushed texts; setting a classification benchmark according to the plurality of text keywords; classifying each preset business card based on the classification benchmark to obtain a plurality of business card classification clusters; calculating a matching degree of each to-be-pushed text and the business card classification cluster; selecting at least one matched business card classification cluster for each to-be-pushed text according to the matching degree; pushing the to-be-pushed text to a user of the corresponding business card classification cluster according to the selection result.

2. The text push method of claim 1, wherein, The step of calculating the matching degree of each to-be-pushed text and the business card classification cluster comprises: obtaining the text keywords and the corresponding keyword quantity of each to-be-pushed text; setting a corresponding dimension threshold value according to each text keyword; detecting a target business card quantity of each business card classification cluster greater than or equal to the dimension threshold value; According to the formula The matching degree of each said to be pushed text and each business card classification cluster is calculated; wherein A represents said business card classification cluster, B represents to be pushed text, cos(A, B) represents said matching degree, s i represents the target number of business cards corresponding to the i th text keyword in said business card classification cluster, t i represents the keyword number corresponding to the i th text keyword of said to be pushed text, n represents the total number of text keywords, w i represents the weight corresponding to the i th text keyword.

3. The text push method of claim 2, wherein, The formula is Before the step of calculating the matching degree of each of the to-be-pushed texts and each of the business card classification clusters, the method further comprises: dividing the text keywords into a plurality of levels according to a preset rule. The weight of the text keyword with the lowest setting level is set as w c ; According to the formula The weights of the text keywords of the remaining levels are set, wherein w c represents the lowest weight, w x represents the weight of the xth level, R x represents the preset parameter of the xth level, n x represents the number of text keywords corresponding to the xth level, the level of x+1 is lower than the xth level, and c represents the number of set levels.

4. The text push method of claim 1, wherein, The step of extracting a plurality of text keywords in the plurality of to-be-pushed texts comprises: performing word segmentation processing on the plurality of to-be-pushed texts to obtain a plurality of word segments of the plurality of to-be-pushed texts; converting the word segments into corresponding word vectors respectively; extracting a target word vector from the word vectors according to a preset word database; obtaining front and rear word vectors of each target word vector in the to-be-pushed text and splicing them to obtain a comprehensive word vector of the target word vector; inputting the comprehensive word vector into a preset keyword judgment model to obtain whether the target word vector is the text keyword; wherein the keyword judgment model is generated by training a deep neural network model with the comprehensive word vector as input and the corresponding result of whether it is a text keyword as output.

5. The text push method of claim 1, wherein, The step of setting a classification benchmark according to the plurality of text keywords comprises: converting the plurality of text keywords into text vectors according to a preset vector conversion relationship; Setting multiple classification functions using a preset linear classifier wherein b t = b t-1 + m t , and b1 = m1, m t represents a relevant constant corresponding to the text vector, b t represents a bias, t is a positive integer, w represents a preset weight vector, f t (x) represents the tthclassification function, x represents the text vector, and W is a preset parameter using each classification function as a classification benchmark.

6. The text push method of claim 1, wherein, Before the step of classifying each preset business card based on the classification benchmark to obtain a plurality of business card classification clusters, the method further comprises: scanning paper business cards through a preset scanning device to obtain first business card information of each paper business card; comparing unique landmark information in the first business card information with unique landmark information in a preset information library to obtain corresponding second business card information in the preset information library; integrating the first business card information and the second business card information and performing deduplication processing to obtain the preset business card.

7. The text push method of claim 1, wherein, The step of selecting at least one matched business card classification cluster for each to-be-pushed text according to the matching degree comprises: judging whether the matching degree of each to-be-pushed text and the business card classification cluster is greater than a matching degree threshold value; performing first round distribution on to-be-pushed texts and the business card classification cluster with a matching degree greater than the matching degree threshold value; judging whether there is a target to-be-pushed text without a matched business card classification cluster after the first round distribution; If there is no business card classification cluster matching the target text to be pushed, a preset number of business card classification clusters are selected according to the matching degrees of the target text to be pushed and each business card classification cluster for second round distribution.

8. A text push apparatus, characterized by comprising: The apparatus comprises: a text keyword extraction module configured to acquire a plurality of texts to be pushed and extract a plurality of text keywords from the plurality of texts to be pushed; a classification reference setting module configured to set a classification reference according to the plurality of text keywords; a business card classification cluster acquisition module configured to classify each preset business card based on the classification reference to obtain a plurality of business card classification clusters; a matching degree calculation module configured to calculate a matching degree of each text to be pushed and the business card classification cluster; a business card classification cluster selection module configured to select at least one matching business card classification cluster for each text to be pushed according to the matching degree; a text pushing module configured to push the text to be pushed to a user of a corresponding business card classification cluster according to the selection result.

9. A computer-readable storage medium, characterized in that, The computer program is stored in the memory and is executed by the processor to enable the processor to perform the steps of the text pushing method according to any one of claims 1 to 7.

10. An electronic device, comprising: The device comprises a memory and a processor, and the memory stores a computer program which is executed by the processor to enable the processor to perform the steps of the text pushing method according to any one of claims 1 to 7.