Service pushing method, system and equipment based on public opinion analysis and storage medium
By extracting corporate entities, events and time from public opinion data, building an event chain and predicting the operating status, the problems of incomplete information and poor pushing of financial institutions when expanding customer resources are solved, and high-accurate service data push is achieved.
Patent Information
- Application Number
- CN202510210050.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-13
AI Technical Summary
When financial institutions expand customer resources, the existing technology has problems such as low efficiency, incomplete information, and inability to update in real time. When actively pushing service project information to enterprises, it is difficult for enterprises to accurately locate the services they need, resulting in poor push results.
By obtaining public opinion data, extracting corporate entities, events and time, building an event chain, predicting the business status of the company, and pushing the description information of the corresponding service items to the company based on the corresponding relationship between the pre-set business status and the service items.
It realizes accurate prediction of the business status of the enterprise based on the event chain, improves the accuracy of service data push, and makes the pushed service items more in line with the actual needs of the enterprise.
Smart Images

Figure CN120146050A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data processing, and particularly relates to a service push method, system, device and storage medium based on public opinion analysis. Background Art
[0002] Currently, when financial institutions expand customer resources, they mainly rely on traditional means such as market research, customer referrals, and historical data. Although these methods are effective to a certain extent, they have problems such as low efficiency, incomplete information, and inability to be updated in real time.
[0003] Some active push methods often push a list of information of all applicable service items to the enterprise side according to the business operation type of the enterprise. This method is relatively extensive, and it is difficult for the enterprise side to accurately locate the services it needs from a large amount of information, so the push effect is not good. Summary of the Invention
[0004] In view of the above deficiencies of the prior art, the present invention provides a service push method, system, device and storage medium based on public opinion analysis to solve the above technical problems.
[0005] In a first aspect, the present invention provides a service push method based on public opinion analysis, including: Obtaining public opinion data; Extracting enterprise entities, events and time from the public opinion data, and arranging the events in chronological order as an event chain; Predicting the business status of the enterprise entity according to the event chain; According to the pre-set correspondence between the business status and service items, and the business status of the enterprise entity, pushing the description information of the corresponding service items to the enterprise entity.
[0006] In an optional embodiment, the method further includes: Generating a hash value for each piece of public opinion data, and removing duplicate data by verifying the uniqueness of the hash value; Correcting the dates and locations in the public opinion data; Using regular expressions to remove special characters in the public opinion data and normalizing the numerical values.
[0007] In an optional embodiment, extracting enterprise entities, events and time from the public opinion data, and arranging the events in chronological order as an event chain includes: Using a pre-trained named entity recognition model to extract enterprise entities, object entities and relationships from the public opinion data; Integrating the enterprise entity, object entity and relationship into an event; Arranging multiple events containing the same enterprise entity in chronological order as an event chain.
[0008] In an optional embodiment, the pre-trained named entity recognition model includes: A BERT layer for extracting dynamic context word vectors; A CRF layer for generating an entity label sequence according to the legal transition rules between entity labels; A fusion layer for using a binary classification network to determine whether there is a decorative relationship between entity pairs in the entity label sequence and the position encoding of the entities, where the decorative relationship is used to determine the attributes of the enterprise body; One or more fully connected layers for classifying the relationships between entities to obtain relationship categories.
[0009] In an optional embodiment, predicting the operating state of an enterprise entity according to the event chain includes: Performing credibility verification on the events of multiple event chains of multiple enterprise entities, clearing the events that do not pass the verification, and obtaining credible event chains; Converting the credible event chains of the enterprise entity into a vector sequence, with each event corresponding to a vector; Inputting the vector sequence into a pre-trained long short-term memory neural network model to obtain the operating state of the enterprise entity.
[0010] In an optional embodiment, pushing the description information of the corresponding service item to the enterprise entity according to the pre-set correspondence between the operating state and the service item and the operating state of the enterprise entity includes: Establishing a correspondence between the operating state and the service item according to the service requirements corresponding to the operating state; Screening out associated service items for the corresponding enterprise entity according to the operating state and the correspondence; Obtaining the attributes of the enterprise entity, matching the attributes with the object restriction information of the associated service items, and outputting the associated service items with successful matching as high-quality service items; Pushing the description information of the high-quality service items to the email of the enterprise entity.
[0011] In an optional embodiment, the method further includes: Tracking the feedback behavior of the enterprise entity, where the feedback behavior includes adopting the service item, adopting other service items, or not adopting any service items; Fine-tuning the correspondence according to the feedback behavior.
[0012] In a second aspect, the present invention provides a service push system based on public opinion analysis, including: An acquisition module for acquiring public opinion data; A processing module, configured to extract enterprise entities, events, and time from the public opinion data, and arrange the events in chronological order into an event chain; A prediction module, configured to predict the business status of the enterprise entity according to the event chain; A push module, configured to push the description information of the corresponding service item to the enterprise entity according to the correspondence between the preset business status and the service item and the business status of the enterprise entity.
[0013] In a third aspect, a device is provided, including: A memory, configured to store a service push program based on public opinion analysis; A processor, configured to implement the steps of the service push method based on public opinion analysis provided in the first aspect when executing the service push program based on public opinion analysis.
[0014] In a fourth aspect, a computer-readable storage medium is provided, on which a service push program based on public opinion analysis is stored. When the service push program based on public opinion analysis is executed by a processor, the steps of the service push method based on public opinion analysis provided in the first aspect are implemented.
[0015] The beneficial effects of the present invention are as follows. The service push method, system, device, and storage medium based on public opinion analysis provided by the present invention extract enterprise entities, events, and time from public opinion data, and then arrange the events in chronological order into an event chain. This event chain not only represents the state characteristics of the enterprise entity, but also can represent the chronological characteristics of state changes. In this way, the business status of the enterprise entity can be accurately predicted based on the event chain, and then the description information of the adapted service item can be pushed to it specifically, greatly improving the accuracy of service data push.
[0016] In addition, the design principle of the present invention is reliable, the structure is simple, and it has a very wide application prospect. Description of the Drawings
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can also be obtained according to these drawings without creative efforts.
[0018] Figure 1 is a schematic flowchart of the method in an embodiment of the present invention.
[0019] Figure 2 is a schematic flowchart of the public opinion data processing of the method in an embodiment of the present invention.
[0020] Figure 3It is a schematic block diagram of a system according to an embodiment of the present invention.
[0021] Figure 4 It is a schematic structural diagram of a device provided by an embodiment of the present invention. Detailed implementation manners
[0022] In order to enable those skilled in the art of the present technology to better understand the technical solutions in the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.
[0023] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field of the present invention. The terms used in the description of the present invention herein are only for the purpose of describing specific embodiments, and are not intended to limit the present invention.
[0024] The following explains the key terms that appear in the present invention.
[0025] BERT (Bidirectional Encoder Representations from Transformers) is a pre-trained language representation model developed by Google.
[0026] Core architecture: Based on the encoder part of Transformer, composed of multiple stacked identical layers, each layer contains a multi-head self-attention mechanism and a feed-forward neural network.
[0027] Input representation: Fusing word embeddings, paragraph embeddings, and position embeddings, splitting the text into sub-word units, and also introducing special tokens such as [CLS], [SEP], [MASK].
[0028] Pre-training tasks: Masked Language Model (MLM): Randomly mask 15% of the tokens in the input text, and let the model predict the masked words according to the context to achieve bidirectional semantic understanding.
[0029] Next Sentence Prediction (NSP): Receive a pair of sentences and predict whether the second sentence is the next sentence of the first sentence in the original text to help the model learn the logical relationship between sentences.
[0030] Technical advantages: Bidirectional Context Understanding: Different from unidirectional language models, it can consider the context information before and after a word simultaneously, better capturing semantic relationships.
[0031] Pre-training and Fine-tuning Framework: First, pre-train on large-scale unsupervised text data to learn general language representations, and then fine-tune for specific tasks, reducing training costs and improving efficiency.
[0032] Deep Semantic Understanding: It can understand the grammar and semantic relationships of long texts and performs well in modeling sentence relationships.
[0033] Application Scenarios: Many natural language processing tasks such as text classification, named entity recognition, question answering systems, machine translation, sentiment analysis, etc.
[0034] CRF, namely Conditional Random Field, is an undirected graph model.
[0035] Principle: Used to model problems such as sequence labeling and classification, considering the dependencies between input data, and realizing the modeling and prediction of sequence data by learning the conditional probability distribution between features, which can be expressed as the conditional probability distribution of the labeling sequence Y given the observation sequence X.
[0036] Application Scenarios: Named Entity Recognition (NER): Using sequence labeling combined with context information and feature functions to capture entity relationships and improve recognition accuracy.
[0037] Part-of-Speech Tagging (POS tagging): Using features such as vocabulary and syntax to learn the relationships between part-of-speech in the context and achieve accurate tagging.
[0038] Syntactic Analysis: Modeling and annotating the sentence structure to help understand the grammar structure and meaning of the sentence.
[0039] Information Extraction: Extracting information such as entity relationships and events in the text, and combining context features and constraints to achieve accurate information extraction and relationship recognition.
[0040] The service push method based on public opinion analysis provided by the embodiments of the present invention is executed by a computer device. Correspondingly, the service push system based on public opinion analysis runs in the computer device.
[0041] Figure 1 It is a schematic flowchart of the method of an embodiment of the present invention. Among them, Figure 1 The execution subject can be a service push system based on public opinion analysis. According to different requirements, the order of the steps in this flowchart can be changed, and some can be omitted.
[0042] Such as Figure 1As shown in the figure, the method includes: S1. Obtain public opinion data.
[0043] Collect publicly available news public opinion data through web crawler technology, search engine collection, and API interfaces of social media.
[0044] S2. Extract enterprise entities, events, and time from the public opinion data, and arrange the events in chronological order into an event chain.
[0045] Apply the named entity recognition (NER) technology in natural language processing, using a deep learning-based model such as BERT+LSTM+CRF. First, preprocess the text data, including word segmentation, stop word removal, etc. Then, input the processed text into the pre-trained BERT model to obtain the semantic representation of the text. Next, model the sequence information through the LSTM network to capture context features. Finally, use the CRF layer to learn the dependencies between labels, so as to accurately identify enterprise entities in the text.
[0046] Adopt a method combining rules and machine learning. First, define a series of event trigger words and templates, and initially identify events through pattern matching. Then, use classification algorithms such as support vector machine (SVM) to classify and refine the initially identified events. For example, for major events of enterprises, such as acquisitions, listings, etc., construct corresponding feature vectors and train the SVM model to improve the accuracy of event extraction.
[0047] Use a specialized time expression recognition tool, such as the time parser in Stanford CoreNLP. This tool can recognize various time formats, including absolute time (such as "February 24, 2025") and relative time (such as "yesterday", "next week"). Normalize the recognized time into a unified time format for subsequent chronological arrangement of events.
[0048] Associate the extracted events with the corresponding time and enterprise entities, sort the events in chronological order, and construct an event chain. Use a database (such as MySQL) to store the event chain information for convenient subsequent query and analysis.
[0049] S3. Predict the operating status of enterprise entities based on the event chain.
[0050] Extract various features from the event chain, including the type, frequency, sentiment tendency, etc. of the events. For example, count the positive events (such as winning awards, launching new products) and negative events (such as legal disputes, executive departures) of the enterprise separately, and calculate the proportion of positive events and negative events as an important feature.
[0051] Adopt models such as decision trees and random forests in machine learning, or recurrent neural networks (RNN) and their variants (such as LSTM, GRU) in deep learning. Taking the random forest as an example, the extracted features are used as the input, and the operating status of the enterprise (such as good, average, poor) is used as the label to train the random forest model. Through the prediction of the model, the future operating status of the enterprise is obtained.
[0052] S4. According to the corresponding relationship between the preset operating status and service items, and the operating status of the enterprise entity, push the description information of the corresponding service items to the enterprise entity.
[0053] Create a mapping table in the database to store the corresponding relationship between different operating statuses and service items. For example, when the operating status of the enterprise is "poor", the corresponding service items may be financial consulting, crisis public relations, etc.; when the operating status is "good", the corresponding service items may be market expansion consulting, brand building services, etc.
[0054] Adopt a message queue (such as Kafka) and push notification technology (such as SMS, email). After predicting the operating status of the enterprise, obtain the description information of the corresponding service items according to the mapping table, and send the information to the Kafka message queue. Then, through the SMS gateway or email server, push the description information of the service items to the enterprise entity.
[0055] In an embodiment of the present invention, based on step S1, a possible embodiment will be given below to non-restrictively elaborate on its specific implementation.
[0056] Write a crawler program using Python's Scrapy framework. First, construct a list of URLs according to the structure of the target news website to determine the page range to be crawled. For example, for common news websites, analyze the URL rules of their news list pages and detail pages, and use regular expressions and other methods to generate complete URLs. During the crawling process, set a reasonable crawling interval time to avoid putting too much pressure on the target website and prevent the IP from being blocked by the website. Use Scrapy middleware to process the request header information and simulate real browser access to improve the crawling success rate. For some websites with strong anti-crawler mechanisms, adopt the IP proxy pool technology to dynamically change the IP address for crawling.
[0057] Call the APIs of search engines such as Baidu and Google. For example, use the interface of Baidu Search Open Platform. When making the call, construct the query keywords reasonably, combining industry terms, company names, etc., to ensure that relevant public opinion news can be obtained. At the same time, utilize the advanced search syntax of the search engine, such as limiting the time range, file type, etc., to accurately screen the data. To improve the collection efficiency, perform pagination on the search results and batch obtain the data.
[0058] Preprocess the obtained public opinion data. The specific preprocessing methods include: (1) Generate a hash value for each piece of public opinion data, and remove duplicate data by verifying the uniqueness of the hash value.
[0059] Select a mature hash algorithm, such as SHA - 256 (256 - bit version of the Secure Hash Algorithm). In Python, it can be implemented with the help of the hashlib library. First, convert each piece of public opinion data into a byte stream form to ensure the data format is unified. For example, if the public opinion data is in text form, it needs to be encoded first, such as data.encode('utf - 8'). Then, call the hashlib.sha256() method, passing in the encoded data to generate the corresponding hash value.
[0060] Maintain a set of hash values. For each generated hash value of the data, check whether the hash value already exists in the set. In Python, the set data structure can be used to efficiently perform member checks. If the hash value already exists, it means the data is duplicate and can be directly discarded; if the hash value does not exist, add it to the set and retain the corresponding public opinion data.
[0061] (2) Correct the dates and locations in the public opinion data.
[0062] Utilize a dedicated date parsing library, such as dateutil. First, use the dateutil.parser.parse() method to attempt to parse the date string in the public opinion data. If the parsing fails, it indicates that there may be a problem with the date format. For example, for ambiguous date expressions, such as "February 32, 2025", custom correction logic can be written. By judging the reasonable range of the month and date, make corrections. For common date format errors, such as the year lacking a leading zero, it can also be completed. If the date parsing is successful but is significantly different from the actual situation (such as outside the reasonable time range), it can be judged and corrected in combination with the context information.
[0063] With the help of an address resolution library, such as geopy. First, standardize the location names in the public opinion data by removing extra spaces, special symbols, etc. Then, use a geocoder in geopy.geocoders, such as Nominatim, to resolve the locations. If the resolution result is empty or does not match the expectation, it is possible to perform fuzzy matching and correction by querying a place name database (such as a publicly available administrative division database). For example, for some abbreviations and aliases, they can be converted to standard place names through a mapping table.
[0064] (3) Use regular expressions to remove special characters in the public opinion data and normalize the numerical values.
[0065] Use the re module in Python to write regular expressions to match and remove special characters. For example, to remove all characters other than letters, numbers, and Chinese characters, the regular expression [^\w\u4e00-\u9fff]+ can be used.
[0066] For numerical values in the public opinion data, such as the number of likes and comments, first extract them. If the numerical types are inconsistent, uniformly convert them to floating-point types. Then, according to the data distribution, select a suitable normalization method. If the data distribution is relatively uniform, the min-max normalization method can be used.
[0067] In an embodiment of the present invention, based on step S2, a possible embodiment will be given below to non-restrictively elaborate on its specific implementation.
[0068] S201. Use a pre-trained named entity recognition model to extract enterprise entities, object entities, and relationships from the public opinion data.
[0069] Among them, the named entity recognition model includes: The BERT layer is used to extract dynamic context word vectors; The CRF layer is used to generate an entity label sequence according to the legal transition rules between entity labels; The fusion layer is used to determine whether there is a decorative relationship between entity pairs according to the candidate entity pairs and the position encoding of the entities in the entity label sequence by using a binary classification network, and the decorative relationship is used to determine the attributes of the enterprise body; One or more fully connected layers, and the fully connected layer is used to classify the relationships between entities to obtain relationship categories.
[0070] The data processing flow of this model can be referred to Figure 2 , including: 1. Dynamic context encoding (BERT layer) Input: Original news text (e.g., "CATL announced a cooperation with Tesla to expand the Berlin Gigafactory and accelerate its new energy layout in Europe.")
[0071] Processing: Generate context-sensitive semantic vectors through an improved BERT model (MuLER).
[0072] Utilize the attention mechanism (weight > 0.9) to capture the interaction between entities (such as the cooperation intention between "CATL - Tesla").
[0073] Reinforce the association between the action ("expand") and the object ("Berlin Gigafactory") through positional encoding.
[0074] Output: Enhanced context word vectors (including attention weights and positional encoding information).
[0075] Interaction features between entities (such as high-weight signals of cooperation intention).
[0076] The specific data processing principle is as follows: Improved attention mechanism: On the basis of the multi-head attention of the standard BERT, introduce Intent-Aware Attention. Dynamically weight the interaction between entities (such as "CATL - Tesla") (attention weight > 0.9), and filter out irrelevant contexts through a gating mechanism. For example, the word "cooperation" in the sentence will be given a high weight, while "announced" may be down-weighted.
[0077] Positional encoding enhancement: Perform relative positional encoding on the positional relationship between the action verb (such as "expand") and its direct object (such as "Berlin Gigafactory"). Formula: Map the positional difference (Δpos) through a sine function to strengthen the association between the action and the object.
[0078] Semantic vector generation: Output the context vector of each Token (such as h_i ∈ R^d), which contains global semantic information (from the Transformer layer of BERT) and local interaction features (from attention weights and positional encoding).
[0079] Compared with the original BERT, dynamically adjust the semantic focus range through attention weights to avoid noise interference from long-distance dependencies. Specifically capture the cooperation intention between enterprises through high-weight attention heads (such as Head-5), and low-weight heads handle conventional semantics.
[0080] 2. CRF Entity Label Generation Input: Context word vectors output by the BERT layer.
[0081] Processing: Generate entity labels based on conditional random field (CRF) decoding and combined with the BIOES annotation system.
[0082] Example annotation: CATL → B-Enterprise; Expansion → B-Action; Berlin Gigafactory → B-Object.
[0083] Output: Structured entity list (including text, type, location): {"text": "CATL", "type": "Enterprise", "span": (0,4)}; {"text": "Expansion", "type": "Action", "span": (10,12)}; {"text": "Berlin Gigafactory", "type": "Object", "span": (13,19)}.
[0084] Its data processing principle is as follows: Conditional random field (CRF) modeling: Input: Token vector sequence [h 1 , h 2 ,..., h n output by BERT.
[0085] Emission Matrix: The scores of each Token on different labels (B / I / O / E / S) are calculated by a linear layer.
[0086] Transition Matrix: The transition probabilities between labels (e.g., the score from B-Enterprise to I-Enterprise is higher, and the score from B-Enterprise to B-Action is lower).
[0087] BIOES annotation system: Five types of labels: B (Begin), I (Inside), O (Outside), E (End), S (Single).
[0088] For example: "CATL" is annotated as B-Enterprise → I-Enterprise → E-Enterprise (assuming character splitting).
[0089] Viterbi decoding: Select the globally optimal label sequence through dynamic programming to ensure legal label transitions (e.g., avoid directly following B-Enterprise with B-Action).
[0090] CRF explicitly models the dependencies between tags (such as the continuity of enterprise names) through the transition matrix. Compared with direct classification by Softmax, CRF can correct local misjudgments (such as misclassification of a single Token).
[0091] 3. Decorative Relationship Fusion Layer Input: Vector concatenation of entity pairs (such as the vectors of "CATL" and "Berlin Gigafactory").
[0092] Relative position embedding (indicating the distance between entities in a sentence).
[0093] Processing: Binary classification to determine whether there is a relationship between entity pairs (such as probability P = 0.98).
[0094] If there is a relationship, trigger attribute attachment (such as adding the attribute "leading power battery company" to "CATL").
[0095] Output: Relationship existence marker (yes / no).
[0096] Entity attribute attachment information: {"entity": "CATL", "attribute": "leading power battery company"}.
[0097] Its data processing principle is: Entity pair feature concatenation: Take the average or concatenate the vectors of candidate entity pairs (such as "CATL" and "Berlin Gigafactory"): For example, h_CATL ⊕ h_Berlin Gigafactory ∈ R^{2d}.
[0098] Add relative position embedding (Relative Position Embedding): Calculate the position difference between the two entities in the sentence (Δpos = j - i) and map it to a low-dimensional vector.
[0099] Binary classification logic: Input: Concatenated feature vector [h_pair; pos_emb].
[0100] Output the probability of relationship existence through a fully connected layer + Softmax: Formula: P = σ(W · [h_pair; pos_emb] + b), where σ is the Sigmoid function.
[0101] If P > 0.5 (P = 0.98 in the example), it is determined that there is a relationship.
[0102] Attribute attachment: If the relationship exists, add attributes according to predefined rules or an external knowledge base (such as an enterprise encyclopedia): For example, when "CATL" is detected as an entity, automatically append the attribute of the leading power battery company (from knowledge graph matching).
[0103] Quickly filter out irrelevant entity pairs through binary classification to reduce subsequent computational complexity. Integrate external knowledge to enhance the context understanding of relationship classification.
[0104] 4. Fully Connected Relationship Classifier Input: The features of entity pairs with confirmed relationships (such as the vector concatenation of "CATL - Tesla").
[0105] The output of the decorative relationship fusion layer (such as attribute information).
[0106] Processing: Classify according to predefined relationship types (cooperation, expansion, etc.).
[0107] Use a fully connected network to calculate the confidence level (such as 96.3%).
[0108] Output: The specific relationship type and confidence level between entities: {"subject": "CATL", "relation": "cooperation", "object": "Tesla", "confidence": 0.963}.
[0109] Its data processing principle includes: Feature Engineering: Input: The features of entity pairs with confirmed relationships + decorative attributes (such as h_CATL ⊕ h_Tesla ⊕ h_leading power battery company).
[0110] Syntactic features may be added (such as dependency paths, verb phrases between entities).
[0111] Multi - classification Logic: Fully connected network structure: Input layer → Hidden layer (ReLU) → Output layer (Softmax).
[0112] The output dimension corresponds to predefined relationship types (such as 5 types: cooperation, acquisition, expansion, etc.).
[0113] Loss function: Cross - Entropy Loss.
[0114] Confidence Calculation: Determine the final relationship type through Softmax probability values (such as "cooperation" corresponding to a probability of 0.963).
[0115] Formula: P(relation=k)=exp(z_k) / Σ(exp(z_i)), where z_i are the output layer logits.
[0116] At the same time, semantic vectors, attributes, and syntactic features are used to improve the classification accuracy. Analyze the certainty of relation determination through a confidence analysis model (such as 96.3% vs. marginal case 70%).
[0117] In this named entity recognition model, the attention weights of BERT directly affect the quality of the emission matrix of CRF. For example, Token vectors in high attention weight regions (such as "cooperation") are more likely to be correctly labeled as B-action by CRF. The entity positions (spans) output by CRF are used to extract the vectors and relative position embeddings of entity pairs. If CRF misses labeling "Tesla", subsequent relation classification will surely fail. Decorative attributes (such as "leader in power batteries") provide additional semantic clues for the classifier through feature splicing. For example, enterprises with this attribute are more likely to trigger "cooperation" rather than "acquisition". The parameters of each layer can be fine-tuned through joint training (JointTraining). For example: weighted sum of the label loss of CRF and the cross-entropy loss of the relation classifier, and backpropagation to update the underlying parameters of BERT-MuLER.
[0118] S202. Integrate enterprise entities, object entities, and relations into events.
[0119] The relation can be a competitive relation, an acquisition relation, a cooperation relation, etc. Integrating enterprise entities, object entities, and relations into events means forming a triple with enterprise entities, object entities, and relations, and this triple represents an event.
[0120] In a specific example, it includes the following steps: In memory, use Python's dictionary to temporarily store the extracted enterprise entities, object entities, and relations. For example, represent each entity as a dictionary object containing attributes such as entity name and type; relations are also represented as a dictionary containing relation type, relevant descriptions, etc. In terms of the database, it is more appropriate to choose a graph database (such as Neo4j) because a graph database can well represent the complex associations between entities and relations. Store enterprise entities and object entities as nodes and relations as edges in the graph database, which is convenient for subsequent querying and analysis.
[0121] Traverse the extracted enterprise entities, object entities, and relationship data. For each set of matching data, construct a triple. For example, in Python, the tuple data structure can be used to represent a triple, such as (enterprise_entity, object_entity, relationship). During the construction process, data validation is required to ensure that the enterprise entity and object entity are valid and uniquely identified, and the relationship is also clear and conforms to predefined types (competitive relationship, acquisition relationship, cooperation relationship, etc.).
[0122] When encountering entities or relationships that cannot be clearly matched, record the relevant data in the error log for subsequent manual review. For example, if the extracted relationship is ambiguous and it is impossible to accurately determine whether it is a competitive relationship or a cooperation relationship, record the relevant data and mark the construction of this triple as failed.
[0123] S203. Arrange multiple events containing the same enterprise entity in chronological order as an event chain.
[0124] Extract time information from the public opinion data related to each event. Use a dedicated time parsing tool, such as the parser module in the dateutil library. For example, for the time string "March 15, 2025", it can be parsed into a Python datetime object through dateutil.parser.parse("March 15, 2025"). Normalize all different time representations into datetime objects for subsequent comparison and sorting.
[0125] Maintain a dictionary with the enterprise entity as the key and the value as a list of events related to that enterprise. Traverse all the constructed event triples and add the events to the corresponding list according to the enterprise entity. For example, in Python: event_chain_dict = {}; for event in all_events: enterprise = event[0]; if enterprise not in event_chain_dict: event_chain_dict[enterprise] = []; event_chain_dict[enterprise].append(event).
[0126] Then, sort the event list corresponding to each enterprise entity. Use Python's sorted function, combined with a comparison function for the time information in the events, and sort them in chronological order. For example: for enterprise, events in event_chain_dict.items(): event_chain_dict[enterprise] = sorted(events, key=lambda x: x[3]) # Assume the time information is in the fourth position of the triple.
[0127] Store the constructed event chains in a database, such as a relational database (MySQL, PostgreSQL). To improve query efficiency, create an index on the enterprise entity field so that when querying the event chain of a certain enterprise, relevant data can be quickly located. At the same time, optimize the database regularly, clean up useless data, and ensure the efficiency of data storage and query.
[0128] In an embodiment of the present invention, based on step S3, the following will give a possible embodiment to non-restrictively elaborate on its specific implementation.
[0129] S301. Perform credibility verification on the events of multiple event chains of multiple enterprise entities, clear the events that fail the verification, and obtain credible event chains.
[0130] Establish a media credibility database to record the credibility scores of major news media and social media platforms. For example, give a higher credibility score to authoritative official news media, and a lower score to some unproven gossip platforms. When an event comes from a low-credibility platform, conduct key reviews.
[0131] For the same event, if the same or similar information can be obtained from multiple independent high-credibility sources, the credibility of the event increases. For example, an enterprise acquisition event reported simultaneously on multiple mainstream news media has higher credibility than a message published only on a small niche platform.
[0132] Check whether the internal logical relationship of the event is reasonable. For example, in an enterprise cooperation event, whether there is complementarity in the business fields of the two cooperating parties, and whether the expected goals of the cooperation conform to market logic. If a logical contradiction is found, such as a traditional manufacturing enterprise and an Internet finance enterprise having an unrelated technical cooperation, the credibility of the event is in doubt.
[0133] Write a program to perform a preliminary automated screening of events according to preset verification rules. For example, use the pandas library in Python to read event chain data, and quickly screen out events that clearly do not meet the rules through data matching and logical judgment.
[0134] For events that are still in doubt after automated screening, organize professionals for manual review. The reviewers comprehensively judge the authenticity and credibility of the events by referring to relevant materials and industry knowledge. For example, for events involving complex technical fields, invite experts in the relevant fields to participate in the review.
[0135] S302. Convert the trusted event chain of the enterprise entity into a vector sequence, with each event corresponding to a vector.
[0136] Extract the key attributes of the enterprise entity and object entity in the event, such as enterprise scale (number of employees, total assets), industry category, market share, etc., and use these numericalized attributes as part of the vector.
[0137] Quantify the relationships in the event, such as competitive relationships, acquisition relationships, cooperation relationships, etc. For example, use one - hot encoding to convert different relationship types into corresponding binary vectors.
[0138] Convert the time when the event occurs into a timestamp and perform normalization processing to make it match other features in the numerical range, as the time - dimension feature of the vector.
[0139] Concatenate the extracted feature vectors of various types in a certain order to form a complete event vector. For example, in Python, use the numpy library to concatenate the entity feature vector, relationship feature vector, and time feature vector into a one - dimensional array as the vector corresponding to the event. Finally, form a vector sequence by arranging the vectors corresponding to all trusted events of an enterprise entity in the order of event occurrence.
[0140] S303. Input the vector sequence into a pre - trained long short - term memory neural network model to obtain the operating state of the enterprise entity.
[0141] Select a suitable deep - learning framework, such as TensorFlow or PyTorch, and load the pre - trained long short - term memory neural network (LSTM) model. For example, in PyTorch, define a model structure containing an LSTM layer and a fully connected layer, and load the pre - trained model parameters.
[0142] Batch process the vector sequences to meet the model input requirements. For example, form a batch of vector sequences of multiple enterprise entities, where each batch contains a fixed number of vector sequences. At the same time, pad or truncate the vector sequences within the batch to make their lengths consistent. When padding or truncating, pay attention to retaining the chronological information of the events.
[0143] Input the processed vector sequences into the LSTM model for prediction. The model outputs a numerical value or category representing the business operation status of the enterprise. For example, it outputs a numerical value between 0 and 1, representing the probability of good business operation status of the enterprise; or it outputs specific categories such as "start-up stage", "growth stage", "maturity stage", "special operation status". Interpret and analyze the model output results, and combine with the actual business situation to provide a valuable business operation status evaluation report for the enterprise.
[0144] In an embodiment of the present invention, based on step S4, a possible embodiment will be given below to non-restrictively elaborate on its specific implementation scheme.
[0145] S401. Establish the corresponding relationship between the business operation status and service items according to the service requirements corresponding to the business operation status.
[0146] When formulating service items, the present application has considered the targeted service groups. For example, study industry authoritative reports and successful enterprise cases, and analyze the service solutions adopted by enterprises in different business operation states. For example, refer to how excellent enterprises in the same industry achieve turning losses into profits by means of professional financial consulting services when facing financial difficulties, so as to determine the corresponding service items.
[0147] According to the collected and analyzed data, formulate the corresponding rules between the business operation status and service items. For example, if the business operation status of an enterprise is "difficult capital turnover", the corresponding service items may include "short-term financing service", "financial consulting and cost optimization service"; if the operation status is "rapid business expansion", it corresponds to "foreign currency conversion", etc.
[0148] Use a relational database (such as MySQL, PostgreSQL) to store the corresponding relationship. Design a table that includes fields for business operation status, service item ID, and brief description of service items, etc. Through foreign key association, ensure the consistency and integrity of the data.
[0149] S402. Screen out the associated service items for the corresponding enterprise entities according to the business operation status and the corresponding relationship.
[0150] Read the business status information of the enterprise entity from the database, as well as the pre-established correspondence table between the business status and service items. In Python, use the pandas library to read the data, and use the merge function to match according to the business status field to filter out the list of service item IDs corresponding to the enterprise's business status.
[0151] In addition to the basic matching based on the business status, further filtering and optimization can be performed according to other attributes of the enterprise (such as industry category, enterprise scale, etc.). For example, for small enterprises, when matching service items, services with lower costs and simpler operations are preferably recommended; for large enterprises, more comprehensive and customized service items can be provided.
[0152] S403. Obtain the attributes of the enterprise entity, match the attributes with the object restriction information of the associated service items, and output the successfully matched associated service items as high-quality service items.
[0153] Standardize the attributes from the enterprise entity attributes extracted by the pre-trained named entity recognition model. For example, unify different expressions of the industry to which the enterprise belongs into the standard industry classification code; represent the enterprise scale with specific values (such as the number of employees, total assets).
[0154] For the object restriction information of the associated service items, perform parsing and structuring processing. For example, for the service item "R&D subsidy application service for high-tech enterprises", its object restriction information can be parsed as "Industry = High-tech industry".
[0155] Use conditional judgment and logical operations to match the attributes with the object restriction information. In Python, it is implemented by writing a function, such as: Def match_attributes(enterprise_attributes, service_restrictions): for key, value in service_restrictions.items(): if key not in enterprise_attributes or enterprise_attributes[key]!=value: return False; return True.
[0156] Traverse all associated service items, call the matching function for matching. Mark the successfully matched service items as high-quality service items, and organize and output information including the service item name, detailed description, service provider, etc.
[0157] S404. Push the description information of the high-quality service items to the email of the enterprise entity.
[0158] Use the enterprise's own mail server or select a third-party mail service provider such as SendGrid, Mailgun, etc. Evaluate and select based on factors such as the volume of emails sent, security, and cost.
[0159] When sending emails using Python, use the smtplib library (for the enterprise's own mail server) or a third-party library (such as the sendgrid library). Configure parameters such as the sender's email address, authorization password (or API key), SMTP server address, and port number.
[0160] Generate the email content in HTML format or plain text format according to the description information of the high-quality service items. Clearly display information such as the advantages of the service items, applicable scenarios, and expected effects in the email. For example, use Python's email library to construct the email content and set the email subject, body, attachments (if any), etc.
[0161] Traverse the email list of the enterprise entity and send the generated emails one by one. Add an error handling mechanism during the sending process, such as capturing exceptions when the email sending fails and recording error logs for subsequent troubleshooting and handling.
[0162] In addition, add a feedback mechanism: When sending emails with service item description information to the enterprise entity, utilize the tracking functions of the mail service provider (such as email open and click tracking of SendGrid) to record whether the enterprise has opened the email and which service item links have been clicked. If the enterprise replies to the email to express its views or requirements on the service items, use natural language processing technology (such as using the NLTK library for text classification and keyword extraction) to determine whether its feedback behavior is adoption, partial adoption, or rejection.
[0163] Create a dedicated online feedback form and guide the enterprise to fill it out in the email or on the service platform page. The form content covers whether to adopt the recommended service items, the reasons if not adopted, and whether there are other service requirements, etc. Use a Web development framework (such as Django or Flask) to build a form collection system and store the data in a database.
[0164] If the enterprise communicates with the customer service team by phone, instant messaging, etc., the customer service staff shall record the communication content in detail and enter the relevant feedback information into the customer relationship management system (CRM). Regularly extract data from the CRM system and organize it into a feedback behavior record.
[0165] Design a feedback record table in a relational database (such as MySQL), including fields such as enterprise entity ID, feedback time, feedback behavior type (adopt the described service item, adopt other service items, do not adopt any service items), and specific feedback content. Associate through the enterprise entity ID with the correspondence table of business status and service items to facilitate subsequent analysis.
[0166] Regularly clean the collected feedback data to remove duplicate records and invalid data (such as form submissions with incorrect formats). Use data processing tools (such as the pandas library) to preprocess the data to ensure the accuracy and consistency of the data, providing reliable data support for subsequent analysis and fine-tuning of the correspondence.
[0167] If a large number of enterprises adopt the recommended service items, it indicates that the correspondence between the current business status and service items is relatively accurate, and the weight of this correspondence can be appropriately increased. For example, in the recommendation algorithm, increase the recommendation priority of such service items.
[0168] If an enterprise adopts other service items, analyze the potential connection between these service items and the current business status. If a new association pattern is found, such as enterprises with the business status of "declining market share" generally adopting the "competitor analysis and market strategy adjustment service", then add this service item to the correspondence table and establish corresponding weights and recommendation rules.
[0169] When an enterprise does not adopt any service items, deeply analyze the reasons. If it is because the service items do not match the business status, re-evaluate and adjust the correspondence based on the enterprise feedback and market research. For example, for an enterprise with the business status of "technical innovation bottleneck", if the feedback shows that the recommended technical training service is not applicable, through market research, it is found that the enterprise needs more technical cooperation and introduction services, so the correspondence is modified.
[0170] In some embodiments, the service push system based on public opinion analysis may include multiple functional modules composed of computer program segments. The computer programs of each program segment in the service push system based on public opinion analysis can be stored in the memory of the computer device and executed by at least one processor to execute (see details in Figure 1 the description) the functions of service push based on public opinion analysis.
[0171] In this embodiment, the service push system based on public opinion analysis can be divided into multiple functional modules according to the functions it performs, such as Figure 3As shown. The functional modules of the system may include: an acquisition module, a processing module, a prediction module, and a push module. The modules referred to in the present invention refer to a series of computer program segments that can be executed by at least one processor and can complete fixed functions, and are stored in a memory. In this embodiment, the functions of each module will be described in detail in subsequent embodiments.
[0172] The acquisition module is used to acquire public opinion data; The processing module is used to extract enterprise entities, events, and time from the public opinion data, and arrange the events in chronological order as an event chain; The prediction module is used to predict the operating status of the enterprise entity according to the event chain; The push module is used to push the description information of the corresponding service item to the enterprise entity according to the correspondence between the preset operating status and the service item, and the operating status of the enterprise entity.
[0173] Figure 4 The service push method based on public opinion analysis provided in the embodiments of the present application can be applied to a device. Those skilled in the art can understand that the device structure involved in the embodiments of the present invention does not constitute a limitation on the device. The device may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements. In the embodiments of the present invention, the device includes, but is not limited to, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The device may also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the embodiments of the present application described herein and / or claimed.
[0174] Among them, the device 400 may include: a processor 410, a memory 420, and a communication unit 430. These components communicate through one or more buses. Those skilled in the art can understand that the structure of the server shown in the figure does not constitute a limitation on the present invention. It can be a bus structure, a star structure, and may also include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0175] Among them, the memory 420 can be used to store the execution instructions of the processor 410. The memory 420 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disc. When the execution instructions in the memory 420 are executed by the processor 410, the device 400 can execute some or all of the steps in the above method embodiments.
[0176] The processor 410 is the control center of the storage device, connecting various parts of the entire electronic device through various interfaces and lines. By running or executing the software programs and / or modules stored in the memory 420, and by calling the data stored in the memory, it executes various functions of the electronic device and / or processes data. The processor can be composed of an integrated circuit (IC). For example, it can be composed of a single packaged IC, or can be composed of multiple packaged ICs with the same or different functions connected together. For example, the processor 410 can include only a central processing unit (CPU). In the embodiment of the present invention, the CPU can be a single operation core or can include multiple operation cores.
[0177] The communication unit 430 is used to establish a communication channel so that the storage device can communicate with other devices. It receives user data sent by other devices or sends user data to other devices.
[0178] The present invention also provides a computer storage medium. Among them, the computer storage medium can store a program, and when the program is executed, it can include some or all of the steps in the various embodiments provided by the present invention. The storage medium can be a magnetic disk, an optical disc, a read-only memory (ROM) or a random access memory (RAM), etc.
[0179] Those skilled in the art can clearly understand that the technology in the embodiments of the present invention can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solutions in the embodiments of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk, or an optical disc, etc., various media that can store program codes, including several instructions for causing a computer device (which can be a personal computer, a server, or a second device, a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0180] For the same or similar parts among the various embodiments in this specification, reference can be made to each other. In particular, for the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the descriptions in the method embodiments.
[0181] In several embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are only illustrative. For example, the division of the modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of systems or modules can be in electrical, mechanical, or other forms.
[0182] The modules described as separate components may or may not be physically separated. The components displayed as modules may or may not be physical modules, that is, they can be located in one place, or they can be distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0183] In addition, the various functional modules in the various embodiments of the present invention can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module.
[0184] Although the present invention has been described in detail by reference to the accompanying drawings and in conjunction with the preferred embodiments, the present invention is not limited thereto. Without departing from the spirit and essence of the present invention, those of ordinary skill in the art can make various equivalent modifications or substitutions to the embodiments of the present invention, and all such modifications or substitutions should be within the scope of the present invention. / Any person skilled in the art within the technical scope disclosed by the present invention can easily conceive of changes or substitutions, which should all be covered within the protection scope of the present invention.
Claims
1. A service push method based on public opinion analysis, characterized in that: include: Obtain public opinion data; Extracting enterprise entities, events and time from the public opinion data, and arranging the events in chronological order into an event chain; predicting the operating status of the business entity based on the chain of events; According to the preset correspondence between the business status and the service items, and the business status of the business entity, the description information of the corresponding service items is pushed to the business entity.
2. The method according to claim 1, characterized in that: The method further comprises: Generate a hash value for each piece of public opinion data and remove duplicate data by verifying the uniqueness of the hash value; Correct the dates and locations in public opinion data; Use regular expressions to remove special characters in public opinion data and normalize the values.
3. The method according to claim 1, characterized in that: Extracting enterprise entities, events and time from the public opinion data, and arranging the events in chronological order into an event chain, including: Use pre-trained named entity recognition models to extract corporate entities, object entities, and relationships from public opinion data; Integrate enterprise entities, object entities, and relationships into events; Arrange multiple events involving the same business entity in chronological order as a chain of events.
4. The method according to claim 3, characterized in that The pre-trained named entity recognition model includes: BERT layer, used to extract dynamic context word vectors; The CRF layer is used to generate entity tag sequences based on the legal jump rules between entity tags; A fusion layer, used to use a binary classification network to determine whether there is a decorative relationship between entity pairs according to candidate entity pairs and entity position encoding in the entity label sequence, wherein the decorative relationship is used to determine the attributes of the enterprise entity; One or more fully connected layers, where the fully connected layers are used to classify the relationships between entities to obtain relationship categories.
5. The method according to claim 1, characterized in that Predicting the operating status of the business entity based on the event chain includes: Conduct trustworthy verification on events in multiple event chains of multiple enterprise entities, remove events that fail verification, and obtain a trustworthy event chain; Convert the trusted event chain of enterprise entities into a vector sequence, with each event corresponding to a vector; The vector sequence is input into a pre-trained long short-term memory neural network model to obtain the operating status of the business entity.
6. The method according to claim 4, characterized in that According to the preset correspondence between the business status and the service items, and the business status of the business entity, the description information of the corresponding service items is pushed to the business entity, including: According to the service requirements corresponding to the operating status, establish the corresponding relationship between the operating status and the service items; Filter out related service items for the corresponding business entity according to the business status and the corresponding relationship; Acquire the attributes of the enterprise entity, match the attributes with the object restriction information of the associated service items, and output the successfully matched associated service items as high-quality service items; The description information of the premium service item is pushed to the mailbox of the corporate entity.
7. The method according to claim 6, characterized in that The method further comprises: Tracking the feedback behavior of the business entity, including adoption of the service item, adoption of other service items, or non-adoption of any service item; The corresponding relationship is fine-tuned according to the feedback behavior.
8. A service push system based on public opinion analysis, characterized in that: include: Acquisition module, used to obtain public opinion data; A processing module, used to extract enterprise entities, events and time from the public opinion data, and arrange the events in chronological order into an event chain; a prediction module, for predicting the operating status of the business entity based on the event chain; The push module is used to push the description information of the corresponding service items to the enterprise entity according to the preset correspondence between the business status and the service items and the business status of the enterprise entity.
9. A device, characterized in that: include: A memory device for storing a service push program based on public opinion analysis; A processor, used to implement the steps of the service push method based on public opinion analysis as described in any one of claims 1-7 when executing the service push program based on public opinion analysis.
10. A computer-readable storage medium storing a computer program, characterized in that: The readable storage medium stores a service push program based on public opinion analysis, and when the service push program based on public opinion analysis is executed by the processor, the steps of the service push method based on public opinion analysis as described in any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Enterprise innovation power cloud evaluation and resource adaptation system and method
CN120765110A
Enterprise Innovation Cloud Assessment and Resource Matching System and Methodology
CN120765110B
Information sorting method and device, electronic equipment, storage medium and product
CN120833173A