Real-time processing method of automobile data based on artificial intelligence
Through the improved Transformer deep learning model and self-attention mechanism, combined with timing analysis and reinforcement learning technology, the problem of multimodal data fusion is solved, high-accurate emotion recognition and public opinion warning are achieved, and the company's market decision-making ability is improved.
Patent Information
- Application Number
- CN202510282783.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-05-20
- Estimated Expiration
- 2045-03-11
AI Technical Summary
The existing technology is difficult to effectively integrate text, voice, images and videos in multimodal data in the automotive industry, resulting in low accuracy of emotional recognition, affecting the accuracy of public opinion analysis and corporate decision-making.
The improved Transformer deep learning model, self-attention mechanism and multimodal feature fusion technology are used to conduct cross-context correlation analysis, combined with timing analysis and reinforcement learning technology to achieve accurate emotion recognition and public opinion warning of multimodal data.
It improves the accuracy of multimodal emotion recognition, can analyze emotional change trends across modalities, detect potential negative emotions or satirical information, provide real-time emotional trend prediction and public opinion response strategies, and enhance the company's brand management and market decision-making capabilities.
Smart Images

Figure CN119783051B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a real-time processing method for automotive data based on artificial intelligence. Background Art
[0002] In recent years, with the rapid development of the automotive industry and the popularization of Internet information dissemination, the market dynamics, user feedback, and brand public opinion information in the automotive industry have shown exponential growth. Channels such as social media, news media, short video platforms, forums, and blogs have become important platforms for consumers to express opinions, share experiences, and discuss hot events. However, due to the fragmentation, diversification, and large volume of information, traditional manual monitoring and analysis methods can no longer meet the needs of enterprises for market trend identification, consumer sentiment analysis, and competitor dynamics tracking. Therefore, a big data insight system for the automotive industry based on artificial intelligence has emerged, which takes natural language processing (NLP), deep learning, and big data analysis technologies as the core to achieve efficient collection, accurate analysis, and intelligent early warning of public opinion data in the automotive industry, so as to provide scientific decision-making support for enterprises.
[0003] The existing technologies have the following deficiencies:
[0004] In the process of big data analysis of automotive industry public opinion, the problem of cross-context sentiment recognition of multi-modal data is particularly serious and there are few effective solutions. Since automotive public opinion involves multi-modal data such as text (news, social media comments), voice (user feedback), images (accident photos, advertisements), and videos (evaluations, live broadcasts), there are huge differences in context understanding among different modal data, and the emotional expression methods are also different. For example, a user posts a post containing text and pictures on social media. Analyzing the text alone may show a positive evaluation, but in combination with the picture (such as a photo of a damaged vehicle), the actual emotion may be negative. However, existing natural language processing (NLP) and computer vision (CV) technologies are difficult to accurately integrate these heterogeneous data, making the system unable to accurately identify the true user emotions, resulting in distorted public opinion analysis results, thus affecting the brand management and market strategy formulation of enterprises. Summary of the Invention
[0005] The purpose of the present invention is to provide a real-time processing method for automotive data based on artificial intelligence to solve the deficiencies in the background art.
[0006] To achieve the above purpose, the present invention provides the following technical solutions: A real-time processing method for automotive data based on artificial intelligence, including the following steps:
[0007] S1: Real-time obtain multi-modal data related to the automotive industry from multiple data sources, including text, voice, images, and videos;
[0008] S2: Extract features from the collected multi-modal data. Among them, perform word segmentation, denoising, and sentiment word recognition processing on the text data, convert the speech data into text through speech recognition, and perform object detection and sentiment feature extraction on the image and video data;
[0009] S3: Based on the improved Transformer deep learning model, combined with the self-attention mechanism and multi-modal feature fusion technology, perform cross-contextual correlation analysis on the different features extracted from the multi-modal data, and evaluate the accuracy of the recognition of the user's true sentiment tendency according to the analysis results;
[0010] S4: For the accurately recognized true sentiment tendency of the user, construct a real-time sentiment trend prediction mechanism through a time series analysis model, and automatically send a warning signal to the preset enterprise management terminal when an abnormal event is detected;
[0011] S5: Utilize knowledge graph and reinforcement learning technology, combined with historical data and the accuracy of the recognition of the user's true sentiment tendency, generate coping strategies for public opinion events, and provide a visual analysis report to assist enterprises in formulating accurate market coping plans.
[0012] Preferably, the S1 multi-modal data collection adopts a combination of Web crawler, API interface call, and streaming data collection to obtain real-time and comprehensive data related to the automotive industry from social media, news platforms, video platforms, user forums, and e-commerce platforms.
[0013] Preferably, the S2 multi-modal data feature extraction includes the joint processing of text, speech, image, and video data, where: the text data is processed by BERT-WordPiece word segmentation, denoising, and sentiment dictionary matching; the speech data is subjected to sentiment recognition through the Wav2Vec2.0 and CNN+LSTM models; the image data uses YOLO, Faster R-CNN, and ResNet for object detection and OCR recognition; the video data uses TimeSformer and CNN-LSTM for key frame extraction and semantic sentiment analysis.
[0014] Preferably, the cross-contextual sentiment analysis based on the improved Transformer model extracts the correlation features between text, speech, image, and video through the self-attention mechanism, and maps different modal data to a unified sentiment representation space through modal alignment technology to improve the accuracy of cross-modal sentiment recognition.
[0015] Preferably, the sentiment drift index is used to calculate the sentiment change trend of the user in different time periods, and the context contradiction degree is used to evaluate the sentiment consistency of the multi-modal data.
[0016] Preferably, the calculation method of the emotional drift index is as follows: collect the review data of users on the same brand / product at multiple time points, and use the BERT+LSTM time series model to calculate the emotional score at each time point , and the expression is: ; where is the user text at time t, and f is the output function of the LSTM model; calculate the emotional change rate between adjacent time points as the emotional drift index: ; where Δt is the time interval, is the emotional drift index;
[0017] The calculation method of the context contradiction degree CDI is as follows: perform emotional recognition on multi-modal data, calculate the text emotion ST, speech emotion SV, image emotion SI, and video emotion SF respectively, and the range of each emotion score is [-1,1]; calculate the emotional consistency EC, and the expression is: ; calculate the context contradiction degree CDI, and the expression is: .
[0018] Preferably, use the Transformer encoding layer to jointly encode the text, speech, image, and video features; use the modal alignment technology to map different modal data to the same semantic space; use the self-attention mechanism to calculate the weights of different modal information , and the calculation formula: ; where Q, K, and V are the query, key, and value matrices respectively, is the dimension of the key, and T is the matrix transpose;
[0019] Measure the change relationship between SDC and CDI: If Corr(SDC, CDI) > 0.5, it means that the emotional drift is large, and the emotional contradiction between modalities is serious, and there is an implicit negative emotion or ironic tone; if Corr(SDC, CDI)< , it means that the emotional change is stable, the emotional expression is consistent, and the user's evaluation is credible.
[0020] Preferably, calculate the matching degree between the emotional label predicted by the model and the manually labeled emotional label, that is, calculate the emotional recognition accuracy SPA: ; in the formula, Predicted Sentiment is the emotional prediction value of the model for the sample, and the range is [-1,1]; Ground Truth is the true emotional label manually labeled, and the range is [-1,1], and N is the total number of samples. If SPA is lower than the threshold, the model is optimized;
[0021] Adopt the reinforcement learning training strategy to optimize the emotional recognition model through the reward mechanism: ; Use the policy gradient optimization method to adjust the parameters of the Transformer model to improve the accuracy of future predictions.
[0022] Preferably, the S4 time-series sentiment trend prediction mechanism is based on an improved Transformer structure, combined with a long short-term memory network, to construct a real-time public opinion change prediction model, and when an abnormal sentiment fluctuation value exceeds a set threshold, an early warning signal is automatically sent to the enterprise management terminal.
[0023] In the above technical solution, the technical effects and advantages provided by the present invention are:
[0024] 1. Through multi-modal data fusion, an improved Transformer deep learning model, self-attention mechanism, time-series analysis, and reinforcement learning technology, the present invention can collect real-time data related to the automotive industry from various channels such as social media, news, videos, and voice comments, and perform joint feature extraction on text, voice, image, and video data to accurately identify the true sentiment tendency of users. At the same time, through the calculation of the sentiment drift index (SDC) and context contradiction degree (CDI), the system can analyze the sentiment change trend across modalities, detect potential negative emotions or sarcastic information, and avoid misjudgment problems in single-modal analysis.
[0025] 2. The present invention also predicts the development trend of public opinion through a time-series analysis model (TimeSformer + LSTM), and when an abnormal event (such as a sudden increase in negative emotions, breaking news) is detected, an early warning signal is automatically sent to the enterprise management terminal to ensure that the enterprise can respond quickly. In addition, using knowledge graph and reinforcement learning technology, combined with historical data and user sentiment change patterns, an optimal public opinion response strategy is intelligently generated, and visual analysis reports are used to assist enterprises in making accurate market decisions. The present invention not only improves the accuracy of cross-modal sentiment recognition, but also optimizes the enterprise's brand management, market public relations, and risk control, and improves the enterprise's public opinion management ability and user trust in the highly competitive automotive market. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings.
[0027] Figure 1 It is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0028] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0029] For the embodiments, please refer to Figure 1 As shown, the real-time automotive data processing method based on artificial intelligence in this embodiment includes the following steps:
[0030] S1: Real-time obtain multi-modal data related to the automotive industry from multiple data sources, including text, speech, images, and videos;
[0031] S2: Extract features from the collected multi-modal data. Among them, perform word segmentation, denoising, and sentiment word recognition processing on text data, perform speech recognition on speech data to convert it into text, and perform object detection and sentiment feature extraction on image and video data;
[0032] S3: Based on an improved Transformer deep learning model, combine self-attention mechanism and multi-modal feature fusion technology to perform cross-contextual correlation analysis on different features extracted from multi-modal data, and evaluate the accuracy of identifying the true sentiment tendency of users according to the analysis results;
[0033] S4: For the accurately identified true sentiment tendency of users, construct a real-time sentiment trend prediction mechanism through a time series analysis model, and automatically send a warning signal to a preset enterprise management terminal when an abnormal event is detected;
[0034] S5: Utilize knowledge graph and reinforcement learning technology, combine historical data and the accuracy of identifying the true sentiment tendency of users to generate coping strategies for public opinion events, and provide a visual analysis report to assist enterprises in formulating accurate market response plans.
[0035] In S1, based on multiple heterogeneous data sources, ensure that the collected information is comprehensive and accurate, covering the following types:
[0036] Social media platforms: Obtain automotive-related comments, short videos, live content, voice discussions, etc. published by users.
[0037] News media: Collect automotive market dynamics, industry policies, brand evaluations, recall information, etc.
[0038] Video platforms: Obtain video content such as automotive evaluations, user sharing, and fault feedback, and perform parsing in combination with automatic caption / speech analysis.
[0039] User forums and communities: Capture users' long-form car-buying experiences, problem consultations, discussion posts, and other textual content.
[0040] E-commerce and after-sales platforms: obtain users’ car purchase evaluations, maintenance feedback, parts quality reviews, etc.
[0041] In order to collect the above multimodal data in real time and efficiently, the present invention adopts the following technical means: Web crawler technology: use crawler programs to regularly or in real time crawl news, forums, social media and other web page data. API interface call: obtain real-time content through the official API interface provided by social media and video platforms, streaming data collection: use streaming data processing frameworks such as Kafka and Flink to achieve continuous collection of real-time published content. Speech to text (ASR, Automatic Speech Recognition): Automatic speech recognition of voice content in audio or video, and transcription of speech into text for further sentiment analysis. Image and video analysis (OCR, target detection, sentiment analysis): perform target recognition (such as brand logo, vehicle model recognition), OCR text extraction, video summary analysis, etc. on the acquired pictures and video frames.
[0042] The collected multimodal data will be pre-processed through the ETL (Extract, Transform, Load) process and stored in a distributed database (such as Elasticsearch, MongoDB) or a big data platform (such as Hadoop, Spark) for subsequent deep learning modeling and public opinion analysis.
[0043] Since multimodal data may contain a large amount of redundant or irrelevant information, this invention ensures the accuracy and validity of the data through technologies such as text deduplication algorithms (such as SimHash), spam filtering (such as blacklists, stop words), and format standardization (such as timestamp conversion).
[0044] S2: Extract features from the collected multimodal data, including word segmentation, denoising and sentiment word recognition for text data, speech recognition and conversion of speech data into text, and target detection and sentiment feature extraction for image and video data.
[0045] Text data mainly comes from social media, news websites, user comments, forum posts, policy announcements, etc. The processing process includes:
[0046] Since Chinese text has no natural word boundaries, a word segmentation algorithm is needed to split sentences into independent words to improve the model's understanding ability: rule-based word segmentation (such as regular expression matching), statistical word segmentation (such as TF-IDF weight calculation), deep learning word segmentation (such as BERT-WordPiece, LSTM word segmenter).
[0047] To improve data quality, it is necessary to remove invalid information: HTML tags, special symbols, advertising content, stop words (such as meaningless words like "de", "le", "ne", etc.), and duplicate character detection (e.g., "aaaaaa" → "a").
[0048] Using the combination of sentiment dictionary and sentiment analysis model, identify positive, negative, or neutral sentiment in the text: keyword matching (such as "positive review", "fault", "complaint"), dependency syntactic analysis (e.g., "vehicle + [has a fault]" indicates negative sentiment), and combine with the deep learning BERT sentiment analysis model for sentiment classification.
[0049] The voice data mainly comes from user voice comments, telephone customer service recordings, video audio, etc. The processing process includes: using automatic speech recognition (ASR) technology to convert speech into text. Common methods include: traditional HMM-GMM (Hidden Markov Model + Gaussian Mixture Model); modern deep learning models. Example: User voice: "This car is really too fuel-consuming, simply unaffordable!" Converted text: "This car is really too fuel-consuming, simply unaffordable!" Analyze the emotional characteristics of the voice (such as anger, satisfaction, doubt). Methods include: Mel Frequency Cepstral Coefficients (MFCC) to extract emotional features; pitch, volume, and speech rate analysis (e.g., high pitch and fast speech rate when angry); neural network emotion classification (such as CNN + LSTM); finally, the emotional tendency (positive / negative / neutral) and emotion category (such as angry / happy / anxious) can be obtained.
[0050] The image data mainly comes from social media pictures, news photos, accident photos, vehicle advertisements, etc. The processing process includes: used to identify the car brand, model, license plate, damaged parts, etc. in the picture. Methods include: CNN (Convolutional Neural Network) + Faster R-CNN; YOLO (You Only Look Once) series; lightweight models such as ResNet and MobileNet.
[0051] Identify the text information in the image, such as license plate number, announcement, user evaluation screenshot, etc. Adopt: Tesseract OCR (open source), PaddleOCR (high-precision), Google Vision OCR API; Example: Input: A screenshot of a recall announcement, OCR result: Identify "Due to a fault in the battery management system, this batch of vehicles will be recalled for repair."
[0052] Analyze the main body of the image (such as accident scene, product promotion picture), color emotion analysis (e.g., red tends to warn, blue tends to be safe), and deep learning expression recognition (such as whether the user shows anger / satisfaction in the picture).
[0053] Video data mainly comes from car reviews, user experience sharing, accident videos, advertising, etc. The processing process includes: using frame difference method, optical flow method, and CNN-LSTM method to extract key frames and remove redundant frames. For example, screen out key frames related to vehicle damage from "user complaint videos". Detect vehicle brands, models, road conditions, etc. through algorithms such as YOLO and Faster R-CNN; perform OCR recognition on the caption or annotation information in the video to extract text information; for example, extract vehicle parameters from "car advertisement videos". Extract the voice content in the video and convert it to text using ASR; perform sentiment analysis on the comments of narrators or users.
[0054] S3: Based on the improved Transformer deep learning model, combined with self-attention mechanism and multi-modal feature fusion technology, perform cross-contextual correlation analysis on different features extracted from multi-modal data, and evaluate the accuracy of identifying users' true sentiment tendencies according to the analysis results.
[0055] Collect real-time data related to the automotive industry generated by users from different channels, including: text (social media, forums, news comments, etc.); voice (phone customer service, user video comments); images (accident photos uploaded by users, product advertisements); videos (car reviews, news reports, user experience sharing).
[0056] Data preprocessing and feature extraction, text: word segmentation (BERT-WordPiece), denoising, sentiment word recognition; voice: automatic speech recognition (ASR), speech emotion feature extraction (Mel-frequency cepstral coefficients, MFCC); images: object detection (YOLO, Faster R-CNN), OCR text extraction; videos: key frame extraction (CNN-LSTM), speech emotion analysis.
[0057] Perform cross-contextual correlation analysis on different features extracted from multi-modal data, including:
[0058] Measure the emotional change trend of users towards the same brand / event at different time periods. Traditional sentiment analysis is usually based on static text or sentiment snapshots at a single time point. By introducing an emotional drift index, calculate the change rate of sentiment tendency over time to identify long-term effects or short-term fluctuations.
[0059] Among them, the calculation method of the emotional drift index is:
[0060] Collect users' comment data on the same brand / product at multiple time points (such as Weibo posts, forum discussions, news reports, etc.). Use the BERT+LSTM time series model to calculate the sentiment score at each time point , the expression is: ; among them, is the user text at time t, and f is the output function of the LSTM model.
[0061] Calculate the emotional change rate at adjacent time points as the emotional drift index: ; where Δt is the time interval, is the emotional drift index.
[0062] Negative SDC (emotional decline trend): If the user's emotional drift coefficient drops sharply in a short period (e.g., from 0.8 to -0.5), it indicates that a sudden negative event may have occurred for the brand / product (such as vehicle recall, accident exposure).
[0063] Positive SDC (emotional rise trend): If the SDC grows rapidly, it means that the user's mood has changed positively, which may be related to factors such as corporate public relations activities and product upgrades.
[0064] Stable SDC (low change rate): The emotion remains stable, indicating that the brand's reputation is relatively persistent.
[0065] Measure the consistency or contradiction degree of multi-modal information such as text, speech, images, and videos in emotional expression. Existing sentiment analysis usually analyzes a single modality separately. By obtaining the Contextual Discrepancy Index (CDI), the conflict situation between modalities can be detected to avoid misjudgment.
[0066] The calculation method of the Contextual Discrepancy Index CDI is:
[0067] Perform sentiment recognition on multi-modal data, and calculate the text sentiment ST, speech sentiment SV, image sentiment SI, and video sentiment SF respectively. Each sentiment score ranges from [-1, 1].
[0068] Calculate the Emotional Consistency EC, and the expression is: ; where, A value close to 1 indicates high consistency, A value close to 0 indicates a high degree of conflict.
[0069] Calculate the Contextual Discrepancy Index CDI, and the expression is: ; The higher the CDI value, the higher the degree of contradiction in emotion between different modalities.
[0070] High CDI (large degree of contradiction): It means that there is an emotional conflict in the information expressed by the user. For example, the user's comment content is "This car is great!", but the accompanying picture is a photo of the car breaking down, indicating that the user may be being sarcastic or have potential negative emotions.
[0071] Low CDI (small degree of contradiction): It means that the multi-modal data is relatively consistent in emotion, indicating that the user's emotional expression is true and credible.
[0072] The Transformer Encoder is used to jointly encode text, speech, image, and video features.
[0073] The Modality Alignment technique is adopted to map different modality data into the same semantic space.
[0074] The Self-Attention mechanism is used to calculate the weights of different modality information, ensuring that the most relevant information receives higher attention.
[0075] Calculation formula (Self-Attention calculation): ; where Q, K, and V are the Query, Key, and Value matrices respectively, is the dimension of the key, and T is the matrix transpose.
[0076] Calculate SDC-Correlation to measure the change relationship between SDC and CDI: If Corr(SDC, CDI) > 0.5, it indicates a large emotional drift, serious emotional contradictions between modalities, and the existence of implicit negative emotions or a sarcastic tone; if Corr(SDC, CDI) < , it indicates a stable emotional change, consistent emotional expression, and the credibility of the user's evaluation.
[0077] Calculate the matching degree between the emotional label predicted by the model and the manually annotated emotional label, that is, calculate the Sentiment Recognition Accuracy SPA: ; In the formula, is the emotional prediction value of the model for the sample, with a range of [-1, 1] (-1 represents strong negative, 0 represents neutral, and 1 represents strong positive). GroundTruth is the true emotional label manually annotated, also with a range of [-1, 1], and N is the total number of samples. If SPA is lower than the threshold (such as 85%), the model is optimized.
[0078] Adopt the Reinforcement Learning (RL) training strategy to optimize the emotion recognition model through the Reward Function: ; Use the Policy Gradient Optimization method to adjust the parameters of the Transformer model to improve the accuracy of future predictions.
[0079] If a decrease in SDC and an increase in CDI are detected, an enterprise public opinion warning is triggered to remind the brand side to take measures. Combining historical data, use the knowledge graph to generate automated response strategies, such as: issuing an official statement in advance, adjusting the marketing strategy, and improving the product design.
[0080] S4: For the accurately identified true emotional tendency of users, construct a real-time emotional trend prediction mechanism through a time series analysis model, and automatically send a warning signal to the preset enterprise management terminal when an abnormal event is detected.
[0081] Emotional trend prediction requires data streams based on continuous time, mainly including the following information:
[0082] Text emotion (ST): Social media comments, news, forum posts (calculate emotion scores using BERT / BART);
[0083] Voice emotion (SV): User voice feedback, complaint calls (calculate emotion scores using Wav2Vec2.0);
[0084] Image emotion (SI): Accident photos, promotional advertisements uploaded by users (identify visual emotions using ResNet + OCR);
[0085] Video emotion (SF): User evaluations, news reports, short videos (calculate video emotional trends using TimeSformer).
[0086] These emotional data are stored in a time series and standardized: Among them, S(t) represents the comprehensive emotional state at time t.
[0087] Construct a time series emotional dataset: Aggregate data by time windows (such as 1 hour, 1 day, 1 week), calculate the average emotional value, emotional fluctuation range, and abnormal emotional deviation.
[0088] Traditional time series models (such as ARIMA) cannot effectively capture non-linear, long-term dependent emotional changes. Therefore, this method adopts a Transformer prediction model:
[0089] For large-scale time series data, this method uses TimeSformer (temporal Transformer) for prediction and uses self-attention mechanisms to model long-term emotional changes: ; Among them, Q, K, and V represent the query, key, and value matrices of emotional data respectively, and softmax calculates the correlation to ensure that important information receives more attention.
[0090] Abnormal events usually manifest as sudden and drastic changes in emotional values. Therefore, this method calculates the emotional anomaly degree (AE, Anomaly Score) to determine whether to trigger a warning: ; Among them: S(t) is the current emotional score, is the normal emotional trend predicted by the model, is the historical emotional fluctuation range; if If > θ (set threshold, such as 2.5), it is considered that abnormal emotional fluctuations occur, and a warning signal is automatically sent to the preset enterprise management terminal.
[0091] S5: Using knowledge graph and reinforcement learning technologies, combined with the accuracy of historical data and user real emotional tendency recognition, generate response strategies for public opinion events, and provide visual analysis reports to assist enterprises in formulating accurate market response plans.
[0092] The knowledge graph is used to represent the multi-level semantic relationships of public opinion events, facilitating situation analysis, risk assessment, and response strategy reasoning.
[0093] Knowledge graph construction includes: Entities (Nodes): brands, models, events (spontaneous combustion, recall), markets, etc.; Relationships (Edges): vehicle failure → recall, product upgrade → user praise, negative news; Using natural language processing (NLP) + relationship extraction, extract triples from news, social media, and industry reports to construct a dynamic knowledge base.
[0094] Use reinforcement learning (RL) to train an intelligent policy model so that it can automatically optimize enterprise response measures in different public opinion situations. The core Agent-Environment interaction architecture of reinforcement learning: Agent (intelligent agent): public opinion management system; State (status): current public opinion situation (event type, emotional trend, market impact, etc.); Action (action): possible enterprise response strategies (such as official statement, product recall, price adjustment, etc.). Reward (reward): Measure the effectiveness of each strategy based on historical data, such as user emotional recovery degree and brand trust improvement.
[0095] Reinforcement learning optimization strategy process: State perception (State Representation): Input current public opinion information (negative emotion index, event severity, etc.); Policy selection (Policy Selection): Use algorithms such as Q-learning / DDPG / PPO to select the best response measures; Feedback learning (Reward Calculation): Update the policy based on user emotional recovery degree and market impact analysis; Continuous optimization (Policy Improvement): Optimize the policy recommendation system through multiple rounds of interactive training.
[0096] Generate precise public opinion response strategies through reinforcement learning historical decision backtracking + real user sentiment data. Historical similar event mining: Query based on the knowledge graph to find past similar public opinion events. Effect evaluation: Calculate the public opinion recovery rate (RR) and user sentiment change (SDC, CDI) of historical strategies. Automatically learn the optimal strategy: The reinforcement learning model compares the effects of different strategies and selects the optimal solution example.
[0097] Real user sentiment analysis: Text analysis (social media, forum comments): Extract negative sentiment keywords (such as "out of control", "recall"); Voice sentiment analysis (customer service, complaint recordings): Extract angry / anxious emotion features; Multimodal sentiment fusion (pictures / videos): Detect emotional signals such as user expressions and accident photos; The system adjusts the response plan according to the real sentiment data of users: If users are mainly worried about safety → Strengthen technical transparency and release safety improvement measures; If the angry emotion of users dominates → Prioritize compensation / recall strategies; If the public opinion has not fully spread → Respond low-key to avoid expanding negative impacts. Finally, the system displays the public opinion dynamics through data visualization and provides intelligent decision-making support. Generate a public opinion visualization report, including: Public opinion heat trend (changes in negative news & social discussion volume); Sentiment tendency change (changes in user praise rate, complaint rate); Competitor comparison analysis (Has the competing product also encountered similar public opinion? How to respond?).
[0098] Example visualization content: Timeline graph (Timeline): Show the whole process of the outbreak, spread, and subsidence of public opinion; Sentiment trend curve (Sentiment Trend): The change of users' positive and negative emotions over time; Impact range map (Geo-Impact Map): Analyze the public opinion distribution of users in different regions; Strategy simulation comparison (Scenario Simulation): The expected impact of different response strategies.
[0099] Generate intelligent response strategy suggestions, including: Optimal strategy recommendation (calculate the best solution based on the knowledge graph + reinforcement learning); Automatically generate public relations manuscripts (generate official response texts in combination with NLP); Marketing adjustment suggestions (such as adjusting advertising investment to reduce losses).
[0100] In this embodiment, first, the system collects multi-source data from social media, news, voice comments, images, videos, etc. in real time and performs feature extraction, including text tokenization, speech transcription, image / video object detection, and sentiment analysis. Subsequently, an improved Transformer deep learning model is adopted, combined with the self-attention mechanism and multi-modal feature fusion technology, to perform cross-context association analysis and improve the accuracy of identifying the true emotional tendencies of users. Based on the accurately identified emotional data, the system predicts the emotional trend through a time series analysis model and automatically sends a warning to the enterprise management end when an abnormal event is detected. Finally, using knowledge graph and reinforcement learning technologies, combined with historical data and user emotional trends, an optimal public opinion response strategy is automatically generated and a visual analysis report is provided to help enterprises make precise market decisions. This system can improve the enterprise's public opinion monitoring, risk warning, and brand management capabilities and enhance its market competitiveness.
[0101] The above formulas are all dimensionless and take their numerical values for calculation. The formulas are obtained by collecting a large amount of data for software simulation to obtain a formula that is closest to the actual situation. The preset parameters in the formulas are set by those skilled in the art according to the actual situation.
[0102] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless (such as infrared, wireless, microwave, etc.) manner. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains a collection of one or more available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.
[0103] It should be understood that the term "and / or" in this text is merely a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent: the situation where A exists alone, the situation where both A and B exist simultaneously, and the situation where B exists alone. Here, A and B can be singular or plural. Additionally, the character " / " in this text generally represents an "or" relationship between the associated objects before and after, but it may also represent an "and / or" relationship, and the specific meaning can be understood by referring to the context before and after. Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this text can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0104] As described above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, and all should be covered within the protection scope of this application.
Claims
1. A real-time processing method for automobile data based on artificial intelligence, characterized in that: The following steps are involved: S1: Real-time acquisition of automotive-related multimodal data from multiple data sources, including text, voice, images, and videos; S2: Extract features from the collected multimodal data, including word segmentation, denoising and sentiment word recognition for text data, speech recognition and conversion of speech data into text, and object detection and sentiment feature extraction for image and video data; S3: Based on the improved Transformer deep learning model, combined with the self-attention mechanism and multimodal feature fusion technology, cross-context correlation analysis is performed on different features extracted from multimodal data, and the accuracy of identifying the user's true emotional tendency is evaluated based on the analysis results; The S3 is based on cross-context sentiment analysis of the improved Transformer deep learning model. It extracts correlation features between text, speech, images and videos through the self-attention mechanism, and maps different modal data to a unified sentiment representation space through modal alignment technology to improve the accuracy of cross-modal sentiment recognition. It uses the sentiment drift index to calculate the sentiment change trend of users in different time periods, and uses the context contradiction degree to evaluate the sentiment consistency of multimodal data. The calculation method of sentiment drift index is as follows: collect users’ comments on the same brand or product at multiple time points, use BERT+LSTM time series model to calculate the sentiment score at each time point , the expression is: ;in, is the user text at time t, f is the output function of the LSTM model; the sentiment change rate at adjacent time points is calculated as the sentiment drift index: ; where Δt is the time interval, is the sentiment drift index; The calculation method of contextual contradiction degree CDI is as follows: perform emotion recognition on multimodal data, calculate text emotion ST, voice emotion SV, image emotion SI, and video emotion SF respectively, and the score range of each emotion is [-1,1]; calculate emotion consistency EC, the expression is: ; Calculate the contextual contradiction degree CDI, the expression is: ; The Transformer encoding layer is used to jointly encode text, speech, image, and video features; the modality alignment technology is used to map different modal data to the same semantic space; the self-attention mechanism is used to calculate the weights of different modal information , calculation formula: ; where Q, K, and V are query, key, and value matrices, respectively. is the dimension of the key, T is the matrix transpose; Measuring the changing relationship between SDC and CDI: If Corr(SDC, CDI) > 0.5, it means that the emotional drift is large, and the emotional contradiction between the modes is serious, and there is implicit negative emotion or sarcastic tone; if Corr(SDC, CDI)< , indicating that the emotional changes are stable, the emotional expressions are consistent, and the user's evaluation is credible; Calculate the matching degree between the emotion labels predicted by the model and the manually annotated emotion labels, that is, calculate the emotion recognition accuracy SPA: ; Where Predicted Sentiment is the sentiment prediction value of the model for the sample, ranging from [−1, 1]; Ground Truth is the true sentiment label manually annotated, ranging from [−1, 1], N is the total number of samples, if SPA is lower than the threshold, the model is optimized; Adopt reinforcement learning training strategy and optimize emotion recognition model through reward mechanism: ; Use policy gradient optimization to adjust the parameters of the Transformer model to improve the accuracy of future predictions; S4: For the accurate identification of the user's true emotional tendency, a real-time emotional trend prediction mechanism is built through a time series analysis model, and when an abnormal event is detected, an early warning signal is automatically sent to the preset enterprise management end; The S4 time series sentiment trend prediction mechanism is based on the improved Transformer deep learning model, combined with the long short-term memory network, to build a real-time public opinion change prediction model. The input of the model is the current sentiment score, and the output is the sentiment abnormality. When the abnormal sentiment fluctuation value exceeds the set threshold, it automatically sends a warning signal to the enterprise management end. S5: Utilize knowledge graphs and reinforcement learning technology, combined with historical data and the accuracy of identifying users’ real emotional tendencies, to generate response strategies for public opinion events and provide visual analysis reports to assist companies in formulating accurate market response plans.
2. The method for real-time processing of automobile data based on artificial intelligence according to claim 1, characterized in that: The S1 multimodal data collection adopts a combination of Web crawlers, API interface calls and streaming data collection to obtain real-time and comprehensive automotive industry-related data from social media, news platforms, video platforms, user forums and e-commerce platforms.
3. The method for real-time processing of automobile data based on artificial intelligence according to claim 1, characterized in that: The S2 multimodal data feature extraction includes the joint processing of text, speech, image and video data, wherein: text data is processed by BERT-WordPiece word segmentation, denoising and sentiment dictionary matching; speech data is processed by Wav2Vec2.0 and CNN+LSTM model for sentiment recognition; image data is processed by YOLO, Faster R-CNN and ResNet for target detection and OCR recognition; video data is processed by TimeSformer and CNN-LSTM for key frame extraction and semantic sentiment analysis.
Citation Information
Patent Citations
Network information analysis method based on privacy grouping and emotion recognition
CN112464281A
Self-supervised multi-modal sentiment analysis method based on Transform feature learning
CN118861980A