Sales behavior intelligent analysis and early warning system
Patent Information
- Application Number
- CN202610826207.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-09
- Publication Date
- 2026-09-08
AI Technical Summary
[0005]有鉴于此,本申请的实施例致力于提供一种销售行为智能分析与预警系统及方法,旨在克服现有技术中数据采集依赖手动上传导致的操作繁琐、数据滞后、信息遗漏、主观性强,数据分析依赖人工经验导致的客观性无法保证、分析效率低、隐性风险难以发现,以及数据处理延迟高、数据隐私风险大、预警精度低、多模态数据利用不充分、边缘端计算资源受限等技术问题中的至少一个
1. 自动化数据采集,消除手动操作负担:通过数据采集模块实时自动采集多模态原始数据,销售人员无需在客户沟通结束后手动录入跟进记录和任务状态,节省30-60分钟的每日录入时间,同时避免信息遗漏和数据滞后,确保数据的完整性、及时性和客观性。
Smart Images

Figure CN122714065A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent sales management technology, and more specifically, to an intelligent analysis and early warning system for sales behavior. Background Technology
[0002] In the field of sales management, real-time monitoring, analysis, and early warning of sales personnel's customer follow-up behavior are key means to improve sales efficiency and conversion rates.
[0003] In existing sales management systems, the entry of data such as task status, follow-up records, and customer feedback mainly relies on manual operation by sales staff. Sales staff need to spend extra time filling in follow-up records, updating task status, and uploading communication summaries in the CRM system after customer communication. This manual uploading method has the following drawbacks: Inconvenient operation and wasted time: Sales staff need to spend 30-60 minutes daily on data entry, consuming valuable customer communication time and reducing work efficiency. Data lag: Manual entry is usually done after communication, making real-time data collection and analysis impossible, resulting in sales data seen by managers being delayed by hours or even days. Information omission: Sales staff may forget or selectively enter some information, leading to the loss of key data and affecting the completeness of subsequent analysis. High subjectivity: Sales staff's summaries and generalizations of communication content are influenced by personal subjective judgment; different sales staff may describe the follow-up status of the same type of customer differently, resulting in inconsistent data standards.
[0004] Current systems primarily rely on the human experience of sales managers to analyze sales data. Managers identify problematic salespeople, customers, and processes by reviewing CRM reports, listening to call recordings, and reading chat logs, based on their personal experience. This experience-dependent approach has several drawbacks: Subjectivity: Different managers may interpret the same data differently, leading to inconsistent results and a lack of objective, unified evaluation standards. Low efficiency: In larger teams (50+ people), managers cannot thoroughly analyze the follow-up quality of each salesperson, resulting in a focus on the big picture and overlooking many potential problems. Difficulty in detecting hidden risks: Manual analysis struggles to identify subtle, gradual anomalies in massive datasets (such as a slow decline in sales script consistency or subtle changes in customer emotions), often only becoming apparent after a significant drop in performance. High reliance on experience: The experience of excellent managers is difficult to replicate and pass on; when managers change or leave, the team's analytical capabilities and management level may decline significantly. Summary of the Invention
[0005] In view of this, the embodiments of this application are committed to providing a sales behavior intelligent analysis and early warning system and method, aiming to overcome at least one of the following technical problems in the prior art: data collection relies on manual uploading, resulting in cumbersome operation, data lag, information omission, and strong subjectivity; data analysis relies on human experience, resulting in unreliable objectivity, low analysis efficiency, and difficulty in discovering hidden risks; as well as high data processing latency, high data privacy risks, low early warning accuracy, insufficient utilization of multimodal data, and limited computing resources at the edge.
[0006] This application provides a sales behavior intelligent analysis and early warning system, including:
[0007] The edge processing unit deployed on each salesperson's local terminal includes: The data acquisition module is configured to collect multimodal raw data generated by sales personnel during customer follow-up in real time; the multimodal raw data includes: text modal data, voice modal data, structured behavioral modal data, and time-series behavioral modal data; A multimodal encoder group, connected to the data acquisition module, is configured to extract features from each modal data and generate corresponding modal feature vectors. A multimodal fusion module, connected to the multimodal encoder group, is configured to fuse the feature vectors of each modality through a cross-modal attention mechanism to generate a unified multimodal fusion feature; An analysis module, connected to the multimodal fusion module, is configured to analyze the actual state of a task based on the multimodal fusion features, wherein the actual state of a task includes at least the task execution progress. The prediction module, connected to the multimodal fusion module, is configured to perform at least one prediction task based on the multimodal fusion features to obtain a prediction result; the prediction task includes at least one of: order probability prediction, performance amount prediction, and customer churn risk assessment. And, the server, communicating with the edge processing units of each local terminal, including: An incremental data receiving interface is configured to receive incremental data uploaded by the edge processing unit; the incremental data includes at least: information on changes in the actual state of the task and information on updates to the prediction results; The real-time dashboard service module is connected to the incremental data receiving interface and is configured to perform real-time aggregation and statistics on local historical data and received incremental data, obtain statistical data, and push it to a preset administrator terminal. The management terminal displays sales performance data through a visual interface based on the statistical data.
[0008] Optionally, the multimodal raw data includes data specific to different tasks; The edge processing unit also includes a priority scheduling module, configured to prioritize processing data corresponding to high-priority tasks based on a preset task priority configuration.
[0009] Optionally, the multimodal encoder group includes a speech preprocessing encoder, which is configured as follows: The speech modal data is segmented, and low-information speech data is identified and removed. The low-information voice data includes at least one of the following: Small talk and social opening remarks; Filler words and catchphrases; Silence and pauses; Repetitive information; Non-business related information.
[0010] Optionally, the analysis module is further configured as follows: Analyze the difficulties encountered during task processing and advancement; When the difficult information meets the preset alarm rules, an alarm prompt message is generated.
[0011] Optionally, the analysis module is specifically configured as follows: Construct or obtain a sales persona for a salesperson, wherein the sales persona includes at least: job title tags, competency tags, and historical performance tags; The threshold parameters of the preset alarm rules are dynamically adjusted based on the salesperson profile.
[0012] Optionally, the analysis module is further configured as follows: When the preset triggering conditions are met, the salesperson profile is updated, and the preset alarm rules corresponding to the salesperson are updated simultaneously. The preset triggering conditions include at least one of the following: Received a preset character portrait adjustment command; Sales personnel have completed a preset threshold number of tasks. The time interval since the last portrait update has reached the preset duration.
[0013] Optionally, the edge processing unit is a lightweight deployment unit that satisfies at least one of the following characteristics: The multimodal encoder group, multimodal fusion module, analysis module, and prediction module are lightweight models that have undergone model quantization, knowledge distillation, or structured pruning. The edge processing unit operates independently when the local terminal is offline, without relying on a real-time network connection with the server.
[0014] Optionally, the manager terminal is communicatively connected to the edge processing units of each local terminal; The administrator terminal is configured to receive the actual task status and prediction results uploaded by the edge processing unit. It provides an interactive visualization interface that supports filtering and drill-down analysis of the actual status and predicted results of tasks by time, personnel, and team dimensions.
[0015] Optionally, the edge processing unit further includes a local storage module, configured as follows: The system stores the multimodal raw data collected by the data acquisition module.
[0016] Optionally, the local storage module is further configured to asynchronously upload the stored multimodal raw data to the server when the local terminal is detected to be in a preset idle state; The server is equipped with a pre-trained general processing model for offline batch processing of uploaded multimodal raw data to correct or optimize the analysis and prediction results of the edge processing unit.
[0017] Compared with the prior art, the sales behavior intelligent analysis and early warning system and method provided in this application have the following beneficial effects: 1. Automated data collection eliminates the burden of manual operation: The data collection module automatically collects multimodal raw data in real time, eliminating the need for sales personnel to manually enter follow-up records and task status after customer communication, saving 30-60 minutes of daily data entry time. At the same time, it avoids information omissions and data lag, ensuring the integrity, timeliness and objectivity of the data.
[0018] 2. Objective data analysis, eliminating reliance on human experience: Through multimodal encoders, multimodal fusion modules, analysis modules, and prediction modules, automated and standardized analysis of sales behavior is achieved, outputting objective analysis and prediction results, unaffected by personal subjective judgment. Managers no longer need to analyze sales data piecemeal based on experience, significantly improving analysis efficiency and objectivity. Hidden risks (such as a slow decline in the standardization of sales scripts or subtle changes in customer emotions) can be automatically identified and warned of, preventing them from being discovered only after a significant drop in performance.
[0019] 3. Low-latency real-time response: By deploying multimodal encoding, fusion, analysis, and prediction all on the local terminal's edge processing unit, real-time local processing of sales activities is achieved. From the occurrence of a sales activity to the output of analysis results, the end-to-end latency can be controlled within 100 milliseconds, supporting real-time alerts and guidance at critical moments in sales and customer communication.
[0020] 4. High privacy protection: All raw multimodal data (especially sensitive voice and text data) is processed locally on the terminal and not uploaded to the server. Only anonymized incremental data (such as status changes and prediction results) is uploaded, meeting data compliance requirements such as GDPR and personal information protection laws, and is applicable to highly regulated industries such as finance, education, and healthcare.
[0021] 5. High Early Warning Accuracy: By fusing multimodal data and comprehensively utilizing multi-dimensional information such as voice, text, behavior, and time series, the prediction accuracy can be improved by 15-25% compared to single-modal analysis. Personalized early warnings are achieved by adaptively adjusting alarm thresholds based on user profiles, resulting in a 30-50% reduction in false alarm rate compared to fixed threshold solutions. This significantly increases sales personnel's trust in and acceptance of early warnings. Attached Figure Description
[0022] The above and other objects, features, and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the embodiments of this application to explain this application and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.
[0023] Figure 1 This is a schematic diagram of the structure of a sales behavior intelligent analysis and early warning system provided in one embodiment of this application. Detailed Implementation
[0024] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0025] Figure 1 This is a schematic diagram of the structure of a sales behavior intelligent analysis and early warning system provided in one embodiment of this application; see reference. Figure 1 The sales behavior intelligent analysis and early warning system provided in this application includes: The edge processing unit deployed on each salesperson's local terminal includes: The data acquisition module is configured to collect multimodal raw data generated by sales personnel during customer follow-up in real time; the multimodal raw data includes: text modal data, voice modal data, structured behavioral modal data, and time-series behavioral modal data; A multimodal encoder group, connected to the data acquisition module, is configured to extract features from each modal data and generate corresponding modal feature vectors. A multimodal fusion module, connected to the multimodal encoder group, is configured to fuse the feature vectors of each modality through a cross-modal attention mechanism to generate a unified multimodal fusion feature; An analysis module, connected to the multimodal fusion module, is configured to analyze the actual state of a task based on the multimodal fusion features, wherein the actual state of a task includes at least the task execution progress. The prediction module, connected to the multimodal fusion module, is configured to perform at least one prediction task based on the multimodal fusion features to obtain a prediction result; the prediction task includes at least one of: order probability prediction, performance amount prediction, and customer churn risk assessment. And, the server, communicating with the edge processing units of each local terminal, including: An incremental data receiving interface is configured to receive incremental data uploaded by the edge processing unit; the incremental data includes at least: information on changes in the actual state of the task and information on updates to the prediction results; The real-time dashboard service module is connected to the incremental data receiving interface and is configured to perform real-time aggregation and statistics on local historical data and received incremental data, obtain statistical data, and push it to a preset administrator terminal. The management terminal displays sales performance data through a visual interface based on the statistical data.
[0026] This embodiment adopts a layered architecture of "real-time processing at the edge + aggregation and statistics on the server side", and its design considerations are as follows: First, sales behavior data involves customer privacy and is subject to strict compliance requirements. Customer call recordings, chat logs, and other similar content are highly sensitive information. Uploading raw data to the cloud not only poses a risk of data leakage but may also violate regulations such as GDPR and the Personal Data Protection Act. For heavily regulated industries such as finance, education, and healthcare, data export or uploading to third-party cloud platforms is often strictly restricted or even prohibited. Therefore, this embodiment retains all multimodal raw data for processing on the local terminal, uploading only anonymized incremental data, thus ensuring data privacy and security from an architectural perspective.
[0027] Secondly, sales management requires a holistic view and statistical analysis. While edge processing units can provide real-time individual guidance to sales personnel, managers need to grasp the overall sales dynamics at the team level, including team conversion rate trends, individual performance rankings, and the distribution of abnormal tasks. These statistical tasks do not require millisecond-level real-time performance, but they do require strong data aggregation capabilities. Therefore, this embodiment deploys statistical, aggregation, and display tasks on the server, forming a reasonable division of labor: "edge processing for individual guidance, server processing for team management."
[0028] The edge processing unit 100 is deployed on the local terminals of each salesperson. The local terminal can be a smartphone, tablet, AI smart badge, smart voice recorder, or wearable device specifically for sales. The edge processing unit 100 is deployed on the local terminal in the form of an SDK (Software Development Kit) or an APP built-in module, and runs automatically after the terminal is powered on, without requiring manual startup or configuration by the salesperson.
[0029] The core function of the edge processing unit 100 is to collect multimodal raw data in real time, complete feature extraction, multimodal fusion, task status analysis and prediction tasks locally, output analysis results and prediction results in real time, and upload incremental data to the server 200.
[0030] Server 200 can be one or more physical servers, cloud servers, or virtual servers, deployed on an enterprise intranet or public cloud platform. Server 200 establishes communication connections with each edge processing unit 100 via the Internet or the enterprise intranet. The core function of server 200 is to receive incremental data uploaded by each edge processing unit 100, aggregate and statistically analyze it with locally stored historical data, generate team-level statistical data, and push it to the administrator terminal 300.
[0031] The administrator terminal 300 can be a manager's personal computer, tablet, or smartphone. The administrator terminal 300 communicates with the server 200 and accesses real-time dashboard services via a browser or dedicated app. The core function of the administrator terminal 300 is to display sales performance data through a visual interface based on statistical data pushed by the server 200. It supports managers in filtering and drilling down into data by time, personnel, and team dimensions, assisting them in making objective sales management decisions.
[0032] Furthermore, the edge processing unit 100 includes a data acquisition module 110. The data acquisition module 110 is configured to acquire multimodal raw data generated by sales personnel during customer follow-up in real time.
[0033] The data acquisition module 110 adopts a real-time automatic data acquisition method, unlike traditional CRM systems that require sales personnel to manually enter data. Specifically: First, the data acquisition module 110 automatically starts after the local terminal is powered on and runs continuously in the background without requiring manual triggering or intervention from sales personnel. Sales personnel communicate with customers according to their daily workflow, and the data acquisition module 110 automatically collects various types of data during the process. Sales personnel do not need to spend extra time filling in follow-up records, updating task status, or uploading communication summaries after the communication ends, saving 30-60 minutes of data entry time per day.
[0034] Second, the data acquisition module 110 adopts an event-driven acquisition strategy, collecting data only when specific events occur, rather than continuously recording indiscriminately, to reduce resource consumption. For example, when the system detects a call is connected, it starts collecting voice data; when it detects a call is disconnected, it stops collecting data. When it detects that a WeChat or Enterprise WeChat chat window is active, it starts collecting text data; when the window is closed or there is no activity for a long time, it stops collecting data.
[0035] Third, the data acquisition module 110 uses a circular buffer for temporary storage. The size of the circular buffer is configurable, and by default it retains the most recent 30 minutes of multimodal raw data. When the buffer is full, the oldest data is automatically overwritten. This design ensures the data window required for real-time analysis while avoiding storage space exhaustion caused by unlimited storage.
[0036] Definition of multimodal raw data The multimodal raw data includes the following four modalities: (a) Text modal data Text modal data refers to textual information generated during sales and customer communication. Specifically, it includes: Instant messaging text: Chat messages sent and received through instant messaging tools such as WeChat Work and WeChat. The data acquisition module 110 obtains the text content of the chat interface through system API or accessibility services, including the message sender (salesperson or customer), message content, and message timestamp.
[0037] CRM Notes Text: Text content such as customer follow-up records, customer needs descriptions, and customer feedback entered by sales personnel in the CRM system. Data acquisition module 110 obtains this text data through the CRM system's open interface or direct database connection.
[0038] Email text: The content of emails exchanged between sales and customers. Data acquisition module 110 obtains the text content of email bodies and attachments through the email system's IMAP protocol or API interface.
[0039] The value of text modal data lies in the fact that it directly reflects the content of information exchange between sales and customers, including sales tactics, customer needs, objections, and purchase intention signals. It is a core data source for analyzing the quality of sales communication.
[0040] (ii) Speech modal data Voice modal data refers to audio information generated during sales and customer communication. Specifically, it includes: Call recording: Two-way call audio generated through the local terminal's telephone system or VoIP calling software. After obtaining the necessary permissions, the data acquisition module 110 acquires the raw audio stream of the call through the system's audio interface.
[0041] Voice messages: Voice messages sent and received via instant messaging tools. The data acquisition module 110 obtains the audio files of the voice messages through the system API.
[0042] The value of speech modality data lies in its inclusion of paralinguistic information lost during text conversion, including tone of voice, speech rate, emotional state (e.g., positive, negative, hesitant, excited), pause patterns (e.g., thinking pauses, hesitant pauses), interruptions, and overlaps. This paralinguistic information is crucial for determining a customer's true intentions and the effectiveness of sales communication. For example, when a customer says "I'll think about it," a relaxed tone might simply be a polite excuse, while a heavy tone could indicate genuine concern.
[0043] (iii) Structured behavioral modal data Structured behavioral modality data refers to structured operation records generated by sales personnel in CRM systems or sales tools. Specifically, it includes: Customer tags: Tag information added by sales staff to customers, such as "high-intent customer", "price-sensitive", "complex decision chain", etc.
[0044] Order status: The current status of a customer's order, such as "Pending payment", "Paid", "Shipped", "Completed", "Cancelled", etc.
[0045] Product type: Information such as the product category, specifications, and quantity that customers are interested in or purchasing.
[0046] Follow-up methods: The follow-up methods used by sales personnel, such as "telephone follow-up", "in-person visit", "online demonstration", etc.
[0047] Task type: The type of task created or completed by sales personnel, such as "demand mining", "product introduction", "quotation sending", "contract signing", etc.
[0048] The value of structured behavioral modal data lies in its ability to record key business information during the sales process in a structured format, facilitating statistical analysis and rule-based judgment. Compared to unstructured data such as text and voice, structured data has lower processing costs, higher interpretability, and provides clear foundational features for multimodal fusion.
[0049] (iv) Temporal behavioral modal data Temporal behavioral modal data refers to the sequential information generated by the changes in sales personnel's behavior over time. Specifically, it includes: Response time sequence: The sequence of time intervals between when a customer sends a message and when the salesperson replies. Response time reflects the salesperson's responsiveness and the customer's level of importance; excessively long response times may lead to customer churn.
[0050] Follow-up frequency sequence: The sequence of the number of times a salesperson follows up with a specific customer within a unit of time (e.g., daily, weekly). Follow-up frequency reflects the intensity of sales follow-up and customer activity; too high or too low a follow-up frequency may affect conversion rates.
[0051] Operation timestamp sequence: The timestamp sequence of various operations performed by sales personnel in the CRM system (such as logging in, viewing customers, editing information, and creating tasks). The operation timestamp sequence reflects the salesperson's work rhythm and time allocation pattern.
[0052] Communication round sequence: The sequence of time and content length for each round of dialogue between sales and customer (sales statement + customer response). The communication round sequence reflects the depth of the dialogue and the quality of the interaction.
[0053] The value of time-series behavioral modal data lies in its ability to capture dynamic patterns of sales behavior, rather than simply providing static statistics. For example, a gradually lengthening response time may indicate a loss of confidence in the customer or a decline in customer interest; a sudden increase in follow-up frequency may suggest that the sales team has identified an urgent closing opportunity. These time-series patterns are invaluable for predicting customer churn risk and the probability of closing a sale.
[0054] The acquisition frequency of the data acquisition module 110 varies depending on the modality type: For speech modal data, a continuous acquisition mode is used, with a sampling rate of 16kHz to ensure the quality of speech recognition.
[0055] For text modal data, an event-triggered mode is used, and data collection is triggered each time a new message arrives.
[0056] For structured behavioral modal data, a change-triggered mode is adopted, and data is collected only when the data changes.
[0057] For time-series behavioral modal data, a timed sampling mode is adopted, recording the current state every 10 seconds to form a time-series sequence.
[0058] By employing this differentiated acquisition strategy, the data acquisition module 110 ensures data integrity while controlling the consumption of system resources by data acquisition.
[0059] The edge processing unit 100 further includes a multimodal encoder group 120, which is connected to the data acquisition module 110. The multimodal encoder group 120 is configured to extract features from each modal data and generate corresponding modal feature vectors.
[0060] The design principle of the multimodal encoder group 120 is as follows: data of different modalities have different data structures, statistical characteristics, and semantic representations, requiring specially designed encoders for feature extraction in order to effectively convert the raw data into vector representations that computers can understand and process. At the same time, considering the limited computing resources at the edge, each encoder adopts a lightweight design to minimize computational load and memory consumption while ensuring feature representation capabilities.
[0061] The following sections describe the specific implementation methods of four modal encoders.
[0062] A text encoder is used to convert text modal data into text feature vectors. This embodiment uses the lightweight pre-trained language model TinyBERT as the core of the text encoder.
[0063] The reasons for choosing TinyBERT are as follows: While large-scale pre-trained language models like BERT perform excellently on natural language understanding tasks, their parameter count is enormous (BERT-base has approximately 110M parameters and a model file of approximately 400MB), making them unsuitable for edge computing. TinyBERT, through knowledge distillation, compresses BERT's knowledge into a smaller network, with approximately 4-6M parameters and a model file of approximately 15-20MB. This allows it to run in real-time on edge devices such as mobile phones, while retaining over 90% of BERT's semantic understanding capabilities.
[0064] The training process of TinyBERT is as follows: First, pre-training is performed on general corpora. General corpora such as BookCorpus and English Wikipedia are used to learn general language representation capabilities.
[0065] Second, fine-tuning was performed on the sales-related corpus. A large number of sales call records and chat logs (after anonymization) were collected to build a sales-related corpus. Based on the general pre-trained model, further pre-training was performed using the sales-related corpus, enabling the model to learn specific terms in the sales field (such as "closing," "objection handling," and "value delivery"), communication patterns, and customer feedback patterns.
[0066] Third, fine-tune the downstream tasks. For the specific tasks in this embodiment (speech classification, objection identification, purchase intention identification, etc.), fine-tune the model using labeled data so that the model's output features can better serve subsequent analysis and prediction tasks.
[0067] The input processing flow of the text encoder is as follows: First, text preprocessing. This involves cleaning the raw text, including removing special characters and emojis, standardizing English letter case, and processing URLs and phone numbers.
[0068] Second, tokenization. The text is segmented into tokens using the TinyBERT-compatible tokenizer. The maximum length of the token sequence is set to truncate any lengths exceeding this limit and padding is used for any short lengths.
[0069] Third, encoding. The word sequence is input into the TinyBERT network, processed by a 12-layer Transformer encoder (TinyBERT-base has 12 layers, TinyBERT-small has 4 layers), and outputs a 768-dimensional feature vector for each word.
[0070] Fourth, pooling. Pooling is performed on the word-level feature vectors to obtain sentence-level feature vectors. Pooling methods can be chosen from the output of the CLS-tagged feature vectors (the output of the [CLS] tag in BERT is typically used for classification tasks) or average pooling (averaging the features of all words). This embodiment uses the output of the CLS-tagged feature vectors as the text feature vectors. It has 768 dimensions.
[0071] A speech encoder is used to convert speech modal data into speech feature vectors. The speech encoder in this embodiment includes two sub-modules: a speech preprocessing encoder and a speech feature extractor.
[0072] The function of a speech preprocessing encoder is to preprocess the raw speech stream, identifying and removing low-information speech segments. Low-information speech refers to speech segments that contribute little or no to sales performance prediction and real-time intervention. Specifically, this includes: greetings and social openings (such as "Hello," "Good morning," "How have you been lately," filler words and verbal tics (such as "um," "uh," "that"), silences and pauses, repetitive information, and non-business-related information.
[0073] The processing flow of the speech preprocessing encoder is as follows: First, speech activity detection. A lightweight VAD model (such as WebRTC VAD) is used to identify valid speech segments and silence / pause segments in the speech stream. Silence / pause segments are directly marked as low-information speech.
[0074] Second, speech-to-text conversion. A lightweight edge-based ASR model is used to convert effective speech segments into text. The ASR model adopts an RNN-T or Conformer architecture, and the model size after quantization is approximately 10-15MB.
[0075] Third, text filtering. The transcribed text is subjected to keyword matching, semantic deduplication, and topic classification to identify and remove low-information speech segments.
[0076] By processing the speech preprocessor encoder, approximately 50-70% of low-information speech can be eliminated, significantly reducing the computational load on the subsequent speech feature extractor.
[0077] The function of a speech feature extractor is to extract acoustic features from valid speech segments. This embodiment uses an architecture combining MFCC (Mel-frequency cepstral coefficients) and TinyLSTM.
[0078] The MFCC feature extraction process is as follows: First, pre-emphasis. This involves using a high-pass filter to boost the high-frequency components of the speech signal, compensating for the effects of glottal pulses and lip radiation.
[0079] Second, framing and windowing. The speech signal is divided into short frames of 20-40ms, with a frame shift typically of 10ms. A Hamming window is applied to each frame to reduce spectral leakage.
[0080] Third, Fourier transform. Perform a fast Fourier transform on each frame to obtain the spectrum.
[0081] Fourth, the Mel filter bank. This simulates the human ear's perception of different frequencies by passing the spectrum through a set of triangular bandpass filters (Mel filter bank). The number of filters is typically set to 40 or 80.
[0082] Fifth, logarithmic operations and discrete cosine transform. Take the logarithm of the filter bank output and then perform a discrete cosine transform to obtain the MFCC coefficients. Usually, the first 13 coefficients are taken as features.
[0083] Finally, each frame of speech is converted into a 13-dimensional or 40-dimensional MFCC feature vector. For a speech segment of length T frames, a feature sequence is obtained; The time series modeling process is as follows: MFCC feature sequences are time-series data, and their dynamic change patterns need to be captured through time-series models. This embodiment uses TinyLSTM as the time-series modeling module. TinyLSTM is a lightweight version of the standard LSTM, reducing computational cost by decreasing the hidden layer dimension and the number of gate units.
[0084] The speech feature extractor also outputs auxiliary features, including: Sentiment classification: By connecting a softmax classifier to a fully connected layer, the probability distribution of three sentiments—positive, negative, and neutral—is output.
[0085] Customer intent classification: Output the probability distribution of intents such as interest, hesitation, rejection, and no clear intent.
[0086] The output of the speech feature extractor is denoted as With auxiliary features, the total dimension is approximately 200.
[0087] Structured behavior encoders are used to convert structured behavior modal data into structured feature vectors.
[0088] The characteristics of structured behavioral data are: the data consists of multiple fields, each of which may be categorical (such as customer tags, product types) or numerical (such as order amount, number of follow-ups). For categorical fields, an embedding layer is used for encoding; for numerical fields, normalization is performed before direct use.
[0089] Taking customer tags as an example, suppose there are 500 possible customer tags. The structured behavior encoder maintains a 500×32500×32 embedding matrix, mapping the one-hot encoding (500-dimensional) of customer tags to a 32-dimensional dense vector. The initial values of the embedding matrix are randomized and learned end-to-end with other modules during model training, so that the vector representation of the tags can reflect the semantic similarity between tags (e.g., the embedding vectors of "high-intent customer" and "imminent customer" are close).
[0090] Taking product types as an example, suppose there are 100 product types in total. Similarly, using a 100×16 embedding matrix, the product types are encoded as 16-dimensional vectors.
[0091] For numeric fields, such as order amounts (range 0-100,000 yuan), maximum and minimum normalization is used to map the values to the [0,1] interval. The same normalization process is applied to count data such as the number of follow-ups.
[0092] The encoded results of all fields are concatenated to obtain a structured feature vector. In this embodiment, the structured feature vector is designed to have 128 dimensions.
[0093] The output of the structured behavior encoder is denoted as .
[0094] A temporal behavior encoder is used to convert temporal behavior modal data into temporal feature vectors.
[0095] The characteristics of temporal behavioral data are: the data is a sequence that changes over time and has dynamic evolutionary properties. This embodiment uses a temporal convolutional network (TCN) as the core of the temporal behavioral encoder.
[0096] The reasons for choosing TCN instead of LSTM are as follows: TCN uses a convolutional structure, which allows for highly parallel computation and inference speeds 2-5 times faster than LSTM; TCN obtains an exponentially larger receptive field through dilated convolution, enabling it to capture long-term dependencies; TCN has shorter gradient paths, does not suffer from the gradient vanishing problem, and has more stable training.
[0097] The core components of TCN include: First, causal convolution. Causal convolution ensures that the model's output at time step t depends only on the input at time step t and before, and does not "see" information from the future. This is a necessary constraint for time series prediction tasks.
[0098] Second, dilated convolution. Dilated convolution expands the receptive field exponentially without increasing the number of parameters by inserting holes between the elements of the convolution kernel (i.e., skipping parts of the input). For a sequence of length L, using stacked dilated convolutional layers, the receptive field size increases exponentially with the number of layers.
[0099] Third, residual connections. Each TCN module consists of two dilated convolutional layers. Each convolutional layer is followed by weight normalization and a ReLU activation function, and then residual connections are added to directly add the input to the output. Residual connections alleviate the gradient vanishing problem in deep networks, allowing TCNs to be stacked to a relatively deep number of layers.
[0100] In this embodiment, the input to the time-series behavior encoder is multiple time-series sequences, including response duration sequences, follow-up frequency sequences, operation timestamp sequences, etc. Each time-series sequence is encoded independently, and then the encoded results are concatenated.
[0101] Taking a response time sequence as an example: Assume the response time (in minutes) of the most recent 10 customer messages is recorded, and the sequence length is 10. The sequence is input into a TCN network, which contains 4 dilated convolutional layers with dilation rates of 1, 2, 4, and 8, and a receptive field size of [missing value]. The sequence length exceeds 10, thus capturing the complete temporal dependencies. After TCN encoding, a 256-dimensional temporal feature vector is output. For the follow-up frequency sequence and the operation timestamp sequence, the same TCN structure is used for independent encoding, each outputting a 256-dimensional feature. Then, the encoding results of the three sequences are concatenated to obtain a 768-dimensional temporal feature vector.
[0102] Considering the resource constraints at the edge, the time-series behavior encoder undergoes the following lightweighting process: Reduce the number of TCN channels from 128 to 64; Reduce the number of TCN layers from 6 to 4; The weights are quantized using INT8.
[0103] The lightweight temporal behavior encoder has approximately 1MB of parameters, a model file of approximately 3-5MB, and a single inference time of approximately 5-10ms.
[0104] The output of the timing behavior encoder is denoted as After being lightweighted, the dimensions can be reduced to 256.
[0105] The four encoders in the multimodal encoder group 120 operate in parallel, each processing its corresponding modal data. Since different modalities have different data generation frequencies and processing times, the encoder group employs an asynchronous processing mechanism: each encoder runs independently, and after a particular encoder completes feature extraction, it stores the feature vector in a shared feature buffer for use by the multimodal fusion module 130. If data for a particular modality is temporarily missing (e.g., no voice data in a non-call scenario), the corresponding encoder outputs a zero vector, and the multimodal fusion module 130 automatically adjusts the attention weights to reduce the impact of the missing modality.
[0106] The edge processing unit 100 further includes a multimodal fusion module 130, which is connected to the multimodal encoder group 120. The multimodal fusion module 130 is configured to fuse the feature vectors of each modality through a cross-modal attention mechanism to generate a unified multimodal fusion feature.
[0107] The necessity of multimodal fusion Data from different modalities describes the same sales communication process from different perspectives, and there is a wealth of complementary and correlated information among them. For example: When a customer says "I'll think about it some more," if the voice emotion recognition is positive, it may mean that the customer is interested in the product but needs time; if the voice emotion recognition is negative, it may mean that the customer has concerns about the price or the product and needs further investigation to uncover their objections.
[0108] There is a temporal correlation between the type of sales pitch (text modality) and changes in customer emotions (voice modality): whether the customer's emotions improve or worsen when a salesperson uses a certain pitch reflects the effectiveness of the pitch.
[0109] There is a correlation between sales actions (structured modality) and response time (temporal modality): a quick reply after viewing customer details in the CRM indicates that the salesperson values the customer highly; a long delay in replying after viewing may mean that the salesperson has encountered difficult information and needs to think about a response strategy.
[0110] The goal of the multimodal fusion module 130 is to capture these cross-modal correlations and generate a comprehensive fusion feature, providing a richer and more accurate information foundation for subsequent analysis and prediction.
[0111] Cross-modal attention mechanism This embodiment employs a cross-modal attention mechanism to achieve multimodal fusion. The core idea of the attention mechanism is to dynamically focus on the parts of other modalities that are most relevant to a particular modality when processing features of that modality.
[0112] Specifically, this embodiment uses text modality features as the query and speech modality features and temporal modality features as the key and value. The reason for choosing text modality as the query is as follows: text modality contains the core semantic information of sales communication, while other modalities (speech, temporal) are more of a supplement and modification to text modality. Using text as the anchor for cross-modal alignment is more in line with the semantic structure of sales communication. A cross-modal attention mechanism is adopted; the calculation process of the attention mechanism can be understood as follows: for each semantic unit in the text features, its relevance score with each part of the speech and temporal features is calculated (using QKT), and then the speech and temporal features are weighted and summed according to the relevance scores (using softmax and V) to obtain the fused features.
[0113] Specifically, to capture cross-modal associations across different dimensions, this embodiment employs a multi-head attention mechanism. Multi-head attention uses multiple sets of independent attention parameters (multiple sets of WQ, WK, WV), each set referred to as an "attention head." Each attention head focuses on different association patterns, for example: The first point of attention might focus on the association between dissenting keywords in the text and negative emotions in the speech; The second point of attention might focus on the connection between the closing statements in the text and the customer's excitement in the voice; The third point of focus might be the correlation between product descriptions in the text and response times in the timeline.
[0114] Each attention head outputs a 128-dimensional fused feature. The outputs of multiple attention heads are concatenated to obtain a multi-head fused feature. This embodiment uses four attention heads, resulting in a 512-dimensional feature after concatenation.
[0115] In addition to fusing text, speech, and temporal modalities through cross-modal attention, the multimodal fusion module 130 also needs to incorporate structured features into the fusion process. Structured features differ in nature from the other three: the former three are sequential data (text is a sequence of words, speech is a sequence of frames, and temporal data is a sequence of events), while structured features are static key-value pairs. Therefore, the fusion of structured features employs a different approach.
[0116] This embodiment employs a gated fusion mechanism to merge structured features and attention-based fusion features. The core idea of this mechanism is to allow the model to dynamically learn the appropriate weights for attention-based and structured features within the multimodal fusion feature set. For example, when structured features (such as customer tags like "high-intent customer") have strong predictive power, the gate value g will be smaller, giving structured features a higher weight; conversely, when structured features lack sufficient information, the gate value g will be larger, giving attention-based fusion features a higher weight.
[0117] After the above processing, the unified multimodal fusion feature output by the multimodal fusion module has the following characteristics: First, information completeness. The fusion feature integrates information from four modalities: text, speech, structured, and temporal. Information loss from a single modality will not lead to severe degradation of the overall feature. Second, cross-modal alignment. Through a cross-modal attention mechanism, the information from different modalities in the fusion feature has been aligned at the semantic level, making it easy for subsequent analysis and prediction modules to use directly. Third, moderate dimensionality. The 256-dimensional fusion feature achieves a balance between expressive power and computational efficiency, retaining sufficient information without causing excessive computational burden on subsequent modules. Fourth, lightweight design. The multimodal fusion module 130 itself is also lightweight, with approximately 0.5M parameters and a single fusion time of approximately 10-20ms.
[0118] The edge processing unit 100 further includes an analysis module 140, which is connected to the multimodal fusion module 130. The analysis module 140 is configured to analyze the actual state of the task based on multimodal fusion features.
[0119] The actual status of the task includes at least the task execution progress. Task execution progress refers to the current stage of the salesperson's customer follow-up process and the percentage of work completed. The task of analysis module 140 is to determine, based on multimodal fusion features, which stage the current task is in, and which actions have been completed and which actions remain in that stage.
[0120] The analysis module 140 automatically analyzes the task status, eliminating the need for sales personnel to manually update task progress or for managers to rely on experience to judge the task's stage. The analysis results, based on multimodal fusion features and deep learning models, are objective and consistent, unaffected by subjective judgment. The task execution progress reflects the salesperson's current position in the customer follow-up process. This embodiment divides the sales follow-up process into seven standard stages: early concept phase, late concept phase, needs assessment phase, product introduction phase, objection handling phase, closing phase, and after-sales follow-up phase. These seven stages are arranged in the natural order of sales progress, forming a complete path from initial customer contact to final transaction and after-sales service.
[0121] The analysis module's task is to determine which stage the current task is in and how much work has been completed within that stage. Specifically, the analysis module outputs two core metrics: the current stage name and the percentage of progress completed in that stage.
[0122] The current stage name is determined using a stage classification model. This model maps multimodal fusion features to seven stage categories, outputs the probability distribution for each stage, and takes the stage with the highest probability as the determination result. For example, when the system detects that a salesperson is introducing product features to a customer and answering product-related questions, it will determine that the current stage is the "product introduction stage"; when it detects that a customer is raising questions about price or service and the salesperson is explaining and persuading, it will determine that the current stage is the "objection handling stage".
[0123] The percentage of progress completed in a phase is a more granular indicator. Within the same phase, the actual progress of different tasks may differ. For example, two salespeople may both be in the "objection handling phase," but one might have just begun handling their first objection, while the other may have already resolved all objections and is about to enter the closing phase. The analysis module outputs a progress value between 0 and 100 through a regression model, reflecting the completion rate within the current phase. Combining the progress value with phase information provides a more precise description of the actual status of the task—for example, "objection handling phase, progress 65%" provides more information than simply stating "objection handling phase."
[0124] This automated analysis does not rely on manual updates from sales personnel; it is entirely based on multimodal fusion features for judgment. The dialogue content in speech, the semantic information in text, and the operation records in behavior together provide rich evidence for stage judgment and progress assessment.
[0125] The prediction module 150 enables objective prediction of sales results without relying on the subjective experience and judgment of sales or management personnel. The prediction results are based on multimodal fusion features and deep learning models, featuring data-driven and objective quantification, avoiding the subjectivity and inconsistency of human experience judgment.
[0126] The prediction module, also based on multimodal fusion features, performs the task of predicting future sales results. This embodiment supports three prediction tasks: order probability prediction, sales amount prediction, and customer churn risk assessment. These three tasks share the underlying multimodal fusion features and are jointly trained using a multi-task learning architecture, allowing for the simultaneous output of multiple prediction results.
[0127] Sales conversion probability prediction refers to forecasting the likelihood of a customer making a purchase within the current follow-up period. The prediction module outputs a probability value between 0 and 1, with higher values indicating a greater probability of a sale. This probability is calculated based on a combination of factors: the level of intent reflected in the customer profile, the quality of communication and interaction during the communication process, follow-up frequency, and response speed. Sales conversion probability helps salespeople determine which customers deserve priority in their efforts. For example, if a customer's sales conversion probability increases from 0.6 to 0.8, it indicates that the current sales communication strategy is effective and can be maintained; if the probability continues to decline, it is necessary to adjust the communication approach promptly or seek management assistance.
[0128] Sales revenue forecasting refers to predicting the potential transaction amount from a current customer. Unlike the probability of closing a deal, sales revenue forecasting outputs a specific monetary value. This forecast is based on the customer's historical spending power, the scale of demand and budget range demonstrated in current communication, and the distribution of transaction amounts among similar customers. For salespeople simultaneously working with multiple customers, sales revenue forecasting can help them allocate time and effort effectively—prioritizing customers with high expected transaction amounts and a high probability of closing a deal. For managers, the aggregated value of all sales revenue forecasts can serve as a reference for team performance forecasting.
[0129] Customer churn risk assessment predicts the probability that a current customer will churn within a certain period (default 7 days). Customer churn is defined as: a customer ceasing to respond to sales messages, explicitly stating they will no longer consider purchasing, or switching to a competitor. The churn risk assessment outputs a probability between 0 and 1, along with a risk level (high / medium / low). Unlike sales conversion probability, churn risk assessment focuses more on identifying "red flags," such as decreased customer response frequency, increased negative emotions, making unreasonable demands, and actively inquiring about competitor information. When the churn risk is high, the system will prompt sales to intervene promptly and take retention measures.
[0130] The server acts as a central data aggregation hub. Each salesperson's local terminal's edge processing unit establishes a persistent connection with the server, uploading incremental data to the server in real time. Upon receiving this data, the server aggregates, statistically analyzes, and stores it, then pushes the statistical data to the administrator's terminal via a real-time dashboard service module.
[0131] This three-tier architecture of "edge-server-manager" has the following advantages: First, the data transmitted between the edge processing unit and the server is incremental rather than raw data, and the data volume is extremely small (a single upload is usually less than 1KB), which reduces the network bandwidth requirements; second, the server undertakes the calculation tasks of data aggregation and statistics, while the manager terminal only needs to be responsible for display, which reduces the computational burden on the manager terminal; third, all data is stored uniformly on the server, which facilitates historical query and trend analysis.
[0132] The incremental data receiving interface serves as the entry point for communication between the server and each edge processing unit. This interface uses the HTTP / 2 or WebSocket protocol and supports high-concurrency, low-latency data reception.
[0133] The incremental data uploaded by the edge processing unit has an "incremental" characteristic, meaning that only information that has changed is uploaded, rather than the entire dataset. Specifically, incremental data includes two types of information: The first category is information about changes in the actual status of the task. When the analysis module detects a change in the task status (e.g., moving from the "product introduction phase" to the "objection handling phase," or updating the stage completion progress from 45% to 52%), the edge processing unit packages and uploads this change information. The uploaded content includes: task identifier, change time, status before the change, and status after the change. Because only the change is uploaded, rather than the full status, the data volume is extremely small.
[0134] The second category is update information for prediction results. When the output of the prediction module changes significantly (e.g., the order probability changes from 0.6 to 0.75, or the churn risk level changes from low to medium), the edge processing unit uploads the update information. The uploaded content includes: task identifier, update time, prediction metric name, value before update, and value after update. The system can be configured with a threshold for "significant change"—for example, only triggering an upload when the order probability changes by more than 0.05, avoiding unnecessary network overhead caused by overly frequent uploads.
[0135] After receiving data, the incremental data receiving interface first performs data verification (checking data format and signature verification), then writes the data to a message queue for asynchronous processing, and finally returns an acknowledgment response to the edge processing unit. This asynchronous processing mechanism ensures the interface's high throughput capacity, preventing congestion even when a large number of edge processing units upload data simultaneously.
[0136] The real-time dashboard service module is the core component of the server for data aggregation and statistics. This module connects to the incremental data receiving interface, continuously consumes incremental data from the message queue, and performs real-time aggregation and statistics.
[0137] The aggregated statistics include calculations across multiple dimensions. In terms of time, the system performs rolling aggregations of data at granularities such as seconds, minutes, hours, and days, allowing managers to view trends at different time granularities. In terms of personnel, the system summarizes data at levels such as individuals, teams, and departments, allowing managers to drill down from macro to micro levels. In terms of metrics, the system calculates key indicators such as predicted total performance (the sum of the probability of all tasks closing a deal × predicted amount), the number of high-risk customers (the number of customers with a churn risk > 0.7), the team's average script standardization score, and the distribution of tasks at each stage.
[0138] The real-time dashboard service module uses a streaming computing framework (such as Apache Flink or a self-developed lightweight streaming computing engine) to achieve real-time aggregation. When incremental data arrives, the system updates the relevant statistical results within milliseconds, ensuring that managers always see the latest data. For example, when a salesperson uploads an update stating that "the probability of closing a deal has increased from 0.6 to 0.8," the system's predicted total performance will immediately increase by the corresponding amount, and managers can see the change without refreshing the page.
[0139] In addition to real-time aggregation, the real-time dashboard service module is also responsible for pushing statistical data to the administrator's terminal. The push method uses a WebSocket long connection; the server proactively pushes updates to the administrator's terminal, and the administrator's terminal automatically refreshes the interface. This "server-proactive push" mode has lower latency and consumes fewer resources than the traditional "client-side timed fetch" mode.
[0140] The administrator terminal displays sales performance data through a visual interface based on server-pushed statistical data. The administrator terminal can be a PC-based web application or a mobile app, allowing administrators to choose flexibly according to their usage scenarios.
[0141] In this embodiment, the server and the aforementioned edge processing unit form a clear division of labor: the edge processing unit is responsible for real-time analysis and prediction of individuals, while the server is responsible for the aggregation, statistics, and display of the team; the edge processing unit processes real-time data at the millisecond level, while the server processes aggregated data at the second to minute level; the output of the edge processing unit serves the salesperson, while the output of the server serves the managers and the team.
[0142] This division of labor fully leverages the advantages of both ends: the edge, being close to the data source, enables the lowest possible latency response; the server, with its stronger computing power and larger storage capacity, is capable of performing complex aggregation analysis and long-term data storage. Working together, they form a complete sales intelligence analysis and early warning system.
[0143] The server and its real-time dashboard service address the issue of "data analysis relying on human experience and being inefficient" mentioned in the background technology. Managers no longer need to ask each salesperson "how's the follow-up going?" or guess team performance based on experience. The server automatically aggregates data from each edge processing unit, providing managers with an objective, real-time, and comprehensive view of the team through real-time aggregation statistics and visualization. This data-driven management approach is more efficient, accurate, and traceable than traditional manual management that relies on personal experience.
[0144] In some embodiments, the multimodal raw data includes data for different tasks; the edge processing unit further includes a priority scheduling module configured to prioritize processing data corresponding to high-priority tasks according to a preset task priority configuration.
[0145] In real-world sales scenarios, salespeople typically follow up with multiple clients and handle multiple tasks simultaneously. The importance and urgency of these different tasks vary significantly. For example: Customers in the "closing phase" may signal a deal or raise objections at any time, requiring immediate response from sales. If the system fails to analyze the voice and text data of this task in real time, sales may miss closing opportunities or fail to resolve objections in a timely manner.
[0146] Clients in the "early concept stage" are still in the initial understanding phase and have relatively low requirements for the real-time nature of analysis and processing. Even if the system's analysis results are delayed by a few seconds or even tens of seconds, the impact on the final transaction is far less than that of the tasks in the closing stage.
[0147] However, the computing resources (CPU, memory, NPU) of the edge processing unit 100 are limited. When sales personnel simultaneously follow up with multiple customers, or when a large amount of multimodal data is generated in a short period of time, the edge processing unit 100 may face a shortage of computing resources. If a first-in-first-out equal processing strategy is adopted for all tasks' data, the following problems may occur: Data for high-priority tasks is queued in long queues, significantly increasing processing latency and making it impossible to provide real-time alerts and guidance at critical moments. Low-priority tasks consume a lot of computing resources, which leads to the processing resources of high-priority tasks being squeezed out. In extreme cases, exhaustion of computing resources can cause system lag or crashes, preventing all tasks from being processed in a timely manner.
[0148] Therefore, a priority scheduling mechanism is needed to ensure that high-priority tasks are processed first when computing resources are limited, thus guaranteeing the real-time requirements of critical sales scenarios.
[0149] The priority scheduling module is deployed between the data acquisition module 110 and the multimodal encoder group 120, serving as the front-end scheduler for the data flow. Its core functions include: task identification, priority mapping, queue management, and resource allocation.
[0150] When collecting data, the data acquisition module 110 records the session context information of the data. For voice modal data, the corresponding customer is identified by the phone number of the caller or the VoIP session ID; for text modal data, the corresponding customer is identified by the session ID of the chat window or the contact information; for structured behavioral data, the corresponding customer is identified by the customer ID in the CRM record. Each customer corresponds to an independent sales task.
[0151] For data that cannot be directly linked to a specific customer (such as system-level operation logs), the priority scheduling module aggregates data based on time windows. For example, all data generated within the past 5 minutes is grouped into the same task batch, and the priority of the entire batch is determined based on the highest priority of the identifiable task within that batch.
[0152] The salesperson's terminal interface provides a manual task priority marking function. Salespeople can mark currently being processed customers as "high priority" or "top," and the priority scheduling module adjusts the data processing priority of the task based on the salesperson's manual marking.
[0153] The priority scheduling module maintains a priority mapping table internally, mapping various attributes of sales tasks to specific priority levels. In this embodiment, the priorities are divided into four levels, from highest to lowest: P0 (highest), P1 (high), P2 (medium), and P3 (low).
[0154] The priority mapping table supports dynamic adjustment. The system can automatically increase or decrease task priority based on real-time analysis results. For example: When the analysis module 140 detects that the probability of customer churn exceeds 0.7, it automatically raises the priority of the task to P1.
[0155] When the probability of an order being generated output by the prediction module 150 exceeds 0.8, the priority of the task is automatically raised to P1.
[0156] If the same task fails to generate any data for 10 consecutive minutes, its priority is reduced to P3 to avoid consuming unnecessary resources.
[0157] In addition, the priority mapping table supports manual configuration by administrators. Administrators can use the server's management interface to set custom priority rules for specific sales, customers, or task types. For example, an administrator can set all tasks for "VIP customers" to P0, ensuring that these customers' data always receives the highest priority processing.
[0158] The priority scheduling module maintains four first-in-first-out (FIFO) queues, corresponding to priority levels P0, P1, P2, and P3. After the data acquisition module 110 collects multimodal raw data, the priority scheduling module places the data into the corresponding queue based on its task priority. The core strategy of queue management is strict priority scheduling. That is, data in low-priority queues will only be processed when all high-priority queues are empty. To prevent tasks in low-priority queues from being "starved" (i.e., not processed for a long time), the priority scheduling module also implements an aging mechanism. Specifically: data waiting in queue P3 for more than 30 minutes is promoted to P2; data waiting in queue P2 for more than 15 minutes is promoted to P1; and data waiting in queue P1 for more than 5 minutes is promoted to P0. The time settings are merely illustrative. The aging mechanism ensures that even with a continuous influx of high-priority data, low-priority data will eventually get a chance to be processed, preventing tasks from being completely ignored in extreme cases.
[0159] In some embodiments, the analysis module is further configured to: analyze the difficulty information in the task processing and advancement process; and generate alarm prompt information when the difficulty information meets the preset alarm rules.
[0160] In the sales process, salespeople often encounter various difficult problems that hinder the progress of their tasks. For example: Customer objections: If a customer raises questions or objections regarding price, features, or services, and the sales team fails to handle these issues properly, it may lead to customer loss.
[0161] Customer silence: If a customer does not respond for an extended period of time at a critical juncture (such as after closing the deal), the salesperson may be unable to determine the customer's true intentions and may be forced to wait passively.
[0162] Inappropriate communication techniques: Salespeople used inappropriate techniques (such as over-promising, piling up technical jargon, or failing to address customer concerns), which reduced the effectiveness of communication.
[0163] Delayed follow-up: Sales staff fail to follow up with customers as planned, leading to a decline in customer enthusiasm or the loss of customers to competitors.
[0164] Emotional fluctuations: If the customer exhibits obvious negative emotions (impatience, anger, disappointment), and the salesperson fails to soothe them in time, it may lead to a breakdown in communication.
[0165] Additional requirements: The customer raised additional requirements, and the salesperson may not have the authority to make these requirements.
[0166] It should be noted that there are many difficulties in practical applications, which will not be listed here.
[0167] At this point, the system will determine whether the sales staff can handle these issues based on the preset alarm rules. If it determines that the sales staff cannot handle them, an alarm will be triggered. Furthermore, it also includes: the analysis module is specifically configured as follows: Construct or obtain a sales person profile for a salesperson, the sales person profile including at least: job title tags, ability tags, and historical performance tags; dynamically adjust the threshold parameters of the preset alarm rules based on the sales person profile.
[0168] Salespeople vary significantly in ability, experience, and performance. Using a uniform alarm threshold would inevitably lead to excessive alarms for newcomers and insufficient alarms for experienced salespeople. Therefore, the analytics module creates a dynamically updated salesperson profile for each salesperson.
[0169] Job tags describe the salesperson's job attributes and work background, including job level (intern, junior, intermediate, senior, expert), length of service, team, and business line of responsibility. These tags are usually synchronized from the company's HR system, are relatively stable, and reflect the company's expectations for the sales position.
[0170] Competency tags describe a salesperson's performance level in various sales skills, including scores for sales script skills, objection handling skills, and closing skills. Unlike job title tags, competency tags are automatically calculated by the analysis module based on multimodal fusion features and are updated daily, reflecting dynamic changes in sales capabilities. For example, if a salesperson's sales script skills score improves from 65 to 80, the system will automatically detect this progress.
[0171] Historical performance tags describe a salesperson's past performance, including conversion rate, average order value, number of orders completed, and performance trends over the past 30 days. These tags are synchronized from the CRM system and updated daily, reflecting the actual output level of sales.
[0172] With salesperson profiles, the analysis module can dynamically adjust alarm thresholds based on the characteristics of different salespeople. The core principle of adjustment is: corresponding to the difficulties in standardizing themselves: the more capable, higher-ranking, and better-performing the salesperson, the stricter the threshold (higher requirements) should be applied to address these difficulties; conversely, a relatively lenient threshold should be applied (to protect confidence and proceed gradually). The opposite applies to the difficulties in reassuring customers.
[0173] Taking a standardized sales script alert as an example, the baseline threshold is 60 points. For an intern or a newcomer with a low sales script score, the threshold might be lowered to 50 points—meaning that an alert won't be triggered as long as a score of 50 is reached, preventing newcomers from feeling frustrated due to frequent alerts. However, for a senior salesperson or a high-performing salesperson with a sales script score of 85, the threshold might be raised to 70 points—meaning the company has higher expectations, and scores below 70 require attention and improvement. Taking a delayed response alert as another example, the baseline threshold is 3 minutes. For high-priority tasks or senior salespeople, the threshold might be tightened to 1.5 minutes, requiring faster customer responses; for low-priority tasks or new salespeople, the threshold might be relaxed to 5 minutes, providing more buffer time.
[0174] For example, regarding emotional fluctuations, the baseline threshold is 60 points. For an intern or a newcomer with a low sales skills score, the threshold might be lowered to 50 points—meaning that reaching 50 points would trigger an alarm, alerting relevant personnel to intervene, and the salesperson might not be able to handle the situation. However, for a senior salesperson or a high-performing salesperson with a high sales skills score, the threshold might be raised to 70 points—meaning the company has higher expectations of them; for emotional fluctuations below 70 points, the salesperson can handle the situation independently without assistance.
[0175] Salesperson profiles are not static. Competency tags are updated daily to capture subtle changes in sales abilities; historical performance tags are synchronized daily to reflect the latest performance; while job tags are relatively stable, they are updated promptly when a salesperson is promoted or transferred. As the profile is continuously updated, alert thresholds are also dynamically adjusted. As a salesperson progresses from a newcomer to a senior professional, the alert thresholds gradually transition from lenient to strict, always maintaining a level commensurate with their current abilities and expectations. Managers can also manually adjust certain tags through the system backend, for example, prematurely marking high-performing newcomers as "near-senior" level, providing them with higher expectations and stricter guidance standards.
[0176] This embodiment defines three conditions for triggering profile updates, corresponding to update needs in different scenarios: receiving a preset profile adjustment instruction; sales personnel completing a preset threshold number of tasks; and the time interval since the last profile update reaching a preset duration.
[0177] The first trigger condition is receiving a preset personnel profile adjustment instruction. This type of update is triggered by external input and mainly includes two scenarios: manual adjustments by managers and automatic reception of adjustment instructions by the system. For example, when a new salesperson performs exceptionally well, the manager can adjust their job level label from "Junior" to "Intermediate" in the system backend. Upon receiving this instruction, the system immediately triggers a profile update and simultaneously tightens the alarm threshold. Another example is when the job information of sales personnel changes in the company's HR system; this system automatically receives the adjustment instruction through an interface and updates the profile label accordingly. Manual adjustments by managers have the highest priority and can override the results of automatic calculations because managers possess first-hand information such as performance evaluations and job adjustments.
[0178] The second trigger condition is when the number of tasks performed by the salesperson reaches a preset threshold. This type of update is triggered by the accumulation of tasks, reflecting the principle of "quantitative change leading to qualitative change." Salespeople's skill improvement usually requires sufficient practical experience; without enough customer communication sessions, skill improvement lacks data support. In this embodiment, the preset threshold can be set to 10, 30, or 100 customer communications. Whenever a salesperson reaches these milestones, the system automatically triggers a profile update, reassessing their competency tags based on the latest accumulated data. For example, if a new salesperson completes 15 customer communications in their first week, exceeding the 10-time threshold, the system will trigger an update and find that their sales skills score has increased from 45 to 52. Therefore, the system will automatically tighten the alarm threshold and impose higher requirements on them.
[0179] The third trigger condition: the time interval since the last profile update reaches a preset duration. This type of update is triggered by the passage of time and serves as a fallback mechanism, ensuring that the profile is refreshed regularly even if the sales target is insufficient. In this embodiment, the preset duration can be set to 7 days, 14 days, or 30 days. The role of regular updates is reflected in two aspects: Firstly, for salespeople with a small sales target, updates cannot be triggered by the task quantity threshold, making regular updates the only way to maintain the timeliness of their profiles; secondly, regular updates can capture gradual changes over time. For example, the script standardization score may slowly decrease from 75 points to 68 points in the past 30 days. The changes may not be obvious on a daily basis, but a monthly comparison clearly shows the downward trend, thus allowing for timely detection of problems.
[0180] Three triggering conditions work together; as long as any one of them is met, the system will trigger a profile update and simultaneously adjust the alarm rules. Through this mechanism, the salesperson profile can remain accurate as sales grow and change, and the alarm thresholds will always match the actual performance level of the salesperson.
[0181] In some embodiments, the edge processing unit is a lightweight deployment unit that satisfies at least one of the following characteristics: the multimodal encoder group, multimodal fusion module, analysis module, and prediction module are lightweight models that have undergone model quantization, knowledge distillation, or structured pruning; the edge processing unit operates independently in the offline state of the local terminal and does not depend on a real-time network connection with the server.
[0182] One of the core innovations of this application is deploying sales behavior analysis and early warning capabilities on sales personnel's local terminals (such as smartphones) to achieve millisecond-level real-time response and protect data privacy. However, there is a significant gap between the computing resources of local terminals and cloud servers: the computing power of a mobile phone's CPU is far weaker than that of a server's GPU, memory and storage space are limited, and battery life is also severely restricted. Large-scale deep learning models like BERT and GPT (often hundreds of megabytes or even gigabytes in size) simply cannot run on mobile phones. Even if they are forced to run, it will cause the phone to overheat severely, the battery to drain rapidly, and the application to lag or even crash. Therefore, it is necessary to lightweight compress deep learning models so that they can run smoothly on resource-constrained edge terminals while maintaining sufficient analytical accuracy.
[0183] Taking a speech encoder as an example: First, knowledge distillation is used to compress a large model into a smaller model. Then, structured pruning is used to remove redundant parameters from the smaller model. Finally, INT8 quantization is used to convert the parameters from floating-point numbers to integers. After this series of processes, the original 80MB speech model is finally compressed to 5-8MB, which can run in real time on a mobile device.
[0184] In addition to the lightweight model, another important feature of this embodiment is that the edge processing unit has the ability to operate independently offline, without relying on a real-time network connection with the server.
[0185] In traditional cloud processing models, local terminals collect data and upload it to the cloud, then wait for the cloud to process it before receiving the results. This process is not only highly delayed but also completely dependent on the network—if the network signal is weak or completely offline, the system becomes unusable. When sales personnel communicate with customers in locations with poor signal, such as high-speed trains, subways, and underground parking garages, the system becomes virtually useless.
[0186] This embodiment's lightweight deployment solution completely solves this problem. All analysis and prediction models are deployed on local terminals, and data acquisition, feature extraction, multimodal fusion, state analysis, and result prediction are all completed locally on the terminal. Sales personnel can obtain complete system services from any location and under any network conditions—whether in the office with Wi-Fi or in the subway with a weak signal.
[0187] Offline, independent operation also offers advantages in privacy protection. Since all raw data (call recordings, chat logs, etc.) is processed locally without uploading to the cloud, the risk of data leakage during transmission and storage is fundamentally avoided. This is particularly important for heavily regulated industries such as finance, education, and healthcare.
[0188] Of course, the edge processing unit is not completely isolated from the server. When it detects that the device is idle (such as charging and connected to Wi-Fi), it asynchronously uploads anonymized statistical data and analysis results to the server for real-time dashboard display by administrators and subsequent model optimization. However, even when never connected to the internet, the core analysis and alerting functions of the edge processing unit remain completely unaffected.
[0189] In some embodiments, the administrator terminal is communicatively connected to the edge processing units of each local terminal; the administrator terminal is configured to receive the actual task status and prediction results uploaded by the edge processing units; and provides an interactive visualization interface that supports filtering and drill-down analysis of the actual task status and prediction results by time dimension, personnel dimension, and team dimension.
[0190] Edge processing units provide sales personnel with real-time, personalized analytics and alerts, helping them gain immediate guidance in customer communications. However, sales managers also need to grasp the overall sales dynamics of the team in order to promptly identify team-level issues, distinguish between high-performing and struggling sales, and make informed management decisions.
[0191] The administrator terminal does not communicate directly with the edge processing units of each local terminal; instead, a server acts as an intermediary. The edge processing units upload anonymized incremental data (information on changes in the actual state of tasks and updates to prediction results) to the server. The server then aggregates and statistically analyzes this data before pushing the statistical data back to the administrator terminal. The advantage of this indirect communication model is that the server handles data aggregation and storage, eliminating the need for the administrator terminal to establish direct connections with hundreds or thousands of edge processing units, resulting in a clearer architecture and greater scalability.
[0192] The administrator's terminal can be the administrator's personal computer (accessed via a web browser), tablet, or smartphone (accessed via a dedicated app). After logging into the system with their own account, administrators can only view sales data within their authorized scope (such as data for their own department or team).
[0193] Sometimes the server may disconnect. In this case, the administrator terminal can communicate directly with the edge processing units of each local terminal.
[0194] After obtaining the data processed by the server or the data directly sent by the edge processing units of each local terminal, the administrator terminal can perform visualization display locally: For example, after logging in, managers first see a global dashboard of team sales data. The dashboard displays key metrics in the form of cards, charts, and tables, including: the team's projected total performance (the sum of all sales conversion probabilities × projected amount), the number of high-risk customers (customers with a churn probability > 0.7), the number of pending alarms, the team's average sales script standardization score, and the team's average response time. These metrics are dynamically updated over time; when the edge processing unit uploads new incremental data, the numbers on the dashboard are updated in real time.
[0195] In addition to numerical metrics, the dashboard also displays trend charts. For example, it shows the team's conversion rate over the past 30 days, the salesperson's sales skills ranking, and the distribution of various alerts (such as price objection alerts being the most frequent, followed by response delay alerts). Managers can quickly understand the overall health of the team as soon as they log into the system.
[0196] The global dashboard displays aggregated team-level data, but managers often need to view data from specific ranges. For example, they might only want to see data from this week, not all historical data, or only data from a specific team, not the entire department, or only data from difficult tasks with a success rate of less than 0.3.
[0197] The visual interface provides filters in three dimensions: Time-based filtering: Supports viewing data by daily, weekly, monthly, and quarterly time granularities. Managers can select "Today" to view real-time dynamics, "This Week" to view recent trends, and "This Month" for monthly reviews. The time range also supports customization, such as viewing data from the last 7 days or the last 30 days.
[0198] Personnel-based filtering: Supports viewing data by individual. Managers can select a specific salesperson from the drop-down list, and the interface will switch to that salesperson's dedicated view, displaying their personal profile tags, task status, prediction results, and historical alert records. Personnel filtering also supports multi-selection, allowing for comparison of data from multiple salespeople simultaneously.
[0199] Team-based filtering: Supports viewing data by organizational structure. Managers can choose to view data for the entire department or drill down to a specific team or region. This is especially useful for managers of large teams—first see the big picture, identify problems, and then drill down to specific teams for analysis.
[0200] The three-dimensional filters can be used in combination. For example, a manager can select "This Week + Zhang San + Sales Department 1" to view all of Zhang San's task data for this week within Sales Department 1. The filter conditions are related by "AND," meaning that all conditions must be met simultaneously.
[0201] In some embodiments, the filtering function helps managers narrow down the scope of data, while the drill-down function helps managers delve deeper along the path of "team → individual → task → details" to pinpoint the root cause of the problem.
[0202] The visualization interface can adopt a front-end and back-end separation architecture. The server provides a RESTful API interface, which supports querying aggregated statistical data by time, personnel, team, and other dimensions; the management terminal (web front-end or mobile APP) calls these APIs to obtain data and uses chart libraries such as ECharts and D3.js for visualization rendering.
[0203] To ensure real-time performance, a persistent WebSocket connection is established between the administrator terminal and the server. When the edge processing unit uploads new incremental data, the server proactively pushes an update notification via WebSocket, and the administrator terminal automatically refreshes the relevant charts, allowing the administrator to see the latest data without manually refreshing the page.
[0204] Managers no longer need to meticulously review each salesperson's CRM records or listen to call recordings. Instead, they can quickly grasp team dynamics and identify problematic salespeople and tasks through visual dashboards, filtering, and drill-down functions. This data-driven management approach is more efficient, comprehensive, and provides a more evidence-based basis for decision-making compared to traditional manual analysis that relies on personal experience.
[0205] It should be noted that if the edge processing unit generates an alarm message that requires assistance from other personnel, this alarm message should be proactively displayed on the administrator's terminal as soon as possible so that the administrator can intervene and resolve the issue.
[0206] In some embodiments, the edge processing unit further includes a local storage module configured to store the multimodal raw data collected by the data acquisition module. The local storage module is also configured to asynchronously upload the stored multimodal raw data to the server when a preset idle state is detected in the local terminal; the server is deployed with a pre-trained general processing model for offline batch processing of the uploaded multimodal raw data to correct or optimize the analysis and prediction results of the edge processing unit.
[0207] While lightweight models in edge processing units can run in real time on local terminals, their analytical accuracy lags behind that of large-scale cloud models due to limitations in terminal computing resources and model size. This gap is primarily reflected in the following: lightweight models approximate the original large models, resulting in a slight loss of accuracy during compression; and edge models can only be trained on limited historical data, unlike cloud models which can utilize full datasets for optimization.
[0208] However, sales behavior analysis and forecasting require high accuracy. Managers want the system to not only respond in real time but also continuously improve the accuracy of the analysis. How to maintain real-time performance at the edge while continuously improving the accuracy of lightweight models is a technical problem that needs to be solved.
[0209] Furthermore, multimodal raw data (especially call recordings and chat logs) contains a wealth of information. Edge processing units only extract key features during real-time processing, meaning the raw data itself may contain untapped value. How to fully extract the value of raw data without compromising real-time performance is also a problem that needs to be solved.
[0210] The local storage module in the edge processing unit is responsible for managing the local storage of multimodal raw data. The voice, text, structured behavior, and time-series behavior data collected in real time by the data acquisition module are not discarded immediately after real-time processing, but are written to the circular buffer of the local storage module.
[0211] A circular buffer is a fixed-size circular storage structure. In this embodiment, the buffer size is set to store the most recent 7 days of multimodal raw data. When the buffer is full, the oldest data is automatically overwritten. This design balances storage space and data retention time—a 7-day window is sufficient to cover the analysis needs of most sales cycles, while not consuming excessive terminal storage space (default allocation 2-3GB).
[0212] The local storage module is also responsible for data compression and encryption. The raw voice data is compressed using Opus or AAC, reducing its size to 10-20% of its original size; all stored data is locally encrypted using AES-256, ensuring that the original data cannot be easily accessed even if the terminal is lost or stolen.
[0213] The local storage module does not upload data immediately after collection; instead, it waits until the terminal is in a "preset idle state" before initiating asynchronous uploads. This design aims to prevent upload tasks from interfering with sales staff's daily work and to avoid consuming the network bandwidth required for sales operations.
[0214] The preset idle state includes the following conditions, all of which must be met simultaneously to trigger an upload: Charging status detection: The device must be charging. This ensures that uploading tasks does not consume battery power, avoiding disruption to subsequent mobile work for sales.
[0215] Network status detection: The terminal must be connected to a Wi-Fi network to avoid using cellular data, which would incur data charges or consume mobile network bandwidth. For certain enterprise scenarios, it can also be configured to allow uploading over 5G networks, but explicit user authorization is required.
[0216] Time window detection: The current time is within the preset upload window, which defaults to 2:00 AM to 5:00 AM. During this period, sales staff are usually not working, and the terminal is idle, making it an ideal time to perform background tasks.
[0217] Load status detection: The terminal CPU utilization is below 30%, and there is sufficient available memory. This ensures that the upload task does not compete for computing resources with other applications that the user is using.
[0218] When all the above conditions are met, the local storage module starts the upload task. If any condition is no longer met during the upload process (for example, the user unplugs the charger, or the user starts using the phone causing an increase in CPU usage), the upload task will be paused and will resume from where it was interrupted once the condition is restored (supports resuming interrupted uploads).
[0219] The server is deployed with a pre-trained general-purpose processing model. Compared to the lightweight models of edge processing units, the general-purpose processing model has the following characteristics: The model is larger in size. The general processing model uses standard BERT (approximately 400MB) instead of TinyBERT, and standard LSTM (256-dimensional hidden states) instead of TinyLSTM. The larger model capacity allows it to learn more complex feature representations.
[0220] The training data is more comprehensive. The general processing model is trained using the full historical data, including all task data from all sales. In contrast, the edge model can only be trained and fine-tuned based on local data from a single sales.
[0221] Higher accuracy. Due to the advantages of model size and training data, the analytical and predictive accuracy of general-purpose processing models is typically 5-10 percentage points higher than that of lightweight edge models.
[0222] The general processing model and the lightweight model in the edge processing unit maintain a homogeneous structure—that is, they have the same network hierarchy, only differing in the width (number of neurons) and depth (number of layers) of each layer. This isomorphic design allows the edge model to be viewed as a "compressed version" of the cloud model, enabling knowledge transfer between the two.
[0223] Offline batch processing and correction workflow Once the server receives the multimodal raw data asynchronously uploaded by the edge processing unit, it initiates the offline batch processing process.
[0224] Step 1: Data Preprocessing and Alignment. The server cleans, formats, and aligns the uploaded raw data to ensure data quality.
[0225] Step 2: General Model Inference. The server runs a general processing model to perform complete forward inference on the uploaded data, obtaining high-precision analysis and prediction results.
[0226] Step 3: Result Comparison. The results of the general model are compared line by line with the real-time results generated by the edge processing unit. Comparison dimensions include: consistency in stage classification, differences in predicted order probability, consistency in churn risk assessment levels, and consistency in alarm triggering.
[0227] Step 4: Discrepancy Analysis and Correction. For inconsistent samples identified during the comparison, the system conducts in-depth analysis. If there is a systematic deviation between the edge-end results and the cloud-based results (e.g., the edge-end systematically underestimates the probability of generating a single order), it indicates a calibration problem in the edge-end model, requiring adjustment. The system generates correction data, including: correct analysis results, correct prediction results, and a measure of the difference between the edge-end and cloud-based results.
[0228] Step 5: Model Optimization. The calibration data is used to optimize the lightweight model at the edge. There are two optimization methods: one is to periodically retrain the edge model using the calibration data as training samples to improve model accuracy; the other is to use knowledge distillation to allow the edge model to learn the output distribution of the cloud model. The optimized model is deployed to the edge via silent updates, without affecting daily sales operations.
[0229] Through a closed-loop mechanism of "real-time edge processing + offline cloud correction," the system achieves continuous self-optimization. Below is a complete example of this closed-loop mechanism: In the first week, a new salesperson, Xiao Li, joined the company, and the edge processing unit used an initial lightweight model to serve him. The initial model had limited accuracy, and some judgments were not accurate enough.
[0230] At the end of the first week, Xiao Li's local storage module asynchronously uploaded the week's raw data to the server during the night when it was idle.
[0231] The server used a general processing model to batch process this week's data and discovered a bias in the edge model's customer objection identification—misclassifying some "customer inquiries" as "customer objections," leading to unnecessary alerts. The system generated corrective data and adjusted the parameters of the objection identification module.
[0232] At the start of the second week, the optimized model was silently updated and sent to Xiao Li's phone. The new model was more accurate in identifying objections, the number of alarms decreased, and Xiao Li's trust in the system increased.
[0233] As the closed-loop system continues to operate, the accuracy of the edge model gradually approaches that of the cloud model. Although the edge accuracy cannot completely reach the cloud level due to model size limitations, the gap will gradually narrow from the initial 10 percentage points to within 3-5 percentage points. Simultaneously, due to continuous model updates, the edge model can adapt to changes in sales operations (such as new product launches and new sales strategies), consistently maintaining good analytical performance.
[0234] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this application to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.
Claims
1. A sales behavior intelligent analysis and early warning system, characterized in that, include: The edge processing unit deployed on each salesperson's local terminal includes: The data acquisition module is configured to collect multimodal raw data generated by sales personnel during customer follow-up in real time; the multimodal raw data includes: text modal data, voice modal data, structured behavioral modal data, and time-series behavioral modal data; A multimodal encoder group, connected to the data acquisition module, is configured to extract features from each modal data and generate corresponding modal feature vectors. A multimodal fusion module, connected to the multimodal encoder group, is configured to fuse the feature vectors of each modality through a cross-modal attention mechanism to generate a unified multimodal fusion feature; An analysis module, connected to the multimodal fusion module, is configured to analyze the actual state of a task based on the multimodal fusion features, wherein the actual state of a task includes at least the task execution progress. The prediction module, connected to the multimodal fusion module, is configured to perform at least one prediction task based on the multimodal fusion features to obtain a prediction result; the prediction task includes at least one of: order probability prediction, performance amount prediction, and customer churn risk assessment. And, the server, communicating with the edge processing units of each local terminal, including: An incremental data receiving interface is configured to receive incremental data uploaded by the edge processing unit; the incremental data includes at least: information on changes in the actual state of the task and information on updates to the prediction results; The real-time dashboard service module is connected to the incremental data receiving interface and is configured to perform real-time aggregation and statistics on local historical data and received incremental data, obtain statistical data, and push it to a preset administrator terminal. The administrator terminal displays the statistical data through a visual interface based on the statistical data.
2. The system according to claim 1, characterized in that, The multimodal raw data includes data for different tasks; The edge processing unit further includes a priority scheduling module, configured to prioritize processing data corresponding to high-priority tasks according to a preset task priority configuration; wherein each task is assigned a priority identifier.
3. The system according to claim 1, characterized in that, The multimodal encoder group includes a speech preprocessing encoder, which is configured as follows: The speech modal data is segmented, and low-information speech data is identified and removed. The low-information voice data includes at least one of the following: Small talk and social opening remarks; Filler words and catchphrases; Silence and pauses; Repetitive information; Non-business related information.
4. The system according to claim 1, characterized in that, The analysis module is also configured to: Analyze the difficulties encountered during task processing and advancement; When the difficult information meets the preset alarm rules, an alarm prompt message is generated.
5. The system according to claim 4, characterized in that, The analysis module is specifically configured as follows: Construct or obtain a sales persona for a salesperson, wherein the sales persona includes at least: job title tags, competency tags, and historical performance tags; The threshold parameters of the preset alarm rules are dynamically adjusted based on the salesperson profile.
6. The system according to claim 5, characterized in that, The analysis module is also configured to: When the preset triggering conditions are met, the salesperson profile is updated, and the preset alarm rules corresponding to the salesperson are updated simultaneously. The preset triggering conditions include at least one of the following: Received a preset character portrait adjustment command; Sales personnel have completed a preset threshold number of tasks. The time interval since the last portrait update has reached the preset duration.
7. The system according to claim 1, characterized in that, The edge processing unit is a lightweight deployment unit that meets at least one of the following characteristics: The multimodal encoder group, multimodal fusion module, analysis module, and prediction module are lightweight models that have undergone model quantization, knowledge distillation, or structured pruning. The edge processing unit operates independently when the local terminal is offline, without relying on a real-time network connection with the server.
8. The system according to claim 1, characterized in that, The administrator terminal is communicatively connected to the edge processing units of each local terminal; The administrator terminal is configured to receive the actual task status and prediction results uploaded by the edge processing unit. The management terminal provides an interactive visual interface that supports filtering of task status and predicted results by time, personnel, and team dimensions.
9. The system according to claim 1, characterized in that, The edge processing unit further includes a local storage module configured as follows: The system stores the multimodal raw data collected by the data acquisition module.
10. The system according to claim 9, characterized in that, The local storage module is also used to asynchronously upload the stored multimodal raw data to the server when the local terminal is detected to be in a preset idle state. The server is equipped with a pre-trained general processing model for offline batch processing of uploaded multimodal raw data to correct or optimize the analysis and prediction results of the edge processing unit.