Data processing method and device, electronic equipment and storage medium

By performing multimodal data expansion and multidimensional feature fusion on the raw data of communication events, and combining it with conditions such as user identification for clustering, the problem of low accuracy in identifying user satisfaction risks has been solved, and accurate identification and personalized support for user satisfaction risks have been achieved.

CN121842012APending Publication Date: 2026-04-10中国移动通信集团云南有限公司 +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-26
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

In existing technologies, the fixed weight information corresponding to each multimodal data leads to low accuracy in identifying user satisfaction risks.

Method used

By acquiring raw data of communication events and performing multimodal data expansion, multidimensional feature extraction and feature fusion are carried out. Based on clustering conditions such as user identifier, time information, business type, scenario and location, the multidimensional fused feature vector is clustered to identify user satisfaction risks.

Benefits of technology

It enables accurate identification of user satisfaction risks, providing a reliable basis for subsequent decision-making and improving the accuracy and personalized support of identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121842012A_ABST
    Figure CN121842012A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method and device, electronic equipment and a storage medium. The data processing method comprises the following steps: acquiring original data of a communication event, and performing multi-modal data expansion on the original data of the communication event to obtain a multi-modal data entry of the communication event; performing multi-dimensional feature extraction and feature fusion on the multi-modal data items to obtain a multi-dimensional fusion feature vector, multiple dimensions including at least one of an emotion dimension, a network dimension and a terminal dimension; based on at least one of the user identifier, the time information, the service type, the scene and the position as a clustering condition, performing clustering processing on the multi-dimensional fusion feature vector to obtain a data cluster; according to the user satisfaction risk identification method and the user satisfaction risk identification device, user satisfaction risk identification is performed on any user identifier based on the multi-dimensional feature value corresponding to the user identifier in the data cluster to obtain the user satisfaction risk identification result, so that accurate determination of the user satisfaction risk identification result is realized, and accurate data support is provided for subsequent decision making.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] Embodiments of the present application relate to the technical field of data processing, and in particular to a data processing method and device, electronic equipment and storage medium. BACKGROUND

[0002] With the rapid development of communication technology, more and more attention is paid to the quality of communication.

[0003] In the prior art, multi-modal data such as user recording data, network data and log data are obtained, and a user satisfaction risk level is determined according to weight information corresponding to the multi-modal data, thereby realizing identification of the user satisfaction risk. In the prior art, since the weight information corresponding to the multi-modal data is fixed, there is a problem of low accuracy of user satisfaction risk identification. SUMMARY

[0004] Embodiments of the present application provide a data processing method and device, electronic equipment and storage medium to realize accurate identification of user satisfaction risk.

[0005] According to an aspect of the present application, a data processing method is provided, which comprises:

[0006] obtaining original data of a communication event, performing multi-modal data expansion on the original data of the communication event, and obtaining a multi-modal data entry of the communication event;

[0007] performing multi-dimensional feature extraction and feature fusion on the multi-modal data entry to obtain a multi-dimensional fusion feature vector, the multi-dimensions including at least one of an emotion dimension, a network dimension and a terminal dimension;

[0008] performing clustering processing on the multi-dimensional fusion feature vector based on at least one of a user identifier, time information, a service type, a scene and a location as a clustering condition, and obtaining a data cluster;

[0009] performing user satisfaction risk identification on the multi-dimensional feature values corresponding to the data cluster based on the user identifier for any user identifier, and obtaining a user satisfaction risk identification result.

[0010] According to another aspect of the present application, a data processing device is provided, which comprises:

[0011] a multi-modal data entry obtaining module configured to obtain original data of a communication event, perform multi-modal data expansion on the original data of the communication event, and obtain a multi-modal data entry of the communication event;

[0012] The multi-dimensional fusion feature vector determination module is configured to perform multi-dimensional feature extraction and feature fusion on the multi-modal data entries to obtain multi-dimensional fusion feature vectors, wherein the multi-dimensions include at least one of an emotion dimension, a network dimension and a terminal dimension;

[0013] The clustering module is configured to perform clustering processing on the multi-dimensional fusion feature vectors based on at least one of a user identifier, time information, a service type, a scene and a location as a clustering condition to obtain data clusters.

[0014] The recognition module is configured to perform user satisfaction risk recognition on multi-dimensional feature values corresponding to the data clusters based on a user identifier for any user identifier to obtain a user satisfaction risk recognition result.

[0015] According to another aspect of the present application, an electronic device is provided, which comprises:

[0016] at least one processor; and

[0017] a memory connected in communication with the at least one processor; wherein

[0018] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to perform the data processing method of any embodiment of the present application.

[0019] According to another aspect of the present application, a computer readable storage medium is provided, which stores computer instructions for enabling a processor to implement the data processing method of any embodiment of the present application when executed by the processor.

[0020] The technical scheme of the embodiment of the application obtains original data of a communication event, performs multi-modal data expansion on the original data of the communication event, obtains a multi-modal data entry of the communication event, realizes accurate determination of the multi-modal data entry of the communication event, and provides diversified data support for subsequent analysis and processing; multi-dimensional feature extraction and feature fusion are performed on the multi-modal data entry to obtain a multi-dimensional fusion feature vector, the multi-dimensions include at least one of an emotion dimension, a network dimension and a terminal dimension, determination of the multi-dimensional fusion feature vector is realized, and comprehensive data support is provided for subsequent analysis and recognition; clustering processing is performed on the multi-dimensional fusion feature vector based on at least one of a user identifier, time information, a service type, a scene and a location as a clustering condition to obtain a data cluster, clustering processing of the multi-dimensional fusion feature vector is realized, and personalized data support is provided for subsequent analysis and recognition; for any user identifier, user satisfaction risk recognition is performed on the multi-dimensional feature values corresponding to the data cluster based on the user identifier to obtain a user satisfaction risk recognition result, the problem of low accuracy of user satisfaction risk recognition in the prior art is solved, accurate recognition of user satisfaction risk is realized, and accurate basis is provided for subsequent decision-making.

[0021] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the application, nor is it used to limit the scope of the application. Other features of the application will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0022] In order to more clearly illustrate the technical solutions of the embodiments of the application, the following will briefly introduce the drawings needed to be used in the embodiments. It should be understood that the following drawings only show some of the embodiments of the application, and therefore should not be regarded as limiting the scope. For those skilled in the art, other related drawings can also be obtained without creative labor on the basis of these drawings.

[0023] Figure 1 is a flow chart of a data processing method provided by the first embodiment of the application;

[0024] Figure 2 is a structural schematic diagram of an emotion profiler provided by the embodiment of the application;

[0025] Figure 3 is a structural schematic diagram of a network profiler provided by the embodiment of the application;

[0026] Figure 4 is a structural schematic diagram of a terminal profiler provided by the embodiment of the application;

[0027] Figure 5 is a structural schematic diagram of a data processing method provided by the second embodiment of the application;

[0028] Figure 6 is a flow chart of a data processing method provided by an embodiment of the present application;

[0029] Figure 7 is a schematic diagram of a multi-dimensional fusion feature vector determination process provided by an embodiment of the present application;

[0030] Figure 8 is a structural schematic diagram of a data processing device provided by the third embodiment of the present application;

[0031] Figure 9 is a structural schematic diagram of an electronic device provided by the fourth embodiment of the present application. DETAILED DESCRIPTION

[0032] In order to make the personnel in the technical field better understand the present application scheme, the technical scheme in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by the person skilled in the art without creative labor should belong to the scope of protection of the present application.

[0033] It should be noted that the terms "first", "second" and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily have to describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0034] It can be understood that before using the technical solutions disclosed in the embodiments of the present application, the type of personal information involved in the present application, the use range, the use scene and the like should be informed to the user and the authorization of the user should be obtained according to relevant laws and regulations.

[0035] Embodiment one

[0036] Figure 1This is a flowchart of a data processing method provided in Embodiment 1 of the present invention. This embodiment is applicable to situations involving the identification of user satisfaction risks. The data processing method can be executed by a data processing device in this embodiment, which can be implemented in software and / or hardware. This data processing device can be configured in an electronic device provided in this embodiment, such as a server, computer, or mobile terminal, for example, a mobile terminal can be a mobile phone, tablet computer, etc. Figure 1 As shown, the method specifically includes the following steps:

[0037] S110. Obtain the original data of the communication event, perform multimodal data expansion on the original data of the communication event, and obtain the multimodal data entries of the communication event.

[0038] Communication events refer to business interaction behaviors occurring within the communication system. Communication events include, but are not limited to, call events. Raw data refers to unprocessed data generated by business interaction behaviors occurring within the communication system. Optionally, raw data for communication events includes audio recordings, terminal logs, and wireless extension record data. Audio recordings are audio information recording business interaction behaviors. Terminal logs are operational records of communication terminals. Wireless extension record data is data from wireless communication scenarios. Wireless extension record data includes, but is not limited to, the type and speed of the accessed network. Raw data for communication events may also include timestamps and latitude / longitude, where the timestamp represents the occurrence time of the communication event corresponding to the raw data, and the latitude / longitude represents the location information of the raw data of the communication event. Raw data for communication events can be obtained from a raw database, which can store raw data from multiple users under different communication events. The preset number of users' raw data under different communication events in the raw database is determined according to requirements. Multimodal data entries are data obtained by multi-dimensional expansion of the raw data of communication events. For example, data associated with the raw data is determined based on the raw data, and multimodal data entries for communication events are determined based on the raw data and the data associated with the raw data.

[0039] Specifically, based on the requirements, a preset number of users' raw data under different communication events are determined in the original database. Based on the raw data, data associated with the raw data is determined. Based on the raw data and the data associated with the raw data, multimodal data entries for communication events are determined, thus achieving accurate determination of multimodal data entries for communication events and providing diversified data support for subsequent analysis and processing.

[0040] For example, after obtaining the raw data of the communication event, the raw data is sorted according to the ascending order of the timestamps, and the sorted raw data is adjusted according to the minimum misalignment algorithm to obtain the matrix corresponding to the rearranged raw data, thus realizing the time sequence correction of the raw data.

[0041] Based on the above embodiments, after time-series correction of the original data, corrected original data is obtained. The corrected original data also includes a user's unique identifier. The corrected original data is divided according to a preset time interval to obtain multiple data segments. Data segments with the same user's unique identifier are aggregated to obtain data segments corresponding to different users.

[0042] Optionally, the original data of the communication event is extended with multimodal data to obtain multimodal data entries for the communication event, including: querying related information based on the timestamp and latitude and longitude in the original data of the communication event to obtain related data, which includes at least one of meteorological data, holiday status data and event broadcast data; converting the latitude and longitude in the original data of the communication event into raster encoding; forming environmental description information based on the related data and converting the environmental description information into scene labels; and generating multimodal data entries for the communication event based on at least one of the original data, business fields, timestamp, related data, raster encoding and scene labels.

[0043] The associated data refers to data related to the original data of the communication event. Associated data includes at least one of meteorological data, holiday status data, and sports broadcast data. Associated data can be determined based on the timestamp and latitude / longitude of the original data of the communication event. For example, the corresponding meteorological data, holiday status data, and sports broadcast data can be obtained by searching the meteorological database, holiday status database, and sports broadcast data respectively based on the timestamp and latitude / longitude of the original data of the communication event. Raster encoding is information obtained by mapping latitude / longitude according to a preset mapping rule. For example, latitude / longitude can be mapped according to a raster mapping function to obtain raster encoding. Different latitude / longitude correspond to different raster encodings. Environmental description information is information characterizing the environmental features of the associated data. For example, the associated data can be input into a trained environmental description information determination model for processing to obtain environmental description information. The environmental description information determination model includes, but is not limited to, neural network models. The environmental description information determination model is selected according to requirements, and this invention is not limited to this. Scene labels are used to characterize the scene information of the data associated with the original data of the communication event. Scene labels include, but are not limited to, subways, parks, and tunnels. Scene labels can be determined based on environmental description information. For example, environmental description information can be mapped to corresponding scene labels according to preset label mapping rules. Alternatively, environmental description information can be input into a trained scene label conversion model for conversion to obtain corresponding scene labels. Scene label conversion models include, but are not limited to, neural network models; this invention does not impose any limitations. The raw data of the communication event can also include business fields, which represent the business type of the raw data. Different business fields correspond to different business types. Multimodal data entries for the communication event can also be generated based on the raw data, business fields, timestamps, associated data, raster encoding, and scene labels of the communication event. For example, the raw data, business fields, timestamps, associated data, raster encoding, and scene labels of the communication event can be arranged according to a preset data order to obtain multimodal data entries for the communication event.

[0044] Specifically, based on the timestamp and latitude / longitude of the raw data of the communication event, the corresponding meteorological data, holiday status data, and sports broadcast data are searched in the meteorological database, holiday status data, and sports broadcast data, respectively. The latitude and longitude are mapped using a raster mapping function to obtain raster codes. The associated data is input into a trained environmental description information determination model for processing to obtain environmental description information. This environmental description information is then mapped into corresponding scene labels according to preset label mapping rules. Finally, the raw data, business fields, timestamp, associated data, raster codes, and scene labels of the communication event are arranged according to a preset data order to obtain multimodal data entries for the communication event. This achieves accurate determination of multimodal data entries for the communication event, providing diverse data support for subsequent analysis and processing.

[0045] S120. Perform multi-dimensional feature extraction and feature fusion on the multimodal data entries to obtain a multi-dimensional fused feature vector, wherein the multi-dimensional features include at least one of the following: emotion dimension, network dimension, and terminal dimension.

[0046] The multi-dimensional fusion feature vector is the data obtained by extracting and fusing multi-dimensional features from multi-modal data entries. For example, multi-modal data entries can be input into a trained multi-dimensional fusion feature vector determination model for processing to obtain a multi-dimensional fusion feature vector. The multi-dimensional fusion feature vector determination model includes, but is not limited to, a neural network model. The multi-dimensional fusion feature vector determination model can be selected according to the requirements, and this invention does not impose any restrictions.

[0047] Specifically, multimodal data entries are input into a pre-trained multidimensional fusion feature vector determination model for processing to obtain multidimensional fusion feature vectors, thus realizing the determination of multidimensional fusion feature vectors and providing comprehensive data support for subsequent analysis and recognition.

[0048] Optionally, multi-dimensional feature extraction and feature fusion are performed on the multi-modal data entries to obtain a multi-dimensional fused feature vector, including: inputting the recording data from the multi-modal data entries into the emotion profiler to obtain an emotion vector; inputting the wireless extension recording data from the multi-modal data entries into the network profiler to obtain a network vector; inputting the terminal logs from the multi-modal data entries into the terminal profiler to obtain a terminal vector; and performing fusion processing based on the emotion vector, network vector, and terminal vector to obtain a multi-dimensional fused feature vector.

[0049] The emotion profiler is a tool for analyzing emotions in recorded data. Emotion profilers include, but are not limited to, machine learning models; the choice of emotion profiler is based on requirements and is not limited in this invention. An emotion vector is data representing changes in emotion within the recorded data. The emotion vector includes an emotion score. Inputting recorded data from multimodal data entries into the emotion profiler to obtain emotion vectors involves: inputting recorded data from multimodal data entries into the emotion profiler; the emotion profiler identifying multiple sentences from the recorded data; processing each sentence according to the emotion discrimination operator in the emotion profiler to obtain the corresponding emotion features for each sentence; and arranging the emotion features corresponding to multiple sentences in chronological order to obtain the emotion vector. By using the emotion profiler to analyze and process the recorded data to obtain the emotion vector, accurate determination of the emotion vector is achieved, providing multi-dimensional data support for subsequent analysis and recognition. For example, see [link to example]. Figure 2 , Figure 2 This is a schematic diagram of an emotion profiler provided in an embodiment of the present invention, wherein the emotion profiler includes a bidirectional gated loop unit, a self-attention mechanism layer, a fully connected layer and a Dropout layer.

[0050] A network profiler is a tool for analyzing wireless extended record data. Network profilers include, but are not limited to, machine learning models; the choice of network profiler is based on requirements, and this invention is not restrictive. A network vector is data representing the network state. Wireless extended record data may include network uplink / downlink rates, jitter counts, and packet loss identifiers. Inputting wireless extended record data from multimodal data entries into the network profiler to obtain network vectors includes: inputting wireless extended record data from multimodal data entries into the network profiler; the network profiler determining the instantaneous network rate, average jitter rate, and cumulative packet loss value based on the network uplink / downlink rates, jitter counts, and packet loss identifiers in each wireless extended record data; and arranging the instantaneous network rate, average jitter rate, and cumulative packet loss value corresponding to multiple wireless extended record data in ascending chronological order to obtain the network vector. By analyzing and processing wireless extended record data through the network profiler to obtain network vectors, accurate determination of network vectors is achieved, providing multi-dimensional data support for subsequent analysis and identification. For example, see [link to example]. Figure 3 , Figure 3 This is a schematic diagram of a network profiler provided in an embodiment of the present invention, wherein the network profiler includes a one-dimensional convolutional layer, two max pooling layers, a gated recurrent unit, and a fully connected layer.

[0051] A terminal profiler is a tool for analyzing terminal logs. Terminal profilers include, but are not limited to, machine learning models; the choice of terminal profiler is based on requirements, and this invention is not restrictive. A terminal vector is data representing the operating status of a terminal device. The terminal vector may include the cumulative number of terminal device anomalies. Terminal logs may include the terminal device's battery level, temperature, and restart alarms. Inputting the terminal logs from multimodal data entries into the terminal profiler to obtain terminal vectors includes: inputting the terminal logs from multimodal data entries into the terminal profiler; the terminal profiler determining the terminal device's battery decay rate and anomaly trigger frequency based on the battery level, temperature, and restart alarms in the terminal logs; and arranging the battery decay rate and anomaly trigger frequency corresponding to multiple terminal logs in ascending chronological order to obtain the terminal vectors. By analyzing and processing the terminal logs through the terminal profiler to obtain terminal vectors, accurate determination of terminal vectors is achieved, providing multi-dimensional data support for subsequent analysis and identification. For example, see [link to example]. Figure 4 , Figure 4 This is a schematic diagram of a terminal profiler provided in an embodiment of the present invention, wherein the terminal profiler includes a one-dimensional convolutional layer, two max pooling layers, a gated recurrent unit, and a fully connected layer.

[0052] It should be noted that the emotion profiler, network profiler, and terminal profiler can have the same structure or different structures, depending on the requirements. This invention does not impose any restrictions.

[0053] The multi-dimensional fusion feature vector can also be determined based on the emotion vector, network vector, and terminal vector. For example, the emotion vector, network vector, and terminal vector can be concatenated to obtain the multi-dimensional fusion feature vector. Alternatively, the emotion vector, network vector, and terminal vector can be input into a trained feature fusion model for processing to obtain the multi-dimensional fusion feature vector. The feature fusion model includes, but is not limited to, neural network models. The feature fusion model is selected according to the requirements, and this invention does not impose any limitations.

[0054] Specifically, the audio recording data from the multimodal data entries is input into the emotion profiler, which analyzes the recording data to obtain an emotion vector. The wireless extension recording data from the multimodal data entries is input into the network profiler to obtain a network vector. The terminal logs from the multimodal data entries are input into the terminal profiler to obtain a terminal vector. The emotion vector, network vector, and terminal vector are concatenated to obtain a multidimensional fusion feature vector, which achieves accurate determination of the multidimensional fusion feature vector and provides accurate and comprehensive data support for subsequent analysis and recognition.

[0055] Optionally, a multi-dimensional fused feature vector is obtained by fusing the emotion vector, network vector, and terminal vector, including: adding scene labels to the emotion vector, network vector, and terminal vector respectively; and fusing the emotion vector, network vector, and terminal vector with added scene labels based on multi-dimensional weights to obtain the multi-dimensional fused feature vector.

[0056] Adding scene labels to the emotion vector, network vector, and terminal vector can be achieved by adding the scene label as a prefix to the first position of each vector. Multi-dimensional weights represent the weights corresponding to the emotion, network, and terminal dimensions, respectively. The multi-dimensional fusion feature vector can also be obtained by fusing the emotion vector, network vector, and terminal vector with added scene labels based on the multi-dimensional weights. For example, the product of the weights corresponding to the emotion dimension and the emotion vector with added scene labels, the product of the weights corresponding to the network dimension and the network vector with added scene labels, and the product of the weights corresponding to the terminal dimension and the terminal vector with added scene labels can be calculated. The sum of these products—the weights corresponding to the emotion dimension and the emotion vector with added scene labels, the weights corresponding to the network dimension and the network vector with added scene labels, and the weights corresponding to the terminal dimension and the terminal vector with added scene labels—can then be used as the multi-dimensional fusion feature vector.

[0057] Specifically, scene labels are added as prefixes to the first position of the emotion vector, network vector, and terminal vector. The products of the weights corresponding to the emotion dimension and the emotion vector with the added scene label, the weights corresponding to the network dimension and the network vector with the added scene label, and the weights corresponding to the terminal dimension and the terminal vector with the added scene label are calculated. The sum of these products—the weights corresponding to the emotion dimension and the emotion vector with the added scene label, the weights corresponding to the network dimension and the network vector with the added scene label, and the weights corresponding to the terminal dimension and the terminal vector with the added scene label—is then used as the multi-dimensional fusion feature vector. This achieves accurate determination of the multi-dimensional fusion feature vector, providing accurate and comprehensive data support for subsequent analysis and recognition.

[0058] For example, the formula for calculating the multi-dimensional fused feature vector is as follows:

[0059] ;

[0060] in, This represents a multi-dimensional fused feature vector; Indicates scene label; Indicates multi-dimensional weights; Represents an emotion vector; Represents a network vector; This represents the terminal vector.

[0061] Based on the above embodiments, before adding scene labels to the emotion vector, network vector, and terminal vector respectively, the data processing method further includes: segmenting the emotion vector, network vector, and terminal vector, specifically including: performing time axis alignment on the emotion vector, network vector, and terminal vector to determine the emotion vector, network vector, and terminal vector with common time intersection, and compressing the emotion vector, network vector, and terminal vector with common time intersection according to the interval summarization operator, so that the compressed emotion vector, network vector, and terminal vector with common time intersection have the same dimension, thereby achieving time alignment and dimension alignment of the emotion vector, network vector, and terminal vector with common time intersection, providing unified data support for subsequent analysis.

[0062] S130. Based on at least one of the following as clustering conditions—user identifier, time information, business type, scenario, and location—the multi-dimensional fused feature vectors are clustered to obtain data clusters.

[0063] The raw data also includes user identifiers. User identifiers are identification information used to distinguish different users. User identifiers can be determined based on at least one of the following: unique user identification information and business type. Clustering conditions are used to determine vectors with the same characteristics in the multi-dimensional fused feature vectors. Data clusters are feature vectors obtained by clustering the multi-dimensional fused feature vectors according to the clustering conditions. A data cluster may include multiple feature vectors. Multiple data clusters can correspond to multiple multi-dimensional fused feature vectors. Data clusters can be obtained by clustering the multi-dimensional fused feature vectors according to the clustering conditions. For example, using user identifier, scenario, and business type as clustering conditions, multi-dimensional fused feature vectors with the same user identifier can be grouped into one data cluster, multi-dimensional fused feature vectors with the same scenario label can be grouped into another data cluster, and multi-dimensional fused feature vectors with the same business type can be grouped into another data cluster, resulting in multiple data clusters.

[0064] Specifically, based on the clustering conditions, multi-dimensional fused feature vectors with the same user identifier are grouped into one data cluster, multi-dimensional fused feature vectors with the same scene label are grouped into one data cluster, and multi-dimensional fused feature vectors with the same business type are grouped into one data cluster. This achieves clustering processing of multi-dimensional fused feature vectors, providing personalized data support for subsequent analysis and identification.

[0065] S140. For any user identifier, perform user satisfaction risk identification based on the multi-dimensional feature values ​​corresponding to the user identifier in the data cluster, and obtain the user satisfaction risk identification result.

[0066] In this context, multi-dimensional feature values ​​are the feature values ​​of the multi-dimensional fused feature vector corresponding to the user identifier in the data cluster. Different user identifiers correspond to different multi-dimensional feature values. The user satisfaction risk identification result represents whether the user corresponding to the user identifier faces a risk of satisfaction loss. The risk identification result can be either no satisfaction risk or a satisfaction risk present. The user satisfaction risk identification result can be determined based on the multi-dimensional feature values ​​corresponding to the user identifier in the data cluster. For example, by searching for the user identifier in the data cluster, the multi-dimensional feature values ​​corresponding to the user identifier are determined. These multi-dimensional feature values ​​are then input into a trained user satisfaction risk identification model for identification, yielding the user satisfaction risk identification result. The user satisfaction risk identification model includes, but is not limited to, neural network models. The choice of user satisfaction risk identification model is based on requirements, and this invention does not impose any limitations.

[0067] Specifically, the system searches within the data cluster based on user identifiers to determine the corresponding multi-dimensional feature values. These feature values ​​are then input into a trained user satisfaction risk identification model for identification, resulting in accurate identification of user satisfaction risks and providing a reliable basis for subsequent decision-making.

[0068] Optionally, user satisfaction risk identification is performed based on the multi-dimensional feature values ​​corresponding to the user identifier in the data cluster to obtain the user satisfaction risk identification result, including: extracting multi-dimensional feature values ​​within a set time window based on the user identifier in the data cluster, and determining the amount of change of the multi-dimensional features; obtaining a multi-dimensional weight vector, and determining the user satisfaction risk value corresponding to the user identifier based on the multi-dimensional weight vector and the amount of change of the multi-dimensional features; and determining the user satisfaction risk level corresponding to the user satisfaction risk value based on the risk level mapping relationship of the preset value, as the user satisfaction risk identification result.

[0069] The time window can be set to 15 seconds or 20 seconds, depending on the requirements; this invention does not impose any limitations. Multi-dimensional feature values ​​include, but are not limited to, sentiment scores, network uplink and downlink rates, and the cumulative number of terminal device anomalies. Correspondingly, the multi-dimensional feature change represents the changes in multi-dimensional feature values ​​within the set time window. Multi-dimensional feature change includes, but is not limited to, sentiment fluctuation amplitude, network fluctuation, and the cumulative increase in terminal device anomalies. The multi-dimensional weight vector represents the weight information corresponding to each multi-dimensional feature change. The user satisfaction risk value is used to characterize the degree of user dissatisfaction with the communication event experience. The user satisfaction risk value can be determined based on the multi-dimensional weight vector and the multi-dimensional feature change. For example, the weighted sum between the multi-dimensional weight vector and the multi-dimensional feature change is calculated, and this weighted sum is used as the user satisfaction risk value. The user satisfaction risk level characterizes whether there is a risk of low user satisfaction. The user satisfaction risk level can be determined based on a preset risk level mapping relationship. For example, the user satisfaction risk value is matched against a preset risk level mapping relationship table to determine the user satisfaction risk level corresponding to the user satisfaction risk value.

[0070] Specifically, a pre-set time window is used. Within the data cluster, based on the user identifier, the emotional score, network uplink and downlink speeds, and cumulative number of terminal device anomalies within the set time window are extracted. The amplitude of emotional fluctuations, network drops, and cumulative increases in terminal device anomalies within the set time window are calculated to obtain a multi-dimensional weight vector. The weighted sum between the multi-dimensional weight vector and the changes in multi-dimensional features is calculated, and this weighted sum is used as the user satisfaction risk value. The user satisfaction risk value is matched against a preset risk level mapping table to determine the user satisfaction risk level corresponding to the user satisfaction risk value. This user satisfaction risk level is used as the user satisfaction risk identification result, achieving accurate identification of user satisfaction risk and providing an accurate basis for subsequent decision-making.

[0071] For example, the time window is set as The formula for calculating the user satisfaction risk value is as follows:

[0072] ;

[0073] ;

[0074] in, This represents the risk value of user satisfaction. Represents a multi-dimensional weight vector; Indicates the amount of change in multi-dimensional features; Indicates the quantity of multiple dimensions; express The emotional score corresponding to each moment; express The emotional score corresponding to each moment; express The network uplink and downlink rates at any given moment; express The network uplink and downlink rates at any given moment; express The cumulative number of terminal device anomalies corresponding to each moment; express The cumulative number of terminal device anomalies corresponding to each moment.

[0075] For example, the risk level mapping relationship of the preset values ​​is as follows:

[0076] ;

[0077] in, Indicates the risk level of user satisfaction; This represents the risk value of user satisfaction. This represents the first user satisfaction risk threshold; This represents the second user satisfaction risk threshold; This represents the third user satisfaction risk threshold; This represents the fourth user satisfaction risk threshold; This represents the fifth user satisfaction risk threshold.

[0078] Based on the above embodiments, the mapping relationship between user satisfaction risk levels and preset values ​​can also be updated. The mapping relationship between the user satisfaction risk value corresponding to a user identifier and the preset value can be updated according to a preset time interval. The preset time interval can be 10 minutes or 20 minutes, depending on the requirements; this invention is not limited by this. For example, the preset time interval can be 10 minutes. When the user satisfaction risk level corresponding to a user identifier is greater than the preset maximum user satisfaction risk level for the first preset number of consecutive times, the multi-dimensional weight vector can be increased and / or the corresponding user satisfaction risk threshold can be increased; when the user satisfaction risk level corresponding to a user identifier is less than or equal to the preset minimum user satisfaction risk level for the first preset number of consecutive times, the multi-dimensional weight vector can be decreased and / or the corresponding user satisfaction risk threshold can be decreased.

[0079] Optionally, the data processing method further includes: determining the maximum component of the Hadamard operation for the multi-dimensional weight vector and the multi-dimensional feature changes; performing attribution mapping on the maximum component based on the attribution mapping table to determine the attribution label corresponding to the user satisfaction risk identification result.

[0080] The maximum component of the Hadamard operation for the multi-dimensional weight vector and the multi-dimensional feature changes can be determined by performing a Hadamard operation on the multi-dimensional weight vector and the multi-dimensional feature changes to obtain multiple components, and then identifying the maximum component with the largest value among these components. Attribution labels are used to characterize the reasons for the user's low satisfaction risk. Attribution labels include, but are not limited to, network quality, pricing packages, service attitude, and terminal malfunction. Attribution labels can be determined by attribution mapping the maximum component according to an attribution mapping table. For example, matching the maximum component in the attribution mapping table yields the attribution label corresponding to the user satisfaction risk identification result.

[0081] Specifically, Hadamard operation is performed on the multi-dimensional weight vector and the multi-dimensional feature change to obtain multiple components. The largest component with the largest value among the multiple components is determined. The largest component is then matched in the attribution mapping table to obtain the attribution label corresponding to the user satisfaction risk identification result. This realizes the attribution of the user's low satisfaction risk and provides a basis for subsequent decision-making.

[0082] The technical solution of this embodiment obtains the original data of communication events, expands the original data of communication events into multimodal data, and obtains multimodal data entries of communication events. This achieves accurate determination of multimodal data entries of communication events, providing diversified data support for subsequent analysis and processing. Multi-dimensional feature extraction and feature fusion are performed on the multimodal data entries to obtain multi-dimensional fused feature vectors. The multi-dimensional features include at least one of emotion dimension, network dimension, and terminal dimension, achieving determination of the multi-dimensional fused feature vectors and providing comprehensive data support for subsequent analysis and identification. Based on at least one of user identifier, time information, service type, scenario, and location as clustering conditions, the multi-dimensional fused feature vectors are clustered to obtain data clusters, achieving clustering of multi-dimensional fused feature vectors and providing personalized data support for subsequent analysis and identification. For any user identifier, user satisfaction risk identification is performed based on the multi-dimensional feature values ​​corresponding to the user identifier in the data cluster, obtaining user satisfaction risk identification results. This achieves accurate identification of user satisfaction risk and provides an accurate basis for subsequent decision-making.

[0083] Example 2

[0084] Figure 5 This is a schematic diagram of a data processing method provided in Embodiment 2 of the present invention. This embodiment is a refinement of the above embodiments. Based on the foregoing embodiments, it provides a detailed explanation of clustering data clusters by using at least one of user identifier, time information, business type, scenario, and location as clustering conditions, and performing multi-dimensional fused feature vector clustering. For specific implementation details, please refer to the technical solution of this embodiment. Technical terms that are the same as or corresponding to those in the above embodiments will not be repeated here. Figure 5 As shown, the method specifically includes the following steps:

[0085] S210. Obtain the original data of the communication event, perform multimodal data expansion on the original data of the communication event, and obtain the multimodal data entries of the communication event.

[0086] S220. Perform multi-dimensional feature extraction and feature fusion on the multimodal data entries to obtain a multi-dimensional fused feature vector, wherein the multi-dimensional features include at least one of the following: emotion dimension, network dimension, and terminal dimension.

[0087] S230. Construct a user grid. Based on the user identifier and timestamp corresponding to the multi-dimensional fused feature vector, map the multi-dimensional fused feature vector to the slots of the user grid to form experience atoms in the slots. Among them, one slot in the user grid corresponds to one user identifier and one unit time range. Based on at least one of business type, scenario and location, perform clustering processing on the experience atoms in each slot to obtain multiple nodes. Based on at least one of business type, scenario and location, perform clustering processing on each node to obtain multiple data clusters.

[0088] The user grid stores the multi-dimensional fused feature vectors of users. Different users correspond to different user grids. The user grid can be constructed based on user identifiers. For example, the multi-dimensional fused feature vectors corresponding to a user identifier can be matched against the user identifier. The multi-dimensional fused feature vectors corresponding to that user identifier are then sorted in ascending order of time to obtain the user grid corresponding to the user identifier. A slot is the smallest data unit position of the user grid. The same user grid includes multiple slots, and different slots store multi-dimensional fused feature vectors within a unit time range. The unit time range can be 1 second or 3 seconds, and is set according to requirements; this invention does not impose any restrictions. An experience atom is a multi-dimensional fused feature vector within a unit time range stored in a slot. A single multi-dimensional fused feature vector can be stored in different slots. Multiple multi-dimensional fused feature vectors can also be stored in the same slot. Experience atoms can be determined based on multi-dimensional fused feature vectors. For example, the multi-dimensional fused feature vectors can be mapped to slots in the user grid based on the user identifier and timestamp to obtain experience atoms. A node is a set of multiple experience atoms. Different clustering criteria correspond to different nodes. For example, using business type as the clustering criterion, experience atoms in each slot are clustered, and experience atoms with the same business type are clustered together to obtain the nodes corresponding to that business type. Data clusters can also be obtained by clustering nodes. For example, clustering nodes according to business type, and clustering nodes with the same business type together, yields data clusters corresponding to that business type.

[0089] Specifically, the process involves matching user identifiers within multi-dimensional fused feature vectors to determine the corresponding multi-dimensional fused feature vector. This vector is then sorted in ascending order by time to obtain the user grid. The multi-dimensional fused feature vectors are mapped to slots within the user grid based on the user identifier and timestamp, resulting in experience atoms. Using business type as the clustering criterion, experience atoms in each slot are clustered, with those sharing the same business type forming a cluster of nodes corresponding to that business type. Finally, nodes are clustered according to their business type, resulting in data clusters corresponding to that business type. This precise clustering of data clusters provides accurate data support for subsequent analysis and processing.

[0090] Optionally, if the slot does not include a stored experience atom, the multi-dimensional fused feature vector can be used as a new experience atom.

[0091] Specifically, when the slot does not contain any stored experience atoms, the multi-dimensional fusion feature vector corresponding to the user identifier is used as a new experience atom and stored in the slot, thus realizing the addition of experience atoms.

[0092] Optionally, if the slot includes a stored experience atom, the slot centroid is determined based on the stored experience atom. The determination is made based on the distance between the slot centroid and the multi-dimensional fusion feature vector. If the multi-dimensional fusion feature vector and the experience atom belong to the same local experience, the slot centroid is updated based on the multi-dimensional fusion feature vector. If the multi-dimensional fusion feature vector and the experience atom do not belong to the same local experience, the multi-dimensional fusion feature vector is treated as an independent experience atom.

[0093] The slot centroid is used to represent the mean feature vector of the experience atoms in the slot. The slot centroid can be determined based on stored experience atoms. For example, the multi-dimensional fusion feature vector corresponding to the experience atom can be determined based on the stored experience atoms. This multi-dimensional fusion feature vector is then input into a trained slot centroid determination model for processing to obtain the slot centroid. The slot centroid determination model includes, but is not limited to, neural network models and mathematical models. The appropriate model is selected based on requirements; this invention does not impose any restrictions. Local experiences are used to represent experience atoms with consistent slot centroids. Whether they belong to the same local experience can be determined based on the distance between the slot centroid and the multi-dimensional fusion feature vector. For example, the Euclidean distance between the slot centroid and the multi-dimensional fusion feature vector can be calculated. When the Euclidean distance between the slot centroid and the multi-dimensional fusion feature vector is less than a preset distance threshold, it indicates that the multi-dimensional fusion feature vector and the experience atom belong to the same local experience; when the Euclidean distance between the slot centroid and the multi-dimensional fusion feature vector is greater than or equal to the preset distance threshold, it indicates that the multi-dimensional fusion feature vector and the experience atom do not belong to the same local experience.

[0094] Specifically, based on the stored experience atoms, the multi-dimensional fusion feature vector corresponding to the experience atom is determined. This multi-dimensional fusion feature vector is then input into a trained slot centroid determination model for processing to obtain the slot centroid. The Euclidean distance between the slot centroid and the multi-dimensional fusion feature vector is calculated. When the Euclidean distance is less than a preset distance threshold, it indicates that the multi-dimensional fusion feature vector and the experience atom belong to the same local experience. The multi-dimensional fusion feature vector is then input into a trained slot centroid update model for processing to obtain the updated slot centroid. When the Euclidean distance is greater than or equal to the preset distance threshold, it indicates that the multi-dimensional fusion feature vector and the experience atom do not belong to the same local experience. The multi-dimensional fusion feature vector is then treated as an independent experience atom and stored within the slot. This achieves the updating of experience atoms within the slot and the updating of the slot centroid, providing data support for subsequent decision-making.

[0095] For example, the formula for updating the center of gravity of a slot is as follows:

[0096] ;

[0097] in, Indicates the center of gravity of the updated slot; Indicates the center of gravity of the slot before the update; This represents a multi-dimensional fused feature vector; This represents the adaptive smoothing coefficient. The update formula is as follows:

[0098] ;

[0099] in, Represents the normalization constant; Indicates the number of experience atoms in the slot; This indicates the correction parameters corresponding to the scene label, which can be preset; This indicates the correction parameters corresponding to the business type, which can be preset.

[0100] Optionally, the data processing method further includes: obtaining the update time of nodes in the data cluster, determining the decay index of the data cluster at the current time based on the update time of nodes in the data cluster, and discarding the data cluster if the decay index of the data cluster at the current time is zero.

[0101] The update time refers to the time when a node in the data cluster last updated its experience atom. The update time includes, but is not limited to, the time when a new experience atom is added and the time when a multi-dimensional fusion feature vector is added to an experience atom. For example, when a new experience atom is added, the time when the most recent multi-dimensional fusion feature vector, as an independent experience atom, was stored in the slot is used as the node's update time. Similarly, when a multi-dimensional fusion feature vector is added to an experience atom, the time when the multi-dimensional fusion feature vector was most recently added to the experience atom can be used as the node's update time. The decay index is used to characterize the activity state of the data cluster. The decay index of the data cluster at the current time can be determined based on the update time of the nodes in the data cluster. For example, the update time of the nodes in the data cluster can be input into a trained decay index determination model for processing to obtain the decay index of the data cluster at the current time. The decay index determination model includes, but is not limited to, neural network models and mathematical models. The decay index determination model is selected according to requirements, and this invention does not impose any restrictions. When the decay index of a data cluster at the current time is zero, it indicates that the data cluster is in an inactive state. To reduce the impact of inactive data clusters, data clusters with a decay index of zero at the current time can be discarded.

[0102] Specifically, the update time of the most recent updated experience atom of the node in the data cluster is obtained. The update time of the node in the data cluster is input into the trained decay index determination model for processing, so as to obtain the decay index of the data cluster at the current time. When the decay index of the data cluster at the current time is zero, the data cluster is discarded, which can reduce the impact of inactive data clusters and help improve the efficiency and accuracy of data processing.

[0103] Based on the above embodiments, the decay index of a data cluster at the current time can be determined according to the scene label and location. The formula for calculating the decay index of a data cluster at the current time is as follows:

[0104] ;

[0105] in, This represents the decay index of the data cluster at the current time; This represents the decay exponent of a data cluster at update time; This represents the time difference between the current moment and the update time of a node in the data cluster; Indicates the rate of decay over time; Indicates the weight when the scene label changes; This indicates the scene label corresponding to the current moment; This indicates the scene label corresponding to the update time; Indicates an indicator function; This indicates the weight corresponding to the change in position; Indicates the position corresponding to the current moment; This indicates the position corresponding to the update time.

[0106] Based on the above embodiments, the slot center of gravity can also be updated according to scene tags and business types. The formula for updating the slot center of gravity can also be:

[0107] ;

[0108] in, Indicates the center of gravity of the updated slot; Indicates the center of gravity of the slot before the update; This indicates the cumulative number of experienced atoms in the slot; This represents a multi-dimensional fused feature vector; This represents the weight adjustment value corresponding to the scene label; This indicates the weight adjustment value corresponding to the business type; Indicates hyperparameters; This represents the time difference between the current moment and the update time of a node in the data cluster.

[0109] S240. For any user identifier, perform user satisfaction risk identification based on the multi-dimensional feature values ​​corresponding to the user identifier in the data cluster, and obtain the user satisfaction risk identification result.

[0110] Optionally, the data processing method further includes: if the user satisfaction risk identification result indicates the existence of satisfaction risk, determining the satisfaction gap corresponding to the user identifier, and determining the target intervention action based on the return valuation of the intervention action and the satisfaction gap.

[0111] The satisfaction gap characterizes the actionable shortcoming between the current user experience level and the target user experience level for a given user identifier. Different users correspond to different target satisfaction levels. Intervention actions are optimization measures taken for users at satisfaction risk, aimed at improving their satisfaction. Intervention actions include, but are not limited to, fee compensation, customer service follow-up, and network parameter adjustments. The return estimate is the user satisfaction level for users at satisfaction risk after implementing optimization measures. The target intervention action is the optimization measure adapted to the user corresponding to the user identifier. The target intervention action can be determined based on the return estimate and satisfaction gap of the intervention action. For example, the return estimate and satisfaction gap of the intervention action can be input into a trained target intervention action determination model for processing to obtain the target intervention action. The target intervention action determination model includes, but is not limited to, a neural network model.

[0112] Specifically, when the user satisfaction risk identification result indicates the existence of satisfaction risk, the difference between the user satisfaction risk identification result corresponding to the user identifier and the target satisfaction is calculated. This difference is used as the satisfaction gap. The estimated return of the intervention action and the satisfaction gap are input into the trained target intervention action determination model for processing to obtain the target intervention action, thus providing a basis for improving user satisfaction.

[0113] Optionally, determining the satisfaction gap corresponding to the user identifier includes: obtaining the target satisfaction corresponding to the user identifier, and determining the satisfaction difference based on the target satisfaction and the user satisfaction risk identification result; obtaining the scene sensitivity and emotional sensitivity corresponding to the user identifier, determining the first gap item based on the scene sensitivity and the satisfaction difference, and determining the second gap item based on the emotional sensitivity and the user satisfaction risk identification result; and determining the satisfaction gap corresponding to the user identifier based on the first gap item and the second gap item.

[0114] Different users have different user levels, and different user levels correspond to different target satisfaction levels. The target satisfaction level corresponding to a user identifier can be determined according to a preset target satisfaction level mapping relationship. For example, the target satisfaction level corresponding to a user identifier can be matched in the target satisfaction level mapping table to obtain the target satisfaction level corresponding to the user identifier. The satisfaction difference is used to characterize the gap between the target satisfaction level and the user satisfaction risk identification result. The satisfaction difference can be determined based on the target satisfaction level and the user satisfaction risk identification result. For example, the difference between the target satisfaction level and the user satisfaction risk identification result is calculated, and this difference is used as the satisfaction difference. Scene sensitivity is data that characterizes the degree of influence of a scene on the user experience. Different scenes correspond to different scene sensitivities. Scene sensitivity can be determined based on scene labels. For example, the scene sensitivity under a scene can be obtained by matching the scene label in the scene sensitivity mapping table. The first gap term is used to characterize the amount of experience shortcomings under the scene dimension. The first gap term can be determined based on scene sensitivity and satisfaction difference. For example, the product between scene sensitivity and satisfaction difference can be calculated, and this product is used as the first gap term. Emotional sensitivity is data that characterizes the degree of influence of emotions on the user experience. Emotional sensitivity can be matched against an emotional sensitivity lookup table based on the emotional vector to obtain the corresponding emotional sensitivity. The second gap term represents the quantity of experience shortcomings under the emotional dimension. The second gap term can be determined based on emotional sensitivity and user satisfaction risk identification results. For example, the product between emotional sensitivity and user satisfaction risk identification results can be calculated, and this product can be used as the second gap term. The satisfaction gap corresponding to a user identifier can be determined based on the first and second gap terms. For example, the weighted sum between the first and second gap terms can be calculated, and this weighted sum can be used as the satisfaction gap corresponding to the user identifier.

[0115] Specifically, the target satisfaction level corresponding to the user identifier is matched in the target satisfaction mapping table to obtain the target satisfaction level corresponding to the user identifier; the scene sensitivity is matched in the scene sensitivity mapping table according to the scene label to obtain the scene sensitivity in that scene; and the emotion sensitivity is matched in the emotion sensitivity lookup table according to the emotion vector to obtain the emotion sensitivity corresponding to that emotion vector. The product between the scene sensitivity and the satisfaction difference is calculated, and this product is used as the first gap term. The product between the emotion sensitivity and the user satisfaction risk identification result is calculated, and this product is used as the second gap term. The weighted sum between the first gap term and the second gap term is calculated, and this weighted sum is used as the satisfaction gap corresponding to the user identifier. This achieves accurate determination of the satisfaction gap corresponding to the user identifier, providing accurate data support for subsequent decision-making.

[0116] For example, the formula for calculating the satisfaction gap corresponding to a user identifier is as follows:

[0117]

[0118] ;

[0119] in, This indicates the satisfaction gap corresponding to the user identifier; Indicates target satisfaction; This indicates the results of user satisfaction risk identification; Indicates scene sensitivity; Indicates emotional sensitivity; This represents the risk value of user satisfaction.

[0120] Optionally, the method for determining the return valuation of intervention actions includes: identifying risk nodes corresponding to user identifiers in the data cluster, and forming risk trajectories based on the risk nodes; obtaining intervention execution logs corresponding to user identifiers, which include historical intervention actions; inserting historical intervention actions into the risk trajectory to form a hybrid trajectory; forming a causal link of historical intervention actions based on the preceding and following risk nodes of each historical intervention action in the hybrid trajectory; and determining the return valuation of intervention actions based on the causal link of historical intervention actions.

[0121] In this data set, risk nodes are nodes corresponding to user identifiers within the data cluster that represent a risk to satisfaction. Risk nodes can be determined based on user identifiers. For example, risk nodes can be identified within the data cluster based on user identifiers. Risk trajectories characterize the changes in risk nodes. Risk trajectories can be formed based on risk nodes. For example, risk nodes can be arranged in ascending chronological order to form risk trajectories. Intervention execution logs record executed intervention actions. Intervention execution logs can be determined based on user identifiers. For example, an intervention execution log database can be searched based on a user identifier to obtain the corresponding intervention execution log. The database can store intervention execution logs for different users. Intervention execution logs include historical intervention actions, which are the intervention actions that have already been executed. Hybrid trajectories are sequences formed by inserting historical intervention actions into risk trajectories chronologically. Causal links characterize the relationship between historical intervention actions and their corresponding risk nodes. Causal links can be determined based on historical intervention actions and the preceding and following risk nodes of those historical intervention actions. For example, a causal chain of historical intervention actions can be formed by following the order of the preceding risk node, the historical intervention action, and the following risk node. The return estimate of the intervention action can be determined based on the causal chain of historical intervention actions. For example, the causal chain of historical intervention actions can be input into a trained return estimate determination model for processing to obtain the return estimate of the intervention action. The return estimate determination model includes, but is not limited to, a neural network model. The return estimate determination model can be selected according to the needs, and this invention is not limited thereto.

[0122] Specifically, risk nodes corresponding to user identifiers are identified within the data cluster. These risk nodes are then arranged in ascending chronological order to form risk trajectories. Intervention execution logs corresponding to the user identifier are retrieved from the intervention execution log database. Historical intervention actions are inserted into the risk trajectories to form hybrid trajectories. Within these hybrid trajectories, causal links of historical intervention actions are established in the order of the preceding risk node, the historical intervention action itself, and the following risk node. These causal links are then input into a pre-trained reward estimation model for processing, yielding the reward estimation for each intervention action. This process ensures accurate determination of the reward estimation for intervention actions and provides data support for determining subsequent target intervention actions.

[0123] For example, the formula for calculating the return on an intervention action is as follows:

[0124] ;

[0125] in, Indicates return valuation; This indicates user satisfaction before the intervention was implemented; Indicates target satisfaction; Weighting parameters representing the results of user satisfaction risk identification; Indicates an intervention action; Indicates the cost of the intervention; Weighted parameters representing the cost of intervention actions; Indicates user satisfaction after the intervention action was performed; This indicates the change in user satisfaction before and after the intervention was implemented; Indicates network quality; Weight parameters that represent network quality.

[0126] Optionally, the target intervention action is determined based on the return valuation and satisfaction gap of the intervention action, including: screening the optional intervention actions based on at least one of the scenario action whitelist and the action whitelist corresponding to the user identifier; for the screened intervention actions, the comprehensive benefit of the screened intervention actions is determined based on the return valuation, cost, historical effect representation value and preference penalty item of the intervention action; and the target intervention action is determined based on the comprehensive benefit of each screened intervention action.

[0127] The system comprises several key components: a scenario-based action whitelist (which lists executable intervention actions for each scenario), a target intervention action (which specifies actions applicable to users with different user identifiers), a historical effect representation (representing the effect achieved after an intervention), a preference penalty (representing data on intervention actions rejected by users with different user identifiers), and a comprehensive benefit (representing the benefit information after an intervention). The comprehensive benefit can be determined based on the estimated return, cost, historical effect representation, and preference penalty of the intervention action. For example, these factors can be input into a trained comprehensive benefit determination model to obtain the comprehensive benefit. The comprehensive benefit determination model can include, but is not limited to, neural network models and mathematical models. Target intervention actions can also be determined based on the comprehensive benefit of the selected intervention actions. For example, the intervention action with the highest comprehensive benefit among multiple comprehensive benefits can be used as the target intervention action.

[0128] Specifically, a scene action whitelist is determined based on scene tags, and an action whitelist suitable for users corresponding to user identifiers is determined based on user identifiers. Common intervention actions in the scene action whitelist and the action whitelist are identified and used as the filtered intervention actions. For the filtered intervention actions, the estimated return, cost, historical effect representation value, and preference penalty term of the intervention action are determined and input into a trained comprehensive return determination model for processing to obtain the comprehensive return. The intervention action corresponding to the largest comprehensive return among multiple comprehensive returns is used as the target intervention action, realizing the personalized determination of the target intervention action, which is conducive to improving user satisfaction.

[0129] For example, the formula for calculating comprehensive income is as follows:

[0130]

[0131] ;

[0132] in, This indicates the overall benefit of the intervention actions after screening; This indicates the intervention action taken after screening; Indicates the rate of risk reduction; This indicates a satisfaction gap; This indicates the time lag before the intervention measures take effect after screening. The weights representing the time lag before the intervention actions take effect after screening; This indicates the cost of the intervention action after screening; The weights representing the costs of intervention actions after screening; Indicates historical performance values; Indicates scene label; Indicates the business type; This indicates the user preferences corresponding to the user identifier; This represents the weight corresponding to the user preference associated with the user identifier.

[0133] For example, see Figure 6 and Figure 7 , Figure 6 This is a flowchart of a data processing method provided in an embodiment of the present invention. Figure 7 This is a schematic diagram illustrating the process of determining a multi-dimensional fusion feature vector according to an embodiment of the present invention.

[0134] The technical solution of this embodiment obtains the raw data of communication events, expands the raw data of communication events into multimodal data, and obtains multimodal data entries of communication events, thereby achieving accurate determination of multimodal data entries of communication events and providing diversified data support for subsequent analysis and processing; it then performs multi-dimensional feature extraction and feature fusion on the multimodal data entries to obtain multi-dimensional fused feature vectors, with the multi-dimensional features including at least one of emotion dimension, network dimension, and terminal dimension, thereby determining the multi-dimensional fused feature vectors and providing comprehensive data support for subsequent analysis and identification; finally, it constructs a user grid and maps the multi-dimensional fused feature vectors to the user grid based on the user identifier and timestamp corresponding to the multi-dimensional fused feature vectors. Within each slot, experience atoms are formed. Each slot in the user grid corresponds to a user identifier and a unit time range. Based on at least one of business type, scenario, and location, the experience atoms in each slot are clustered to obtain multiple nodes. Based on at least one of business type, scenario, and location, each node is clustered to obtain multiple data clusters. This achieves accurate clustering of data clusters, providing accurate data support for subsequent analysis and processing. For any user identifier, user satisfaction risk is identified based on the multi-dimensional feature values ​​corresponding to the user identifier in the data cluster, resulting in accurate identification of user satisfaction risk and providing an accurate basis for subsequent decision-making.

[0135] Example 3

[0136] Figure 8 This is a schematic diagram of the structure of a data processing device provided in Embodiment 3 of the present invention. Figure 8 As shown, the data processing device specifically includes a multimodal data entry acquisition module 310, a multidimensional fusion feature vector determination module 320, a clustering module 330, and a recognition module 340.

[0137] The system includes the following modules: a multimodal data entry acquisition module 310, which acquires the original data of a communication event and performs multimodal data expansion on the original data to obtain multimodal data entries for the communication event; a multidimensional fusion feature vector determination module 320, which extracts and fuses multidimensional features from the multimodal data entries to obtain a multidimensional fusion feature vector, wherein the multidimensional features include at least one of the following dimensions: emotion dimension, network dimension, and terminal dimension; a clustering module 330, which performs clustering processing on the multidimensional fusion feature vector based on at least one of the following clustering conditions: user identifier, time information, business type, scenario, and location, to obtain a data cluster; and an identification module 340, which identifies user satisfaction risk based on the multidimensional feature values ​​corresponding to the user identifier in the data cluster for any user identifier, to obtain a user satisfaction risk identification result.

[0138] The technical solution of this embodiment acquires the original data of communication events through a multimodal data entry acquisition module, expands the original data of communication events into multimodal data, and obtains multimodal data entries of communication events, thus achieving accurate determination of multimodal data entries of communication events and providing diversified data support for subsequent analysis and processing. Through a multidimensional fusion feature vector determination module, multidimensional features are extracted and fused from the multimodal data entries to obtain multidimensional fusion feature vectors. The multidimensional features include at least one of emotion, network, and terminal dimensions, thus determining the multidimensional fusion feature vectors and providing comprehensive data support for subsequent analysis and identification. Through a clustering module, the multidimensional fusion feature vectors are clustered based on at least one of user identifier, time information, service type, scenario, and location as clustering conditions to obtain data clusters, thus achieving clustering processing of multidimensional fusion feature vectors and providing personalized data support for subsequent analysis and identification. Through an identification module, for any user identifier, user satisfaction risk is identified based on the multidimensional feature values ​​corresponding to the user identifier in the data cluster, thus obtaining user satisfaction risk identification results and achieving accurate identification of user satisfaction risks, providing accurate basis for subsequent decision-making.

[0139] Based on the above embodiments, optionally, the multimodal data entry acquisition module 310 is further configured to: perform a correlation information query based on the timestamp and latitude and longitude in the original data of the communication event to obtain correlation data, wherein the correlation data includes at least one of meteorological data, holiday status data and event broadcast data; convert the latitude and longitude in the original data of the communication event into raster encoding; form environmental description information based on the correlation data, and convert the environmental description information into scene labels; and generate multimodal data entries for the communication event based on at least one of the original data, business fields, timestamp, correlation data, raster encoding and scene labels of the communication event.

[0140] Optionally, the raw data of the communication event includes recording data, terminal logs, and wireless extension recording data.

[0141] Optionally, the multi-dimensional fusion feature vector determination module 320 is further configured to: input the recording data from the multimodal data entries into the emotion profiler to obtain the emotion vector; input the wireless extended recording data from the multimodal data entries into the network profiler to obtain the network vector; input the terminal logs from the multimodal data entries into the terminal profiler to obtain the terminal vector; and perform fusion processing based on the emotion vector, network vector, and terminal vector to obtain the multi-dimensional fusion feature vector.

[0142] Optionally, the multi-dimensional fusion feature vector determination module 320 is also used to: add scene labels to the emotion vector, network vector and terminal vector respectively; and perform fusion processing on the emotion vector, network vector and terminal vector with added scene labels based on multi-dimensional weights to obtain a multi-dimensional fusion feature vector.

[0143] Optionally, the clustering module 330 is also used to: construct a user grid, map the multi-dimensional fusion feature vector to slots in the user grid based on the user identifier and timestamp corresponding to the multi-dimensional fusion feature vector, forming experience atoms in the slots; wherein, one slot in the user grid corresponds to one user identifier and one unit time range; cluster the experience atoms in each slot based on at least one of business type, scenario and location to obtain multiple nodes; and cluster each node based on at least one of business type, scenario and location to obtain multiple data clusters.

[0144] Optionally, the clustering module 330 is also used to: treat the multi-dimensional fused feature vector as a new experience atom when the slot does not include a stored experience atom.

[0145] Optionally, the clustering module 330 is further configured to: determine the slot centroid based on the stored experience atoms when the slot includes stored experience atoms, and make a judgment based on the distance between the slot centroid and the multi-dimensional fusion feature vector; if the multi-dimensional fusion feature vector and the experience atom belong to the same local experience, then update the slot centroid based on the multi-dimensional fusion feature vector; if the multi-dimensional fusion feature vector and the experience atom do not belong to the same local experience, then treat the multi-dimensional fusion feature vector as an independent experience atom.

[0146] Optionally, the data processing device further includes a discard module, used to: obtain the update time of nodes in the data cluster, determine the decay index of the data cluster at the current time based on the update time of nodes in the data cluster, and discard the data cluster if the decay index of the data cluster at the current time is zero.

[0147] Optionally, the identification module 340 is further configured to: extract multi-dimensional feature values ​​within a set time window based on the user identifier in the data cluster, and determine the amount of change of the multi-dimensional features; obtain a multi-dimensional weight vector, and determine the user satisfaction risk value corresponding to the user identifier based on the multi-dimensional weight vector and the amount of change of the multi-dimensional features; and determine the user satisfaction risk level corresponding to the user satisfaction risk value based on the risk level mapping relationship of the preset value, as the user satisfaction risk identification result.

[0148] Optionally, the data processing device further includes an attribution mapping module, used to: determine the maximum component of the Hadamard operation of the multi-dimensional weight vector and the multi-dimensional feature changes; perform attribution mapping on the maximum component based on the attribution mapping table, and determine the attribution label corresponding to the user satisfaction risk identification result.

[0149] Optionally, the data processing device further includes a target intervention action determination module, used to: determine the satisfaction gap corresponding to the user identifier when the user satisfaction risk identification result indicates that there is a satisfaction risk, and determine the target intervention action based on the return valuation of the intervention action and the satisfaction gap.

[0150] Optionally, the target intervention action determination module is also used to: identify risk nodes corresponding to user identifiers in the data cluster, and form risk trajectories based on risk nodes; obtain intervention execution logs corresponding to user identifiers, which include historical intervention actions; insert historical intervention actions into the risk trajectory to form a hybrid trajectory; in the hybrid trajectory, form a causal link of historical intervention actions based on the preceding and following risk nodes of each historical intervention action; and determine the return valuation of the intervention action based on the causal link of historical intervention actions.

[0151] Optionally, the target intervention action determination module is also used to: obtain the target satisfaction corresponding to the user identifier, and determine the satisfaction difference based on the target satisfaction and the user satisfaction risk identification result; obtain the scene sensitivity and emotional sensitivity corresponding to the user identifier, determine the first gap item based on the scene sensitivity and the satisfaction difference, determine the second gap item based on the emotional sensitivity and the user satisfaction risk identification result; and determine the satisfaction gap corresponding to the user identifier based on the first gap item and the second gap item.

[0152] Optionally, the target intervention action determination module is also used to: filter the selectable intervention actions based on at least one of the scenario action whitelist and the action whitelist corresponding to the user identifier; for the filtered intervention actions, determine the comprehensive benefit of the filtered intervention actions based on the return valuation, cost, historical effect representation value and preference penalty item of the intervention actions; and determine the target intervention action based on the comprehensive benefit of each filtered intervention action.

[0153] The data processing apparatus provided in this embodiment of the invention can execute a data processing method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of executing the method.

[0154] Example 4

[0155] Figure 9This is a schematic diagram of the structure of an electronic device provided in Embodiment 4 of the present invention. The electronic device 10 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described herein and claimed by the author.

[0156] like Figure 9 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0157] Multiple components in electronic device 10 are connected to input / output (I / O) interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of monitors, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and various telecommunications networks.

[0158] Processor 11 can be a variety of general-purpose and special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as a data processing method.

[0159] In some embodiments, a data processing method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and mounted on electronic device 10 via read-only memory (ROM) 12 and communication unit 19. When the computer program is loaded into random access memory (RAM) 13 and executed by processor 11, one or more steps of a data processing method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform a data processing method by any other suitable means (e.g., by means of firmware).

[0160] Various implementations of the systems and techniques described above herein can be realized in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and combinations thereof. These various implementations may include: implementations in one or more computer programs that can be executed and interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0161] A computer program for implementing a data processing method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer program causes the functions / operations specified in the flowcharts and block diagrams to be implemented. The computer program can be executed entirely on the machine, partially on the machine, as a standalone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0162] Example 5

[0163] Embodiment 5 of the present invention also provides a computer-readable storage medium storing computer instructions for causing a processor to execute a data processing method, the method comprising:

[0164] The process involves acquiring raw data of communication events, performing multimodal data expansion on the raw data to obtain multimodal data entries for the communication events, extracting and fusing multi-dimensional features from the multimodal data entries to obtain multi-dimensional fused feature vectors, where the multi-dimensional features include at least one of the following dimensions: emotion, network, and terminal. Based on at least one of the following clustering conditions—user identifier, time information, service type, scenario, and location—the multi-dimensional fused feature vectors are clustered to obtain data clusters. For any user identifier, user satisfaction risk is identified based on the multi-dimensional feature values ​​corresponding to the user identifier in the data cluster, resulting in a user satisfaction risk identification result.

[0165] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0166] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0167] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0168] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0169] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0170] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A data processing method, characterized by, The method comprises: obtaining original data of a communication event, performing multi-modal data expansion on the original data of the communication event to obtain a multi-modal data entry of the communication event; performing multi-dimensional feature extraction and feature fusion on the multi-modal data entry to obtain a multi-dimensional fusion feature vector, the multi-dimensions including at least one of an emotion dimension, a network dimension and a terminal dimension; performing clustering processing on the multi-dimensional fusion feature vector based on at least one of a user identifier, time information, a service type, a scene and a location as a clustering condition to obtain a data cluster; performing user satisfaction risk identification based on the user identifier on the multi-dimensional feature values corresponding to the data cluster to obtain a user satisfaction risk identification result for any user identifier.

2. The method of claim 1, wherein, The multi-modal data expansion on the original data of the communication event to obtain a multi-modal data entry of the communication event comprises: performing associated information query based on a timestamp and a longitude and latitude in the original data of the communication event to obtain associated data, the associated data including at least one of meteorological data, holiday state data and event broadcast data; converting the longitude and latitude in the original data of the communication event into a grid code; forming environment description information based on the associated data and converting the environment description information into a scene label; generating the multi-modal data entry of the communication event based on at least one of the original data of the communication event, a service field, a timestamp, the associated data, the grid code and the scene label.

3. The method of claim 1, wherein, The original data of the communication event includes audio data, terminal logs and wireless extension record data; The multi-dimensional feature extraction and feature fusion on the multi-modal data entry to obtain a multi-dimensional fusion feature vector comprises: inputting the audio data in the multi-modal data entry into an emotion profiler to obtain an emotion vector; inputting the wireless extension record data in the multi-modal data entry into a network profiler to obtain a network vector; inputting the terminal logs in the multi-modal data entry into a terminal profiler to obtain a terminal vector; performing fusion processing based on the emotion vector, the network vector and the terminal vector to obtain a multi-dimensional fusion feature vector.

4. The method of claim 3, wherein, The fusion processing based on the emotion vector, the network vector and the terminal vector to obtain a multi-dimensional fusion feature vector comprises: adding a scene label to the emotion vector, the network vector and the terminal vector respectively; performing fusion processing on the emotion vector, the network vector and the terminal vector to which the scene label is added based on multi-dimensional weights to obtain the multi-dimensional fusion feature vector.

5. The method of claim 1, wherein, The clustering processing on the multi-dimensional fusion feature vector based on at least one of a user identifier, time information, a service type, a scene and a location as a clustering condition to obtain a data cluster comprises: constructing a user grid, mapping the multi-dimensional fusion feature vector into a slot in the user grid based on a user identifier and a timestamp corresponding to the multi-dimensional fusion feature vector to form an experience atom in the slot; wherein one slot in the user grid corresponds to one user identifier and one unit time range. cluster the experience atoms in each of the slots based on at least one of the business type, the scene, and the location, to obtain a plurality of nodes; cluster each of the nodes based on at least one of the business type, the scene, and the location, to obtain a plurality of data clusters.

6. The method of claim 5, wherein, The method further comprises: in the case that the slot does not include the stored experience atom, taking the multi-dimensional fusion feature vector as a new experience atom; in the case that the slot includes the stored experience atom, determining a slot barycenter based on the stored experience atom, determining based on a distance between the slot barycenter and the multi-dimensional fusion feature vector, if the multi-dimensional fusion feature vector and the experience atom belong to the same local experience, updating the slot barycenter based on the multi-dimensional fusion feature vector, if the multi-dimensional fusion feature vector and the experience atom do not belong to the same local experience, taking the multi-dimensional fusion feature vector as an independent experience atom.

7. The method of claim 5, wherein, The method further comprises: obtaining an update time of the nodes in the data cluster, determining an attenuation index of the data cluster at the current time based on the update time of the nodes in the data cluster; in the case that the attenuation index of the data cluster at the current time is zero, discarding the data cluster.

8. The method of claim 1, wherein, The user satisfaction risk identification based on the user identifier on the multi-dimensional feature values corresponding to the data cluster, to obtain a user satisfaction risk identification result, comprises: in the data cluster, extracting the multi-dimensional feature values within a set time window based on the user identifier, and determining a multi-dimensional feature change amount; obtaining a multi-dimensional weight vector, determining a user satisfaction risk value corresponding to the user identifier based on the multi-dimensional weight vector and the multi-dimensional feature change amount; determining a user satisfaction risk level corresponding to the user satisfaction risk value based on a risk level mapping relationship of a preset value, as the user satisfaction risk identification result.

9. The method of claim 8, wherein, The method further comprises: determining a maximum component of a Hadamard operation of the multi-dimensional weight vector and the multi-dimensional feature change amount; performing attribution mapping on the maximum component based on an attribution mapping table, to determine an attribution label corresponding to the user satisfaction risk identification result.

10. The method of claim 1, wherein, The method further comprises: in the case that the user satisfaction risk identification result is that there is a satisfaction risk, determining a satisfaction gap corresponding to the user identifier, and determining a target intervention action based on a return value of the intervention action and the satisfaction gap.

11. The method of claim 10, wherein, The determination manner of the return value of the intervention action comprises: identifying a risk node corresponding to the user identifier in the data cluster, forming a risk trajectory based on the risk node; obtaining an intervention execution log corresponding to the user identifier, the intervention execution log including historical intervention actions; inserting the historical intervention actions into the risk trajectory to form a mixed trajectory; in the mixed trajectory, forming a causal link of the historical intervention action based on a previous risk node and a next risk node of each of the historical intervention actions; determining the return value of the intervention action based on the causal link of the historical intervention action.

12. The method of claim 10, wherein, The determination of the satisfaction gap corresponding to the user identifier comprises: obtain a target satisfaction degree corresponding to the user identifier, determine a satisfaction difference value based on the target satisfaction degree and the user satisfaction degree risk identification result; obtain a scene sensitivity and an emotion sensitivity corresponding to the user identifier, determine a first gap item based on the scene sensitivity and the satisfaction difference value, determine a second gap item based on the emotion sensitivity and the user satisfaction degree risk identification result; determine a satisfaction gap corresponding to the user identifier based on the first gap item and the second gap item.

13. The method of claim 10, wherein, determine a target intervention action based on a reward estimation value of an intervention action and the satisfaction gap, including: screen the optional intervention action based on at least one of a scene action whitelist and an action whitelist corresponding to the user identifier; for the screened intervention action, determine a comprehensive benefit of the screened intervention action based on a reward estimation value of the intervention action, a cost, a historical effect representation value and a preference penalty item; and determine the target intervention action based on the comprehensive benefits of the screened intervention actions.

14. A data processing apparatus, characterized by including: a multi-modal data item obtaining module, configured to obtain original data of a communication event, perform multi-modal data expansion on the original data of the communication event, and obtain multi-modal data items of the communication event; a multi-dimensional fusion feature vector determining module, configured to perform multi-dimensional feature extraction and feature fusion on the multi-modal data items, and obtain a multi-dimensional fusion feature vector, the multi-dimensions including at least one of an emotion dimension, a network dimension and a terminal dimension; a clustering module, configured to perform clustering processing on the multi-dimensional fusion feature vector based on at least one of a user identifier, time information, a service type, a scene and a location as clustering conditions, and obtain data clusters; an identification module, configured to, for any user identifier, perform user satisfaction degree risk identification on multi-dimensional feature values corresponding to the data clusters based on the user identifier, and obtain a user satisfaction degree risk identification result.

15. An electronic device, comprising: The electronic device includes: at least one processor; and a memory connected with the at least one processor in communication; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the data processing method in any one of claims 1-13.

16. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for enabling the processor to implement the data processing method in any one of claims 1-13 when executed.