Call content dynamic auditing method
By using a graph neural network model to perform deep feature learning on user call audio data and dynamically adjusting risk assessment parameters, the problem of policy mismatch caused by changes in user risk levels in the 5G content review system is solved, achieving efficient and accurate user risk assessment and review.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-31
- Publication Date
- 2026-03-27
AI Technical Summary
The existing 5G content moderation system cannot respond to the dynamic changes in users' risk levels in real time, resulting in a mismatch between the moderation strategy and the actual risk of users, which affects the accuracy and efficiency of the moderation.
By acquiring users' call audio data, deep feature learning is performed using a graph neural network model to generate a high-dimensional embedded feature representation of the user. Combined with risk assessment parameters, the sampling frequency, sampling duration, and risk judgment threshold are dynamically adjusted to achieve dynamic review.
It has increased the rigor of the review process for high-risk users and optimized the review efficiency for low-risk users, thus achieving a reasonable allocation of resources and greater accuracy in the review process.
Smart Images

Figure CN121750784A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of artificial intelligence review, in particular to a call content dynamic review method. BACKGROUND
[0002] Media resource security review is a technology that monitors and analyzes the transmitted media content in real time during communication, aiming to ensure the compliance and safety of the content and prevent the spread of harmful content. With the development of 5G technology, the transmission speed and quality of audio resources have been significantly improved, which makes the security review of audio resources particularly important.
[0003] Current solutions generally adopt a two-step approach to process and analyze audio content. First, through voice recognition technology (ASR, Automatic Speech Recognition), audio data is converted into text format. Next, natural language processing (NLP, Natural Language Processing) algorithms are used to analyze the generated text in depth.
[0004] In this process, NLP algorithms can screen for sensitive words, inappropriate content, promotional information, or political topics in the text. It can also automatically mark or filter out specific types of risks or inappropriate content according to a pre-set rule set. In addition, with the help of sentiment analysis, a branch of NLP technology, algorithms can assess the emotional tendencies in audio, such as anger, excitement, sadness, and other emotional states, to help identify potential risk signals such as extreme speech or harassment behavior.
[0005] Through this combination of voice recognition and natural language processing technology, audio content can be effectively monitored and managed to ensure compliance with relevant regulations and platform policies, while providing a safer and more harmonious communication environment for users.
[0006] Media resource security review is a technology that monitors and analyzes the transmitted media content in real time during communication, aiming to ensure the compliance and safety of the content and prevent the spread of harmful content. With the development of 5G technology, the transmission speed and quality of audio resources have been significantly improved, which makes the security review of audio resources particularly important. However, the existing 5G content review mainly has the following defects: the risk level of users may change dynamically over time and behavior patterns (such as users changing from low risk to high risk), but static review parameters cannot respond to such changes in real time, resulting in a mismatch between review strategies and actual user risks, affecting review accuracy and efficiency. SUMMARY
[0007] To achieve the above purpose, the present disclosure adopts the following technical solutions: One aspect of this disclosure provides a method for dynamic review of call content, comprising the following steps: Obtain the user's call audio data; Based on call audio data and risk assessment parameters, a user's risk score is determined; and Adjust risk assessment parameters based on risk scores; The risk assessment parameters include one or more of the following: the sampling frequency of the call audio data, the sampling duration of the call audio data, and the risk judgment threshold.
[0008] In one alternative implementation, determining a user's risk score includes: Collect user data, which includes basic attribute data, online behavior data, and historical interaction data; A user relationship network is constructed based on user data, and basic attribute data and network behavior data are used as attributes of user nodes in the relationship network. By using a graph neural network model to perform deep feature learning on user relationship networks and node attributes, a high-dimensional embedded feature representation of users is generated. By matching high-dimensional embedded feature representations with a pre-built user vector library, a user's risk score can be determined.
[0009] In one alternative implementation, determining a user's risk score further includes: Convert call audio data into text information; Perform semantic understanding and feature extraction on text information to generate text embedding vectors; Input the text embedding vector into the violation content classification model to obtain the content violation probability; The probability of content violations is compared with the risk assessment threshold to determine the risk score corresponding to the call audio data.
[0010] In one optional implementation, adjusting risk assessment parameters based on risk scores includes: Based on risk scoring, the user's risk level is determined; Adjust risk assessment parameters based on risk level.
[0011] In one optional implementation, risk assessment parameters are adjusted based on the risk level, including: When the risk level is a first predetermined level, a risk assessment parameter is set as the first assessment parameter. When the risk level is a second predetermined level, a risk assessment parameter is set as the second assessment parameter. The first predetermined level is higher than the second predetermined level. The first assessment parameter includes a first sampling frequency, a first sampling duration, and a first risk judgment threshold. The second assessment parameter includes a second sampling frequency, a second sampling duration, and a second risk judgment threshold. The first sampling frequency is higher than the second sampling frequency, and / or the first sampling duration is longer than the second sampling duration, and / or the first risk judgment threshold is lower than the second risk judgment threshold.
[0012] In one alternative implementation, the user data may also include one or more of the following: historical violation records, correlation with high-risk areas / users, device and IP address stability, social network structural risks, and user feedback and complaint records.
[0013] In one optional implementation, it further includes: User data is updated at predetermined time intervals based on the user's risk score.
[0014] Another aspect of this disclosure provides a dynamic call content review device, comprising: The audio data acquisition module is configured to acquire the user's call audio data. The risk scoring determination module is configured to determine the user's risk score based on the call audio data and risk assessment parameters; and The risk assessment parameter adjustment module is configured to adjust the risk assessment parameters based on the risk score. The risk assessment parameters include one or more of the following: the sampling frequency of the call audio data, the sampling duration of the call audio data, and the risk judgment threshold.
[0015] In another aspect of this disclosure, an electronic device is provided, comprising: At least one memory stores computer-executable instructions non-transiently; At least one processor configured to run computer-executable instructions, The aforementioned method for dynamic review of call content is implemented by the execution of computer-executable instructions by the processor.
[0016] In another aspect, this disclosure provides a computer-readable storage medium storing computer-executable instructions that, when executed by at least one processor, implement the aforementioned dynamic call content review method.
[0017] Technical effects of this disclosure: According to the dynamic call content review method disclosed herein, a user's risk score is determined based on the user's call audio data and risk assessment parameters. Furthermore, the risk assessment parameters are adjusted based on the risk score, thereby achieving dynamic adjustment of the review mechanism. This increases the rigor of review for high-risk users and optimizes the review efficiency for low-risk users, achieving a more rational allocation of resources. Attached Figure Description
[0018] The accompanying drawings are provided to further illustrate the present disclosure and form part of the specification. They are used together with the embodiments of the present disclosure to explain the disclosure and do not constitute a limitation thereof. In the drawings: Figure 1 A flowchart of a method for dynamic review of call content provided in this embodiment of the present disclosure; Figure 2 A framework diagram of a 5G call content dynamic review system that integrates user profiling and risk assessment, provided in an embodiment of this disclosure; Figure 3 A block diagram of an electronic device provided in an embodiment of this disclosure; Figure 4 A block diagram of a computer-readable storage medium provided for embodiments of this disclosure. Detailed Implementation
[0019] The technical solutions of the present disclosure will be described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present disclosure, and not all embodiments.
[0020] In the following description, the terms "first," "second," etc., are used for descriptive convenience only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined with "first," "second," etc., may explicitly or implicitly include one or more of that feature. In the description of this disclosure, unless otherwise stated, "a plurality of" means two or more.
[0021] In this disclosure, unless otherwise expressly specified and limited, the term "connection" should be interpreted broadly. For example, "connection" can be a fixed mechanical connection, a detachable mechanical connection, or an integral part; or, "connection" can be a direct connection or an indirect connection through an intermediate medium. Furthermore, unless otherwise expressly specified and limited, the term "coupling" should be interpreted broadly. For example, "coupling" can be a direct electrical connection, such as physical contact and electrical conduction between two components; it can also be understood as an electrical connection between different components in a circuit structure through physical lines capable of transmitting electrical signals, such as copper foil or wires on a printed circuit board (PCB), to transmit electrical signals; or, "coupling" can be an indirect electrical connection between two components through an intermediate medium; or, "coupling" can be an electrical connection between two components in a non-contact manner, such as an electrical connection between two components using capacitive coupling to transmit electrical signals.
[0022] In this embodiment of the disclosure, directional terms such as "up," "down," "left," and "right" may be defined relative to the orientation in which the components are schematically placed in the accompanying drawings. It should be understood that these directional terms can be relative concepts, used for relative description and clarification, and can change accordingly depending on the orientation in which the components are placed in the accompanying drawings.
[0023] like Figure 1 As shown in the figure, this disclosure provides a method for dynamic review of call content, which includes the following steps: S100: Acquire user's call audio data; S200: Determines the user's risk score based on call audio data and risk assessment parameters; S300: Adjust risk assessment parameters based on risk scores.
[0024] Specifically, the risk assessment parameters include one or more of the following: the sampling frequency of the call audio data, the sampling duration of the call audio data, and the risk determination threshold.
[0025] In the above embodiments, user data collection and feature vector generation involve the system acquiring relevant data about the target user, including: Basic attribute data: age, real-name authentication information, and package type; Network behavior data: commonly used IP addresses, device models; Historical interaction data: call records for the past 6 months, historical violation records, and association with high-risk users; Supplementary data: social network structure, user feedback and complaint records.
[0026] Based on the above data, a user relationship network is constructed. A graph neural network model is used to perform deep feature learning on the network structure and node attributes, generating a 128-dimensional high-dimensional embedding feature vector for each user.
[0027] Once the risk score is determined, the generated 128-dimensional feature vector is matched against a pre-built user vector library, and the cosine similarity is calculated. Assuming the matching result shows a 92% similarity between the user and the low-risk user vector library, the system outputs an initial risk score of 60 points (out of 100), with higher scores indicating higher risk.
[0028] The review parameters are dynamically adjusted, and the risk level is classified as low risk based on the risk score, and a preset low-risk review strategy is matched accordingly. Sampling frequency: Random sampling once every 100 calls; Sampling duration: Each sampling extracts the first 30 seconds of audio after the start of the call; Risk assessment threshold: 0.8, meaning that a text is considered to be in violation when the probability of violation is ≥0.8.
[0029] Call content review and risk score update, call sampling and text conversion: When a user initiates a call, the system randomly triggers sampling according to a low-risk strategy, such as sampling the 50th call, collecting the audio of the first 30 seconds of the call, and converting it into text using ASR technology. Violation identification: The text was input into a binary classification model, and the model output a violation probability of 0.15. Since 0.15 is less than the risk assessment threshold of 0.8, the call content was determined not to be in violation.
[0030] User data and risk score updates: User data is updated every 24 hours. If a user continues to have no violations, their risk score is updated to 55 points, maintaining a low-risk level, and the review parameters remain unchanged.
[0031] Dynamic adjustments in high-risk scenarios, if the user subsequently exhibits the following behaviors: Frequent calls to high-risk overseas IP addresses; Two new user complaint records have been added to the historical interaction data; After updating user data, feature vectors are regenerated using a graph neural network. After matching with the user vector library, the risk score rises to 85 points, and the risk level is adjusted to high risk. At this point, the review parameters are dynamically adjusted as follows: Sampling frequency: 100% full call sampling; Sampling duration: The entire call was recorded; Risk assessment threshold: 0.3, meaning that a text is considered to be in violation when the probability of violation is ≥0.3.
[0032] If the probability of a violation in a certain call text detected by the model is 0.4 (≥0.3), the system will directly determine it as a violation and trigger a manual review process.
[0033] like Figure 1 As shown, based on Embodiment 1, the steps provided in this disclosure embodiment include: determining the user's risk score, which includes: Collect user data, which includes basic attribute data, network behavior data, and historical interaction data; A user relationship network is constructed based on the user data, and the basic attribute data and network behavior data are used as attributes of user nodes in the relationship network; A graph neural network model is used to perform deep feature learning on the user relationship network and node attributes to generate a high-dimensional embedded feature representation of the user. The high-dimensional embedded feature representation is matched with a pre-built user vector library to determine the user's risk score.
[0034] Specifically, determining the user's risk score further includes: Convert the call audio data into text information; The text information is semantically understood and its features are extracted to generate a text embedding vector; The text embedding vector is input into the violation content classification model to obtain the content violation probability; The probability of content violation is compared with the risk assessment threshold to determine the risk score corresponding to the call audio data.
[0035] Specifically, adjusting the risk assessment parameters based on the risk score includes: Based on the risk score, the user's risk level is determined; Based on the risk level, adjust the risk assessment parameters.
[0036] Specifically, adjusting the risk assessment parameters based on the risk level includes: When the risk level is a first predetermined level, the risk assessment parameter is set as a first assessment parameter; when the risk level is a second predetermined level, the risk assessment parameter is set as a second assessment parameter. The first predetermined level is higher than the second predetermined level. The first assessment parameter includes a first sampling frequency, a first sampling duration, and a first risk judgment threshold. The second assessment parameter includes a second sampling frequency, a second sampling duration, and a second risk judgment threshold. The first sampling frequency is higher than the second sampling frequency, and / or the first sampling duration is longer than the second sampling duration, and / or the first risk judgment threshold is lower than the second risk judgment threshold.
[0037] Specifically, the user data also includes one or more of the following: historical violation records, correlation with high-risk areas / users, device and IP address stability, social network structural risks, and user feedback and complaint records.
[0038] Specifically, it also includes: The user data is updated at predetermined time intervals based on the user's risk score.
[0039] The dynamic assessment is based on the user's real-time risk score, which is used to classify risk levels. Sampling strategies are set for different risk levels: high-risk users have higher sampling frequency and longer sampling duration, while low-risk users have lower sampling frequency and shorter sampling duration. Set differentiated violation thresholds: low thresholds for high-risk users and high thresholds for low-risk users. The sampling strategy and violation threshold are matched with the corresponding risk level based on the user's current risk score in real time, and the call content review process is dynamically adjusted. By implementing differentiated review processes, we can achieve strict monitoring of high-risk users and reduce redundant reviews for low-risk users, thereby optimizing the efficiency of review resource allocation.
[0040] In one optional implementation, the step of setting the risk assessment threshold includes establishing a risk score and parameter mapping rule, and presetting the correspondence between the risk score range and the sampling frequency, sampling duration, and risk assessment threshold. Based on the current range in which the user's risk score falls, retrieve the corresponding sampling frequency; Based on the matched risk score range, set the sampling duration for call audio data; Configure risk assessment thresholds based on risk scoring ranges; The matching sampling frequency, sampling duration, and risk judgment threshold are integrated into a dynamic review parameter group for the real-time review process of the current call content.
[0041] The graph neural network model adaptively learns the attention weights of user nodes in the relational network, thereby more accurately aggregating neighbor node information to update its own embedded features.
[0042] Specifically, in the process of generating the dynamic high-dimensional embedded feature representation, a joint loss function that integrates node attribute similarity and network structure similarity is used for model training to enhance the ability of embedded features to distinguish user risk characteristics.
[0043] Specifically, the user risk level assessment includes: Construct a user vector sample library containing labels of different risk levels. The user vectors in the sample library are obtained by training through manual annotation combined with multi-dimensional compliance data. The dynamic high-dimensional embedded feature representation of the user to be evaluated is matched with the sample vectors in the user vector sample library for similarity. Based on the similarity matching results and the risk level labels of samples in the sample library, a weighted calculation is performed to obtain the real-time risk score of the user to be evaluated.
[0044] Specifically, the multi-dimensional compliance data includes, but is not limited to, historical violation records, correlation with high-risk areas / users, device and IP address stability, social network structural risks, and user feedback and complaint records.
[0045] Specifically, the content sampling parameters include sampling frequency and sampling duration, which are positively correlated with the user's real-time risk score; the violation judgment threshold is negatively correlated with the user's real-time risk score, that is, the higher the risk score, the lower the judgment threshold, so as to improve the sensitivity of reviewing high-risk users.
[0046] Specifically, the intelligent review of 5G call content includes: Adaptive sampling of the call audio stream is performed based on the adjusted content sampling parameters; Convert the sampled audio segments into text information; The text information is semantically understood and features are extracted using a pre-trained natural language processing model to generate text embedding vectors; The text embedding vector is input into the violation content classification model to obtain the content violation probability; The probability of content violation is compared with the adjusted violation judgment threshold to determine whether the call content violates regulations.
[0047] Specifically, the automated intervention measures include, but are not limited to, real-time call blocking, marking and isolating illegal segments, sending risk warnings to users, and reporting high-risk user information and call records to the regulatory platform; The feedback update mechanism includes: when a user's call content is judged to be in violation, the system automatically updates the user's historical violation record attributes and triggers the graph neural network model to perform incremental or periodic recalculation of the embedded feature representation of the user and its associated users, so as to realize the dynamic evolution of user profiles and continuous optimization of risk assessment.
[0048] Specifically, it also includes a cold start mechanism for newly registered users or users with sparse data: using the user's initial registration information, device fingerprint, and preliminary characteristics of the first call, combined with a preset general risk rule base, a temporary risk score is generated and gradually optimized in subsequent interactions.
[0049] The above embodiments will introduce the following aspects: user profile data collection, graph neural network construction based on user profile data, embedded feature representation and risk assessment of user profiles, user call content review based on review parameters, automated decision-making and intervention mechanism, and dynamic adjustment of user profile features.
[0050] User profile data includes user characteristic data, user IP data, and user historical call data.
[0051] The system extracts caller profile data from the user database, including age, gender, occupation, and past violation records. This foundational data helps the system initially identify basic user characteristics, providing a basis for subsequent user profiling and risk scoring.
[0052] Obtain IP address and geolocation: The system determines a caller's geographical location using their IP address. It obtains the user's specific geographical location, such as city and country, through IP address mapping.
[0053] Extracting users' historical call data: Historical call data of users is extracted from the call log database to build a relationship network. This data includes information such as the recipients, frequency, and duration of each call.
[0054] The acquired data is used as user attribute data (node attributes), including age, gender, occupation, historical violation records, and IP geolocation. Extracted historical call data is used as relationship network data (edge attributes). This data includes information such as the recipient, frequency, and duration of each call, forming an interaction graph between users.
[0055] Graph network model construction, definition of graph network: Node definition: Each user is a node in the graph network, and the node's attribute data (such as age, gender, occupation, and historical violation records) are used as input features.
[0056] Edge definition: The call actions between users constitute the edges in the graph, and the weight of the edge is the normalized result of the total call duration between the call pairs. This results in higher edge weights for users who frequently have long calls.
[0057] Graph structure representation: Construct an undirected weighted graph G=(V,E), where V represents the set of user nodes and E represents the set of edges between users.
[0058] Model selection: Graph Attention Network (GAT) is used to analyze graph data. GAT can capture the complex relationships between node features and graph structure, making it very suitable for processing relational data.
[0059] For graph network model training, the input data includes: a node feature matrix X, containing attribute data for each user, such as age, gender, occupation, and historical violation records; and an adjacency matrix A, representing the relationships between users, where elements Aij represent the strength of the relationship between user i and user j, i.e., the edge weights in the graph.
[0060] Model training: Graph Convolutional Layer (GCN Layer): This layer performs a graph convolution operation on the node feature matrix X and the adjacency matrix A to learn the latent feature representation of each node. The formula for the graph convolutional layer operation is as follows:
[0061] Where H(l+1) represents the node feature matrix of the (l+1)th layer, and H(l) represents the node feature matrix of the th layer. When l=0, H(0)=X, which is the original node feature matrix. W(l) is the trainable weight matrix of the th layer, D is the degree matrix, and σ is a non-linear activation function such as the sigmoid function.
[0062] Multi-layer graph convolution: Multi-layer graph convolutional networks can stack multiple layers of GCN, with each layer extracting deeper feature representations, thereby improving the model's ability to learn complex relationships.
[0063] The domain similarity loss method is employed: maximizing the feature similarity between similar nodes (such as neighboring nodes) while minimizing the feature similarity between dissimilar nodes. This can be optimized using negative sampling techniques, as shown in the following formula:
[0064] in, and Let be the embedding vectors of nodes i and j, respectively, σ be the sigmoid function, and E represent the true neighbor pair. E represents a negative sample pair.
[0065] User profile embedding feature representation and risk assessment, user embedding feature representation: Through multi-layer graph convolution operations, the system generates the final node embedding matrix H(L), where each row represents a user's high-dimensional embedding vector, i.e., the user's embedding feature representation. This embedding feature representation comprehensively considers the user's attribute features (such as age, gender, occupation, and historical violation records) as well as the user's relationship network information with other users. This high-dimensional vector can capture the user's position and characteristics in the entire network, providing a crucial foundation for subsequent risk assessment.
[0066] To assess the risk level of new users, the system matches the embedded feature representation of the new users with a pre-built user vector library to generate a risk score.
[0067] Building a User Vector Library: To improve the accuracy of the system's user risk assessment, 1% of user data was first selected as a representative sample. These user nodes contain rich historical data, representing user groups with different risk levels. Next, a small-scale but high-quality labeled database will be built based on the information from these nodes, providing crucial support for subsequent similarity matching. First, multi-dimensional data related to user compliance is extracted; this data can reflect the risk of illegal or irregular behavior by users during calls. For example: Historical violation record: The number of times the user has violated regulations in historical calls, directly reflecting the user's risk.
[0068] User call correlation with high-risk areas: Frequent calls by users to phone numbers in specific high-risk areas (such as areas with high fraud rates) may also be a warning sign of potential violations.
[0069] Device information and IP address change records: Frequent changes in devices or IP addresses, especially cross-border or abnormal geographical jumps, may indicate that the user is circumventing system monitoring or attempting to engage in illegal activities.
[0070] Social network activity and relationship chains: If a user has frequent contact or interaction with multiple known high-risk users, it may indicate potential compliance risks.
[0071] Historical complaints and user feedback records: Whether there are other users' complaints about this user in the system, especially complaints about violations or illegal behavior during calls, can serve as an important reference for risk assessment.
[0072] Based on the collected multi-dimensional compliance data, manual annotation was performed. The annotation standard was a risk score from 0 to 1, where 0 indicates an extremely low risk of violation and 1 indicates an extremely high risk of violation. Annotators were trained to combine different dimensions to comprehensively score the overall compliance of users. This comprehensive score serves as the risk score label for the sample users.
[0073] Next, based on the collected user profile data of the sample users, feature engineering will be used to construct multidimensional feature data of the sample users.
[0074] For example, user profile data includes the following characteristics: Number of violations: 10; Is it related to calls from high-risk areas: Yes; Do you frequently change devices or IP addresses? Yes; Are there any risks associated with social network activity and relationship chains? No; Number of historical complaints and user feedback records: 8; Feature engineering can be used to construct multidimensional data, which can be represented as [10,1,1,0,8].
[0075] The risk score labels and multidimensional feature data of the labeled sample users are stored in a user vector database, forming a small sample labeled database. This database will be used for subsequent similarity matching to complete the risk assessment of new users.
[0076] The system calculates the similarity between the embedded feature representation of a new user and each feature vector in the user vector database. To measure the similarity between two vectors, the system uses cosine similarity. Cosine similarity measures the angle between two vectors in vector space; the closer the value is to 1, the more similar the two vectors are. The formula for calculating cosine similarity is:
[0077] in, This represents the embedded feature representation of a new user. This represents the feature vector in the user's vector library.
[0078] The system sorts the new user's similarity to all users in the user vector database and selects the top three users with the highest similarity, i.e., the top-3 users. The feature vectors of these three users are most similar to the new user's embedded feature representation, providing valuable reference for the new user's risk scoring.
[0079] The system averages the risk scores of the top 3 users and uses this average as the risk score ruser for new users.
[0080] In this way, the system can generate a reasonable risk score for new users based on information from the existing user vector database. This score comprehensively considers the user's attribute characteristics, relationship network, and similarity with other users, thus providing a basis for subsequent review strategies.
[0081] In the content moderation system, we can dynamically adjust the moderation parameters based on each user's risk score, thereby influencing the moderation process and achieving more accurate and personalized content moderation. The formula is expressed as follows:
[0082]
[0083] Where fs and ts are the sampling frequency and sampling duration, respectively, and ruser is the user's risk score. This represents the embedded text features of the call content, pviolation represents the probability that the text is judged as a violation, and Tviolation is the threshold for judging whether the content is compliant. fs, ts, and Tviolation are all functions of Russel.
[0084] The following are the specific implementation steps and a detailed explanation of the formula: The system first acquires audio data from the user's call, with the sampling frequency and duration dynamically adjusted based on the user's risk score. The sampling frequency determines the number of samples acquired per second from the analog audio signal. A higher sampling frequency means more data points, resulting in a larger file size. For example, a sampling rate of 44.1 kHz means 44,100 samples per second, while a sampling rate of 96 kHz means 96,000 samples per second. The sampling duration refers to the length of time the audio signal is captured each time; this parameter is typically measured in milliseconds (ms) or microseconds (μs).
[0085] For users with higher risk scores, the system will use a higher sampling frequency and a longer sampling duration to ensure that more audio data is captured for review. For users with lower risk scores, the system can reduce the sampling frequency and duration to improve review efficiency.
[0086] The formulas for adjusting the sampling frequency fs and the sampling duration ts are as follows:
[0087] Where fbase and tbase are the base sampling frequency and base sampling duration, respectively, ruser is the user's risk score, and α and β are adjustment factors.
[0088] After acquiring the audio data, the system performs speech-to-text conversion using the Whisper model. The Whisper model can process speech data and efficiently convert it into text. The text generated in this step provides input for subsequent analysis.
[0089] After being converted into text, the system uses a pre-trained BERT model to embed the text, transforming it into a high-dimensional vector. This high-dimensional vector effectively captures the semantic and contextual information of the text.
[0090] The training data uses a labeled text dataset, where samples are labeled as "violation" or "compliance." For example, texts involving sensitive, terrorist, political, or pornographic content are considered violations. The text data is processed using Whisper speech-to-text and BERT embedding, and the resulting high-dimensional vectors are used as input to the classification model.
[0091] This binary classification model employs a linear classifier structure, as shown in the following formula:
[0092] Where W is the model's weight matrix, b is the bias, σ is the activation function (such as the sigmoid function), and pviolation represents the probability that the text is judged as a violation. This indicates that the call content is embedded with text features.
[0093] During training, labeled data is used to optimize W and b by minimizing the cross-entropy loss to improve the model's prediction accuracy.
[0094] The system uses a pre-trained binary classification model to analyze text embedding vectors and determine whether the content is in violation of regulations. This binary classification model receives the input embedding vector and outputs a probability value representing the probability that the text belongs to a violation category.
[0095] Based on users' risk scores, the system dynamically adjusts the threshold for determining inappropriate content. For users with higher risk scores, the system uses a lower threshold to increase the rigor of the review process. The formula for adjusting the threshold Tviolation can be expressed as:
[0096] Where Tbase is the basic decision threshold, and γ is the adjustment factor. The exponential decay function causes the threshold to decrease rapidly as the risk score increases, thereby strengthening the screening of high-risk users.
[0097] The system makes a review decision based on the output probability pviolation of the classification model and the dynamically adjusted judgment threshold Tviolation. If pviolation ≥ Tviolation, the content is judged to be in violation, and the system will take measures such as blocking or reporting; otherwise, the content passes the review.
[0098] Based on the above content compliance determination results, the system takes further action. When a user's call content is determined to be non-compliant, the following intervention measures are taken: Content blocking: Once the call content is determined to be non-compliant, the system immediately blocks the relevant content to prevent its spread, such as interrupting the call, prohibiting recording and external transmission, etc. Warning notification: The system can automatically send warning notifications to users, informing them that the relevant content has been marked as high-risk.
[0099] Reporting Management: For recurring violations, the system can report relevant records and logs to the management department or service provider for further review and processing.
[0100] The system periodically updates user attribute data, acquiring the latest user profile data, including IP address, geographic location, and call data, at set fixed time intervals (e.g., monthly or bi-weekly). By periodically updating the user's node attributes in the graph neural network, the system can promptly reflect dynamic changes in the user. This update process integrates the latest user attribute information (such as new call records and changed geographic locations) into the graph network, allowing the embedded feature representation of the user profile to be recalculated based on the latest information. Through such periodic updates, it is ensured that the embedded feature representation of the user profile always accurately reflects the user's current state and behavior.
[0101] For users newly added during the update cycle, the system temporarily uses the default threshold Tviolation for processing. The initial review of these new users may not be as accurate as the personalized review of long-term users, but the default parameter settings ensure the effectiveness of the review to a certain extent. Upon the next data update, the attribute data of the new users will be incorporated into the graph neural network model. The system will then recalculate the feature vectors of these users based on the latest information and further optimize the review parameters and risk assessment strategies.
[0102] This method of periodically updating user attribute data saves system resources, avoids the computational overhead of frequent updates, and also utilizes the latest user information to improve the accuracy of the review process. Through reasonable resource allocation and an efficient update mechanism, the system can ensure both efficient review and accurate results.
[0103] The feedback mechanism for review results impacts a user's violation record with each review, dynamically adjusting the user's risk score. Specifically, when a user's call content is deemed a violation during the review process, the system records this violation in the user's "Historical Violation Count" attribute. This attribute directly serves as a feature of nodes in the graph network, affecting the node's attribute data. Therefore, at the next update time, changes in the user's violation record will cause updates to the user node features in the graph network, leading to changes in the user profile feature vector. As the number of user violations increases, the user's risk score will also be adjusted accordingly, further influencing review parameters and strategies. This feedback mechanism ensures that the user profile dynamically adapts to changes in user behavior, reflecting the latest risk level.
[0104] like Figure 2 As shown in the figure, this disclosure provides a call content dynamic review device 200, including: The audio data acquisition module 201 is configured to acquire the user's call audio data; Risk scoring module 202 is configured to determine a user's risk score based on call audio data and risk assessment parameters; and The risk assessment parameter adjustment module is configured to adjust risk assessment parameters based on risk scores.
[0105] Specifically, the risk assessment parameters include one or more of the following: the sampling frequency of the call audio data, the sampling duration of the call audio data, and the risk determination threshold.
[0106] Specifically, the audio data acquisition module 201 may include: The user attribute feature submodule collects the user's age, gender, occupation, historical violation records, and device fingerprint information; The network connectivity characteristics submodule collects users' IP address, geographical location, IP address change frequency, network access method, and stability. The historical interaction behavior submodule collects the user's historical call object identifier, call frequency, call duration, call time period, risk correlation of the call object, and historical complaint records.
[0107] Specifically, the risk scoring determination module 202 may include: The graph structure definition submodule defines each user as a node in the graph network, with node attributes including user attribute features and network connection features; and defines the call interaction between users as an edge in the graph network, with the edge weight dynamically calculated and normalized by integrating call frequency, call duration and risk correlation of the call object. The GAT model training submodule employs a multi-layer graph attention network. Each layer calculates the attention weights of a node to its neighboring nodes and aggregates neighbor features to update its own feature representation. During training, an adaptive negative sampling mechanism for domain similarity loss function is introduced to optimize the discriminativeness and robustness of node embeddings.
[0108] Specifically, the risk scoring and determination module 202 may further include: The risk indicator library construction submodule builds a multi-dimensional risk indicator system that includes historical violation records, high-risk area call correlation, device and IP abnormal change index, social network risk propagation coefficient, and user complaint weighting value. The representative user vector library construction submodule uses a combination of density peak clustering and active learning to select and label representative user samples from all users and build a labeled user vector library. The dynamic risk score calculation submodule performs multi-scale similarity calculations between the high-dimensional embedding vectors of new users or users to be evaluated and the sample vectors in the user vector library. It then combines the real-time weights of risk indicators in each dimension and generates a real-time dynamic risk score for the user through weighted voting or neural network regression.
[0109] Specifically, the risk assessment parameter adjustment module 203 may include: The audio sampling parameter adjustment submodule dynamically adjusts the audio sampling frequency and single sampling duration based on the user's risk score. Higher-risk users are assigned a higher sampling frequency and a longer sampling duration. The speech-to-text accuracy control submodule dynamically selects different accuracy levels of speech-to-text models or adjusts model parameters based on the user's risk score. Higher-risk users trigger higher-precision speech recognition. The semantic understanding depth adjustment submodule dynamically adjusts the depth of text semantic analysis based on the user's risk score. For example, for high-risk users, a more complex contextual understanding model and entity relationship extraction are enabled. The dynamic threshold generation submodule dynamically generates or adjusts violation judgment thresholds based on user risk scores, current network environment risk levels, and historical false judgment rates using an adaptive algorithm.
[0110] Specifically, the risk scoring determination module 202 may also include: The real-time speech-to-text submodule uses a speech recognition model that supports low latency and high accuracy to convert speech into text streams in real time. The text embedding and semantic understanding submodule uses a pre-trained language model to perform deep semantic embedding on the converted text, and combines it with a domain knowledge base to perform entity recognition, intent recognition and sensitive information detection. The dynamic threshold comparison and judgment submodule compares the text semantic understanding results with the dynamic judgment thresholds from the dynamic review strategy engine, and combines them with contextual semantic coherence analysis to output the final content compliance judgment result.
[0111] Specifically, the risk scoring determination module 202 may also include: The tiered intervention strategy submodule implements tiered intervention measures, including but not limited to real-time content blocking, call interruption, warning notifications, temporary account restrictions, archiving of violation records, and reporting to the management department, based on the severity of the violation, the user's historical risk records, and the level of confidence. The closed-loop feedback learning submodule uses the results of this review, intervention measures, and subsequent user behavior data as new training samples or update signals, feeding them back to the dynamic graph network construction and deep feature learning module and the user risk intelligent assessment module. This triggers incremental updates or periodic retraining of the user profile embedded features, enabling continuous self-optimization of the system.
[0112] Figure 3 A block diagram of an exemplary electronic device suitable for implementing embodiments of the present disclosure is shown.
[0113] The electronic device may include a central processing unit / microprocessor / main control chip, etc. 4; and a storage medium 5, coupled to the central processing unit / microprocessor / main control chip, etc. 4, and storing computer-executable instructions therein for performing the steps of the various methods of the embodiments of this disclosure when executed by the processor.
[0114] The central processing unit / microprocessor / main control chip, etc., can include, but are not limited to, one or more processors or microprocessors.
[0115] Storage medium 5 may include, but is not limited to, random access memory (RAM), read-only memory (ROM), flash memory, EPROM memory, EEPROM memory, registers, computer storage media (e.g., hard disk, floppy disk, solid-state drive, removable disk, CD-ROM, DVD-ROM, Blu-ray disc, etc.).
[0116] In addition, the electronic device may also include (but is not limited to) a data bus 6, an input / output bus / external bus / device bus 7, a display 8, and input / output devices 9 (e.g., keyboard, mouse, speaker, etc.).
[0117] The central processing unit / microprocessor / main control chip, etc. 4 can communicate with external devices (8, 9, etc.) via I / O bus 7 through wired or wireless network (not shown).
[0118] The storage medium 5 may also store at least one computer-executable instruction for performing the steps of various functions and / or methods in the embodiments described herein when the central processing unit / microprocessor / main control chip, etc., 4 is running.
[0119] In one embodiment, the at least one computer-executable instruction may also be compiled into or comprise a software product, wherein one or more computer-executable instructions are executed by a processor to perform the steps of the various functions and / or methods in the embodiments described herein.
[0120] Figure 4 A schematic diagram of a computer-readable storage medium according to an embodiment of the present disclosure is shown.
[0121] like Figure 4As shown, the non-transitory computer-readable storage medium 11 stores instructions, such as computer-readable instructions 10. When the computer-readable instructions 10 are executed by a processor, the various methods described above can be performed. The non-transitory computer-readable storage medium includes, but is not limited to, volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-transitory non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. For example, the non-transitory computer-readable storage medium 11 can be connected to a computing device such as a computer, and then, when the computing device executes the computer-readable instructions 10 stored on the computer-readable storage medium 11, the various methods described above can be performed.
[0122] In the several embodiments provided in this disclosure, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0123] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0124] Furthermore, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0125] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for executing all or part of the steps of the methods of the various embodiments of this disclosure through a computer device (which may be a personal computer, server, or network device, etc.). The aforementioned storage medium includes: USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, optical disks, and other media capable of storing program code.
[0126] The above embodiments are only used to illustrate the technical solutions of this disclosure, and are not intended to limit it. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this disclosure.
Claims
1. A method for dynamic verification of call content, characterized in that, include: Obtain the user's call audio data; Based on the call audio data and risk assessment parameters, the user's risk score is determined; as well as Based on the risk score, adjust the risk assessment parameters; The risk assessment parameters include one or more of the following: the sampling frequency of the call audio data, the sampling duration of the call audio data, and the risk judgment threshold.
2. The method for dynamic review of call content as described in claim 1, characterized in that, Determining the user's risk score includes: Collect user data, which includes basic attribute data, network behavior data, and historical interaction data; A user relationship network is constructed based on the user data, and the basic attribute data and network behavior data are used as attributes of user nodes in the relationship network; A graph neural network model is used to perform deep feature learning on the user relationship network and node attributes to generate a high-dimensional embedded feature representation of the user. The high-dimensional embedded feature representation is matched with a pre-built user vector library to determine the user's risk score.
3. The method for dynamic review of call content as described in claim 2, characterized in that, Determining the user's risk score further includes: Convert the call audio data into text information; The text information is semantically understood and its features are extracted to generate a text embedding vector; The text embedding vector is input into the violation content classification model to obtain the content violation probability; The probability of content violation is compared with the risk assessment threshold to determine the risk score corresponding to the call audio data.
4. The method for dynamic review of call content as described in claim 1, characterized in that, The adjustment of the risk assessment parameters based on the risk score includes: Based on the risk score, the user's risk level is determined; Based on the risk level, adjust the risk assessment parameters.
5. The method for dynamic review of call content as described in claim 4, characterized in that, The adjustment of the risk assessment parameters based on the risk level includes: When the risk level is a first predetermined level, the risk assessment parameter is set as a first assessment parameter; when the risk level is a second predetermined level, the risk assessment parameter is set as a second assessment parameter. The first predetermined level is higher than the second predetermined level. The first assessment parameter includes a first sampling frequency, a first sampling duration, and a first risk judgment threshold. The second assessment parameter includes a second sampling frequency, a second sampling duration, and a second risk judgment threshold. The first sampling frequency is higher than the second sampling frequency, and / or the first sampling duration is longer than the second sampling duration, and / or the first risk judgment threshold is lower than the second risk judgment threshold.
6. The method for dynamic review of call content as described in claim 2, characterized in that, The user data also includes one or more of the following: historical violation records, correlation with high-risk areas / users, device and IP address stability, social network structure risks, and user feedback and complaint records.
7. The method for dynamic review of call content as described in claim 6, characterized in that, Also includes: The user data is updated at predetermined time intervals based on the user's risk score.
8. A dynamic call content verification device, characterized in that, include: The audio data acquisition module is configured to acquire the user's call audio data. The risk score determination module is configured to determine the user's risk score based on the call audio data and risk assessment parameters; as well as The risk assessment parameter adjustment module is configured to adjust the risk assessment parameters based on the risk score. The risk assessment parameters include one or more of the following: the sampling frequency of the call audio data, the sampling duration of the call audio data, and the risk judgment threshold.
9. An electronic device, comprising: At least one memory stores computer-executable instructions non-transiently; At least one processor, configured to run the computer-executable instructions, The characteristic feature is that the computer-executable instructions are executed by the processor to implement the dynamic call content review method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by at least one processor, implement the dynamic call content review method according to any one of claims 1-7.