Communication content identification method and device, electronic equipment and storage medium
By obtaining and analyzing communication voice data and resource scheduling maps in real time, and dynamically determining the target server, the problem of static computing resource selection strategies in the existing technology is solved, rapid response and efficient adaptation are achieved, and the quality of voice transcription and user experience are improved.
Patent Information
- Application Number
- CN202510443039.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-06-24
AI Technical Summary
The existing technology lacks real-time monitoring and feedback mechanisms in the analysis and processing of voice-transfer data, resulting in static computing resource selection strategies, which are not flexible and intelligent enough, cannot respond quickly to high loads or emergencies, and is difficult to adapt to changing communication environments and task requirements, resulting in low resource utilization efficiency and affecting service quality, efficiency and user experience.
By obtaining the communication voice data and resource scheduling map at the current moment, using the Digestella algorithm and real-time resource status monitoring, the target cloud server and edge server are dynamically determined, real-time voice translator is realized, and response speed and adaptability are improved.
It realizes rapid response in high load or emergency situations, improves adaptability to changing communication environments and task requirements, improves the quality, efficiency and user experience of voice transcription, and realizes flexible and stable communication content recognition.
Smart Images

Figure CN120199251A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of network communication technologies, and in particular, to a method, apparatus, electronic device, and computer-readable storage medium for identifying communication content. Background Art
[0002] With the development of communication technologies, especially the popularization of mobile Internet and 5G (5th Generation Mobile Communication Technology) technologies, the analysis and processing of voice-to-text data in the communication industry are crucial for extracting valuable information.
[0003] Currently, the analysis and processing of voice-to-text data usually rely solely on a central server or an edge server, and there is a lack of a mechanism for real-time monitoring and feedback on the performance of the central server or the edge server. The selection strategy of computing resources is static, not flexible and intelligent enough, resulting in the inability of voice-to-text to respond quickly under high load or urgent tasks; in addition, it is difficult to adapt to the ever-changing communication environment and task requirements, leading to excessive consumption of computing resources when facing simple voice-to-text tasks, and possible inability to ensure the processing effect due to insufficient resources when processing complex voice-to-text tasks, thus failing to achieve efficient utilization of resources and affecting the quality, efficiency, and user experience of the overall service. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide, in view of the above deficiencies of the prior art, a method, apparatus, electronic device, and computer-readable storage medium for identifying communication content, which can achieve flexible and stable identification of communication content, improve the response speed of communication content identification under high load or urgent tasks, improve the adaptability of communication content identification to the ever-changing communication environment and task requirements, and further improve the quality, efficiency, and user experience of the overall service of communication content identification.
[0005] In a first aspect, the present invention provides a method for identifying communication content, including: obtaining communication voice data and a resource scheduling diagram at a current moment, where the communication voice data includes key communication voice data; determining a target cloud server based on the resource scheduling diagram at the current moment and the Dijkstra algorithm; and performing voice-to-text conversion on the key communication voice data at the current moment based on a first model of the target cloud server.
[0006] Preferably, the communication voice data further includes non-critical communication voice data. The obtaining of the communication voice data and the resource scheduling diagram at the current moment specifically includes: obtaining the communication data, the first operating state, and the second operating state at the current moment, where the first operating state refers to the operating states of all cloud servers, and the second operating state refers to the operating states of all edge servers; performing data annotation and classification on the communication data to obtain the critical communication voice data and the non-critical communication voice data at the current moment; and constructing the resource scheduling diagram at the current moment based on the first operating state, the second operating state, the critical communication voice data, and the non-critical communication voice data.
[0007] Preferably, the resource scheduling diagram includes a first task node, a second task node, a first resource node, a second resource node, a first path, a second path, a first weight, and a second weight. The constructing of the resource scheduling diagram at the current moment based on the first operating state, the second operating state, the critical communication voice data, and the non-critical communication voice data specifically includes: respectively evaluating the priorities corresponding to the critical communication voice data and the non-critical communication voice data, and classifying the critical communication voice data and the non-critical communication voice data based on the priorities, where the priorities include the sorting of urgency and the sorting of importance; generating a first task node, a second task node, a first resource node, a second resource node, a first path, and a second path, where the first task node refers to the critical communication voice data corresponding to different priorities, the second task node refers to the non-critical communication voice data corresponding to different priorities, the first resource node refers to the cloud servers for processing the critical communication voice data corresponding to different priorities, the second resource node refers to the edge servers for processing the non-critical communication voice data corresponding to different priorities, the first path refers to the directed edge from the first task node to the first resource node, and the second path refers to the directed edge from the second task node to the second resource node; calculating the first weight and the second weight based on the first operating state, the second operating state, the critical communication voice data, and the non-critical communication voice data to obtain the resource scheduling diagram at the current moment, where the first weight refers to the execution time, energy consumption, and resource load of the first path, and the second weight refers to the execution time, energy consumption, and resource load of the second path.
[0008] Preferably, determining the target cloud server based on the resource scheduling graph and Dijkstra's algorithm at the current moment specifically includes: Based on Dijkstra's algorithm, determining whether there are conflicting resource nodes in the resource scheduling graph at the current moment, where a conflicting resource node refers to a first resource node that has an optimal path with multiple first task nodes, and the optimal path refers to the first path corresponding to the minimum value of the first weight; in response to the absence of conflicting resource nodes in the resource scheduling graph at the current moment, determining that the target cloud server includes all the first resource nodes corresponding to the optimal paths in the resource scheduling graph at the current moment; in response to the presence of conflicting resource nodes in the resource scheduling graph at the current moment, predicting the resource scheduling graph within a preset time period, and determining that the target cloud server includes all the first resource nodes corresponding to the optimal paths in the resource scheduling graph within the preset time period.
[0009] Preferably, predicting the resource scheduling graph within a preset time period specifically includes: determining non-optimized resource nodes and resource nodes to be optimized, where a non-optimized resource node refers to a conflicting resource node corresponding to a first task node with a high priority in the resource scheduling graph at the current moment, and a resource node to be optimized refers to a conflicting resource node corresponding to a first task node with a low priority in the resource scheduling graph at the current moment; predicting the third operating state within a preset time period, where the third operating state refers to the operating state of the non-optimized resource nodes; based on the third operating state within the preset time period, adjusting the first weight of the resource nodes to be optimized in the resource scheduling graph at the current moment to obtain the resource scheduling graph within the preset time period.
[0010] Preferably, after obtaining the communication voice data and the resource scheduling graph at the current moment, the communication content recognition method further includes: determining the target edge server based on the resource scheduling graph and Dijkstra's algorithm at the current moment; performing speech transcription on the non-critical communication voice data at the current moment based on the second model of the target edge server.
[0011] Preferably, before performing speech transcription on the non-critical communication voice data at the current moment based on the second model of the target edge server, the communication content recognition method further includes: obtaining historical communication voice data; constructing and training the first model of the target cloud server based on the historical communication voice data and the adaptive learning rate adjustment strategy; performing distillation processing on the first model of the target cloud server based on the teacher-student training method and the combined loss function of hard labels and soft labels to obtain the second model of the target edge server.
[0012] In a second aspect, the present invention further provides a recognition device for communication content, including a first acquisition module, a first determination module, and a first speech transcription module. The first acquisition module is configured to acquire communication voice data and a resource scheduling diagram at the current moment. The communication voice data includes key communication voice data. The first determination module is connected to the first acquisition module and is configured to determine a target cloud server based on the resource scheduling diagram at the current moment and the Dijkstra algorithm. The first speech transcription module is connected to the first determination module and is configured to perform speech transcription on the key communication voice data at the current moment based on a first model of the target cloud server.
[0013] In a third aspect, the present invention further provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to implement the recognition method for communication content provided in the first aspect above.
[0014] In a fourth aspect, the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the recognition method for communication content provided in the first aspect above is implemented.
[0015] The recognition method, device, electronic device, and computer-readable storage medium provided by the present invention, through a real-time updated resource scheduling diagram, realize the performance status monitoring and feedback of cloud servers and edge servers at the second level, and further provide real-time and effective data support for allocating the target cloud server corresponding to key communication voice data subsequently, improving the transcription quality, efficiency, and user experience of high-load or emergency key communication voice data in a changing communication environment and task requirements. Therefore, the present invention can realize flexible and stable communication content recognition, improve the response speed of communication content recognition in high-load or task-emergency situations, improve the adaptability of communication content recognition to changing communication environments and task requirements, and further improve the quality, efficiency, and user experience of the overall service of communication content recognition. Description of the Drawings
[0016] Figure 1 It is a flowchart of a recognition method for communication content according to Embodiment 1 of the present invention;
[0017] Figure 2 It is a flowchart of a recognition method for communication content according to Embodiment 2 of the present invention;
[0018] Figure 3 It is a structural schematic diagram of a recognition device for communication content according to Embodiment 3 of the present invention. Detailed Embodiments
[0019] To enable those skilled in the art to better understand the technical solutions of the present invention, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.
[0020] It can be understood that the specific embodiments and the accompanying drawings described herein are only for explaining the present invention, rather than limiting the present invention.
[0021] It can be understood that, without conflict, the various embodiments in the present invention and the various features in the embodiments can be combined with each other.
[0022] It can be understood that, for the convenience of description, only the parts related to the present invention are shown in the accompanying drawings of the present invention, and the parts unrelated to the present invention are not shown in the accompanying drawings.
[0023] It can be understood that each unit and module involved in the embodiments of the present invention may correspond to only one entity structure, or may be composed of multiple entity structures, or multiple units and modules may also be integrated into one entity structure.
[0024] It can be understood that, without conflict, the functions and steps marked in the flowcharts and block diagrams of the present invention may occur in an order different from that marked in the accompanying drawings.
[0025] It can be understood that in the flowcharts and block diagrams of the present invention, the possible architectures, functions, and operations of the systems, devices, equipment, and methods according to the embodiments of the present invention are shown. Among them, each block in the flowchart or block diagram may represent a unit, module, program segment, or code, which contains executable instructions for implementing the specified function. Moreover, each block or combination of blocks in the block diagram and flowchart can be implemented by a hardware-based system for implementing the specified function, or can be implemented by a combination of hardware and computer instructions.
[0026] It can be understood that the units and modules involved in the embodiments of the present invention can be implemented in software or in hardware. For example, the units and modules can be located in the processor.
[0027] Embodiment 1:
[0028] As Figure 1 shown, this embodiment provides a method for identifying communication content. The method for identifying communication content includes:
[0029] S101, obtaining the communication voice data and the resource scheduling diagram at the current moment, where the communication voice data includes key communication voice data.
[0030] Optionally, the communication voice data further includes non-key communication voice data.
[0031] In this embodiment, communication voice data refers to communication data that needs to be transcribed into text. Critical communication voice data refers to communication voice data with strong real-time requirements and high priority, such as communication voice data involved in emergency calls, customer support calls, or critical tasks. Critical communication voice data usually needs to be transmitted in real time, and any delay may seriously affect the call quality. When allocating network resources, critical communication voice data needs to be processed prior to non-critical communication voice data to ensure clear and coherent calls. Non-critical communication voice data refers to communication voice data with low priority and strong retransmissibility, such as product content, non-real-time calls, and chatting communication voice data. Non-critical communication voice data can usually tolerate a certain degree of delay or loss, with less impact on communication quality. A resource scheduling diagram refers to a visual representation of the allocation and usage of network resources, where network resources include, but are not limited to, the first model of the cloud server and the second model of the edge server.
[0032] Specifically, S101: Obtain the communication voice data and the resource scheduling diagram at the current moment, including steps S1011 - S1013:
[0033] S1011, obtain the communication data, the first operating state, and the second operating state at the current moment, where the first operating state refers to the operating states of all cloud servers, and the second operating state refers to the operating states of all edge servers.
[0034] In this embodiment, through the cloud intelligent data collection program of the operator's large network and the interfaces provided by the operator, according to preset rules and conditions, the required communication data is quickly screened and extracted. Among them, the communication data includes, but is not limited to, the call records and text message contents of users. This embodiment uses the cloud intelligent data collection program of the operator's large network, which can automatically adapt to changes in data interfaces of different operators and efficiently process a large amount of communication data. The collected data is cleaned, formatted, and normalized for subsequent analysis and model training.
[0035] The operating status includes, but is not limited to, the usage of the CPU (Central Processing Unit), GPU (Graphics Processing Unit), memory, and storage space. In this embodiment, by deploying monitoring tools, the usage of the CPU, GPU, memory, and storage space of the edge server and the cloud server is periodically collected to obtain the first operating status and the second operating status. Among them, the monitoring tools include, but are not limited to, sensors and monitoring centers, and the period includes, but is not limited to, every millisecond. The sensors include, but are not limited to, Prometheus and Nagios. Therefore, by collecting the usage of the CPU, GPU, memory, and storage space of the edge server and the cloud server every millisecond through high-precision sensors and quickly transmitting the usage of the CPU, GPU, memory, and storage space to the monitoring center through a low-latency network protocol, the resource usage data can be obtained in real time and accurately, ensuring the timeliness and accuracy of the data.
[0036] It should be noted that after receiving the usage of the CPU, GPU, memory, and storage space, the monitoring center converts the usage of the CPU, GPU, memory, and storage space into intuitive charts and concise resource usage reports through data visualization techniques and intelligent analysis algorithms. Among them, the data visualization techniques include, but are not limited to, Tableau and Power BI, and the intelligent analysis algorithms include, but are not limited to, regression analysis, clustering analysis, and association rule learning. In this embodiment, by integrating and displaying the resource usage, it is convenient to quickly understand and analyze the resource status, and then it is convenient to generate a reasonable resource scheduling diagram subsequently.
[0037] S1012, perform data annotation and classification on the communication data to obtain the key communication voice data and non-key communication voice data at the current moment.
[0038] In this embodiment, the communication data includes, in addition to the communication voice data that needs to be voice-transcribed, various texts, signaling, metadata, and information on network performance and security related to communication. Therefore, after quickly screening and extracting the required communication data through the cloud intelligent data collection program of the operator's large network and the interfaces provided by the operator according to preset rules and conditions, this embodiment also performs data preprocessing on the communication data at the current moment to obtain the communication voice data at the current moment. Among them, the data preprocessing includes, but is not limited to, pattern recognition, standardization conversion, and data cleaning.
[0039] Perform data preprocessing on the communication data at the current moment, specifically including: segmenting the continuous speech signal in the communication data into short frames of 20 - 40 ms (with 50% overlap) to eliminate signal non-stationarity; using spectral subtraction or a deep learning model (such as RNNoise) to remove background noise; using MFCC (Mel Frequency Cepstrum Coefficient) to extract the spectral envelope features (12 - 20 dimensions) of the speech, where the spectral envelope features include but are not limited to Prosodic Features, such as: fundamental frequency (F0), energy, and speech rate. MFCC includes the following steps: STFT (Short-Time Fourier Transform) → Mel filter bank → log → DCT (Discrete Cosine Transform); using a pre-trained speech encoder (such as: Wav2Vec 2.0) to extract context-aware speech embeddings.
[0040] By performing data preprocessing on the communication data at the current moment, this embodiment can identify and process communication data in different formats and qualities, perform unified conversion on data in different formats, ensure data consistency and availability, effectively filter out noise in the communication data, remove noise, error information, and redundant data in the communication data, identify the communication speech data in the communication data, lay a foundation for identifying communication speech data with strong real-time performance and high priority as well as communication speech data with low priority and strong retransmissibility, and improve the accuracy and efficiency of data annotation and classification for communication data.
[0041] After performing data preprocessing on the communication data at the current moment to obtain the communication speech data at the current moment, this embodiment also performs data annotation on the communication speech data at the current moment. Performing data annotation on the communication speech data at the current moment specifically includes: using a pre-trained ASR (Automatic Speech Recognition) model (such as Whisper) to transcribe the speech into text, for example:
[0042] import whisper
[0043] model=whisper.load_model("base")
[0044] text = model.transcribe("audio.wav")["text"]; Identify keywords in the communication data using the keyword list of the operator's tag library (such as "emergency", "transfer", "authorization code"), and calculate the keyword density KeyScore; Use the keyword density KeyScore as a tag to perform data annotation on the communication data; In addition, have professionals review and correct the tags of the communication data. In this embodiment, by using a machine learning model to perform preliminary annotation on the communication data and then having professionals review and correct it, the annotation efficiency and accuracy are improved.
[0045] After performing data annotation on the communication voice data at the current moment, this embodiment also classifies the communication data at the current moment. Classifying the communication data at the current moment specifically includes: Obtaining a large number of sample data from different industries and application scenarios, and through the sample data, constructing and training a binary classification model so that the binary classification model can learn the characteristic patterns of key communication voice data in different communication scenarios. For example: In a call in the financial field, the amount and account information involved are usually regarded as key communication voice data, and in a call in the logistics field, the goods transportation status, address, etc. are key communication voice data; Accurately identify the tags in the communication data according to the binary classification model, and calculate FinalScore; According to FinalScore, classify the communication data into key communication voice data and non-key communication voice data, that is: If FinalScore is greater than the preset threshold, classify the communication data as key communication voice data; If FinalScore is less than or equal to the preset threshold, classify the communication data as non-key communication voice data. In this embodiment, through an efficient binary classification model, it is possible to monitor and classify communication data in real time, thereby accelerating the response speed to key communication voice data.
[0046] It should be noted that removing background noise specifically includes: According to formula (1): S^(f) = max(|Y(f)|^2 - α·|N(f)|^2, β·|Y(f)|^2), where S^(f) represents the speech spectrum after removing background noise, Y(f) represents the noisy speech spectrum, N(f) represents the noise spectrum estimated through the silent segment, α represents the over-subtraction factor (usually 1.0 - 1.5), and β represents the guaranteed attenuation coefficient (usually 0.01 - 0.1).
[0047] Performing data annotation on the communication data at the current moment also includes: Using a voiceprint recognition model to verify the identity of the speaker, for example:
[0048] import speechbrain as sb
[0049] verifier = sb.pretrained.SpeakerRecognition.from_hparams("speechbrain / spkrec-ecapa-voxceleb")
[0050] score, _ = verifier.verify_files("ref.wav", "test.wav"), and mark the speaker identity (such as executive, risk control personnel) in the communication data. Among them, the speaker recognition model includes but is not limited to: ECAPA-TDNN (Emphasized Channel Attention-based Propagation-Time-Delay Neural Network); input the communication data into a multimodal sentiment analysis model (such as Multimodal Transformer), output and mark the sentiment polarity (Angry / Urgent) of the communication data. Among them, the multimodal sentiment analysis model includes a speech branch and a text branch. The speech branch is used to input and process the MFCC features of the communication data, and the text branch is used to input and process the BERT (Bidirectional Encoder Representations from Transformers, a language representation model with a bidirectional Transformer structure) embeddings of the ASR transcription text.
[0051] Calculate FinalScore, specifically including: qualitatively determine the values of all labels of the communication data, such as the label purchase intention (high purchase intention, medium purchase intention, low purchase intention); at the same time, support information extraction, such as: extract the communication keywords of both parties (financial insurance, bank app, renewal operation), extract the call score and basis (call score 88 points, this employee has followed up the customer well) paragraphs, as the basis for label weights; calculate FinalScore according to the qualitative determination of label values and the basis of label weights, such as: FinalScore = 0.4×KeyScore + 0.3×SpeakerRisk + 0.3×AngerLevelFinalScore, where SpeakerRisk refers to the qualitative determination of the speaker identity value, and AngerLevelFinalScore refers to the qualitative determination of the sentiment polarity value.
[0052] If the confidence level of the binary classification model in calculating the FinalScore is low, a lightweight LSTM (Long Short-Term Memory) classification model is called for a second judgment. Additionally, in this embodiment, the precision and recall rate of the binary classification model are also evaluated. If the precision and recall rate do not meet the expectations, the binary classification model is optimized and updated based on the confusion matrix and misclassified samples.
[0053] S1013. Based on the first operating state, the second operating state, the critical communication voice data, and the non-critical communication voice data, construct a resource scheduling graph at the current moment.
[0054] Specifically, the resource scheduling graph includes a first task node, a second task node, a first resource node, a second resource node, a first path, a second path, a first weight, and a second weight.
[0055] In this embodiment, by using the graph database technology, it is possible to efficiently store and manage large-scale first operating state, second operating state, critical communication voice data, and non-critical communication voice data, quickly construct a resource scheduling graph to visually display the relationships among the first operating state, the second operating state, the critical communication voice data, and the non-critical communication voice data, and support dynamic update and query operations.
[0056] Specifically, S1013: Based on the first operating state, the second operating state, the critical communication voice data, and the non-critical communication voice data, construct a resource scheduling graph at the current moment, including: respectively evaluate the priorities corresponding to the critical communication voice data and the non-critical communication voice data, and classify the critical communication voice data and the non-critical communication voice data based on the priorities, where the priorities include the urgency ranking and the importance ranking; generate a first task node, a second task node, a first resource node, a second resource node, a first path, and a second path, where the first task node refers to the critical communication voice data corresponding to different priorities, the second task node refers to the non-critical communication voice data corresponding to different priorities, the first resource node refers to the cloud server for processing the critical communication voice data corresponding to different priorities, the second resource node refers to the edge server for processing the non-critical communication voice data corresponding to different priorities, the first path refers to the directed edge from the first task node to the first resource node, and the second path refers to the directed edge from the second task node to the second resource node; calculate the first weight and the second weight based on the first operating state, the second operating state, the critical communication voice data, and the non-critical communication voice data to obtain the resource scheduling graph at the current moment, where the first weight refers to the execution time, energy consumption, and resource load of the first path, and the second weight refers to the execution time, energy consumption, and resource load of the second path.
[0057] In this embodiment, a priority evaluation model based on task characteristics and business rules is used to evaluate the urgency ranking and importance ranking of critical communication voice data and non-critical communication voice data, and different priorities are assigned to critical communication voice data and non-critical communication voice data based on the urgency ranking and importance ranking of financial transactions of critical communication voice data and non-critical communication voice data to ensure the efficient processing of critical communication voice data and non-critical communication voice data, as shown in Table 1.
[0058] Table 1 Priorities of Critical Communication Voice Data and Non-Critical Communication Voice Data
[0059]
[0060] Generate the first task node, the second task node, the first resource node, the second resource node, the first path, and the second path, specifically including:
[0061]
[0062]
[0063]
[0064] Based on the first operating state, the second operating state, critical communication voice data, and non-critical communication voice data, calculate the first weight and the second weight to obtain the resource scheduling graph at the current moment, specifically including: First weight / Second weight = α * (Execution time) + β * (Energy consumption cost) + γ * (Load penalty factor); Execution time = (Packet size / Processing speed) + (Current queue length * Average processing time); Energy consumption cost = (Base power consumption 150W + Dynamic power consumption 0.2W / Mbps * 2Mbps) * Processing time; Load penalty factor = max(0, Current load rate - Safety threshold) / 100. In this embodiment, taking the first operating state including the processing speed of the cloud server: 200Mbps / core and the second operating state including the remaining power of the edge server: 78% as an example, the first weight = 0.6 * (2Mbps / (200Mbps * 0.35)) + (8 tasks * 5ms) + 0.3 * 150.4W * 0.06857s + 0.1 * ((65% - 70%) → 0) = 43.24. In this embodiment, by calculating the first weight and the second weight based on the first operating state, the second operating state, critical communication voice data, and non-critical communication voice data, factors such as the complexity of the task, the size of the data volume, the processing efficiency of the model for different tasks, and the load situation of the current resources are comprehensively considered, enabling the scheduling graph to accurately reflect the matching relationship between tasks and resources and providing a reasonable basis for the selection of the optimal path.
[0065] S102. Based on the resource scheduling graph at the current moment and Dijkstra's algorithm, determine the target cloud server.
[0066] Specifically, S102: Based on the resource scheduling graph at the current moment and Dijkstra's algorithm, determine the target cloud server, including steps S1021 - S1023:
[0067] S1021. Based on Dijkstra's algorithm, determine whether there are conflicting resource nodes in the resource scheduling graph at the current moment. Here, a conflicting resource node refers to a first resource node that has an optimal path with multiple first task nodes. Here, the optimal path refers to the first path corresponding to the minimum value of the first weight.
[0068] In this embodiment, taking the first task nodes including: critical communication voice data 1 and critical communication voice data 2, the first resource nodes including: cloud server A and cloud server B, the first paths including critical communication voice data 1 -> cloud server A, critical communication voice data 2 -> cloud server A, communication voice data 1 -> cloud server B, critical communication voice data 2 -> cloud server B, and the first weights including: 20, 30, 30, 52 as an example, it can be preliminarily determined that the optimal path of critical communication voice data 1 is critical communication voice data 1 -> cloud server A, and the optimal path of critical communication voice data 2 is critical communication voice data 2 -> cloud server A. Since cloud server A has an optimal path with both critical communication voice data 1 and critical communication voice data 2, cloud server A is a conflicting resource node.
[0069] S1022. In response to the absence of conflicting resource nodes in the resource scheduling graph at the current moment, determine that the target cloud server includes all the first resource nodes corresponding to the optimal paths in the resource scheduling graph at the current moment.
[0070] In this embodiment, taking the first task nodes including: critical communication voice data 1 and critical communication voice data 2, the first resource nodes including: cloud server A and cloud server B, the first paths including critical communication voice data 1 -> cloud server B, critical communication voice data 2 -> cloud server A, communication voice data 1 -> cloud server B, critical communication voice data 2 -> cloud server B, and the first weights including: 40, 30, 30, 52 as an example, it can be preliminarily determined that the optimal path of critical communication voice data 1 is critical communication voice data 1 -> cloud server B, and the optimal path of critical communication voice data 2 is critical communication voice data 2 -> cloud server A. Therefore, determine that the target cloud server includes cloud server A and cloud server B.
[0071] S1023. In response to the existence of conflicting resource nodes in the resource scheduling graph at the current moment, predict the resource scheduling graph within a preset time period, and determine that the target cloud server includes the first resource nodes corresponding to all the optimal paths in the resource scheduling graph within the preset time period.
[0072] Specifically, predicting the resource scheduling graph within a preset time period includes: determining non-optimized resource nodes and resource nodes to be optimized. Among them, non-optimized resource nodes refer to the conflicting resource nodes corresponding to the first task nodes with high priority in the resource scheduling graph at the current moment, and resource nodes to be optimized refer to the conflicting resource nodes corresponding to the first task nodes with low priority in the resource scheduling graph at the current moment; predicting the third running state within the preset time period, where the third running state refers to the running state of the non-optimized resource nodes; based on the third running state within the preset time period, adjust the first weight of the resource nodes to be optimized in the resource scheduling graph at the current moment to obtain the resource scheduling graph within the preset time period.
[0073] In this embodiment, taking the priority of critical communication voice data 1 being higher than that of critical communication voice data 2 as an example, determine that cloud server A corresponding to critical communication voice data 1 is a non-optimized resource node, and cloud server A corresponding to critical communication voice data 2 is a resource node to be optimized. That is, it is determined that the optimal path of critical communication voice data 1 is still critical communication voice data 1 -> cloud server A, and the optimal path of critical communication voice data 2: in critical communication voice data 2 -> cloud server A, cloud server A may need to be adjusted. The preset time period refers to the time period required to send critical communication voice data 1 to cloud server A and for cloud server A to perform speech transcription on critical communication voice data 1, and predict the running state of cloud server A within the preset time period. Similarly to calculating the first weight and the second weight based on the first running state, the second running state, critical communication voice data, and non-critical communication voice data, based on the third running state within the preset time period, adjust the first weight of the resource nodes to be optimized in the resource scheduling graph at the current moment to obtain the resource scheduling graph within the preset time period. In this embodiment, based on the Dijkstra algorithm, combined with the priority of tasks and the real-time status information of resources, the search direction is guided to avoid path failure caused by changes in resource status, significantly reducing the calculation time and resource consumption, improving the applicability and efficiency of the Dijkstra algorithm in complex communication environments, and being able to quickly find the optimal resource allocation path.
[0074] S103. Based on the first model of the target cloud server, perform speech transcription on the critical communication voice data at the current moment.
[0075] Optionally, after S101: obtaining the communication voice data and the resource scheduling graph at the current moment, the communication content recognition method further includes:
[0076] S104. Determine the target edge server based on the resource scheduling graph at the current moment and the Dijkstra algorithm.
[0077] In this embodiment, taking the second task node including non-critical communication voice data 1 and non-critical communication voice data 2, the second resource node including edge server a1, edge server a2, edge server b1, and edge server b2, the second path including non-critical communication voice data 1 -> edge server a1, non-critical communication voice data 1 -> edge server a2, non-critical communication voice data 1 -> edge server b1, non-critical communication voice data 1 -> edge server b2, non-critical communication voice data 2 -> edge server a1, non-critical communication voice data 2 -> edge server a2, non-critical communication voice data 2 -> edge server b1, non-critical communication voice data 2 -> edge server b2, and the second weight including 60, 82, 75, 53, 62, 53, 61, 22 as an example, the optimal path of non-critical communication voice data 1 can be initially determined as non-critical communication voice data 1 -> edge server b2, and the optimal path of non-critical communication voice data 2 can be initially determined as non-critical communication voice data 2 -> edge server b2. Since there are optimal paths between edge server b2 and both non-critical communication voice data 1 and non-critical communication voice data 2, edge server b2 is a conflict resource node.
[0078] Optionally, before S104: Determine the target edge server based on the resource scheduling graph at the current moment and the Dijkstra algorithm, the communication content recognition method further includes:
[0079] S106. Obtain historical communication voice data.
[0080] In this embodiment, the historical communication voice data refers to the critical communication voice data and non-critical communication voice data before the current moment. After obtaining the historical communication voice data, this embodiment further includes: constructing an initial first model based on the historical communication voice data; respectively dividing the historical critical communication voice data and non-critical communication voice data into the training data set of the first model and the training data set of the second model; using the DistributedDataParallel module of PyTorch to allocate the initial first model and the training data set of the first model to multiple cloud servers, and allocate the training data set of the second model to multiple edge servers, for example:
[0081] import torch
[0082] import torch.distributed as dist from torch.nn.parallel import DistributedDataParallel as DDP
[0083] # Initialize the distributed environment
[0084] dist.init_process_group(backend='nccl')
[0085] # Construct the model and allocate it to the device
[0086] model = MyModel().cuda()
[0087] model = DDP(model, device_ids=[torch.cuda.current_device()])
[0088] # Use DistributedSampler to allocate data sampler = torch.utils.data.distributed.DistributedSampler(dataset) dataloader = torch.utils.data.DataLoader(dataset, sampler=sampler)
[0089] S107. Build and train the first model of the target cloud server based on historical communication voice data and an adaptive learning rate adjustment strategy.
[0090] In this embodiment, after the target cloud server receives the allocated initial first model and the training dataset of the first model, it trains the initial first model based on the training dataset of the first model and the adaptive learning rate adjustment strategy to obtain the first model of the target cloud server. In this embodiment, taking the adaptive learning rate adjustment strategy as the polynomial decay strategy as an example, it trains the initial first model based on the training dataset of the first model and the polynomial decay strategy to obtain the first model of the target cloud server, where the polynomial decay strategy is as follows:
[0091] def decay_lr_poly(base_lr, epoch_i, batch_i, total_epochs, total_batches, warm_up, power = 1.0):
[0092] if warm_up > 0 and epoch_i < warm_up:
[0093] rate = (epoch_i * total_batches + batch_i) / (warm_up * total_batches)
[0094] else:
[0095] rate = (1.0 - ((epoch_i - warm_up) * total_batches + batch_i) / ((total_epochs - warm_up) * total_batches)) ** power
[0096] return rate * base_lr. During the training process of the first model at the beginning of training, the learning rate gradually decreases as the number of training epochs increases, thus avoiding the training instability caused by an overly large learning rate. Therefore, in this embodiment, by training the initial first model based on the training dataset of the first model and the polynomial decay strategy, the learning rate can be dynamically adjusted, the training process can be optimized, and the accuracy of the first model of the target cloud server can be improved.
[0097] It should be noted that before training the initial first model based on the training dataset of the first model and the polynomial decay strategy, this embodiment also uses data augmentation technology to augment the training dataset of the first model to increase the robustness of the initial first model to different inputs, thereby improving the generalization ability. Among them, the data augmentation technology is as follows: from tensorflow, keras.preprocessing.image import ImageDataGenerator datagen = ImageDataGenerator(
[0098] rotation_range = 10,
[0099] zoom_range = 0.1,
[0100] width_shift_range = 0.1,
[0101] height_shift_range = 0.1).
[0102] S108. Based on the teacher-student training method and the combined loss function of hard labels and soft labels, perform distillation processing on the first model of the target cloud server to obtain the second model of the target edge server.
[0103] In this embodiment, taking the model distillation technology as the teacher-student training method and the target loss function as the combined loss function of hard labels and soft labels as an example, the teacher-student training method is adopted. By selecting appropriate temperature parameters (such as the best performance when T = 2) and the combined loss function of hard labels and soft labels, the knowledge of the first model of the target cloud server is effectively transferred to the second model of the target edge server. While reducing the model complexity and computational resource consumption, high performance is maintained. In this embodiment, by distilling the first model of the target cloud server, more than 80% of the performance improvement is retained after distillation.
[0104] Adopting the teacher-student training method, selecting appropriate temperature parameters and the combined loss function of hard labels and soft labels, and effectively transferring the knowledge of the first model of the target cloud server to the second model of the target edge server specifically includes: using the first model of the target cloud server as the teacher model and the second model of the target edge server as the student model; using the output probability distribution of the teacher model as the soft target and calculating the KL divergence (Kullback-Leibler Divergence) between the output of the student model and the soft target:
[0105] where T represents the temperature parameter, T = 2, f teacher (x) represents the output of the teacher model, f student (x) represents the output of the student model, x represents the training dataset of the first model, Softmax(·) represents the function that converts a real number vector into a probability distribution, KL(·) represents the asymmetric measure that measures the difference between two probability distributions, and L soft represents the KL divergence between the output of the student model and the soft target; In this embodiment, taking the hard target loss as the cross-entropy loss as an example, calculate the cross-entropy loss between the output of the student model and the target output:
[0106] L hard = CrossEntropy(f student (x), y), where y represents the target output, CrossEntropy(·) represents the cross-entropy loss function, and L hard represents the cross-entropy loss between the output of the student model and the target output; Weightedly sum the KL divergence between the output of the student model and the soft target and the cross-entropy loss between the output of the student model and the target output: L total = a·L soft +(1 - a)·L hard , where a represents the hyperparameter used to balance the two losses.
[0107] It should be noted that in this embodiment, the knowledge of the first model of the target cloud server can be effectively transferred to the second model of the target edge server directly through the following code:
[0108] # Calculate the soft target loss
[0109] soft loss = kl divergence(softmax(teacher_output / T) softmax(studentoutput / T)) / T**2
[0110] # Calculate the hard target loss
[0111] hard loss = cross_entropy(student_output, true labels)
[0112] # Total loss total loss = alpha * soft loss + (1 - alpha) * hard loss.
[0113] The temperature parameter T is a key hyperparameter in the distillation process. A higher temperature will make the soft target distribution smoother, which helps the student model learn more extensive knowledge. Experiments show that when T = 2, the distillation effect is better.
[0114] S105, based on the second model of the target edge server, perform speech transcription on the non-critical communication voice data at the current moment.
[0115] A method for identifying communication content provided by this embodiment realizes the performance status monitoring and feedback of the cloud server and the edge server at the second level through a real-time updated resource scheduling diagram, and then provides real-time and effective data support for subsequent allocation of the target cloud server corresponding to the critical communication voice data, improving the transcription quality, efficiency, and user experience of high-load or emergency critical communication voice data under changing communication environments and task requirements, realizing flexible and stable communication content recognition, improving the response speed of communication content recognition in high-load or task-emergency situations, improving the adaptability of communication content recognition to changing communication environments and task requirements, and further improving the quality, efficiency, and user experience of the overall service of communication content recognition.
[0116] Embodiment 2:
[0117] As Figure 2 shown, this embodiment provides a method for identifying communication content. The method for identifying communication content includes:
[0118] S201. Obtain the communication data at the current moment, and perform data annotation and classification on the communication data to obtain the key communication voice data and non-key communication voice data at the current moment.
[0119] In this embodiment, performing data annotation and classification on the communication data is Figure 2 the data preparation in. Through the cloud intelligent data acquisition program of the operator's large network and the interfaces provided by the operator, according to the preset rules and conditions, quickly screen and extract the required communication data. Among them, the communication data includes but is not limited to: the call records and text message contents of users. In addition to the communication voice data that needs to be voice-transcribed, the communication data also includes various texts, signaling, metadata, and information on network performance and security related to communication. Therefore, after quickly screening and extracting the required communication data through the cloud intelligent data acquisition program of the operator's large network and the interfaces provided by the operator according to the preset rules and conditions, this embodiment also performs data preprocessing on the communication data at the current moment to obtain the communication voice data at the current moment. Among them, the data preprocessing includes but is not limited to: pattern recognition, standardization conversion, and data cleaning. After performing data preprocessing on the communication data at the current moment to obtain the communication voice data at the current moment, this embodiment also performs data annotation on the communication voice data at the current moment. After performing data annotation on the communication voice data at the current moment, this embodiment also classifies the communication data at the current moment.
[0120] S202. Obtain the historical communication voice data; based on the historical communication voice data and the adaptive learning rate adjustment strategy, construct and train the first model of the target cloud server; based on the teacher-student training method and the combined loss function of hard labels and soft labels, perform distillation processing on the first model of the target cloud server to obtain the second model of the target edge server.
[0121] In this embodiment, the first model of the target cloud server is Figure 2 the large cloud model in, and the second model of the target edge server is Figure 2 the small edge model in.
[0122] S203. Obtain the first operating state and the second operating state, where the first operating state refers to the operating states of all cloud servers, and the second operating state refers to the operating states of all edge servers.
[0123] S204, respectively evaluate the priorities corresponding to the critical communication voice data and the non-critical communication voice data, and classify the critical communication voice data and the non-critical communication voice data based on the priorities, where the priorities include the urgency ranking and the importance ranking; generate a first task node, a second task node, a first resource node, a second resource node, a first path, and a second path, where the first task node refers to the critical communication voice data corresponding to different priorities, the second task node refers to the non-critical communication voice data corresponding to different priorities, the first resource node refers to the cloud server for processing the critical communication voice data corresponding to different priorities, the second resource node refers to the edge server for processing the non-critical communication voice data corresponding to different priorities, the first path refers to the directed edge from the first task node to the first resource node, and the second path refers to the directed edge from the second task node to the second resource node; calculate a first weight and a second weight based on the first operating state, the second operating state, the critical communication voice data, and the non-critical communication voice data to obtain the resource scheduling graph at the current moment, where the first weight refers to the execution time, energy consumption, and resource load of the first path, and the second weight refers to the execution time, energy consumption, and resource load of the second path.
[0124] In this embodiment, the resource scheduling graph at the current moment is the Figure 2 scheduling graph in.
[0125] S205, based on the Dijkstra algorithm, determine whether there are conflicting resource nodes in the resource scheduling graph at the current moment, where the conflicting resource node refers to the first resource node that has an optimal path with multiple first task nodes, and the optimal path refers to the first path corresponding to the minimum value of the first weight; in response to the absence of conflicting resource nodes in the resource scheduling graph at the current moment, determine that the target cloud server includes all the first resource nodes corresponding to the optimal paths in the resource scheduling graph at the current moment; in response to the existence of conflicting resource nodes in the resource scheduling graph at the current moment, determine the non-optimized resource nodes and the resource nodes to be optimized, where the non-optimized resource nodes refer to the conflicting resource nodes corresponding to the first task nodes with higher priorities in the resource scheduling graph at the current moment, and the resource nodes to be optimized refer to the conflicting resource nodes corresponding to the first task nodes with lower priorities in the resource scheduling graph at the current moment, predict the third operating state within a preset time period, where the third operating state refers to the operating state of the non-optimized resource nodes, and based on the third operating state within the preset time period, adjust the first weight of the resource nodes to be optimized in the resource scheduling graph at the current moment to obtain the resource scheduling graph within the preset time period, and determine that the target cloud server includes all the first resource nodes corresponding to the optimal paths in the resource scheduling graph within the preset time period.
[0126] S206, based on the first model of the target cloud server, perform speech transcription on the critical communication voice data at the current moment.
[0127] A method for identifying communication content provided in this embodiment realizes the monitoring and feedback of the performance status of the cloud server and the edge server at the second level through a real-time updated resource scheduling graph, and further provides real-time and effective data support for allocating the target cloud server corresponding to the critical communication voice data subsequently, improving the transcription quality, efficiency, and user experience of the critical communication voice data with high load or urgency under the changing communication environment and task requirements, realizing flexible and stable communication content recognition, improving the response speed of communication content recognition in the case of high load or urgent tasks, and improving the adaptability of communication content recognition to the changing communication environment and task requirements, thereby improving the quality, efficiency, and user experience of the overall service of communication content recognition.
[0128] Embodiment 3:
[0129] As Figure 3 shown, this embodiment provides an apparatus for identifying communication content, including: a first acquisition module 31, a first determination module 32, and a first speech transcription module 33. The first acquisition module 31 is configured to acquire the communication voice data and the resource scheduling graph at the current moment, where the communication voice data includes critical communication voice data. The first determination module 32 is connected to the first acquisition module 31 and is configured to determine the target cloud server based on the resource scheduling graph at the current moment and the Dijkstra algorithm. The first speech transcription module 33 is connected to the first determination module 32 and is configured to perform speech transcription on the critical communication voice data at the current moment based on the first model of the target cloud server.
[0130] Specifically, the first acquisition module 31 includes: an acquisition unit 311, a classification unit 312, and a construction unit 313. The acquisition unit 311 is configured to acquire the communication data, the first operating state, and the second operating state at the current moment, where the first operating state refers to the operating states of all cloud servers, and the second operating state refers to the operating states of all edge servers. The classification unit 312 is configured to perform data annotation and classification on the communication data to obtain the critical communication voice data and the non-critical communication voice data at the current moment. The construction unit 313 is configured to construct the resource scheduling graph at the current moment based on the first operating state, the second operating state, the critical communication voice data, and the non-critical communication voice data.
[0131] Specifically, the construction unit 313 includes: an evaluation subunit, a generation subunit, and a calculation subunit. The evaluation subunit is configured to evaluate the priorities corresponding to the critical communication voice data and the non-critical communication voice data respectively, and classify the critical communication voice data and the non-critical communication voice data based on the priorities. Among them, the priorities include the urgency ranking and the importance ranking. The generation subunit is configured to generate a first task node, a second task node, a first resource node, a second resource node, a first path, and a second path. Among them, the first task node refers to the critical communication voice data corresponding to different priorities, the second task node refers to the non-critical communication voice data corresponding to different priorities, the first resource node refers to the cloud server for processing the critical communication voice data corresponding to different priorities, the second resource node refers to the edge server for processing the non-critical communication voice data corresponding to different priorities, the first path refers to the directed edge from the first task node to the first resource node, and the second path refers to the directed edge from the second task node to the second resource node. The calculation subunit is configured to calculate a first weight and a second weight based on the first operating state, the second operating state, the critical communication voice data, and the non-critical communication voice data, so as to obtain the resource scheduling graph at the current moment. Among them, the first weight refers to the execution time, energy consumption, and resource load of the first path, and the second weight refers to the execution time, energy consumption, and resource load of the second path.
[0132] Specifically, the first determination module 32 includes: a judgment unit 321, a first determination unit 322, and a second determination unit 323. The judgment unit 321 is configured to judge whether there are conflicting resource nodes in the resource scheduling graph at the current moment based on the Dijkstra algorithm. Among them, the conflicting resource node refers to the first resource node that has an optimal path with multiple first task nodes. Among them, the optimal path refers to the first path corresponding to the minimum value of the first weight. The first determination unit 322 is configured to determine that the target cloud server includes all the first resource nodes corresponding to the optimal paths in the resource scheduling graph at the current moment in response to the absence of conflicting resource nodes in the resource scheduling graph at the current moment. The second determination unit 323 is configured to predict the resource scheduling graph within a preset time period and determine that the target cloud server includes all the first resource nodes corresponding to the optimal paths in the resource scheduling graph within the preset time period in response to the presence of conflicting resource nodes in the resource scheduling graph at the current moment.
[0133] Specifically, the second determination unit 323 includes: a determination subunit, a prediction subunit, and an adjustment subunit. The determination subunit is configured to determine a non-optimized resource node and an optimized resource node to be optimized. Among them, the non-optimized resource node refers to a conflicting resource node corresponding to a first task node with a high priority in the resource scheduling graph at the current moment, and the optimized resource node to be optimized refers to a conflicting resource node corresponding to a first task node with a low priority in the resource scheduling graph at the current moment. The prediction subunit is configured to predict a third operating state within a preset time period. Among them, the third operating state refers to the operating state of the non-optimized resource node. The adjustment subunit is configured to adjust the first weight of the optimized resource node to be optimized in the resource scheduling graph at the current moment based on the third operating state within the preset time period, so as to obtain a resource scheduling graph within the preset time period.
[0134] Optionally, the communication content recognition device further includes: a second determination module 34 and a second speech transcription module 35. The second determination module 34 is configured to determine a target edge server based on the resource scheduling graph at the current moment and the Dijkstra algorithm. The second speech transcription module 35 is configured to perform speech transcription on the non-critical communication voice data at the current moment based on the second model of the target edge server.
[0135] Optionally, the communication content recognition device further includes: a second acquisition module 36, a construction module 37, and a distillation module 38. The second acquisition module 36 is configured to acquire historical communication voice data. The construction module 37 is configured to construct and train the first model of the target cloud server based on the historical communication voice data and the adaptive learning rate adjustment strategy. The distillation module 38 is configured to perform distillation processing on the first model of the target cloud server based on the teacher-student training method and the combined loss function of hard labels and soft labels to obtain the second model of the target edge server.
[0136] It can be understood that the above-provided communication content recognition device executes the communication content recognition method corresponding to Embodiment 1 provided above. Therefore, the beneficial effects it can achieve can refer to the beneficial effects of the solution corresponding to the communication content recognition method in Embodiment 1 above, which will not be elaborated here.
[0137] Embodiment 4:
[0138] This embodiment provides an electronic device, including a memory and a processor. A computer program is stored in the memory, and the processor is configured to run the computer program to implement the communication content recognition method in Embodiment 1 or Embodiment 2 above.
[0139] Embodiment 5:
[0140] This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method for identifying communication content in the above-mentioned Embodiment 1 or Embodiment 2 is implemented.
[0141] It can be understood that the above embodiments are merely exemplary embodiments adopted to illustrate the principles of the present invention. However, the present invention is not limited thereto. For those of ordinary skill in the art, various modifications and improvements can be made without departing from the spirit and essence of the present invention, and these modifications and improvements are also regarded as the protection scope of the present invention.
Claims
1. A method for identifying communication content, characterized in that: include: Acquire communication voice data and a resource scheduling diagram at the current moment, wherein the communication voice data includes key communication voice data; Determine the target cloud server based on the current resource scheduling graph and Dijkstra algorithm; Based on the first model of the target cloud server, the key communication voice data at the current moment is voice-transcribed.
2. The method for identifying communication content according to claim 1, characterized in that: The communication voice data also includes non-critical communication voice data. The obtaining of the communication voice data and resource scheduling diagram at the current moment specifically includes: Acquire the communication data, the first operation status and the second operation status at the current moment, wherein the first operation status refers to the operation status of all cloud servers, and the second operation status refers to the operation status of all edge servers; Performing data labeling and classification on the communication data to obtain critical communication voice data and non-critical communication voice data at the current moment; A resource scheduling diagram at the current moment is constructed based on the first operating state, the second operating state, the critical communication voice data, and the non-critical communication voice data.
3. The method for identifying communication content according to claim 2, characterized in that: The resource scheduling graph includes a first task node, a second task node, a first resource node, a second resource node, a first path, a second path, a first weight, and a second weight. The constructing a resource scheduling diagram at the current moment based on the first operating state, the second operating state, the critical communication voice data and the non-critical communication voice data specifically includes: Evaluate the priorities corresponding to the critical communication voice data and the non-critical communication voice data respectively, and classify the critical communication voice data and the non-critical communication voice data based on the priorities, wherein the priorities include urgency ranking and importance ranking; Generate a first task node, a second task node, a first resource node, a second resource node, a first path, and a second path, wherein the first task node refers to the critical communication voice data corresponding to different priorities, the second task node refers to the non-critical communication voice data corresponding to different priorities, the first resource node refers to a cloud server that processes the critical communication voice data corresponding to different priorities, the second resource node refers to an edge server that processes the non-critical communication voice data corresponding to different priorities, the first path refers to a directed edge from the first task node to the first resource node, and the second path refers to a directed edge from the second task node to the second resource node; Based on the first operating state, the second operating state, the critical communication voice data and the non-critical communication voice data, the first weight and the second weight are calculated to obtain the resource scheduling diagram at the current moment, wherein the first weight refers to the execution time, energy consumption and resource load of the first path, and the second weight refers to the execution time, energy consumption and resource load of the second path.
4. The method for identifying communication content according to claim 3, characterized in that: The target cloud server is determined based on the resource scheduling diagram and Dijkstra algorithm at the current moment, specifically including: Based on the Dijkstra algorithm, determine whether there is a conflicting resource node in the resource scheduling graph at the current moment, wherein the conflicting resource node refers to a first resource node that has an optimal path with multiple first task nodes, wherein the optimal path refers to the first path corresponding to the minimum first weight value; In response to the absence of conflicting resource nodes in the resource scheduling graph at the current moment, determining that the target cloud server includes first resource nodes corresponding to all optimal paths in the resource scheduling graph at the current moment; In response to the existence of conflicting resource nodes in the resource scheduling diagram at the current moment, the resource scheduling diagram within a preset time period is predicted, and the target cloud server is determined to include the first resource node corresponding to all optimal paths in the resource scheduling diagram within the preset time period.
5. The method for identifying communication content according to claim 4, characterized in that: The resource scheduling diagram within the predicted preset time period specifically includes: Determine a non-optimized resource node and a resource node to be optimized, wherein the non-optimized resource node refers to a conflicting resource node corresponding to a first task node with a high priority in a resource scheduling diagram at the current moment, and the resource node to be optimized refers to a conflicting resource node corresponding to a first task node with a low priority in a resource scheduling diagram at the current moment; Predicting a third operating state within a preset time period, wherein the third operating state refers to an operating state of a non-optimized resource node; Based on the third operating state within the preset time period, the first weight of the resource node to be optimized in the resource scheduling graph at the current moment is adjusted to obtain the resource scheduling graph within the preset time period.
6. The method for identifying communication content according to claim 2, characterized in that: After acquiring the communication voice data and resource scheduling diagram at the current moment, the method further includes: Determine the target edge server based on the current resource scheduling graph and Dijkstra algorithm; Based on the second model of the target edge server, voice transcription is performed on the non-critical communication voice data at the current moment.
7. The method for identifying communication content according to claim 6, characterized in that: Before the second model based on the target edge server performs voice transcription on the non-critical communication voice data at the current moment, the method further includes: Acquire historical communication voice data; Based on historical communication voice data and adaptive learning rate adjustment strategy, a first model of the target cloud server is constructed and trained; Based on the teacher-student training method and the combined loss function of hard labels and soft labels, the first model of the target cloud server is distilled to obtain the second model of the target edge server.
8. A communication content identification device, characterized in that: It includes a first acquisition module, a first determination module and a first speech transcription module, The first acquisition module is used to acquire the communication voice data and resource scheduling diagram at the current moment, wherein the communication voice data includes key communication voice data, The first determination module is connected to the first acquisition module and is used to determine the target cloud server based on the resource scheduling diagram at the current moment and the Dijkstra algorithm. The first speech transcription module is connected to the first determination module and is used to perform speech transcription on the key communication speech data at the current moment based on the first model of the target cloud server.
9. An electronic device, characterized in that: It comprises a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to implement a method for identifying communication content as claimed in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, a method for identifying communication content according to any one of claims 1 to 7 is implemented.