Abstract generation method and device for electric power communication dispatching voice dialogue
Through the improved LSTM unit and multi-layer perceptron model to generate a power communication scheduling voice dialogue summary, the problem of low efficiency in the generation of automated abstracts in the prior art is solved, and efficient and accurate voice recording processing is achieved.
Patent Information
- Application Number
- CN202510632984.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-08-15
AI Technical Summary
The prior art lacks an effective automated voice recording summary generation method in power communication scheduling, resulting in inefficient dispatcher information retrieval and decision-making.
Using a two-layer improved summary generation model of long and short-term memory network unit and multi-layer perceptron, a speech dialogue summary is generated by performing feature extraction and importance scoring of the power communication scheduling voice dialogue. The model includes attention mechanism neural network architecture with horizontal and vertical structures.
It realizes the automatic summary generation of voice conversations for power communication scheduling, improves efficiency and accuracy, reduces manual intervention, ensures the relevance and integrity of the summary content, and adapts to the needs of different scenarios.
Smart Images

Figure CN120496504A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electric power communication dispatching, and in particular to a method and device for generating a summary of an electric power communication dispatching voice dialogue. Background Art
[0002] As the complexity of power communication networks increases, the number of voice recordings generated during power communication dispatching continues to rise. Traditional voice recording processing methods often rely on manual recording and organization, which is inefficient and prone to errors. Voice recording summarization is a crucial component of intelligent communication dispatching. It not only helps dispatchers quickly obtain key information, avoiding the need to review lengthy voice recordings one by one, saving time and effort and improving overall work efficiency, but also allows dispatchers to quickly obtain relevant information during power communication dispatching, helping them to monitor power equipment status in real time and adjust dispatching strategies in a timely manner in response to emergencies. This provides fundamental data support for subsequent intelligent analysis and decision-making.
[0003] Research on summarization of voice conversation records is still in its infancy, with most existing research focusing on speech recognition and text summary generation. However, existing technologies lack effective automated summary generation methods for power communication dispatch voice records, resulting in difficulties for dispatchers in information retrieval and decision-making. Summary of the Invention
[0004] The embodiments of the present invention provide a method and device for summarizing voice conversations in electric power communication dispatching, which realizes the summary generation of voice conversations in electric power communication dispatching based on semantic guidance, and improves the efficiency and accuracy of electric power communication dispatching by automatically processing and summarizing voice records of electric power calls.
[0005] In a first aspect, this embodiment provides a method for generating a summary of a power communication scheduling voice conversation, comprising:
[0006] Extracting features from an original voice conversation of the electric power communication dispatching to obtain a voice feature sequence of the original voice conversation;
[0007] Inputting the speech feature sequence into a pre-trained optimal summary generation model to determine the importance score of each speech segment in the original speech conversation, wherein the optimal summary generation model includes a two-layer improved long short-term memory network unit and a multi-layer perceptron, the improved long short-term memory network unit includes a horizontal structure and a vertical structure, and both the horizontal structure and the vertical structure are neural network architectures based on an attention mechanism embedded in the long short-term memory network unit;
[0008] A speech dialogue summary of the original speech dialogue is determined according to the importance score of each speech segment.
[0009] In a second aspect, this embodiment provides a device for generating a summary of a power communication scheduling voice conversation, including:
[0010] A feature extraction module is used to extract features from the original voice dialogue of the electric power communication dispatching to obtain a voice feature sequence of the original voice dialogue;
[0011] a score prediction module, configured to input the speech feature sequence into a pre-trained optimal summary generation model to determine the importance score of each speech segment in the original speech conversation, wherein the optimal summary generation model includes a two-layer improved long short-term memory network unit and a multi-layer perceptron, the improved long short-term memory network unit including a horizontal structure and a vertical structure, both of which are neural network architectures based on an attention mechanism embedded in the long short-term memory network unit;
[0012] A summary generation module is used to generate a speech dialogue summary of the original speech dialogue according to the importance score of each speech segment.
[0013] An embodiment of the present invention provides a method and device for summarizing voice conversations in electric power communication dispatching. The method comprises: extracting features from an original voice conversation in electric power communication dispatching to obtain a voice feature sequence of the original voice conversation; inputting the voice feature sequence into a pre-trained optimal summary generation model to determine the importance score of each voice segment in the original voice conversation, wherein the optimal summary generation model comprises a two-layer improved long short-term memory network unit and a multi-layer perceptron, wherein the improved long short-term memory network unit comprises a horizontal structure and a vertical structure, each of which is a neural network architecture based on an attention mechanism embedded in the long short-term memory network unit; and determining a voice conversation summary of the original voice conversation based on the importance score of each voice segment. The above technical solution, based on the optimal summary generation model of the improved LSTM unit, integrates the Transformer architecture into the LSTM unit, and realizes bidirectional modeling of the voice conversation sequence through two layers of improved LSTM units. Based on the optimal summary generation model, semantically guided summary generation of electric power communication dispatching voice conversations is achieved. By automatically processing and summarizing voice recordings of electric power calls, the efficiency and accuracy of electric power communication dispatching are improved.
[0014] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0016] Figure 1 A flowchart of a method for generating a summary of a power communication scheduling voice dialogue provided in the first embodiment of the present invention;
[0017] Figure 2 A schematic structural diagram of an optimal model for summarization generation in the execution of a method for summarization generation of a power communication scheduling voice dialogue provided in the first embodiment of the present invention;
[0018] Figure 3 A schematic diagram of the structure of an improved long short-term memory network unit in the optimal model for summary generation provided in Example 1 of the present invention;
[0019] Figure 4 A schematic diagram of the horizontal structure of an improved long short-term memory network unit in the optimal model for summary generation provided in the first embodiment of the present invention;
[0020] Figure 5 A schematic diagram of the vertical structure of an improved long short-term memory network unit in the optimal model for summary generation provided in the first embodiment of the present invention;
[0021] Figure 6 A flowchart of another method for generating a summary of a power communication scheduling voice dialogue provided in the second embodiment of the present invention;
[0022] Figure 7 This is a schematic diagram of the structure of a summary generation device for power communication scheduling voice dialogue provided by the third embodiment of the present invention;
[0023] Figure 8 This is a structural diagram of an electronic device provided in Example 4 of the present invention. DETAILED DESCRIPTION
[0024] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0025] It should be noted that the terms "current", "next", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0026] Example 1
[0027] Figure 1 This is a flow chart of a method for generating summaries of power communication dispatching voice conversations provided in a first embodiment of the present invention. This method is applicable to situations where summaries of power communication dispatching voice conversations are automatically generated. This method can be performed by a summary generation device for power communication dispatching voice conversations. The summary generation device for power communication dispatching voice conversations can be implemented in the form of hardware and / or software and is generally integrated into an electronic device.
[0028] like Figure 1 As shown, the method for generating a summary of a power communication scheduling voice dialogue provided in the first embodiment may specifically include the following steps:
[0029] S101. Extract features from an original voice dialogue of electric power communication dispatching to obtain a voice feature sequence of the original voice dialogue.
[0030] In this embodiment, the original voice dialogue of power communication dispatch can be specifically understood as the voice dialogue generated when staff communicate about power communication dispatch during the power communication dispatch process. For example, if a maintenance worker calls a dispatcher at the dispatch center to confirm information related to power communication dispatch, the voice dialogue generated during this process is the original voice dialogue of power communication dispatch. For each single speech segment in the original voice dialogue, the single speech segment is voice segmented into several speech segments. Considering that the input of the model is the features corresponding to the original voice dialogue, after voice segmenting each single speech segment into several speech segments, a speech feature extraction model can be used to extract features from each speech segment, thereby obtaining the speech features corresponding to each speech segment in the single speech segment. The speech features corresponding to each speech segment in a single speech segment constitute the speech feature sequence of the single speech segment. The speech feature sequences of each single speech segment are combined to form the speech feature sequence of the original voice dialogue.
[0031] S102: Input the speech feature sequence into a pre-trained optimal summary generation model to determine the importance score of each speech segment in the original speech conversation.
[0032] Among them, the optimal model for summary generation includes a two-layer improved long short-term memory network unit and a multi-layer perceptron. The improved long short-term memory network unit includes a horizontal structure and a vertical structure. Both the horizontal structure and the vertical structure are neural network architectures based on the attention mechanism embedded in the long short-term memory network unit.
[0033] In this embodiment, the optimal model for summarization generation consists of two layers of improved Long Short-Term Memory (LSTM) units. The improved LSTM units are integrated into the Transformer structure and are called Trans_LSTM units. The speech feature sequences are modeled in the forward and backward directions respectively to capture temporal dependencies. For example, Figure 2 This is a structural diagram of an optimal model for summarizing generation in the execution of a method for summarizing a power communication scheduling voice conversation provided in the first embodiment of the present invention, as shown in FIG. Figure 2 As shown, the T speech segments of each single speech segment are input into each Trans_LSTM unit, and each speech segment is represented as S1, S2, S3, ... S T This model uses a two-layer LSTM unit structure, with the lower layer being the forward sequence and the upper layer being the backward sequence. Forward refers to the association calculation between each speech segment and its previous speech segment, and backward refers to the association calculation between each speech segment and its subsequent speech segment. Each speech segment is input into the corresponding Trans_LSTM unit in the forward and backward sequences respectively, and feature reconstruction and importance score output are achieved through the multi-layer perceptron (MLP). The output importance scores are represented as Y1, Y2, Y3, ... Y T .
[0034] Following the above description, Figure 3 This is a schematic diagram of the structure of the improved long short-term memory network unit in the optimal model for summary generation provided in the first embodiment of the present invention, as shown in FIG. Figure 3 As shown in the figure, each Trans_LSTM unit is composed of two components: vertical and horizontal, corresponding to inter-segment and inter-dialogue relational operations, respectively. The horizontal structure takes as input the input feature vector (i.e., the current speech features) and the previous state feature vector as input, and outputs the next state feature vector as output. The vertical structure takes as input the input feature vector (i.e., the current speech features) and the previous state feature vector as input, and outputs importance features.
[0035] In this embodiment, each Trans_LSTM unit is composed of two directions, vertical and horizontal, and the main components of these two directions include self-attention modules and cross-attention modules. Figure 4 This is a schematic diagram of the horizontal structure of the improved long short-term memory network unit in the optimal model for summary generation provided in the first embodiment of the present invention, as shown in FIG. Figure 4 As shown, the horizontal structure includes the Transformer's self-attention and criss-cross attention modules, as well as a forget gate, to process long sequences. The self-attention and criss-cross attention modules are similar to those in the vertical structure, but are used to process relationships within a single speech segment. The input to the self-attention module is the feature vector K, V, and Q obtained by linearly transforming it from the previous Trans_LSTM horizontal unit. Q represents the query, K represents the key, and V represents the value. This is the core component of the self-attention mechanism in the Transformer model. The forget gate allows the model to selectively retain or discard past information, thereby processing an entire speech sequence.
[0036] For example, Figure 5 This is a schematic diagram of the vertical structure of the double-layer improved long short-term memory network unit in the optimal model for summary generation provided in the first embodiment of the present invention, as shown in FIG. Figure 5 As shown in the figure, the vertical direction includes a self-attention module and a cross-attention module, which process the input speech segment features. The self-attention module in the Transformer calculates the correlation between the speech segment features. The cross-attention module in the Transformer combines the speech features of the previous speech segment with the speech features of the current speech segment to integrate historical information. The output is a new feature, denoted as the output importance feature.
[0037] The input of each layer of the model is the speech features extracted from the speech segment, which are then passed vertically through the Trans_LSTM module to obtain the output. The output is then passed through a multi-layer perceptron to obtain a predicted importance score, which indicates the likelihood of the segment being selected as a summary.
[0038] S103: Determine a speech dialogue summary of the original speech dialogue based on the importance score of each speech segment.
[0039] In this embodiment, several speech segments with the highest scores are extracted to form extractive summaries of the original input speech audio. Specifically, based on a preset speech conversation summary length, the number of speech segments required to generate a speech conversation summary of the original speech conversation is determined, denoted as the target number of required speech segments. The speech segments are ranked from high to low according to their predicted importance scores, and a target number of speech segments are selected from high to low to form the speech conversation summary of the original speech conversation. For example, the length of the generated summary can be specified by a predefined time length or a predefined number of words, thereby determining a true speech conversation summary of a power communication dispatch speech conversation sample. Unlike existing technologies, this approach, by introducing semantic guidance, can effectively improve the efficiency and accuracy of summarization of power communication dispatch speech recordings. This technical solution has a high degree of automation, reduces manual intervention, and improves work efficiency. Precise semantic analysis ensures the relevance and completeness of the summary content. It is also highly adaptable and can be adjusted to different power communication dispatch scenarios and needs. Furthermore, it can address issues such as vanishing gradients and excessive computational resource consumption that can occur when processing multiple speech summaries.
[0040] The above technical solution is based on the optimal model for summary generation of improved LSTM units, integrates the Transformer architecture into the LSTM units, realizes bidirectional modeling of voice dialogue sequences through two layers of improved LSTM units, and realizes the summary generation of semantically guided power communication dispatch voice dialogues based on the optimal model for summary generation. By automatically processing and summarizing power call voice records, the efficiency and accuracy of power communication dispatch are improved.
[0041] As an optional embodiment of the present invention, based on the above embodiment, the training steps of the optimal model for summarization generation may be optimized, including:
[0042] a1) Obtain a training sample set of electric power communication dispatch voice dialogues.
[0043] In this embodiment, the electric power communication scheduling voice dialogue training sample set includes at least a set number of training sample pairs, where the training sample pairs include electric power communication scheduling voice dialogue samples and real voice dialogue summaries of the electric power communication scheduling voice dialogue samples.
[0044] Optionally, based on the above embodiment, the steps for constructing the electric power communication scheduling voice dialogue training sample set may be optimized, including:
[0045] a11) obtaining a set number of groups of electric power communication dispatching voice dialogue samples, and performing voice segmentation on the electric power communication dispatching voice dialogue samples to obtain voice segment samples of the electric power communication dispatching voice dialogue samples.
[0046] In this embodiment, voice conversation data related to power communication and dispatch, including multiple categories such as site, optical cable, and service names, is collected. Data is then sorted to remove data with excessive noise or that cannot be correctly labeled. This data is recorded as power communication and dispatch voice conversation samples. The speech in each set of power communication and dispatch voice conversation samples is then segmented into voice segments, which are recorded as voice segment samples. In this embodiment, there is no specific limit on the set number; a large number of power communication and dispatch voice conversation samples can be obtained. For example, the set number can be set to 500. For example, the speech in each set of power communication and dispatch voice conversation samples is segmented into 1-second sample voice segments.
[0047] a12) performing importance scoring on each speech segment sample to obtain a true importance score of each speech segment sample.
[0048] In this embodiment, after segmenting the speech into sample segments, each segment is assigned an importance score. Professional speech proofreaders and power communication dispatch staff collaborate to annotate the segments, obtaining the true importance score for each segment. For example, annotators are asked to rate the importance of each segment. Each segment is assigned an importance score by 20 annotators, and the average score is calculated as the importance score for the segment.
[0049] a13) determining a true voice dialogue summary of the electric power communication dispatch voice dialogue sample based on the true importance score of each voice segment sample.
[0050] Specifically, based on the importance scores of each speech segment sample, the highest-scoring segments are extracted to form an extractive summary of the original input speech audio. The length of the generated summary can be specified by pre-defining the time length or the number of words to determine a true speech dialogue summary for a power communication dispatch speech dialogue sample.
[0051] a14) The electric power communication dispatching voice dialogue sample and the real voice dialogue summary corresponding to the electric power communication dispatching voice dialogue sample are used as a set of training sample pairs, and each training sample pair constitutes an electric power communication dispatching voice dialogue training sample set.
[0052] Specifically, the electric power communication dispatching voice dialogue sample and the real voice dialogue corresponding to the electric power communication dispatching voice dialogue sample are taken as a group of training sample pairs, and multiple training sample pairs constitute the electric power communication dispatching voice dialogue training sample set.
[0053] The above technical solution specifies the steps for constructing a training sample set for electric power communication dispatching voice dialogues, providing basic data for subsequent model training.
[0054] b1) extracting features from the electric power communication dispatching voice dialogue sample to obtain a voice feature sequence sample of the electric power communication dispatching voice dialogue sample.
[0055] In this example, a speech segmentation process is performed on a sample of electric power communication dispatching voice conversations to obtain multiple speech segment samples. A speech feature extraction model is then used to convert each sample speech segment into discrete features, recorded as speech feature sequence samples. These speech feature sequence samples serve as input for model training.
[0056] c1) Inputting the speech feature sequence sample into the power dialogue speech summary generation model based on the improved long short-term memory network unit to obtain the predicted importance score of each speech segment sample in the speech feature sequence sample.
[0057] Specifically, a speech feature sequence sample is input into a speech summary generation model for power conversations based on an improved long short-term memory network unit. The model then outputs a predicted importance score for each sample speech segment. The sample speech feature sequence serves as the input to the initial network model, and the predicted importance score for each sample speech segment serves as the output of the speech summary generation model.
[0058] d1) determining a predicted speech dialogue summary of the electric power communication dispatching speech dialogue sample based on the predicted importance score of each speech segment sample.
[0059] In this embodiment, the highest-scoring speech segment samples are extracted to form an extractive summary of the speech conversation sample. The length of the generated summary can be specified by pre-defining the time length of the summary or by a pre-defined number of words. Based on this, a predicted speech conversation summary of the speech conversation sample is obtained.
[0060] e1) Determine a comprehensive loss value of the predicted speech conversation summary and the actual speech conversation summary based on a preset loss function.
[0061] In this embodiment, model training uses three loss functions: classification loss, reconstruction loss, and diversity loss. Based on the predicted and actual speech conversation summaries, the three loss functions are combined to determine three loss values for each. A combined loss value for the predicted and actual speech conversation summaries is then determined based on these three loss values.
[0062] As a specific implementation method, the step of determining the comprehensive loss value of the predicted speech conversation summary and the actual speech conversation summary based on the predicted speech conversation summary and the actual speech conversation summary in combination with a preset loss function can be optimized, including:
[0063] d11) Substituting the predicted importance score and the true importance score of the speech segment sample into a pre-built classification loss function to determine a first loss value of the predicted importance score and the true importance score of the speech segment sample.
[0064] Among them, the classification loss function is used to quantify the probability of overlap between the predicted speech dialogue summary and the real speech dialogue summary. Considering that the classification loss function is used to train the model to identify whether the speech segment belongs to the summary, the binary cross entropy loss function is directly used. Where T is the total number of speech segments, Score the true importance of the predicted i-th speech segment, that is, whether it is included in the summary. i The model predicts the probability of the i-th speech segment in the cross entropy. This probability value can be calculated from the importance score through the activation (sigmoid) function. Lc is the first loss value. The classification loss function can be expressed as:
[0065]
[0066] Specifically, the predicted importance score and the true importance score of the sample speech segment are substituted into a pre-built classification loss function to determine a first loss value of the predicted importance score and the true importance score of the sample speech segment.
[0067] d12) Substituting the speech features of the speech segments and the speech features in the speech feature sequence samples contained in the predicted speech conversation summary into a pre-constructed reconstruction loss function, and determining a second loss value of the speech features of the speech segments and the speech features in the speech feature sequence samples contained in the predicted speech conversation summary.
[0068] In this embodiment, the reconstruction loss function is used to quantify the similarity between the predicted speech dialogue summary and the real speech dialogue summary. Considering that the reconstruction loss function is used to ensure the similarity between the generated summary and the original audio content, the mean square error is used to calculate the similarity. i is the feature vector of the speech segment of the i-th segment, is the speech feature of the reconstructed speech segment in the summary, Lr represents the second loss value, and the reconstruction loss function is expressed as: Specifically, the speech features of the speech segments contained in the predicted speech conversation summary and the speech features in the speech feature sequence are substituted into the pre-constructed reconstruction loss function to determine the second loss value of the speech features of the speech segments contained in the predicted speech conversation summary and the speech features in the speech feature sequence sample.
[0069] d13) Substituting the speech features of each speech segment in the predicted speech conversation summary into a pre-constructed diversity loss function to determine a third loss value of the speech features of each speech segment in the predicted speech conversation summary.
[0070] The diversity loss function is used to quantify the diversity of the predicted speech dialogue summary. Considering that the diversity loss function is used to penalize the similarity of the selected audio segments, thereby ensuring the diversity of the speech segments in the summary, the pairwise cosine similarity is used to calculate their correlation. M is the number of speech segments finally selected, and is the speech feature of the i-th and j-th speech segments reconstructed by the model, L d Represents the third loss value, and the diversity loss function is expressed as:
[0071] Specifically, the speech features of each speech segment in the predicted speech conversation summary are substituted into a pre-constructed diversity loss function to determine a third loss value of the speech features of each speech segment in the predicted speech conversation summary.
[0072] d14) performing a weighted summation on the first loss value, the second loss value, and the third loss value, and using the weighted summation result as the comprehensive loss value of the predicted speech conversation summary and the actual speech conversation summary.
[0073] Specifically, a weighted sum is performed on the first loss value, the second loss value, and the third loss value, and the result is used as the comprehensive loss value.
[0074] The above technical solution concretizes the steps for determining the comprehensive loss value. The loss value is confirmed to train the model by identifying whether the voice clip belongs to a summary, the similarity between the generated summary and the original audio content, and the similarity of the selected audio clips, thereby ensuring the accuracy and diversity of the summaries generated by the trained optimal summary generation model.
[0075] e1) Based on the comprehensive loss value, adjust the parameters of the electric power conversation speech summary generation model, and return to re-execute the step of extracting features from the electric power communication scheduling speech conversation sample until the electric power conversation speech summary generation model that meets the iteration termination condition is obtained as the optimal summary generation model.
[0076] Specifically, based on the comprehensive loss value, the parameters of the power conversation speech summary generation model are adjusted, and steps a1) to e1) are repeated to train the model until an iteration termination condition is met. The trained initial network model can then be used as the optimal model for summary generation. For example, the iteration termination condition can be when the comprehensive loss value falls below a set loss threshold.
[0077] The above technical solution specifies the training steps of the optimal model for summary generation and provides a basis for the subsequent summary generation.
[0078] Example 2
[0079] Figure 6A flow chart of another method for generating a summary of a voice conversation of electric power communication dispatching provided in the second embodiment of the present invention. This embodiment is a further optimization of the above embodiment. In this embodiment, the method further limits and optimizes "performing feature extraction on the original voice conversation of electric power communication dispatching to obtain a voice feature sequence of the original voice conversation", as well as "inputting the voice feature sequence into the pre-trained optimal model for summary generation to determine the importance score of each voice segment in the original voice conversation", and "determining the voice conversation summary of the original voice conversation based on the importance score of each voice segment".
[0080] like Figure 6 As shown, this embodiment 2 provides a method for generating a summary of a power communication scheduling voice dialogue, which specifically includes the following steps:
[0081] S201: For each single speech segment in the original speech conversation, perform speech segmentation on the single speech segment to obtain at least one speech segment of the single speech segment.
[0082] In this embodiment, the original voice conversation contains several single-segment voices. For example, it is assumed that during the power communication dispatching process, a voice conversation occurs between a maintenance worker and a dispatcher. This voice conversation is recorded as the original voice conversation. The original voice conversation contains multiple single-segment voices of the maintenance worker and multiple single-segment voices of the dispatcher. The multiple voice segments generated each time the maintenance worker or the dispatcher has a conversation are recorded as single-segment voices. For each single-segment voice in the original voice conversation, the single-segment voice is segmented into several voice segments. For example, each single-segment voice is segmented at intervals of 1s to obtain multiple voice segments of the single-segment voice. It should be noted that the method of segmenting the original voice conversation is the same as the method of segmenting the voice conversation samples during model training.
[0083] S202: Using a preset speech feature extraction model, extract features from each speech segment to obtain a speech feature sequence for a single speech segment.
[0084] In this embodiment, considering that the model input is the features corresponding to the original speech conversation, after each single speech segment is segmented into several speech segments, the speech feature extraction model can be used to extract features from each speech segment, thereby obtaining the speech features corresponding to each speech segment in the single speech segment. The speech features corresponding to each speech segment in a single speech segment constitute the speech feature sequence of the single speech segment. For each single speech segment, its corresponding speech feature sequence is determined.
[0085] For example, the speech feature extraction model may be a Wav2Vec model, which is a speech pre-training model primarily used for speech recognition tasks. The Wav2Vec model is used to convert raw speech dialogue into discrete features, forming a speech feature sequence for a single speech segment.
[0086] S203: Combining the speech feature sequences of the individual speech segments into a speech feature sequence of the original speech dialogue.
[0087] Specifically, the speech feature sequences of each single speech segment are combined to form the speech feature sequence of the original speech dialogue. The speech feature sequence of the original speech dialogue formed can be expressed as: v ={f v1 ,f v2 ,……,f vT ,}. T is the number of individual single-segment speech in the conversation. Each single-segment speech is divided into t speech segments, recorded as As input to the model, the output is a set of importance scores in
[0088] S204: Input the current speech feature and the previous state feature in the speech feature sequence into the horizontal structure of the double-layer improved long short-term memory network unit in the optimal model for summary generation to determine the current state feature.
[0089] In this embodiment, the horizontal structure of the double-layer improved long short-term memory network unit in the summary generation optimal model is used to generate state features, which is used to consider the relationship between this speech segment and the previous speech segments. The speech feature sequence is composed of speech features corresponding to multiple speech segments. Each speech feature in the traversal of the speech feature sequence can be regarded as the current speech feature. Based on the current speech feature and the previous state feature, the next state feature obtained through horizontal structure processing is recorded as the current state feature. The current state feature contains the relationship between the current speech feature and the previous speech feature. The previous state feature can be specifically understood as the result of horizontal structure processing of the previous speech feature and the previous state feature.
[0090] As a specific implementation method, the current speech feature and the previous state feature in the speech feature sequence can be optimized to input the horizontal structure of the two-layer improved long short-term memory network unit in the optimal summary generation model to determine the current state feature, including:
[0091] The current speech features and the previous state features of the speech feature sequence are input into the horizontal structure of the double-layer improved long short-term memory network unit, the current speech features and the previous state features are linearly transformed, and the new speech features are output through the attention mechanism. After being processed by the forget gate, the current state features are output.
[0092] Continue to refer Figure 4 , the input of the horizontal structure is the input feature vector (i.e. the current speech feature) and the previous state feature vector (i.e. the previous state feature), and the output is the next state feature vector. The horizontal structure contains the Transformer's self-attention module and cross-attention module, as well as the forget gate, to process long sequence information. The self-attention module and the cross-attention module are similar to the vertical direction, but are used to process the relationship between single-segment speech. The input of the self-attention module is K, V, and Q obtained by linearly changing the feature vector transmitted by the previous Trans_LSTM horizontal unit. The forget gate allows the model to selectively retain or discard past information, thereby processing an entire speech sequence.
[0093] S205: Input the current speech feature and the current state feature into the vertical structure of the improved long short-term memory network unit in the optimal model for summary generation, and determine the output importance feature of the current speech feature.
[0094] In this embodiment, the vertical structure of the improved long short-term memory network unit in the optimal summary generation model is used to combine state features and current speech features to output importance features. The output importance features can be represented in the form of a matrix. In addition to considering the current input speech segment, the importance of a speech segment can also consider the relationship between this speech segment and previous speech segments. This relationship is provided by the previous state feature, and how this state feature is obtained needs to be determined by the horizontal structure of the two-layer improved long short-term memory network unit in the optimal summary generation model.
[0095] As a specific implementation method, the vertical structure of the improved long short-term memory network unit in the optimal model for summarization generation can be optimized by inputting the current speech features and the current state features into the optimal model for summarization generation to determine the output speech features, including:
[0096] The speech feature sequence and the current state feature are input into the horizontal structure of the improved long short-term memory network unit, the current speech feature and the current state feature are linearly transformed, and processed through the attention mechanism to obtain the output importance feature of the current speech feature.
[0097] Continue to refer Figure 5 , Vertical direction: Contains self-attention module and cross-attention module to process the input speech segment features. The self-attention module in Transformer is used to calculate the relationship between speech segment features. The cross-attention module in Transformer combines the features of the previous speech segment with the features of the current speech segment to achieve the fusion of historical information. The input of the vertical structure is the input feature vector (i.e., the current speech feature) and the feature vector of the previous state, and the output is the importance feature.
[0098] S206: reconstruct the output importance features of the current speech features through the multi-layer perceptron in the summary generation optimal model, and obtain the importance score of each speech segment in the original speech dialogue.
[0099] Specifically, a multi-layer perceptron is used to reconstruct the importance features of the multi-layer output and perform score prediction to obtain the importance score of each voice segment in the original voice conversation. Figure 3 The input to the multi-layer perceptron includes the original input speech features and the output importance features of each layer of Trans_LSTM units. The multi-layer speech features are reconstructed and score predicted to obtain the importance score of each speech segment in the original speech conversation.
[0100] S207: Determine the target number of required voice segments based on the preset voice dialogue summary length.
[0101] In this embodiment, several speech segments with the highest scores are extracted to form an extractive summary of the original input speech audio. Specifically, based on a preset speech dialogue summary length, a number of speech segments required to generate a speech dialogue summary of the original speech dialogue is determined, denoted as the target number of speech segments required. Exemplarily, the length of the generated summary is specified by a predefined time length or a predefined number of words.
[0102] S208 : Select a target number of voice segments from each voice segment according to the predicted importance scores from high to low to form a voice dialogue summary of the original voice dialogue.
[0103] Specifically, each voice segment is arranged from high to low according to the predicted importance score, and a target number of voice segments are selected from high to low as the voice dialogue summary of the original voice dialogue.
[0104] The above technical solution specifies the steps of extracting features from the original voice conversation, scoring each voice segment in the voice feature sequence based on the improved summary generation optimal model, and selecting the voice conversation summary based on the scoring results. First, the original voice conversation is divided into multiple voice segments, and features are extracted from these segments using the voice feature extraction model to obtain a voice feature sequence. Then, using the improved summary generation optimal model, the Transformer architecture is integrated into the LSTM unit. Two layers of improved LSTM units are used to achieve bidirectional modeling of the voice feature sequence. The relationship between voice segments is processed vertically, and the relationship between single voice segments is processed horizontally. This realizes the semantic-guided summary generation of power communication dispatch voice conversations. By automatically processing and summarizing power call voice records, the efficiency and accuracy of power communication dispatch are improved.
[0105] Example 3
[0106] Figure 7 This is a schematic diagram of the structure of a summary generation device for power communication and dispatch voice dialogues provided in the third embodiment of the present invention. The device is applicable to the case of generating summaries for power communication and dispatch voice dialogues. The summary generation device for power communication and dispatch voice dialogues can be implemented in the form of hardware and / or software and is generally integrated into electronic devices. Figure 7 As shown, the device includes: a feature extraction module 31, a score prediction module 32 and a summary generation module 33, wherein:
[0107] A feature extraction module 31 is used to extract features from the original voice dialogue of the electric power communication dispatching to obtain a voice feature sequence of the original voice dialogue;
[0108] a score prediction module 32 for inputting the speech feature sequence into a pre-trained optimal summary generation model to determine the importance score of each speech segment in the original speech conversation, wherein the optimal summary generation model includes a two-layer improved long short-term memory network unit and a multi-layer perceptron, wherein the improved long short-term memory network unit includes a horizontal structure and a vertical structure, and both the horizontal structure and the vertical structure are neural network architectures based on an attention mechanism embedded in the long short-term memory network unit;
[0109] The summary generation module 33 is used to generate a speech dialogue summary of the original speech dialogue according to the importance score of each speech segment.
[0110] The above technical solution is based on the optimal model for summary generation of improved LSTM units, integrates the Transformer architecture into the LSTM units, realizes bidirectional modeling of voice dialogue sequences through two layers of improved LSTM units, and realizes the summary generation of semantically guided power communication dispatch voice dialogues based on the optimal model for summary generation. By automatically processing and summarizing power call voice records, the efficiency and accuracy of power communication dispatch are improved.
[0111] Optionally, the feature extraction module 31 is specifically configured to:
[0112] For each single speech segment in the original speech conversation, performing speech segmentation on the single speech segment to obtain at least one speech segment of the single speech segment;
[0113] Using the preset speech feature extraction model, feature extraction is performed on each speech segment to obtain the speech feature sequence of a single speech segment;
[0114] The speech feature sequences of each single speech segment are combined into the speech feature sequence of the original speech dialogue.
[0115] Optionally, the score prediction module 32 includes:
[0116] A first feature determination unit is configured to input the current speech feature and the previous state feature in the speech feature sequence into the horizontal structure of the improved long short-term memory network unit in the summary generation optimal model to determine the current state feature;
[0117] a second feature determination unit, configured to input the current speech feature and the current state feature into the vertical structure of the improved long short-term memory network unit in the optimal model for summary generation, and determine an output importance feature of the current speech feature;
[0118] The scoring output unit is used to reconstruct the output importance features of the current speech features through the multi-layer perceptron in the summary generation optimal model, and obtain the importance score of each speech segment in the original speech dialogue.
[0119] Optionally, the first feature determination unit is specifically configured to:
[0120] The current speech features and the previous state features of the speech feature sequence are input into the horizontal structure of the improved long short-term memory network unit, the current speech features and the previous state features are linearly transformed, and the new speech features are output through the attention mechanism. After being processed by the forget gate, the current state features are output.
[0121] Optionally, the second feature determination unit is specifically configured to:
[0122] The current speech features and current state features are input into the vertical structure of the improved long short-term memory network unit, the current speech features and current state features are linearly transformed, and processed through the attention mechanism to obtain the output importance features of the current speech features.
[0123] Optionally, the summary generation module 33 is specifically configured to:
[0124] Determine the target number of required voice segments based on the preset voice dialogue summary length;
[0125] A target number of speech segments are selected from each speech segment according to the predicted importance scores from high to low to form a speech dialogue summary of the original speech dialogue.
[0126] Optionally, the device further includes a model training module, the model training module including:
[0127] a sample acquisition unit, configured to acquire a power communication dispatching voice dialogue training sample set, the power communication dispatching voice dialogue training sample set including at least a set number of training sample pairs, the training sample pairs including power communication dispatching voice dialogue samples and real voice dialogue summaries of the power communication dispatching voice dialogue samples;
[0128] A sample feature extraction unit is used to extract features from the electric power communication dispatching voice dialogue sample to obtain a voice feature sequence sample of the electric power communication dispatching voice dialogue sample;
[0129] A sample scoring output unit is used to input the speech feature sequence sample into the power conversation speech summary generation model based on the improved long short-term memory network unit to obtain the predicted importance score of each speech segment sample in the power communication scheduling speech conversation sample;
[0130] a sample summary determination unit, configured to determine a predicted voice dialogue summary of the electric power communication dispatch voice dialogue sample based on the predicted importance score of each voice segment sample;
[0131] a sample loss value determination unit, configured to determine a comprehensive loss value of the predicted speech conversation summary and the actual speech conversation summary based on the predicted speech conversation summary and the actual speech conversation summary in combination with a preset loss function;
[0132] The model determination unit is used to adjust the parameters of the power conversation speech summary generation model based on the comprehensive loss value, return to re-execute the step of extracting features from the power communication scheduling speech conversation sample, and obtain the power conversation speech summary generation model that meets the iteration termination condition as the optimal summary generation model.
[0133] Optionally, the sample acquisition unit is specifically configured to:
[0134] Acquire a set number of groups of electric power communication dispatch voice dialogue samples, and perform voice segmentation on the electric power communication dispatch voice dialogue samples to obtain voice segment samples of the voice dialogue samples;
[0135] Score the importance of each speech segment sample to obtain the true importance score of each speech segment sample;
[0136] Determine the true voice dialogue summary of the electric power communication dispatch voice dialogue sample based on the true importance score of each voice segment sample;
[0137] The electric power communication dispatching voice dialogue samples and the real voice dialogue summaries corresponding to the electric power communication dispatching voice dialogue samples are taken as a group of training sample pairs, and each training sample pair constitutes an electric power communication dispatching voice dialogue training sample set.
[0138] Optionally, the sample loss value determining unit is specifically configured to:
[0139] Substituting the predicted importance score and the true importance score of the speech segment sample into a pre-built classification loss function to determine a first loss value of the predicted importance score and the true importance score of the speech segment sample. The classification loss function is used to quantify the probability of overlap between the predicted speech conversation summary and the true speech conversation summary.
[0140] Substituting the speech features of the speech segment and the speech features in the speech feature sequence sample contained in the predicted speech conversation summary into a pre-constructed reconstruction loss function, and determining a second loss value of the speech features of the speech segment and the speech features in the speech feature sequence sample contained in the predicted speech conversation summary. The reconstruction loss function is used to quantify the similarity between the predicted speech conversation summary and the real speech conversation summary;
[0141] Substituting the speech features of each speech segment in the predicted speech conversation summary into a pre-constructed diversity loss function to determine a third loss value of the speech features of each speech segment in the predicted speech conversation summary. The diversity loss function is used to quantify the diversity of the predicted speech conversation summary.
[0142] The first loss value, the second loss value, and the third loss value are weightedly summed, and the weighted summation result is used as the comprehensive loss value of the predicted speech conversation summary and the actual speech conversation summary.
[0143] The summary generation device for electric power communication dispatching voice dialogue provided by the embodiment of the present invention can execute the summary generation method for electric power communication dispatching voice dialogue provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0144] Example 4
[0145] Figure 8 A schematic diagram of the structure of an electronic device provided for embodiment four of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein.
[0146] like Figure 8As shown, the electronic device 40 includes at least one processor 41 and a memory, such as a read-only memory (ROM) 42, a random access memory (RAM) 43, etc., which is communicatively connected to the at least one processor 41. The memory stores a computer program that can be executed by the at least one processor, and the processor 41 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 42 or the computer program loaded from the storage unit 48 into the random access memory (RAM) 43. Various programs and data required for the operation of the electronic device 40 can also be stored in the RAM 43. The processor 41, ROM 42, and RAM 43 are connected to each other via a bus 44. An input / output (I / O) interface 45 is also connected to the bus 44.
[0147] Multiple components in the electronic device 40 are connected to the I / O interface 45, including an input unit 46, such as a keyboard, a mouse, etc.; an output unit 47, such as various types of displays, speakers, etc.; a storage unit 48, such as a magnetic disk, an optical disk, etc.; and a communication unit 49, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 49 allows the electronic device 40 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0148] Processor 41 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of processor 41 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any other suitable processors, controllers, microcontrollers, etc. Processor 41 executes the various methods and processes described above, such as the method for generating summaries of power communication scheduling voice conversations.
[0149] In some embodiments, the method for generating a summary of a power communication scheduling voice conversation can be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 48. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 40 via the ROM 42 and / or the communication unit 49. When the computer program is loaded into the RAM 43 and executed by the processor 41, one or more steps of the method for generating a summary of a power communication scheduling voice conversation described above can be performed. Alternatively, in other embodiments, the processor 41 can be configured to perform the method for generating a summary of a power communication scheduling voice conversation by any other appropriate means (e.g., by means of firmware).
[0150] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0151] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0152] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0153] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0154] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0155] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.
[0156] An embodiment of the present invention further provides a computer program product, including a computer program, which, when executed by a processor, implements the method for generating a summary of a power communication scheduling voice dialogue as provided in any embodiment of the present invention.
[0157] The computer program product may be implemented by writing computer program code for performing the operations of the present disclosure in one or more programming languages, or a combination thereof, including, but not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0158] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.
[0159] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A method for generating a summary of a power communication dispatch voice dialogue, characterized in that: include: Extracting features from an original voice conversation of the electric power communication dispatching to obtain a voice feature sequence of the original voice conversation; Inputting the speech feature sequence into a pre-trained optimal summary generation model to determine the importance score of each speech segment in the original speech conversation, wherein the optimal summary generation model includes a two-layer improved long short-term memory network unit and a multi-layer perceptron, the improved long short-term memory network unit includes a horizontal structure and a vertical structure, and both the horizontal structure and the vertical structure are neural network architectures based on an attention mechanism embedded in the long short-term memory network unit; A speech dialogue summary of the original speech dialogue is determined according to the importance score of each speech segment.
2. The method according to claim 1, characterized in that The feature extraction of the original voice dialogue of the electric power communication scheduling to obtain the voice feature sequence of the original voice dialogue includes: For each single speech segment in the original speech conversation, performing speech segmentation on the single speech segment to obtain at least one speech segment of the single speech segment; Using a preset speech feature extraction model, feature extraction is performed on each of the speech segments to obtain a speech feature sequence of the single speech segment; The speech feature sequences of the single speech segments are combined into a speech feature sequence of the original speech dialogue.
3. The method according to claim 1, characterized in that Inputting the speech feature sequence into a pre-trained optimal summary generation model to determine the importance score of each speech segment in the original speech conversation includes: Inputting the current speech feature and the previous state feature in the speech feature sequence into the horizontal structure of the improved long short-term memory network unit in the summary generation optimal model to determine the current state feature; Inputting the current speech feature and the current state feature into the vertical structure of the improved long short-term memory network unit in the optimal model for summary generation, and determining an output importance feature of the current speech feature; The output importance feature of the current speech feature is reconstructed through a multi-layer perceptron in the summary generation optimal model, and an importance score of each speech segment in the original speech dialogue is obtained.
4. The method according to claim 3, characterized in that The step of inputting the current speech feature and the previous state feature in the speech feature sequence into the horizontal structure of the improved long short-term memory network unit in the summary generation optimal model to determine the current state feature includes: The current speech feature and the previous state feature of the speech feature sequence are input into the horizontal structure of the improved long short-term memory network unit, the current speech feature and the previous state feature are linearly transformed, and the new speech feature is output through the attention mechanism, and the current state feature is output after being processed by the forget gate.
5. The method according to claim 3, characterized in that Inputting the current speech feature and the current state feature into the vertical structure of the improved long short-term memory network unit in the summary generation optimal model, and determining the output importance feature of the current speech feature, includes: The current speech feature and the current state feature are input into the vertical structure of the improved long short-term memory network unit, the current speech feature and the current state feature are linearly transformed, and the output importance feature of the current speech feature is obtained through attention mechanism processing.
6. The method according to claim 1, characterized in that Generating a speech dialogue summary of the original speech dialogue according to the predicted importance score of each speech segment includes: Determine the target number of required voice segments based on the preset voice dialogue summary length; The target number of speech segments are selected from the speech segments according to the predicted importance scores from high to low to form a speech dialogue summary of the original speech dialogue.
7. The method according to claim 1, characterized in that The training steps of the optimal summary generation model include: Acquire a power communication dispatch voice dialogue training sample set, wherein the power communication dispatch voice dialogue training sample set includes at least a set number of training sample pairs, wherein the training sample pairs include power communication dispatch voice dialogue samples and real voice dialogue summaries of the power communication dispatch voice dialogue samples; Performing feature extraction on the electric power communication dispatching voice dialogue sample to obtain a voice feature sequence sample of the electric power communication dispatching voice dialogue sample; Inputting the speech feature sequence sample into a power conversation speech summary generation model based on an improved long short-term memory network unit to obtain a predicted importance score of each speech segment sample in the power communication scheduling speech conversation sample; Determining a predicted voice dialogue summary of the electric power communication scheduling voice dialogue sample based on the predicted importance score of each of the voice segment samples; Determining a comprehensive loss value of the predicted voice conversation summary and the real voice conversation summary based on the predicted voice conversation summary and the real voice conversation summary in combination with a preset loss function; Based on the comprehensive loss value, the parameters of the electric power conversation speech summary generation model are adjusted, and the step of extracting features from the electric power communication scheduling speech conversation sample is returned to be re-executed until an electric power conversation speech summary generation model that meets the iteration termination condition is obtained as the optimal summary generation model.
8. The method according to claim 7, characterized in that The steps of constructing the electric power communication scheduling voice dialogue training sample set include: Acquire at least a set number of groups of electric power communication scheduling voice dialogue samples, and perform voice segmentation on the electric power communication scheduling voice dialogue samples to obtain voice segment samples of the electric power communication scheduling voice dialogue samples; Performing importance scoring on each of the speech segment samples to obtain a true importance score of each of the speech segment samples; Determining a true voice dialogue summary of the electric power communication scheduling voice dialogue sample according to the true importance score of each voice segment sample; The electric power communication scheduling voice dialogue sample and the real voice dialogue summary corresponding to the electric power communication scheduling voice dialogue sample are used as a group of training sample pairs, and each of the training sample pairs constitutes an electric power communication scheduling voice dialogue training sample set.
9. The method according to claim 7, characterized in that The step of determining a comprehensive loss value of the predicted voice conversation summary and the real voice conversation summary based on the predicted voice conversation summary and the real voice conversation summary in combination with a preset loss function includes: Substituting the predicted importance score and the true importance score of the speech segment sample into a pre-constructed classification loss function to determine a first loss value of the predicted importance score and the true importance score of the speech segment sample, wherein the classification loss function is used to quantify the probability of overlap between the predicted speech conversation summary and the true speech conversation summary; Substituting the speech features of the speech segment and the speech features in the speech feature sequence sample included in the predicted speech conversation summary into a pre-constructed reconstruction loss function to determine a second loss value for the speech features of the speech segment and the speech features in the speech feature sequence sample included in the predicted speech conversation summary, wherein the reconstruction loss function is used to quantify the similarity between the predicted speech conversation summary and the real speech conversation summary; Substituting the speech features of each speech segment in the predicted speech conversation summary into a pre-constructed diversity loss function to determine a third loss value of the speech features of each speech segment in the predicted speech conversation summary, wherein the diversity loss function is used to quantify the diversity of the predicted speech conversation summary; A weighted sum is performed on the first loss value, the second loss value, and the third loss value, and the weighted sum result is used as the comprehensive loss value of the predicted voice conversation summary and the real voice conversation summary.
10. A device for generating a summary of a power communication dispatch voice dialogue, characterized in that: include: A feature extraction module is used to extract features from the original voice dialogue of the electric power communication dispatching to obtain a voice feature sequence of the original voice dialogue; a score prediction module, configured to input the speech feature sequence into a pre-trained optimal summary generation model to determine the importance score of each speech segment in the original speech conversation, wherein the optimal summary generation model includes a two-layer improved long short-term memory network unit and a multi-layer perceptron, the improved long short-term memory network unit including a horizontal structure and a vertical structure, both of which are neural network architectures based on an attention mechanism embedded in the long short-term memory network unit; A summary generation module is used to generate a speech dialogue summary of the original speech dialogue according to the importance score of each speech segment.