Electric power customer service voice data processing method and system, terminal equipment and computer readable storage medium

By performing speech recognition and part-of-speech matching analysis on the voice data of the power customer service, text statements with high context relevance are selected, which solves the problem of part-of-speech understanding deviation in the power customer service system, realizes accurate identification of customer intentions and business processing, and improves service quality.

CN120356472APending Publication Date: 2025-07-22GUANGDONG POWER GRID CO LTD CUSTOMER SERVICE CENT
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510489462.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

When the power customer service system processes customer voice data, there are problems such as misunderstanding and delay in service requests caused by part-of-speech understanding, which affects the customer experience.

Method used

Text statements are generated through speech recognition, and the preset vocabulary matches the part of speech is used to generate text statements with different combinations of speeches. The context relevance degree is determined through vocabulary correlation analysis, and the text statements closest to the customer's intentions are selected for business processing.

Benefits of technology

Accurately identify customers' true intentions, improve the accuracy and efficiency of power customer service business processing, and improve customer service experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356472A_ABST
    Figure CN120356472A_ABST
Patent Text Reader

Abstract

The invention discloses an electric power customer service voice data processing method and system, terminal equipment and a computer readable storage medium. The method comprises the following steps: acquiring to-be-processed electric power customer service voice data; performing voice recognition on the power customer service voice data to generate a text statement; for each vocabulary in the character statement, matching the current vocabulary with a preset vocabulary, and taking the part-of-speech of all the successfully matched preset vocabulary as the candidate part-of-speech of the current vocabulary; according to the candidate part-of-speech of each vocabulary, generating a plurality of character statements to be selected under different part-of-speech combinations; according to the part-of-speech of each vocabulary in each to-be-selected character statement, performing vocabulary correlation analysis, and determining a context correlation degree; and comparing the context association degree of each to-be-selected character statement with a preset association degree threshold value, and taking the to-be-selected character statement with the maximum context association degree greater than the preset association degree threshold value as a target character statement. According to the invention, the real intention of the client can be accurately reflected.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of signal processing, and particularly to a method, system, terminal device and computer-readable storage medium for processing power customer service voice data. Background Art

[0002] As a core service link in the power industry, the power customer service system plays a crucial role. It is specifically responsible for handling diverse customer service needs, including but not limited to consultations, complaint handling, and repair requests. In the early service mode, power customer service mainly relied on traditional telephone communication lines and manual labor. Through the way of manual answering, patiently and meticulously record and process various service requests from customers.

[0003] However, with the continuous expansion of the power customer group and the increasing diversification and complexity of service demands, the traditional power customer service mode has begun to show limitations. To address this challenge, the power industry has gradually introduced electronic customer service systems, aiming to improve service quality and response speed through more advanced and efficient technical means.

[0004] Although the power customer service system has achieved remarkable results in improving service efficiency, in the actual application process, it still faces a series of technical problems. Especially in the service link of electronic automatic reply, the power customer service system often has inaccurate problems for descriptions of different parts of speech. For example, when a customer issues an instruction "Is the equipment grounding good?", if "grounding" is used as a noun, the meaning of the sentence is to check whether the grounding device of the equipment is installed correctly, firmly connected, and whether the grounding resistance is within the specified range, etc., to ensure good grounding performance of the equipment and guarantee the safety of personnel and equipment; if "grounding" is used as a verb, this sentence may be understood as performing a grounding operation on the equipment, that is, establishing an electrical connection between the equipment and the ground. This deviation in the understanding of parts of speech can easily lead to misunderstandings or delays in handling service requests, thereby having an adverse impact on the customer service experience. Summary of the Invention

[0005] Embodiments of the present invention provide a method, system, terminal device and computer-readable storage medium for processing power customer service voice data, which can accurately reflect the true intentions of customers and improve the experience of customers in the process of business handling.

[0006] An embodiment of the present invention provides a method for processing power customer service voice data, including:

[0007] Obtain the power customer service voice data to be processed;

[0008] Perform speech recognition on the power customer service voice data to generate a text statement;

[0009] For each word in the text statement, match the current word with each preset vocabulary list, and use the parts of speech of all successfully matched preset vocabulary lists as the candidate parts of speech corresponding to the current word. Among them, the preset vocabulary lists include: noun vocabulary list, verb vocabulary list, adjective vocabulary list, adverb vocabulary list, and conjunction vocabulary list;

[0010] Generate several first candidate text statements under different combinations of parts of speech according to the candidate parts of speech of each word;

[0011] For each first candidate text statement, perform lexical relevance analysis on each word according to the parts of speech of the words in the first candidate text statement to determine the context relevance degree of each first candidate text statement;

[0012] Compare the context relevance degree of each first candidate text statement with the preset relevance degree threshold respectively, and use the first candidate text statement whose context relevance degree is greater than the preset relevance degree threshold as the second candidate text statement;

[0013] Use the second candidate text statement with the highest context relevance degree as the target text statement;

[0014] Perform power customer service business processing according to the words and corresponding parts of speech combinations in the target text statement.

[0015] Further, perform speech recognition on the power customer service voice data to generate a text statement, including:

[0016] Select the speech recognition model of the corresponding language from the preset language library according to the language corresponding to the power customer service voice data as the target speech recognition model;

[0017] Perform segmentation processing on the power customer service voice data to obtain power customer service voice segmented data;

[0018] For each power customer service voice segmented data, input the power customer service voice segmented data into the target speech recognition model, so that the target speech recognition model performs speech recognition on the power customer service voice segmented data to generate a text statement in the target language of the corresponding segment;

[0019] Combine all the text statements in the target language of each segment to generate a complete text statement.

[0020] Further, the language corresponding to the power customer service voice data is confirmed by the following method:

[0021] Extract acoustic features from the power customer service voice data to obtain the first feature vector;

[0022] Select voice data in different languages from the preset language library, and perform acoustic feature extraction on the selected voice data in each language to obtain a second feature vector for each language;

[0023] For each language, calculate the cosine similarity between the second feature vector and the first feature vector;

[0024] Take the language corresponding to the maximum cosine similarity as the language corresponding to the power customer service voice data.

[0025] Further, before segmenting the power customer service voice data, it also includes:

[0026] Perform discretization processing on the power customer service voice data to obtain discretized power customer service voice data;

[0027] Pass the discretized power customer service voice data through a first-order high-pass filter for pre-emphasis processing of high-frequency signals to obtain preprocessed power customer service voice data;

[0028] Perform frequency domain conversion on the preprocessed power customer service voice data to obtain the spectrum information of the preprocessed power customer service voice data;

[0029] Determine the noise frequency according to the spectrum information of the preprocessed power customer service voice data;

[0030] Determine the passband frequency range of the band-pass filter according to the noise frequency;

[0031] Pass the preprocessed power customer service voice data through a band-pass filter for noise reduction processing to obtain enhanced power customer service voice data.

[0032] Further, segment the power customer service voice data to obtain power customer service voice segment data, including:

[0033] Calculate the zero-crossing rate of the enhanced power customer service voice data based on the data value at each sampling point of the enhanced power customer service voice data, and determine the corresponding zero-crossing position;

[0034] Compare the zero-crossing rate of the enhanced power customer service voice data with a preset segmentation threshold;

[0035] If the zero-crossing rate of the enhanced power customer service voice data is greater than or equal to the preset segmentation threshold, segment the enhanced power customer service voice data according to the zero-crossing position to obtain several power customer service voice segment data;

[0036] If the zero-crossing rate of the enhanced power customer service voice data is less than the preset segmentation threshold, directly use the enhanced power customer service voice data as the power customer service voice segment data.

[0037] Further, according to the data value of the enhanced power customer service voice data at each sampling point, the zero-crossing rate of the enhanced power customer service voice data is calculated, including:

[0038] According to the data value of the enhanced power customer service voice data at each sampling point, the zero-crossing rate of the enhanced power customer service voice data is calculated through the following formula:

[0039]

[0040] Where Z represents the zero-crossing rate of the enhanced power customer service voice data, N represents the length of the enhanced power customer service voice data, X(n) represents the data value of the enhanced power customer service voice data at the nth sampling point, and sgn(·) represents the sign function.

[0041] Based on the above method embodiment, the present invention correspondingly provides a system embodiment, including: a voice data acquisition module, a voice recognition module, a part-of-speech matching module, a part-of-speech combination module, an association analysis module, a first screening module, a second screening module, and a service processing module;

[0042] The voice data acquisition module is used to acquire the power customer service voice data to be processed;

[0043] The voice recognition module is used to perform voice recognition on the power customer service voice data to generate a text statement;

[0044] The part-of-speech matching module is used to match each word in the text statement with each preset vocabulary list, and use the part-of-speech of all successfully matched preset vocabulary lists as the candidate part-of-speech corresponding to the current word; among them, the preset vocabulary lists include: a noun vocabulary list, a verb vocabulary list, an adjective vocabulary list, an adverb vocabulary list, and a conjunction vocabulary list;

[0045] The part-of-speech combination module is used to generate several first candidate text statements under different part-of-speech combinations according to the candidate part-of-speech of each word;

[0046] The association analysis module is used to perform vocabulary association analysis on each word in each first candidate text statement according to the part-of-speech of each word in the first candidate text statement, and determine the context association degree of each first candidate text statement;

[0047] The first screening module is used to compare the context association degree of each first candidate text statement with a preset association degree threshold respectively, and use the first candidate text statement with a context association degree greater than the preset association degree threshold as the second candidate text statement;

[0048] The second screening module is used to use the second candidate text statement with the largest context association degree as the target text statement;

[0049] A service processing module, configured to perform power customer service processing according to each word and the corresponding part-of-speech combination in the target text statement.

[0050] Further, the speech recognition module includes: a speech model confirmation sub-module, a speech data segmentation sub-module, a segmented speech recognition sub-module, and a text statement combination sub-module;

[0051] The speech model confirmation sub-module is configured to select a speech recognition model corresponding to the language of the power customer service speech data from a preset language library as the target speech recognition model;

[0052] The speech data segmentation sub-module is configured to segment the power customer service speech data to obtain segmented power customer service speech data;

[0053] The segmented speech recognition sub-module is configured to input the segmented power customer service speech data into the target speech recognition model for each segmented power customer service speech data, so that the target speech recognition model performs speech recognition on the segmented power customer service speech data to generate a text statement in the target language corresponding to the segment;

[0054] The text statement combination sub-module is configured to combine all the text statements in the target language of the segments to generate a complete text statement.

[0055] Based on the above method item embodiments, the present invention correspondingly provides a terminal device item embodiment, including: a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the steps of the method for processing power customer service speech data as described in the present invention are implemented.

[0056] Based on the above method item embodiments, the present invention correspondingly provides a computer-readable storage medium item embodiment, including: a stored computer program. When the computer program runs, it controls the device where the computer-readable storage medium is located to execute the steps of the method for processing power customer service speech data as described in the present invention.

[0057] Compared with the prior art, the beneficial effects of the embodiments of the present solution are as follows:

[0058] The present invention obtains the power customer service voice data to be processed, performs speech recognition on the power customer service voice data to generate corresponding text sentences, and initially converts the voice into a processable text form. Then, since some words in the text sentences may have multiple parts of speech themselves, in order to accurately understand the meaning expressed by the customer, each word in the text sentences is respectively matched with a preset vocabulary list, and the parts of speech of all the successfully matched preset vocabulary lists are used as the candidate parts of speech corresponding to the current word. According to the candidate parts of speech of each word, several first candidate text sentences under different combinations of parts of speech are generated, listing the possibilities of different parts of speech. Among them, the preset vocabulary list includes vocabulary lists such as nouns, verbs, adjectives, and adverbs, and each vocabulary list details different words under the corresponding parts of speech, which helps to judge the meaning of each word in a specific context. Then, according to the parts of speech of each word in each first candidate text sentence, a lexical relevance analysis is performed on each word to determine the context relevance degree of each first candidate text sentence, that is, the logical relationship and semantic coherence of the words in the text sentence under different parts of speech. Finally, the context relevance degree of each first candidate text sentence is respectively compared with a preset relevance degree threshold, and the first candidate text sentences with a context relevance degree greater than the preset relevance degree threshold are used as second candidate text sentences, initially screening out the candidate text sentences with a relatively large relevance degree, and the second candidate text sentence with the largest context relevance degree is used as the target text sentence to ensure that the finally recognized text sentence can accurately reflect the true intention of the customer. Subsequently, according to the words and the corresponding part of speech combinations in the target text sentence, power customer service business processing is performed. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 is a schematic flowchart of a method for processing power customer service voice data provided by an embodiment of the present invention;

[0060] Figure 2 is a schematic structural diagram of a system for processing power customer service voice data provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0061] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without making creative efforts shall fall within the protection scope of the present invention.

[0062] In the description of the present invention, it should be understood that the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features.

[0063] Such as Figure 1As shown in the figure, an embodiment of the present invention provides a method for processing power customer service voice data, and the method at least includes the following steps:

[0064] Step S1: Obtain the power customer service voice data to be processed;

[0065] For step S1, in the power customer service system, when customers contact the power customer service to handle business or consult questions, they usually express their needs and questions through voice. In order to effectively understand and respond to customers, the power customer service system needs to first obtain the voice data of customers.

[0066] Step S2: Perform speech recognition on the power customer service voice data to generate text sentences;

[0067] In a preferred embodiment, performing speech recognition on the power customer service voice data to generate text sentences includes:

[0068] According to the language type corresponding to the power customer service voice data, select the speech recognition model of the corresponding language type from the preset language type library as the target speech recognition model;

[0069] Perform segmentation processing on the power customer service voice data to obtain power customer service voice segmented data;

[0070] For each power customer service voice segmented data, input the power customer service voice segmented data into the target speech recognition model, so that the target speech recognition model performs speech recognition on the power customer service voice segmented data to generate text sentences of the target language type corresponding to the segment;

[0071] Combine all the text sentences of the target language type of the segments to generate complete text sentences.

[0072] In a preferred embodiment, performing segmentation processing on the power customer service voice data to obtain power customer service voice segmented data includes:

[0073] According to the data value of the enhanced power customer service voice data at each sampling point, calculate the zero-crossing rate of the enhanced power customer service voice data and determine the corresponding zero-crossing position;

[0074] Compare the zero-crossing rate of the enhanced power customer service voice data with the preset segmentation threshold;

[0075] If the zero-crossing rate of the enhanced power customer service voice data is greater than or equal to the preset segmentation threshold, perform segmentation processing on the enhanced power customer service voice data according to the zero-crossing position to obtain several power customer service voice segmented data;

[0076] If the zero-crossing rate of the enhanced power customer service voice data is less than the preset segmentation threshold, directly use the enhanced power customer service voice data as the power customer service voice segmented data.

[0077] For step S2, the collected power customer service voice data is subjected to speech recognition and converted into text sentences. However, since customers may communicate in different languages, such as Mandarin, Cantonese, Hakka, or even languages of other countries like English, this increases the complexity and challenge of speech recognition.

[0078] To solve this problem, first, it is necessary to identify the language type in the power customer service voice data. This is because there are significant differences in the speech characteristics, vocabulary, and grammatical structures of different languages, which directly affect the accuracy and efficiency of recognition.

[0079] Preferably, the language type corresponding to the power customer service voice data is confirmed by the following method:

[0080] Extract the acoustic features of the power customer service voice data to obtain the first feature vector;

[0081] Select the voice data of different language types from the preset language type library, and extract the acoustic features of the selected voice data of each language type to obtain the second feature vector of each language type;

[0082] For each language type, calculate the cosine similarity between the second feature vector and the first feature vector;

[0083] Take the language type corresponding to the maximum cosine similarity as the language type corresponding to the power customer service voice data.

[0084] Specifically, extract the acoustic features of the power customer service voice data. Acoustic features refer to the physical attributes in the voice signal, such as pitch, intensity, duration, timbre, etc. These features can reflect the pronunciation method and language characteristics, and form the first feature vector with the obtained acoustic features. Next, select the voice data of different language types from the preset language type library. This language type library should contain voice data samples of multiple languages to cover all possible languages that may be encountered. Extract the acoustic features of the selected voice data of each language type to obtain the second feature vector of each language type. These vectors represent the acoustic features of the voices of each language type. For each language type, calculate the cosine similarity between its second feature vector and the first feature vector of the power customer service voice data. Cosine similarity is an index to measure the similarity of the directions of two vectors, and its value range is between -1 and 1. When the directions of the two vectors are exactly the same, the cosine similarity is 1; when the directions are exactly opposite, it is -1; when the two vectors are perpendicular, it is 0. In this embodiment, the closer the cosine similarity is to 1, the more similar the power customer service voice data is to the voice data of a certain language type in terms of acoustic features. Finally, take the language type corresponding to the maximum cosine similarity as the language type corresponding to the power customer service voice data. This means that the language type that is most similar to the acoustic features of the power customer service voice data is determined as the customer's communication language.

[0085] After determining the language corresponding to the power customer service voice data, according to the language corresponding to the power customer service voice data, a voice recognition model of the corresponding language is selected from a preset language library as the target voice recognition model. The preset language library also includes language recognition models of different languages, which are trained according to the voice characteristics, vocabulary, and grammar structures of different languages and can accurately recognize the voice information of the corresponding language.

[0086] Since there may be sentence breaks or pauses in the voice data of the customer, for better recognition, it is necessary to segment the pauses and analyze the segmented voice, which reduces the subsequent calculation amount. However, before that, in order to improve the recognition accuracy and efficiency, it is necessary to first perform enhancement processing on the voice data.

[0087] Preferably, before segmenting the power customer service voice data, it also includes:

[0088] Perform discretization processing on the power customer service voice data to obtain discretized power customer service voice data;

[0089] Perform pre-emphasis processing on the high-frequency signal of the discretized power customer service voice data through a first-order high-pass filter to obtain preprocessed power customer service voice data;

[0090] Perform frequency domain conversion on the preprocessed power customer service voice data to obtain the spectrum information of the preprocessed power customer service voice data;

[0091] Determine the noise frequency according to the spectrum information of the preprocessed power customer service voice data;

[0092] Determine the passband frequency range of the band-pass filter according to the noise frequency;

[0093] Perform denoising processing on the preprocessed power customer service voice data through a band-pass filter to obtain enhanced power customer service voice data.

[0094] Specifically, first, convert the continuous power customer service voice data into a discrete digital signal through the sampling theorem for processing and analysis on a computer. It should be noted that the sampling rate should be selected high enough to retain the important information in the voice signal.

[0095] It should be noted that the construction formula of the first-order high-pass filter is:

[0096] H(z) = 1 - az -1 ;

[0097] where H(z) represents the transfer function of the first-order high-pass filter, a represents the pre-emphasis coefficient, and z represents the input data.

[0098] Therefore, in this embodiment, the discretized power customer service voice data is pre-emphasized for high-frequency signals through the following calculation formula using a first-order high-pass filter to obtain the pre-processed power customer service voice data:

[0099] y(n) = x(n) - ax(n - 1);

[0100] where y(n) represents the pre-processed power customer service voice data at time n, x(n) represents the sampling value of the discretized power customer service voice data at time n, a represents the pre-emphasis coefficient, and x(n - 1) represents the sampling value of the discretized power customer service voice data at time n - 1.

[0101] It should be noted that pre-emphasizing the power customer service voice data is to enhance the high-frequency components of the voice signal, improve the spectral characteristics of the signal, and facilitate subsequent spectral analysis and feature extraction. In addition, during the transmission of the voice signal, the high-frequency components tend to be attenuated, and pre-emphasis can compensate for this attenuation, making the high-frequency components more prominent in subsequent processing.

[0102] Next, the fast Fourier transform (FFT) algorithm is usually used to transform the pre-processed voice data from the time domain to the frequency domain. After the transformation, the spectrogram of the voice signal can be obtained, which contains the intensity information of each frequency component. By analyzing the spectral information, the noise components in the voice signal and their frequency ranges are determined. These noise components appear as irregular fluctuations or random noise in the spectrogram. Subsequently, according to the determined noise frequency range, the passband frequency range of the band-pass filter is set to filter out these noise components while retaining the useful information in the voice signal. Finally, the pre-processed voice data is input into the band-pass filter, and the filter will filter out the noise components according to the set passband frequency range. After filtering, the obtained voice data will be clearer and the noise interference will be reduced.

[0103] After enhancement and denoising, the power customer service voice data is still a continuous signal. Therefore, in order to reduce the computational complexity of subsequent voice analysis, this continuous signal needs to be segmented. The key to segmentation lies in identifying the pauses in the voice because segmenting at the pauses can naturally divide different voice paragraphs, and each paragraph contains relatively complete semantic information.

[0104] Preferably, according to the data value of the enhanced power customer service voice data at each sampling point, the zero-crossing rate of the enhanced power customer service voice data is calculated through the following formula:

[0105]

[0106] Among them, Z represents the zero-crossing rate of the enhanced power customer service voice data, N represents the length of the enhanced power customer service voice data, X(n) represents the data value of the enhanced power customer service voice data at the nth sampling point, and sgn(·) represents the sign function.

[0107] Specifically, when a person is speaking and needs to pause or make a stop, their voice signal will change from high or low sound to silence. From the perspective of energy, it changes from high energy to zero energy, that is, the zero-crossing rate. The sign function is 1 when greater than zero and -1 when less than zero. Therefore, by calculating the sign function at the previous moment and the current moment, it is determined whether there is a pause during the customer's speech. Then, segmentation is performed at the pause, and the segmented speech is analyzed, reducing the subsequent computational amount.

[0108] To determine when to perform segmentation, a preset segmentation threshold is set. This threshold represents the standard for judging whether a pause is significant. When the zero-crossing rate of the enhanced power customer service voice data is greater than or equal to this threshold, it is considered that a significant pause has occurred, and segmentation processing should be performed at this time. The segmentation position is the position corresponding to the peak of the zero-crossing rate.

[0109] After the segmentation processing is completed, several pieces of segmented power customer service voice data are obtained. Next, these segmented data are input into the target speech recognition model selected previously one by one. This model has been trained to accurately recognize speech in a specific target language and convert it into corresponding text sentences. After the target speech recognition model processes each piece of segmented data, it will generate the text sentences in the target language for the corresponding segment. To construct a complete text sentence, these text sentences generated from the segments are combined in an orderly manner, ensuring that all information extracted from the original voice data is accurately and completely converted into text form.

[0110] Step S3: For each word in the text sentence, match the current word with each preset vocabulary list, and use the parts of speech of all the preset vocabulary lists that match successfully as the candidate parts of speech corresponding to the current word; among them, the preset vocabulary lists include: noun vocabulary list, verb vocabulary list, adjective vocabulary list, adverb vocabulary list, and conjunction vocabulary list;

[0111] For step S3, the text sentence obtained by step S2 is composed of a series of words. For each word in the text sentence, it is matched with each preset vocabulary list. In this embodiment, the preset vocabulary list includes a noun vocabulary list, a verb vocabulary list, an adjective vocabulary list, an adverb vocabulary list and a conjunction vocabulary list. The noun vocabulary list contains all words that are considered nouns, the verb vocabulary list contains all words that are considered verbs, the adjective vocabulary list contains all words that are considered adjectives, and the adverb vocabulary list contains all words that are considered adverbs. The matching process is completed by searching the vocabulary list, that is, checking whether the current word exists in a certain vocabulary list. If a word finds a match in a certain preset vocabulary list, then the part of speech of the vocabulary list is used as a candidate part of speech for the current word. It should be noted that a word may have matches in multiple vocabulary lists, so it may have multiple candidate parts of speech. For example: a customer issues an instruction "Is the equipment grounded well?", in which the vocabulary includes: equipment, grounding, whether, good; after matching, the candidate parts of speech of each vocabulary can be determined: equipment (noun and adjective), grounding (verb and noun), whether (conjunction), good (adjective), the word "equipment" when used as a noun means a device or a machine, and at the same time, "equipment" can also be used as an adjective, omitting the word "of", indicating a noun phrase of the equipment; the word "grounding" when used as a verb means the grounding action, and at the same time, "grounding" can also be used as a noun to indicate the grounding terminal.

[0112] Step S4: generating a plurality of first candidate text sentences under different part-of-speech combinations according to the candidate parts-of-speech of each word;

[0113] For step S4, different part-of-speech combinations can be generated according to the candidate parts-of-speech of each vocabulary word. For example, following the above example, "equipment" can be a noun or an adjective, and "grounding" can be a verb or a noun, so four first candidate text sentences of part-of-speech combinations can be generated, the first type: equipment (noun), grounding (verb), whether (conjunction), good (adjective); the second type: equipment (noun), grounding (noun), whether (conjunction), good (adjective); the third type: equipment (adjective), grounding (verb), whether (conjunction), good (adjective); the fourth type: equipment (adjective), grounding (noun), whether (conjunction), good (adjective).

[0114] Step S5: for each first candidate text sentence, perform a word relevance analysis on each word according to the part of speech of each word in the first candidate text sentence to determine the context relevance of each first candidate text sentence;

[0115] For step S5, for the first candidate text statement generated in step S4, lexical relevance analysis is performed one by one through the trained natural language model. The natural language model has learned the lexical relevance, grammar rules, and semantic information in a large amount of text data, and can understand the meaning of words in a specific context and how words are related to each other. Therefore, the natural language model can perform lexical relevance analysis on each word according to the part of speech of each word in the first candidate text statement, and output the context relevance degree of each first candidate text statement. In this embodiment, the context relevance degree is in the form of a score, which reflects the matching degree of the statement with the given context or instruction. The higher the context relevance degree, the more the statement meets the requirements of the context and the more accurately it can convey the intention of the instruction.

[0116] Continuing with the above example, assume that the context relevance degree of the first candidate text statement under the first part-of-speech combination is 8, the context relevance degree of the first candidate text statement under the second part-of-speech combination is 4, the context relevance degree of the first candidate text statement under the third part-of-speech combination is 2, and the context relevance degree of the first candidate text statement under the fourth part-of-speech combination is 10.

[0117] Step S6: Compare the context relevance degree of each first candidate text statement with a preset relevance threshold respectively, and take the first candidate text statements with a context relevance degree greater than the preset relevance threshold as the second candidate text statements;

[0118] For step S6, compare the context relevance degree of each first candidate text statement with a preset relevance threshold respectively. The preset relevance threshold is used to measure whether the matching degree of the statement with the given context or instruction is high enough. All the first candidate text statements that meet the requirements of the preset relevance threshold are preliminarily screened and taken as the second candidate text statements.

[0119] Step S7: Take the second candidate text statement with the largest context relevance degree as the target text statement;

[0120] For step S7, there may still be multiple part-of-speech combinations among the second candidate text statements screened by step S6. That is, when the preset relevance threshold is 7, two groups of first candidate text statements under the four part-of-speech combinations in the example are selected as the second candidate text statements. Therefore, in order to select the text statement under the part-of-speech combination that is closest to the instruction meaning or has the best context consistency, take the second candidate text statement with the largest context relevance degree as the target text statement. Corresponding to the example, take the first candidate text statement under the first part-of-speech combination with the largest context relevance degree as the target text statement, that is, device (adjective), grounded (noun), whether (conjunction), good (adjective).

[0121] Step S8: Perform power customer service business processing based on each word and its corresponding part-of-speech combination in the target text statement.

[0122] For Step S8, based on each word and its corresponding part-of-speech combination in the target text statement parsed in Step S7, match through a preset power customer service business logic and perform corresponding business operations, such as querying the database, calling the API interface to transfer to manual service, or generating a reply, etc.

[0123] The present invention performs language identification on power customer service voice data, thereby identifying text statements according to different languages, being able to identify and process multiple languages, effectively breaking the language communication barrier, and greatly expanding the service scope of power customer service. On this basis, match the part of speech of each word in the text statement with a preset vocabulary list to obtain text statements under different part-of-speech combinations. Finally, use a natural language processing model to evaluate the context relevance of these combinations and screen out the text statement under the most logically coherent and semantically appropriate set of part-of-speech combinations. This process not only ensures the accuracy and rationality of the statement but also can accurately identify and understand the true intention of the customer, thereby providing more personalized and intelligent services for the customer. Ultimately, the customer can smoothly and efficiently handle the required business, and their service experience has been significantly improved.

[0124] As Figure 2 shown, based on the above method item embodiments, corresponding system item embodiments are provided;

[0125] An embodiment of the present invention provides a processing system for power customer service voice data, including: a voice data acquisition module, a speech recognition module, a part-of-speech matching module, a part-of-speech combination module, a correlation analysis module, a first screening module, a second screening module, and a business processing module;

[0126] The voice data acquisition module is used to acquire power customer service voice data to be processed;

[0127] The speech recognition module is used to perform speech recognition on the power customer service voice data to generate a text statement;

[0128] The part-of-speech matching module is used for each word in the text statement to match the current word with each preset vocabulary list and use the parts of speech of all successfully matched preset vocabulary lists as the candidate parts of speech corresponding to the current word; among them, the preset vocabulary lists include: a noun vocabulary list, a verb vocabulary list, an adjective vocabulary list, an adverb vocabulary list, and a conjunction vocabulary list;

[0129] The part-of-speech combination module is used to generate several first candidate text statements under different part-of-speech combinations according to the candidate parts of speech of each word;

[0130] An association analysis module is used to perform lexical association analysis on each vocabulary in the first candidate text statement according to the part of speech of each vocabulary in the first candidate text statement, and determine the context correlation degree of each first candidate text statement;

[0131] A first screening module is used to compare the context correlation degree of each first candidate text statement with a preset correlation degree threshold respectively, and use the first candidate text statement with a context correlation degree greater than the preset correlation degree threshold as the second candidate text statement;

[0132] A second screening module is used to use the second candidate text statement with the largest context correlation degree as the target text statement;

[0133] A business processing module is used to perform power customer service business processing according to each vocabulary and the corresponding part of speech combination in the target text statement.

[0134] In a preferred embodiment, the speech recognition module includes: a speech model confirmation sub-module, a speech data segmentation sub-module, a segmented speech recognition sub-module, and a text statement combination sub-module;

[0135] The speech model confirmation sub-module is used to select a speech recognition model corresponding to the language of the power customer service speech data from a preset language library as the target speech recognition model;

[0136] The speech data segmentation sub-module is used to perform segmentation processing on the power customer service speech data to obtain power customer service segmented speech data;

[0137] The segmented speech recognition sub-module is used to input the power customer service segmented speech data into the target speech recognition model for each power customer service segmented speech data, so that the target speech recognition model performs speech recognition on the power customer service segmented speech data and generates a text statement in the target language corresponding to the segment;

[0138] The text statement combination sub-module is used to combine all the text statements in the target language of the segments to generate a complete text statement.

[0139] It can be understood that the above system item embodiments correspond to the method item embodiments of the present invention, and can implement the processing method of the power customer service speech data provided by any one of the above method item embodiments of the present invention.

[0140] It should be noted that the system embodiments described above are merely illustrative, and some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the system embodiments provided by the present invention, the connection relationships between modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines. Those of ordinary skill in the art can understand and implement this without creative efforts.

[0141] Based on the above embodiments of the method for processing power customer service voice data, another embodiment of the present invention provides a terminal device, which includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the method for processing power customer service voice data according to any embodiment of the present invention is implemented.

[0142] Exemplarily, in this embodiment, the computer program can be divided into one or more modules. The one or more modules are stored in the memory and executed by the processor to complete the present invention. The one or more module elements can be a series of computer program instruction segments capable of completing specific functions, and the instruction segments are used to describe the execution process of the computer program in the terminal device.

[0143] The terminal device can be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The terminal device may include, but is not limited to, a processor and a memory.

[0144] The so-called processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the terminal device, and connects various parts of the entire terminal device through various interfaces and lines.

[0145] Based on the above method embodiment, another embodiment is provided: A computer-readable storage medium provided by another embodiment of the present invention includes a stored computer program, wherein when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the processing method of the power customer service voice data described in any one of the above method embodiments of the present invention.

[0146] Among them, the module / unit integrated in the processing system / terminal device of the power customer service voice data, if implemented in the form of a software functional unit and sold or used as an independent product, can be stored in a computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above embodiment methods of the present invention, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or system capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc.

[0147] The above is the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present invention.

Claims

1. A method for processing power customer service voice data, characterized in that Including: Obtain the power customer service voice data to be processed; Perform speech recognition on the power customer service voice data to generate text sentences; For each word in the text sentence, match the current word with each preset vocabulary list, and use the part-of-speech of all successfully matched preset vocabulary lists as the candidate part-of-speech corresponding to the current word; among them, the preset vocabulary lists include: noun vocabulary list, verb vocabulary list, adjective vocabulary list, adverb vocabulary list, and conjunction vocabulary list; Generate several first candidate text sentences under different part-of-speech combinations according to the candidate part-of-speech of each word; For each first candidate text sentence, perform lexical relevance analysis on each word according to the part-of-speech of the words in the first candidate text sentence to determine the context relevance degree of each first candidate text sentence; Compare the context relevance degree of each first candidate text sentence with the preset relevance degree threshold respectively, and use the first candidate text sentence with a context relevance degree greater than the preset relevance degree threshold as the second candidate text sentence; Use the second candidate text sentence with the largest context relevance degree as the target text sentence; Perform power customer service business processing according to the words and corresponding part-of-speech combinations in the target text sentence.

2. The processing method of power customer service voice data according to claim 1, wherein, Performing speech recognition on the power customer service voice data to generate text sentences includes: Select a speech recognition model of the corresponding language from the preset language library according to the language corresponding to the power customer service voice data as the target speech recognition model; Perform segmentation processing on the power customer service voice data to obtain segmented power customer service voice data; For each segmented power customer service voice data, input the segmented power customer service voice data into the target speech recognition model, so that the target speech recognition model performs speech recognition on the segmented power customer service voice data to generate text sentences in the target language of the corresponding segment; Combine the text sentences in the target language of all segments to generate a complete text sentence.

3. The method for processing power customer service voice data according to claim 2, wherein The language corresponding to the power customer service voice data is confirmed by the following method: Extract acoustic features from the power customer service voice data to obtain a first feature vector; Select speech data in different languages from the preset language library, and extract acoustic features from the selected speech data in each language to obtain a second feature vector for each language; For each language, calculate the cosine similarity between the second feature vector and the first feature vector; Use the language corresponding to the maximum cosine similarity as the language corresponding to the power customer service voice data.

4. The processing method of power customer service voice data according to claim 2, characterized in that, Before performing segmentation processing on the power customer service voice data, it also includes: Perform discretization processing on the power customer service voice data to obtain discretized power customer service voice data; Perform pre-emphasis processing on the high-frequency signal of the discretized power customer service voice data through a first-order high-pass filter to obtain preprocessed power customer service voice data; Perform frequency domain conversion on the preprocessed power customer service voice data to obtain the frequency spectrum information of the preprocessed power customer service voice data; Determine the noise frequency according to the frequency spectrum information of the preprocessed power customer service voice data; Determine the passband frequency range of the band-pass filter according to the noise frequency; The preprocessed power customer service voice data is denoised by the band-pass filter to obtain enhanced power customer service voice data.

5. The processing method of power customer service voice data according to claim 4, characterized in that, The power customer service voice data is segmented to obtain power customer service voice segmented data, including: According to the data value of the enhanced power customer service voice data at each sampling point, the zero-crossing rate of the enhanced power customer service voice data is calculated, and the corresponding zero-crossing position is determined; The zero-crossing rate of the enhanced power customer service voice data is compared with a preset segmentation threshold; If the zero-crossing rate of the enhanced power customer service voice data is greater than or equal to the preset segmentation threshold, the enhanced power customer service voice data is segmented according to the zero-crossing position to obtain several power customer service voice segmented data; If the zero-crossing rate of the enhanced power customer service voice data is less than the preset segmentation threshold, the enhanced power customer service voice data is directly used as the power customer service voice segmented data.

6. The method for processing power customer service voice data according to claim 5, wherein According to the data value of the enhanced power customer service voice data at each sampling point, the zero-crossing rate of the enhanced power customer service voice data is calculated, including: According to the data value of the enhanced power customer service voice data at each sampling point, the zero-crossing rate of the enhanced power customer service voice data is calculated by the following formula: where Z represents the zero-crossing rate of the enhanced power customer service voice data, N represents the length of the enhanced power customer service voice data, X(n) represents the data value of the enhanced power customer service voice data at the nth sampling point, and sgn(·) represents the sign function.

7. A processing system for power customer service voice data, characterized in that, Including: A voice data acquisition module, a speech recognition module, a part-of-speech matching module, a part-of-speech combination module, a correlation analysis module, a first screening module, a second screening module, and a service processing module; The voice data acquisition module is used to acquire the power customer service voice data to be processed; The speech recognition module is used to perform speech recognition on the power customer service voice data to generate a text statement; The part-of-speech matching module is used to match each word in the text statement with each preset vocabulary list, and use the part-of-speech of all successfully matched preset vocabulary lists as the candidate part-of-speech corresponding to the current word; where the preset vocabulary lists include: a noun vocabulary list, a verb vocabulary list, an adjective vocabulary list, an adverb vocabulary list, and a conjunction vocabulary list; The part-of-speech combination module is used to generate several first candidate text statements under different part-of-speech combinations according to the candidate part-of-speech of each word; The correlation analysis module is used to perform vocabulary correlation analysis on each word in each first candidate text statement according to the part-of-speech of each word in the first candidate text statement, and determine the context correlation degree of each first candidate text statement; The first screening module is used to compare the context correlation degree of each first candidate text statement with a preset correlation degree threshold respectively, and use the first candidate text statement with a context correlation degree greater than the preset correlation degree threshold as the second candidate text statement; The second screening module is used to use the second candidate text statement with the largest context correlation degree as the target text statement; The business processing module is used to perform power customer service business processing according to each word and the corresponding part-of-speech combination in the target text statement.

8. The processing system for power customer service voice data according to claim 7, wherein The speech recognition module includes: a speech model confirmation sub-module, a speech data segmentation sub-module, a segmented speech recognition sub-module, and a text statement combination sub-module; The speech model confirmation sub-module is used to select a speech recognition model of the corresponding language from a preset language library according to the language corresponding to the power customer service speech data as the target speech recognition model; The speech data segmentation sub-module is used to perform segmentation processing on the power customer service speech data to obtain segmented power customer service speech data; The segmented speech recognition sub-module is used to input the segmented power customer service speech data into the target speech recognition model for each segmented power customer service speech data, so that the target speech recognition model performs speech recognition on the segmented power customer service speech data to generate a text statement in the target language of the corresponding segment; The text statement combination sub-module is used to combine all the text statements in the target language of the segments to generate a complete text statement.

9. A terminal device, characterized in that, Including: A processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the processing method of the power customer service speech data as described in any one of claims 1-6 is implemented.

10. A computer-readable storage medium, characterized in that, Including: A stored computer program, wherein when the computer program runs, it controls the device where the computer-readable storage medium is located to execute the processing method of the power customer service speech data as described in any one of claims 1-6.