A regional power intelligent customer service system supporting dialect recognition
By constructing a regional intelligent customer service system for the power industry, the problem of the system's inability to adapt to regional dialects was solved, achieving efficient dialect recognition and business processing, and improving user service experience and system efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHUHAI POWER SUPPLY BUREAU GUANGDONG POWER GIRD CO
- Filing Date
- 2026-03-10
- Publication Date
- 2026-06-09
AI Technical Summary
Existing intelligent customer service systems for the power industry cannot effectively adapt to regional dialects, leading to communication barriers for middle-aged and elderly users and rural users. They also suffer from low recognition accuracy and high manual transfer rates, failing to meet the requirement of full coverage for power business scenarios.
A regional intelligent power customer service system supporting dialect recognition is constructed, including a multimodal interactive access layer, a regional dialect recognition engine, a power business dialect corpus, a hierarchical intent understanding module, and a dynamic optimization module. Through preprocessing, dialect feature extraction, corpus mapping, and dynamic iterative optimization, the system realizes the conversion of dialect text to Mandarin and business processing.
It improved the accuracy of dialect recognition from 68% to 92%, reduced the manual transfer rate to below 20%, improved the user service experience and system operation efficiency, adapted to the rapid switching of multi-regional dialect models, and reduced promotion costs.
Smart Images

Figure CN122173625A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power customer service technology, and more specifically, to a regional intelligent power customer service system that supports dialect recognition. Background Technology
[0002] In the field of public power services, intelligent customer service systems have become a core infrastructure for improving service efficiency and reducing labor costs. Currently, most mainstream intelligent customer service systems for the power industry are built on Mandarin speech recognition technology, assuming that users will interact in standard Mandarin. While this can provide efficient service in urban areas where Mandarin is widely spoken, it has significant limitations in adaptability in rural areas, county-level cities, and multi-ethnic areas where dialects are prevalent.
[0003] On the one hand, regional dialect users face widespread communication barriers. my country's dialect system is complex, encompassing ten major dialect regions within the Chinese language alone, including Mandarin, Wu, Cantonese, and Min. Multiple dialect branches also exist within some counties. Many middle-aged and elderly users, as well as rural residents left behind in rural areas, rely on their dialects for daily communication and struggle to use standard Mandarin to interact with intelligent customer service. This leads to intelligent customer service failing to accurately understand user requests, forcing users to repeatedly adjust their expressions or be transferred to human agents. Data shows that in some dialect-heavy areas, the human transfer rate for power company intelligent customer service exceeds 60%, increasing the workload of human agents and reducing the user experience.
[0004] On the other hand, existing general dialect recognition solutions are not suitable for power business scenarios. Most common dialect recognition models on the market focus on everyday spoken language interactions and are not optimized for dialectal expressions of power-related technical terms. The power business has a large number of proprietary terms, such as tripping, line loss, and peak-valley electricity pricing, whose dialectal expressions have strong scenario-specific characteristics. For example, in some regions, tripping is described as "turn-off power failure," which general models easily misinterpret as everyday speech, leading to biased intent recognition. Furthermore, general dialect models have limited coverage, supporting only a few mainstream dialects. Their accuracy in recognizing less common regional dialects such as Xiang, Gan, and Hakka is less than 70%, failing to meet the requirement of comprehensive coverage for power services.
[0005] Furthermore, existing systems lack regional dynamic optimization capabilities. Traditional intelligent customer service systems often use general pre-trained models for recognition, making it difficult to iterate based on actual voice data from users in different regions after deployment. This results in an inability to adapt to changes in regional dialect pronunciation, the evolution of new vocabulary, and other factors. As users' language habits change within a region, the system's recognition accuracy gradually declines, making it difficult to maintain service quality in the long term.
[0006] Existing intelligent customer service systems for the power industry cannot effectively address the service adaptation issues for users with regional dialects, thus hindering the equitable coverage of public power services. There is a need to construct an intelligent customer service system optimized for regional dialects and adapted to power business scenarios. Therefore, this paper proposes a regional intelligent customer service system for the power industry that supports dialect recognition. Summary of the Invention
[0007] The purpose of this invention is to address the problems raised in the existing background technology. To achieve the above-mentioned objective, this invention provides the following technical solution: a regional intelligent power customer service system supporting dialect recognition, comprising a multimodal interactive access layer for receiving power service requests initiated by users via telephone voice, online voice, and text input, and performing noise reduction, speech rate correction, and dialect audio feature enhancement preprocessing on voice requests; The regional dialect recognition engine is based on a pre-trained general speech model. It combines the dialect data of the target region with fine-tuning to build a dialect-specific recognition model. The dialect acoustic feature extraction module captures the differences in tone, initial consonant and final vowel variants of regional dialects and converts the pre-processed speech into the corresponding dialect text. The power business dialect corpus stores dialect expressions of power business scenarios within the target area, and establishes a three-dimensional mapping relationship between dialect expressions, common power terms, and business scenarios, covering mainstream power business scenarios such as fault reporting, business consultation, and payment inquiry. The hierarchical intent understanding module first converts the recognized dialect text into standard Mandarin text, and then uses a two-level architecture of basic intent matching and subdivided intent recognition, combined with a power business dialect corpus, to match the user's specific business needs. The business processing and response optimization layer calls the power back-end business interface to obtain the corresponding business data based on the intent recognition results, and optimizes and adapts the voice or text response content to the dialect of the target area. The dynamic optimization module collects user interaction voice data and recognition results, which are then manually verified and added to the speech corpus of the power business. The speech recognition model is also fine-tuned and iterated regularly based on the new speech data. The speech rate correction formula is as follows: in, To correct the speech rate, Original speech speed The average pronunciation duration of a single word in the target region's dialect is the baseline value. The average pronunciation duration of a single word in the original speech.
[0008] The formula for extracting dialect tone features is as follows: in, For the first The feature weights of each tone. For the first The first dialect data The actual fundamental frequency value of each tone. For the target region dialect The average fundamental frequency of each tone For the target region dialect The standard deviation of the fundamental frequency of each tone The number of dialect corpora involved in feature calculation.
[0009] As a preferred technical solution of the present invention, the dialect acoustic feature extraction module uses Mel-frequency cepstral coefficients (MFCC) combined with regional dialect tone features for modeling, and constructs independent classifiers for special tone variants of the target region dialect to enhance the recognition accuracy of dialect pronunciation differences: in, To determine the tone matching degree between the input speech and the dialect of the target region, For the first in the input voice The feature weights of each tone In the regional dialect model, the first Standard feature weights for each tone This represents the total number of tones in the dialect of the target region.
[0010] As a preferred technical solution of the present invention, the power business dialect corpus also includes dialect abnormal expression correction rules, which automatically map colloquial omissions, inversions and non-standard dialect expressions of users in the region to standardized power business terms.
[0011] As a preferred technical solution of the present invention, the hierarchical intent understanding module includes an associated scenario matching unit, which determines the priority of the core request by combining the contextual semantics when the user's expression contains multiple business keywords. in, The intent weight for keyword k. Keywords Word frequency in the input text, This represents the total number of power business scenarios. For keywords The number of business scenarios.
[0012] As a preferred technical solution of the present invention, the business processing and response optimization layer has a built-in dialect speech synthesis engine, which supports converting the general Mandarin business results into target regional dialect speech output. During the synthesis process, the speech rate and tone characteristics of the regional dialect are matched to ensure the naturalness of the response speech.
[0013] As a preferred technical solution of the present invention, the dynamic optimization module sets up a user feedback interface, providing two quick evaluation entry points: accurate recognition and incorrect recognition. Voice data marked as incorrectly recognized is automatically pushed to a manual verification queue, and corpora with validity exceeding a threshold are automatically added to the corpus for model iteration. in, For the first The effectiveness of the corpus The number of power business keywords contained in the corpus. The total number of words in the corpus. Rate the accuracy of user feedback, with 1 for accuracy and 0 for error.
[0014] As a preferred technical solution of the present invention, the regional dialect recognition engine supports multi-region model switching, and presets independent dialect recognition model packages for different target regions, which can automatically match the recognition model of the corresponding region according to the user's caller location and IP address.
[0015] As a preferred technical solution of the present invention, the multimodal interactive access layer also supports dialect text input adaptation. When the user inputs a dialect expression through text, the power business dialect corpus is automatically called to complete the conversion of dialect text into general business terms without going through the speech recognition process.
[0016] As a preferred technical solution of the present invention, the system further includes a dialect recognition accuracy monitoring unit, which statistically analyzes the dialect recognition accuracy in different regions and business scenarios in real time, optimizes the visualization report, and automatically triggers an emergency model fine-tuning process when the recognition accuracy in a certain region is continuously lower than a preset threshold. in, The overall accuracy of the regional model. For the first The recognition accuracy for each business scenario For the first The percentage of user requests per business scenario.
[0017] As a preferred technical solution of the present invention, the power business dialect corpus supports linkage with the power customer service human agent system. When the human agent handles the dialect user's request, the dialect dialogue between the agent and the user is automatically recorded and added to the corpus after being desensitized, so as to realize the continuous accumulation of corpus data.
[0018] Compared with existing technologies, the beneficial effects of this invention are as follows: This system effectively covers niche dialect areas by constructing a targeted regional dialect recognition engine, solving the communication difficulties between dialect-speaking groups such as the elderly and rural users and intelligent customer service. Actual testing shows that the accuracy rate of dialect recognition in the target area has increased from 68% in the general model to over 92%, and the manual transfer rate has been reduced to less than 20%. This allows dialect-speaking users to efficiently complete operations such as power business consultation and fault reporting without needing to switch to Mandarin, breaking down language barriers and facilitating the extension of public power services to grassroots areas.
[0019] By leveraging a dedicated corpus of dialects used in the power industry, the system accurately binds dialect expressions with professional power terminology, resolving the issue of misrecognition of business terms by general dialect models. For example, for the dialect expression "meters are running fast in the region," the system can directly match it to a meter reading anomaly scenario, avoiding misinterpretation as everyday speech, thus improving intent matching accuracy to over 95%. Simultaneously, by optimizing responses to power business processes and using dialect phrases familiar to regional users, the system ensures users clearly understand service outcomes and reduces communication costs.
[0020] The system constructs a data closed-loop optimization mechanism, continuously iterating the model and corpus based on real user interaction data from the region. As usage time increases, the system automatically adapts to changes in local dialect pronunciation habits and the evolution of dialect expressions for new business terms, with recognition accuracy gradually improving over time. For example, after six months of operation in a certain county, the system automatically learned the dialect expressions for newly added photovoltaic grid connection inquiries, achieving accurate recognition without requiring manual model retraining.
[0021] The system effectively reduces the transfer pressure on human agents, increasing the proportion of dialect-based business requests that can be handled by the intelligent customer service from 30% to over 80%, and significantly reducing repetitive communication work for human customer service representatives. Simultaneously, it automatically handles dialect-based user requests through standardized business processes, such as automatically optimizing fault repair work orders and pushing power outage notifications, shortening business processing time. The average service time per customer is reduced from 5 minutes to less than 1.5 minutes, improving overall customer service operational efficiency.
[0022] The system supports rapid switching between multi-regional dialect models. For different counties and cities, local dialect data can be directly imported and the model fine-tuned without rebuilding the underlying framework. A power grid company has already deployed the system in three dialect-dense areas, with each region's model building cycle taking only two weeks. This significantly reduces the cost of regional rollout and provides a replicable solution for dialect adaptation in intelligent power customer service nationwide.
[0023] By adopting service methods tailored to users' language habits, the service satisfaction score for users speaking local dialects increased from 58 to 91, effectively resolving complaints caused by communication difficulties. For example, a user in a rural area reported: "Before, it would take forever to explain things clearly when calling customer service, but now I can report the fault directly in my local dialect. It's so convenient!" This significantly improved the user reputation of electricity services and the image of public services. Attached Figure Description
[0024] Figure 1 A hierarchical fine-tuning data diagram for the regional dialect recognition model provided by this invention; Figure 2 This is a schematic diagram of the iterative process of the dynamic optimization module provided by the present invention; Figure 3 This is a data flowchart of the multi-region model switching adaptation rules provided by the present invention. Detailed Implementation
[0025] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention.
[0026] Therefore, the following detailed description of the embodiments of the present invention is not intended to limit the scope of the claimed invention, but merely illustrates some embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without inventive effort are within the scope of protection of the present invention. It should be noted that, in the absence of conflict, the embodiments and features and technical solutions in the embodiments of the present invention can be combined with each other. It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0027] Example 1: A regional intelligent power customer service system supporting dialect recognition, including a multimodal interactive access layer, used to receive power service requests initiated by users through telephone voice, online voice, and text input, and to perform noise reduction, speech rate correction, and dialect audio feature enhancement preprocessing on voice requests; The regional dialect recognition engine is based on a pre-trained general speech model. It combines the dialect data of the target region with fine-tuning to build a dialect-specific recognition model. The dialect acoustic feature extraction module captures the differences in tone, initial consonant and final vowel variants of regional dialects and converts the pre-processed speech into the corresponding dialect text. The power business dialect corpus stores dialect expressions of power business scenarios within the target area, and establishes a three-dimensional mapping relationship between dialect expressions, common power terms, and business scenarios, covering mainstream power business scenarios such as fault reporting, business consultation, and payment inquiry. The hierarchical intent understanding module first converts the recognized dialect text into standard Mandarin text, and then uses a two-level architecture of basic intent matching and subdivided intent recognition, combined with a power business dialect corpus, to match the user's specific business needs. The business processing and response optimization layer calls the power back-end business interface to obtain the corresponding business data based on the intent recognition results, and optimizes and adapts the voice or text response content to the dialect of the target area. The dynamic optimization module collects user interaction voice data and recognition results, which are then manually verified and added to the speech corpus of the power business. The speech recognition model is also fine-tuned and iterated regularly based on the new speech data. The speech rate correction formula is: in, To correct the speech rate, Original speech speed The average pronunciation duration of a single word in the target region's dialect is the baseline value. The average pronunciation duration of a single word in the original speech.
[0028] The formula for extracting dialect tone features is: in, For the first The feature weights of each tone For the first The first dialect data The actual fundamental frequency value of each tone. For the target region dialect The average fundamental frequency of each tone For the target region dialect The standard deviation of the fundamental frequency of each tone The number of dialect corpora involved in feature calculation.
[0029] The dialect acoustic feature extraction module uses Mel-frequency cepstral coefficients (MFCC) combined with regional dialect tone features for modeling. It constructs independent classifiers for specific tone variants of the target region's dialects, enhancing the accuracy of dialect pronunciation difference recognition. in, To determine the tone matching degree between the input speech and the dialect of the target region, For the first in the input voice The feature weights of each tone In the regional dialect model, the first The standard feature weights of each tone, M is the total number of tones in the target region dialect.
[0030] The power business dialect corpus also includes dialect abnormal expression correction rules, which automatically map colloquial omissions, inversions and non-standard dialect expressions of users in the region to standardized power business terms.
[0031] The hierarchical intent understanding module includes a context matching unit. When a user's statement contains multiple business keywords, it combines the semantic context to determine the priority of the core request. in, The intent weight for keyword k. Keywords Word frequency in the input text, This represents the total number of power business scenarios. For keywords The number of business scenarios.
[0032] The business processing and response optimization layer has a built-in dialect speech synthesis engine that supports converting general Mandarin business results into target regional dialect speech output. During the synthesis process, it matches the speech rate and tone characteristics of the regional dialect to ensure the naturalness of the response speech.
[0033] The dynamic optimization module includes a user feedback interface, providing quick evaluation entry points for both accurate and incorrect recognition. Voice data marked as incorrectly recognized is automatically pushed to a manual verification queue, while corpora with validity scores exceeding a threshold are automatically added to the corpus for model iteration. in, For the first The effectiveness of the corpus The number of power business keywords contained in the corpus. The total number of words in the corpus. Rate the accuracy of user feedback, with 1 for accuracy and 0 for error.
[0034] The regional dialect recognition engine supports switching between multiple regional models and presets independent dialect recognition model packages for different target regions. It can automatically match the corresponding regional recognition model based on the user's caller location and IP address.
[0035] The multimodal interaction access layer also supports dialect text input adaptation. When a user inputs a dialect expression, the system automatically calls the power business dialect corpus to convert the dialect text into general business terms without going through a speech recognition process.
[0036] The system also includes a dialect recognition accuracy monitoring unit, which provides real-time statistics on dialect recognition accuracy in different regions and business scenarios, and optimizes visualization reports. When the recognition accuracy in a certain region continuously falls below a preset threshold, it automatically triggers an emergency model fine-tuning process. in, The overall accuracy of the regional model. For the first The recognition accuracy for each business scenario For the first The percentage of user requests per business scenario.
[0037] The dialect corpus for power business supports linkage with the power customer service human agent system. When human agents handle requests from users with dialects, the dialect dialogue between the agent and the user is automatically recorded, and after being anonymized, it is added to the corpus to achieve continuous accumulation of corpus data.
[0038] The working principle of a regional intelligent power customer service system supporting dialect recognition: This system addresses the communication needs of users speaking regional dialects regarding power business. Through a closed-loop process of voice access, dialect recognition, intent matching, business processing, and iterative optimization, it achieves accurate adaptation between dialects and intelligent power customer service. The specific working principle is as follows: Multimodal access and preprocessing: The system receives user requests through a voice gateway and online interactive interface, and performs layered preprocessing for voice interaction: Noise reduction: Through an adaptive noise suppression algorithm, background noise, current noise and other interference in the environment are filtered out, while retaining the core signal of the user's dialect voice.
[0039] Dialect feature enhancement: Based on the pronunciation characteristics of the target region's dialect, highlighting key acoustic features such as tone and stress, weakening the default weight of Mandarin pronunciation in general speech recognition, and reducing the impact of dialect pronunciation deviation on recognition.
[0040] Multi-format adaptation: Automatically converts audio formats with different sampling rates, such as telephone voice and online voice, into a unified standard format to ensure compatibility in subsequent recognition processes.
[0041] Regional dialect recognition engine working logic: The dialect recognition engine is based on a general pre-trained speech model and combines it with the target region dialect data for fine-tuning to build localized recognition capabilities: Dialect feature matching: The engine has a built-in target region dialect acoustic feature library. By extracting the tone curve and pronunciation characteristics of the initials and finals of the input speech, it matches them with typical dialect features in the feature library, prioritizing the activation of the regional dialect recognition path rather than the general Mandarin recognition logic.
[0042] Adaptive error-tolerant mechanism: To address issues such as ambiguous pronunciation and word omissions in dialects, the mechanism uses contextual semantic association to perform inference and completion. For example, when a user says the meter is running fast, the mechanism automatically completes the sentence as "the meter is running fast" based on the electricity business context, avoiding recognition errors caused by colloquial omissions.
[0043] Confidence verification: After recognition is completed, the confidence score of the result is output. If the confidence score is lower than the preset threshold, the secondary recognition process is automatically triggered, and the regional backup dialect model is called for re-recognition, or the process is switched to a human agent as a backup.
[0044] Collaborative Mechanism for Power Speech Corpus: The power speech corpus serves as the core knowledge base of the system, providing localized support for recognition and intent understanding. Three-dimensional mapping and matching: The corpus stores the association between dialect expressions, general terms, and business scenarios. When dialect text is identified, it is first matched with the corresponding general power term, and then mapped to the specific business scenario. For example, the dialect term "electricity tripped" is matched with the general term "circuit trip," and further associated with the fault reporting and repair business scenario.
[0045] Contextualized Expansion: For diverse dialect expressions related to the same business, full coverage is achieved through synonym association. For example, for payment scenarios, it covers various dialect variations such as paying electricity bills, charging fees, and so on.
[0046] Real-time support: During the intent recognition and response optimization stages, the corpus provides dynamic data support. For example, when optimizing dialect responses, the corpus prioritizes expressions frequently used by users in the region to ensure that the response content conforms to users' language habits.
[0047] Layered intent understanding process: The layered intent understanding module enables accurate conversion from dialect text to business requirements. Standardized conversion: Based on the mapping relationship of the corpus, the identified dialect spoken text is converted into general Mandarin business text, which solves the problems of colloquialism and non-standardization of dialect expressions.
[0048] Basic intent recognition: Identify the core business direction through keyword matching. For example, when words such as "power outage" or "no power" are detected, it is determined to be the basic intent of a power outage query.
[0049] Detailed intent matching: By combining contextual semantics, specific requests can be further clarified. For example, if a user asks when the power will be restored at the village entrance, based on the basic intent of power outage query, a more detailed intent query of power restoration time in the region can be matched to accurately locate the user's needs.
[0050] V. Optimization of Business Processing and Dialect-Based Responses The business processing and response optimization layer realizes the transformation from intent to service output: Business Interface Calls: Based on the intent recognition results, automatically call the power backend business interfaces to obtain data such as power outage information, payment records, and work order progress. For example, when a meter fault repair intent is matched, the repair work order creation interface is automatically triggered, synchronizing user address, contact information, and other information.
[0051] Dialectal Response Adaptation: The response optimization module uses dialect expression templates from a corpus to convert business data into response content that conforms to regional language habits. It supports both voice and text output formats: voice responses use synthesized speech in the regional dialect to match user auditory habits; text responses directly output dialect expressions to adapt to online interaction scenarios.
[0052] Personalized adjustments: Adjust the level of detail in responses based on the user's historical interaction records. For example, automatically extend the duration of voice responses and add repetitive reminders for elderly users to improve information reception efficiency.
[0053] The dynamic optimization module enables continuous evolution of system capabilities: closed-loop data accumulation: real-time collection of user interaction data, recognition results, and content supplemented by human agents, which are then stored in the corpus after anonymization. Emphasis is placed on collecting recognition error cases and low-confidence results as core training data for model iteration.
[0054] Periodic model fine-tuning: The dialect recognition engine is fine-tuned every quarter based on the new corpus, the acoustic feature library weights are updated, the recognition priority of high-frequency dialect expressions is strengthened, and the recognition accuracy is gradually improved.
[0055] User feedback-driven: User feedback is collected through evaluation of the accuracy of recognition during the interaction process. Cases marked as having recognition errors are automatically pushed to a manual verification queue, corrected, and added to the corpus, allowing user feedback to directly drive system optimization. Regional adaptation updates: New dialect data is regularly collected to update the feature library for subtle changes in regional dialect pronunciation, ensuring the system's adaptability to changes in regional language habits.
[0056] The above embodiments are only used to illustrate the present invention and are not intended to limit the technical solutions described in the present invention. Although the present invention has been described in detail with reference to the above embodiments, the present invention is not limited to the specific embodiments described above. Therefore, any modifications or substitutions to the present invention, and all technical solutions and improvements that do not depart from the spirit and scope of the invention, are covered within the scope of the claims of the present invention.
Claims
1. A regional intelligent power customer service system supporting dialect recognition, characterized in that, include: The multimodal interaction access layer is used to receive power service requests initiated by users via telephone voice, online voice, and text input, and to perform noise reduction, speech rate correction, and dialect audio feature enhancement preprocessing on voice requests. The regional dialect recognition engine is based on a pre-trained general speech model. It combines the dialect data of the target region with fine-tuning to build a dialect-specific recognition model. The dialect acoustic feature extraction module captures the differences in tone, initial consonant and final vowel variants of regional dialects and converts the pre-processed speech into the corresponding dialect text. The power business dialect corpus stores dialect expressions of power business scenarios within the target area, and establishes a three-dimensional mapping relationship between dialect expressions, common power terms, and business scenarios, covering mainstream power business scenarios such as fault reporting, business consultation, and payment inquiry. The hierarchical intent understanding module first converts the recognized dialect text into standard Mandarin text, and then uses a two-level architecture of basic intent matching and subdivided intent recognition, combined with a power business dialect corpus, to match the user's specific business needs. The business processing and response optimization layer calls the power back-end business interface to obtain the corresponding business data based on the intent recognition results, and optimizes and adapts the voice or text response content to the dialect of the target area. The dynamic optimization module collects user interaction voice data and recognition results, which are then manually verified and added to the speech corpus of the power business. The speech recognition model is also fine-tuned and iterated regularly based on the new speech data. The speech rate correction formula is as follows: in, To correct the speech rate, Original speech speed The average pronunciation duration of a single word in the target region's dialect is the baseline value. The average pronunciation duration of a single word in the original speech. The formula for extracting dialect tone features is as follows: in, For the first The feature weights of each tone For the first The first dialect data The actual fundamental frequency value of each tone. For the target region dialect The average fundamental frequency of each tone For the target region dialect The standard deviation of the fundamental frequency of each tone This represents the number of dialect corpora involved in feature calculation.
2. The regional intelligent power customer service system supporting dialect recognition according to claim 1, characterized in that, The dialect acoustic feature extraction module uses Mel-frequency cepstral coefficients (MFCC) combined with regional dialect tone features for modeling. It constructs independent classifiers for specific tone variants of the target region's dialect, thereby enhancing the accuracy of dialect pronunciation difference recognition. in, To determine the tone matching degree between the input speech and the dialect of the target region, For the first in the input voice The feature weights of each tone In the regional dialect model, the first Standard feature weights for each tone This represents the total number of tones in the dialect of the target region.
3. The regional intelligent power customer service system supporting dialect recognition according to claim 1, characterized in that, The power business dialect corpus also includes dialect abnormal expression correction rules, which automatically map colloquial omissions, inversions and non-standard dialect expressions of users in the region to standardized power business terms.
4. The regional intelligent power customer service system supporting dialect recognition according to claim 1, characterized in that, The hierarchical intent understanding module includes an associated scenario matching unit. When a user's statement contains multiple business keywords, it combines the contextual semantics to determine the priority of the core request. in, The intent weight for keyword k. Keywords Word frequency in the input text, This represents the total number of power business scenarios. For keywords The number of business scenarios.
5. The regional intelligent power customer service system supporting dialect recognition according to claim 1, characterized in that, The business processing and response optimization layer has a built-in dialect speech synthesis engine, which supports converting general Mandarin business results into target regional dialect speech output. During the synthesis process, the speech rate and tone characteristics of the regional dialect are matched to ensure the naturalness of the response speech.
6. The regional intelligent power customer service system supporting dialect recognition according to claim 1, characterized in that, The dynamic optimization module includes a user feedback interface, providing two quick evaluation entry points: accurate recognition and incorrect recognition. Voice data marked as incorrectly recognized is automatically pushed to a manual verification queue, while corpora with validity exceeding a threshold are automatically added to the corpus for model iteration. in, For the first The effectiveness of the corpus The number of power business keywords contained in the corpus. The total number of words in the corpus. Rate the accuracy of user feedback, with 1 for accuracy and 0 for error.
7. The regional intelligent power customer service system supporting dialect recognition according to claim 1, characterized in that, The regional dialect recognition engine supports multi-region model switching, presets independent dialect recognition model packages for different target regions, and automatically matches the corresponding regional recognition model based on the user's caller location and IP address.
8. The regional intelligent power customer service system supporting dialect recognition according to claim 1, characterized in that, The multimodal interactive access layer also supports dialect text input adaptation. When a user inputs a dialect expression, the system automatically calls the power business dialect corpus to convert the dialect text into general business terms without going through a speech recognition process.
9. The regional intelligent power customer service system supporting dialect recognition according to claim 1, characterized in that, The system also includes a dialect recognition accuracy monitoring unit, which calculates the dialect recognition accuracy in different regions and business scenarios in real time, and optimizes the visualization reports. When the recognition accuracy in a certain region is continuously lower than a preset threshold, the system automatically triggers an emergency model fine-tuning process. in, The overall accuracy of the regional model. For the first The recognition accuracy for each business scenario For the first The percentage of user requests per business scenario.
10. The regional intelligent power customer service system supporting dialect recognition according to claim 1, characterized in that, The power business dialect corpus supports linkage with the power customer service human agent system. When a human agent handles a dialect user's request, the dialect dialogue between the agent and the user is automatically recorded and added to the corpus after being anonymized.