Method, apparatus and device for responding to service based on voice geographical type and medium
Patent Information
- Application Number
- CN202610962714.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-30
- Publication Date
- 2026-08-21
AI Technical Summary
[0005]本发明的主要目的在于提供一种基于语音地缘类型的业务响应方法、装置、设备及存储介质,旨在解决现有语音交互方式在多地缘发音和非标准口语表达并存时,难以将用户语音稳定转化为可供意图分类使用的标准化业务文本,导致业务处理节点确定偏差的技术问题
[0010] Beneficial Effects: This invention relates to the field of speech and semantic technology, and discloses a business response method, apparatus, device, and medium based on speech geographical type. The method includes: acquiring a user's speech sequence and determining the speech geographical type; calling the corresponding acoustic transcription model to perform speech transcription, generating a preliminary text sequence and transcription confidence data; determining the standardization of expression based on the preliminary text sequence and transcription confidence data; generating standardized business text under standard expression conditions, and performing semantic normalization processing based on a retrieval-enhanced knowledge base and a preset language model under non-standard expression conditions to generate standardized business text; classifying intent based on the standardized business text, determining business processing nodes, and generating business response data. This invention can be applied to business scenarios such as fintech and healthcare. By matching the speech geographical type with the acoustic transcription model, corresponding transcription results can be obtained for speech from multiple geographical locations; by distinguishing between direct processing and semantic normalization processing through expression standardization determination, non-standard spoken expressions can be converted into standardized business text; and by classifying intent and calling business processing nodes based on standardized business text, intent classification bias and business processing node selection bias can be reduced.
Smart Images

Figure CN122619002A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of speech semantics technology, and in particular to a service response method, apparatus, device and medium based on speech geography. Background Technology
[0002] With the increasing application of speech recognition and natural language understanding technologies in intelligent customer service, business processing, and automated interaction processes, existing voice interaction methods typically rely on general acoustic transcription models to convert user speech into text, and then classify intent and schedule business nodes based on the text content. This approach is more suitable for scenarios with clear pronunciation, standard vocabulary, and regular word order. When user speech has obvious regional pronunciation characteristics, localized expressions, or natural colloquial inversions, existing methods are prone to producing transcribed text that deviates from the user's true meaning, making it difficult to obtain stable and standardized business semantics in the subsequent intent classification stage.
[0003] In the fintech sector, intelligent customer service and digital business processing workflows are commonly used in scenarios such as underwriting consultation, insurance claims, claims progress inquiries, policy services, account anomaly handling, payment dispute resolution, and risk verification. When describing an incident, claims, account anomalies, or service requests, users may use regional accents, local slang, and colloquial expressions. Existing voice interaction methods primarily cater to standard pronunciation and written expression, making it difficult to reliably convert such non-standard spoken content into standardized semantics recognizable by financial business processes. This can easily lead to biases in claim category identification and business processing node selection.
[0004] In the healthcare sector, voice interaction is increasingly being incorporated into processes such as intelligent triage, online consultations, health management follow-ups, medical expense reimbursement, and health insurance services. When describing symptoms, medications, test results, medical locations, reimbursement requests, or follow-up feedback, users often use localized descriptions of symptoms, colloquial language, and incomplete sentence structures. Existing voice interaction methods, when faced with these non-standard expressions, are prone to interpreting symptom descriptions, medical behaviors, expense types, or service requests as textual content deviating from their original meaning, thus affecting the accurate correspondence between subsequent intent classification results and business response content. Summary of the Invention
[0005] The main objective of this invention is to provide a business response method, apparatus, device, and storage medium based on voice geographical type, aiming to solve the technical problem that existing voice interaction methods are unable to stably convert user voice into standardized business text that can be used for intent classification when multiple geographical pronunciations and non-standard spoken expressions coexist, leading to deviations in the determination of business processing nodes.
[0006] To achieve the above objectives, the present invention provides a service response method based on voice geography type, comprising: Obtain a user's speech sequence, extract the pronunciation fingerprint features of the user's speech sequence, and determine the speech geography type based on the pronunciation fingerprint features; The acoustic transcription model corresponding to the speech geography type is invoked, and the user speech sequence is input into the acoustic transcription model for speech transcription to generate a preliminary text sequence and transcription confidence data; Based on the preliminary text sequence and the transcription confidence data, the expression standardization is determined, and a standardization determination result is generated; When the standardization determination result indicates that the preliminary text sequence belongs to the standard expression situation, the preliminary text sequence is determined as standardized business text; When the standardization determination result indicates that the preliminary text sequence belongs to the non-standard expression case, the preliminary text sequence is input into the preset semantic normalization layer, and the semantic normalization processing of the preliminary text sequence is performed based on the retrieval enhancement knowledge base and the preset language model to generate standardized business text; The standardized business text is input into the intent classification engine for intent classification, and intent classification results are generated. Based on the intent classification results, business processing nodes are determined, the business processing nodes are called to obtain node processing results, and business response data is generated based on the node processing results.
[0007] Furthermore, to achieve the above objectives, the present invention provides a service response device based on voice geography type, comprising: The voice geography recognition module is used to acquire user voice sequences, extract pronunciation fingerprint features of the user voice sequences, and determine the voice geography type based on the pronunciation fingerprint features; The acoustic transcription processing module is used to call the acoustic transcription model corresponding to the speech geography type, input the user speech sequence into the acoustic transcription model for speech transcription, and generate a preliminary text sequence and transcription confidence data; The expression standardization determination module is used to determine the expression standardization based on the preliminary text sequence and the transcription confidence data, and generate a standardization determination result; The direct text generation module is used to determine the preliminary text sequence as standardized business text when the standardization determination result indicates that the preliminary text sequence belongs to the standard expression situation; The semantic normalization processing module is used to input the preliminary text sequence into a preset semantic normalization layer when the standardization judgment result indicates that the preliminary text sequence belongs to a non-standard expression case, and perform semantic normalization processing on the preliminary text sequence based on the retrieval enhancement knowledge base and the preset language model to generate standardized business text. The intent scheduling and response module is used to input the standardized business text into the intent classification engine for intent classification, generate intent classification results, determine business processing nodes based on the intent classification results, call the business processing nodes to obtain node processing results, and generate business response data based on the node processing results.
[0008] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and a voice-based geographical type-based service response program stored in the memory and executable on the processor, wherein the voice-based geographical type-based service response program, when executed by the processor, implements the steps of the voice-based geographical type-based service response method as described above.
[0009] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing a service response program based on voice geography type, wherein when the service response program based on voice geography type is executed by a processor, it implements the steps of the service response method based on voice geography type as described above.
[0010] Beneficial Effects: This invention relates to the field of speech and semantic technology, and discloses a business response method, apparatus, device, and medium based on speech geographical type. The method includes: acquiring a user's speech sequence and determining the speech geographical type; calling the corresponding acoustic transcription model to perform speech transcription, generating a preliminary text sequence and transcription confidence data; determining the standardization of expression based on the preliminary text sequence and transcription confidence data; generating standardized business text under standard expression conditions, and performing semantic normalization processing based on a retrieval-enhanced knowledge base and a preset language model under non-standard expression conditions to generate standardized business text; classifying intent based on the standardized business text, determining business processing nodes, and generating business response data. This invention can be applied to business scenarios such as fintech and healthcare. By matching the speech geographical type with the acoustic transcription model, corresponding transcription results can be obtained for speech from multiple geographical locations; by distinguishing between direct processing and semantic normalization processing through expression standardization determination, non-standard spoken expressions can be converted into standardized business text; and by classifying intent and calling business processing nodes based on standardized business text, intent classification bias and business processing node selection bias can be reduced. Attached Figure Description
[0011] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings: Figure 1 This is a schematic diagram of an application environment for a service response method based on voice geography in one embodiment of the present invention; Figure 2 This is a flowchart illustrating an embodiment of the service response method based on voice geography of the present invention; Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the service response device based on voice geography of the present invention; Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention; Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation
[0012] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.
[0013] The service response method based on voice geography provided in this invention can be applied to, for example... Figure 1 In this application environment, the client communicates with the server via a network. The server can obtain the user's speech sequence from the client and determine the speech geography; it then calls the corresponding acoustic transcription model to perform speech transcription, generating a preliminary text sequence and transcription confidence data; based on the preliminary text sequence and transcription confidence data, it determines the standardization of expression; in the case of standard expression, it generates standardized business text; in the case of non-standard expression, it performs semantic normalization processing based on a retrieval-enhanced knowledge base and a preset language model to generate standardized business text; based on the standardized business text, it performs intent classification, determines business processing nodes, and generates business response data. This invention can be applied to business scenarios such as fintech and healthcare. By matching the speech geography to the acoustic transcription model, it enables speech from multiple geography regions to obtain corresponding transcription results; by distinguishing between direct processing and semantic normalization processing through expression standardization determination, it enables non-standard spoken expressions to be converted into standardized business text; and by performing intent classification and business processing node invocation based on standardized business text, it can reduce intent classification bias and business processing node selection bias. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will be described in detail below through specific embodiments.
[0014] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the service response method based on voice geography provided by the present invention. It should be noted that although a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.
[0015] like Figure 2 As shown, the service response method based on voice geography proposed in this invention includes the following steps: S10, acquire the user's voice sequence, extract the pronunciation fingerprint features of the user's voice sequence, and determine the voice geographical type based on the pronunciation fingerprint features; In this embodiment, the user's voice sequence is obtained by organizing continuous voice content accessed through telephone channels, mobile terminal voice entry points, web page voice entry points, or smart terminal voice entry points. The continuous voice content is collected according to session identifiers, time stamps, and voice source channels, and through voice boundary resolution and sound source identification, prompts, ambient sounds, and bypass sounds are removed, retaining the user's continuous spoken content.
[0016] Pronunciation fingerprint features are used to characterize geographical pronunciation differences in user speech sequences, and can include initial consonant pronunciation shifts, final vowel pronunciation shifts, tone direction, and prosodic pauses. During generation, the user speech sequence is divided into pronunciation segments, syllable boundary markings and tone change position markings are performed on the pronunciation segments, and then matched with a baseline pronunciation template set and a geographical acoustic template set to obtain pronunciation fingerprint features.
[0017] Speech geolocation is obtained by matching pronunciation fingerprint features with a geolocation index table. The geolocation index table records geolocation identifiers, pronunciation offset ranges, tone patterns, and prosodic pause patterns. Matching methods can include template similarity matching, vector retrieval, or geolocation discrimination models. Speech geolocation is used to represent pronunciation attributes in speech, not to represent the user's actual geographical location.
[0018] In cases where background announcements and other voices are present in telephone conversations, user speech sequences are generated by session aggregation and speech source identification. Then, speech fingerprint features are extracted by segmenting speech segments, marking syllable boundaries, and marking tone change positions. Finally, speech geographic type is determined by template matching.
[0019] When short voice messages are submitted in multiple parts on a mobile device, a short sentence merging and geo-geographical discrimination model is used to determine the speech geo-type. Multiple short voice segments are merged into a user speech sequence according to time stamps and session identifiers. The model performs acoustic coding, syllable boundary marking, tone change marking, and feature fusion on the user speech sequence, and outputs pronunciation fingerprint features and speech geo-type.
[0020] In the fintech business, when users describe insurance claims, policy service needs, or account anomalies through voice input, the voice content is aggregated through conversation and the source of the voice is identified to form a user voice sequence. Then, the pronunciation fingerprint features are extracted and the geographical type of the voice is determined.
[0021] In the healthcare business, when users submit symptom descriptions, medication feedback, or health follow-up voice messages through health service platforms, multiple voice segments can be aggregated into a user voice sequence, and then the geographical type of the voice can be determined through pronunciation fingerprint features.
[0022] This embodiment generates user speech sequences by session aggregation, speech boundary parsing, and speech source identification, which can reduce interference from non-target sounds; it obtains speech fingerprint features by syllable boundary marking, tone change position marking, and speech offset extraction, which can convert geographical speech differences into matchable acoustic representations; and it determines speech geographical types by using speech fingerprint features, which can improve the stability of speech attribute recognition.
[0023] S20, invoke the acoustic transcription model corresponding to the speech geography type, input the user speech sequence into the acoustic transcription model for speech transcription, and generate a preliminary text sequence and transcription confidence data; In this embodiment, the speech geography type is used to select an acoustic transcription model that matches the user's pronunciation attributes. After matching the speech geography type with the acoustic model index data, an acoustic transcription model identifier and a geography pronunciation parameter package are obtained. The geography pronunciation parameter package may include initial consonant offset compensation parameters, final vowel pronunciation compensation parameters, tone change compensation parameters, and business word pronunciation weight data, which are used to enhance the recognition weight of business words such as insurance claims, policy services, account anomalies, symptom descriptions, and medication feedback under different geography pronunciations.
[0024] After the user's speech sequence is input into the acoustic transcription model, the acoustic coding layer converts the speech frames into a geosound representation sequence. The candidate word decoding layer generates a set of candidate word paths based on the geosound representation sequence and rearranges the paths by combining the business word pronunciation weight data to obtain the target word path. The target word path generates a preliminary text sequence after text merging and duplicate segment removal. The confidence output layer extracts path confidence values and word segment confidence values from the path confidence distribution and combines them to generate transcription confidence data.
[0025] In cases where telephone voice recordings suffer from compression distortion and geographical pronunciation shifts, a geographical pronunciation parameter package loading method is used to complete the speech transcription. The acoustic transcription model loads initial consonant offset compensation parameters, final vowel pronunciation compensation parameters, and business word pronunciation weight data according to the geographical type of the speech. After the user's speech sequence is encoded and candidate words are decoded, a preliminary text sequence and transcription confidence data are output.
[0026] When short, intermittent voice messages are submitted on mobile devices, a segmented encoding and path merging method is used to complete the speech transcription. The user's voice sequence is divided into multiple short voice segments according to the pause positions. Each short voice segment generates candidate word paths, which are then merged into target word paths according to time stamps, and transcription confidence data is generated.
[0027] In the fintech business, when users describe insurance claims, policy service needs, or account anomalies, the acoustic transcription model enhances the transcription weights of business terms such as claims, policies, and accounts based on speech geography, generating preliminary text sequences and transcription confidence data.
[0028] In the healthcare business, when users submit symptom descriptions, medication feedback, or health follow-up audio recordings, the acoustic transcription model processes localized pronunciations based on the geographical type of the speech and rearranges candidate words for business terms such as symptoms, medication, and follow-up.
[0029] This embodiment calls the corresponding acoustic transcription model by voice geography type and loads the geography pronunciation parameter package, which can transcribe the user's voice sequence according to pronunciation attributes; by rearranging the candidate word paths by business word pronunciation weight data, the stability of business word transcription under geography pronunciation conditions can be improved; by synchronously generating preliminary text sequence and transcription confidence data, text content and credibility can be provided for subsequent text judgment.
[0030] S30, based on the preliminary text sequence and the transcription confidence data, a standardization determination is made, and a standardization determination result is generated; In this embodiment, the preliminary text sequence is the acoustically transcribed text content, which may contain accurately transcribed standard expressions, or it may contain low-confidence words, inverted word order content, or localized expressions. Transcription confidence data is used to represent the credibility of each text segment in the preliminary text sequence, and may include overall confidence values, word segment confidence values, and markers for consecutive low-confidence segments. The expression standardization determination determines whether the preliminary text sequence can be directly used as standard expression content by simultaneously analyzing the text content and confidence state. During the determination, the preliminary text sequence can be divided into a set of text segments to be determined, and then low-confidence consecutive segments are marked based on the transcription confidence data; simultaneously, word order structure recognition is performed on the set of text segments to be determined, and inverted word order segments are marked; expression hit recognition can also be performed in a preset multi-geographical expression index data, and non-standard expression segments are marked. When no low-confidence consecutive segments, inverted word order segments, or non-standard expression segments appear, a standardization determination result indicating that the preliminary text sequence belongs to a standard expression is generated; when any marker appears, a standardization determination result indicating that the preliminary text sequence belongs to a non-standard expression is generated.
[0031] When the transcription result is long and contains multiple business descriptions, sentence segmentation and segment-level confidence labeling are used to determine the standardization of expression. The initial text sequence is divided into a set of text segments to be judged based on pause positions, punctuation positions, and business phrase boundaries. Transcription confidence data is aligned with the set of text segments to be judged based on segment position; segments with a confidence level below a preset value and continuous distribution are marked as low-confidence continuous segments. The set of text segments to be judged then undergoes word order structure recognition and multi-geographical expression hit recognition, marking inverted word order segments and non-standard expression segments respectively, ultimately generating the standardization judgment result.
[0032] In situations involving numerous short sentence interactions on mobile devices, a holistic approach using merged short sentences is employed to reduce misjudgments of individual sentences. After merging the transcribed content of multiple short sentences into a preliminary text sequence, consecutive low-confidence segments are identified based on transcription confidence data, and inverted word order segments are identified according to the positional relationships between business action words, object words, and state words. Data segments that match the localization expression index are marked as non-standard expression segments. All marking results are used together to generate a standardized judgment result.
[0033] In the fintech business, when users describe insurance claims, policy service needs, payment disputes, or account anomalies via voice, the initial text sequence may contain localized terms or inverted expressions. By transcribing confidence data and judging text fragments, low-confidence continuous segments, inverted word order segments, and non-standard expression segments can be identified, generating standardized judgment results.
[0034] In the healthcare field, when users submit symptom descriptions, medication feedback, or health follow-up audio recordings, the initial text sequence may contain everyday descriptions of symptoms and incomplete sentence order. By segmenting, using confidence markers, and identifying expression hits, it can be determined whether the initial text sequence belongs to a standard or non-standard expression.
[0035] This embodiment can identify text regions with insufficient transcription reliability by dividing the initial text sequence into a set of text segments to be judged and marking low-confidence continuous segments with transcription confidence data; it can identify sentence order abnormalities and localized expressions by recognizing word order structure and multi-geographical expression hits; and it can improve the accuracy of separating standard and non-standard expressions by generating standard judgment results from the above judgment results.
[0036] S40, when the standardization determination result indicates that the preliminary text sequence belongs to the standard expression situation, the preliminary text sequence is determined as standardized business text; In this embodiment, the standardization judgment result is used to indicate whether the preliminary text sequence can directly enter the standardized text generation path. A standard description indicates that the business statement structure in the preliminary text sequence is complete, the transcription reliability meets the requirements, and it does not contain localized expressions requiring semantic rewriting. In this case, the preliminary text sequence does not require deep semantic rewriting and can obtain standardized business text through statement boundary location, business statement fragment extraction, colloquial content removal, and standardization of business terminology expressions.
[0037] The initial text sequence can be boundary-located according to pause positions, punctuation positions, business action phrases, and business object phrases, resulting in multiple standard candidate text fragments. These standard candidate text fragments retain their original word order, only removing repetitive words, modal particles, and redundant colloquialisms without business meaning. The refined direct-access text sequence is then matched against a pre-defined business terminology list to unify different expressions under the same business meaning into standard business terminology, forming a standard direct-access text sequence. Since the standard direct-access text sequence is complete and its semantics have not been rewritten, it can be identified as standardized business text.
[0038] Assuming the telephone voice transcription results are complete in word order and have stable confidence, a direct processing method is used to generate standardized business text. The initial text sequence is segmented into multiple standard candidate text fragments according to sentence boundaries, removing filler text such as "um," "ah," and repetitive confirmations, while preserving the user's original business expression order. The direct processing text sequence is then matched with a preset business terminology table, unifying expressions such as claims progress, policy services, account anomalies, and payment disputes into corresponding business terminology.
[0039] When the results of short-sentence speech-to-text transcription on mobile devices are relatively short, a segment merging and terminology unification approach is used to generate standardized business text. After multiple short sentences form a preliminary text sequence, standard candidate text segments are extracted according to business actions and business objects. Adjacent segments of the same business object are merged to generate a passable text sequence. After terminology consistency matching, the passable text sequence yields a standard passable text sequence, which is then identified as the standardized business text.
[0040] In the fintech business, when users transcribe their speech to inquire about policy services and claims progress, if the initial text sequence is complete and the confidence level is stable, the spoken language can be directly removed, and the business expression can be standardized into a standardized business text.
[0041] In the healthcare business, when users transcribe their voice messages to provide feedback on medication use or inquire about health follow-up arrangements, if the initial text sequence is a standard expression, the original sentence structure can be retained, and only the expression of business terminology can be standardized to generate standardized business text.
[0042] This embodiment identifies standard expressions by determining the standardization results, which can avoid unnecessary semantic rewriting of preliminary text sequences that already meet the standard expression requirements. By locating sentence boundaries, eliminating colloquial filler content, and unifying the expression of business terms, standardized business text can be generated while maintaining the original semantics, thereby improving the stability of standard expression content in the subsequent business understanding process.
[0043] S50, when the standardization determination result indicates that the preliminary text sequence belongs to the non-standard expression case, the preliminary text sequence is input into the preset semantic normalization layer, and the semantic normalization processing of the preliminary text sequence is performed based on the retrieval enhancement knowledge base and the preset language model to generate standardized business text; In this embodiment, non-standard expressions refer to the presence of localized vocabulary, colloquialisms, omitted components, inverted word order, or low-confidence text fragments in the initial text sequence. A preset semantic normalization layer is used to standardize and transform this type of text, enabling the initial text sequence to be organized into standardized business text. The semantic normalization layer may include a non-standard expression recognition unit, a retrieval enhancement unit, a prompt data generation unit, a language model rewriting unit, and a consistency verification unit. The non-standard expression recognition unit identifies localized vocabulary, inverted sentences, and omitted semantic fragments from the initial text sequence, generating a set of non-standard expression units. The retrieval enhancement knowledge base stores the mapping content between non-standard expressions and standard business terms, as well as synonyms, scenario phrases, and business action phrases. The retrieval enhancement unit performs term index retrieval and semantic vector retrieval based on the set of non-standard expression units, generating a set of candidate mapping terms. The prompt data generation unit combines the set of candidate mapping terms with the initial text sequence to form normalized prompt data, which is used to constrain the rewriting range of the preset language model. The pre-defined language model, based on normalized prompt data, performs business term replacement, semantic component completion, and word order adjustment on the initial text sequence to generate candidate normalized text. After the candidate normalized text undergoes mapping consistency verification, inconsistent term fragments are corrected to obtain standardized business text.
[0044] When localized vocabulary accounts for a high proportion, a combination of term index retrieval and semantic vector retrieval is used to generate normalized prompt data. Lexical fragments from the non-standard expression unit set are input into the retrieval enhancement knowledge base. Term index retrieval is used to match explicitly registered localized expressions, while semantic vector retrieval is used to match synonyms and colloquial phrases. The candidate mapping term set contains standard business terms, candidate semantic categories, and corresponding text fragment positions. After receiving the normalized prompt data, the pre-defined language model replaces localized vocabulary in the initial text sequence with business terms, retains the user's original request content, and outputs candidate normalized text.
[0045] In cases involving numerous word order inversions and omissions, a combination of structural completion and word order reordering is employed to generate standardized business text. The semantic normalization layer identifies semantic gaps based on business action phrases, business object phrases, and state description phrases in the initial text sequence, and then writes standard terms from the candidate mapping term set into the normalized prompt data. The pre-defined language model completes the missing business objects, action relationships, and state descriptions based on the normalized prompt data, and reorders the inverted fragments into standard business word order. The mapping consistency verification result checks whether the business terms in the candidate normalized text are consistent with the candidate mapping term set; if inconsistent, terminology correction is performed.
[0046] In the fintech business, when users describe insurance claims, policy services, account anomalies, or payment disputes via voice, the initial text sequence may contain localized expressions and inverted sentences. After the semantic normalization layer identifies non-standard expression units, it retrieves standard business terms from the enhanced knowledge base. Based on these terms, the pre-defined language model completes the business semantics and restores the word order, generating standardized business text.
[0047] In the healthcare business, when users submit symptom descriptions, medication feedback, or health follow-up audio recordings, the initial text sequence may contain colloquial expressions of symptoms, omitted subjects, or incomplete sentence order. The semantic normalization layer, based on a retrieval-enhanced knowledge base and a pre-defined language model, performs terminology replacement and semantic completion on content such as symptoms, medication, and follow-up to generate standardized business text.
[0048] This embodiment identifies non-standard expression units through a semantic normalization layer, which can locate localized words, inverted sentences, and omitted segments in the initial text sequence; by retrieving and enhancing the knowledge base to generate a set of candidate mapping terms, it can provide clear standard terms and semantic constraints for the preset language model; by performing business term replacement, semantic component completion, and word order adjustment through the preset language model, non-standard expressions can be converted into standardized business text, thereby improving the stability of non-standard spoken content in the subsequent business understanding process.
[0049] S60, input the standardized business text into the intent classification engine for intent classification, generate intent classification results, determine the business processing node based on the intent classification results, call the business processing node to obtain the node processing results, and generate business response data based on the node processing results.
[0050] In this embodiment, the standardized business text has been unified in its expression and can be used as input text for the intent classification engine. The intent classification engine may include a text encoding unit, a phrase recognition unit, an intent matching unit, and a slot extraction unit. The text encoding unit converts the standardized business text into a semantic representation; the phrase recognition unit identifies business action phrases, business object phrases, and state description phrases; the intent matching unit determines the intent category based on the above phrases; and the slot extraction unit extracts fields such as object, action, state, and time required for business processing, ultimately forming the intent classification result. The intent classification engine can be trained using business text with intent category annotations and slot boundary annotations. The input is standardized business text, and the output is intent category and slot data. The training process can combine intent classification loss and slot boundary recognition loss to update the model parameters.
[0051] Business processing nodes are used to handle the business processing steps corresponding to the intent classification results. The intent classification results can be matched with the business node configuration data, which records the intent category, node identifier, node input parameter structure, and response content format. After a business processing node is identified, its input parameter structure can be populated based on standardized business text and the intent classification results to generate node call data. Upon receiving the node call data, the business processing node returns the node processing result, which may include processing status, business content, prompts, and next interaction steps. After status recognition and response content extraction, the node processing result generates business response data.
[0052] When standardized business text contains explicit business actions and objects, a combination of intent template matching and slot extraction is used to determine business processing nodes. The intent classification engine identifies action phrases, object phrases, and state phrases in the standardized business text, generates intent classification results, and then determines the business processing nodes based on the business node configuration data. After the node input parameter structure is filled in according to the intent classification results and the standardized business text, it forms node invocation data. The business processing node returns the node processing result and converts it into business response data.
[0053] When business expressions are complex or contain multiple state descriptions, a combination of text encoding and intent classification models is used to generate intent classification results. The text encoding model converts standardized business text into a semantic representation, the intent classification model outputs candidate intent categories, and the slot extraction unit extracts business objects and state fields. The candidate intent categories and slot fields together determine the business processing nodes, and the node processing results are used to generate business response data after processing state identification and response content extraction.
[0054] In the fintech business, after user voice is converted into standardized business text, the intent classification engine can identify request categories such as insurance claims, policy services, payment disputes, or account anomalies, and determine the corresponding business processing node. After the business processing node returns the processing status and prompt content, business response data is generated.
[0055] In the healthcare business, after user voice is converted into standardized business text, the intent classification engine can identify request categories such as symptom consultation, medication feedback, health follow-up, or expense reimbursement, and determine the corresponding business processing node. After the business processing node returns inquiry prompts, follow-up arrangements, or expense explanations, business response data is generated.
[0056] This embodiment inputs standardized business text into the intent classification engine, which can identify business actions, business objects, and status descriptions based on the unified text content; determines business processing nodes through intent classification results, which can make the business processing process correspond to user requests; and generates business response data through node processing results, which can organize the node return content into feedback-friendly response content, thereby reducing the matching deviation between intent classification results and business processing nodes.
[0057] In one embodiment, step S10 above includes: S101, Receive a voice interaction request, and extract the continuous voice stream, the conversation voice identifier, and the voice source channel identifier from the voice interaction request; S102, based on the conversation speech identifier and the speech source channel identifier, the continuous speech stream is segmented to generate a candidate speech segment set; S103, Perform speech boundary parsing on the candidate speech segment set to generate speech start and end positions and pause interval positions; S104, Identify the source of speech in the candidate speech segment set and generate non-user speech markers; S105, based on the speech start and end positions, the pause interval positions, and the non-user speech markers, filter user speech segments from the candidate speech segment set to generate a user speech sequence; S106, Based on the user's speech sequence, divide the speech segments to generate a speech segment sequence; S107, mark the syllable boundaries and tone change positions of each syllable segment in the syllable segment sequence to generate a syllable marking sequence; S108, Based on the syllable marker sequence and the preset reference pronunciation template set, extract the initial consonant pronunciation offset features, final vowel pronunciation offset features, tone direction features and rhythmic pause features to generate a candidate pronunciation feature set; S109, perform segment-level matching between the candidate pronunciation feature set and the preset geoacoustic template set to generate a segment geo-tag set; S110, based on the arrangement order of the pronunciation segment sequence, the set of geographical markers of the segments is continuously aggregated to generate pronunciation fingerprint features; S111, the pronunciation fingerprint features are matched with the geopolitical type index table, and the matched geopolitical type identifier is determined as the pronunciation geopolitical type.
[0058] In this embodiment, the voice interaction request carries the voice data and source information generated during a single voice interaction. The continuous voice stream is the audio data separated from the voice interaction request, retaining the user's continuous vocalizations during the interaction. The session voice identifier distinguishes different voice interaction processes, preventing the mixing of voice segments from different sessions. The voice source channel identifier distinguishes between telephone channels, mobile entry points, web page entry points, or smart terminal entry points, enabling subsequent segment aggregation to process sampling frequency, noise levels, and voice segment formats according to channel characteristics. After the continuous voice stream is aggregated using the session voice identifier and the voice source channel identifier, a candidate voice segment set is formed. This candidate voice segment set retains voice segments belonging to the same source channel within the same session, reducing the mixing of cross-session and cross-channel segments.
[0059] Speech boundary resolution is used to locate valid speech boundaries from a set of candidate speech segments, generating speech start and end positions and pause interval positions. Speech start and end positions can be determined based on energy changes, zero-crossing rate changes, spectral stability, and speech activity detection results, used to mark the beginning and end positions of each speech segment. Pause interval positions can be determined based on the duration of silence between adjacent speech segments, used to distinguish between continuous expressions, short pauses, and sentence intervals. Speech source identification is used to determine whether there are prompts, background announcements, bypass voices, or other non-target speech content in the candidate speech segment set, and generates non-user speech markers. Based on speech start and end positions, pause interval positions, and non-user speech markers, user speech segments can be filtered, preserving continuous spoken content from the same user and generating user speech sequences.
[0060] The user's speech sequence is used to carry the vocal content required for subsequent pronunciation attribute analysis. Pronunciation segment division can be performed by dividing the user's speech sequence into multiple pronunciation segments based on syllable changes, short-term spectral changes, and pause intervals, generating a pronunciation segment sequence. Each pronunciation segment, after being marked with syllable boundaries and tone change positions, forms a syllable marker sequence. Syllable boundary markers determine the position range of each syllable within the pronunciation segment, while tone change position markers determine the locations of pitch rises, falls, transitions, or extensions. After comparing the syllable marker sequence with a preset set of reference pronunciation templates, initial consonant pronunciation offset features, final vowel pronunciation offset features, tone direction features, and prosodic pause features can be extracted to generate a candidate pronunciation feature set. Initial consonant pronunciation offset features reflect the deviation of the initial consonant's pronunciation position, aspiration state, and retroflex state from the reference pronunciation template set; final vowel pronunciation offset features reflect changes in mouth opening, nasalization, and coda extension; tone direction features reflect pitch rises, falls, and transitions; and prosodic pause features reflect pauses between short sentences, stress distribution, and connected speech rhythm.
[0061] The candidate pronunciation feature set is matched with a pre-defined set of geo-acoustic templates at the segment level to obtain a set of segment geo-tags for each pronunciation segment. The geo-acoustic template set can store initial consonant offset patterns, vowel change patterns, tone direction patterns, and prosodic pause patterns corresponding to different geo-pronunciation types. Segment-level matching can employ template similarity matching, vector distance matching, or classification models. The segment geo-tag set is continuously aggregated according to the order of the pronunciation segment sequence, which can reduce the impact of mismatches in individual segments on the overall judgment and generate pronunciation fingerprint features. After the pronunciation fingerprint features are matched with the geo-type index table, the matched geo-type identifier is determined as the speech geo-type. The geo-type index table can record the geo-type identifier, geo-acoustic template identifier, pronunciation offset range, tone direction pattern, and prosodic pause pattern, enabling the pronunciation fingerprint features to be converted into speech geo-types that can be used subsequently.
[0062] This embodiment uses continuous speech streams, conversational speech identifiers, and speech source channel identifiers in voice interaction requests to group segments, reducing the mixing of speech from different conversations and different source channels. By resolving speech boundaries and identifying the source of speech, it generates speech start and end positions, pause intervals, and non-user speech markers, allowing for the selection of user speech segments from the candidate speech segment set, reducing interference from prompts, ambient sounds, and bypass sounds on pronunciation analysis. Through speech segment division, syllable boundary marking, tone change position marking, and comparison with a baseline pronunciation template set, a candidate pronunciation feature set reflecting geographical pronunciation differences can be obtained. By performing segment-level matching between the candidate pronunciation feature set and the geographical acoustic template set, and continuously aggregating them according to the sequence of speech segments, stable pronunciation fingerprint features can be formed. By matching the pronunciation fingerprint features with a geographical type index table, the geographical type of speech can be determined, improving the stability of pronunciation attribute recognition in multi-geographical speech environments.
[0063] In one embodiment, step S20 above includes: S201, based on the speech geography type, match the acoustic transcription model identifier and the geography pronunciation parameter package containing the pronunciation weight data of the business words from the preset acoustic model index data; S202, based on the acoustic transcription model identifier, the acoustic transcription model is invoked, and the geopolitical pronunciation parameter package is loaded into the acoustic transcription model to generate a geopolitical transcription processing instance; S203, input the user speech sequence into the geopolitical transcription processing instance for geopolitical pronunciation enhancement encoding to generate a geopolitical acoustic representation sequence; S204, Based on the geoacoustic representation sequence, candidate words are decoded to generate a candidate word path set and a path confidence distribution; S205, Based on the business word pronunciation weight data in the geographical pronunciation parameter package, the candidate word path set is rearranged to generate the target word path; S206, convert the target word path into a preliminary text sequence; S207, extract the path confidence value and word fragment confidence value corresponding to the target word path from the path confidence distribution, and generate transcription confidence data based on the path confidence value and the word fragment confidence value.
[0064] In this embodiment, the speech geography type is used to select an acoustic transcription model that can handle the corresponding pronunciation attributes from the acoustic model index data. The acoustic model index data can store the speech geography type, the acoustic transcription model identifier, the geography pronunciation parameter package, and the model version tag. After the speech geography type is matched, the acoustic transcription model identifier and the geography pronunciation parameter package containing the pronunciation weight data of the business words are obtained. The geography pronunciation parameter package may include initial consonant offset compensation parameters, final vowel pronunciation compensation parameters, tone change compensation parameters, prosodic pause compensation parameters, and business word pronunciation weight data. The initial consonant offset compensation parameters are used to correct acoustic differences caused by initial consonant pronunciation position, aspiration intensity, and tongue roll state; the final vowel pronunciation compensation parameters are used to correct acoustic differences caused by mouth opening degree, nasalization degree, and coda extension; the tone change compensation parameters are used to correct pitch direction and tone transition differences; and the business word pronunciation weight data is used to increase the probability that specific business words are identified as target words under different pronunciation forms.
[0065] The acoustic transcription model may include an acoustic feature access layer, a geophonic parameter loading layer, a geophonic enhancement coding layer, a candidate word decoding layer, a path rearrangement layer, and a confidence output layer. The acoustic feature access layer receives the user's speech sequence and converts it into a frame-level acoustic representation. The geophonic parameter loading layer writes geophonic parameter packets into the model's inference parameters, enabling subsequent encoding processes to utilize geophonic compensation information. The geophonic enhancement coding layer encodes the frame-level acoustic representation, generating a geophonic representation sequence. This sequence preserves information such as syllable boundaries, pitch variations, pronunciation duration, and pause distribution, and compensates for pronunciation shifts by incorporating the geophonic parameter packets.
[0066] The candidate word decoding layer generates a candidate word path set and a path confidence distribution based on the geosonic representation sequence. The candidate word path set contains multiple possible word combination paths, and the path confidence distribution records the confidence state of different candidate paths and word fragments. The path reordering layer sorts and adjusts the candidate word path set according to the business word pronunciation weight data, giving business words a higher path selection priority under geosonic pronunciation conditions. After text merging, duplicate fragment removal, and sentence boundary adjustment, the target word path generates a preliminary text sequence. The confidence output layer extracts the path confidence value and word fragment confidence value corresponding to the target word path from the path confidence distribution and combines them into transcription confidence data. The path confidence value reflects the overall credibility of the entire target word path, and the word fragment confidence value reflects the recognition credibility of local words.
[0067] The acoustic transcription model can be trained using speech data with annotations for geographical speech type, word transcription, and business term. After the training data is input into the acoustic feature access layer, the geographical speech enhancement coding layer generates an acoustic encoded representation, and the candidate word decoding layer outputs candidate word paths. Training objectives may include word transcription loss, business term recognition loss, and confidence calibration loss. The word transcription loss is used to constrain candidate word paths to be close to the annotated text, the business term recognition loss is used to enhance the recognition ability of business terms under different geographical speech types, and the confidence calibration loss is used to ensure that the path confidence value and word segment confidence value are consistent with the actual transcription reliability. After the model is updated, the acoustic model index data can save the model version tag and applicable geographical speech type, enabling the calling process to select the appropriate acoustic transcription model based on the geographical speech type.
[0068] This embodiment utilizes a speech geography matching acoustic transcription model and a geography pronunciation parameter package to allow user speech sequences to enter the corresponding acoustic transcription process according to pronunciation attributes. The geography pronunciation parameter package compensates for initial consonant shifts, final vowel shifts, tone changes, and prosodic pauses, reducing the impact of geography pronunciation differences on acoustic coding. By rearranging the candidate word path set using business word pronunciation weight data, the recognition stability of business words in multi-geographical pronunciation environments can be improved. Finally, by generating transcription confidence data using path confidence values and word fragment confidence values, the reliability of the transcription can be provided while outputting the initial text sequence.
[0069] In one embodiment, step S30 above includes: S301, Based on the preliminary text sequence, perform sentence segmentation to generate a set of text segments to be judged; S302, based on the transcription confidence data, mark the low-confidence continuous segments in the set of text segments to be judged that are lower than the preset confidence value, and generate a confidence judgment mark set; S303, based on the set of text fragments to be determined, perform word order structure recognition, mark the word order inverted fragments, and generate a word order structure mark set; S304, Based on the set of text fragments to be determined, expression hit identification is performed in the preset multi-geographical expression index data, and non-standard expression fragments are marked to generate a non-standard expression tag set; S305, when the confidence determination mark set does not contain low-confidence continuous segment marks, the word order structure mark set does not contain word order inversion segment marks, and the non-standard expression mark set does not contain non-standard expression segment marks, a standard determination result indicating that the preliminary text sequence belongs to the standard expression case is generated; S306, when the confidence determination mark set contains low-confidence continuous segment marks, the word order structure mark set contains word order inversion segment marks, or the non-standard expression mark set contains non-standard expression segment marks, a standard determination result indicating that the preliminary text sequence belongs to the non-standard expression situation is generated.
[0070] In this embodiment, the initial text sequence is divided into sentence segments to form a set of text segments to be judged. Sentence segments can be divided according to pause positions, punctuation positions, semantic breaks, action word positions, and object word positions. Each segment retains its content, position, and adjacency relationships, facilitating the subsequent labeling of confidence states, word order states, and multi-geographical expression states to the corresponding segments. The set of text segments to be judged does not alter the textual content of the initial text sequence; it only organizes the initial text sequence into segments, enabling subsequent judgments to locate specific segments without requiring a single judgment of the entire text.
[0071] Transcription confidence data can include overall confidence values, word / fragment confidence values, and path confidence values. When labeling low-confidence contiguous segments based on transcription confidence data, each segment in the set of text segments to be judged can be aligned with the word / fragment confidence value, and consecutive words with confidence values below a preset threshold can be grouped into low-confidence contiguous segments. Low-confidence contiguous segments are used to indicate that a continuous text has insufficient transcription reliability. The confidence judgment label set records the location, content, and confidence status of low-confidence contiguous segments, enabling the standard judgment results to reflect the degree of transcription reliability.
[0072] Word order structure recognition is used to determine the arrangement relationships between action phrases, object phrases, state phrases, and positional phrases in a set of text fragments to be judged. During recognition, predicate structures, object components, state components, and complementary components are extracted from the text fragments and compared with a preset word order structure template. When object phrases and action phrases are inverted, state phrases are located in positions lacking a sequential relationship, or positional phrases are inserted, resulting in incomplete action-object relationships, they are marked as word order inverted fragments. The word order structure tag set records word order inverted fragments and their corresponding structure types, used to determine whether the initial text sequence can directly form a standard business expression.
[0073] The multi-geographical expression index data is used to store localized vocabulary, colloquial phrases, and spoken abbreviations, along with their corresponding standard expression categories. When identifying expression matches based on a set of text fragments to be judged, word segmentation, phrase extraction, and semantic vector retrieval can be performed on the text fragments, and then matched against the multi-geographical expression index data. Fragments that match localized vocabulary, colloquial phrases, or spoken abbreviations are marked as non-standard expression fragments. The non-standard expression tag set records the location of the non-standard expression fragments, the matched words, and the candidate standard expression categories, used to determine whether the initial text sequence needs further normalization.
[0074] The standardization judgment result is generated jointly by the confidence judgment mark set, the word order structure mark set, and the non-standard expression mark set. When no corresponding abnormal mark is found in any of the three mark sets, it indicates that the preliminary text sequence has a relatively stable transcription confidence, complete word order structure, and standard expression content, generating a standardization judgment result indicating that the preliminary text sequence belongs to the standard expression situation. When any mark set contains a low-confidence continuous segment mark, a word order inversion segment mark, or a non-standard expression segment mark, it indicates that the preliminary text sequence has transcription instability, word order abnormalities, or localized expression content, generating a standardization judgment result indicating that the preliminary text sequence belongs to the non-standard expression situation.
[0075] This embodiment generates a set of text segments to be judged by segmenting sentence fragments, which can locate the subsequent judgment to a specific text segment; it identifies text regions with insufficient reliability of continuous transcription by marking low-confidence continuous segments by transcribing confidence data; it identifies text structures that affect the semantic continuity of business by marking inverted segments by word order structure identification; it identifies localized words and colloquial abbreviations by marking non-standard expression segments by multi-geographical expression index data; and it generates a standard judgment result by generating three types of mark sets, which can improve the accuracy of distinguishing between standard and non-standard expressions.
[0076] In one embodiment, step S40 above includes: S401, when the standardization determination result indicates that the preliminary text sequence belongs to the standard expression case, a pass-through processing mark is generated based on the standardization determination result; S402, based on the preliminary text sequence, perform statement boundary localization and business statement segment extraction to generate a standard candidate text segment set; S403, based on the pass-through processing marker, the standard candidate text segment set is subjected to original word order preservation and redundant colloquial text removal to generate a pass-through text sequence; S404, perform term consistency matching between the direct text sequence and the preset business terminology list to generate a term consistency tag set; S405, Based on the terminology consistency tag set, the expression form of business terms in the direct text sequence is unified to generate a standard direct text sequence, and the standard direct text sequence is used as standardized business text.
[0077] In this embodiment, after the standardization judgment result indicates that the preliminary text sequence belongs to the standard expression situation, a pass-through processing marker can be generated based on the standardization judgment result. The pass-through processing marker is used to identify that the preliminary text sequence has the conditions for direct processing and to control the subsequent text processing to maintain the original word order and avoid semantic rewriting of the preliminary text sequence. The pass-through processing marker may include a text source identifier, a judgment status identifier, and a pass-through processing status identifier, so that the same preliminary text sequence remains traceable in the subsequent processes of segment extraction, redundant content removal, and terminology unification.
[0078] Statement boundary localization is used to determine the boundary positions of different statement segments in the initial text sequence. Boundary positions can be determined based on pause markers, punctuation marks, business action phrases, business object phrases, and status description phrases. Business statement segment extraction is used to extract text segments with business meaning from the initial text sequence and filter out blank segments, repeated pause segments, and short response segments that have no business implications. After statement boundary localization and business statement segment extraction, a standard candidate text segment set is generated. The standard candidate text segment set retains the business semantic units in the initial text sequence and records the segment arrangement positions.
[0079] After the pass-through processing marker is applied to the standard candidate text fragment set, the original word order is preserved and redundant colloquial text is removed. Preserving the original word order maintains the arrangement relationship between business actions, business objects, and status descriptions in the standard candidate text fragment set. Redundant colloquial text can include tone pauses, repeated confirmation words, and conjunctions without business meaning. After these are removed, the remaining text fragments are combined into a pass-through text sequence according to the original arrangement relationship. Compared to the initial text sequence, the pass-through text sequence reduces invalid colloquial content without changing the business semantics and sentence structure.
[0080] A pre-defined business terminology list records standard terms, synonyms, and permitted business phrases within the same business meaning. When performing terminology consistency matching between the direct-access text sequence and the pre-defined business terminology list, a terminology consistency tag set can be determined based on term matching, phrase matching, and semantic similarity matching. The terminology consistency tag set records the positions of business terms requiring standardization, candidate standard terms, and matching status within the direct-access text sequence. Based on the terminology consistency tag set, the expressions of business terms in the direct-access text sequence can be unified to the standard expressions in the pre-defined business terminology list, generating a standard direct-access text sequence. The standard direct-access text sequence retains the semantic content that already meets the standard expression criteria in the initial text sequence and organizes the business terminology expressions into a unified text format, thus serving as standardized business text.
[0081] This embodiment generates pass-through processing markers based on standard judgment results, which can distinguish the pass-through processing status of the initial text sequence; by locating sentence boundaries and extracting business sentence fragments, it can extract text fragments with business meaning from the initial text sequence; by preserving the original word order and removing redundant colloquial text, it can reduce invalid colloquial content without changing the business semantic relationship; by unifying the expression form of business terms through a preset business terminology list and terminology consistency marker set, it can generate standardized business text with consistent text form, improving the stability of text processing under standard expression.
[0082] In one embodiment, step S50 above includes: S501, when the standardization determination result indicates that the preliminary text sequence belongs to the non-standard expression case, the preliminary text sequence is input into the preset semantic normalization layer to generate normalization processing tags; S502, based on the normalized processing marker, non-standard expression units are identified in the preliminary text sequence to generate a set of non-standard expression units; S503, based on the set of non-standard expression units, perform term index retrieval and semantic vector retrieval in the retrieval enhancement knowledge base to generate a set of candidate mapping terms; S504, Generate normalized prompt data based on the candidate mapping term set and the preliminary text sequence; S505, the normalized prompt data is input into a preset language model, and the preset language model performs business term replacement, semantic component completion and word order adjustment on the preliminary text sequence based on the normalized prompt data to generate candidate normalized text; S506, Based on the candidate mapping term set, perform mapping consistency verification on the candidate normalized text, and generate mapping consistency verification result; S507, Based on the mapping consistency verification result, perform term correction on the candidate normalized text to generate standardized business text.
[0083] In this embodiment, after the standardization judgment result indicates that the preliminary text sequence belongs to the non-standard expression category, the preliminary text sequence is sent to a preset semantic normalization layer. The semantic normalization layer can include a non-standard expression recognition unit, a retrieval organization unit, a prompt construction unit, a model rewriting unit, and a consistency verification unit. A normalization processing marker, generated from the standardization judgment result, is used to indicate that the preliminary text sequence needs semantic normalization and records the triggering reason, such as low-confidence continuous segments, inverted word order segments, or non-standard expression segments. The normalization processing marker can carry the segment position, trigger type, and processing range, enabling subsequent recognition processes to focus on processing text segments containing non-standard expressions.
[0084] Non-standard expression unit identification is based on normalized markers and a preliminary text sequence. Identification targets may include localized vocabulary, colloquial phrases, abbreviations, omitted semantic components, inverted sentences, and low-confidence text fragments. During identification, the preliminary text sequence is analyzed according to word segmentation, phrase extraction, word order relationship identification, and fragment confidence status. The identified non-standard expression fragments are then organized into a set of non-standard expression units. This set records the non-standard expression content, fragment location, expression type, and semantic category to be mapped.
[0085] The enhanced retrieval knowledge base stores the mapping between non-standard expressions and standard business terms, as well as synonyms, scenario phrases, business action phrases, and state description phrases. When performing term index retrieval based on the set of non-standard expression units, registered terms can be matched according to literal content and phrase structure. When performing semantic vector retrieval, non-standard expression units can be converted into semantic vectors, and semantically similar standard expressions can be retrieved. The results of term index retrieval and semantic vector retrieval are merged into a candidate mapping term set, which can contain non-standard expression fragments, candidate standard terms, semantic categories, matching sources, and matching states.
[0086] The normalized prompt data is obtained by combining a set of candidate mapping terms and a preliminary text sequence. The normalized prompt data may include the original preliminary text sequence, a set of non-standard expression units, candidate standard terms, semantic categories, word order requirements, and prohibited rewriting business objects. After receiving the normalized prompt data, the pre-defined language model performs business term replacement, semantic component completion, and word order adjustment on the preliminary text sequence. Business term replacement replaces localized words or colloquial phrases with candidate standard terms; semantic component completion completes omitted action objects, state descriptions, or business relationships; word order adjustment restructures inverted segments into a text structure with clear business semantics. After model rewriting, candidate normalized text is generated.
[0087] Mapping consistency verification is used to check the terminology mapping relationship between candidate normalized text and the candidate mapping terminology set. During verification, it compares whether business terms in the candidate normalized text originate from the candidate mapping terminology set, checks whether non-standard expression units have been replaced, and identifies whether there are terms in the candidate normalized text that are not supported by the candidate mapping terminology set. The mapping consistency verification results record consistent segments, segments to be corrected, and the basis for correction. After terminology correction of the candidate normalized text based on the mapping consistency verification results, standardized business text is obtained.
[0088] This embodiment uses normalization processing to locate non-standard expression fragments, which can reduce the participation of irrelevant text in semantic normalization; it uses a set of non-standard expression units and a retrieval-enhanced knowledge base for term index retrieval and semantic vector retrieval, which can provide clear standard terms and semantic categories for the preset language model; it uses the preset language model to perform business term replacement, semantic component completion, and word order adjustment, which can convert non-standard expressions into semantically clear candidate normalized text; and it reduces terminology deviations during model rewriting by verifying mapping consistency and correcting terms, making the generated standardized business text more stable.
[0089] In one embodiment, step S60 above includes: S601, The standardized business text is input into the intent classification engine, and the intent classification engine performs business action phrase recognition, business object phrase recognition and state description phrase recognition on the standardized business text to generate an intent element set; S602, Based on the set of intent elements, perform intent matching in the preset intent template data to generate a set of candidate intent tags; S603, Generate an intent classification result based on the candidate intent tag set and the intent element set; S604, Based on the intent classification result, determine the business processing node in the preset business node configuration data, and extract the node input parameter structure corresponding to the business processing node; S605, based on the standardized business text and the intent classification result, fill in the node input parameter structure to generate node call data; S606, Invoke the business processing node based on the node call data and obtain the node processing result; S607, perform processing status identification and response content extraction on the node processing result to generate node response content; S608, Generate business response data based on the node response content and the intent classification result.
[0090] In this embodiment, the standardized business text has undergone terminology standardization and semantic organization, and can be used as input content for the intent classification engine. After receiving the standardized business text, the intent classification engine can perform word segmentation, phrase boundary recognition, and semantic role labeling on the text to identify business action phrases, business object phrases, and state description phrases. Business action phrases are used to indicate the type of operation that the user wants to trigger, business object phrases are used to indicate the operation object, and state description phrases are used to indicate the current state, abnormal state, or service demand state of the object. These phrases are combined into an intent element set, which can record phrase content, phrase position, phrase category, and dependencies between phrases.
[0091] Intent template data is used to store action combinations, object combinations, state combinations, and slot structures corresponding to different intent categories. When performing intent matching based on the intent element set, business action phrases, business object phrases, and state description phrases can be matched against template fields in the intent template data to generate a candidate intent tag set. The candidate intent tag set may include candidate intent categories, hit elements, missing elements, and matching states. When generating intent classification results based on the candidate intent tag set and intent element set, the final intent category can be determined according to the coverage relationship between candidate intent categories and intent elements, and the business action, business object, and state description can be written as slot data into the intent classification result.
[0092] The business node configuration data records the mapping relationship between intent categories and business processing nodes, and records the node input parameter structure corresponding to each business processing node. When determining a business processing node based on the intent classification result, the node identifier can be queried in the business node configuration data according to the intent category, and the node input parameter structure bound to the node identifier can be extracted. The node input parameter structure may include an action field, an object field, a status field, a time field, and supplementary fields. When populating the node input parameter structure based on standardized business text and intent classification results, the slot data in the intent classification results can be written into the corresponding fields, and supplementary content can be extracted from the standardized business text to generate node call data.
[0093] Node call data is used to submit structured requests to business processing nodes. After receiving the node call data, the business processing node returns the node processing result. The node processing result may include processing status, business return content, exception prompts, fields to be supplemented, and prompts for the next interaction. When identifying the processing status of the node processing result, it is possible to identify statuses such as successful processing, fields to be supplemented, inability to process, or needing to be transferred; when extracting response content, feedback content, supplementary prompts, and status descriptions can be extracted from the node processing result to generate node response content. When generating business response data based on node response content and intent classification results, the response content can be combined with intent categories, slot data, and processing status to ensure that the business response data reflects user intent, node return content, and subsequent interaction requirements.
[0094] This embodiment uses an intent classification engine to identify business action phrases, business object phrases, and status description phrases in standardized business text, converting the text content into a set of matchable intent elements. By matching the set of intent elements with preset intent template data, intent classification results containing intent categories and slot data can be generated. The intent classification results are used to determine business processing nodes in preset business node configuration data and populate the node input parameter structure, ensuring that the node call data is consistent with user requests. By identifying the processing status of the node processing results and extracting the response content, the content returned by the business processing nodes can be organized into business response data, thereby reducing the matching deviation between intent classification results and business processing nodes.
[0095] In one embodiment, in a car insurance underwriting and claims service scenario, a user initiates a claim through the insurance institution's voice customer service portal. During the call, the user speaks with a Cantonese accent and describes the accident using natural spoken language. The user says their car was hit, what to do, the car is currently stuck on the roadside, whether third-party liability insurance will cover it, when the claims adjuster will arrive, and also wants to inquire about the claims progress. After receiving the voice interaction request, the voice portal extracts the continuous voice stream, conversation voice identifier, and voice source channel identifier from the request. Based on the conversation voice identifier, multiple voice segments within the same call are grouped into the same business interaction. Based on the voice source channel identifier, the voice is confirmed to originate from the telephone customer service channel. The telephone channel contains prompts, agent transfer prompts, and ambient human voices. The platform performs voice boundary analysis on the grouped candidate voice segment set to obtain the start and end positions and pause intervals of the user's voice. It also identifies non-user voice content by voice source identification and then filters out the user's voice segments from the candidate voice segment set to generate the user's voice sequence.
[0096] The user's voice sequence is divided into multiple pronunciation segments, such as "车撞咗" (the car crashed), "点样搞" (how to do), "而家趴窝" (now broken down), "三者险" (third-party liability insurance), "定损师傅" (damage assessor), "赔付进度" (claim progress). The platform performs syllable boundary marking and tone change position marking on the pronunciation segments to generate a syllable marking sequence, and extracts initial consonant pronunciation offset features, final vowel pronunciation offset features, tone trend features, and prosodic pause features based on a preset benchmark pronunciation template set. Due to the Cantonese tone and final sound extension features in the pronunciation of words such as "点样", "而家", "趴窝", etc., after the candidate pronunciation feature set is matched with the preset geopolitical acoustic template set at the segment level, multiple segment geopolitical markers with Cantonese geopolitical features are generated. The platform continuously aggregates the segment geopolitical marker set in the order of the pronunciation segment sequence to generate pronunciation fingerprint features, and matches the pronunciation fingerprint features with the geopolitical type index table, and determines the matched geopolitical type identifier as the voice geopolitical type.
[0097] The platform matches the acoustic transcription model identifier and the geopolitical pronunciation parameter package from the preset acoustic model index data according to the voice geopolitical type. The geopolitical pronunciation parameter package contains business word pronunciation weight data for the car insurance underwriting and claim settlement scenarios, such as "报案" (report a case), "查勘" (survey), "定损" (damage assessment), "三者险" (third-party liability insurance), "道路救援" (road rescue), "进厂跟踪" (factory entry tracking), "单证上传" (document upload), "赔付进度" (claim progress). The platform calls the acoustic transcription model based on the acoustic transcription model identifier and loads the geopolitical pronunciation parameter package into the acoustic transcription model to generate a geopolitical transcription processing instance. After the user's voice sequence is input into the geopolitical transcription processing instance, the geopolitical transcription processing instance performs geopolitical pronunciation enhancement encoding on the voice frames to generate a geopolitical acoustic representation sequence. After the geopolitical acoustic representation sequence undergoes candidate word decoding, a candidate word path set and a path confidence distribution are formed. For phrases such as "点样搞", "趴窝", "车撞咗", etc., there may be different candidate results in the candidate word path set. For example, "点样搞" corresponds to "怎么办" (how to do) or "如何处理" (how to handle), "趴窝" corresponds to "车辆无法行驶" (the vehicle cannot move), and "车撞咗" corresponds to "车辆发生碰撞事故" (a collision accident occurred to the vehicle). Due to the geopolitical pronunciation parameter package increasing the recognition weight of underwriting and claim settlement business words, the platform rearranges the paths in the candidate word path set, determines the path that is more matched with "道路救援", "三者险", "定损", "赔付进度" as the target word path, and converts the target word path into a preliminary text sequence. At the same time, the path confidence value and the word segment confidence value are extracted from the path confidence distribution to generate transcription confidence data.
[0098] The platform divides the statement fragments based on the preliminary text sequence to generate a set of text fragments to be judged. For content such as "The car crashed. What should I do? The car is now stalled by the roadside. Is there any compensation under the third-party liability insurance? When will the damage assessor come? I also want to ask about the compensation progress", the platform marks the low-confidence continuous fragments with a confidence data marker lower than the preset confidence value to generate a confidence judgment marker set; based on the set of text fragments to be judged, it identifies the word order structure, and marks natural oral inversion content such as "The car is now stalled by the roadside. When will the damage assessor come?" as word order inversion fragments to generate a word order structure marker set; based on the set of text fragments to be judged, it performs expression hit recognition in the preset multi-regional expression index data, and marks content such as "What should I do?", "now", "stalled", "when will... come" as non-standard expression fragments to generate a non-standard expression marker set. When the confidence judgment marker set contains low-confidence continuous fragment markers, or the word order structure marker set contains word order inversion fragment markers, or the non-standard expression marker set contains non-standard expression fragment markers, the platform generates a standardness judgment result indicating that the preliminary text sequence belongs to a non-standard expression situation.
[0099] In the case of non-standard expressions, the preliminary text sequence is input into the preset semantic normalization layer. The platform generates a normalization processing marker and identifies non-standard expression units based on the normalization processing marker. The non-standard expression unit set includes expressions such as "What should I do?", "now", "stalled", "when will... come", "Is there any compensation". The platform performs entry index retrieval and semantic vector retrieval in the retrieval-enhanced knowledge base based on the non-standard expression unit set to generate a candidate mapping entry set. The retrieval-enhanced knowledge base maps "What should I do?" to "Request processing suggestions", "now" to "current", "stalled" to "The vehicle is inoperable", "The car crashed" to "The vehicle has a collision accident", "Is there any compensation under the third-party liability insurance" to "Consultation on the scope of compensation for the third-party liability insurance", and "When will the damage assessor come" to "Query the arrival time of the survey and damage assessment personnel". The platform generates normalization prompt data based on the candidate mapping entry set and the preliminary text sequence, and inputs the normalization prompt data into the preset language model. The preset language model performs business term replacement, semantic component completion, and word order regularization on the preliminary text sequence based on the normalization prompt data to generate candidate normalized texts, such as "The vehicle has a collision accident. The vehicle is currently inoperable and located by the roadside. Request road rescue. Consult the scope of compensation for the third-party liability insurance, the arrival time of the survey and damage assessment personnel, and the compensation progress". The platform then performs mapping consistency verification on the candidate normalized texts based on the candidate mapping entry set, confirms that terms such as "The vehicle is inoperable", "road rescue", "third-party liability insurance", "survey and damage assessment", "compensation progress", etc. all come from the candidate mapping entry set, and performs term correction on the candidate normalized texts based on the mapping consistency verification result to generate standardized business texts.
[0100] In the same call, the user later added that they wanted to check the policy claim progress and see if the loss assessment had been completed. The initial text sequence generated after acoustic transcription of this speech segment showed stable confidence, complete sentence structure, and did not match any non-standard expression segments in the multi-geographical expression index data. The platform generated a standard judgment result indicating that the initial text sequence belonged to the standard expression category, and generated a pass-through processing tag based on the standard judgment result. The platform performed sentence boundary localization and business sentence segment extraction on the initial text sequence, generating a set of standard candidate text segments, retaining the two business sentence segments: querying the policy claim progress and confirming the loss assessment status. Based on the pass-through processing tag, the platform maintained the original word order and removed redundant colloquialisms such as "I want to," "just a moment," and "let's see," generating a pass-through text sequence. The pass-through text sequence was matched with a preset business terminology list for terminology consistency, generating a set of terminology consistency tags. The platform unified expressions such as claim progress and whether the loss assessment was completed into "claim progress query" and "loss assessment status query," generating a standard pass-through text sequence, which was then used as the standardized business text.
[0101] After standardized business text is input into the intent classification engine, the engine identifies business action phrases, business object phrases, and state description phrases, generating an intent element set. For example, a standardized business text might include phrases like "vehicle collision occurred, vehicle is currently unable to move, requesting roadside assistance, inquiring about third-party liability insurance coverage, arrival time of the claims adjuster, and claims progress." Business action phrases include "request," "inquire," and "query." Business object phrases include "vehicle," "roadside assistance," "third-party liability insurance," "claims adjuster," and "claims progress." State description phrases include "collision occurred," "currently unable to move," and "located on the roadside." The platform performs intent matching within a pre-set intent template data based on the intent element set, generating a candidate intent tag set. The candidate intent tag set is then combined with the intent element set to generate intent classification results, such as "roadside assistance dispatch," "third-party liability insurance claims inquiry," "claims adjustment progress inquiry," and "claims progress inquiry."
[0102] Based on intent classification results, the platform determines business processing nodes from pre-defined business node configuration data and extracts the corresponding node input parameter structures. For example, roadside assistance dispatch corresponds to a roadside assistance business processing node, with input parameters including accident location, vehicle status, contact number, and assistance type. Third-party liability insurance claims consultation corresponds to an insurance type claims consultation node, with input parameters including insurance type, accident liability information, and case identifier. Damage assessment progress query corresponds to a damage assessment status query node, with input parameters including report number, vehicle information, and investigation status fields. Claims progress query corresponds to a claims progress query node, with input parameters including policy number, case number, and claims stage fields. The platform fills the node input parameter structures with standardized business text and intent classification results to generate node call data. After the node call data calls the corresponding business processing node, the platform obtains the node processing result, identifies the processing status, extracts response content, and generates node response content. The node processing results may include information such as roadside assistance accepted, claims assessment pending dispatch, third-party liability insurance requiring verification based on accident liability determination, and claims cases in the document review stage. The platform generates business response data based on the node response content and intent classification results. For example, it may indicate that a roadside assistance dispatch has been initiated for you, the current case is in the document review stage, claims assessment personnel are pending dispatch, and third-party liability insurance coverage requires verification based on accident liability determination and policy liability confirmation; please provide the accident location or confirm your current location.
[0103] In one embodiment, a service response device based on voice geographic type is provided, which corresponds one-to-one with the service response method based on voice geographic type in the above embodiments. (Refer to...) Figure 3 , Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the service response device based on voice geography of the present invention. The module includes a voice geography recognition module 10, an acoustic transcription processing module 20, a representation standardization determination module 30, a direct text generation module 40, a semantic normalization processing module 50, and an intent scheduling response module 60. Detailed descriptions of each functional module are as follows: The voice geography recognition module 10 is used to acquire a user's voice sequence, extract the pronunciation fingerprint features of the user's voice sequence, and determine the voice geography type based on the pronunciation fingerprint features; The acoustic transcription processing module 20 is used to call the acoustic transcription model corresponding to the speech geography type, input the user speech sequence into the acoustic transcription model for speech transcription, and generate a preliminary text sequence and transcription confidence data; The expression standardization determination module 30 is used to determine the expression standardization based on the preliminary text sequence and the transcription confidence data, and generate a standardization determination result; The direct text generation module 40 is used to determine the preliminary text sequence as standardized business text when the standardization determination result indicates that the preliminary text sequence belongs to the standard expression situation; The semantic normalization processing module 50 is used to input the preliminary text sequence into a preset semantic normalization layer when the standardization judgment result indicates that the preliminary text sequence belongs to a non-standard expression case, and perform semantic normalization processing on the preliminary text sequence based on the retrieval enhancement knowledge base and the preset language model to generate standardized business text. The intent scheduling and response module 60 is used to input the standardized business text into the intent classification engine for intent classification, generate intent classification results, determine business processing nodes based on the intent classification results, call the business processing nodes to obtain node processing results, and generate business response data based on the node processing results.
[0104] In one embodiment, the voice geography recognition module 10 is specifically used for: Receive a voice interaction request, and extract the continuous voice stream, the conversation voice identifier, and the voice source channel identifier from the voice interaction request; Based on the conversation speech identifier and the speech source channel identifier, the continuous speech stream is segmented to generate a candidate speech segment set; The candidate speech segment set is parsed to generate speech start and end positions and pause interval positions; The candidate speech segment set is subjected to speech source identification to generate non-user speech tags; Based on the start and end positions of the speech, the pause interval positions, and the non-user voice markers, user voice segments are selected from the candidate speech segment set to generate a user speech sequence; Based on the user's speech sequence, pronunciation segments are divided to generate a pronunciation segment sequence; Each pronunciation segment in the pronunciation segment sequence is marked with syllable boundaries and tone change positions to generate a syllable marking sequence; Based on the syllable marker sequence and the preset reference pronunciation template set, extract the initial consonant pronunciation offset features, final vowel pronunciation offset features, tone direction features and prosodic pause features to generate a candidate pronunciation feature set; The candidate pronunciation feature set is matched with a preset geoacoustic template set at the segment level to generate a segment geospatial tag set; Based on the arrangement order of the pronunciation segment sequence, the set of geographical markers of the segments is continuously aggregated to generate pronunciation fingerprint features; The pronunciation fingerprint features are matched with the geopolitical type index table, and the matched geopolitical type identifier is determined as the pronunciation geopolitical type.
[0105] In one embodiment, the acoustic transcription processing module 20 is specifically used for: Based on the speech geography type, the acoustic transcription model identifier and the geography pronunciation parameter package containing business word pronunciation weight data are matched from the preset acoustic model index data; Based on the acoustic transcription model identifier, the acoustic transcription model is invoked, and the geopolitical pronunciation parameter package is loaded into the acoustic transcription model to generate a geopolitical transcription processing instance; The user's speech sequence is input into the geopolitical transcription processing instance for geopolitical pronunciation enhancement encoding to generate a geopolitical acoustic representation sequence; Based on the geoacoustic representation sequence, candidate words are decoded to generate a candidate word path set and a path confidence distribution; Based on the business word pronunciation weight data in the geo-pronunciation parameter package, the candidate word path set is rearranged to generate the target word path; Convert the target word path into a preliminary text sequence; Extract the path confidence value and word fragment value corresponding to the target word path from the path confidence distribution.
[0106] In one embodiment, the statement standardization determination module 30 is specifically used for: Based on the preliminary text sequence, sentence segments are divided to generate a set of text segments to be judged; Based on the transcription confidence data, low-confidence continuous segments in the set of text segments to be judged that are lower than the preset confidence value are marked to generate a confidence judgment mark set; Based on the set of text fragments to be determined, word order structure is identified, and fragments with inverted word order are marked to generate a set of word order structure tags; Based on the set of text fragments to be judged, expression hit identification is performed in the preset multi-geographic expression index data, and non-standard expression fragments are marked to generate a non-standard expression tag set; When the confidence determination tag set does not contain low-confidence continuous segment tags, the word order structure tag set does not contain word order inversion segment tags, and the non-standard expression tag set does not contain non-standard expression segment tags, a standard determination result indicating that the preliminary text sequence belongs to the standard expression case is generated. When the set of confidence determination markers contains low-confidence continuous segment markers, the set of word order structure markers contains word order inversion segment markers, or the set of non-standard expression markers contains non-standard expression segment markers, a standard determination result indicating that the preliminary text sequence belongs to a non-standard expression is generated.
[0107] In one embodiment, the direct text generation module 40 is specifically used for: When the standardization determination result indicates that the preliminary text sequence belongs to the standard expression case, a pass-through processing mark is generated based on the standardization determination result; Based on the preliminary text sequence, sentence boundary location and business sentence fragment extraction are performed to generate a standard candidate text fragment set; Based on the pass-through processing markers, the original word order is preserved and redundant colloquial text is removed from the standard candidate text fragment set to generate a pass-through text sequence. The direct text sequence is matched with a preset business terminology list to generate a terminology consistency tag set. Based on the terminology consistency tag set, the expression forms of business terms in the direct text sequence are unified to generate a standard direct text sequence, which is then used as standardized business text.
[0108] In one embodiment, the semantic normalization processing module 50 is specifically used for: When the standardization determination result indicates that the preliminary text sequence belongs to a non-standard expression, the preliminary text sequence is input into a preset semantic normalization layer to generate a normalization processing label. Based on the normalized processing markers, non-standard expression units are identified in the preliminary text sequence to generate a set of non-standard expression units; Based on the set of non-standard expression units, term index retrieval and semantic vector retrieval are performed in the enhanced retrieval knowledge base to generate a set of candidate mapped terms. Normalized prompt data is generated based on the candidate mapping term set and the preliminary text sequence; The normalized prompt data is input into a preset language model, and the preset language model performs business term replacement, semantic component completion and word order adjustment on the preliminary text sequence based on the normalized prompt data to generate candidate normalized text; Based on the candidate mapping term set, the candidate normalized text is subjected to mapping consistency verification to generate mapping consistency verification results. Based on the mapping consistency verification results, the candidate normalized text is corrected for terms to generate standardized business text.
[0109] In one embodiment, the intent scheduling response module 60 is specifically used for: The standardized business text is input into the intent classification engine, which then performs business action phrase recognition, business object phrase recognition, and state description phrase recognition on the standardized business text to generate a set of intent elements. Based on the set of intent elements, intent matching is performed in the preset intent template data to generate a set of candidate intent tags; An intent classification result is generated based on the candidate intent tag set and the intent element set; Based on the intent classification results, the business processing node is determined in the preset business node configuration data, and the node input parameter structure corresponding to the business processing node is extracted. Based on the standardized business text and the intent classification result, the node input parameter structure is filled in to generate node call data; Based on the node call data, the business processing node is invoked to obtain the node processing result; The processing status of the node processing results is identified and the response content is extracted to generate node response content; Business response data is generated based on the node response content and the intent classification results.
[0110] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When executed by the processor, the computer program implements server-side functions or steps of a voice-based geographic type-based service response method.
[0111] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When executed by the processor, the computer program implements client-side functions or steps of a voice-based geographic type-based service response method.
[0112] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: Obtain a user's speech sequence, extract the pronunciation fingerprint features of the user's speech sequence, and determine the speech geography type based on the pronunciation fingerprint features; The acoustic transcription model corresponding to the speech geography type is invoked, and the user speech sequence is input into the acoustic transcription model for speech transcription to generate a preliminary text sequence and transcription confidence data; Based on the preliminary text sequence and the transcription confidence data, the expression standardization is determined, and a standardization determination result is generated; When the standardization determination result indicates that the preliminary text sequence belongs to the standard expression situation, the preliminary text sequence is determined as standardized business text; When the standardization determination result indicates that the preliminary text sequence belongs to the non-standard expression case, the preliminary text sequence is input into the preset semantic normalization layer, and the semantic normalization processing of the preliminary text sequence is performed based on the retrieval enhancement knowledge base and the preset language model to generate standardized business text; The standardized business text is input into the intent classification engine for intent classification, and intent classification results are generated. Based on the intent classification results, business processing nodes are determined, the business processing nodes are called to obtain node processing results, and business response data is generated based on the node processing results.
[0113] In one embodiment, a computer-readable storage medium is provided, which may be non-volatile or volatile, and a computer program is stored thereon, which, when executed by a processor, performs the following steps: Obtain a user's speech sequence, extract the pronunciation fingerprint features of the user's speech sequence, and determine the speech geography type based on the pronunciation fingerprint features; The acoustic transcription model corresponding to the speech geography type is invoked, and the user speech sequence is input into the acoustic transcription model for speech transcription to generate a preliminary text sequence and transcription confidence data; Based on the preliminary text sequence and the transcription confidence data, the expression standardization is determined, and a standardization determination result is generated; When the standardization determination result indicates that the preliminary text sequence belongs to the standard expression situation, the preliminary text sequence is determined as standardized business text; When the standardization determination result indicates that the preliminary text sequence belongs to the non-standard expression case, the preliminary text sequence is input into the preset semantic normalization layer, and the semantic normalization processing of the preliminary text sequence is performed based on the retrieval enhancement knowledge base and the preset language model to generate standardized business text; The standardized business text is input into the intent classification engine for intent classification, and intent classification results are generated. Based on the intent classification results, business processing nodes are determined, the business processing nodes are called to obtain node processing results, and business response data is generated based on the node processing results.
[0114] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0115] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0116] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0117] It should be noted that any AI models, software tools, or components not belonging to this company appearing in the embodiments of this application are merely illustrative examples and do not represent actual use. All user personal information involved in the embodiments of this application has been authorized (with the knowledge and consent) by the relevant parties or has been fully authorized by all parties, and the executing entity may obtain it through various legal and compliant means. The collection, storage, use, processing, transmission, provision, and disclosure of the information, data, and signals involved all comply with relevant laws and regulations and do not violate public order and good morals.
[0118] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A service response method based on voice geographic type, characterized in that, Includes the following steps: Obtain a user's speech sequence, extract the pronunciation fingerprint features of the user's speech sequence, and determine the speech geography type based on the pronunciation fingerprint features; The acoustic transcription model corresponding to the speech geography type is invoked, and the user speech sequence is input into the acoustic transcription model for speech transcription to generate a preliminary text sequence and transcription confidence data; Based on the preliminary text sequence and the transcription confidence data, the expression standardization is determined, and a standardization determination result is generated; When the standardization determination result indicates that the preliminary text sequence belongs to the standard expression situation, the preliminary text sequence is determined as standardized business text; When the standardization determination result indicates that the preliminary text sequence belongs to a non-standard expression, the preliminary text sequence is input into a preset semantic normalization layer, and the preliminary text sequence is semantically normalized based on the retrieval enhancement knowledge base and the preset language model to generate standardized business text. The standardized business text is input into the intent classification engine for intent classification, and intent classification results are generated. Based on the intent classification results, business processing nodes are determined, the business processing nodes are called to obtain node processing results, and business response data is generated based on the node processing results.
2. The service response method based on voice geography as described in claim 1, characterized in that, Acquire a user's speech sequence, extract the pronunciation fingerprint features of the user's speech sequence, and determine the speech geography type based on the pronunciation fingerprint features, including: Receive a voice interaction request, and extract the continuous voice stream, the conversation voice identifier, and the voice source channel identifier from the voice interaction request; Based on the conversation speech identifier and the speech source channel identifier, the continuous speech stream is segmented to generate a candidate speech segment set; The candidate speech segment set is parsed to generate speech start and end positions and pause interval positions; The candidate speech segment set is subjected to speech source identification to generate non-user speech tags; Based on the start and end positions of the speech, the pause interval positions, and the non-user voice markers, user voice segments are selected from the candidate speech segment set to generate a user speech sequence; Based on the user's speech sequence, pronunciation segments are divided to generate a pronunciation segment sequence; Each pronunciation segment in the pronunciation segment sequence is marked with syllable boundaries and tone change positions to generate a syllable marking sequence; Based on the syllable marker sequence and the preset reference pronunciation template set, extract the initial consonant pronunciation offset features, final vowel pronunciation offset features, tone direction features and prosodic pause features to generate a candidate pronunciation feature set; The candidate pronunciation feature set is matched with a preset geoacoustic template set at the segment level to generate a segment geospatial tag set; Based on the arrangement order of the pronunciation segment sequence, the set of geographical markers of the segments is continuously aggregated to generate pronunciation fingerprint features; The pronunciation fingerprint features are matched with the geopolitical type index table, and the matched geopolitical type identifier is determined as the pronunciation geopolitical type.
3. The service response method based on voice geography as described in claim 1, characterized in that, The acoustic transcription model corresponding to the stated speech geography is invoked, and the user's speech sequence is input into the acoustic transcription model for speech transcription, generating a preliminary text sequence and transcription confidence data, including: Based on the speech geography type, the acoustic transcription model identifier and the geography pronunciation parameter package containing business word pronunciation weight data are matched from the preset acoustic model index data; Based on the acoustic transcription model identifier, the acoustic transcription model is invoked, and the geopolitical pronunciation parameter package is loaded into the acoustic transcription model to generate a geopolitical transcription processing instance; The user's speech sequence is input into the geopolitical transcription processing instance for geopolitical pronunciation enhancement encoding to generate a geopolitical acoustic representation sequence; Based on the geoacoustic representation sequence, candidate words are decoded to generate a candidate word path set and a path confidence distribution; Based on the business word pronunciation weight data in the geo-pronunciation parameter package, the candidate word path set is rearranged to generate the target word path; Convert the target word path into a preliminary text sequence; Extract the path confidence value and word fragment confidence value corresponding to the target word path from the path confidence distribution, and generate transcription confidence data based on the path confidence value and the word fragment confidence value.
4. The service response method based on voice geography as described in claim 1, characterized in that, Based on the preliminary text sequence and the transcription confidence data, a standardization determination is performed to generate a standardization determination result, including: Based on the preliminary text sequence, sentence segments are divided to generate a set of text segments to be judged; Based on the transcription confidence data, low-confidence continuous segments in the set of text segments to be judged that are lower than the preset confidence value are marked to generate a confidence judgment mark set; Based on the set of text fragments to be determined, word order structure is identified, and fragments with inverted word order are marked to generate a set of word order structure tags; Based on the set of text fragments to be judged, expression hit identification is performed in the preset multi-geographic expression index data, and non-standard expression fragments are marked to generate a non-standard expression tag set; When the confidence determination tag set does not contain low-confidence continuous segment tags, the word order structure tag set does not contain word order inversion segment tags, and the non-standard expression tag set does not contain non-standard expression segment tags, a standard determination result indicating that the preliminary text sequence belongs to the standard expression case is generated. When the set of confidence determination markers contains low-confidence continuous segment markers, the set of word order structure markers contains word order inversion segment markers, or the set of non-standard expression markers contains non-standard expression segment markers, a standard determination result indicating that the preliminary text sequence belongs to a non-standard expression is generated.
5. The service response method based on voice geography as described in claim 1, characterized in that, When the standardization determination result indicates that the preliminary text sequence belongs to the standard expression situation, the preliminary text sequence is determined as standardized business text, including: When the standardization determination result indicates that the preliminary text sequence belongs to the standard expression case, a pass-through processing mark is generated based on the standardization determination result; Based on the preliminary text sequence, sentence boundary location and business sentence fragment extraction are performed to generate a standard candidate text fragment set; Based on the pass-through processing markers, the original word order is preserved and redundant colloquial text is removed from the standard candidate text fragment set to generate a pass-through text sequence. The direct text sequence is matched with a preset business terminology list to generate a terminology consistency tag set. Based on the terminology consistency tag set, the expression forms of business terms in the direct text sequence are unified to generate a standard direct text sequence, which is then used as standardized business text.
6. The service response method based on voice geography as described in claim 1, characterized in that, When the standardization determination result indicates that the preliminary text sequence belongs to a non-standard expression, the preliminary text sequence is input into a preset semantic normalization layer. Based on the retrieval enhancement knowledge base and a preset language model, the preliminary text sequence is semantically normalized to generate standardized business text, including: When the standardization determination result indicates that the preliminary text sequence belongs to a non-standard expression, the preliminary text sequence is input into a preset semantic normalization layer to generate a normalization processing label. Based on the normalized processing markers, non-standard expression units are identified in the preliminary text sequence to generate a set of non-standard expression units; Based on the set of non-standard expression units, term index retrieval and semantic vector retrieval are performed in the enhanced retrieval knowledge base to generate a set of candidate mapped terms. Normalized prompt data is generated based on the candidate mapping term set and the preliminary text sequence; The normalized prompt data is input into a preset language model, and the preset language model performs business term replacement, semantic component completion and word order adjustment on the preliminary text sequence based on the normalized prompt data to generate candidate normalized text; Based on the candidate mapping term set, the candidate normalized text is subjected to mapping consistency verification to generate mapping consistency verification results. Based on the mapping consistency verification results, the candidate normalized text is corrected for terms to generate standardized business text.
7. The service response method based on voice geography as described in claim 1, characterized in that, The standardized business text is input into an intent classification engine for intent classification, generating intent classification results. Based on the intent classification results, business processing nodes are determined, and the processing results of these nodes are obtained. Based on the processing results, business response data is generated, including: The standardized business text is input into the intent classification engine, which then performs business action phrase recognition, business object phrase recognition, and state description phrase recognition on the standardized business text to generate a set of intent elements. Based on the set of intent elements, intent matching is performed in the preset intent template data to generate a set of candidate intent tags; An intent classification result is generated based on the candidate intent tag set and the intent element set; Based on the intent classification results, the business processing node is determined in the preset business node configuration data, and the node input parameter structure corresponding to the business processing node is extracted. Based on the standardized business text and the intent classification result, the node input parameter structure is filled in to generate node call data; Based on the node call data, the business processing node is invoked to obtain the node processing result; The processing status of the node processing results is identified and the response content is extracted to generate node response content; Business response data is generated based on the node response content and the intent classification results.
8. A service response device based on voice geographic type, characterized in that, The service response device based on voice geography includes: The voice geography recognition module is used to acquire user voice sequences, extract pronunciation fingerprint features of the user voice sequences, and determine the voice geography type based on the pronunciation fingerprint features; The acoustic transcription processing module is used to call the acoustic transcription model corresponding to the speech geography type, input the user speech sequence into the acoustic transcription model for speech transcription, and generate a preliminary text sequence and transcription confidence data; The expression standardization determination module is used to determine the expression standardization based on the preliminary text sequence and the transcription confidence data, and generate a standardization determination result; The direct text generation module is used to determine the preliminary text sequence as standardized business text when the standardization determination result indicates that the preliminary text sequence belongs to the standard expression situation; The semantic normalization processing module is used to input the preliminary text sequence into a preset semantic normalization layer when the standardization judgment result indicates that the preliminary text sequence belongs to a non-standard expression case, and perform semantic normalization processing on the preliminary text sequence based on the retrieval enhancement knowledge base and the preset language model to generate standardized business text. The intent scheduling and response module is used to input the standardized business text into the intent classification engine for intent classification, generate intent classification results, determine business processing nodes based on the intent classification results, call the business processing nodes to obtain node processing results, and generate business response data based on the node processing results.
9. A computer device, characterized in that, The computer device includes a memory, a processor, and a voice-based geopolitical type-based service response program stored in the memory and executable on the processor. When executed by the processor, the voice-based geopolitical type-based service response program implements the steps of the voice-based geopolitical type-based service response method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The storage medium stores a service response program based on voice geography type, which, when executed by a processor, implements the steps of the service response method based on voice geography type as described in any one of claims 1-7.