Alarm receiving and handling foreign language virtual simultaneous transmission method and system based on intelligent voice interaction service

By collecting, converting, identifying and splitting voice signals, real-time and efficient translation is achieved when foreign personnel report an alarm, solving the delay in handling alarms caused by language barriers and improving the real-time and accuracy of translation.

CN120708616AInactive Publication Date: 2025-09-26XINZHI DAOSHU (SHANGHAI) TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202511220420.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2025-09-26
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In existing technologies, when foreign personnel report a crime, communication barriers caused by language barriers lead to delays in handling the case, and manual translation is unstable, making it difficult to achieve all-weather coverage and efficient translation.

Method used

By collecting voice signals from the alarm and receiving ends, converting them into electrical signals, identifying the language type, performing text conversion and translation, generating voice packets, building a data transmission link, and splitting it into independent call links, recording and synthesizing voice strips to achieve real-time translation.

Benefits of technology

It has improved the real-time and accuracy of language translation, increased the efficiency of handling foreign-related police incidents, and solved the problem of communication delays caused by language barriers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708616A_ABST
    Figure CN120708616A_ABST
Patent Text Reader

Abstract

The invention discloses an alarm receiving and handling foreign language virtual simultaneous transmission method and system based on intelligent voice interaction service, and relates to the technical field of voice interaction, and the method comprises the steps: converting alarm input voice and alarm receiving input voice into an alarm electric signal and an alarm receiving electric signal; extracting a language type to obtain an alarm input character and an alarm receiving input character; extracting language data to obtain alarm receiving output characters and alarm output characters; converting the alarm receiving output character to obtain an alarm receiving output voice, and converting the alarm output character to obtain an alarm output voice; constructing a data transmission link, and protecting communication data generated in each stage; splitting an original call into an alarm person-virtual simultaneous alarm receiver call link and an alarm receiver-virtual simultaneous alarm person call link; and recording and synthesizing the alarm input voice, the alarm receiving input voice, the alarm receiving output voice and the alarm output voice, and generating four sound buttons. The method has the effect of improving the real-time performance, the accuracy and the processing efficiency of language translation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of voice interaction, and in particular to a method and system for virtual simultaneous interpretation in foreign languages ​​for receiving and handling police calls based on intelligent voice interaction services. Background Art

[0002] As the process of globalization accelerates, the number of foreigners in cities continues to rise, and the demand for alarms in different languages ​​has exploded. The existing traditional model can no longer meet the needs of efficiently and accurately handling foreign-related police incidents and standardizing call record management.

[0003] In current police response and handling, when foreign nationals report a crime, language barriers often create significant communication barriers between the responding officer and the caller. Traditionally, this approach relies on calling a third party, seeking a foreign language translator to participate in a three-way call to facilitate communication. However, this approach has significant drawbacks: Finding a suitable foreign language translator is time-consuming and laborious, easily delaying the golden opportunity for handling the crime; the limited number of foreign language call-handling volunteers makes it difficult to provide uninterrupted service around the clock, and the language coverage is limited by the volunteers' abilities; and human translators are easily affected by factors such as fatigue and emotion, resulting in unstable translation accuracy and efficiency. Summary of the Invention

[0004] The purpose of the present invention is to provide a method and system for virtual simultaneous interpretation in foreign languages ​​for receiving and handling police calls based on intelligent voice interaction services, so as to solve the problems raised in the above-mentioned background technology.

[0005] In a first aspect, the present application provides a method for virtual simultaneous interpretation of foreign languages ​​for receiving and handling police calls based on an intelligent voice interaction service, the method comprising: During the alarm receiving process, the alarm input voice of the alarm end and the alarm receiving voice of the alarm receiving end are collected respectively, and the alarm input voice and the alarm receiving voice are converted into an alarm electric signal and an alarm receiving electric signal respectively; Extracting the language types of the alarm signal and the alarm receiving signal, and selecting different preset recognition models according to the language types to perform text conversion to obtain alarm input text and alarm receiving input text; Extracting language data of the alarm input text and the alarm receiving input text, identifying the language data to obtain a translation rule, and translating the alarm input text and the alarm receiving input text into each other according to the translation rule to obtain an alarm receiving output text and an alarm output text; Extracting timbre data, pitch data, and speech rate data from the alarm input voice and the alarm receiving input voice, generating an alarm person voice package and an alarm receiving person voice package, converting the alarm receiving output text according to the alarm person voice package to obtain an alarm receiving output voice, and converting the alarm output text according to the alarm receiving person voice package to obtain an alarm output voice; Build a data transmission link, protect the communication data generated in each stage, and send it to the next stage; Split the original call into two independent call links, extract the data transmission information of the data transmission link, and reconnect the call links according to the data transmission information to obtain the call link between the alarmer and the virtual simultaneous caller and the call link between the caller and the virtual simultaneous caller; The alarm input voice, the alarm input voice, the alarm output voice and the alarm output voice are recorded to generate four voice strips, the four voice strips are synthesized according to the time axis, and four sound buttons are generated to control the sound switch state of each voice strip.

[0006] Preferably, the step of collecting the alarm input voice of the alarm terminal and the alarm receiving input voice of the alarm receiving terminal, and converting the alarm input voice and the alarm receiving input voice into the alarm electric signal and the alarm receiving electric signal respectively, is specifically: During the alarm reception process, the alarm terminal collects the voice of the alarm caller to obtain the alarm input voice, and the alarm receiving terminal collects the voice of the alarm receiver to obtain the alarm input voice; Extracting the alarm audio track data and the alarm receiving audio track data of the alarm input voice and the alarm receiving input voice respectively; Obtaining alarm sound wave data and alarm receiving sound wave data according to the alarm audio track data and the alarm receiving audio track data; According to the alarm sound wave data and the alarm receiving sound wave data, electric signal conversion is performed to obtain an alarm electric signal and an alarm receiving electric signal.

[0007] Preferably, the steps of extracting the language types of the alarm signal and the alarm receiving signal, and selecting different preset recognition models for text conversion according to the language types, are specifically: Converting the alarm electrical signal and the alarm receiving electrical signal to obtain discrete digital sampling signals; Segmenting the digital sampling signal according to a preset time frame to obtain an initial signal frame, and performing windowing processing on the initial signal frame to obtain a signal frame; Performing speech recognition on the signal frame to obtain a plurality of speech features, obtaining a language type based on the speech features, and establishing a speech feature vector sequence based on the plurality of speech features; Mapping the speech feature vector sequence with a preset pronunciation dictionary according to the language type to obtain a pronunciation sequence; Grammatical information, semantic information and contextual information are extracted according to the pronunciation sequence, language modeling is performed on the pronunciation sequence according to the grammatical information, the semantic information and the contextual information to obtain initial text data, text normalization and punctuation insertion are performed on the initial text data to obtain the alarm input text corresponding to the alarm electric signal and the alarm input text corresponding to the alarm receiving electric signal.

[0008] Preferably, the steps of extracting the language data of the alarm input text and the alarm receiving input text, identifying the language data, obtaining translation rules, and translating the alarm input text and the alarm receiving input text into each other according to the translation rules are specifically as follows: Performing text cleaning on the alarm input text and the alarm receiving input text, and performing text segmentation and word segmentation to obtain an alarm text sequence and an alarm receiving text sequence; Obtaining language data of the input text according to the alarm text sequence and the alarm receiving text sequence, and obtaining the alarm language type and the alarm receiving language type according to the language data; Searching a preset language database according to the alarm language type and the alarm receiving language type to obtain a translation rule; generating dense vectors for the alarm text sequence and the alarm receiving text sequence according to the translation rules, and extracting alarm grammar information and alarm receiving grammar information according to the dense vectors; Encoding the alarm grammar information and the alarm reception grammar information to obtain an alarm context vector and an alarm reception context vector; Based on the alarm context vector and the alarm receiving context vector, text translation is performed, and sequence optimization and domain adaptation are performed on the translated text to obtain alarm receiving output text and alarm output text.

[0009] Preferably, the steps of extracting the timbre data, pitch data and speech speed data from the alarm input voice and the alarm receiving input voice, generating an alarm person voice package and an alarm receiving person voice package, converting the alarm receiving output text according to the alarm person voice package to obtain the alarm receiving output voice, and converting the alarm output text according to the alarm receiving person voice package to obtain the alarm output voice are specifically as follows: Extracting timbre data, pitch data and speech speed data from the alarm input voice and the alarm receiving input voice respectively; Classifying and packaging the timbre data, the pitch data, and the speech speed data according to alarm and alarm reception, and generating an alarm person voice package and an alarm receiver voice package; According to the alarm person's voice package, the alarm output text is initially converted to obtain a first pronunciation data segment of each text, and the first pronunciation data segments are integrated and sorted to obtain an initial alarm output voice; Based on the alarm input voice, first tone information and first emotion information of the alarm person are obtained, and the initial alarm output voice is emotion-optimized according to the first tone information and the first emotion information to obtain the alarm output voice; According to the alarm receiver voice package, the alarm output text is initially converted to obtain a second pronunciation data segment of each text, and the second pronunciation data segments are integrated and sorted to obtain an initial alarm output voice; Based on the alarm input voice, the second voice information and second emotional information of the alarm receiver are obtained, and the initial alarm output voice is emotionally optimized according to the second voice information and the second emotional information to obtain the alarm output voice.

[0010] Preferably, the steps of building a data transmission link, protecting the communication data generated in each stage, and sending it to the next stage are specifically: Building a data transmission link, the data transmission link extracts the communication data generated in each stage, and identifies the communication data to obtain the data type and data length of the communication data; According to the data type, a data protection mode is obtained, and according to the data length, a data protection range is obtained; According to the data protection method and the data protection scope, a data key is generated, and the communication data is protected according to the data key and sent to the next stage.

[0011] Preferably, the steps of splitting the original call into two independent call links, extracting data transmission information of the data transmission links, and reconnecting the call links according to the data transmission information are specifically as follows: Split the original call into two independent call links, including the first call link being the original voice communication and the second call link being the virtual simultaneous interpretation voice communication; extracting data transmission information in the first call link and the second call link, and obtaining a data transfer node according to the data transmission information; The first call link and the second call link are interrupted according to the data transfer node to obtain the first alarm call link, the second alarm call link, the first alarm receiving call link, and the second alarm receiving call link; Connecting the first call link for alarm and the second call link for receiving alarm to obtain a call link between the alarmer and the virtual simultaneous call taker; The second alarm call link and the first alarm receiving call link are connected to obtain an alarm receiver-virtual simultaneous alarm person call link.

[0012] Preferably, the steps of synthesizing the four voice strips according to the time axis and generating four sound buttons to control the sound on / off state of each voice strip are specifically as follows: Extracting four time axes of the four voice strips, and aligning the four time axes head-on to obtain a time axis track; Placing the four voice strips on the timeline track, and aligning the voice strips following the head according to the same timestamp position to obtain a call recording; A separate sound button is provided for each of the voice bars, and the sound button controls the sound on / off state of each of the voice bars.

[0013] In a second aspect, the present application provides a foreign language virtual simultaneous interpretation system for receiving and handling police calls based on an intelligent voice interaction service, the system comprising: Voice collection module: used to collect the alarm input voice of the alarm end and the alarm receiving voice of the alarm receiving end during the alarm receiving process, and convert the alarm input voice and the alarm receiving voice into alarm electrical signals and alarm receiving electrical signals respectively; Speech recognition module: used to extract the language type of the alarm signal and the alarm receiving signal, and select different preset recognition models according to the language type to perform text conversion to obtain the alarm input text and the alarm receiving input text; Large model translation module: used to extract the language data of the alarm input text and the alarm reception input text, identify the language data, obtain translation rules, and translate the alarm input text and the alarm reception input text according to the translation rules to obtain the alarm reception output text and the alarm output text; Speech synthesis module: used to extract timbre data, pitch data and speech speed data from the alarm input voice and the alarm receiving input voice, generate an alarm person voice package and an alarm receiving person voice package, convert the alarm receiving output text according to the alarm person voice package to obtain the alarm receiving output voice, and convert the alarm output text according to the alarm receiving person voice package to obtain the alarm output voice; Communication transmission module: used to build a data transmission link, protect the communication data generated in each stage, and send it to the next stage; Call management module: used to split the original call into two independent call links, extract the data transmission information of the data transmission link, and reconnect the call links according to the data transmission information to obtain the call link between the alarmer and the virtual simultaneous caller and the call link between the caller and the virtual simultaneous caller; Recording synthesis module: record the alarm input voice, the alarm input voice, the alarm output voice and the alarm output voice to generate four voice strips, synthesize the four voice strips according to the time axis, and generate four sound buttons to control the sound switch state of each voice strip.

[0014] In summary, this application includes at least one of the following beneficial technical effects: The system collects the alarm input and alarm reception voice inputs from the alarm and alarm reception terminals, converts them into alarm and alarm reception signals, and then processes and identifies the two voice signals to determine the language type corresponding to each signal. Based on the language type, different pre-set recognition models are then selected to perform text conversion, resulting in the alarm input and alarm reception text. The system then extracts the language data for the two input texts, identifies the language data, and determines the translation rules for converting between the two languages. The system then translates the alarm input and alarm reception texts into each other based on the translation rules, resulting in the alarm reception output and alarm output texts. Specifically, the alarm input text is translated into the alarm reception output text, and the alarm reception text is translated into the alarm output text. The system then extracts the timbre, pitch, and speech rate data for the two input voices, and based on these data, performs virtual text-to-sound conversion on the alarm reception output and alarm output texts, respectively, to obtain the alarm reception output and alarm output voices. A data transmission link is then established, connecting the various modules and protecting the data transmitted between them. The original call is then split into two independent call links, which are then reassembled to create a call link between the caller and the virtual simultaneous caller, and another link between the caller and the virtual simultaneous caller. The entire call is then recorded, generating four audio strips. These strips are arranged on a single audio track, and a separate audio button is set up for each strip to turn it on or off. This improves the real-time and accuracy of language translation, as well as the efficiency of handling foreign-related police incidents. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 This is a flowchart of the steps of the method for receiving and handling police calls in a foreign language using a virtual simultaneous interpretation service based on an intelligent voice interaction service provided in an embodiment of the present application; Figure 2 This is a module block diagram of a foreign language virtual simultaneous interpretation system for receiving and handling police calls based on intelligent voice interaction services provided in an embodiment of the present application.

[0016] Explanation of the accompanying symbols: 1. Voice acquisition module; 2. Voice recognition module; 3. Large model translation module; 4. Voice synthesis module; 5. Communication transmission module; 6. Call management module; 7. Recording synthesis module. DETAILED DESCRIPTION

[0017] The following is combined with Figure 1-Figure 2 This application is further described in detail, but the embodiments of the present invention are not limited thereto.

[0018] The embodiments of the present application disclose a method and system for virtual simultaneous interpretation in foreign languages ​​for receiving and handling police calls based on intelligent voice interaction services.

[0019] In this embodiment, a foreign language virtual simultaneous interpretation method for receiving and handling police calls based on an intelligent voice interaction service includes: S100: During the alarm receiving process, the alarm input voice of the alarm end and the alarm receiving voice of the alarm receiving end are collected respectively, and the alarm input voice and the alarm receiving voice are converted into an alarm electric signal and an alarm receiving electric signal respectively; S200: extracting the language type of the alarm signal and the alarm receiving signal, and selecting different preset recognition models according to the language type to perform text conversion to obtain the alarm input text and the alarm receiving input text; S300: extracting language data of the alarm input text and the alarm reception input text, identifying the language data, obtaining a translation rule, and translating the alarm input text and the alarm reception input text according to the translation rule to obtain the alarm reception output text and the alarm output text; S400: extracting timbre data, pitch data, and speech rate data from the alarm input voice and the alarm receiving input voice, generating an alarmer voice package and an alarm receiving voice package, converting the alarm receiving output text according to the alarmer voice package to obtain the alarm receiving output voice, and converting the alarm output text according to the alarm receiving voice package to obtain the alarm output voice; S500: Build a data transmission link, protect the communication data generated in each stage, and send it to the next stage; S600: Split the original call into two independent call links, extract data transmission information of the data transmission link, and reconnect the call links based on the data transmission information to obtain a call link between the caller and the virtual simultaneous caller, and a call link between the caller and the virtual simultaneous caller. S700: Record the alarm input voice, alarm input voice, alarm output voice and alarm output voice to generate four voice strips, synthesize the four voice strips according to the time axis, and generate four sound buttons to control the sound switch state of each voice strip.

[0020] It should be noted that the above process is only the basic steps of this embodiment. During the specific implementation process, some steps can be appropriately added, reduced or modified without affecting the overall implementation effect.

[0021] The steps of collecting the alarm input voice of the alarm end and the alarm receiving voice of the alarm receiving end, and converting the alarm input voice and the alarm receiving voice into the alarm electric signal and the alarm receiving electric signal respectively, are specifically as follows: During the alarm reception process, the alarm terminal collects the voice of the alarm caller to obtain the alarm input voice, and the alarm receiving terminal collects the voice of the alarm receiver to obtain the alarm input voice; Extract the alarm audio track data and the alarm audio track data of the alarm input voice and the alarm input voice respectively; According to the alarm audio track data and the alarm receiving audio track data, the alarm sound wave data and the alarm receiving sound wave data are obtained; According to the alarm sound wave data and the alarm receiving sound wave data, electric signal conversion is performed to obtain the alarm electric signal and the alarm receiving electric signal.

[0022] During the alarm response process, the alarm terminal collects the caller's voice to generate the alarm input audio, while the alarm receiving terminal collects the caller's voice to generate the call-receiving audio. For example, a foreign tourist suddenly falls ill while visiting a scenic spot and calls the police using their mobile phone. The alarm terminal (the tourist's mobile phone) collects the tourist's English plea for help: "I need help, I feel very dizzy!" Simultaneously, the alarm receiving terminal (the equipment at the alarm reception center) collects the caller's Chinese response: "Don't worry, please tell me your exact location." Next, the system extracts the alarm audio track data for the alarm input audio and the call-receiving audio track data for the call-receiving audio. For example, the alarm audio track data includes the amplitude curve and spectral characteristics of the tourist's voice, while the call-receiving audio track data includes the waveform distribution and frequency range of the call-receiving audio. The system then generates the alarm sound wave data and the call-receiving sound wave data based on the alarm and call-receiving audio track data. The alarm sound wave data is an analog signal with an amplitude range of 0.5V to 2.3V, and the alarm reception sound wave data is an analog signal with an amplitude range of 0.3V to 2.0V. Finally, the system converts these sound wave data into electrical signals, converting the alarm sound wave data into a digitized alarm electrical signal (such as the sequence 01010101) and the alarm reception sound wave data into a digitized alarm electrical signal (such as the sequence 10101010).

[0023] The steps of extracting the language type of the alarm signal and the alarm receiving signal, and selecting different preset recognition models for text conversion according to the language type are as follows: Convert the alarm electrical signal and the alarm receiving electrical signal to obtain discrete digital sampling signals; The digital sampling signal is divided into a predetermined time frame to obtain an initial signal frame, and the initial signal frame is subjected to a windowing process to obtain a signal frame; Performing speech recognition on the signal frame to obtain a plurality of speech features, obtaining a language type based on the speech features, and establishing a speech feature vector sequence based on the plurality of speech features; Mapping the speech feature vector sequence to a preset pronunciation dictionary according to the language type to obtain a pronunciation sequence; Grammatical information, semantic information and contextual information are extracted based on the pronunciation sequence, language modeling is performed on the pronunciation sequence based on the grammatical information, semantic information and contextual information to obtain initial text data, text normalization and punctuation insertion are performed on the initial text data to obtain the alarm input text corresponding to the alarm signal and the alarm input text corresponding to the alarm signal.

[0024] During operation, the alarm and alarm receiving signals are converted into discrete digital sampled signals. For example, the alarm signal (01010101 sequence) is converted into a digital signal with a sampling frequency of 16kHz, and the alarm receiving signal (10101010 sequence) is converted into a digital signal with a sampling frequency of 16kHz. Next, the system segments these digital sampled signals into pre-set time frames (e.g., 20 milliseconds per frame) to generate initial signal frames. For example, the voice signal from the alarm end is segmented into 50 initial signal frames, and the voice signal from the alarm receiving end is segmented into 45 initial signal frames. The system then performs windowing on each initial signal frame to generate a signal frame. This windowing reduces spectral leakage in the signal frame and improves signal boundaries. For example, the signal energy of a frame from the alarm end is concentrated between 500Hz and 2000Hz, while the signal energy of a frame from the alarm receiving end is concentrated between 300Hz and 1500Hz. The system then performs speech recognition on the signal frames to extract multiple speech features. The speech features of the alarm speech frame include a fundamental frequency of 120Hz and a formant frequency of 800Hz; the speech features of the alarm-receiving speech frame include a fundamental frequency of 110Hz and a formant frequency of 700Hz. Based on these speech features, the system determines that the language type of the alarm speech is English and the language type of the alarm-receiving speech is Chinese. The system then creates a speech feature vector sequence: the feature sequence for the alarm speech is 120, 800, 0.8, and the feature sequence for the alarm-receiving speech is 110, 700, 0.7. Next, the system maps the speech feature vector sequence to a preset pronunciation dictionary (such as an English pronunciation dictionary and a Chinese pronunciation dictionary) based on the language type to obtain a pronunciation sequence. For example, the pronunciation sequence of the alarm speech is " / aɪ ni:d hɛlp / ," and the pronunciation sequence of the alarm-receiving speech is " / qǐng bù yào zhāo jí / ." Finally, the system extracts grammatical information (such as subject-verb-object structure), semantic information (such as "call for help"), and contextual information (such as "emergency medical situation") from the pronunciation sequence, performs language modeling, and generates initial text data. The initial text data for the alarm voice is "I need help," and the initial text data for the answering voice is "Please don't worry." After text standardization (such as capitalization correction) and punctuation insertion (such as adding commas), the final text data for the alarm input is "Need help," and the text for the answering voice is "Please don't worry."

[0025] The steps of extracting the language data of the alarm input text and the alarm receiving input text, identifying the language data, obtaining translation rules, and translating the alarm input text and the alarm receiving input text according to the translation rules are as follows: Perform text cleaning on the alarm input text and the alarm receiving input text, and perform text segmentation and word segmentation to obtain the alarm text sequence and the alarm receiving text sequence; Obtain the language data of the input text according to the alarm text sequence and the received alarm text sequence, and obtain the alarm language type and the received alarm language type according to the language data; Search in a preset language database according to the alarm language type and the received alarm language type to obtain a translation rule; Generate dense vectors for the alarm text sequence and the received alarm text sequence according to the translation rule, and extract alarm grammar information and received alarm grammar information according to the dense vectors; Encode according to the alarm grammar information and the received alarm grammar information to obtain an alarm context vector and a received alarm context vector; Perform text translation based on the alarm context vector and the received alarm context vector, and perform sequence optimization and domain adaptation on the translated text to obtain the received alarm output text and the alarm output text.

[0026] In operation, the alarm input text "I need help," is cleaned to "I need help," and the received alarm input text "Please don't worry," is cleaned to "Please don't worry." Next, the system performs text segmentation and word segmentation. For example, the alarm text "Ineed help" is segmented into the word sequence "I", "need", "help", and the received alarm text "Please don't worry" is segmented into the word sequence "Please", "don't", "worry". Then, the system determines the language data according to the word sequence. The alarm text is in English, and the received alarm text is in Chinese. According to the language data, the system obtains that the alarm language type is "English" and the received alarm language type is "Chinese". Next, the system searches for a translation rule in a preset language database (such as a database containing English-Chinese translation rules). "help" in English corresponds to "帮助" in Chinese, and "着急" in Chinese corresponds to "worry" in English. Then, the system generates dense vectors for the alarm text sequence and the received alarm text sequence according to the translation rule. The dense vector of the alarm text sequence is 0.8, -0.2, 0.5, and the dense vector of the received alarm text sequence is -0.3, 0.7, -0.1. According to the dense vectors, the system extracts alarm grammar information (such as subject-predicate structure) and received alarm grammar information (such as imperative sentence structure). Next, the system encodes the grammar information to generate an alarm context vector (such as 0.9, 0.1, -0.3) and a received alarm context vector (such as -0.2, 0.8, 0.4). Finally, the system performs text translation based on the context vectors. The alarm text "I need help" is translated into Chinese to get "我需要帮助"; the received alarm text "Please don't worry" is translated into English to get "Don't worry". After sequence optimization (such as adjusting the word order) and domain adaptation (such as adapting to police terms), the received alarm output text is finally generated as "我需要帮助" and the alarm output text is "Don't worry".

[0027] The steps of extracting timbre data, pitch data, and speech speed data from the alarm input voice and the alarm receiving input voice, generating an alarm person voice package and an alarm receiving person voice package, converting the alarm receiving output text according to the alarm person voice package to obtain the alarm receiving output voice, and converting the alarm output text according to the alarm receiving person voice package to obtain the alarm output voice are specifically as follows: Extracting timbre data, pitch data and speech speed data from the alarm input voice and the alarm receiving input voice respectively; The timbre data, pitch data and speech speed data are classified and packaged according to the alarm and alarm reception, and the alarm person voice package and the alarm receiver voice package are generated; According to the alarm person's voice package, the alarm output text is initially converted to obtain the first pronunciation data segment of each text, and the first pronunciation data segments are integrated and sorted to obtain the initial alarm output voice; Based on the alarm input voice, first tone information and first emotion information of the alarm person are obtained, and the initial alarm output voice is emotionally optimized according to the first tone information and the first emotion information to obtain the alarm output voice; According to the voice package of the alarm receiver, the alarm output text is initially converted to obtain the second pronunciation data segment of each text, and the second pronunciation data segments are integrated and sorted to obtain the initial alarm output voice; Based on the alarm input voice, the second voice information and second emotional information of the alarm receiver are obtained, and the initial alarm output voice is emotionally optimized according to the second voice information and the second emotional information to obtain the alarm output voice.

[0028] During operation, the timbre data, pitch data, and speech rate data are respectively extracted from the alarm input speech and the dispatcher input speech. The timbre data of the alarm input speech is a deep male timbre (fundamental frequency 120 Hz), the pitch data is a rising tone (amplitude +20%), and the speech rate data is 120 words per minute; the timbre data of the dispatcher input speech is a clear female timbre (fundamental frequency 220 Hz), the pitch data is a flat tone (amplitude ±5%), and the speech rate data is 100 words per minute. Next, the system classifies and packages these data according to alarm and dispatching, generating an alarm voice package (including timbre, pitch, and speech rate parameters) and a dispatcher voice package (including timbre, pitch, and speech rate parameters). Then, the system performs an initial conversion on the dispatching output text ("I need help") according to the alarm voice package. The text is segmented into a single-word sequence "I", "need", "help", "assist", and the first pronunciation data segment of each word is generated according to the pitch parameters of the alarm voice package (for example, the pronunciation duration of "I" is 0.3 seconds, and the fundamental frequency is 120 Hz). Then, all the pronunciation data segments are integrated and sorted to obtain the initial dispatching output speech. Next, the system analyzes the first tone information (urgent) and the first emotion information (anxious) of the alarm caller based on the alarm input speech, and optimizes the emotion of the initial dispatching output speech according to this information. Add a vibrato effect at the end of "I need help", and increase the speech rate to 140 words per minute to generate the final dispatching output speech. At the same time, the system performs an initial conversion on the alarm output text ("Don't worry") according to the dispatcher voice package. The text is segmented into a single-word sequence "Don′t", "worry", and the second pronunciation data segment of each word is generated according to the pitch parameters of the dispatcher voice package (for example, the pronunciation duration of "Don't" is 0.4 seconds, and the fundamental frequency is 220 Hz). After integration and sorting, the initial alarm output speech is obtained. Finally, the system analyzes the second voice information (steady) and the second emotion information (calm) of the dispatcher based on the dispatcher input speech, and optimizes the emotion of the initial alarm output speech. Add a steady ending sound at the end of "Don't worry", and adjust the speech rate to 90 words per minute to generate the final alarm output speech.

[0029] Steps to build a data transmission link, protect the communication data generated in each stage, and send it to the next stage, specifically: Build a data transmission link. The data transmission link extracts the communication data generated in each stage and identifies the communication data to obtain the data type and data length of the communication data; According to the data type, obtain the data protection method, and according to the data length, obtain the data protection range; According to the data protection method and the data protection range, generate a data key, and protect the communication data according to the data key and send it to the next stage.

[0030] During operation, the system establishes a data transmission link consisting of a voice acquisition node (Node1), a voice recognition node (Node2), a translation node (Node3), and a voice synthesis node (Node4). The data transmission link extracts communication data generated by each module (such as voice signals, text data, and translation results). Node1 generates alarm input voice data (2MB in size), and Node2 generates alarm input text data (0.1MB in size). Next, the system identifies the communication data, determining its data type (e.g., audio format for voice data, text format for text data) and data length (e.g., 2MB for voice data, 0.1MB for text data). The system then determines the data protection method based on the data type. Voice data is encrypted using AES-256, and text data is encoded using Base64. Furthermore, the data protection scope is determined based on the data length. Voice data is encrypted for the entire file, while text data is encoded for key fields. Finally, the system generates a data key based on the protection method and scope (e.g., a 32-bit random string for voice data) and uses this key to protect the communication data. After encryption, the alarm input voice data becomes a garbled sequence, and the alarm input text data becomes "SSBuZWVkIGhlbHA=" after encoding. The processed data is sent to the next node through a secure link (for example, Node2 receives encrypted voice data, and Node3 receives encoded text data).

[0031] The steps of splitting the original call into two independent call links, extracting data transmission information of the data transmission link, and reconnecting the call links according to the data transmission information are specifically as follows: Split the original call into two independent call links, including the first call link being the original voice communication and the second call link being the virtual simultaneous interpretation voice communication; Extracting data transmission information in the first call link and the second call link, and obtaining a data transfer node according to the data transmission information; The first call link and the second call link are interrupted according to the data transfer node to obtain the first call link for alarm, the second call link for alarm, the first call link for receiving alarm, and the second call link for receiving alarm; Connect the first call link of the alarm and the second call link of the alarm reception to obtain the call link between the alarmer and the virtual simultaneous call taker; The second call link of the alarm and the first call link of the alarm receiver are connected to obtain the call link of the alarm receiver-virtual simultaneous alarm person.

[0032] During operation, the original call is split into two independent call links. The first call link is the original audio communication, such as the direct conversation between the visitor (caller) and the police dispatcher; the second call link is a virtual simultaneous voice communication, such as the system-generated translation voice link. Next, the system extracts data transmission information (such as packet size and transmission timestamp) for both links. The packet size of the first call link is 5MB, with a transmission timestamp of 2023-10-01 10:00:00; the packet size of the second call link is 3MB, with a transmission timestamp of 2023-10-01 10:00:05. The system then determines the data handover node based on the data transmission information (for example, the point with overlapping timestamps is the handover node). When the first call link reaches the 10th second, the second call link begins transmission simultaneously, and a handover node is generated. Next, the system breaks the two call links at the handover node, generating four sub-links: the first call link for alarm (the original voice of the caller), the second call link for alarm (the virtual simultaneous interpretation of the caller), the first call link for receiving the call (the original voice of the caller), and the second call link for receiving the call (the virtual simultaneous interpretation of the caller). Finally, the system connects the first call link for alarm with the second call link for receiving the call to generate a call link between the caller and the virtual simultaneous caller (i.e., the caller hears the system-translated voice of the caller). Simultaneously, the second call link for alarm is connected to the first call link for receiving the call to generate a call link between the caller and the virtual simultaneous caller (i.e., the caller hears the system-translated voice of the caller). For example, a tourist hears the caller's voice translated into English: "Don't worry, help is on the way." The caller hears the tourist's voice translated into Chinese: "I am at the east gate of the scenic spot."

[0033] The steps to synthesize four audio strips according to the timeline and generate four sound buttons to control the sound on / off status of each audio strip are as follows: Extract the four time axes of the four audio strips and align the four time axes at the head to obtain the time axis track; Place the four audio clips on the timeline track and align the audio clips following the head according to the same timestamp position to obtain the call recording; Set a separate sound button for each voice bar, and the sound button controls the sound on / off status of each voice bar.

[0034] During use, the alarm input voice (the original voice of the tourist), the alarm input voice (the original voice of the police officer), the alarm output voice (the system-generated police officer's translated voice) and the alarm output voice (the system-generated police officer's translated voice) are recorded to generate four voice strips. For example: Audio strip 1: Tourist's original voice: "I need help!" (3 seconds) Voice strip 2: Police officer's original voice: "Please don't worry!" (2 seconds) Voice bar 3: Alarm output voice "I need help!" (system synthesis, duration 3 seconds) Voice strip 4: Alarm output voice "Don't worry!" (system synthesis, duration 2 seconds) Next, the system extracts the timelines of the four audio strips (e.g., audio strip 1 starts at 0 seconds and ends at 3 seconds; audio strip 2 starts at 3 seconds and ends at 5 seconds). It then aligns the tops of all the timelines (e.g., starting at 0 seconds) to generate a timeline track. For example: Audio bar 1: 0s → 3s Voice strip 2: 3s → 5s Audio bar 3: 0s → 3s Audio bar 4: 3s → 5s Then, place the audio strips on the timeline track according to the timestamps and align the subsequent parts (for example, audio strip 3 is aligned with audio strip 1, and audio strip 4 is aligned with audio strip 2). Finally, a call recording file is generated with a length of 5 seconds. Finally, the system sets an independent sound button for each audio strip: Button 1: Controls the switch of voice bar 1 (tourist's original voice) Button 2: Controls the switch of voice bar 2 (police officer’s original voice) Button 3: Controls the switch of voice bar 3 (alarm translation voice) Button 4: Controls the switch of voice bar 4 (alarm translation voice) For example, users can click buttons 3 and 4 to turn off the police translation voice (voice bars 3 and 4), retaining only the original voice of the tourist (voice bar 1) and the original voice of the police officer (voice bar 2), enabling flexible listening.

[0035] An embodiment of the present invention provides a foreign language virtual simultaneous interpretation system for receiving and handling police calls based on an intelligent voice interaction service, using any of the above-mentioned foreign language virtual simultaneous interpretation methods for receiving and handling police calls based on an intelligent voice interaction service. The system includes the following contents: Voice collection module 1: used to collect the alarm input voice of the alarm end and the alarm receiving voice of the alarm receiving end during the alarm receiving process, and convert the alarm input voice and the alarm receiving voice into alarm electrical signals and alarm receiving electrical signals respectively; Speech recognition module 2: used to extract the language type of the alarm signal and the alarm receiving signal, and select different preset recognition models according to the language type to perform text conversion to obtain the alarm input text and the alarm receiving input text; Large model translation module 3: used to extract the language data of the alarm input text and the alarm reception input text, identify the language data, obtain translation rules, and translate the alarm input text and the alarm reception input text according to the translation rules to obtain the alarm reception output text and the alarm output text; Speech synthesis module 4: used to extract the timbre data, pitch data and speech speed data from the alarm input voice and the alarm receiving input voice, generate the alarm person voice package and the alarm receiving person voice package, convert the alarm receiving output text according to the alarm person voice package to obtain the alarm receiving output voice, and convert the alarm output text according to the alarm receiving person voice package to obtain the alarm output voice; Communication transmission module 5: used to build a data transmission link, protect the communication data generated in each stage, and send it to the next stage; Call management module 6: used to split the original call into two independent call links, extract the data transmission information of the data transmission link, and reconnect the call links according to the data transmission information to obtain the call link between the alarmer and the virtual simultaneous caller and the call link between the caller and the virtual simultaneous caller; Recording synthesis module 7: Record the alarm input voice, alarm input voice, alarm output voice and alarm output voice to generate four voice strips, synthesize the four voice strips according to the time axis, and generate four sound buttons to control the sound switch status of each voice strip.

[0036] The above are all preferred embodiments of the present application, and are not intended to limit the scope of protection of the present application. Therefore, any equivalent changes made based on the structure, shape, and principle of the present application should be included in the scope of protection of the present application.

Claims

1. A foreign language virtual simultaneous interpretation method for police call reception and handling based on intelligent voice interaction service, characterized by: The following steps are involved: During the alarm receiving process, the alarm input voice of the alarm end and the alarm receiving voice of the alarm receiving end are collected respectively, and the alarm input voice and the alarm receiving voice are converted into an alarm electric signal and an alarm receiving electric signal respectively; Extracting the language types of the alarm signal and the alarm receiving signal, and selecting different preset recognition models according to the language types to perform text conversion to obtain alarm input text and alarm receiving input text; Extracting language data of the alarm input text and the alarm receiving input text, identifying the language data to obtain a translation rule, and translating the alarm input text and the alarm receiving input text into each other according to the translation rule to obtain an alarm receiving output text and an alarm output text; Extracting timbre data, pitch data, and speech rate data from the alarm input voice and the alarm receiving input voice, generating an alarm person voice package and an alarm receiving person voice package, converting the alarm receiving output text according to the alarm person voice package to obtain an alarm receiving output voice, and converting the alarm output text according to the alarm receiving person voice package to obtain an alarm output voice; Build a data transmission link, protect the communication data generated in each stage, and send it to the next stage; Split the original call into two independent call links, extract the data transmission information of the data transmission link, and reconnect the call links according to the data transmission information to obtain the call link between the alarmer and the virtual simultaneous caller and the call link between the caller and the virtual simultaneous caller; The alarm input voice, the alarm input voice, the alarm output voice and the alarm output voice are recorded to generate four voice strips, the four voice strips are synthesized according to the time axis, and four sound buttons are generated to control the sound switch state of each voice strip.

2. The method for virtual simultaneous interpretation of foreign languages ​​for receiving and handling police calls based on intelligent voice interaction services according to claim 1 is characterized in that: The steps of collecting the alarm input voice of the alarm terminal and the alarm receiving input voice of the alarm receiving terminal, and converting the alarm input voice and the alarm receiving input voice into the alarm electric signal and the alarm receiving electric signal respectively, are specifically: During the alarm reception process, the alarm terminal collects the voice of the alarm caller to obtain the alarm input voice, and the alarm receiving terminal collects the voice of the alarm receiver to obtain the alarm input voice; Extracting the alarm audio track data and the alarm receiving audio track data of the alarm input voice and the alarm receiving input voice respectively; Obtaining alarm sound wave data and alarm receiving sound wave data according to the alarm audio track data and the alarm receiving audio track data; According to the alarm sound wave data and the alarm receiving sound wave data, electric signal conversion is performed to obtain an alarm electric signal and an alarm receiving electric signal.

3. The method for virtual simultaneous interpretation of foreign languages ​​for receiving and handling police calls based on intelligent voice interaction services according to claim 2 is characterized in that: The steps of extracting the language types of the alarm signal and the alarm receiving signal, and selecting different preset recognition models for text conversion according to the language types, are specifically as follows: Converting the alarm electrical signal and the alarm receiving electrical signal to obtain discrete digital sampling signals; Segmenting the digital sampling signal according to a preset time frame to obtain an initial signal frame, and performing windowing processing on the initial signal frame to obtain a signal frame; Performing speech recognition on the signal frame to obtain a plurality of speech features, obtaining a language type based on the speech features, and establishing a speech feature vector sequence based on the plurality of speech features; Mapping the speech feature vector sequence with a preset pronunciation dictionary according to the language type to obtain a pronunciation sequence; Grammatical information, semantic information and contextual information are extracted according to the pronunciation sequence, language modeling is performed on the pronunciation sequence according to the grammatical information, the semantic information and the contextual information to obtain initial text data, text normalization and punctuation insertion are performed on the initial text data to obtain the alarm input text corresponding to the alarm electric signal and the alarm input text corresponding to the alarm receiving electric signal.

4. The method for virtual simultaneous interpretation of foreign languages ​​for receiving and handling police calls based on intelligent voice interaction services according to claim 3 is characterized in that: The steps of extracting the language data of the alarm input text and the alarm receiving input text, identifying the language data, obtaining a translation rule, and translating the alarm input text and the alarm receiving input text into each other according to the translation rule are specifically as follows: Performing text cleaning on the alarm input text and the alarm receiving input text, and performing text segmentation and word segmentation to obtain an alarm text sequence and an alarm receiving text sequence; Obtaining language data of the input text according to the alarm text sequence and the alarm receiving text sequence, and obtaining the alarm language type and the alarm receiving language type according to the language data; Searching a preset language database according to the alarm language type and the alarm receiving language type to obtain a translation rule; generating dense vectors for the alarm text sequence and the alarm receiving text sequence according to the translation rules, and extracting alarm grammar information and alarm receiving grammar information according to the dense vectors; Encoding the alarm grammar information and the alarm reception grammar information to obtain an alarm context vector and an alarm reception context vector; Based on the alarm context vector and the alarm receiving context vector, text translation is performed, and sequence optimization and domain adaptation are performed on the translated text to obtain alarm receiving output text and alarm output text.

5. The method for virtual simultaneous interpretation of foreign languages ​​for receiving and handling police calls based on intelligent voice interaction services according to claim 4 is characterized in that: The steps of extracting timbre data, pitch data, and speech speed data from the alarm input voice and the alarm receiving input voice, generating an alarm person voice package and an alarm receiving person voice package, converting the alarm receiving output text according to the alarm person voice package to obtain the alarm receiving output voice, and converting the alarm output text according to the alarm receiving person voice package to obtain the alarm output voice are specifically as follows: Extracting timbre data, pitch data and speech speed data from the alarm input voice and the alarm receiving input voice respectively; Classifying and packaging the timbre data, the pitch data, and the speech speed data according to alarm and alarm reception, and generating an alarm person voice package and an alarm receiver voice package; According to the alarm person's voice package, the alarm output text is initially converted to obtain a first pronunciation data segment of each text, and the first pronunciation data segments are integrated and sorted to obtain an initial alarm output voice; Based on the alarm input voice, first tone information and first emotion information of the alarm person are obtained, and the initial alarm output voice is emotion-optimized according to the first tone information and the first emotion information to obtain the alarm output voice; According to the alarm receiver voice package, the alarm output text is initially converted to obtain a second pronunciation data segment of each text, and the second pronunciation data segments are integrated and sorted to obtain an initial alarm output voice; Based on the alarm input voice, the second voice information and second emotional information of the alarm receiver are obtained, and the initial alarm output voice is emotionally optimized according to the second voice information and the second emotional information to obtain the alarm output voice.

6. The method for virtual simultaneous interpretation of foreign languages ​​for receiving and handling police calls based on intelligent voice interaction services according to claim 5 is characterized in that: The steps to build a data transmission link, protect the communication data generated in each stage, and send it to the next stage are as follows: Building a data transmission link, the data transmission link extracts the communication data generated in each stage, and identifies the communication data to obtain the data type and data length of the communication data; According to the data type, a data protection mode is obtained, and according to the data length, a data protection range is obtained; According to the data protection method and the data protection scope, a data key is generated, and the communication data is protected according to the data key and sent to the next stage.

7. The method for virtual simultaneous interpretation of foreign languages ​​for receiving and handling police calls based on intelligent voice interaction services according to claim 6 is characterized in that: The steps of splitting the original call into two independent call links, extracting data transmission information of the data transmission links, and reconnecting the call links according to the data transmission information are specifically as follows: Split the original call into two independent call links, including the first call link being the original voice communication and the second call link being the virtual simultaneous interpretation voice communication; extracting data transmission information in the first call link and the second call link, and obtaining a data transfer node according to the data transmission information; The first call link and the second call link are interrupted according to the data transfer node to obtain the first alarm call link, the second alarm call link, the first alarm receiving call link, and the second alarm receiving call link; Connecting the first call link for alarm and the second call link for receiving alarm to obtain a call link between the alarmer and the virtual simultaneous call taker; The second alarm call link and the first alarm receiving call link are connected to obtain an alarm receiver-virtual simultaneous alarm person call link.

8. The method for virtual simultaneous interpretation of foreign languages ​​for receiving and handling police calls based on intelligent voice interaction services according to claim 7 is characterized in that: The steps of synthesizing the four voice strips according to the time axis and generating four sound buttons to control the sound on / off state of each voice strip are specifically as follows: Extracting four time axes of the four voice strips, and aligning the four time axes head-on to obtain a time axis track; Placing the four voice strips on the timeline track, and aligning the voice strips following the head according to the same timestamp position to obtain a call recording; A separate sound button is provided for each of the voice bars, and the sound button controls the sound on / off state of each of the voice bars.

9. A virtual simultaneous interpretation system for receiving and handling police calls based on intelligent voice interaction services, the system using a virtual simultaneous interpretation method for receiving and handling police calls based on intelligent voice interaction services as claimed in any one of claims 1 to 8, characterized in that: The system comprises: Voice collection module: used to collect the alarm input voice of the alarm end and the alarm receiving voice of the alarm receiving end during the alarm receiving process, and convert the alarm input voice and the alarm receiving voice into alarm electrical signals and alarm receiving electrical signals respectively; Speech recognition module: used to extract the language type of the alarm signal and the alarm receiving signal, and select different preset recognition models according to the language type to perform text conversion to obtain the alarm input text and the alarm receiving input text; Large model translation module: used to extract the language data of the alarm input text and the alarm reception input text, identify the language data, obtain translation rules, and translate the alarm input text and the alarm reception input text according to the translation rules to obtain the alarm reception output text and the alarm output text; Speech synthesis module: used to extract timbre data, pitch data and speech speed data from the alarm input voice and the alarm receiving input voice, generate an alarm person voice package and an alarm receiving person voice package, convert the alarm receiving output text according to the alarm person voice package to obtain the alarm receiving output voice, and convert the alarm output text according to the alarm receiving person voice package to obtain the alarm output voice; Communication transmission module: used to build a data transmission link, protect the communication data generated in each stage, and send it to the next stage; Call management module: used to split the original call into two independent call links, extract the data transmission information of the data transmission link, and reconnect the call links according to the data transmission information to obtain the call link between the alarmer and the virtual simultaneous caller and the call link between the caller and the virtual simultaneous caller; Recording synthesis module: record the alarm input voice, the alarm input voice, the alarm output voice and the alarm output voice to generate four voice strips, synthesize the four voice strips according to the time axis, and generate four sound buttons to control the sound switch state of each voice strip.

Citation Information

Patent Citations

  • Simultaneous interpretation system based on speech recognition technology

    CN106486125A

  • Speech recognition system training method and device, speech recognition system recognition method and device and electronic equipment

    CN112885336A

  • Network data information secure transmission method and system

    CN117240627A

  • Inliteration method, device and equipment and storage medium thereof

    CN118609572A

  • Intelligent voice automatic translation system based on AI recognition

    CN119964573A