A positionable intelligent multi-lingual oral practice method, system and device

Wearable devices with dual-mode 4G and Bluetooth communication and positioning modules solve the problems of strong device dependence, disconnect between scene and location, and lack of security functions in existing technologies, and realize flexible learning experience and security protection, which is suitable for a variety of user groups.

CN121661879BActive Publication Date: 2026-04-10SHENZHEN INNOTRIK TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN INNOTRIK TECH
Filing Date
2026-02-09
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing oral language learning devices suffer from problems such as strong device dependence, disconnect between scenarios and locations, and lack of safety features when used independently. They are particularly inconvenient to use in scenarios such as sports and driving, and lack safety designs for children, the elderly, or outbound travelers.

Method used

Wearable devices employing 4G and Bluetooth dual-mode communication and positioning modules, combined with hardware DIP switches and one-button call functions, provide flexible communication mode switching, integrate high-precision positioning and emergency call capabilities, and realize a smart terminal integrating learning, positioning, and communication.

Benefits of technology

It enables flexible learning in different scenarios, provides a dynamic and contextualized learning experience, enhances user security, supports cross-mode learning data synchronization, and broadens application scenarios and user groups.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121661879B_ABST
    Figure CN121661879B_ABST
Patent Text Reader

Abstract

The application discloses a kind of positionable intelligent multilingual oral practice method, system and equipment, belong to the cross field of intelligent language learning equipment and wireless communication technology, first acquisition user voice in external environment, carry out hardware level switching between bluetooth mode and 4G independent mode;In 4G mode, support real-time positioning and recommend oral scene based on position, provide three kinds of practice mode of real-time conversation, preset scene and new word prompt;User learning data generated in the mode is stored synchronously in cloud;When switching to bluetooth mode and connecting mobile phone, mobile phone APP can request and synchronously display complete historical learning data, form continuous learning track.The application combines independent communication, positioning perception, scene learning and security communication function, solves the problem that existing technology equipment is strongly dependent, learning content and scene are disjointed, lacks security design and data continuity is poor, realizes all-weather, full-scene intelligent oral learning and security guarantee.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of intelligent language learning devices and wireless communication technology, and particularly relates to a positionable intelligent multi-lingual oral practice method, system and device. BACKGROUND

[0002] With the acceleration of globalization and the popularity of mobile communication technology, the demand for oral learning anytime, anywhere, personalized and scenario-based is growing. The existing oral practice technical solutions mainly have the following shortcomings:

[0003] The first type is a solution based on a smart phone application, which completely relies on the phone for recording, processing and interaction. Its limitations are: must carry and operate the phone, which is inconvenient and even has safety hazards in motion, driving and other scenarios; the phone microphone and speaker are separated, and the voice interaction quality is poor in noisy environments; the function is limited by the phone network and power, and cannot realize truly "off the phone" learning.

[0004] The second type is a lightweight solution based on a Bluetooth headset, which realizes simple voice interaction by pairing with a mobile phone. Its shortcomings are: the function completely depends on the mobile phone Bluetooth connection, and is invalid once disconnected or the phone is not within reach; the Bluetooth communication distance is limited, and cannot be used independently in a larger range; lacks self-positioning and cellular network communication capabilities, and cannot provide location-based services and emergency communication functions.

[0005] The third type is a dedicated language learning device, which usually has offline content, but lacks interaction intelligence, content update is inconvenient, and most of them do not have modern mobile communication and positioning capabilities, making it difficult to provide dynamic, situational learning experience and necessary safety protection.

[0006] In addition, the existing technology generally lacks safety design for user groups such as children, the elderly or outbound travelers. When users use the learning device independently, they cannot quickly contact the outside world in case of an emergency, and there is a safety blind area.

[0007] Therefore, there is an urgent need for a wearable device and system that integrates independent cellular communication, high-precision positioning, intelligent practice and safety communication to solve the problems of strong device dependence, disconnection of scene and location, and lack of safety functions in existing technology. SUMMARY

[0008] The purpose of the present application is to provide a positionable intelligent multi-lingual oral practice method, system and device. The method and system integrate 4G and Bluetooth dual-mode communication and positioning modules, and innovatively introduce a hardware dial switch and one-key call function, aiming to achieve the following goals:

[0009] Providing two flexible working modes of Bluetooth connection and independent 4G networking, and giving consideration to performance and portability;

[0010] Using real-time positioning information, providing spoken language learning content deeply integrated with user's geographical environment and cultural background;

[0011] In the independent working mode, reliable one-key emergency call capability is provided to enhance the safety of use;

[0012] An intelligent wearable learning terminal integrating learning, positioning and communication is constructed.

[0013] Technical scheme: the method is executed by a wearable device integrated with a 4G and Bluetooth dual-mode communication and positioning module and a mobile terminal, and comprises the following steps:

[0014] Step 1: collecting user's voice in the external environment through the wearable device, and performing voice enhancement processing on the collected voice, wherein the voice enhancement processing comprises echo cancellation AEC, automatic gain control AGC and background noise suppression ANS;

[0015] Step 2: executing 4G mode or Bluetooth mode according to the communication mode selected by the user through the dial switch;

[0016] Step 3: in the 4G mode, executing the corresponding accompanying training process according to the accompanying training mode selected by the user; the accompanying training process comprises a real-time accompanying training mode, a preset scene accompanying training mode and a prompt word scene accompanying training mode.

[0017] Step 4: in the accompanying training process, the user initiates a one-key call through the physical or virtual key on the device, and establishes a voice call connection with the preset matched mobile phone;

[0018] Step 5: recording and analyzing the interactive data of the user, and generating learning suggestions based on the analysis results, wherein the learning suggestions comprise recommendations for jumping or associating between different accompanying training modes;

[0019] Step 6: displaying the position of the wearable device in real time through the mobile phone APP software;

[0020] Step 7: in the 4G mode, the accompanying training process data, learning record and evaluation result of the user are stored in the cloud server synchronously; when the device is switched to the Bluetooth mode and connected with the mobile phone APP, the mobile phone APP initiates a data synchronization request to the cloud server, obtains and displays the complete historical learning data of the user.

[0021] Further, in step 1, the adaptive filter algorithm is adopted for the echo cancellation AEC, and the error signal calculation formula is as follows:

[0022]

[0023] wherein, is the error signal after echo cancellation, is the original input signal, is the prediction output signal of the filter;

[0024] The automatic gain control (AGC) dynamically adjusts the gain based on the root mean square value of the signal, and the gain calculation formula is:

[0025]

[0026] wherein, is the signal after dynamically adjusting the gain, is the target gain value, is the root mean square value of the input signal;

[0027] The background noise suppression (ANS) uses spectral subtraction for noise suppression, and the formula is:

[0028]

[0029] wherein, is the signal spectrum after denoising, is the original signal spectrum, is the estimated noise spectrum.

[0030] Further, in step 2, in the Bluetooth mode, the processed voice data is sent to the mobile APP software through the Bluetooth Low Energy (BLE) protocol, and the cloud server performs voice recognition, dialogue generation and audio synthesis, and returns the generated training content to the wearable device for playback; wherein the specific process of voice recognition, dialogue generation and audio synthesis includes:

[0031] Step 2.1, voice recognition: after the mobile APP receives the voice data, it first uses the end-to-end voice recognition model deployed in the cloud to convert the voice signal into text; the model uses an encoder-decoder structure based on an attention mechanism, the encoder extracts voice features, and the decoder outputs the corresponding text sequence, supports multi-language real-time conversion, and the voice recognition model is represented as:

[0032]

[0033] wherein, is the input voice signal, is the target sentence, is the conditional probability of the voice and the sentence given by the acoustic model, is the language model probability, which is trained from a large amount of corpus, is the marginal probability of the speech signal X, which is a constant during the decoding process and does not affect the maximum a posteriori probability estimation;

[0034] Step 2.2, Natural Language Understanding and Dialogue Generation: The recognized text is analyzed by the natural language understanding module to understand the user's intent and semantics. Combined with the context memory unit and the dialogue state tracking mechanism, the current dialogue topic and user needs are determined. Subsequently, based on the pre-trained generative dialogue model or retrieval dialogue engine, a response text in the target language that is context-appropriate and grammatically correct is generated.

[0035] Step 2.3, Audio Synthesis: The generated response text is input into the neural speech synthesis model. The model uses a Transformer-based acoustic model and vocoder to convert the text into a highly natural, emotionally charged speech waveform in the target language.

[0036] Step 2.4, Content Return: The synthesized audio stream is transmitted back to the wearable device in real time via Bluetooth link for playback, while the dialogue text and pronunciation evaluation prompts are simultaneously displayed on the mobile APP interface.

[0037] Furthermore, in step 2, in the 4G mode, the wearable device works independently of the mobile phone, uploading voice data to the cloud server through the 4G communication module. The cloud server processes the data and sends the generated training content to the wearable device for playback.

[0038] Furthermore, in step 4, the one-click call includes the following steps:

[0039] A call command is triggered by pressing a button on the device;

[0040] Initiate a voice call request to a preset mobile phone number via 4G network;

[0041] The current coaching session is automatically paused during the call and resumed or prompted to continue after the call ends.

[0042] Furthermore, the one-click call also supports voice triggering. After the user says the preset wake-up word, they can say the phone number to make a call. During the call, the device automatically enables voice enhancement processing to ensure call clarity.

[0043] Furthermore, in step 5, the learning suggestions specifically include:

[0044] Step 5.1: Switch from real-time coaching mode to preset scenario mode:

[0045] When the system detects that the user's pronunciation error rate is higher than the preset threshold or the fluency of the conversation is poor during real-time tutoring, it will automatically recommend entering the preset scenario mode to carry out targeted reinforcement training for high-frequency error words or sentence patterns.

[0046] Step 5.2, jump from preset scene mode to real-time accompanying mode:

[0047] When the user completes one or more sets of scene practice in the preset scene mode and the pronunciation accuracy and fluency reach the preset standard, the system recommends entering the real-time accompanying mode, encouraging the user to conduct comprehensive ability training in a conversation environment without fixed scripts;

[0048] Step 5.3, jump from real-time accompanying mode or preset scene mode to prompt word scene mode:

[0049] When the user actively requests personalized practice, or the system analyzes historical data and finds that the user has persistent weaknesses in specific topics, grammar or vocabulary, it is recommended to enter the prompt word scene mode, and customized accompanying content is generated based on user interests or ability short boards;

[0050] Step 5.4, associated recommended content generation mechanism:

[0051] The system intelligently generates cross-mode jump suggestions according to the user's historical training data, real-time performance, personal preferences and geographic location information, dynamically adjusts the difficulty and theme of the recommended content, uses a weighted collaborative filtering model for dynamic personalized recommendation, and combines user historical behavior and real-time location information. The prediction score formula is:

[0052]

[0053] Wherein, is the predicted interest score of the user to the scene , is the behavior similarity between the user and , is the actual interaction score of the user to the scene , is the average score of all scenes for the user u, is the set of similar users for the user u, is the average score of all scenes for the user v.

[0054] Further, in step 7, the data synchronization request is specifically: when the mobile phone APP initiates a synchronization request in Bluetooth mode, after verifying the user's identity, the corresponding historical data is encrypted and transmitted to the mobile phone APP, and is displayed in the form of a chart, a list or a report.

[0055] The application also discloses a positionable intelligent multi-lingual oral accompanying system, which comprises a wearable device integrated with a 4G and Bluetooth dual-mode communication and positioning module, an optionally connected mobile terminal and a cloud server.

[0056] The wearable device comprises:

[0057] a voice collection unit;

[0058] a voice processing unit;

[0059] a 4G and Bluetooth dual-mode communication and positioning unit supporting mobile network connection, Bluetooth data transmission and real-time positioning;

[0060] a state switching module for controlling the device to work in a Bluetooth mode or a 4G mode in response to the operation of a dial switch;

[0061] a call control unit for initiating voice communication with a matched mobile phone in response to a one-key call instruction in the 4G mode;

[0062] a voice playing unit;

[0063] the mobile terminal is installed with an application program for providing complete accompanying training services and interface interaction in the Bluetooth mode;

[0064] the cloud server is used for storing learning data of a user in the 4G mode and responding to a synchronization request of a mobile phone APP to download the historical learning data of the user to the mobile phone APP for visual display.

[0065] The application further discloses a wearable device comprising a shell, a magnetic attraction or clamping structure arranged on the shell, and a voice collection unit, a voice processing unit, a voice playing unit, a 4G and Bluetooth dual-mode communication and positioning unit, a state switching module and a call control unit integrated in the shell.

[0066] Compared with the prior art, the application has the following advantages:

[0067] 1. The application initiates a hardware-level dual-mode seamless switching mechanism: the hardware-level fast and reliable switching between the 4G and Bluetooth modes is realized through a physical dial switch, the operation is intuitive and the response is instant.

[0068] 2. The scene-based learning based on real geographical positions is realized: in the 4G independent mode, the high-precision positioning module integrated in the device can acquire the position of a user in real time.

[0069] 3, Embedded security communication provides core value extension: in the independent mode, the device provides a hardware level one-key emergency call function. When users (especially children, the elderly or travelers) encounter difficulties or dangers, they can quickly establish a voice call with family or emergency contacts through the preset button, which gives the product a vital security attribute and broadens the application scenarios and user groups.

[0070] 4, Build an all-weather, all-scenario intelligent learning closed loop: the system deeply integrates high-quality audio interaction, multi-mode intelligent practice, real-time positioning perception and security communication four functions, creating a wearable smart terminal that integrates learning, navigation and safety. It is not only a language learning tool, but also an intelligent assistant for users in unfamiliar environments.

[0071] 5, Seamless synchronization of cross-mode learning data: learning data generated in 4G independent mode is automatically stored in the cloud. When switching to Bluetooth mode and connecting to the phone, historical data can be synchronized to the phone APP, forming a continuous learning record and growth track, improving user experience and learning continuity. BRIEF DESCRIPTION OF DRAWINGS

[0072] Figure 1 The overall architecture diagram of the positionable intelligent multi-lingual oral practice system provided by the embodiment of the application is provided.

[0073] Figure 2 The Bluetooth mode workflow diagram provided by the embodiment of the application is provided.

[0074] Figure 3 The 4G independent mode workflow diagram provided by the embodiment of the application is provided.

[0075] Figure 4 The one-key call function flowchart provided by the embodiment of the application is provided.

[0076] Figure 5 The flowchart of recommending oral scenes based on location provided by the embodiment of the application is provided.

[0077] Figure 6 The wearable device structure diagram provided by the embodiment of the application is provided.

[0078] Figure 7 The workflow diagram of three practice modes in 4G mode provided by the embodiment of the application is provided.

[0079] Figure 8 The learning suggestion generation and mode jump recommendation flowchart provided by the embodiment of the application is provided.

[0080] Figure 9 The cross-mode learning data synchronization flowchart provided by the embodiment of the application is provided. DETAILED DESCRIPTION

[0081] The technical solutions of the present application are further described below with reference to the accompanying drawings.

[0082] Embodiment 1: System overall architecture

[0083] The positionable intelligent multi-lingual spoken language accompanying training solution provided by the present application is based on a set of "terminal (wearable device)-cloud (mobile phone App)" collaborative system. As shown in the figure, the system 100 mainly includes a wearable device 110 and an accompanying training application (App) running on a mobile terminal 120 (such as a smart phone). The communication control module of the wearable device 110 can control the linking mode of communication, which is divided into two types: 1) connecting with the mobile terminal 120 through wireless communication (such as Bluetooth); 2) directly connecting with the cloud service through the 4G module. The two linking modes are mutually exclusive. Figure 1

[0084] The wearable device 110 is designed to be worn on the chest of the user, for example, fixed by magnetic attraction, clips or lanyards, etc. Its core function is to serve as a high-quality audio interaction front end: the built-in voice collection unit (such as a directional microphone array) collects the user's voice clearly at close range; the voice processing unit performs echo cancellation (AEC), automatic gain control (AGC) and background noise suppression (ANS) processing, etc. to ensure the quality of the uploaded voice data; the voice playback unit (such as a micro speaker) is used to play the accompanying training audio from the App; the communication unit is responsible for data transmission and reception with the mobile phone App; the 4G communication module is responsible for connecting with the cloud service and realizing data transmission and reception; the communication control module is used to switch the communication mode; the call control unit connects with the control mobile phone through the 4G module and realizes the calling function.

[0085] The App on the mobile terminal 120 carries the main intelligent processing and business logic. It includes a mode management interface, a voice recognition / synthesis engine, a machine translation module, a dialogue generation model, a scene construction engine, and a user evaluation module, etc. The App receives clean voice data from the wearable device 110, performs in-depth processing, and sends the generated accompanying training audio or control instructions back to the wearable device 110 for playback.

[0086] Embodiment 2: Bluetooth mode workflow

[0087] As shown in the figure, the workflow of the Bluetooth mode starts with the user switching the physical dial switch on the device to the Bluetooth mode. Figure 2

[0088] ​​Subsequently, the device automatically or manually establishes a BLE Bluetooth connection with the user's smartphone, which has installed a matching application. After a successful connection, the user can begin a spoken language practice. When the user speaks, the device's built-in microphone array collects the voice signal and immediately performs real-time enhancement processing by the local speech processing unit, including echo cancellation, automatic gain control, and background noise suppression, to improve the voice quality.

[0089] The processed pure audio data is transmitted in real time to the application on the phone through the BLE Bluetooth link. After receiving the audio stream, the phone APP further uploads it to the cloud server, which uses its more powerful computing power for deep speech recognition, natural language understanding, and intelligent dialogue generation. The generated speech translation and scene practice content are synthesized into high-quality audio files or streams by the cloud or the local phone.

[0090] Finally, the generated practice audio is transmitted back to the wearable device through Bluetooth connection and played to the user by the device's built-in micro speaker; the generated text is displayed on the APP. During this process, the user can flexibly select different practice modes, view learning progress and analysis reports through the graphical interface of the phone APP.

[0091] In addition, in Bluetooth mode, the phone APP can actively initiate a historical data synchronization request to the cloud server. After verifying the user's identity, the cloud server encrypts and transmits all the user's learning records, practice scores, scene history, and other data generated in 4G independent mode to the phone APP. After receiving the data, the APP displays it to the user in a visual form (such as a learning curve, a transcript, a scene map), achieving complete connection of learning records across modes.

[0092] Embodiment 3: 4G independent mode workflow

[0093] As shown in Figure 3 When the user needs to use independently without the phone, the dial switch can be switched to 4G mode, automatically disconnecting the Bluetooth connection with the phone and entering an independent working state.

[0094] In this mode, the device first obtains the user's real-time geographic location information (such as GPS coordinates) through the integrated 4G communication and positioning module. When the user practices spoken language, the device collects and enhances the voice, and no longer relies on Bluetooth, but directly uploads the voice data along with the location information to the cloud server through the built-in 4G module.

[0095] After receiving the data, the cloud server first performs keyword recognition to determine whether the user is modifying the configuration or performing a practice task. If it is modifying the configuration, it modifies the practice mode, scene selection, etc., and returns a prompt tone indicating success or failure.

[0096] If no keyword is recognized, core processing is performed. In one aspect, the speech is recognized and semantically analyzed to generate personalized coaching dialogue content; in another aspect, based on uploaded real-time location information, a spoken language learning scene related to the geographic location is matched from a scene database (for example, if the user is near an airport, then dialogue scenes such as check-in, security check, boarding, etc. are recommended). The cloud dynamically combines the generated intelligent coaching content with scene recommendations and issues instructions and audio data to the device through a 4G network.

[0097] During this process, the cloud server records the user's coaching process data in real time, including practice time, scene type, dialogue content, pronunciation score, error vocabulary, etc., and stores them in the user learning database according to the user ID. These data are continuously accumulated in 4G mode to establish a complete learning profile for the user.

[0098] Embodiment 4: One-key call function flow

[0099] As shown in Figure 4 , in the 4G independent coaching process, if the user encounters an emergency or needs assistance, the safety call function can be used.

[0100] The user can trigger the call in two ways: one is to press the physical "one-key call" button on the device, which calls the pre-stored mobile phone number; the other is to speak the pre-set voice wake-up word and then speak the phone number to confirm the contact person.

[0101] After triggering the call instruction, the device first suspends any coaching task being performed to ensure the priority of the call. Then, the call control unit initiates a standard voice call request to the emergency contact person's mobile phone through the 4G cellular network.

[0102] After the call is established, the voice processing unit of the device automatically enables the enhanced algorithm for the call scene to ensure clear and smooth communication between the two parties. After the call is over, the device will ask the user "whether to resume learning?" through voice prompts. After the user confirms, the coaching will be seamlessly resumed from the suspended place; if cancelled, the current session will be ended.

[0103] Embodiment 5: Location-based spoken language scene recommendation flow

[0104] As shown in Figure 5 , in the 4G independent mode, the location-based service flow continuously runs in the background to provide dynamic learning content for the user.

[0105] The device actively acquires the accurate location through GPS / Beidou or base station positioning at configurable time intervals (such as every 5 minutes) or after detecting that the user has moved a certain distance. This location information is uploaded to the cloud intelligent recommendation engine in real time.

[0106] The cloud engine matches the user's current location with a large database of geo-tagged and colloquial scenarios. For example, when the system identifies that the user is near a famous museum, it can automatically recommend relevant multilingual conversation exercises such as "buy tickets", "ask about exhibits", "discuss history", etc.

[0107] After receiving the scenario recommendation prompt, the user can choose to accept the recommendation, and the system will automatically guide the user into the scenario for immersive oral practice. Alternatively, the user can ignore the recommendation and continue with the original learning plan. The user's choices and feedback will be recorded for continuous optimization of future recommendation accuracy, achieving a personalized learning path that becomes smarter with use.

[0108] Figure 6 A specific embodiment of the wearable device 110 is shown. The housing 601 is provided with a back clip 602 for attaching to clothing. Inside are integrated a main processing chip (voice processing unit) 603, a microphone (voice acquisition unit) 604, a speaker (voice playback unit) 605, a Bluetooth module (communication unit) 606, a 4G module (communication unit) 607, a Bluetooth / 4G module switch 608, and a one-key wake-up button 609. This device is optimized for clear pickup of chest-front speech and comfortable earphone playback.

[0109] Embodiment 6: 3 practice modes in 4G mode

[0110] In 4G independent working mode, the system supports three practice modes, and the user can select and switch modes through voice wake-up words or device buttons. The following will be described in detail in conjunction with the accompanying Figure 7

[0111] I. Real-time practice process

[0112] This mode simulates a "conversation partner" who is always online and can respond to the user's speech immediately.

[0113] S301: The user speaks a sentence to the wearable device 110 (such as Chinese: "Today the weather is really good"). The voice acquisition unit of the wearable device 110 acquires the voice, and the voice processing unit immediately performs enhancement processing (eliminates environmental noise, adjusts volume).

[0114] S302: The processed voice data is sent to the phone App through Bluetooth.

[0115] S303: After receiving the voice data, the App starts real-time practice processing. The end-to-end voice conversion model takes the source language voice data as direct input, and directly generates the target language practice voice waveform through a deep neural network.

[0116] S304: The generated practice audio data is sent back to the wearable device 110 from the App. ​

[0117] S305: The voice playing unit of the wearable device 110 plays the foreign language audio. The user can hear a natural response as if from a conversation partner.

[0118] II. Scene selection practice flow

[0119] This mode is for in-depth learning of specific topics (such as ordering in a restaurant, checking into a hotel, business meetings).

[0120] S401: The user sets the scene through a wake-up word (such as "restaurant mode").

[0121] S402: The cloud server generates a complete set of learning materials based on the user's selected scene. First, it generates scene description information containing the scene location, character roles, and conversation goals. Then, based on this description, it uses a dialogue generation model to create a logically coherent, idiomatic Chinese target language multi-round dialogue text.

[0122] S403: The user enters the "follow-reading practice" stage. The device automatically plays the generated sentences, and then the user listens to the demonstration pronunciation played by the device.

[0123] S404: After each sentence is played, the wearable device 110 collects the user's follow-reading voice and transmits it back to the cloud server.

[0124] S405: The cloud server recognizes the user's follow-reading voice and performs speech recognition to obtain the user's spoken text, compares it with the played voice text, and evaluates the pronunciation accuracy of each word. If correct, the next sentence is played; otherwise, the sentence is repeated.

[0125] S406: After a round of conversation practice, the user's overall score and pronunciation improvement suggestions can be played.

[0126] III. Cue word practice flow

[0127] This mode is for practicing with the cue words given by the user.

[0128] S501: The user enters this mode through a wake-up word, says a cue word, and initiates "cue word practice".

[0129] S502: The cloud server performs in-depth analysis based on the user's selected cue word: through dictionary API and corpus analysis, it obtains its part of speech (noun), common collocations, applicable context (business, diplomacy, daily bargaining), and cultural background.

[0130] S503: Based on the analysis results, a customized oral practice scene is constructed. For example, for "advanced", the following scene description is generated: "Two children communicate advanced things".

[0131] S504: Similar to the preset scene flow, the App generates the sentence containing "advanced" and related expressions according to the customized description, and pronounces the sentence in turn for the user to learn.

[0132] S505: After each sentence is played, the wearable device 110 collects the follow-up voice of the user and transmits it back to the cloud server.

[0133] S506: The cloud server recognizes the follow-up voice of the user and performs voice recognition to obtain the text spoken by the user, compares it with the played voice text, and evaluates the pronunciation accuracy of each word. If correct, the next sentence is played; otherwise, the sentence is repeated.

[0134] S507: After a round of conversation practice, the user's overall score and pronunciation improvement suggestions can be played.

[0135] Embodiment 7: Cross-mode learning data synchronization flow

[0136] This embodiment describes a seamless synchronization mechanism for learning data between 4G mode and Bluetooth mode, as shown in FIG. 7, and the specific process is as follows: Figure 9

[0137] S701: The user conducts oral practice in the 4G independent mode, and all interaction data (including voice recording, scene selection, pronunciation score, and location information) are uploaded to the cloud server in real time and stored in the user learning database according to the user account ID.

[0138] S702: The user ends the 4G mode and switches the device dial switch to Bluetooth mode and establishes a Bluetooth connection with the phone APP.

[0139] S703: After the phone APP detects the device mode switching and successful connection, it automatically or according to the user's instruction initiates a historical data synchronization request to the cloud server, which includes user identity verification information.

[0140] S704: The cloud server verifies the legality of the user's identity, and retrieves all historical learning data of the user in the 4G mode from the user learning database.

[0141] S705: The cloud server encrypts and packages the retrieved learning data (which can include structured learning reports, original voice clips, score details, scene tracks, etc.) and transmits them to the phone APP through the Internet.

[0142] S706: After receiving and decrypting the data, the phone APP performs analysis and reorganization, and displays them to the user in various visual forms, including but not limited to: learning time curve chart, pronunciation accuracy rate statistics table, scene practice history list, geographic location learning map, and new word mastery progress, etc. ​

[0143] S707: The user can view the detailed learning report through the mobile phone APP, understand the learning achievements in the 4G independent mode, and adjust the subsequent learning plan or select the recommended learning path based on the historical data.

[0144] Embodiment 7 also includes a learning suggestion generation and mode jump recommendation mechanism, as shown in Figure 8 The specific process is as follows:

[0145] When analyzing the user's learning data, the system monitors the user's performance indicators in the current mode in real time, including pronunciation accuracy, fluency, error frequency, scene completion degree, etc. Based on these data, the system generates intelligent learning suggestions according to the preset jump conditions and algorithms, and at the appropriate time, it pushes the user to switch to a more suitable practice mode through voice prompts or mobile phone APP interface, the specific jump logic includes:

[0146] (1) Real-time practice mode → preset scene mode: When the system identifies that the user is continuously making mistakes in a certain type of sentence or vocabulary, it automatically recommends entering the relevant preset scene for specialized practice;

[0147] (2) Preset scene mode → real-time practice mode: When the user reaches the learning goal in this scene, the system encourages the user to enter the real-time mode for real combat dialogue;

[0148] (3) Real-time / preset scene mode → prompt word scene mode: When the user is interested in expressing or the system detects a knowledge gap, it recommends using the prompt word mode for personalized expansion training;

[0149] (4) Dynamic adjustment of recommended content: The system combines user historical performance, geographic location, learning preferences, etc. to update the recommended content in real time, ensuring the personalization and efficiency of the learning path.

[0150] This mechanism works in coordination with the cross-mode data synchronization process ( Figure 9 ), forming a closed loop of "learning-evaluation-recommendation-synchronization" to improve user learning experience and effectiveness.

[0151] Embodiment 8: Details of the voice enhancement processing algorithm and its role in the system

[0152] In the voice enhancement processing, echo cancellation (AEC), automatic gain control (AGC), and background noise suppression (ANS) are the basis for achieving high-quality voice interaction. The processing can be achieved through the following algorithms:

[0153] Echo cancellation (AEC): An adaptive filter algorithm, such as the least mean square error (LMS) algorithm, is used to update the filter weights in real time to eliminate echoes. The error signal calculation formula is:

[0154]

[0155] wherein, is the error signal after echo cancellation, is the original input signal, is the prediction output signal of the filter. This algorithm ensures that while the device plays the accompaniment audio, it can effectively separate and eliminate the echo, avoiding interference of the voice signal in the local loop, providing a clean input source for subsequent speech recognition.

[0156] Automatic Gain Control (AGC): dynamically adjust the gain based on the root mean square value of the signal, ensure that the voice signal is within the appropriate range. The gain calculation formula is:

[0157]

[0158] wherein, is the target gain value, is the root mean square value of the input signal. This algorithm can adapt to different user pronunciation volume and environmental changes, avoiding the influence of weak or strong voice signals on recognition accuracy, improving the robustness of the system in complex environments.

[0159] Ambient Noise Suppression (ANS): uses spectral subtraction for noise suppression, by modeling and subtracting the background noise spectrum, improving speech clarity. The basic formula is:

[0160]

[0161] wherein, is the denoised signal spectrum, is the original signal spectrum, is the estimated noise spectrum. This algorithm is particularly suitable for noisy environments such as outdoors and public places, can significantly reduce the impact of background noise on voice collection and recognition, and ensure the clarity and coherence of the accompaniment conversation.

[0162] In summary, the voice enhancement algorithm described in this embodiment is the key technical basis for the invention to achieve high-precision voice collection and interaction in various environments, directly related to the accuracy and user experience of subsequent speech recognition, dialogue generation and user evaluation.

[0163] Example 9: Details of speech recognition and dialogue generation algorithm implementation and its role in the system

[0164] In the speech recognition process, the cloud server uses a deep neural network-based speech recognition model, such as Convolutional Neural Network (CNN) or Recurrent Neural Network (RNN), to achieve high-accuracy speech-to-text conversion. The model is trained based on the maximum likelihood estimation method, and its probability model can be represented as:

[0165]

[0166] in, The input voice signal. For the target statement, The conditional probabilities of speech and sentences given for the acoustic model. The probabilities of the language model are obtained through training on a large-scale corpus. This model, through end-to-end learning, can effectively capture the temporal and semantic features in speech signals and accurately convert users' speech into text information.

[0167] This recognition result is a prerequisite for all subsequent intelligent coaching functions:

[0168] In real-time coaching mode, the recognized text is directly input into the end-to-end speech conversion model to generate a speech response in the target language.

[0169] In preset scenarios and prompt word scenario modes, the recognized text is used to understand user intent, assess pronunciation accuracy, and drive the dialogue generation engine to build coherent and authentic practice content.

[0170] In a voice-triggered scenario for one-click calling, the recognized text is used to parse the phone number spoken by the user, enabling voice dialing.

[0171] Therefore, the speech recognition and probability model described in this embodiment is the core bridge connecting user voice input and system intelligent feedback, and its performance directly determines the response speed, accuracy and naturalness of the coaching.

[0172] Example 10: Implementation details of a location-based learning-based scene recommendation algorithm and its role in the system

[0173] In location-based learning-based scenario recommendation, the system employs a weighted collaborative filtering model for dynamic personalized recommendations. This algorithm effectively integrates user historical behavior and real-time location information, and its prediction scoring formula is as follows:

[0174]

[0175] in For users For the scene Predicted interest scores For users and Behavioral similarity between them (calculated based on historical practice records, scene preferences, etc.) For users For the scene The algorithm assesses actual interaction scores. By analyzing which scenarios similar users prefer in similar locations, it achieves personalized and contextualized content recommendations.

[0176] This recommendation mechanism is deeply integrated with the location-aware capability of the application:

[0177] Real-time location input: The device acquires GPS / base station positioning data from the positioning unit through 4G / Bluetooth dual-mode communication, which serves as a key contextual input for the recommendation algorithm.

[0178] Scene library matching: The system maintains a library of colloquial scenes tagged with geographic locations (e.g., "airport check-in," "restaurant ordering," "museum tour," etc.). The algorithm filters out a set of candidate scenes with high geographical relevance based on the user's current location.

[0179] Personalized ranking: In the above-mentioned candidate set, a collaborative filtering model is used to combine user historical preference data to score and rank scenes individually, with the highest-scoring scenes being recommended to the user first.

[0180] Dynamic updating: User feedback data such as acceptance, neglect, and practice completion of recommended scenes are used to update the user's profile in real time, which is used to optimize subsequent recommendations, forming a virtuous cycle of "the more you use, the more accurate it becomes."

[0181] Therefore, the recommendation algorithm described in this embodiment is the technical engine that realizes the core innovation of "scenario-based learning based on real geographic location." It transforms static learning content into dynamic, personalized, and deeply integrated learning experiences with the user's environment, significantly improving the immersion, practicality, and user stickiness of learning.

[0182] In summary, the application integrates intelligent colloquial practice, real-time positioning tracking, and emergency communication functions into a wearable device through innovative hardware design and system integration, creating a new generation of safe, convenient, and scenario-based language learning solution. It not only improves learning efficiency but also expands the product's application value and user base.

[0183] The above description is only the preferred embodiment of the application, but the protection scope of the application is not limited to this. Any skilled person in the art can make equivalent replacements or changes to the technical solutions and inventive concepts of the application within the scope of the application, which should be covered by the protection scope of the application.

Claims

1. A location-based intelligent multilingual oral practice method, characterized in that, The method is executed collaboratively by a wearable device and a mobile terminal that integrate 4G and Bluetooth dual-mode communication and positioning modules, and includes the following steps: Step 1: Collect user voice from the external environment through the wearable device, and perform voice enhancement processing on the collected voice. The voice enhancement processing includes echo cancellation (AEC), automatic gain control (AGC), and background noise suppression (ANS). Step 2: Execute 4G mode or Bluetooth mode according to the communication mode selected by the user via the DIP switch; In step 2, in the Bluetooth mode, the processed voice data is sent to the mobile APP via Bluetooth BLE. The cloud server performs speech recognition, dialogue generation, and audio synthesis, and then returns the generated training content to the wearable device for playback. The specific processes for speech recognition, dialogue generation, and audio synthesis include: Step 2.1, Speech Recognition: After receiving the speech data, the mobile app first uses an end-to-end speech recognition model deployed in the cloud to convert the speech signal into text. The model adopts an encoder-decoder structure based on an attention mechanism. The encoder extracts speech features, and the decoder outputs the corresponding text sequence, supporting real-time conversion of multiple languages. The speech recognition model is represented as follows: ; in, The input voice signal. For the target statement, The conditional probabilities of speech and sentences given for the acoustic model. These are the probabilities of the language model, obtained through training on a large-scale corpus. For speech signals The marginal probability; Step 2.2, Natural Language Understanding and Dialogue Generation: The recognized text is analyzed by the natural language understanding module to understand the user's intent and semantics. Combined with the context memory unit and the dialogue state tracking mechanism, the current dialogue topic and user needs are determined. Subsequently, based on the pre-trained generative dialogue model or retrieval dialogue engine, a response text in the target language that is context-appropriate and grammatically correct is generated. Step 2.3, Audio Synthesis: The generated response text is input into the neural speech synthesis model. The model uses a Transformer-based acoustic model and vocoder to convert the text into a target language speech waveform with naturalness and emotional intonation. Step 2.4, Content Return: The synthesized audio stream is transmitted back to the wearable device in real time via Bluetooth link for playback, while the dialogue text and pronunciation evaluation prompts are displayed synchronously on the mobile APP interface; Step 3: In 4G mode, execute the corresponding coaching process according to the coaching mode selected by the user; the coaching process includes real-time coaching mode, preset scenario coaching mode, and prompt word scenario coaching mode. Step 4: During the coaching process, the user initiates a one-click call via physical or virtual buttons on the device to establish a voice call connection with the pre-set paired mobile phone. Step 5: Record and analyze user interaction data, and generate learning suggestions based on the analysis results. The learning suggestions include recommendations for switching or associating between different coaching modes. In step 5, the learning suggestions specifically include: Step 5.1: Switch from real-time coaching mode to preset scenario mode: When the system detects that the user's pronunciation error rate is higher than the preset threshold or the fluency of the conversation is poor during real-time tutoring, it will automatically recommend entering the preset scenario mode to carry out targeted reinforcement training for high-frequency error words or sentence patterns. Step 5.2: Switch from the preset scene mode to the real-time coaching mode: When a user completes one or more sets of scenario exercises in a preset scenario mode and the pronunciation accuracy and fluency reach the preset standards, the system recommends entering the real-time coaching mode, allowing the user to conduct comprehensive ability training in a dialogue environment without a fixed script. Step 5.3: Switch from real-time coaching mode or preset scenario mode to prompt word scenario mode: When a user actively requests personalized practice, or when the system analyzes historical data and finds that the user has persistent weaknesses in specific topics, grammar, or vocabulary, it recommends entering the prompt word scenario mode to generate customized practice content based on the user's interests or ability shortcomings. Step 5.4, Related Recommendation Content Generation Mechanism: Based on the user's historical training data, real-time performance, personal preferences, and geographic location information, the system intelligently generates cross-mode navigation suggestions and dynamically adjusts the difficulty and topic of the recommended content. It employs a weighted collaborative filtering model for dynamic personalized recommendations, integrating user historical behavior and real-time location information. Its prediction scoring formula is as follows: ; in, For users For the scene Predicted interest scores For users and Behavioral similarity between them For users For the scene Actual interaction rating, The average rating of user u across all scenarios. Let u be the set of similar users. The average rating of user v across all scenarios; Step 6: Display the real-time location of the wearable device via the mobile app; Step 7: In 4G mode, the user's practice process data, learning records, and assessment results are synchronously stored on the cloud server. When the device switches to Bluetooth mode and connects to the mobile APP, the mobile APP sends a data synchronization request to the cloud server to obtain and display the user's complete historical learning data.

2. The location-based intelligent multilingual oral practice method according to claim 1, characterized in that, In step 1, the echo cancellation (AEC) employs an adaptive filter algorithm, and its error signal calculation formula is as follows: ; in, This is the error signal after echo cancellation. The original input signal, This is the predicted output signal of the filter; The automatic gain control (AGC) dynamically adjusts the gain based on the root mean square (RMS) value of the signal. The gain calculation formula is as follows: ; in, The signal after dynamic gain adjustment. For the target gain value, The root mean square value of the input signal; The background noise suppression ANS uses spectral subtraction for noise suppression, which is achieved by modeling and subtracting the background noise spectrum, as shown in the formula: ; in, The signal spectrum after denoising. The original signal spectrum, This is the estimated noise spectrum.

3. The location-based intelligent multilingual oral practice method according to claim 1, characterized in that, In step 2, in the 4G mode, the wearable device works independently of the mobile phone, uploading voice data to the cloud server through the 4G communication module. The cloud server processes the data and sends the generated training content to the wearable device for playback.

4. The location-based intelligent multilingual oral practice method according to claim 1, characterized in that, Step 4, the one-click call includes the following steps: Trigger a call command by pressing a button on the device; Initiate a voice call request to a preset mobile phone number via 4G network; The current coaching session is automatically paused during the call and resumed or prompted to continue after the call ends.

5. The location-based intelligent multilingual oral practice method according to claim 4, characterized in that, The one-click call also supports voice triggering. After the user says the preset wake-up word, they can say the phone number to make a call. During the call, the device automatically enables voice enhancement processing to ensure call clarity.

6. The location-based intelligent multilingual oral practice method according to claim 1, characterized in that, In step 7, the data synchronization request specifically involves: when the mobile app initiates a synchronization request in Bluetooth mode, after verifying the user's identity, the corresponding historical data is encrypted and transmitted to the mobile app, and displayed in the form of charts, lists, or reports.

7. A location-based intelligent multilingual oral practice system, used to implement the method as described in claim 1, characterized in that, This includes wearable devices with integrated 4G and Bluetooth dual-mode communication and positioning modules, optional connected mobile terminals, and cloud servers; The wearable device includes: Voice acquisition unit; Voice processing unit; 4G and Bluetooth dual-mode communication and positioning unit, supporting mobile network connection, Bluetooth data transmission and real-time positioning; The status switching module responds to DIP switch operations and controls the device to work in Bluetooth mode or 4G mode. The call control unit is used to respond to one-click call commands in 4G mode and initiate voice calls with the paired mobile phone. Voice playback unit; The mobile terminal has an application installed to provide a complete coaching service and interface interaction in Bluetooth mode; The cloud server is used to store the user's learning data in 4G mode and respond to the synchronization request of the mobile APP in Bluetooth mode, sending the user's historical learning data to the mobile APP for visualization.

8. A wearable device for integrating the system as described in claim 7, characterized in that, It includes a housing, a magnetic or clamping structure mounted on the housing, and a voice acquisition unit, a voice processing unit, a voice playback unit, a 4G and Bluetooth dual-mode communication and positioning unit, a status switching module, and a call control unit integrated within the housing; the device surface is equipped with a DIP switch and a call button.

Citation Information

Patent Citations

  • Language learning auxiliary application system based on speech recognition

    CN120356458A

  • Spoken language pronunciation training correction system based on intelligent equipment

    CN121034351A