Running accompanying voice generation method and device, nonvolatile storage medium and electronic equipment

By integrating multi-dimensional data to generate personalized running voice prompts, the problem of inaccurate feedback and insufficient guidance in existing running voice prompt technologies has been solved, achieving more comprehensive exercise feedback and improved safety.

CN121789637APending Publication Date: 2026-04-03BEIJING CALORIE INFORMATION TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-31
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

In existing technologies, the real-time motion status feedback of voice-guided running is inaccurate and lacks personalized guidance, resulting in insufficient user exercise experience and safety.

Method used

By acquiring exercise physiological data, exercise status data, positioning data, environmental data, and static physiological characteristic data, and combining them with pre-set prompt word templates, the system generates current status text and feedback text based on target exercise ability index values, and finally outputs it in voice form to provide personalized exercise guidance.

Benefits of technology

It improves the comprehensiveness and accuracy of the voice prompts generated during running, enhances the user's exercise experience and safety, and ensures that the feedback content matches the user's specific exercise needs and physical condition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121789637A_ABST
    Figure CN121789637A_ABST
Patent Text Reader

Abstract

The invention discloses a running accompanying voice generation method and device, a nonvolatile storage medium and electronic equipment. The method comprises the following steps: acquiring motion physiological data, motion state data, positioning data, environment data and static physiological feature data of a target object under current motion; generating a current state text of the target object according to a preset cue word template based on the exercise physiological data, the exercise state data, the positioning data, the environment data and the static physiological feature data; based on the current state text and the target motion ability index value of the target object, generating a feedback text according to a prompt word template; and generating running accompanying voice of the target object based on the feedback text. According to the method and the device, the technical problems of inaccurate feedback of the generated running accompanying voice to the real-time motion state of the user and lack of personalized guidance in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of sports science, and more specifically, to a method, apparatus, non-volatile storage medium, and electronic device for generating accompaniment voice. Background Technology

[0002] With the development of smart wearable devices and mobile internet technology, running exercise monitoring devices and applications have become widespread, including smart bracelets, smartwatches, and mobile fitness apps. These technologies utilize GPS (Global Positioning System) modules, accelerometers, and photoelectric heart rate sensors in these devices and applications to collect real-time exercise physiological data such as latitude and longitude, cadence, pace, and heart rate. The data is then used to generate voice prompts at predetermined time intervals. When a specific exercise performance indicator exceeds a pre-set threshold, a simple audio prompt or vibration alarm is emitted. However, these technologies suffer from technical problems such as inaccurate real-time feedback on the user's exercise status and a lack of personalized guidance.

[0003] There is currently no effective solution to the above problems. Summary of the Invention

[0004] This application provides a method, apparatus, non-volatile storage medium, and electronic device for generating accompaniment voice, in order to at least solve the technical problems in the related art where the generated accompaniment voice does not provide accurate feedback on the user's real-time movement status and lacks personalized guidance.

[0005] According to one aspect of the embodiments of this application, a method for generating accompaniment voice is provided, comprising: acquiring the target object's motion physiological data, motion state data, positioning data, environmental data, and static physiological characteristic data under the current motion; generating the target object's current state text based on the motion physiological data, motion state data, positioning data, environmental data, and static physiological characteristic data, according to a pre-set prompt word template; generating feedback text based on the current state text and the target object's target motion ability index value, according to the prompt word template; and generating accompaniment voice for the target object based on the feedback text.

[0006] According to another aspect of the embodiments of this application, a running accompaniment voice generation device is provided, comprising: a data acquisition module, configured to acquire motion physiological data, motion state data, positioning data, environmental data, and static physiological characteristic data of a target object in its current motion; a first generation module, configured to generate current state text of the target object based on the motion physiological data, motion state data, positioning data, environmental data, and static physiological characteristic data, according to a pre-set prompt word template; a second generation module, configured to generate feedback text based on the current state text and the target object's target motion ability index value, according to the prompt word template; and a third generation module, configured to generate running accompaniment voice for the target object based on the feedback text.

[0007] According to another aspect of the embodiments of this application, a non-volatile storage medium is provided, which stores multiple instructions, any one of which is adapted to be loaded by a processor for a voice generation method for accompanying speech.

[0008] According to another aspect of the embodiments of this application, an electronic device is provided, including: one or more processors and a memory, the memory being used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement any one of the accompanying voice generation methods.

[0009] According to another aspect of the embodiments of this application, a computer program product is provided, which, when executed on a data processing device, is adapted to perform the steps of the accompanying voice generation method.

[0010] In this embodiment, the system acquires the target object's motion physiological data, motion state data, location data, environmental data, and static physiological characteristic data during its current movement. Based on these data, it generates the target object's current state text according to a pre-set prompt template. Then, based on the current state text and the target object's target athletic ability index, it generates feedback text according to the prompt template. Finally, it generates a pacing voice for the target object based on the feedback text. This achieves the goal of generating the target object's current state text by integrating its motion physiological data, motion state data, location data, environmental data, and static physiological characteristic data, combining this with the target object's target athletic ability index to generate feedback text, and then outputting it as voice to obtain the pacing voice. This improves the comprehensiveness and accuracy of the generated pacing voice, thereby enhancing the user's exercise experience and safety. It also solves the technical problems in related technologies where the generated pacing voice provides inaccurate feedback on the user's real-time motion state and lacks personalized guidance. Attached Figure Description

[0011] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:

[0012] Figure 1 This is a flowchart of a method for generating voice accompaniment during running, provided according to an embodiment of this application;

[0013] Figure 2 This is a flowchart of an optional voice generation system for accompanying runners, provided according to an embodiment of this application;

[0014] Figure 3 This is a schematic diagram of an optional voice generation device for accompanying a run, provided according to an embodiment of this application. Detailed Implementation

[0015] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present application.

[0016] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0017] According to an embodiment of this application, a method embodiment for generating voice for accompanying runners is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0018] Figure 1 This is a flowchart of a method for generating voice accompaniment during running, provided according to an embodiment of this application. Figure 1 As shown, the method includes the following steps:

[0019] Step S102: Obtain the target object's motion physiological data, motion state data, positioning data, environmental data, and static physiological characteristic data under the current motion.

[0020] It is understood that exercise physiological data may include, but is not limited to, the real-time heart rate (BPM, Beats Per Minute) and current heart rate zone of the user corresponding to the target object; exercise status data may include, but is not limited to, the user's real-time pace and real-time cadence; location data refers to the user's GPS trajectory; environmental data may include, but is not limited to, the weather conditions of the user's area; and static physiological characteristic data may include, but is not limited to, the user's age, gender, height, weight, BMI (Body Mass Index), and past injury history markers. Acquiring the above data provides rich data support for the subsequent generation of accompanying running voice, improving the comprehensiveness and accuracy of the generated results.

[0021] In one optional embodiment, acquiring the target object's motion physiological data, motion state data, location data, environmental data, and static physiological characteristic data under current motion includes: using a smart wearable device to acquire motion physiological data, motion state data, location data, and environmental data; and using a questionnaire survey and intelligent dialogue to acquire static physiological characteristic data.

[0022] It is understandable that the aforementioned exercise physiological data, exercise status data, location data, and environmental data can be acquired using smart wearable devices, such as smartwatches, smart bracelets, and smartphones. The aforementioned static physiological characteristic data can be collected through questionnaires and intelligent dialogue. Smart wearable devices can comprehensively detect physiological and environmental changes during exercise, while the collection of static physiological characteristic data provides a basis for personalized analysis and guidance. Combining this multi-dimensional data can generate feedback text that comprehensively covers the user's exercise status, improving the comprehensiveness and accuracy of the accompanying running voice generation results.

[0023] Optionally, smart wearable devices are equipped with a variety of sensors, such as photoelectric heart rate sensors, accelerometers, gyroscopes, GPS modules, and environmental sensors (such as temperature sensors, humidity sensors, barometric pressure sensors, and light sensors). Through the smart wearable devices worn by the user, accurate collection of the user's exercise physiological data, exercise status data, positioning data, and environmental data of the area where the user is located can be achieved.

[0024] Optionally, users' static physiological characteristics data can be collected through exercise platforms (such as the Keep APP) or specific human-computer interaction interfaces to ensure the privacy and security of user data and data quality.

[0025] Optionally, user exercise physiological data, exercise status data, location data, environmental data, and static physiological characteristic data can be obtained through smart wearable devices (such as smartwatches or smartphones) or mobile terminal sensors (deployed in the smart wearable device worn by the user), questionnaires, and intelligent dialogue. Specifically: smartwatches can collect dynamic physiological data (i.e., exercise physiological data) such as the user's real-time heart rate (BPM) and the current heart rate zone; smartwatches or other smart wearable devices can collect dynamic kinematic data (i.e., exercise status data) such as the user's real-time pace and real-time cadence; smartwatches or other smart wearable devices can collect the user's GPS trajectory (i.e., location data) and the weather conditions of the user's area (i.e., environmental data); and questionnaires and intelligent dialogue can collect basic static data (i.e., static physiological characteristic data) such as the user's age, gender, height, weight, BMI index, and past injury history markers.

[0026] Step S104: Based on exercise physiological data, exercise state data, positioning data, environmental data, and static physiological characteristic data, generate the current status text of the target object according to the pre-set prompt word template;

[0027] It is understandable that by accurately integrating data and generating current status text according to prompt word templates, the real-time movement status of the user corresponding to the target object can be accurately captured, reducing information distortion and false alarms, making the subsequently generated feedback text closer to the user's actual movement status, and improving the accuracy of the accompanying voice.

[0028] In an optional embodiment, before generating the current state text of the target object based on exercise physiological data, exercise state data, positioning data, environmental data, and static physiological characteristic data according to a pre-set prompt word template, the method further includes: determining the role setting and language style corresponding to the prompt word template based on the target object's motor ability; and determining the prompt word template based on the role setting, language style, and output format.

[0029] Understandably, based on the target user's athletic ability (e.g., beginner, intermediate, and professional runners), the role setting and language style corresponding to the prompt word template are determined. For example, for a beginner runner, a gentle coach role setting is adopted, emphasizing companionship and encouragement, with a relaxed and motivating language style; while for a professional runner, a professional coach role setting is adopted, focusing on technical details and training guidance, with a formal and instructive language style. Based on the role setting, language style, and output format corresponding to the prompt word template, the prompt word template for the generated text is determined. The output format may include, but is not limited to, text structure, keyword usage, and text expression. Through the customization of role setting and language style, the generated feedback text not only includes the user's exercise data report but also incorporates professional guidance and emotional support, helping users adjust their exercise strategies, reduce exercise-related injuries, and enhance their exercise motivation and experience.

[0030] Optionally, a large language model can be used to output user feedback. Before using a large language model to output user feedback, the data input to the large language model needs to be converted into text that the model can understand. This involves pre-setting a prompt word template, organizing the data input to the large language model according to the prompt word template format, and obtaining the input text, including text about athletic ability and current status. Simultaneously, the large language model generates feedback text based on the input text, following the prompt word template format. The prompt word template can be determined based on the role persona (i.e., role setting) and language style (e.g., professional and encouraging) of the "professional AI running coach" (i.e., the AI ​​running coach corresponding to the running voice), as well as the output format requirements. For example, the role setting can be determined based on the user's athletic ability: a gentle coach role setting suitable for beginner runners ("few technical terms, more companionship and encouragement"), a guiding coach role setting suitable for intermediate runners ("more guidance and more motivation"), and a professional coach role setting suitable for professional runners ("respect, understanding, support, and recognition").

[0031] Optionally, the collected multi-source data (including exercise physiological data, exercise state data, location data, environmental data, and static physiological characteristic data) can be semantically processed according to the prompt word template to generate current status text. For example, "Today the weather is sunny, the temperature is 18 degrees Celsius, the user is female, 26 years old, the exercise goal is fat loss, the user's current heart rate is 144 beats / minute, which is within the aerobic endurance zone, the current cadence is 155 steps / minute, the pace is 5 minutes and 10 seconds, which is within the easy running zone."

[0032] Step S106: Based on the current status text and the target motion capability index value of the target object, generate feedback text according to the prompt word template;

[0033] It is understandable that feedback text may include, but is not limited to, safety warnings, technical guidance, emotional encouragement, and tool call instructions. By combining the user's current status text with the target athletic ability index value, the feedback text generated can comprehensively cover the user's various needs during exercise. It can not only accurately reflect the user's real-time exercise data, but also provide targeted guidance and suggestions, improve the comprehensiveness of subsequent voice-guided running results, and enhance the user's exercise experience and safety.

[0034] In an optional embodiment, before generating feedback text based on the current state text and the target motion capability index value of the target object according to the prompt word template, the method further includes: obtaining historical motion data of the target object; and determining the target motion capability index value based on the historical motion data and static physiological characteristic data.

[0035] It's understandable that acquiring historical exercise data from the target user—such as average pace, average heart rate, average cadence, and longest running distance over a specific period—and combining this data with the user's static physiological characteristics, allows for the determination of the user's target exercise ability index. By analyzing historical exercise data and combining it with static physiological characteristics, more personalized target exercise ability index values ​​can be set, making the generated coaching voice more closely reflect the user's actual situation. Furthermore, the use of historical exercise data ensures that the target exercise ability index value is set based on the user's long-term exercise performance, rather than a single exercise state, improving the accuracy and rationality of the determined result and avoiding setting the target exercise ability index value too low or too high due to short-term fluctuations.

[0036] In one optional embodiment, determining the target motor ability index value based on historical motion data and static physiological characteristic data includes: determining the initial motor ability index value of the target object based on static physiological characteristic data; and correcting the initial motor ability index value based on historical motion data to obtain the target motor ability index value.

[0037] Understandably, based on a user's static physiological characteristics data, initial exercise capacity indicators are determined, such as baseline heart rate range, suitable exercise intensity range, and expected maximum oxygen uptake. These initial indicators are then adjusted using the user's historical exercise data to arrive at the target exercise capacity indicators. For example, if a user has shown significant improvement in physical fitness over a specific period, the suggested heart rate and pace ranges may be adjusted to reflect their current higher exercise capacity; conversely, if a user shows a decrease in exercise frequency or intensity, they can be advised to reduce exercise intensity and duration to avoid overtraining. Personalized setting of target exercise capacity indicators ensures that the generated coaching voice can provide more targeted guidance based on the user's specific exercise needs and physical condition, increasing user acceptance of the guidance and ensuring exercise safety.

[0038] Optionally, a pre-defined sports science algorithm model can be used to calculate the user's target athletic ability index value by combining the user's historical exercise data and basic static data. First, the acquired user's historical exercise data and basic static data are preprocessed, including data cleaning (removing outliers and filling in missing values), data format conversion (ensuring all data is stored in a uniform format for easy calculation), and data standardization (converting data of different dimensions to the same range, such as standardizing heart rate and pace, to facilitate model understanding and processing). Then, using the preprocessed data, the user's athletic ability model is established through the sports science algorithm model to determine the user's target athletic ability index value. Exercise capacity models may include, but are not limited to, VO2 max prediction models, heart rate zone calculation models, and exercise adaptability assessment models. The VO2 max prediction model uses data such as age, gender, weight, and longest historical running distance to predict a user's suitable VO2 max, reflecting their aerobic exercise capacity. The heart rate zone calculation model uses the user's baseline static data and historical exercise data to calculate the user's heart rate zones, including heart rate ranges corresponding to different intensities of exercise such as fat burning, aerobic endurance, and anaerobic power. The exercise adaptability assessment model uses the user's historical exercise volume, exercise frequency, and recovery data to assess the user's suitable exercise volume range.

[0039] In one optional embodiment, based on the current state text and the target motion capability index value of the target object, feedback text is generated according to the prompt word template, including: generating motion capability text of the target object based on the target motion capability index value and according to the prompt word template; and obtaining feedback text based on the current state text and motion capability text and according to the prompt word template.

[0040] Understandably, based on the target athletic ability index value and following the prompt word template, the system generates text describing the user's athletic ability corresponding to the target object. Combined with the user's current status text, and following the prompt word template again, it generates feedback text for the user. By combining and analyzing the real-time current status text and athletic ability text, a comprehensive assessment of the user's athletic performance can be achieved. This ensures that the generated feedback content not only covers the user's current athletic status but is also closely related to their athletic ability, improving the comprehensiveness and accuracy of the feedback.

[0041] Optionally, the user's target athletic ability index value can be semantically processed according to the prompt word template to generate athletic ability text. For example, "This user has excellent running ability. The cumulative running distance this week is 1km, which is insufficient. The recommended cumulative running distance next week is 18-32km. The current exercise status is declining physical fitness. It is recommended to start training."

[0042] In one optional embodiment, feedback text is obtained based on the current state text and the exercise ability text, according to the prompt word template. This includes: performing deviation analysis and index correlation analysis based on the current state text and the exercise ability text to obtain exercise ability index deviation results and exercise ability index correlation results, wherein the exercise ability index correlation results are used to characterize the mutual influence relationship between exercise ability indicators; and generating feedback text based on the exercise ability index deviation results and the exercise ability index correlation results, according to the prompt word template.

[0043] It is understandable that a deviation analysis is performed between the actual values ​​of the user's current status indicators (included in the current status text) and the target fitness indicator values ​​(included in the fitness indicator text). This deviation result indicates whether a deviation exists and to what extent. Furthermore, a correlation analysis is conducted between the current status text and the fitness indicator text to obtain a correlation result representing the interrelationships between fitness indicators. Based on these deviation and correlation results, specific safety warnings, technical guidance, emotional encouragement, and tool activation instructions are generated. This information is then organized according to a prompt template to generate feedback text. Through deviation and correlation analysis, the feedback text comprehensively considers the user's fitness status, focusing not only on individual fitness indicators but also on the interactions between them, thus providing users with more scientific and comprehensive fitness guidance.

[0044] Optionally, a large language model can be used to generate the aforementioned feedback text. The user's exercise ability text and current status text are input, and the semantic understanding and logical reasoning capabilities of the large language model are used for analysis to generate feedback content. When generating feedback text using the large language model, it performs both reasoning and generation tasks. The reasoning task refers to the large language model comprehensively analyzing the real-time data included in the current status text, its deviation from the exercise ability baseline (i.e., the target exercise ability index value) included in the exercise ability text, and the correlation between various exercise ability indices, to obtain the exercise ability index deviation results and the exercise ability index correlation results. The generation task refers to generating natural language feedback text consistent with the coach's role persona end-to-end based on the reasoning conclusions (including the exercise ability index deviation results and the exercise ability index correlation results). The content of the feedback text includes, but is not limited to, safety warnings, technical guidance, emotional encouragement, and tool call instructions. For example, if a user starts running at a pace far exceeding the pace range of a moderate warm-up or easy run during the warm-up phase, and their heart rate is too high, a message will be generated reminding the user that the pace is too high and instructing them to start the warm-up phase at a slower, more comfortable pace to ensure a more sustained and safer running experience later on. The tool will also send a cadence metronome to the user to guide them on their cadence over a period of time.

[0045] Step S108: Based on the feedback text, generate the accompanying voice for the target object.

[0046] It's understandable that the feedback text is used to generate a voice prompt for the target user during the run. By outputting the feedback text in voice format, users can reduce their reliance on the device screen, allowing them to focus more on the exercise itself, especially outdoors, at night, or in conditions with poor visibility, thus improving safety.

[0047] Optionally, a high-quality text-to-speech (TTS) engine that replicates the voice of Keep's coaches can be used to synthesize speech, generating a running voice with a certain emotional tone, which can then be broadcast in real time through the user's headphones or terminal speaker.

[0048] It should be noted that the data collected and acquired in this application (including but not limited to the user's motion physiological data, motion state data, location data, environmental data, static physiological characteristic data, and historical motion data, etc.) are data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of the relevant data all comply with relevant laws, regulations, and standards, necessary confidentiality measures have been taken, they do not violate public order and good morals, and corresponding operation entry points are provided for users to choose to authorize or refuse.

[0049] Through the above steps S102 to S108, the goal is to generate the target object's current state text by integrating the target object's motion physiological data, motion state data, positioning data, environmental data, and static physiological characteristic data, and to generate feedback text by combining the target object's target motion ability indicators and outputting it as voice, thereby obtaining the target object's accompaniment voice. This achieves the technical effect of improving the comprehensiveness and accuracy of the accompaniment voice generation results, thereby enhancing the user's exercise experience and the safety of the user's exercise, and thus solving the technical problems of inaccurate feedback on the user's real-time motion state and lack of personalized guidance in related technologies.

[0050] Based on the above embodiments and optional embodiments, this application proposes an optional implementation of a pacing voice generation system to implement the above-mentioned pacing voice generation method, generate comprehensive and accurate pacing voice, and improve the user's sports experience and sports safety.

[0051] The relevant technologies employ the following methods to provide feedback on the user's exercise status. First, there is visual feedback, where basic data such as latitude and longitude, cadence, pace, and heart rate are displayed in real-time on the user's smart wearable device, such as a mobile phone screen, for the user to view. Second, there is basic voice broadcasting, which mechanically broadcasts data such as current exercise distance, time, and average pace at preset time intervals (e.g., every kilometer, every 10 minutes). When the data value corresponding to a single exercise ability indicator (e.g., heart rate) exceeds a preset threshold, a simple prompt sound or vibration alarm is emitted.

[0052] The methods for providing feedback on user exercise status in related technologies have the following shortcomings. First, the data processing is limited in scope and lacks comprehensive analytical capabilities. Existing technologies often focus on threshold detection for "single exercise ability indicators" (such as alarms only for excessive heart rate or prompts only for low pace), failing to perform multi-dimensional joint analysis of basic user data (such as age, BMI, exercise background, and recent exercise status) and real-time dynamic data (heart rate, cadence, and pace). Second, the feedback content is mechanical and lacks guidance and motivation. The accompanying voice prompts generated in related technologies often remain at the "counting" stage, simply informing users "how much they've run," without providing incentive mechanisms or specific guidance. For example, when a user's exercise status declines, positive verbal encouragement based on their current level of commitment cannot be provided; when a user's movements become distorted (such as abnormal cadence-to-pace ratio), no actionable adjustment plan can be given, leaving users only with data but not knowing how to better adjust their running status. Finally, there is a high dependence on the screen of smart wearable devices, posing security risks and creating a disconnect between the user experience and the exercise experience. Because the generated accompanying audio has limited information and lacks depth, users need to frequently check the screens of smart wearable devices such as phones or watches to fully understand their own condition. This not only interrupts the user's exercise rhythm, disrupts the coordination of breathing and stride frequency, and affects the continuity and experience of exercise, but also increases safety risks during exercise. For example, due to the user's distraction, the probability of accidents such as collisions and falls increases, making it impossible to achieve an "immersive" screenless exercise experience.

[0053] Figure 2 This is a flowchart of an optional voice generation system for accompanying runners, provided according to an embodiment of this application, such as... Figure 2 As shown, the voice generation system for accompaniment running includes a multi-source data perception module, a user's athletic ability assessment module, a context construction and prompt word generation module, a large model reasoning and generation module, and a speech synthesis and feedback execution module, which will be introduced below.

[0054] The multi-source data sensing module is responsible for real-time collection and integration of multi-dimensional raw data streams from the user during movement.

[0055] User motion physiological data, motion status data, location data, environmental data, and static physiological characteristic data are acquired through smart wearable devices (such as smartwatches or smartphones) or mobile terminal sensors (deployed in the smart wearable devices worn by the user), questionnaires, and intelligent dialogue. Specifically:

[0056] The smartwatch collects dynamic physiological data (i.e., exercise physiological data) such as the user's real-time heart rate (BPM) and the current heart rate zone.

[0057] The system collects users' real-time kinematic dynamic data (i.e., motion status data) such as pace and cadence through smart wearable devices such as smartphones or smartwatches.

[0058] The system collects the user's GPS trajectory (i.e., location data) and the weather conditions (i.e., environmental data) of the user's location through smart wearable devices such as smartphones or smartwatches.

[0059] The system collects basic static data (i.e., static physiological characteristic data) of users, such as age, gender, height, weight, BMI index, and past injury and illness history markers, through questionnaires and intelligent dialogue.

[0060] The user's athletic ability assessment module, as a pre-calculation unit, is used to establish a personalized baseline for the user's athletic ability (i.e., the target athletic ability index value) based on the user's historical athletic data and basic static data, providing a basis for the subsequent judgment of the deviation of the athletic ability index by the large language model.

[0061] The input data for this module consists of the user's basic static data and the user's historical exercise data (including statistical indicators such as average pace, average heart rate, average cadence, and longest running distance over a specific period in the past).

[0062] The processing logic of this module is to use a preset sports science algorithm model, combined with the user's historical sports data and basic static data, to calculate the user's target sports ability index value.

[0063] The context building and prompt generation module acts as a "semantic translator" between data and the large language model, transforming structured numerical data into natural language prompts that the large language model can understand, including motion ability text and current status text.

[0064] The system pre-defines a template (i.e., a prompt word template) for the text generated by the LLM (Large Language Model). The prompt word template is determined based on the role profile (i.e., character setting), language style (e.g., professional and encouraging), and output format requirements of the "professional AI running coach" (i.e., the AI ​​running coach corresponding to the running voice).

[0065] For example, the role setting can be determined according to the user's athletic ability. There is a gentle coach role setting that is suitable for beginner runners with "less technical jargon and more companionship and encouragement", a guiding coach role setting that is suitable for intermediate runners with "more guidance and more motivation", and a professional coach role setting that is suitable for professional runners with "respect, understanding, support and recognition".

[0066] Generate exercise ability text by semantically representing the user's target exercise ability index value output by the user exercise ability assessment module according to the prompt word template. For example, "This user has excellent running ability. The cumulative running distance this week is 1km, which is insufficient. The recommended cumulative running distance next week is 18-32km. The current exercise status is declining physical fitness. It is recommended to start training."

[0067] Generate current status text by collecting multi-source data from the multi-source data perception module, semantically processing it according to the prompt word template, and generating current status text. For example, "Today the weather is sunny, the temperature is 18 degrees Celsius, the user is female, 26 years old, the exercise goal is fat loss, the user's current heart rate is 144 beats / minute, which is in the aerobic endurance zone, the current cadence is 155 steps / minute, the pace is 5 minutes and 10 seconds, which is in the easy running zone."

[0068] The large-scale model reasoning and generation module, serving as the core decision-making hub of the voice generation system, receives prompts from the context construction and prompt generation module, including text on motor ability and current state. It then analyzes these prompts using the semantic understanding and logical reasoning capabilities of the large language model and generates feedback content. When generating feedback text using the large language model, it performs both reasoning and generation tasks.

[0069] The reasoning task refers to the comprehensive analysis of the real-time data included in the current state text by the large language model, the deviation of the data from the baseline of the motor ability included in the motor ability text, and the correlation between various motor ability indicators, to obtain the results of the deviation of motor ability indicators and the correlation of motor ability indicators.

[0070] The generation task refers to the large language model generating natural language feedback text end-to-end that aligns with the coach's role based on inference conclusions (including deviation results and correlation results of exercise ability indicators). The feedback text includes, but is not limited to, safety warnings, technical guidance, emotional encouragement, and tool invocation instructions. For example, if a user starts running at a pace far exceeding the moderate warm-up or easy run pace range during the warm-up phase, and their heart rate is too high, the system will generate feedback text reminding the user that their pace is too high and advising them to start the warm-up phase at a slower, more comfortable pace to ensure a more sustained and safer running experience later. The system will also issue a cadence metronome to the user via a tool invocation instruction to guide their cadence over a period of time.

[0071] The speech synthesis and feedback execution module converts the feedback text generated by the large language model into accompanying speech and plays it back.

[0072] The feedback text output by the large language model inference and generation module is used to synthesize speech by calling a high-quality text-to-speech (TTS) engine that replicates the voice of the Keep coach. This generates a running voice with a certain emotional tone, which is then broadcast in real time through the user's headphones or terminal speaker.

[0073] The above optional implementation methods achieve at least the following effects: Using context-aware semantic reasoning to generate feedback text improves the accuracy and comprehensiveness of the feedback text; personalized setting of target athletic ability index values ​​ensures that the generated accompanying voice can provide more targeted guidance based on the user's specific athletic needs and physical condition, increasing the user's acceptance of the exercise guidance and ensuring the user's exercise safety; LLM generates feedback text in real time based on the AI ​​accompanying coach's persona, featuring a natural tone and specific guidance details, exhibiting high anthropomorphism and richness, thus improving the user's exercise experience.

[0074] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0075] This embodiment also provides a voice generation device for accompanying a runner, which is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the terms "module" and "device" can refer to a combination of software and / or hardware that performs a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, hardware implementations, or a combination of software and hardware, are also possible and contemplated.

[0076] According to an embodiment of this application, an apparatus embodiment for implementing the accompanying voice generation method is also provided. Figure 3 This is a schematic diagram of a voice generation device for accompanying a runner according to an embodiment of this application, such as... Figure 3 As shown, the above-mentioned voice generation device for accompanying runners includes a data acquisition module 302, a first generation module 304, a second generation module 306, and a third generation module 308. The device will be described below.

[0077] The data acquisition module 302 is used to acquire the target object's motion physiological data, motion state data, positioning data, environmental data, and static physiological characteristic data under the current motion.

[0078] The first generation module 304 is connected to the data acquisition module 302 and is used to generate the current status text of the target object based on exercise physiological data, exercise state data, positioning data, environmental data, and static physiological characteristic data, according to a pre-set prompt word template.

[0079] The second generation module 306, connected to the first generation module 304, is used to generate feedback text based on the current state text and the target motion capability index value of the target object, according to the prompt word template.

[0080] The third generation module 308, connected to the second generation module 306, is used to generate accompanying voice for the target object based on the feedback text.

[0081] The accompanying running voice generation device provided in this application embodiment, by setting a data acquisition module 302, a first generation module 304, a second generation module 306, and a third generation module 308, achieves the purpose of integrating the target object's motion physiological data, motion state data, positioning data, environmental data, and static physiological characteristic data to generate the target object's current state text, combining the target object's target motion ability indicators to generate feedback text, and outputting it as voice, thereby obtaining the target object's accompanying running voice. This achieves the technical effect of improving the comprehensiveness and accuracy of the accompanying running voice generation results, thereby enhancing the user's exercise experience and the safety of the user's exercise, and thus solving the technical problems of inaccurate feedback on the user's real-time motion status and lack of personalized guidance in related technologies.

[0082] It should be noted that the above modules can be implemented by software or hardware. For example, for the latter, it can be implemented in the following ways: the above modules can be located in the same processor; or the above modules can be located in different processors in any combination.

[0083] It should be noted that the data acquisition module 302, the first generation module 304, the second generation module 306, and the third generation module 308 mentioned above correspond to steps S102 to S108 in the embodiments. The instances and application scenarios implemented by the above modules and their corresponding steps are the same, but they are not limited to the content disclosed in the above embodiments. It should be noted that the above modules, as part of the device, can run on a computer terminal.

[0084] It should be noted that the optional or preferred implementation methods of this embodiment can be found in the relevant descriptions in the embodiments, and will not be repeated here.

[0085] The aforementioned voice generation device for accompanying runners may also include a processor and a memory. The data acquisition module 302, the first generation module 304, the second generation module 306, the third generation module 308, etc., are all stored in the memory as program units, and the processor executes the aforementioned program units stored in the memory to realize the corresponding functions.

[0086] The processor contains a core that retrieves the corresponding program unit from memory. One or more cores may be configured. Memory may include non-persistent memory in computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory includes at least one memory chip.

[0087] This application provides a non-volatile storage medium storing a program that, when executed by a processor, implements a method for generating accompaniment voice.

[0088] This application provides an electronic device, which includes a processor, a memory, and a program stored in the memory and executable on the processor. When the processor executes the program, it performs the following steps: acquiring motion physiological data, motion state data, positioning data, environmental data, and static physiological characteristic data of a target object in its current motion; generating current state text of the target object based on the motion physiological data, motion state data, positioning data, environmental data, and static physiological characteristic data, according to a pre-set prompt word template; generating feedback text based on the current state text and the target object's target motion ability index value, according to the prompt word template; and generating accompanying voice for the target object based on the feedback text. The device in this document can be a server, PC, etc.

[0089] This application also provides a computer program product, which, when executed on a data processing device, is suitable for executing an initialization program with the following method steps: acquiring the target object's motion physiological data, motion state data, positioning data, environmental data, and static physiological characteristic data under the current motion; generating the target object's current state text according to a pre-set prompt word template based on the motion physiological data, motion state data, positioning data, environmental data, and static physiological characteristic data; generating feedback text according to the prompt word template based on the current state text and the target object's target motion ability index value; and generating the target object's accompaniment voice based on the feedback text.

[0090] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0091] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0092] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0093] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0094] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0095] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0096] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0097] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0098] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0099] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A method for generating voice prompts for accompanying runners, characterized in that, include: Acquire motion physiological data, motion state data, location data, environmental data, and static physiological characteristic data of the target object under its current motion. Based on the aforementioned motion physiological data, motion state data, positioning data, environmental data, and static physiological characteristic data, the current status text of the target object is generated according to a pre-set prompt word template. Based on the current status text and the target motion capability index value of the target object, feedback text is generated according to the prompt word template; Based on the feedback text, a voice message is generated to accompany the target object during the run.

2. The method according to claim 1, characterized in that, The acquisition of the target object's motion physiological data, motion state data, positioning data, environmental data, and static physiological characteristic data under current motion includes: The exercise physiological data, the exercise state data, the positioning data, and the environmental data are acquired using smart wearable devices. The static physiological characteristic data were obtained by using questionnaires and intelligent dialogue.

3. The method according to claim 1, characterized in that, Before generating the current status text of the target object based on the motion physiological data, the motion state data, the positioning data, the environmental data, and the static physiological feature data, according to a pre-set prompt word template, the method further includes: Based on the target object's mobility, determine the character setting and language style corresponding to the prompt word template; Based on the character settings, language style, and output format, the prompt word template is determined.

4. The method according to claim 1, characterized in that, Before generating feedback text based on the current state text and the target motion capability index value of the target object according to the prompt word template, the method further includes: Obtain the historical motion data of the target object; Based on the historical motion data and the static physiological characteristic data, the target motion ability index value is determined.

5. The method according to claim 4, characterized in that, The process of determining the target motor ability index value based on the historical motion data and the static physiological characteristic data includes: Based on the static physiological characteristic data, the initial motor ability index value of the target object is determined; Based on the historical motion data, the initial motion ability index value is corrected to obtain the target motion ability index value.

6. The method according to any one of claims 1 to 5, characterized in that, The step of generating feedback text based on the current state text and the target motion capability index value of the target object, according to the prompt word template, includes: Based on the target mobility index value, the mobility text of the target object is generated according to the prompt word template; Based on the current status text and the mobility text, the feedback text is obtained according to the prompt word template.

7. The method according to claim 6, characterized in that, The feedback text, obtained based on the current state text and the mobility text, and according to the prompt word template, includes: Based on the current status text and the athletic ability text, deviation analysis and index correlation analysis are performed to obtain athletic ability index deviation results and athletic ability index correlation results, wherein the athletic ability index correlation results are used to characterize the mutual influence relationship between athletic ability indicators. Based on the deviation results of the athletic ability indicators and the correlation results of the athletic ability indicators, the feedback text is generated according to the prompt word template.

8. A voice generation device for accompanying runners, characterized in that, include: The data acquisition module is used to acquire the target object's motion physiological data, motion state data, positioning data, environmental data, and static physiological characteristic data under the current motion. The first generation module is used to generate the current status text of the target object based on the exercise physiological data, the exercise state data, the positioning data, the environmental data, and the static physiological feature data, according to a pre-set prompt word template. The second generation module is used to generate feedback text based on the current state text and the target motion capability index value of the target object, according to the prompt word template. The third generation module is used to generate the accompanying voice for the target object based on the feedback text.

9. A non-volatile storage medium, characterized in that, The non-volatile storage medium stores multiple instructions, which are adapted to be loaded by a processor and executed by the accompanying voice generation method according to any one of claims 1 to 7.

10. An electronic device, characterized in that, include: One or more processors and a memory, the memory being used to store one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the accompanying voice generation method according to any one of claims 1 to 7.