A voice-guided interaction method, device, system, and medium

By using voice timbre recognition and artificial intelligence model conversion technology, efficient and visual communication between midwives and mothers during childbirth has been achieved, overcoming the shortcomings of traditional verbal guidance and improving communication efficiency and intuitiveness.

CN120220718BActive Publication Date: 2026-03-06SHUNDE HOSPITAL SOUTHERN MEDICAL UNIV (THE FIRST PEOPLES HOSPITAL OF SHUNDE FOSHAN)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510242888.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2026-03-06
Estimated Expiration
2045-03-03

AI Technical Summary

Technical Problem

Traditional childbirth guidance relies on verbal instructions from doctors, which makes it difficult for pregnant women to understand, inconvenient to operate, and causes them great psychological stress, thus affecting communication efficiency.

Method used

The role of the midwife or mother is determined by recognizing the timbre of the voice signal. Artificial intelligence models are used to convert the guidance or control voice into text information, which controls the playback device to play the corresponding video images or adjust the video status, thus achieving visual guidance and control.

Benefits of technology

It improves the efficiency of communication between midwives and mothers, enhances the intuitiveness and ease of operation of guidance, and reduces the psychological stress of pregnant women.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220718B_ABST
    Figure CN120220718B_ABST
Patent Text Reader

Abstract

This invention discloses a voice-guided interaction method, device, system, and medium. The method includes: acquiring a voice signal; determining the role of the voice sender based on the timbre of the voice signal; when the voice sender's role is determined to be a first preset role, controlling a playback device according to the guidance voice to play the video image corresponding to the guidance voice; when the voice sender's role is determined to be a second preset role, controlling the playback device according to the control voice to control the currently playing video image. By setting two different preset roles and judging the timbre of the voice signal, it is determined whether the spoken language has a guiding or controlling meaning. The playback device is then controlled according to the different meanings. By visually displaying the meaning of the guidance voice, the communication efficiency between different roles is improved. This invention is mainly used in the field of intelligent interaction technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent human-computer interaction technology, specifically to an interaction method, device, system, and medium based on voice guidance. Background Technology

[0002] Intelligent human-computer interaction refers to the information exchange process between humans and computers using certain dialogue voice and interactive methods to complete a specific task. Intelligent human-computer interaction can improve the communication efficiency between two parties. It primarily focuses on the interactive relationship between the system (machine or computerized system) and the user.

[0003] During childbirth, pregnant women may not fully understand the instructions from doctors or midwives due to tension or pain. Traditional childbirth guidance methods mainly rely on verbal instructions from doctors, which have the following problems: 1. Insufficiently intuitive instructions: Verbal descriptions are difficult for pregnant women to quickly understand the correct actions. 2. Inconvenient operation: Midwives or doctors find it difficult to concentrate on operating equipment during the busy childbirth process. 3. High psychological stress: Pregnant women are prone to tension and anxiety during childbirth.

[0004] Therefore, how to utilize intelligent human-computer interaction technology to improve the communication efficiency between midwives (obstetricians) and expectant mothers is a technical issue that urgently needs to be studied in the industry. Summary of the Invention

[0005] This invention provides a voice-guided interaction method, apparatus, system, and medium to solve one or more technical problems existing in the prior art, and at least provide a beneficial option or create conditions.

[0006] This invention provides an interactive method based on voice guidance, comprising: acquiring a voice signal, and determining the role of the voice speaker based on the timbre of the voice signal;

[0007] When the role of the speaker is determined to be a first preset role, the voice signal is recorded as guidance voice; the playback device is controlled according to the guidance voice so that the playback device plays the video image corresponding to the guidance voice;

[0008] When the role of the speaker is determined to be the second preset role, the voice signal is recorded as control voice; the playback device is controlled according to the control voice so that the playback device controls the currently playing video image.

[0009] Furthermore, controlling the playback device according to the guidance voice so that the playback device plays the video image corresponding to the guidance voice specifically includes: converting the guidance voice into text information, wherein the text information records application scenario information and instruction information;

[0010] The text information is semantically recognized by a scene recognition model to obtain the target application scene information.

[0011] The text information is semantically recognized by an indication recognition model to obtain target indication information;

[0012] Based on the target application scenario information, find the video image file whose type tag information contains the target application scenario information, and find the corresponding first target video summary file based on the video image file;

[0013] In this system, video summary text files are bound one-to-one with video image files. The text content of the video summary text file is obtained by content recognition of the video image file it is bound to through a video content recognition model. The type label information of the video image file is preset.

[0014] Based on the target indication information, find the second target video summary file containing the target indication information from the first target video summary file;

[0015] The target video image file corresponding to the second target video summary file is determined;

[0016] The target video file is transmitted to the playback device so that the playback device can play the target video file.

[0017] Furthermore, controlling the playback device according to the control voice so that the playback device controls the currently playing video image specifically includes: converting the control voice into text information, wherein the text information records control parameter information;

[0018] The text information is identified by a control parameter recognition model to obtain control parameters, which are then recorded as target control parameters.

[0019] The target control parameters are transmitted to the playback device so that the playback device controls the video image currently playing on the display interface according to the target control parameters;

[0020] The target control parameters include: increasing volume, decreasing volume, pausing playback, or resuming playback.

[0021] Furthermore, determining the role of the speaker based on the timbre of the speech signal specifically includes: passing the speech signal and the pre-stored speech of the role to a timbre matching model, performing timbre matching on the speech signal and the pre-stored speech of the role through the timbre matching model to obtain a matching result, and determining the role of the speaker based on the matching result.

[0022] Furthermore, the playback device includes: a smart TV, a smart projector, a smart playback device, or VR glasses.

[0023] On the other hand, a voice-guided interactive device is provided, comprising: a processor and a memory, the memory being used to store a computer-readable program; when the computer-readable program is executed by the processor, the processor causes the processor to implement the voice-guided interactive method as described in any of the above technical solutions.

[0024] On the other hand, a voice-guided interactive system is provided, comprising: an acquisition module, a first determination module, and a second determination module;

[0025] The acquisition module is used to: acquire a voice signal and determine the role of the voice speaker based on the timbre of the voice signal;

[0026] The first determining module is used to: when the role of the speaker is determined to be a first preset role, record the voice signal as guidance voice; control the playback device according to the guidance voice so that the playback device plays the video image corresponding to the guidance voice;

[0027] The second determining module is used to: when the role of the voice sender is determined to be a second preset role, record the voice signal as control voice; control the playback device according to the control voice, so that the playback device controls the currently playing video image.

[0028] Furthermore, controlling the playback device according to the guidance voice so that the playback device plays the video image corresponding to the guidance voice specifically includes: converting the guidance voice into text information, wherein the text information records application scenario information and instruction information;

[0029] The text information is semantically recognized by a scene recognition model to obtain the target application scene information.

[0030] The text information is semantically recognized by an indication recognition model to obtain target indication information;

[0031] Based on the target application scenario information, find the video image file whose type tag information contains the target application scenario information, and find the corresponding first target video summary file based on the video image file;

[0032] In this system, video summary text files are bound one-to-one with video image files. The text content of the video summary text file is obtained by content recognition of the video image file it is bound to through a video content recognition model. The type label information of the video image file is preset.

[0033] Based on the target indication information, find the second target video summary file containing the target indication information from the first target video summary file;

[0034] The target video image file corresponding to the second target video summary file is determined;

[0035] The target video file is transmitted to the playback device so that the playback device can play the target video file.

[0036] Furthermore, controlling the playback device according to the control voice so that the playback device controls the currently playing video image specifically includes: converting the control voice into text information, wherein the text information records control parameter information;

[0037] The text information is identified by a control parameter recognition model to obtain control parameters, which are then recorded as target control parameters.

[0038] The target control parameters are transmitted to the playback device so that the playback device controls the video image currently playing on the display interface according to the target control parameters;

[0039] The target control parameters include: increasing volume, decreasing volume, pausing playback, or resuming playback.

[0040] On the other hand, a computer-readable storage medium is provided, characterized in that it stores a processor-executable program, which, when executed by a processor, is used to implement the voice-guided interaction method as described in any of the above technical solutions.

[0041] The present invention has at least the following beneficial effects: The method of the present invention sets up two different preset roles and determines whether the language spoken by the current speaker has a guiding or controlling meaning by judging the timbre of the speech signal. The playback device is then controlled according to the different meanings. Furthermore, by visually displaying the meaning of the guiding speech, the communication efficiency between different roles is improved. This allows for effective guidance and communication between midwives and mothers in certain scenarios, such as childbirth, improving communication efficiency. The present invention also provides corresponding devices, systems, and media, the beneficial effects of which are similar to the method and will not be repeated here. Attached Figure Description

[0042] The accompanying drawings are provided to further understand the technical solutions of the present invention and constitute a part of the specification. They are used together with the embodiments of the present invention to explain the technical solutions of the present invention, and do not constitute a limitation on the technical solutions of the present invention.

[0043] Figure 1 This is a flowchart of the steps involved in a voice-guided interaction method.

[0044] Figure 2 This is a schematic diagram of the device structure of a voice-guided interactive device;

[0045] Figure 3 This is a schematic diagram of the system connection structure of a voice-guided interactive system. Detailed Implementation

[0046] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0047] It should be noted that although functional modules are divided in the system diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the system or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0048] For ease of description, some terms used in the text will be explained below.

[0049] Intelligent devices refer to any device, instrument, or machine that has computing power.

[0050] refer to Figure 1 , Figure 1 This is a flowchart of the steps involved in a voice-guided interaction method.

[0051] This invention primarily utilizes intelligent human-computer interaction to improve the efficiency of guidance and communication between midwives (obstetricians) and mothers during childbirth, making it easier for mothers to understand the guidance provided by midwives (obstetricians).

[0052] This voice-guided interaction method primarily operates through smart devices. When the smart device is running, the steps it can execute include:

[0053] Step 1: Acquire the voice signal and determine the role of the speaker based on the timbre of the voice signal.

[0054] The smart device communicates with a microphone to acquire voice signals. It then processes and recognizes the acquired voice signals to determine the speaker's role based on the timbre. The speaker's role is pre-determined and stored locally on the device.

[0055] In some further specific embodiments, determining the role of the speaker based on the timbre of the speech signal specifically includes: passing the speech signal and the pre-stored speech of the role to a timbre matching model, performing timbre matching on the speech signal and the pre-stored speech of the role through the timbre matching model to obtain a matching result, and determining the role of the speaker based on the matching result.

[0056] When a smart device needs to perform voice matching on a voice signal, it can pass the acquired voice signal and the pre-stored voice of the character to the voice matching model respectively.

[0057] The timbre matching model is a pre-trained artificial intelligence model whose function is to match the timbre similarity of two speech signals. When the timbre similarity of the two speech signals exceeds 90%, the timbre matching model will output a "similar" matching result. When the timbre similarity of the two speech signals does not exceed 90%, the timbre matching model will output a "dissimilar" matching result.

[0058] Smart devices can determine which role the current voice signal belongs to by obtaining the matching results output by the timbre matching model.

[0059] In specific application scenarios, such as childbirth, two roles are typically defined: the midwife and the mother. Because the midwife and mother have different roles during childbirth, the information conveyed by their voice signals differs. To achieve targeted recognition and enhance the efficiency of guidance and communication, in this specific embodiment, the midwife's role is defined as the guidance role, and the mother's role as the control role.

[0060] Both the mother and the midwife had their voice signals pre-recorded and stored as audio files on a local device. For ease of description, the two audio signal files are referred to as the first audio signal and the second audio signal, respectively. The first audio signal is stored in the first audio file, and the second audio signal is stored in the second audio file. Therefore, in practical applications, the smart device matches the acquired voice signals with the first audio signal recorded in the first audio file and the second audio signal recorded in the second audio file, respectively, using a timbre matching model. When the timbre matching model determines that the voice signal matches the first audio signal in the first audio file, the speaker can be considered to be the mother.

[0061] When the timbre matching model determines that the speech signal matches the second speech signal in the second sound file, the speaker can be considered to be a midwife.

[0062] Step 2: When the role of the speaker is determined to be the first preset role, the speech signal is recorded as guidance speech.

[0063] Here, the first preset role is the instructor role.

[0064] When a smart device determines, through a voice matching model, that the speaker is a guide, it can be considered that a midwife is speaking. Since the midwife's voice signal provides guidance, such as instructing the mother on how to prepare for labor, the current voice signal is considered guidance speech.

[0065] Step 21: Control the playback device according to the instruction voice, so that the playback device plays the video image corresponding to the instruction voice.

[0066] Once the smart device identifies the current voice signal as guidance, it can control the playback device based on that guidance. Specific video images are then played through the playback device's display interface to visualize the guidance information contained in the voice message, making it easier for the mother to understand and follow.

[0067] Of course, the video images played by the playback device are pre-recorded, and these video images are formed into a set of video images according to certain rules. To facilitate the device in better identifying video images from the pre-recorded video image set that match the guidance information contained in the guidance voice, in some further specific embodiments, controlling the playback device according to the guidance voice so that the playback device plays the video image corresponding to the guidance voice specifically includes:

[0068] Step 211: Convert the guidance voice into text information, wherein the text information records application scenario information and instruction information.

[0069] Smart devices convert voice instructions into text information. Because the voice instructions follow certain rules, the resulting text information contains application scenario information and instructions.

[0070] The application scenario information is pre-set and expressed through keywords. For example, "Stage 1 of labor preparation," "Stage 2 of labor preparation," etc. The instruction information is also pre-set and expressed through key phrases. For example, "Adjust breathing rhythm to XX times," etc. Of course, these application scenario and instruction information, when combined, will be accompanied by corresponding video footage.

[0071] Step 212: Perform semantic recognition on the text information using a scene recognition model to obtain target application scene information; perform semantic recognition on the text information using an indication recognition model to obtain target indication information.

[0072] The smart device invokes a scene recognition model to perform semantic recognition on the text information. This scene recognition model is a pre-trained artificial intelligence model whose function is to perform semantic recognition on the text information, thereby extracting the application scene information contained within it. For ease of description, this application scene information is referred to as the target application scene information. In addition to extracting the target application scene information from the text information, the smart device also invokes an instruction recognition model to perform semantic recognition on the text information. This instruction recognition model is a pre-trained artificial intelligence model whose function is to perform semantic recognition on the text information, thereby extracting the instruction information contained within it. For ease of description, this instruction information is referred to as the target instruction information.

[0073] Step 213: Find the video image file containing the target application scenario information in the type tag information based on the target application scenario information, and find the corresponding first target video summary file based on the video image file.

[0074] Since the video footage is pre-recorded and stored locally as files, the type tag information of the video footage files contains pre-defined application scenario information for easy retrieval. After obtaining the target application scenario information, the smart device can locate the desired video footage file through the type tag information. To determine the content of the desired video footage file, it will then locate the corresponding first target video summary file based on the found video footage file.

[0075] In this system, a video summary text file is bound to a video image file in a one-to-one correspondence. The text content of the video summary text file is obtained by performing content recognition on the bound video image file using a video content recognition model. The type label information of the video image file is pre-set. The video content recognition model is a pre-trained artificial intelligence model whose function is to perform content recognition on pre-recorded video images, thereby recording the recognized content as text information in the video summary text file.

[0076] Step 214: Locate the second target video summary file containing the target indication information from the first target video summary file according to the target indication information.

[0077] The video images corresponding to the first target video summary file all belong to the same type tag. To find visually representative instruction information within these video images, the smart device also needs to locate the corresponding first target video summary file from the first target video summary file using the target instruction information. For ease of description, the found first video summary file will be referred to as the second target video summary file.

[0078] Step 215: Determine the corresponding target video image file based on the second target video summary file.

[0079] Once a smart device identifies the second target video digest file, it can locate its corresponding video file using that digest. For ease of description, the corresponding video file will be referred to as the target video file.

[0080] Step 216: Transfer the target video file to the playback device so that the playback device can play the target video file.

[0081] Once the smart device has identified the target video file, it can transmit it to the playback device. The smart device can do this via a network connection to the playback device. After receiving the target video file, the playback device can play it, thus visualizing the instruction information.

[0082] Step 3: When the role of the speaker is determined to be the second preset role, the voice signal is recorded as control voice.

[0083] Here, the second preset role is the control role.

[0084] When a smart device determines that the speaker's role is a controller using a voice matching model, it can be assumed that the mother is speaking. Since the mother's voice signal conveys control-related information, such as controlling the current video playback status (e.g., pausing playback, increasing volume), the current voice signal is considered control speech.

[0085] Step 31: Control the playback device according to the control voice, so that the playback device controls the currently playing video image.

[0086] Once a smart device identifies the current voice signal as a control voice, it can control the playback device based on the control voice, thereby enabling the playback device to control the currently playing video image.

[0087] In some further specific embodiments, controlling the playback device according to the control voice so that the playback device controls the currently playing video image specifically includes:

[0088] Step 311: Convert the control voice into text information, wherein the text information records control parameter information.

[0089] Smart devices convert voice commands into text information. Because the issuance of control voice commands follows certain rules, the text information obtained by converting control voice commands contains control parameter information.

[0090] The control parameters are preset and expressed through key phrases. For example, pause playback, increase the volume, etc.

[0091] Step 312: Recognize the text information using the control parameter recognition model to obtain control parameters, and record the control parameters as target control parameters.

[0092] The smart device invokes a control parameter recognition model to identify text information. This model is a pre-trained artificial intelligence model that performs semantic recognition on the text information, thereby extracting the control parameters contained within it. For ease of description, these control parameters are referred to as target control parameters. These target control parameters include: increasing volume, decreasing volume, pausing playback, or resuming playback.

[0093] Step 313: The target control parameters are transmitted to the playback device so that the playback device controls the video image currently playing on the display interface according to the target control parameters.

[0094] Once the smart device has determined the target control parameters, it can transmit these parameters to the playback device. The playback device then controls the currently playing video image based on these target control parameters.

[0095] This invention sets up two different preset roles and determines whether the language spoken by the current speaker has a guiding or controlling meaning by judging the timbre of the speech signal. The playback device is then controlled according to the different meanings. Furthermore, by visually displaying the meaning of the guiding speech, the communication efficiency between the different roles is improved. This allows for effective guidance and communication between midwives and mothers in certain scenarios, such as childbirth, thereby enhancing communication efficiency.

[0096] The playback device can take many forms. In some further specific embodiments, the playback device includes: a smart TV, a smart projector, a smart playback device, or VR glasses.

[0097] On the other hand, reference Figure 2 , Figure 2 This is a schematic diagram of the device structure of a voice-guided interactive device.

[0098] A voice-guided interactive device is provided, comprising: a processor and a memory; wherein the memory is used to store a computer-readable program. When the computer-readable program is executed by the processor, the processor causes the processor to implement the voice-guided interactive method as described in any of the above technical solutions.

[0099] It will be understood by those skilled in the art that all or some of the steps and systems in the methods disclosed above can be implemented as software, firmware, hardware, and suitable combinations thereof. Some or all of the physical components can be implemented as software executed by a processor, such as a central processing unit, digital signal processor, or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit. Such software can be distributed on a computer-readable medium, which can include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and is accessible to a computer. As is known to those skilled in the art, communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.

[0100] On the other hand, reference Figure 3 , Figure 3 This is a schematic diagram of the system connection structure of a voice-guided interactive system.

[0101] A voice-guided interactive system is provided, comprising: an acquisition module, a first determination module, and a second determination module.

[0102] The acquisition module is used to: acquire a voice signal and determine the role of the voice speaker based on the timbre of the voice signal.

[0103] The acquisition module communicates with a microphone to acquire speech signals. It then processes and recognizes the acquired speech signals to determine the speaker's role based on the timbre. The speaker's role is pre-determined and stored locally.

[0104] In some further specific embodiments, determining the role of the speaker based on the timbre of the speech signal specifically includes: passing the speech signal and the pre-stored speech of the role to a timbre matching model, performing timbre matching on the speech signal and the pre-stored speech of the role through the timbre matching model to obtain a matching result, and determining the role of the speaker based on the matching result.

[0105] When the acquisition module needs to perform timbre matching on the speech signal, it can pass the acquired speech signal and the pre-stored speech of the character to the timbre matching model respectively.

[0106] The timbre matching model is a pre-trained artificial intelligence model whose function is to match the timbre similarity of two speech signals. When the timbre similarity of the two speech signals exceeds 90%, the timbre matching model will output a "similar" matching result. When the timbre similarity of the two speech signals does not exceed 90%, the timbre matching model will output a "dissimilar" matching result.

[0107] The acquisition module can determine which role the current voice signal belongs to by obtaining the matching results output by the timbre matching model.

[0108] In specific application scenarios, such as childbirth, two roles are typically defined: the midwife and the mother. Because the midwife and mother have different roles during childbirth, the information conveyed by their voice signals differs. To achieve targeted recognition and enhance the efficiency of guidance and communication, in this specific embodiment, the midwife's role is defined as the guidance role, and the mother's role as the control role.

[0109] Both the mother and the midwife had their voice signals pre-recorded and stored as audio files on a local device. For ease of description, the two audio signal files are referred to as the first audio signal and the second audio signal, respectively. The first audio signal is stored in the first audio file, and the second audio signal is stored in the second audio file. Therefore, in practical applications, the smart device matches the acquired voice signals with the first audio signal recorded in the first audio file and the second audio signal recorded in the second audio file, respectively, using a timbre matching model. When the timbre matching model determines that the voice signal matches the first audio signal in the first audio file, the speaker can be considered to be the mother.

[0110] When the timbre matching model determines that the speech signal matches the second speech signal in the second sound file, the speaker can be considered to be a midwife.

[0111] The first determining module is used to: when the role of the speaker of the voice is determined to be a first preset role, record the voice signal as guidance voice; and control the playback device according to the guidance voice so that the playback device plays the video image corresponding to the guidance voice.

[0112] Here, the first preset role is the instructor role.

[0113] Once the first determining module identifies the speaker's role as a guide through the timbre matching model, it can be considered that the midwife is speaking. Since the midwife's voice signal conveys guiding information, such as instructing the mother on how to prepare for labor, the current voice signal is considered guiding speech.

[0114] Once the first determining module identifies the current voice signal as guidance voice, it can control the playback device based on the guidance voice. The playback device then plays specific video images to visualize the guidance information contained in the voice, making it easier for the mother to understand and follow instructions.

[0115] Of course, the video images played by the playback device are pre-recorded, and these video images are formed into a set of video images according to certain rules. To facilitate the device in better identifying video images from the pre-recorded video image set that match the guidance information contained in the guidance voice, in some further specific embodiments, controlling the playback device according to the guidance voice so that the playback device plays the video image corresponding to the guidance voice specifically includes:

[0116] Step 211: Convert the guidance voice into text information, wherein the text information records application scenario information and instruction information.

[0117] The first determination module converts the guidance speech into text information. Since the guidance speech is issued according to certain rules, the text information obtained by converting the guidance speech contains application scenario information and instruction information.

[0118] The application scenario information is pre-set and expressed through keywords. For example, "Stage 1 of labor preparation," "Stage 2 of labor preparation," etc. The instruction information is also pre-set and expressed through key phrases. For example, "Adjust breathing rhythm to XX times," etc. Of course, these application scenario and instruction information, when combined, will be accompanied by corresponding video footage.

[0119] Step 212: Perform semantic recognition on the text information using a scene recognition model to obtain target application scene information; perform semantic recognition on the text information using an indication recognition model to obtain target indication information.

[0120] The first determining module invokes a scene recognition model to perform semantic recognition on the text information. This scene recognition model is a pre-trained artificial intelligence model whose function is to perform semantic recognition on the text information, thereby extracting the application scene information contained within it. For ease of description, this application scene information is referred to as the target application scene information. In addition to extracting the target application scene information from the text information, the first determining module also invokes an instruction recognition model to perform semantic recognition on the text information. This instruction recognition model is a pre-trained artificial intelligence model whose function is to perform semantic recognition on the text information, thereby extracting the instruction information contained within it. For ease of description, this instruction information is referred to as the target instruction information.

[0121] Step 213: Find the video image file containing the target application scenario information in the type tag information based on the target application scenario information, and find the corresponding first target video summary file based on the video image file.

[0122] Since the video footage is pre-recorded and stored locally as files, the type tag information of the video footage files contains pre-defined application scenario information for easy retrieval. After obtaining the target application scenario information, the first determination module can locate the expected video footage file using the type tag information. To determine the content of the expected video footage file, it will then locate the corresponding first target video summary file using the found video footage file.

[0123] In this system, a video summary text file is bound to a video image file in a one-to-one correspondence. The text content of the video summary text file is obtained by performing content recognition on the bound video image file using a video content recognition model. The type label information of the video image file is pre-set. The video content recognition model is a pre-trained artificial intelligence model whose function is to perform content recognition on pre-recorded video images, thereby recording the recognized content as text information in the video summary text file.

[0124] Step 214: Locate the second target video summary file containing the target indication information from the first target video summary file according to the target indication information.

[0125] The video images corresponding to the first target video summary file all belong to the same type tag. To find visually representative indication information within these video images, the first determining module also needs to locate the corresponding first target video summary file from the first target video summary file using the target indication information. For ease of description, the found first video summary file will be referred to as the second target video summary file.

[0126] Step 215: Determine the corresponding target video image file based on the second target video summary file.

[0127] After determining the second target video summary file, the first determining module can find its corresponding video image file through the second target video summary file. For ease of description, the corresponding video image file is referred to as the target video image file.

[0128] Step 216: Transfer the target video file to the playback device so that the playback device can play the target video file.

[0129] After identifying the target video file, the first determining module can transmit it to the playback device. This first determining module can connect to the playback device via a network. The playback device then transmits the target video file to the playback device through the communication network. Upon receiving the target video file, the playback device can play it, thus visualizing the instruction information.

[0130] The second determining module is used to: when the role of the voice sender is determined to be a second preset role, record the voice signal as control voice; control the playback device according to the control voice, so that the playback device controls the currently playing video image.

[0131] When the second determining module identifies the speaker's role as a control role through the timbre matching model, it can be considered that the current mother is speaking. Since the mother's voice signal conveys control-related information, such as controlling the current video image's state (e.g., pausing playback, increasing volume), the current voice signal is considered control speech.

[0132] Once the second determining module identifies the current voice signal as control voice, it can control the playback device based on the control voice, thereby enabling the playback device to control the currently playing video image.

[0133] In some further specific embodiments, controlling the playback device according to the control voice so that the playback device controls the currently playing video image specifically includes:

[0134] Step 311: Convert the control voice into text information, wherein the text information records control parameter information.

[0135] The second determining module converts the guidance speech into text information. Since the issuance of control speech follows certain rules, the text information obtained by converting the control speech contains control parameter information.

[0136] The control parameters are preset and expressed through key phrases. For example, pause playback, increase the volume, etc.

[0137] Step 312: Recognize the text information using the control parameter recognition model to obtain control parameters, and record the control parameters as target control parameters.

[0138] The second determining module invokes a control parameter recognition model to identify the text information. This control parameter recognition model is a pre-trained artificial intelligence model whose function is to perform semantic recognition on the text information, thereby extracting the control parameters contained within it. For ease of description, these control parameters are referred to as target control parameters. The target control parameters include: increasing volume, decreasing volume, pausing playback, or resuming playback.

[0139] Step 313: The target control parameters are transmitted to the playback device so that the playback device controls the video image currently playing on the display interface according to the target control parameters.

[0140] After determining the target control parameters, the second determining module can transmit the target control parameters to the playback device. The playback device will then control the currently playing video image according to the target control parameters.

[0141] On the other hand, a computer-readable storage medium is provided, wherein a processor-executable program is stored, which, when executed by a processor, is used to implement the voice-guided interaction method as described in any of the above specific embodiments.

[0142] This application also discloses a computer program product, including a computer program or computer instructions, which are stored in a computer-readable storage medium. A processor of a computer device reads the computer program or computer instructions from the computer-readable storage medium and executes the computer program or computer instructions, causing the computer device to perform the voice-guided interaction method as described in any of the preceding embodiments.

[0143] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatuses.

[0144] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0145] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, apparatuses, or units, and may be electrical, mechanical, or other forms.

[0146] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0147] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0148] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0149] Although the description of this application has been quite detailed and particularly focused on several of the described embodiments, it is not intended to limit itself to any of these details or embodiments or any particular embodiment. Rather, it should be considered as effectively covering the intended scope of this application by referring to the appended claims and taking into account the prior art, which provides for a broad possible interpretation of these claims. Furthermore, the foregoing description of this application with respect to embodiments foreseeable by the inventors is intended to provide a useful description, and non-substantial modifications to this application that have not yet been foreseen may still represent equivalent modifications.

Claims

1. A voice-guided based interaction method, characterized in that, The method comprises the following steps: acquiring a voice signal, determining a role of a voice emitter according to a timbre of the voice signal; when it is determined that the role of the voice emitter is a first preset role, recording the voice signal as a guide voice, and controlling a playing device according to the guide voice, so that the playing device plays a video image corresponding to the guide voice; when it is determined that the role of the voice emitter is a second preset role, recording the voice signal as a control voice, and controlling the playing device according to the control voice, so that the playing device controls a currently played video image; controlling the playing device according to the guide voice, so that the playing device plays a video image corresponding to the guide voice specifically comprises the following steps: converting the guide voice into text information, wherein the text information records application scenario information and instruction information; performing semantic recognition on the text information by using a scene recognition model, so as to obtain target application scenario information; performing semantic recognition on the text information by using an instruction recognition model, so as to obtain target instruction information; finding a video image file containing the target application scenario information from type tag information according to the target application scenario information, and finding a corresponding first target video abstract file according to the video image file; wherein a video abstract text file is bound to a video image file one by one, the text content of the video abstract text file is obtained by performing content recognition on the video image file bound thereto by using a video content recognition model, and the type tag information of the video image file is preset; finding a second target video abstract file containing the target instruction information from the first target video abstract file according to the target instruction information; determining a target video image file corresponding to the second target video abstract file; delivering the target video image file to the playing device, so that the playing device plays the target video image file.

2. The voice-guided interaction method of claim 1, wherein, controlling the playing device according to the control voice, so that the playing device controls a currently played video image specifically comprises the following steps: converting the control voice into text information, wherein the text information records control parameter information; performing recognition on the text information by using a control parameter recognition model, so as to obtain a control parameter, and recording the control parameter as a target control parameter; delivering the target control parameter to the playing device, so that the playing device controls a video image currently played in a display interface according to the target control parameter; wherein the target control parameter comprises increasing volume, decreasing volume, pausing playing, or continuing playing.

3. The voice-guided interaction method of claim 1, wherein, determining the role of the voice emitter according to the timbre of the voice signal specifically comprises the following steps: delivering the voice signal and pre-stored voices of roles to a timbre matching model respectively, performing timbre matching on the voice signal and the pre-stored voices of roles by using the timbre matching model, obtaining a matching result, and determining the role of the voice emitter according to the matching result.

4. The voice-guided interaction method of claim 1, wherein, The playing device comprises a smart television, a smart projector, a smart playing device, or a VR glasses.

5. A voice-guided interactive device, characterized by The method comprises the following steps: a processor; a memory for storing a computer readable program; When the computer readable program is executed by the processor, the processor implements the voice guidance based interaction method according to any one of claims 1-4.

6. A voice-guided based interactive system, characterized by, Comprise: The acquisition module, the first determination module and the second determination module; The acquisition module is used for: acquiring a voice signal, and determining a role of a voice issuer according to a timbre of the voice signal; The first determination module is used for: when it is determined that the role of the voice issuer is a first preset role, then the voice signal is recorded as a guidance voice; and controlling a playing device according to the guidance voice, so that the playing device plays a video image corresponding to the guidance voice; The second determination module is used for: when it is determined that the role of the voice issuer is a second preset role, then the voice signal is recorded as a control voice; and controlling a playing device according to the control voice, so that the playing device controls a currently played video image; Controlling a playing device according to the guidance voice, so that the playing device plays a video image corresponding to the guidance voice specifically includes: Converting the guidance voice into text information, wherein the text information records application scenario information and instruction information; Performing semantic recognition on the text information through a scene recognition model, so as to obtain target application scenario information; Performing semantic recognition on the text information through an instruction recognition model, so as to obtain target instruction information; According to the target application scenario information, finding a video image file containing the target application scenario information in type label information, and according to the video image file, finding a corresponding first target video summary file; Wherein, a video summary text file and a video image file are one-to-one bound, the text content of the video summary text file is obtained by performing content recognition on the video image file bound thereto through a video content recognition model, and the type label information of the video image file is pre-set; According to the target instruction information, finding a second target video summary file containing the target instruction information from the first target video summary file; According to the second target video summary file, determining a target video image file corresponding thereto; Transferring the target video image file to the playing device, so that the playing device plays the target video image file.

7. A voice-guided based interactive system as claimed in claim 6, wherein, Controlling a playing device according to the control voice, so that the playing device controls a currently played video image specifically includes: converting the control voice into text information, wherein the text information records control parameter information; Performing recognition on the text information through a control parameter recognition model, obtaining a control parameter, and recording the control parameter as a target control parameter; Transferring the target control parameter to the playing device, so that the playing device controls a video image currently played in a display interface according to the target control parameter; Wherein, the target control parameter includes: increasing volume, decreasing volume, pausing playing or continuing playing.

8. A computer-readable storage medium, characterized in that, Therein, a processor executable program is stored, the processor executable program is executed by a processor to implement the voice guidance based interaction method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Medical terminal response method and device based on voice recognition, equipment and medium

    CN116013288A

  • Intelligent control method and system

    WO2025035715A1