A Human-Computer Voice Interaction Control Method and System Based on Smart TV
By using a voice analysis model and a dual confirmation mechanism on smart TVs, the accuracy problem of TV voice control has been solved, enabling smarter and more accurate voice interaction control and optimizing the user's viewing experience.
Patent Information
- Application Number
- CN202410756474.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-12
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2044-06-12
AI Technical Summary
Existing TV voice control methods require users to get close to the TV or remote control, which cannot accurately recognize regional accent differences, resulting in inaccurate voice control command recognition.
The system analyzes human voice features using a pre-set voice analysis model, extracts wake-up keywords, compares them with a pre-set command library, and combines user confirmation information to evaluate and correct commands. A dual confirmation mechanism is used to improve recognition accuracy, and the system controls TV function switching and adjusts associated devices based on the recognition results.
It improves the intelligence and recognition accuracy of TV voice interaction control, optimizes the user's viewing experience, reduces the interference of program sound on voice command recognition, and enhances the intelligent correction capability of voice interaction.
Smart Images

Figure CN118737141B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of voice interaction, and in particular to a human-computer voice interaction control method and system based on a smart TV. Background Technology
[0002] Currently, with the rapid development of radio and television, and the increasing number of television channels and program sources, the drawbacks of traditional television remote control methods are becoming more and more obvious. Users need to memorize a large number of television channels to make quick selections, or they need to check each television channel one by one to select the appropriate program. Therefore, the demand for voice intelligence technology to be applied to traditional television remote control is becoming increasingly strong.
[0003] Existing TV voice control methods typically involve setting a specific wake-up word or a wake-up button on the remote control. Users input the wake-up word by voice or manually press the wake-up button to input voice commands. A preset program extracts channel switching information from the voice commands and controls the TV to switch channels accordingly, thus achieving the effect of TV voice remote control and reducing the cumbersome process of manually switching multiple channels by pressing buttons. However, existing voice control methods require users to be close to the TV or remote control system to receive voice commands, and differences in accents from different regions make it difficult to accurately recognize voice control commands. Further optimization of the TV's voice interaction method is needed. Summary of the Invention
[0004] To improve the intelligence of voice interaction control in smart TVs, this application provides a human-computer voice interaction control method and system based on a smart TV.
[0005] Firstly, the above-mentioned inventive objective of this application is achieved through the following technical solution:
[0006] A human-computer voice interaction control method based on a smart TV includes:
[0007] Acquire voice data within a preset range, process the voice data to obtain human voice feature data carrying control commands;
[0008] The voice feature data is processed by a preset voice analysis model to extract wake-up keywords and compare them with a preset command library to obtain control command comparison results.
[0009] Based on the control command comparison results, a secondary confirmation request is sent to the user. The user's confirmation voice information is then combined with the command recognition and evaluation, and the command deviation is corrected to obtain the correct control command after correction.
[0010] According to the correct control instructions, the TV is switched to a different function. Based on the program display requirements after the switch, the related devices are adjusted to optimize the display effect, and human-computer voice interaction control data is obtained.
[0011] By adopting the above technical solution and combining voice data within the preset reception range of the television, the voice data is processed to separate the television program sound from the user's voice, obtaining voice feature data carrying control commands. This helps reduce the interference of program sound on the voice command recognition results. A preset voice analysis model is used to analyze the voice feature data, extracting the television's wake-up keywords and comparing them with commands in a preset command library. This provides initial confirmation of the user's voice control of the television. Combined with the user's secondary confirmation, the accuracy of the command recognition is evaluated, and any recognition deviations are corrected. This dual confirmation improves the accuracy of television control command recognition, intelligently correcting the command recognition results. Following the correct control commands, the television performs function switching. Based on the program display requirements after the switch, related devices such as audio and lighting are adjusted, optimizing the program display effect and further enhancing the intelligence of the smart television's voice interaction control, thus improving the user's viewing experience.
[0012] In a preferred embodiment, this application can be further configured as follows: Based on the control command comparison result, a secondary confirmation request is sent to the user, and the user's confirmation voice information is used to perform command recognition evaluation and command deviation correction processing to obtain the corrected control command. Specifically, this includes:
[0013] Based on the control command comparison results, the sound source of the voice control command is located, a preliminary sound source location result is obtained, and a secondary confirmation request is sent to the user.
[0014] Obtain the user's secondary confirmation information, and combine it with the confirmation voice information provided by the user to perform secondary sound source localization processing on the confirmation voice information;
[0015] Based on the overlap between the preliminary and secondary sound source localization results, verify whether the sound source of the confirmed voice information is compliant.
[0016] By combining the sound source judgment result with the compliant confirmation voice information, the accuracy of the model's command recognition result is evaluated, and the deviation command is corrected according to the evaluation result to obtain the correct control command after correction.
[0017] In a preferred embodiment, this application can be further configured such that: the overlap of sound source localization results between the preliminary sound source localization result and the secondary sound source localization result specifically includes:
[0018] The location overlap coefficient between the preliminary sound source localization and the secondary sound source localization is calculated using formula (1). Based on the location overlap coefficient, the sound source localization overlap degree between the preliminary sound source localization result and the secondary sound source localization result is analyzed. Formula (1) is shown below:
[0019] A = cos -1 [tau2*c / (mic_d*2)]*180 / pi2-cos -1 [tau1*c / (mic_d*2)]*180 / pi1 (1)
[0020] Where A represents the location overlap coefficient, tau2 and tau1 represent the sound pickup delay between each microphone array of the TV and the sound source, c represents the sound propagation speed, mic_d represents the distance between the microphone arrays of the TV, and pi2 and pi1 represent the sound pickup coefficient of each microphone array.
[0021] By adopting the above technical solution, the sound source of the voice control command is initially located. Combined with the user's secondary request confirmation information, the confirmation voice information is further processed for secondary sound source localization. The overlap between the two localization results is used to verify whether the sound source of the confirmation voice information is compliant, which helps to improve the accuracy of sound source localization. Furthermore, the accuracy of the command recognition result of the model is evaluated through compliant confirmation voice information, and deviation commands in the initial recognition result are corrected to obtain the correct control command. This helps to improve the accuracy of voice recognition of the control commands of the television.
[0022] In a preferred embodiment, this application can be further configured as follows: The step of performing feature analysis processing on the human voice feature data using a preset voice analysis model, extracting wake-up keywords, and comparing them with a preset instruction library to obtain control instruction comparison results specifically includes:
[0023] The human voice feature data is processed by extracting TV keywords through a preset voice analysis model to obtain the TV wake-up keyword matching result.
[0024] When the wake-up keyword matches, the voice feature data following the wake-up keyword is processed to extract control commands, and the control command data carried in the voice data is obtained.
[0025] The control command data is compared with the preset command library of the television to obtain the control command comparison result.
[0026] By adopting the above technical solution, a preset speech analysis model is used to extract keywords from human voice feature data for television. The captured speech is converted into a binary format that can be recognized by the television and matched with the preset wake-up keywords of the television to obtain the wake-up keyword matching result. When the wake-up keyword matching result is consistent, the voice feature data after the wake-up keyword is processed to extract control commands to obtain the control command data carried in the speech data. The effective command data is extracted by using a dual verification method of wake-up word and command, which reduces the complexity of speech information processing. Combined with the comparison between the preset command library and the control command data, the control command comparison result is obtained, which improves the accuracy of control command comparison.
[0027] In a preferred embodiment, this application can be further configured as follows: Following the correct control command, the television set undergoes a function switching process; based on the displayed program requirements after the switching, related devices are coordinated and adjusted to optimize the display effect, resulting in human-computer voice interaction control data. Specifically, this includes:
[0028] According to the correct control command, the display function of the TV is switched, and the display program information after the switch is obtained;
[0029] Obtain the current display atmosphere data, combine it with historical viewing preference data and the displayed program information, and generate atmosphere adjustment data that matches the user's viewing preferences;
[0030] Based on the ambient adjustment data, the ambient-related devices are activated to adjust parameters and optimize the current display effect, thereby obtaining the human-computer voice interaction control data for the television.
[0031] By adopting the above technical solution, the display function of the TV is switched according to the correct control command. Combined with the information of the program to be displayed after the switch, the atmosphere adjustment parameters are analyzed based on historical viewing preference data and current display atmosphere data. The atmosphere-related devices are then activated accordingly to adjust the parameters, thereby optimizing the current display effect and obtaining the human-computer voice interaction control data of the TV, which helps to improve the viewing experience of the TV.
[0032] Secondly, the above-mentioned inventive objective of this application is achieved through the following technical solutions:
[0033] A human-computer voice interaction control system based on a smart TV, comprising:
[0034] The data acquisition module is used to acquire voice data within a preset range, process the voice data, and obtain human voice feature data carrying control commands.
[0035] The data analysis module is used to perform feature analysis processing on the human voice feature data through a preset voice analysis model, extract wake-up keywords and compare them with a preset instruction library to obtain control instruction comparison results;
[0036] The data correction module is used to send a secondary confirmation request to the user based on the control command comparison results, and to perform command recognition and evaluation and correct command deviation processing in combination with the confirmation voice information fed back by the user, so as to obtain the correct control command after correction.
[0037] The data control module is used to perform function switching on the television set according to the correct control instructions, and, in conjunction with the display requirements of the switched programs, to coordinate and adjust related devices to optimize the display effect, thereby obtaining human-computer voice interaction control data.
[0038] By adopting the above technical solution and combining voice data within the preset reception range of the television, the voice data is processed to separate the television program sound from the user's voice, obtaining voice feature data carrying control commands. This helps reduce the interference of program sound on the voice command recognition results. A preset voice analysis model is used to analyze the voice feature data, extracting the television's wake-up keywords and comparing them with commands in a preset command library. This provides initial confirmation of the user's voice control of the television. Combined with the user's secondary confirmation, the accuracy of the command recognition is evaluated, and any recognition deviations are corrected. This dual confirmation improves the accuracy of television control command recognition, intelligently correcting the command recognition results. Following the correct control commands, the television performs function switching. Based on the program display requirements after the switch, related devices such as audio and lighting are adjusted, optimizing the program display effect and further enhancing the intelligence of the smart television's voice interaction control, thus improving the user's viewing experience.
[0039] Thirdly, the above-mentioned objectives of this application are achieved through the following technical solutions:
[0040] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described human-computer voice interaction control method based on a smart TV.
[0041] Fourthly, the above-mentioned objectives of this application are achieved through the following technical solutions:
[0042] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described human-computer voice interaction control method based on a smart TV.
[0043] In summary, this application includes at least one of the following beneficial technical effects:
[0044] 1. By combining voice data from the TV's preset reception range, the system processes the voice data, separating the TV program audio from the user's voice to obtain voice feature data carrying control commands. This helps reduce interference from program audio on voice command recognition results. A preset voice analysis model is used to analyze the voice feature data, extracting the TV's wake-up keywords and comparing them with commands in a preset command library. This provides initial confirmation of the user's voice control of the TV. A second confirmation from the user further evaluates the accuracy of the command recognition and corrects any discrepancies. This dual confirmation improves the accuracy of TV control command recognition, intelligently correcting the command recognition results. Following the correct control commands, the system controls the TV to switch functions. Based on the program display requirements after the switch, it adjusts related devices such as audio and lighting to optimize the program display effect, further enhancing the intelligence of the smart TV's voice interaction control and improving the user's viewing experience.
[0045] 2. Initially locate the sound source of the voice control command, and combine it with the user's secondary confirmation request information to perform secondary sound source localization processing on the confirmation voice information. Combine the localization overlap between the two sound source localization results to verify whether the sound source of the confirmation voice information is compliant, which helps to improve the accuracy of sound source localization. Through compliant confirmation voice information, the accuracy of the model's command recognition results is evaluated, and the deviation commands in the initial recognition results are corrected to obtain the correct control commands, which helps to improve the voice recognition accuracy of the TV's control commands.
[0046] 3. Using a preset speech analysis model, the voice feature data is processed to extract TV keywords. The captured speech is converted into a binary format that can be recognized by the TV and matched with the TV's preset wake-up keywords to obtain the wake-up keyword matching result. When the wake-up keyword matching result is consistent, the voice feature data after the wake-up keyword is processed to extract control commands to obtain the control command data carried in the speech data. Valid command data is extracted through a dual verification method of wake-up word and command, reducing the complexity of speech information processing. Combined with the comparison between the preset command library and the control command data, the control command comparison result is obtained, improving the accuracy of control command comparison. Attached Figure Description
[0047] Figure 1 This is a flowchart illustrating the implementation of a human-computer voice interaction control method based on a smart TV in this embodiment.
[0048] Figure 2 This is a flowchart illustrating the implementation of step S20 of the human-computer voice interaction control method based on a smart TV in this embodiment.
[0049] Figure 3 This is a flowchart illustrating the implementation of step S30 of the human-computer voice interaction control method based on a smart TV in this embodiment.
[0050] Figure 4 This is a flowchart illustrating the implementation of step S40 of the human-computer voice interaction control method based on a smart TV in this embodiment.
[0051] Figure 5 This is a structural block diagram of a human-computer voice interaction control system based on a smart TV in this embodiment.
[0052] Figure 6 This is a schematic diagram of the internal structure of a computer device used to implement a human-computer voice interaction control method based on a smart TV. Detailed Implementation
[0053] The present application will be further described in detail below with reference to the accompanying drawings.
[0054] In one embodiment, such as Figure 1 As shown, this application discloses a human-computer voice interaction control method based on a smart TV, which specifically includes the following steps:
[0055] S10: Acquire voice data within a preset range, process the voice data, and obtain human voice feature data carrying control commands.
[0056] Specifically, based on the set sound pickup range of the smart TV, voice data within the set sound pickup range is collected in real time through a preset microphone array. The voice data is then processed using voice features such as voiceprint and timbre to remove the human voice from the program being played on the TV and retain the real human voice feature data.
[0057] S20: Through a preset voice analysis model, the voice feature data is processed by feature analysis, wake-up keywords are extracted and compared with the preset command library to obtain the control command comparison result.
[0058] Specifically, such as Figure 2 As shown, step S20 includes:
[0059] S201: Extract keywords from human voice feature data using a preset voice analysis model to obtain the TV wake-up keyword matching result.
[0060] Specifically, combined with a preset speech analysis model, human voice feature data is input into the speech analysis model for sound signal extraction and processing. The human voice feature parameters are analyzed, and the wake-up keywords carried in the human voice feature data are extracted and matched with the preset wake-up keywords of the TV to obtain the wake-up keyword matching result of the TV. This includes two matching results: a successful match that wakes up the TV and a failed match that does not change the current state of the TV.
[0061] S202: When the wake-up keyword matches, the voice feature data following the wake-up keyword is processed to extract control instructions, and the control instruction data carried in the voice data is obtained.
[0062] Specifically, when the wake-up keyword matches, the voice feature data immediately following the wake-up keyword is extracted and converted into text or binary format that the TV can recognize, thus obtaining the control command data carried in the voice data.
[0063] S203: Compare the control command data with the TV's preset command library to obtain the control command comparison result.
[0064] Specifically, the control command data is compared with the TV's preset control command library, such as control commands for program switching, playing songs, and playing news. The control command comparison result includes two results: matching and mismatch. When the comparison result matches, the TV is operated according to the correct control command. When the comparison result does not match, the user is informed of the command input error via voice.
[0065] S30: Based on the control command comparison results, send a secondary confirmation request to the user, and combine the confirmation voice information provided by the user to perform command recognition evaluation and correct command deviation processing to obtain the correct control command after correction.
[0066] Specifically, such as Figure 3 As shown, step S30 specifically includes:
[0067] S301: Based on the control command comparison results, locate the sound source of the voice control command, obtain preliminary sound source location results, and send a secondary confirmation request to the user.
[0068] Specifically, based on the control command comparison results, the voice data that matches the command comparison is processed to locate the sound source of the voice control command. For example, the azimuth and pitch angles between the sound source and the TV are analyzed through a preset program, and the straight-line distance between the sound source and the TV is estimated to obtain a preliminary location result. Then, a secondary confirmation request is sent to the user through the TV's speakers, that is, a confirmation request voice is played.
[0069] S302: Obtain the user's secondary confirmation information, and combine it with the confirmation voice information provided by the user to perform secondary sound source localization processing on the confirmation voice information.
[0070] Specifically, by combining the user's secondary confirmation information, such as picking up the user's voice again through the TV microphone array to obtain the user's confirmation voice information, such as confirming to switch programs or not confirming, when the confirmation voice information carries confirmation keywords for switching programs, secondary sound source localization processing is performed on the confirmation voice information.
[0071] S303: Based on the overlap between the preliminary and secondary sound source localization results, verify whether the sound source of the voice information is compliant.
[0072] Specifically, the location overlap coefficient between the preliminary sound source localization and the secondary sound source localization is calculated using formula (1). Based on the location overlap coefficient, the sound source localization overlap degree between the preliminary sound source localization result and the secondary sound source localization result is analyzed. Formula (1) is shown below:
[0073] A = cos -1 [tau2*c / (mic_d*2)]*180 / pi2-cos -1 [tau1*c / (mic_d*2)]*180 / pi1 (1)
[0074] Where A represents the location overlap coefficient, tau2 and tau1 represent the sound pickup delay between each microphone array of the TV and the sound source, c represents the sound propagation speed, mic_d represents the distance between the microphone arrays of the TV, and pi2 and pi1 represent the sound pickup coefficient of each microphone array.
[0075] If the location overlap coefficient is less than the preset value, it means that the two sound source locations are basically the same within the error range, thus confirming that the sound source of the voice information is compliant, that is, the control commands issued by the same user from the same location. If the location overlap coefficient is greater than or equal to the preset value, it means that there is a difference between the two sound source locations within the error range, thus confirming that the sound source of the voice information is non-compliant, that is, the two voices are issued from different locations. At this time, the voiceprint, timbre and other data of the initial confirmed voice are compared to confirm whether the two voices are issued by the same user. If so, the sound source of the voice information is compliant; if not, the sound source is non-compliant.
[0076] S304: Combining the sound source judgment result with the compliant confirmation voice information, the accuracy of the model's command recognition result is evaluated, and the deviation command is corrected according to the evaluation result to obtain the correct control command after correction.
[0077] Specifically, based on the sound source judgment results, the accuracy of the initial command recognition result is evaluated for the confirmation voice information whose sound source judgment result is compliant. If the confirmation voice information is the user's confirmation of the initial recognition command, it means that the initial recognition result is accurate. If the confirmation voice information is a denial or change of the initial recognition command, it means that there is an error in the initial recognition result. The command information in the confirmation voice information is updated into the recognition result of the initial recognition command, and the deviation command is corrected. That is, the control command in the confirmation voice information is taken as the standard, so as to obtain the correct control command after correction.
[0078] S40: Following the correct control instructions, the TV performs function switching, and in conjunction with the program display requirements after the switch, it coordinates and adjusts related devices to optimize the display effect, thereby obtaining human-computer voice interaction control data.
[0079] Specifically, such as Figure 4 As shown, step S40 includes:
[0080] S401: Following the correct control instructions, switch the display function of the television and obtain the displayed program information after the switch.
[0081] Specifically, following the correct control instructions, the display functions of the television are switched, such as switching the current program from song playback to news playback, and obtaining the displayed program information after the switch.
[0082] S402: Obtain the current display atmosphere data, combine it with historical viewing preference data and displayed program information, and generate atmosphere adjustment data that matches the user's viewing preferences.
[0083] Specifically, the system obtains environmental parameters for the current playback, including lighting, sound, and TV display brightness, to obtain current display atmosphere data. Combined with the user's geographical viewing preferences and the currently displayed program, it performs viewing preference atmosphere analysis to obtain atmosphere adjustment data that matches the user's viewing preferences.
[0084] S403: Based on the atmosphere adjustment data, the system mobilizes the atmosphere-related devices to adjust parameters, optimizes the current display effect, and obtains the human-computer voice interaction control data for the TV.
[0085] Specifically, based on the atmosphere adjustment data, related devices such as room lighting and audio equipment are activated, and the relevant devices are started and their parameters are adjusted according to the atmosphere adjustment data to optimize the current display effect and obtain the human-computer voice interaction control data of the TV.
[0086] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0087] In one embodiment, a human-computer voice interaction control system based on a smart TV is provided, which corresponds one-to-one with the human-computer voice interaction control method based on a smart TV described in the above embodiments. For example... Figure 5 As shown, this human-computer voice interaction control system based on a smart TV includes a data acquisition module, a data analysis module, a data correction module, and a data control module. Detailed descriptions of each functional module are as follows:
[0088] The data acquisition module is used to acquire voice data within a preset range, process the voice data, and obtain human voice feature data carrying control commands.
[0089] The data analysis module is used to perform feature analysis on human voice feature data through a preset voice analysis model, extract wake-up keywords and compare them with a preset command library to obtain control command comparison results.
[0090] The data correction module is used to send a secondary confirmation request to the user based on the control command comparison results, and to perform command recognition evaluation and correction of command deviations by combining the confirmation voice information provided by the user, so as to obtain the correct control command after correction.
[0091] The data control module is used to switch functions of the TV according to the correct control instructions, and, in conjunction with the display requirements of the switched programs, to adjust related devices to optimize the display effect and obtain human-computer voice interaction control data.
[0092] Preferably, the data correction module specifically includes:
[0093] The sound source localization submodule is used to locate the sound source of the voice control command based on the control command comparison results, obtain the preliminary sound source localization result, and send a secondary confirmation request to the user.
[0094] The secondary positioning submodule is used to obtain the user's secondary confirmation information and, in conjunction with the confirmation voice information provided by the user, perform secondary sound source localization processing on the confirmation voice information.
[0095] The sound source assessment submodule is used to verify and confirm whether the sound source of the speech information is compliant based on the overlap between the preliminary sound source localization results and the secondary sound source localization results.
[0096] The instruction correction submodule is used to combine the compliant confirmation voice information with the sound source judgment result to evaluate the accuracy of the instruction recognition result of the model, and to correct the deviation instruction based on the evaluation result to obtain the correct control instruction after correction.
[0097] Preferably, the overlap between the preliminary sound source localization results and the secondary sound source localization results specifically includes:
[0098] The location overlap coefficient between the preliminary sound source localization and the secondary sound source localization is calculated using formula (1). Based on the location overlap coefficient, the sound source localization overlap degree between the preliminary sound source localization result and the secondary sound source localization result is analyzed. Formula (1) is shown below:
[0099] A = cos -1 [tau2*c / (mic_d*2)]*180 / pi2-cos -1 [tau1*c / (mic_d*2)]*180 / pi1 (1)
[0100] Where A represents the location overlap coefficient, tau2 and tau1 represent the sound pickup delay between each microphone array of the TV and the sound source, c represents the sound propagation speed, mic_d represents the distance between the microphone arrays of the TV, and pi2 and pi1 represent the sound pickup coefficient of each microphone array.
[0101] Preferably, the data analysis module specifically includes:
[0102] The data extraction submodule is used to extract TV keywords from human voice feature data using a preset voice analysis model to obtain the TV wake-up keyword matching results.
[0103] The instruction extraction submodule is used to extract control instructions from the voice feature data following the wake-up keyword when the wake-up keyword matches, so as to obtain the control instruction data carried in the voice data.
[0104] The instruction comparison submodule is used to compare the control instruction data with the TV's preset instruction library to obtain the control instruction comparison result.
[0105] Preferably, the data control module specifically includes:
[0106] The function switching submodule is used to switch the display functions of the TV according to the correct control instructions and obtain the display program information after the switch.
[0107] The adjustment and analysis submodule is used to obtain the current display atmosphere data, combine it with historical viewing preference data and displayed program information, and generate atmosphere adjustment data that matches the user's viewing preferences.
[0108] The device control submodule is used to adjust parameters of related devices based on the atmosphere adjustment data, optimize the current display effect, and obtain human-computer voice interaction control data for the TV.
[0109] Specific limitations regarding the human-computer voice interaction control system based on smart TVs can be found in the limitations of the human-computer voice interaction control method based on smart TVs mentioned above, and will not be repeated here. Each module in the aforementioned human-computer voice interaction control system based on smart TVs can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of the computer device in hardware form or independent of it, or they can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0110] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 6 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores intermediate data generated during human-computer voice interaction by the smart TV. The network interface communicates with external terminals via a network connection. When the computer program is executed by the processor, it implements a human-computer voice interaction control method based on a smart TV.
[0111] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program being executed by a processor to implement the steps of a human-computer voice interaction control method based on a smart TV.
[0112] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0113] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the system can be divided into different functional units or modules to complete all or part of the functions described above.
[0114] The above-described embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application, and should all be included within the protection scope of this application.
Claims
1. A human-machine voice interaction control method based on a smart television, characterized in that, The method comprises the following steps: acquiring voice data within a preset range, performing data processing on the voice data, and obtaining human voice feature data carrying control instructions; performing feature analysis processing on the human voice feature data through a preset voice analysis model, extracting a wake-up keyword, and performing instruction comparison with a preset instruction library to obtain a control instruction comparison result, specifically including: performing television keyword extraction processing on the human voice feature data through a preset voice analysis model to obtain a wake-up keyword matching result of the television; when the wake-up keyword matching is consistent, performing control instruction extraction processing on the human voice feature data after the wake-up keyword to obtain control instruction data carried in the voice data; performing instruction comparison between the control instruction data and the preset instruction library of the television to obtain a control instruction comparison result; based on the control instruction comparison result, sending a secondary confirmation request to the user, and performing instruction recognition evaluation and instruction deviation correction processing on the basis of the confirmation voice information fed back by the user to obtain a correct control instruction after correction, specifically including: based on the control instruction comparison result, performing sound source positioning on the voice data with consistent instruction comparison to obtain a preliminary sound source positioning result and send a secondary confirmation request to the user; acquiring user secondary confirmation information, and performing secondary sound source positioning processing on the confirmation voice information fed back by the user; based on the sound source positioning coincidence degree between the preliminary sound source positioning result and the secondary sound source positioning result, verifying whether the sound source of the confirmation voice information is compliant; based on the sound source judgment result that the confirmation voice information is compliant, performing accuracy evaluation on the instruction recognition result of the model, and performing deviation correction processing on the deviation instruction according to the evaluation result to obtain a correct control instruction after correction; based on the correct control instruction, performing function switching processing on the television, and based on the display requirement of the program after switching, performing display effect optimization processing on the associated equipment in linkage to obtain human-computer voice interaction control data. 2.The method of human voice interaction control based on smart TV according to claim 1, characterized in that, based on the sound source positioning coincidence degree between the preliminary sound source positioning result and the secondary sound source positioning result, specifically including: calculating the positioning coincidence coefficient between the preliminary sound source positioning and the secondary sound source positioning through formula (1), and analyzing the sound source positioning coincidence degree between the preliminary sound source positioning result and the secondary sound source positioning result based on the positioning coincidence coefficient, formula (1) being as follows: (1); wherein, denotes a positioning coincidence factor, , denote the pickup time delay between each microphone array of the television set and the sound source, respectively, denotes the sound propagation speed, denotes the distance between the microphone arrays of the television set, , denote the pickup coefficients of each microphone array, respectively. 3.The method of human voice interaction control based on smart TV according to claim 1, characterized in that, based on the correct control instruction, performing function switching processing on the television, and based on the display requirement of the program after switching, performing display effect optimization processing on the associated equipment in linkage to obtain human-computer voice interaction control data, specifically including: based on the correct control instruction, performing function switching processing on the display function of the television, and acquiring display program information after switching; acquiring current display atmosphere data, combining historical viewing preference data and the display program information to generate atmosphere adjustment data conforming to the viewing preference of the user; based on the atmosphere adjustment data, adjusting the parameters of the atmosphere associated equipment to optimize the current display effect, and obtaining human-computer voice interaction control data of the television.
4. A human-machine voice interaction control system based on a smart television, characterized in that, The data acquisition module is configured to acquire voice data within a preset range, perform data processing on the voice data, and obtain human voice feature data carrying a control instruction. The data analysis module is configured to perform feature analysis processing on the human voice feature data through a preset voice analysis model, extract a wake-up keyword, and compare the wake-up keyword with a preset instruction library to obtain a control instruction comparison result, specifically including: The data analysis module is configured to perform feature analysis processing on the human voice feature data through a preset voice analysis model, extract a wake-up keyword, and compare the wake-up keyword with a preset instruction library to obtain a control instruction comparison result, specifically including: The data analysis module is configured to perform feature analysis processing on the human voice feature data through a preset voice analysis model, extract a wake-up keyword, and compare the wake-up keyword with a preset instruction library to obtain a control instruction comparison result, specifically including: The data analysis module is configured to perform feature analysis processing on the human voice feature data through a preset voice analysis model, extract a wake-up keyword, and compare the wake-up keyword with a preset instruction library to obtain a control instruction comparison result, specifically including: The data analysis module is configured to perform feature analysis processing on the human voice feature data through a preset voice analysis model, extract a wake-up keyword, and compare the wake-up keyword with a preset instruction library to obtain a control instruction comparison result, specifically including: The data analysis module is configured to perform feature analysis processing on the human voice feature data through a preset voice analysis model, extract a wake-up keyword, and compare the wake-up keyword with a preset instruction library to obtain a control instruction comparison result, specifically including: The data analysis module is configured to perform feature analysis processing on the human voice feature data through a preset voice analysis model, extract a wake-up keyword, and compare the wake-up keyword with a preset instruction library to obtain a control instruction comparison result, specifically including: The data analysis module is configured to perform feature analysis processing on the human voice feature data through a preset voice analysis model, extract a wake-up keyword, and compare the wake-up keyword with a preset instruction library to obtain a control instruction comparison result, specifically including: The data control module is configured to perform function switching processing on the television set according to the correct control instruction, and perform display effect optimization processing on associated devices in linkage with the display requirements of the switched program to obtain human-computer voice interaction control data. The processor executes the computer program to implement the steps of the human-computer voice interaction control method based on the smart television set according to any one of claims 1 to 3.
5. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The computer program is executed by the processor to implement the steps of the human-computer voice interaction control method based on the smart television set according to any one of claims 1 to 3.
6. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 5.
Citation Information
Patent Citations
Sound control program selection method and system for intelligent television
CN107958668A
Voice interaction method, device and system
CN111402900A
Cited By
Industrial equipment voice control method and system based on large language model
CN122474056A