Method and device for realizing voice control terminal based on large model, equipment and medium
By using a large model on a smart TV to identify user intentions and determine the target coordinates in the display screen, and generating a remote control click signal, the problem that existing smart TV voice control cannot effectively control third-party applications or image presentation modules is solved, and the full-scene voice control is realized, improving the user experience.
Patent Information
- Application Number
- CN202510087250.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-20
- Publication Date
- 2025-05-06
AI Technical Summary
The existing smart TV voice control has great limitations, and can only implement the control of a few built-in modules of the control system, and cannot effectively control third-party applications or image presentation modules.
By collecting the controlled voice data input by the user, using the preset large model to identify the voice data, and obtaining the user's intention information. Then, the display screen of the terminal is intercepted, and the coordinates of the target object are determined based on the intention information using the second large model, and a remote control click signal is generated to realize voice control of any interface of the terminal.
It realizes smooth remote control operation in all scenes based on voice mode in any interface of the terminal, improves the user experience and overcomes the limitations of existing smart TV voice control.
Smart Images

Figure CN119943038A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method, device, equipment and medium for implementing a voice control terminal based on a large model. Background Art
[0002] Existing smart TVs usually have the ability to have a voice conversation, that is, without the need for a remote control, one can wake up the TV with a wake-up word and then use voice conversation to interact with the TV.
[0003] At present, many voice-controlled smart TVs can only control the modules built into the system, but cannot control the image presentation modules and modules in third-party applications. For example, there are many application interfaces installed by third-party application stores on the TV, and there are various modules on the interface that need to be clicked by the remote control. Since the system cannot predict and obtain the modules of third-party applications, when the user says: "Open the second one in the second row", the click operation of the second module in the second row cannot be executed. Also, when a certain module is presented as an image, the user will use voice to express the interaction through the perception of what the eyes see. When the user says: "Open the one in the red skirt", the traditional smart TV cannot know the user's intention and how to help the user click to open the module.
[0004] It can be seen that the existing smart TV voice control has great limitations and can only realize the control of a small part of the built-in modules of the control system. Summary of the invention
[0005] The embodiments of the present invention provide a method, device, equipment and medium for implementing a voice control terminal based on a large model, aiming to solve the problem that the existing smart TV voice control has large limitations and can only realize the control of a small part of the built-in modules of the control system.
[0006] In a first aspect, an embodiment of the present invention provides a method for implementing a voice control terminal based on a large model, which includes:
[0007] Collecting control voice data input by the user;
[0008] Recognize the control voice data by using a preset first large model to obtain intention information;
[0009] Intercepting a display screen of the terminal, and determining the coordinates of the target object in the display screen based on the intention information by using a preset second largest model;
[0010] A remote control click signal is generated at the coordinates of the target object.
[0011] A further technical solution is that the control voice data is recognized by a preset first large model to obtain the intention information, including:
[0012] Converting the voice data into text data;
[0013] Inputting the text data into the first large model;
[0014] Receive the intent information recognized by the first model based on the text data.
[0015] A further technical solution is that the determining the coordinates of the target object in the display screen based on the intention information by using the preset second largest model includes:
[0016] Uploading the intention information and the display screen to a server, wherein the server calls the second large model to determine the coordinates of the target object in the display screen based on the intention information;
[0017] Receive the coordinates of the target object returned by the server.
[0018] A further technical solution is that the method further comprises:
[0019] Determining whether a preset wake-up signal is received;
[0020] If a preset wake-up signal is received, the step of collecting the control voice data input by the user is executed.
[0021] A further technical solution is that before intercepting the display screen of the terminal, the method further includes:
[0022] Determine whether the intention information belongs to a remote control click operation intention;
[0023] If the intention information belongs to a remote control click operation intention, the step of capturing the display screen of the terminal and determining the coordinates of the target object in the display screen based on the intention information through a preset second largest model is executed.
[0024] A further technical solution is that the intention information includes at least one of a person, an object, a color, an object, and a position.
[0025] A further technical solution is that the method further comprises:
[0026] A target operation is performed based on a remote controller click signal at the coordinates of the target object.
[0027] In a second aspect, an embodiment of the present invention further provides a device for implementing a voice control terminal based on a large model, which includes a unit for executing the above method.
[0028] In a third aspect, an embodiment of the present invention further provides a computer device, which includes a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the above method when executing the computer program.
[0029] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, wherein the storage medium stores a computer program, and the computer program can implement the above method when executed by a processor.
[0030] The embodiments of the present invention provide a method, device, equipment and medium for implementing voice control of a terminal based on a large model. The method includes: collecting control voice data input by a user; identifying the control voice data through a preset first large model to obtain intent information; intercepting the display screen of the terminal, and determining the coordinates of a target object in the display screen based on the intent information through a preset second large model; and generating a remote control click signal at the coordinates of the target object. The present invention identifies the user's intention through a large model, determines the click position in the terminal display screen based on the user's intention, and simulates the remote control click operation, thereby realizing full-scene smooth remote control operation based on voice in any interface of the terminal, and improving the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying any creative work.
[0032] Figure 1 A flowchart of a method for implementing a voice control terminal based on a large model provided by an embodiment of the present invention;
[0033] Figure 2 Another flowchart of a method for implementing voice control of a terminal based on a large model provided by an embodiment of the present invention;
[0034] Figure 3 A schematic block diagram of a device for implementing a voice control terminal based on a large model provided by an embodiment of the present invention;
[0035] Figure 4 A schematic block diagram of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0036] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0037] It should be understood that when used in this specification and the appended claims, the terms "include" and "comprises" indicate the presence of described features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.
[0038] It should also be understood that the terms used in this specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the specification of the present invention and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.
[0039] It should be further understood that the term "and / or" used in the present description and the appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0040] As used in this specification and the appended claims, the term "if" may be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if [described condition or event] is detected" may be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.
[0041] See also Figure 1-Figure 2 The embodiment of the present invention provides a method for realizing voice control of a terminal based on a large model. The method can generate a click signal of a remote controller based on the voice data input by the user in any screen, realize convenient voice control of the terminal (smart TV) in all scenarios, and improve the user experience. Figure 1 As shown, the method comprises the following steps:
[0042] S1, collecting control voice data input by the user.
[0043] In a specific implementation, a terminal (such as a smart TV) is configured with a voice collection module, and the voice collection module is used to collect control voice data input by the user.
[0044] For example, in some embodiments, it is determined whether a preset wake-up signal is received; if the preset wake-up signal is received, the step of collecting the control voice data input by the user is executed.
[0045] Specifically, when the terminal receives the wake-up signal, it enters a listening state, waiting for the user to input voice data, and when the user initiates a voice conversation, the terminal collects the control voice data input by the user.
[0046] S2, recognizing the control voice data through a preset first large model to obtain intention information.
[0047] In a specific implementation, the first large model can be specifically a large model intention recognition service, and the large model intention recognition service can recognize the intention of the control voice data and obtain intention information. The intention information includes at least one of a person, an object, a color, an object, and a position. Specifically, the intention information includes the name of the intention and the person, object, color, object, position, etc. described by the user voice.
[0048] In some embodiments, the above step of "recognizing the control voice data through a preset first large model to obtain intention information" specifically includes the following steps: converting the voice data into text data; inputting the text data into the first large model; receiving the intention information recognized by the first large model based on the text data.
[0049] Specifically, in order to improve the accuracy of intent recognition, the voice data is converted into text data to facilitate recognition by the first large model; further, the text data is input into the first large model; and the intent information recognized by the first large model based on the text data is received.
[0050] S3, capturing the display screen of the terminal, and determining the coordinates of the target object in the display screen based on the intention information through a preset second largest model.
[0051] In the specific implementation, it is determined whether the intention information belongs to the remote control click operation intention; if the intention information belongs to the remote control click operation intention, the display screen of the terminal is captured, and the coordinates of the target object in the display screen are determined based on the intention information through a preset second largest model.
[0052] Specifically, the intent information includes the recognized intent. If the recognized intent is a remote control click operation intention, the current display screen of the terminal is captured. If it is another intention, there is no need to capture the current display screen of the terminal.
[0053] In some embodiments, the above step of "determining the coordinates of the target object in the display screen based on the intention information through a preset second large model" specifically includes the following steps: uploading the intention information and the display screen to a server, wherein the server calls the second large model to determine the coordinates of the target object in the display screen based on the intention information; and receiving the coordinates of the target object returned by the server.
[0054] In a specific implementation, the second large model may be a large model image recognition service. The server calls the large model image recognition service to determine the coordinates of the target object in the display screen based on the intent information, and returns the coordinates of the target object to the terminal.
[0055] For example: the control voice data is "I want to see the one wearing red clothes in the upper left corner", and the corresponding intent information is: position = upper left corner, color = red, object = clothes. The large model image recognition service recognizes the coordinates as [15,235], and then returns the corresponding coordinate information to the terminal (smart TV).
[0056] S4, generating a remote control click signal at the coordinates of the target object.
[0057] In a specific implementation, a remote control click signal is generated at the coordinates of the target object, that is, a remote control click signal of the remote control at the coordinates of the target object is simulated, thereby realizing the generation of the remote control click signal through voice control. Further, the terminal performs a target operation based on the remote control click signal at the coordinates of the target object, and the target operation can be, for example, screen jump, which is not specifically limited in the embodiment of the present invention.
[0058] It can be seen that in the embodiment of the present invention, after receiving the coordinates returned by the server, the terminal executes the simulated remote control click event, that is, clicks the coordinate position, thus realizing the voice full-scene any interface, any complex description, can correspond to the remote control click operation. Not limited to the system built-in interface, third-party application interface, web page interface, picture screen or interface generated by dynamic operation, the remote control click operation can be executed through the recognized intention information, realizing the full-scene smooth remote control operation based on voice.
[0059] For example: if the control voice data is "open the black tiger one", the big model intention recognition determines that the user wants to click on a module on the interface. Then the big model performs image analysis on the current display screen, finds the module where the black tiger is located, and determines the coordinates of the remote control click, thus converting the user's complex voice expression into a remote control click behavior.
[0060] The embodiment of the present invention proposes a method for realizing voice control of a terminal based on a large model, the method comprising: collecting control voice data input by a user; identifying the control voice data through a preset first large model to obtain intention information; intercepting the display screen of the terminal, and determining the coordinates of a target object in the display screen based on the intention information through a preset second large model; and generating a remote control click signal at the coordinates of the target object. The present invention recognizes the user's intention through a large model, determines the click position in the terminal display screen based on the user's intention, and simulates the remote control click operation, thereby realizing full-scene smooth remote control operation based on voice in any interface of the terminal, and improving the user experience.
[0061] See also Figure 3 , Figure 3 : is a schematic block diagram of a device 20 for implementing a voice control terminal based on a large model provided by an embodiment of the present invention. Corresponding to the above method for implementing a voice control terminal based on a large model, the present invention also provides a device 20 for implementing a voice control terminal based on a large model. The device 20 for implementing a voice control terminal based on a large model includes a unit for executing the above method for implementing a voice control terminal based on a large model. The device 20 for implementing a voice control terminal based on a large model can be configured in a desktop computer, a tablet computer, a laptop computer, and other terminals. Specifically, the device 20 for implementing a voice control terminal based on a large model includes:
[0062] A collection unit 21, used to collect control voice data input by the user;
[0063] The recognition unit 22 is used to recognize the control voice data through a preset first large model to obtain intention information;
[0064] A determination unit 23, configured to capture a display screen of the terminal and determine the coordinates of a target object in the display screen based on the intention information by using a preset second largest model;
[0065] The generating unit 24 is configured to generate a remote control click signal at the coordinates of the target object.
[0066] In some embodiments, the identifying the control voice data by using a preset first large model to obtain the intention information includes:
[0067] Converting the voice data into text data;
[0068] Inputting the text data into the first large model;
[0069] Receive the intent information recognized by the first model based on the text data.
[0070] In some embodiments, determining the coordinates of the target object in the display screen based on the intention information by using a preset second large model includes:
[0071] Uploading the intention information and the display screen to a server, wherein the server calls the second large model to determine the coordinates of the target object in the display screen based on the intention information;
[0072] Receive the coordinates of the target object returned by the server.
[0073] In some embodiments, the apparatus 20 for implementing voice control of a terminal based on a large model further includes:
[0074] A first determination unit, used to determine whether a preset wake-up signal is received;
[0075] The first jump unit is used to execute the step of collecting the control voice data input by the user if a preset wake-up signal is received.
[0076] In some embodiments, the apparatus 20 for implementing voice control of a terminal based on a large model further includes:
[0077] A second judgment unit, used to judge whether the intention information belongs to a remote control click operation intention;
[0078] The second jump unit is used to execute the step of capturing the display screen of the terminal and determining the coordinates of the target object in the display screen based on the intention information through a preset second large model if the intention information belongs to a remote control click operation intention.
[0079] In some embodiments, the intent information includes at least one of a person, an object, a color, an object, and a location.
[0080] In some embodiments, the apparatus 20 for implementing voice control of a terminal based on a large model further includes:
[0081] An execution unit is used to execute a target operation based on a remote controller click signal at the coordinates of the target object.
[0082] It should be noted that technical personnel in the relevant field can clearly understand that the specific implementation process of the above-mentioned device 20 and each unit for realizing a voice control terminal based on a large model can refer to the corresponding description in the aforementioned method embodiment, and for the convenience and conciseness of the description, it will not be repeated here.
[0083] The above-mentioned device 20 for realizing voice control terminal based on large model can be realized in the form of a computer program, and the computer program can be used in Figure 4 Runs on the computer device shown.
[0084] See also Figure 4 , Figure 4 5 is a schematic block diagram of a computer device provided in an embodiment of the present application. The computer device 500 may be a terminal or a server, wherein the terminal may be an electronic device with communication functions such as a smart phone, a tablet computer, a laptop computer, a desktop computer, a personal digital assistant, and a wearable device. The server may be an independent server or a server cluster composed of multiple servers.
[0085] The computer device 500 includes a processor 502 , a memory, and a network interface 505 connected via a system bus 501 , wherein the memory may include a non-volatile storage medium 503 and an internal memory 504 .
[0086] The non-volatile storage medium 503 can store an operating system 5031 and a computer program 5032. When the computer program 5032 is executed, the processor 502 can execute a method for implementing a voice control terminal based on a large model.
[0087] The processor 502 is used to provide computing and control capabilities to support the operation of the entire computer device 500 .
[0088] The internal memory 504 provides an environment for the operation of the computer program 5032 in the non-volatile storage medium 503. When the computer program 5032 is executed by the processor 502, the processor 502 can execute a method for implementing a voice control terminal based on a large model.
[0089] The network interface 505 is used to communicate with other devices over the network. Those skilled in the art will appreciate that the above structure is only a block diagram of a portion of the structure related to the present application solution, and does not constitute a limitation on the computer device 500 to which the present application solution is applied. The specific computer device 500 may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0090] The processor 502 is used to run a computer program 5032 stored in the memory to implement the steps of a method for implementing a voice control terminal based on a large model provided in any of the above method embodiments.
[0091] It should be understood that in the embodiment of the present application, the processor 502 may be a central processing unit (CPU), and the processor 502 may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0092] It is understood by those skilled in the art that all or part of the processes in the method for implementing the above embodiment can be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a storage medium, which is a computer-readable storage medium. The computer program is executed by at least one processor in the computer system to implement the process steps of the embodiment of the above method.
[0093] Therefore, the present invention also provides a storage medium. The storage medium may be a computer-readable storage medium. The storage medium stores a computer program. When the computer program is executed by a processor, the processor performs the following steps: steps of a method for implementing a voice control terminal based on a large model provided in any of the above method embodiments.
[0094] The storage medium is a physical, non-transient storage medium, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a magnetic disk, or an optical disk, etc., which can store program codes. The computer-readable storage medium can be non-volatile or volatile.
[0095] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0096] In the several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of each unit is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed.
[0097] The steps in the method of the embodiment of the present invention can be adjusted in order, combined and deleted according to actual needs. The units in the device of the embodiment of the present invention can be combined, divided and deleted according to actual needs. In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0098] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, terminal, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention.
[0099] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0100] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalents, the present invention is also intended to include these modifications and variations.
[0101] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed by the present invention, and these modifications or replacements should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be based on the protection scope of the claims.
Claims
1. A method for implementing voice control of a terminal based on a large model, characterized in that: include: Collecting control voice data input by the user; Recognize the control voice data by using a preset first large model to obtain intention information; Intercepting a display screen of the terminal, and determining the coordinates of the target object in the display screen based on the intention information by using a preset second largest model; A remote control click signal is generated at the coordinates of the target object.
2. The method for implementing voice control of a terminal based on a large model according to claim 1, characterized in that: The step of identifying the control voice data by using a preset first large model to obtain the intention information includes: Converting the voice data into text data; Inputting the text data into the first large model; Receive the intent information recognized by the first model based on the text data.
3. The method for implementing voice control of a terminal based on a large model according to claim 1, characterized in that: The determining the coordinates of the target object in the display screen based on the intention information by using the preset second largest model includes: Uploading the intention information and the display screen to a server, wherein the server calls the second large model to determine the coordinates of the target object in the display screen based on the intention information; Receive the coordinates of the target object returned by the server.
4. The method for implementing voice control of a terminal based on a large model according to claim 1, characterized in that: The method further comprises: Determining whether a preset wake-up signal is received; If a preset wake-up signal is received, the step of collecting the control voice data input by the user is executed.
5. The method for implementing voice control of a terminal based on a large model according to claim 1, characterized in that: Before intercepting the display screen of the terminal, the method further includes: Determine whether the intention information belongs to a remote control click operation intention; If the intention information belongs to a remote control click operation intention, the step of capturing the display screen of the terminal and determining the coordinates of the target object in the display screen based on the intention information through a preset second largest model is executed.
6. The method for implementing voice control of a terminal based on a large model according to claim 1, characterized in that: The intention information includes at least one of a person, an object, a color, an object, and a position.
7. The method for implementing voice control of a terminal based on a large model according to claim 1, characterized in that: The method further comprises: A target operation is performed based on a remote controller click signal at the coordinates of the target object.
8. A device for implementing voice control of a terminal based on a large model, characterized in that: The method comprises a unit for executing the method according to any one of claims 1 to 7.
9. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 7 can be implemented.