Display equipment and awakening method and device thereof
By adjusting the parameters of the target wake-up recognition model, combining the wake-up frequency and voice characteristics of the sound source object, the voice wake-up process of the display device is optimized, and the false wake-up and missed wake-up problems of traditional wake-up methods are solved, achieving higher recognition accuracy and adaptability.
Patent Information
- Application Number
- CN202510873906.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-08-26
AI Technical Summary
In traditional voice wake-up methods, based on a single voice template or a simple voiceprint matching method, the accuracy of the display device's wake-up recognition is insufficient, and false wake-up or missed wake-up problems are prone to occur.
By obtaining the pre-trained target wake-up recognition model, combining the wake-up frequency and voice characteristics of the sound source object, adjusting the model parameters, optimizing the wake-up recognition model to adapt to the current application scenario, using the high-frequency characteristics of the sound source object to improve recognition sensitivity, and updating the model parameters under the conditions of meeting.
It improves the accuracy and recognition accuracy of voice wake-up, reduces the situation of false wake-up and missed wake-up, and enhances the adaptability of display devices in different scenarios.
Smart Images

Figure CN120544571A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of display technology, and in particular to a display device and a wake-up method and apparatus thereof. Background Art
[0002] With the development of display technology, voice interaction technology has gradually become a key means of controlling display devices. Among them, the display can be woken up by voice audio, improving the efficiency of human-computer interaction and user experience.
[0003] However, traditional voice wake-up methods, which typically rely on a single, preset voice template or simple voiceprint matching to wake up the display, are prone to false or missed wake-ups in different application scenarios, affecting the accuracy of wake-up recognition. Summary of the Invention
[0004] Based on this, it is necessary to provide a display device and a wake-up method and apparatus thereof to address the above technical issues, so as to improve the accuracy of wake-up recognition.
[0005] In a first aspect, some embodiments provide a display device, including:
[0006] a display configured to display a user interface or a standby interface;
[0007] An audio input interface configured to collect voice interaction data;
[0008] At least one controller is connected to the audio input interface and the display and is configured to:
[0009] Obtain a pre-trained target arousal recognition model;
[0010] When the model parameter update conditions are met, voice attribute data corresponding to the target voice interaction data is obtained; the target voice interaction data is voice interaction data collected by the audio input interface within a preset time period; the voice attribute data includes at least one sound source object and the voice features of the corresponding sound source object;
[0011] Obtaining a target wake-up frequency corresponding to each sound source object; the target wake-up frequency corresponding to the sound source object is the frequency at which the sound source object wakes up the display within a preset time period;
[0012] Adjusting the model parameters in the target awakening recognition model according to the target awakening frequency of each sound source object and the speech characteristics of the corresponding sound source object to update the target awakening recognition model;
[0013] Get the current voice interaction data collected by the audio input interface at the current moment;
[0014] Input the current voice interaction data into the updated target wakeup recognition model to obtain the wakeup recognition result corresponding to the current voice interaction data; the wakeup recognition result includes the untriggered wakeup result and the triggered wakeup result;
[0015] When the wake-up recognition result is a triggered wake-up result, the display is controlled to switch from displaying the standby interface to displaying the user interface.
[0016] In the above embodiment, obtaining a pre-trained target wakeup recognition model provides a foundational framework for subsequent targeted updates to the target wakeup recognition model. By obtaining the speech features of each sound source object corresponding to the target voice interaction data within a preset time period, and obtaining the target wakeup frequency corresponding to the sound source object, provided the model parameters are adapted to the current application scenario, providing the target wakeup recognition model with basic data for updating. By adjusting the model parameters of the target wakeup recognition model based on the target wakeup frequency and speech features of the sound source object, the target wakeup recognition model can be made more suitable for the current application scenario, improving its recognition accuracy in that scenario. Furthermore, based on the target wakeup frequency, the model can prioritize high-frequency speech features, thereby improving the recognition sensitivity of active sound source objects. Current voice interaction data collected by the audio input interface at the current moment is obtained; the current voice interaction data is input into the target wakeup recognition model to obtain a wakeup recognition result corresponding to the current voice interaction data; and if the wakeup recognition result is a triggered wakeup result, the display is controlled to switch from displaying the standby interface to displaying the user interface. This ensures accurate recognition of the current voice interaction data, thereby improving the accuracy of voice wakeup.
[0017] In some embodiments, when the controller adjusts the model parameters in the target wake-up recognition model according to the target wake-up frequency of each sound source object and the voice features of the corresponding sound source object, it is configured as follows: for each sound source object, the feature weight of each voice feature of the sound source object is determined according to the target proportion of the target wake-up frequency of the sound source object in the total target wake-up frequency; the total target wake-up frequency is the sum of the target wake-up frequencies of each sound source object; for each voice feature of each sound source object, the baseline model parameters corresponding to the voice feature are obtained; the baseline model parameters are obtained by training the target wake-up recognition model to be trained using sample voice interaction data; according to the feature weight of the voice feature, the baseline model parameters of the voice feature are weighted to obtain the target model parameters of the voice feature; according to the target model parameters of the different voice features of each sound source object, the model parameters in the target wake-up recognition model are adjusted.
[0018] In the above embodiment, by determining the feature weight of each speech feature of a sound source object based on the target proportion of the target wakeup frequency of the sound source object in the total target wakeup frequency, the wakeup frequency difference corresponding to each speech feature is quantified, forming a feature importance gradient, and providing a quantitative basis for differential parameter optimization. By obtaining the corresponding baseline model parameters for each speech feature of each sound source object, a corresponding benchmark reference is provided for subsequent weighted calculations. Because the baseline model parameters are obtained by training the target wakeup recognition model to be trained using sample speech interaction data with speech features, the matching between the speech features and their corresponding baseline model parameters is improved. By weighting the baseline model parameters of the speech features according to the feature weights of the speech features to obtain the target model parameters of the speech features, and adjusting the model parameters in the target wakeup recognition model based on the target model parameters of the different speech features of each sound source object, the more frequently occurring speech features produce more significant gradient update directions in the parameter space. This allows the updated target wakeup recognition model to adapt to the current application scenario, thereby improving the accuracy of speech wakeup.
[0019] In some embodiments, the baseline model parameters are incremental model parameters corresponding to the fixed model parameters in the target wakeup recognition model; during the training of the target wakeup recognition model, the fixed model parameters in the target wakeup recognition model are kept unchanged, and the incremental model parameters are adjusted; accordingly, when the controller adjusts the model parameters in the target wakeup recognition model according to the target model parameters of the different voice features of each sound source object, it is configured to: for each fixed model parameter, perform parameter fusion on the target model parameters corresponding to the fixed model parameter to obtain a fused model parameter; perform parameter fusion on the fused model parameter and the fixed model parameter to adjust the model parameters in the target wakeup recognition model.
[0020] In the above embodiment, by introducing incremental model parameters and, during the training of the target wakeup recognition model, maintaining the incremental model parameters in the target wakeup recognition model unchanged while adjusting the newly added model parameters, computing resource usage and time costs are reduced. By fusing the target model parameters corresponding to each fixed model parameter, the fused model parameters corresponding to the fixed model parameter can be quickly obtained. By fusing the fused model parameters with the fixed model parameters, the model parameters in the target wakeup recognition model can be adjusted without the need for further model training, which helps improve the model update efficiency of the target wakeup recognition model.
[0021] In some embodiments, the model parameter update condition includes at least one of the following: detecting that the display device is accessing the network for the first time; triggering the target timing task of the timer; the target timing task is used to indicate that the target wake-up recognition model is updated according to a preset period or preset frequency; receiving a model parameter update operation.
[0022] In the above embodiment, the triggering of updating the target wake-up recognition model in different usage scenarios is realized, which is conducive to improving the flexibility of the target wake-up recognition model update. When it is detected that the display device is accessing the network for the first time, that is, for the first use scenario, the model parameters of the target wake-up recognition model are updated, so that the target wake-up recognition model can be adapted to the current application scenario. By setting the target timing task of the timer, it is possible to update the model parameters according to a preset period or preset frequency, avoiding the inability to adapt to changes in application scenarios due to long-term use of fixed model parameters. By responding to the model parameter update operation, the model parameters of the target wake-up recognition model are updated, so that the model parameters can be flexibly updated according to actual needs.
[0023] In some embodiments, the controller is further configured to: when the model parameter update condition is not met, use a pre-trained target wake-up recognition model to recognize the wake-up recognition result of the voice interaction data.
[0024] In the above embodiment, when the model parameter update condition is not met, the pre-trained target wake-up recognition model is used to recognize the wake-up recognition result of the current voice interaction data, thereby realizing full-cycle wake-up recognition for the voice interaction data.
[0025] In some embodiments, when the controller executes to obtain the target wake-up frequency corresponding to each sound source object, it is configured to: for each sound source object, obtain the record items in which the wake-up time of the sound source object is within a preset time period from the interaction log record; obtain the number of each record item to obtain the target wake-up frequency corresponding to the sound source object; wherein the interaction log record records the sound source object and the corresponding wake-up time corresponding to the triggered wake-up result.
[0026] In the above embodiment, by introducing interaction log records, and recording the sound source object and corresponding wake-up time corresponding to the triggering wake-up result in the interaction log records, it is convenient to obtain the target wake-up frequency of different sound source objects within the preset time period.
[0027] In some embodiments, obtaining voice attribute data corresponding to the target voice interaction data includes: calling a voiceprint recognition model to extract voiceprint data from the target voice interaction data, and identifying a sound source object corresponding to the voiceprint data.
[0028] In the above embodiment, by extracting the voiceprint data from the target voice interaction data, it is convenient to extract voiceprint parameters such as fundamental frequency and formant from the voiceprint data, thereby facilitating the improvement of the recognition efficiency of the sound source object.
[0029] In some embodiments, the speech features of the sound source object include basic attribute features and language and regional features; obtaining the speech attribute data corresponding to the target voice interaction data includes: calling a voiceprint recognition model to extract the voiceprint data in the target voice interaction data, and identifying the sound source object and the basic attribute features of the sound source object in the voiceprint data; calling a language recognition model to extract the regional characteristic vocabulary of the sound source object in the target voice interaction data, and identifying the language and regional features corresponding to the regional characteristic vocabulary.
[0030] In the above embodiment, because the speech characteristics of the sound source object include basic attribute characteristics, the updated target wakeup recognition model can improve recognition response across dimensions such as age and gender. Because the speech characteristics of the sound source object include language and regional characteristics, the updated target wakeup recognition model can reduce missed wakeups caused by dialect accents.
[0031] In a second aspect, some embodiments provide a method for waking up a display device, including:
[0032] Obtain a pre-trained target arousal recognition model;
[0033] When the model parameter update conditions are met, voice attribute data corresponding to target voice interaction data is obtained; the target voice interaction data is voice interaction data collected by the audio input interface of the display device within a preset time period; the voice attribute data includes at least one sound source object and the voice features of the corresponding sound source object;
[0034] Obtaining a target wake-up frequency corresponding to each sound source object; the target wake-up frequency corresponding to the sound source object is the frequency at which the sound source object wakes up the display of the display device within a preset time period;
[0035] Adjusting the model parameters in the target awakening recognition model according to the target awakening frequency of each sound source object and the speech characteristics of the corresponding sound source object to update the target awakening recognition model;
[0036] Get the current voice interaction data collected by the audio input interface at the current moment;
[0037] Input the current voice interaction data into the updated target wakeup recognition model to obtain the wakeup recognition result corresponding to the current voice interaction data; the wakeup recognition result includes the untriggered wakeup result and the triggered wakeup result;
[0038] When the wake-up recognition result is a triggered wake-up result, the display is controlled to switch from displaying the standby interface to displaying the user interface.
[0039] In the above embodiment, obtaining a pre-trained target wakeup recognition model provides a foundational framework for subsequent targeted updates to the target wakeup recognition model. By obtaining the speech features of each sound source object corresponding to the target voice interaction data within a preset time period, and obtaining the target wakeup frequency corresponding to the sound source object, provided the model parameters are adapted to the current application scenario, providing the target wakeup recognition model with basic data for updating. By adjusting the model parameters of the target wakeup recognition model based on the target wakeup frequency and speech features of the sound source object, the target wakeup recognition model can be made more suitable for the current application scenario, improving its recognition accuracy in that scenario. Furthermore, based on the target wakeup frequency, the model can prioritize high-frequency speech features, thereby improving the recognition sensitivity of active sound source objects. Current voice interaction data collected by the audio input interface at the current moment is obtained; the current voice interaction data is input into the target wakeup recognition model to obtain a wakeup recognition result corresponding to the current voice interaction data; and if the wakeup recognition result is a triggered wakeup result, the display is controlled to switch from displaying the standby interface to displaying the user interface. This ensures accurate recognition of the current voice interaction data, thereby improving the accuracy of voice wakeup.
[0040] In a third aspect, some embodiments provide a device for waking up a display device, including:
[0041] A first acquisition module is used to acquire a pre-trained target arousal recognition model;
[0042] A second acquisition module, when a model parameter update condition is met, acquires speech attribute data corresponding to target speech interaction data; the target speech interaction data is speech interaction data collected by an audio input interface of a display device within a preset time period; the speech attribute data includes at least one sound source object and speech features of the corresponding sound source object;
[0043] The third acquisition module acquires the target wake-up frequency corresponding to each sound source object; the target wake-up frequency corresponding to the sound source object is the frequency at which the sound source object wakes up the display of the display device within a preset time period;
[0044] An updating module adjusts model parameters in the target awakening recognition model according to the target awakening frequency of each sound source object and the speech characteristics of the corresponding sound source object to update the target awakening recognition model;
[0045] The fourth acquisition module acquires the current voice interaction data collected by the audio input interface at the current moment;
[0046] The input module inputs the current voice interaction data into the updated target wakeup recognition model to obtain the wakeup recognition result corresponding to the current voice interaction data; the wakeup recognition result includes the untriggered wakeup result and the triggered wakeup result;
[0047] The control module controls the display to switch from displaying the standby interface to displaying the user interface when the wake-up recognition result is a trigger wake-up result.
[0048] In the above embodiment, obtaining a pre-trained target wakeup recognition model provides a foundational framework for subsequent targeted updates to the target wakeup recognition model. By obtaining the speech features of each sound source object corresponding to the target voice interaction data within a preset time period, and obtaining the target wakeup frequency corresponding to the sound source object, provided the model parameters are adapted to the current application scenario, providing the target wakeup recognition model with basic data for updating. By adjusting the model parameters of the target wakeup recognition model based on the target wakeup frequency and speech features of the sound source object, the target wakeup recognition model can be made more suitable for the current application scenario, improving its recognition accuracy in that scenario. Furthermore, based on the target wakeup frequency, the model can prioritize high-frequency speech features, thereby improving the recognition sensitivity of active sound source objects. Current voice interaction data collected by the audio input interface at the current moment is obtained; the current voice interaction data is input into the target wakeup recognition model to obtain a wakeup recognition result corresponding to the current voice interaction data; and if the wakeup recognition result is a triggered wakeup result, the display is controlled to switch from displaying the standby interface to displaying the user interface. This ensures accurate recognition of the current voice interaction data, thereby improving the accuracy of voice wakeup.
[0049] In a fourth aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the method of the above-mentioned second aspect in various possible ways when executing the computer program.
[0050] In a fifth aspect, the present application also provides a computer-readable storage medium having a computer program stored thereon, which implements the method of the above-mentioned second aspect in various possible ways when the computer program is executed by a processor.
[0051] In a sixth aspect, the present application also provides a computer program product, comprising a computer program, which, when executed by a processor, implements the method of the above-mentioned second aspect in various possible ways. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following briefly introduces the drawings required for use in the embodiments of the present application or related technical descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying any creative work.
[0053] Figure 1 A schematic diagram of an operation scenario between a display device and a control device provided in some embodiments;
[0054] Figure 2 A schematic diagram of the hardware configuration of a display device provided in some embodiments;
[0055] Figure 3 A schematic diagram of the hardware configuration of a control device provided in some embodiments;
[0056] Figure 4 A schematic diagram of software configuration of a display device provided in some embodiments;
[0057] Figure 5 A schematic flow chart of a method for waking up a display device provided in some embodiments;
[0058] Figure 6 An interactive timing diagram of a wake-up method for a display device provided in some embodiments;
[0059] Figure 7A A schematic flow chart of steps for adjusting model parameters provided in some embodiments;
[0060] Figure 7B A schematic diagram of the structure of a target wakeup recognition model provided in some embodiments;
[0061] Figure 7C A schematic diagram of the structure of an updated target wakeup recognition model provided in some embodiments;
[0062] Figure 8 Interaction sequence diagrams of a method for waking up a display device in some other embodiments provided for some embodiments;
[0063] Figure 9 A schematic structural diagram of a wake-up device for a display device provided in some embodiments;
[0064] Figure 10 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0065] The following embodiments are described in detail, with examples illustrated in the accompanying drawings. When the following description refers to the drawings, identical numbers in different figures represent identical or similar elements unless otherwise indicated. The embodiments described in the following embodiments are not intended to represent all possible implementations consistent with the present application. They are merely examples of systems and methods consistent with certain aspects of the present application, as detailed in the claims.
[0066] It should be noted that the brief descriptions of terms in this application are only for the purpose of facilitating the understanding of the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise specified, these terms should be understood according to their ordinary and usual meanings.
[0067] In the specification and claims of this application and the accompanying drawings, the terms "first," "second," "third," etc. are used to distinguish similar or similar objects or entities, and are not necessarily intended to limit a particular order or sequence, unless otherwise noted. It should be understood that the terms used in this manner are interchangeable under appropriate circumstances.
[0068] The terms "comprise," "comprises," and "having," and any variations thereof, are intended to cover but not exclude inclusion; for example, a product or device comprising a list of components is not necessarily limited to all the components expressly listed but may include other components not expressly listed or inherent to such product or device.
[0069] The term "module" refers to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code that is capable of performing the functionality associated with that element.
[0070] In the embodiments of the present application, the display device 200 generally refers to a device capable of displaying images and processing data. For example, the display device 200 includes but is not limited to a smart TV, a mobile terminal, a computer, a monitor, an advertising screen, a wearable device, a virtual reality device, an augmented reality device, etc.
[0071] Figure 1 This is a schematic diagram of an operation scenario between a display device and a control device provided in some embodiments of the present application. Figure 1 As shown in FIG, a user can operate the display device 200 through touch operation, the mobile terminal 300 and the control device 100. For example, the control device 100 can be a remote controller, a stylus pen, a handle, etc.
[0072] The mobile terminal 300 can function as a control device for performing human-computer interaction between a user and the display device 200. The mobile terminal 300 can also function as a communication device for establishing a communication connection with the display device 200 for data exchange. In some embodiments, the mobile terminal 300 can install software applications with the display device 200, enabling communication via a network communication protocol for one-to-one control and data communication. Audio and video content displayed on the mobile terminal 300 can also be transmitted to the display device 200 for synchronized display.
[0073] like Figure 1 As shown in FIG, the display device 200 also communicates data with the server 400 through various communication methods. The display device 200 may be allowed to communicate via a local area network (LAN), a wireless local area network (WLAN), and other networks.
[0074] The display device 200 may provide a broadcast receiving television function, and may also additionally provide an intelligent network television function with a computer support function, including but not limited to network television, smart TV, Internet Protocol television (IPTV), etc.
[0075] Figure 2 Some embodiments of this application provide Figure 1 2 is a block diagram of the hardware configuration of the display device 200.
[0076] In some embodiments, the display device 200 may include at least one of a tuner 210, a communication device 220, a detector 230, a device interface 240, a controller 250, a display 260, an audio output device 270, a memory, a power supply, and a user input interface.
[0077] In some embodiments, detector 230 is used to collect signals from the external environment or external interactions. For example, detector 230 may include a light receiver, such as a sensor for collecting ambient light intensity; or an image collector, such as a camera, for collecting external environmental scenes, user attributes, or user interaction gestures; or a sound collector, such as a microphone, for receiving external sounds. The sound collector may transmit the collected voice interaction data to controller 250 via an audio input interface.
[0078] In some embodiments, the display 260 includes a display function component for presenting a picture, and a driving component for driving the image display. The display 260 is used to receive an image signal output from the controller 250 for display. For example, the display 260 can be used to display video content, image content, components of a menu control interface, and a user control UI interface. Among them, the display 260 can display a user interface or a standby interface. The user interface can be understood as an interface displayed when the display 260 is in a non-standby state. For example, the user interface can include the main interface of the display 260 or the interface displayed before entering the standby state. The standby interface can be understood as an interface displayed in the standby state. For example, the standby interface can include at least one of a black screen interface, a static prompt interface, a dynamic always interface, or a screen saver animation interface. This application does not impose any restrictions on the specific interface display content of the user interface and the standby interface.
[0079] In some embodiments, the communication device 220 is a component used to communicate with an external device or server 400 according to various communication protocol types. The display device 200 can be provided with multiple communication devices 220 depending on the supported communication methods. For example, if the display device 200 supports wireless network communication, the display device 200 can be provided with a communication device 220 including WiFi functionality. If the display device 200 supports Bluetooth connection communication, the display device 200 needs to be provided with a communication device 220 including Bluetooth functionality.
[0080] The communication device 220 can establish a communication connection between the display device 200 and an external device or server 400 via a wireless or wired connection. A wired connection can connect the display device 200 to an external device via a data cable, an interface, or other components. A wireless connection can connect the display device 200 to an external device via a wireless signal or wireless network. The display device 200 can establish a connection with an external device directly or indirectly through a gateway, router, or connection device.
[0081] In some embodiments, the controller 250 may include at least one of a central processing unit (CPU), a video processor, an audio processor, a graphics processor, and a power processor, and first to nth interfaces for input / output. The controller 250 controls the operation of the display device and responds to user operations using various software control programs stored in a memory. The controller 250 controls the overall operation of the display device 200. The controller 250 may communicate with the server 400 via a network.
[0082] For example, the controller 250 may send the target voice interaction data to the server, so that the server performs feature extraction on the target voice interaction data, obtains voice attribute data corresponding to the interaction data, and feeds the voice attribute data back to the controller 250. The voice attribute data includes at least one sound source object and the voice features of the corresponding sound source object; the target voice interaction data is the voice interaction data collected within a preset time period.
[0083] For example, the controller 250 can obtain baseline model parameters corresponding to each speech feature from the server. The baseline model parameters are incremental model parameters corresponding to fixed model parameters in the target wakeup recognition model. During the training of the target wakeup recognition model, the server maintains the fixed model parameters in the target wakeup recognition model unchanged and adjusts the incremental model parameters.
[0084] In some embodiments, the controller 250 and the tuner 210 may be located in different separate devices, that is, the tuner 210 may also be located in an external device of the main device where the controller 250 is located, such as an external set-top box.
[0085] In some embodiments, the user may input a user command through a graphical user interface (GUI) displayed on the display 260 , and the user input interface receives the user input command through the graphical user interface (GUI).
[0086] In some embodiments, the audio output device 270 may be a local speaker of the display device 200, or an external audio output device connected to the display device 200. For the external audio output device connected to the display device 200, the display device 200 may further be provided with an external audio output terminal, through which the audio output device may be connected to the display device 200 to output the sound of the display device 200.
[0087] In some embodiments, the user input interface 280 may be configured to receive instructions from a user.
[0088] Figure 3 Some embodiments of this application provide Figure 1 The hardware configuration diagram of the control device in the figure. Figure 3 As shown, the control device 100 may include: a controller 110, a communication interface 130, a user input / output interface, a memory, and a power supply.
[0089] The control device 100 is configured to control the display device 200, and can receive the user's input operation instructions, and convert the operation instructions into instructions that the display device 200 can recognize and respond to, playing the role of an interactive intermediary between the user and the display device 200.
[0090] In some embodiments, the control device 100 may be a smart device. For example, the control device 100 may be installed with various applications for controlling the display device 200 according to user needs.
[0091] In some embodiments, as Figure 1 As shown, the mobile terminal 300 or other intelligent electronic devices can play a similar function as the control device 100 after installing the application for controlling the display device 200 .
[0092] The controller 110 includes a processor 112, RAM 113, ROM 114, a communication interface 130, and a communication bus. The controller 110 is used to control the operation and operation of the control device 100, as well as the communication and cooperation between internal components and external and internal data processing functions.
[0093] Under the control of the controller 110, the communication interface 130 communicates control signals and data signals with the display device 200. The communication interface 130 may include at least one of a WiFi chip 131, a Bluetooth module 132, an NFC module 133, or other near field communication modules.
[0094] The user input / output interface 140 includes at least one of a microphone 141 , a touch panel 142 , a sensor 143 , a button 144 and other input interfaces.
[0095] In some embodiments, the control device 100 includes at least one of a communication interface 130 and an input / output interface 140. The control device 100 is configured with the communication interface 130, such as a WiFi, Bluetooth, or NFC module, to encode user input commands via the WiFi protocol, Bluetooth protocol, or NFC protocol and transmit them to the display device 200.
[0096] The memory 190 is used to store various operating programs, data and applications for driving and controlling the control device 100 under the control of the controller. The memory 190 can store various control signal instructions input by the user.
[0097] The power supply 180 is used to provide operating power support for various components of the control device 100 under the control of the controller.
[0098] To facilitate user interaction, in some embodiments, the display device 200 may run an operating system. The operating system is a computer program for managing and controlling the hardware and software resources in the display device 200. The operating system may provide a user interface (control the display device), allow the user to interact with the display device 200, and support the running of various application programs.
[0099] It should be noted that the operating system can be a native operating system based on a specific operating platform, a third-party operating system deeply customized based on a specific operating platform, or an independent operating system specially developed for display devices.
[0100] The operating system can be divided into different modules or layers according to the functions implemented.
[0101] For example, Figure 4 As shown, in some embodiments, the system is divided into four layers, from top to bottom, namely the application layer (abbreviated as "application layer"), the application framework layer (abbreviated as "framework layer"), the system library layer and the kernel layer.
[0102] In some embodiments, the application layer provides services and interfaces for applications, enabling the display device 200 to run applications and interact with the user based on the applications. The application layer can host at least one application, which can include built-in window programs, system settings programs, clock programs, and the like, or applications developed by third-party developers. In specific implementations, the application packages in the application layer are not limited to the examples above.
[0103] The framework layer provides applications with an application programming interface (API) and programming framework. The application framework layer includes predefined functions. The application framework layer acts as a processing center, determining the actions taken by applications in the application layer. Through the API, applications can access system resources and services during execution.
[0104] like Figure 4As shown, in the embodiment of the present application, the application framework layer includes a view system, managers, content providers, etc., wherein the view system can design and implement the interface and interaction of the application, and the view system includes lists, grids, text boxes, buttons, etc. The manager includes at least one of the following modules: an activity manager for interacting with all activities running in the system; a location manager for providing system services or applications with access to the system location service; a package manager for retrieving various information related to the application packages currently installed on the device; a notification manager for controlling the display and clearing of notification messages; and a window manager for managing icons, windows, toolbars, wallpapers, and desktop widgets on the user interface.
[0105] In some embodiments, the activity manager is used to manage the lifecycle of each application and common navigation back functions, such as controlling application exit, opening, and back. The window manager is used to manage all window programs, such as obtaining the display screen size, determining whether there is a status bar, locking the screen, taking screenshots, and controlling changes in display windows, such as shrinking, shaking, or distorting the display window.
[0106] In some embodiments, the system runtime layer can provide support for the framework layer. When the framework layer is used, the operating system will run the instruction library contained in the system runtime layer, such as the C / C++ instruction library, to implement the functions to be implemented by the framework layer.
[0107] In some embodiments, the kernel layer is a functional layer between the hardware and software of the display device 200. The kernel layer can implement functions such as hardware abstraction, multitasking, and memory management. Figure 4 As shown, the kernel layer can be configured with hardware drivers, and the drivers included in the kernel layer can be at least one of the following drivers: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, USB driver, HDMI driver, sensor driver (such as fingerprint sensor, temperature sensor, pressure sensor, etc.), and power driver, etc.
[0108] like Figure 4 As shown, Figure 4 Some embodiments of this application provide Figure 1Schematic diagram of software configuration in the display device. In some embodiments, the system of the display device 200 can be divided into three layers, namely, application layer, middleware layer and hardware layer from top to bottom.
[0109] The application layer mainly includes commonly used applications on the TV and the application framework. Common applications are mainly browser-based applications, such as HTML5 applications (HyperText Markup Language 5 applications, applications based on web technologies such as the fifth edition of Hypertext Markup Language) and native applications.
[0110] The Application Framework is a complete program model that has all the basic functions required by standard application software, such as file access, data exchange, etc., as well as the user interfaces of these functions (toolbars, status bars, menus, dialog boxes).
[0111] Native apps can support online or offline, message push or local resource access.
[0112] The middleware layer includes various TV protocols, multimedia protocols, and system components. Middleware uses the basic services (functions) provided by system software to connect various parts of the application system or different applications on the network, enabling resource and function sharing.
[0113] The hardware layer primarily includes the Hardware Abstraction Layer (HAL) interface, hardware, and drivers. The HAL interface is a unified interface for all TV chipsets, with each chip implementing its own logic. Drivers primarily include audio drivers, display drivers, Bluetooth drivers, camera drivers, Wi-Fi drivers, USB drivers, HDMI drivers, sensor drivers (such as fingerprint sensors, temperature sensors, and pressure sensors), and power drivers.
[0114] It should be noted that the above example is only a simple division of the operating system functions and does not constitute a limitation on the specific operating system form of the display device 200 in the embodiment of the present application. Depending on factors such as the function of the display device and the type of operating system, the number of levels and specific level types contained in the operating system may be expressed in other forms.
[0115] With the development of display technology, voice interaction technology has gradually become a key means of controlling display devices. Among them, the display can be woken up by voice audio, improving the efficiency of human-computer interaction and user experience. However, traditional voice wake-up methods are usually based on a single preset voice template or simple voiceprint matching to wake up the display. In different application scenarios, false wake-ups or missed wake-ups are prone to problems, affecting the accuracy of wake-up recognition.
[0116] For example, the collected voice interaction data can be input into a trained voice wake-up recognition model to identify the voice interaction data and obtain a wake-up recognition result. However, due to differences in language habits and regional dialects in different regions, as well as differences in gender and age among users in different application scenarios, traditional voice wake-up recognition models cannot be accurately adapted to different application scenarios, resulting in misidentification and missed recognition, which affects recognition accuracy.
[0117] In some alternative embodiments, see Figure 5 , provides a display device wake-up method, applied to the controller 250 in the display device 200, including:
[0118] S510: Obtain a pre-trained target wakeup recognition model.
[0119] The target wakeup recognition model can be understood as a recognition model for recognizing voice interaction data and outputting a wakeup recognition result. When the wakeup recognition result is a triggered wakeup result, the display 260 is controlled to switch from displaying the standby interface to displaying the user interface, thereby realizing voice wakeup of the display 260. The wakeup recognition result includes an untriggered wakeup result and a triggered wakeup result.
[0120] In some embodiments, the target wakeup recognition model can be parameterized, thereby iteratively updating the target wakeup recognition model. Therefore, the pre-trained target wakeup recognition model can be understood as the latest version of the target wakeup recognition model.
[0121] It is worth noting that this embodiment does not impose any restrictions on the specific model structure, specific model type, and specific training method of the target wakeup recognition model. Exemplarily, the target wakeup recognition model can include at least one of a traditional machine learning model and a neural network model.
[0122] In some embodiments, a large amount of training sample data and the wake-up recognition results corresponding to each training sample data can be obtained. The training sample data includes sample voice interaction data. Each training sample data is used as the input of a pre-built target wake-up recognition model, and the corresponding wake-up recognition result is used as a training label to train the pre-built target wake-up recognition model, so that the target wake-up recognition model gradually has the ability to determine the wake-up recognition result until the target wake-up recognition model meets the model training termination condition. The model training termination condition can be that the number of training sample data reaches a preset number, the model tends to converge, or the model accuracy reaches a preset accuracy threshold, etc. This embodiment does not impose any restrictions on this.
[0123] For example, the sample voice interaction data may include voice data from different regional accents, different age groups, different genders, and various emotional expressions. The sources of the sample voice interaction data include publicly available voice datasets, actual interaction records with user objects, and specific voice collection activities. This application does not impose any restrictions on the specific voice data categories or specific voice sources of the sample voice interaction data.
[0124] For example, the sample voice interaction data can be preprocessed, including removing background noise, standardizing the audio format, adjusting the audio sampling rate, and other operations to improve data quality and ensure the accuracy of subsequent processing.
[0125] S520. When the model parameter update conditions are met, obtain voice attribute data corresponding to the target voice interaction data; the target voice interaction data is the voice interaction data collected by the audio input interface of the display device within a preset time period; the voice attribute data includes at least one sound source object and the voice features of the corresponding sound source object.
[0126] Exemplarily, the voice attribute data can be obtained by performing feature extraction on the target voice interaction data by the server.
[0127] The model parameter update condition may be understood as a condition for triggering the update of the model parameters of the target wakeup recognition model.
[0128] In some embodiments, the model parameter update condition may include at least one of the following: detecting that the display device is accessing the network for the first time; triggering a target timer task; the target timer task is used to indicate that the target wake-up recognition model is updated according to a preset period or preset frequency; receiving a model parameter update operation. The preset period and preset frequency can be set by technicians based on needs or experience, or determined through extensive experimentation, and this application does not impose any restrictions on this. For example, the preset period can be 30 days.
[0129] The above approach enables triggering the update of the target wakeup recognition model in different usage scenarios, which is conducive to improving the flexibility of the target wakeup recognition model update. It is understandable that in the process of updating the target wakeup recognition model, it is first necessary to obtain the voice attribute data of the voice interaction data within a preset time period to provide data reference for subsequent updates of the target wakeup recognition model.
[0130] The preset period may include a preset historical period before the model parameter update condition is met or a preset future period after the model parameter update condition is met. The preset period may be set by a technician based on needs or experience, or determined through a large number of experiments, and this application does not impose any restrictions on this.
[0131] For example, when the display device is connected to the network for the first time, the display device has not yet acquired the voice interaction data in the current application scenario. Therefore, in some embodiments, the preset time period can be a future time period, for example, acquiring voice attribute data corresponding to voice interaction data within 7 days after the first access to the network.
[0132] For example, in the case of a target timing task that triggers the timer, for example, triggering the timer once every 30 days, the display device can already obtain the voice interaction data in the current application scenario. Therefore, in some embodiments, the preset time period can be a historical time period, such as obtaining the voice attribute data corresponding to the voice interaction data within 7 days before the target timing task that triggers the timer. Of course, in other embodiments, the voice attribute data corresponding to the voice interaction data within 7 days after the target timing task that triggers the timer can also be obtained. It is worth noting that the present application does not impose any restrictions on the specific time length of the preset time period. For example, the time length of the preset time period can also be 10 days, 15 days or 30 days.
[0133] It can be understood that when it is detected that the display device is accessing the network for the first time, that is, for the first use scenario, the model parameters of the target wake-up recognition model are updated, so that the target wake-up recognition model can be adapted to the current application scenario. By setting the target timing task of the timer, the model parameters can be updated according to the preset period or preset frequency, avoiding the inability to adapt to changes in application scenarios due to long-term use of fixed model parameters. By responding to the model parameter update operation, the model parameters of the target wake-up recognition model are updated, so that the model parameters can be flexibly updated according to actual needs.
[0134] The voice interaction data can be understood as data used for voice interaction with the display device 200. For example, the voice interaction data can include at least one of a preset wake-up word or a preset wake-up command, such as "XXX, please wake up the TV." The target voice interaction data is the voice interaction data collected within a preset time period. The preset time period can be set by technicians based on needs or experience, or determined through extensive experiments, and this application does not impose any restrictions on this.
[0135] Speech attribute data can be understood as data extracted from speech interaction data and used to identify the sound source object in the speech interaction data and characterize the speech characteristics of the corresponding sound source object. A sound source object can be understood as an individual utterance in the speech interaction data. Speech characteristics can be understood as data used to describe the specific speech category contained in the sound source object.
[0136] In some embodiments, speech features may include basic attribute features and language and regional features. Basic attribute features can be understood as data describing the basic attribute classification inherent in a sound source object; language and regional features can be understood as data describing the language and regional classification of a sound source object. Optionally, speech features may also include at least one of emotional state features and timbre features.
[0137] Exemplarily, the basic attribute features may include at least one of gender features and age features. The language and regional features may include at least one of language type and dialect type. It is understood that because the speech features of the sound source object include basic attribute features, updating the target wakeup recognition model based on the speech features can improve the target wakeup recognition model's recognition response to dimensions such as age and gender. Because the speech features of the sound source object include language and regional features, updating the target wakeup recognition model based on the speech features can reduce the problem of missed wakeups caused by dialect accents.
[0138] The following is an example of a specific application scenario to illustrate the sound source objects and voice features. It is worth noting that this should not be understood as a limitation on the sound source objects and language features. For example, in a certain family application scenario, the sound source objects corresponding to the voice interaction data within a preset time period may include: sound source object A1, sound source object A2 and sound source object A3. The voice features corresponding to the sound source objects include: sound source object A1, "middle-aged male, Sichuan dialect"; sound source object A2, "middle-aged female, Cantonese Mandarin"; and sound source object A3, "teenage female, standard Mandarin".
[0139] In some embodiments, the controller 250 may perform feature extraction on the target voice interaction data to obtain voice attribute data.
[0140] In other embodiments, the target voice interaction data may be sent to a server so that the server performs feature extraction on the target voice interaction data and feeds back voice attribute data obtained by the feature extraction to the controller.
[0141] For example, as mentioned above, the preset time period may include a preset future time period after the model parameter update condition is satisfied. Accordingly, during the preset time period, the collected voice interaction data can be sent to the server one by one, so that the server can extract features from the voice interaction data and feedback the corresponding voice attribute data to the first terminal. This process continues until the current time reaches the end time of the preset future time period, thereby obtaining the voice attribute data for each voice interaction data, i.e., the voice attribute data for the target voice interaction data.
[0142] For example, as mentioned above, the preset period may include a preset historical period before the model parameter update condition is satisfied, and the target voice interaction data is the voice interaction data within the preset historical period. Accordingly, the target voice interaction data may be sent to a server, so that the server extracts features from the target voice interaction data and feeds back voice attribute data corresponding to the target voice interaction data to the first terminal.
[0143] In some embodiments, a voiceprint recognition model can be invoked to extract voiceprint data from the target voice interaction data and identify the sound source object corresponding to the voiceprint data. For example, the target voice interaction data can be sent to a server, so that the server can extract the voiceprint data from the target voice interaction data, identify the sound source object corresponding to the voiceprint data, and feed it back to the controller.
[0144] For example, the voiceprint data can be extracted from the target voice interaction data, and different sound source objects can be distinguished and identified based on voiceprint parameters such as fundamental frequency and formant in the voiceprint data. In the above method, the introduction of voiceprint data is conducive to improving the efficiency of sound source object identification.
[0145] In some embodiments, the speech features of the sound source object include basic attribute features and language and regional features; accordingly, a voiceprint recognition model can be called to extract the voiceprint data in the target voice interaction data, and identify the sound source object and the basic attribute features of the sound source object in the voiceprint data; and a language recognition model can be called to extract the regional characteristic vocabulary of the sound source object in the target voice interaction data, and identify the language and regional features corresponding to the regional characteristic vocabulary.
[0146] Exemplarily, the target voice interaction data can be sent to a server so that the server extracts the voiceprint data in the target voice interaction data, identifies the sound source object and the basic attribute characteristics of the sound source object in the voiceprint data, and extracts the regional characteristic vocabulary of the sound source object in the target voice interaction data, and identifies the language and regional characteristics corresponding to the regional characteristic vocabulary.
[0147] Exemplarily, the basic attribute features of the sound source object can be identified by extracting the voiceprint parameters from the voiceprint data. The voiceprint parameters may include parameters such as fundamental frequency and formant.
[0148] Exemplarily, the server can perform text conversion on the target voice interaction data to obtain the corresponding target text; extract the regional feature words of the sound source object from the target text, and determine the language regional features according to the regional feature words. Optionally, the target voice interaction data and the target text can be combined to extract the language regional features from the voice dimension and the text dimension respectively.
[0149] For the sake of easy understanding, an exemplary description is given for the above sound source objects A1 - A3. It should be noted that it should not be construed as a limitation on the specific extraction method of the language regional features.
[0150] Exemplarily, the sound source object A1 has "yaode" in Sichuan dialect, so it is determined that the sound source object A1 is in Sichuan dialect; the sound source object A2 has "bao tang" in Mandarin with a Cantonese accent, so it is determined that the sound source object A1 is in Cantonese-accented Mandarin; and the sound source object A3 is in Mandarin and does not have regional feature words, so it is determined that the sound source object A3 is in standard Mandarin.
[0151] S530. Obtain the target wake-up frequency corresponding to each sound source object; the target wake-up frequency corresponding to the sound source object is the frequency of the sound source object waking up the display of the display device within a preset period.
[0152] In some embodiments, in response to the target wake-up frequency input operation, the input user information and the frequency of the corresponding user object waking up the display within a preset period can be obtained, so as to match the user object with the sound source object and determine the frequency of the sound source object waking up the display within a preset period, that is, the target wake-up frequency.
[0153] In other embodiments, for each sound source object, the record items whose wake-up time of the sound source object is within the preset period can be obtained from the interaction log record; the number of each record item is obtained to get the target wake-up frequency of the sound source object; wherein, the interaction log record records the sound source object corresponding to the trigger wake-up result and the corresponding wake-up time. By introducing the interaction log record, and the interaction log record records the sound source object corresponding to the trigger wake-up result and the corresponding wake-up time, it is convenient to obtain the target wake-up frequency of different sound source objects within a preset period.
[0154] Optionally, the voice interaction data can be input into the target wake-up recognition model to obtain a wake-up recognition result; the wake-up recognition result includes an untriggered wake-up result and a triggered wake-up result; when the wake-up recognition result is a triggered wake-up result, the sound source object of the voice interaction data and the corresponding wake-up time are recorded in the interaction log record.
[0155] S540: Adjust the model parameters in the target awakening recognition model according to the target awakening frequency of each sound source object and the speech characteristics of the corresponding sound source object to update the target awakening recognition model.
[0156] In some embodiments, the feature weights of the speech features of the corresponding sound source object can be determined based on the target wake-up frequency of the sound source object; based on the feature weights of different speech features, the target wake-up recognition model is focused on to enhance the sensitivity of the target wake-up recognition model to different speech features, thereby obtaining an updated target wake-up recognition model.
[0157] In other embodiments, the feature weight of the speech feature of the corresponding sound source object can be determined based on the target wakeup frequency of the sound source object; for each speech feature of each sound source object, the corresponding baseline model parameters of the speech feature can be obtained; the baseline model parameters of the speech feature can be weighted according to the feature weight of the speech feature to obtain the target model parameters of the speech feature; and the model parameters in the target wakeup recognition model can be adjusted based on the target model parameters of the different speech features of each sound source object. The baseline model parameters are obtained by training the target wakeup recognition model to be trained using sample speech interaction data.
[0158] S550: Obtain current voice interaction data collected by the audio input interface at the current moment.
[0159] The current voice interaction data can be understood as the voice interaction data collected at the current moment. The current moment can be understood as the moment when the voice interaction data is acquired after the target wakeup recognition model is updated.
[0160] S560: Input the current voice interaction data into the updated target wakeup recognition model to obtain a wakeup recognition result corresponding to the current voice interaction data; the wakeup recognition result includes an untriggered wakeup result and a triggered wakeup result.
[0161] It is understandable that the target wakeup recognition model has been updated in the foregoing. Therefore, inputting the current voice interaction data into the target wakeup recognition model can improve the accuracy of the wakeup recognition result.
[0162] S570: When the wake-up recognition result is a triggered wake-up result, control the display to switch from displaying the standby interface to displaying the user interface.
[0163] In some embodiments, the method further includes: if the model parameter update condition is not met, using a pre-trained target wakeup recognition model to recognize the wakeup recognition result of the voice interaction data. Optionally, if the model parameter update condition is met and before the target wakeup recognition model is updated, the pre-trained target wakeup recognition model can be used to recognize the wakeup recognition result of the voice interaction data.
[0164] It can be understood that the above steps implement the adjustment of the model parameters of the target awakening recognition model, thereby iteratively updating the target awakening recognition model. The pre-trained target awakening recognition model can be understood as the current version of the target awakening recognition model.
[0165] For ease of understanding, the following is an exemplary description of the use process of the target wakeup recognition model. For example, when it is detected that the display device is accessing the network for the first time, the initial version of the target wakeup recognition model M1 is obtained; the initial version of the target wakeup recognition model M1 is used to recognize the wakeup recognition result of the voice interaction data; and based on the above method, the target wakeup recognition model M1 is updated to obtain the target wakeup recognition model M2; the target wakeup recognition model M2 is used to recognize the wakeup recognition result of the voice interaction data; and so on. For example, when the model parameter update conditions are met, the target wakeup recognition model M2 is continued to be updated to obtain the target wakeup recognition model M3.
[0166] In the above embodiment, obtaining a pre-trained target wakeup recognition model provides a foundational framework for subsequent targeted updates to the target wakeup recognition model. By obtaining the speech features of each sound source object corresponding to the target voice interaction data within a preset time period, and obtaining the target wakeup frequency corresponding to the sound source object, provided the model parameters are adapted to the current application scenario, providing the target wakeup recognition model with basic data for updating. By adjusting the model parameters of the target wakeup recognition model based on the target wakeup frequency and speech features of the sound source object, the target wakeup recognition model can be made more suitable for the current application scenario, improving its recognition accuracy in that scenario. Furthermore, based on the target wakeup frequency, the model can prioritize high-frequency speech features, thereby improving the recognition sensitivity of active sound source objects. Current voice interaction data collected by the audio input interface at the current moment is obtained; the current voice interaction data is input into the target wakeup recognition model to obtain a wakeup recognition result corresponding to the current voice interaction data; and if the wakeup recognition result is a triggered wakeup result, the display is controlled to switch from displaying the standby interface to displaying the user interface. This ensures accurate recognition of the current voice interaction data, thereby improving the accuracy of voice wakeup.
[0167] On the basis of the above embodiment, the above display device wake-up method may also be implemented through the interaction between the controller 250 and the server 400.
[0168] refer to Figure 6 The following diagram shows the interactive timing diagram of the wake-up method of the display device, which includes the following steps:
[0169] S610: The controller determines whether a model parameter update condition is satisfied.
[0170] S620: Send the voice interaction data collected by the audio input interface of the display device to the server within a preset period of time.
[0171] S630: The server identifies voice attribute data corresponding to the voice interaction data, where the voice attribute data includes at least one sound source object and voice features of the corresponding sound source object.
[0172] S640: The server sends voice attribute data to the controller.
[0173] S650. The controller obtains a target wake-up frequency corresponding to each sound source object; the target wake-up frequency corresponding to the sound source object is the frequency at which the sound source object wakes up the display of the display device within a preset time period.
[0174] S660: Adjust the model parameters in the target awakening recognition model according to the target awakening frequency of each sound source object and the speech characteristics of the corresponding sound source object to update the target awakening recognition model.
[0175] S670: Obtain current voice interaction data collected by the audio input interface at the current moment.
[0176] S680: Input the current voice interaction data into the updated target wakeup recognition model to obtain a wakeup recognition result corresponding to the current voice interaction data; the wakeup recognition result includes an untriggered wakeup result and a triggered wakeup result.
[0177] S690: When the wake-up recognition result is a triggered wake-up result, control the display to switch from displaying the standby interface to displaying the user interface.
[0178] Based on the technical solutions of the above embodiments, the present application also provides an optional embodiment, in which the steps for adjusting the model parameters are refined.
[0179] See also Figure 7A The steps for adjusting the model parameters shown include:
[0180] S710. For each sound source object, determine the feature weight of each speech feature of the sound source object based on the target proportion of the target wake-up frequency of the sound source object in the total target wake-up frequency; the total target wake-up frequency is the sum of the target wake-up frequencies of each sound source object.
[0181] Among them, feature weight can be understood as an indicator used to quantify the contribution of different audio features to the model decision-making process.
[0182] For ease of understanding, the following is an exemplary description of the process for determining feature weights. It should be noted that this should not be construed as limiting the specific determination process.
[0183] For example, if the target wakeup frequencies for sound source objects A1-A3 are 8, 10, and 5, respectively, the total wakeup frequency is 23. Accordingly, the target proportion for sound source object A1 is 35%, the target proportion for sound source object A2 is 45%, and the target proportion for sound source object A3 is 20%.
[0184] Continuing with the previous example, for example, the speech features of source object A1 include "middle-aged male" and "Sichuan dialect"; the speech features of source object A2 include "middle-aged female" and "Cantonese Mandarin"; and the speech features of source object A3 include "teenage female" and "standard Mandarin." Thus, the feature weights for each speech feature can be obtained, as shown in Table 1.
[0185] Table-1 Feature weights of speech features
[0186] Serial number Voice Features Feature weights 1 middle-aged men 17.5% 2 middle-aged women 22.5% 3 teenage girls 10% 4 Sichuan dialect 17.5% 5 Cantonese Mandarin 22.5% 6 Standard Mandarin 10%
[0187] S720. For each speech feature of each sound source object, obtain a reference model parameter corresponding to the speech feature; the reference model parameter is obtained by training the target wakeup recognition model to be trained using the sample speech interaction data.
[0188] Among them, the baseline model parameters corresponding to the voice features can be understood as model parameters obtained by training the target wakeup recognition model to be trained using the sample voice interaction data corresponding to the voice features.
[0189] Optionally, for each speech feature, the target wakeup recognition model to be trained is trained using sample speech interaction data corresponding to the speech feature to obtain benchmark model parameters.
[0190] For example, the sample voice interaction data may include voice data from different regional accents, different age groups, different genders, and various emotional expressions. The sources of the sample voice interaction data include publicly available voice datasets, actual interaction records with user objects, and specific voice collection activities. This application does not impose any restrictions on the specific voice data categories or specific voice sources of the sample voice interaction data.
[0191] For example, the sample voice interaction data can be preprocessed, including removing background noise, standardizing the audio format, adjusting the audio sampling rate, and other operations to improve data quality and ensure the accuracy of subsequent processing.
[0192] Optionally, the pre-processed sample voice interaction data can be transcribed to obtain the corresponding sample text content. Exemplarily, the sample voice interaction data can be transcribed using a content recognition model. The content recognition model can be a traditional machine learning model or a neural network model to transcribe the sample voice interaction data. This embodiment does not impose any restrictions on the specific model type of the content recognition model.
[0193] Optionally, the sample voice interaction data can be matched with the corresponding sample text content and input into a multimodal large model to obtain feature labels for the sample voice interaction data. Feature labels can include at least one of basic attribute feature labels, language and region feature labels, and emotional state labels.
[0194] Among them, the multimodal large model can be understood as a neural network model that can simultaneously process and understand data in different modalities (such as voice and text).
[0195] For example, basic attribute feature tags may include age feature tags and gender feature tags. Age feature tags may include middle-aged, elderly, and adolescent tags, among others; gender feature tags may include male and female tags. For example, language and region feature tags may include Northern Mandarin, Sichuan dialect, and Cantonese Mandarin. For example, emotional state tags may include happy, sad, and angry tags. For example, timbre tags may include low-pitched male and high-pitched female tags.
[0196] Exemplarily, for each feature label, the target wakeup recognition model to be trained can be trained using sample voice interaction data corresponding to the feature label to obtain benchmark model parameters.
[0197] In some embodiments, the baseline model parameters are incremental model parameters corresponding to fixed model parameters in the target wakeup recognition model. During training of the target wakeup recognition model, the fixed model parameters in the target wakeup recognition model are maintained unchanged, while the incremental model parameters are adjusted. For example, a low-rank adaptive training approach can be employed to train the target wakeup recognition model to obtain the baseline model parameters.
[0198] The incremental model parameters, also known as new low-rank parameters, are the parameters added to each existing model parameter in the target wakeup recognition model. During low-rank adaptive training, by maintaining the existing model parameters (i.e., fixed model parameters) in the target wakeup recognition model unchanged and adjusting the incremental model parameters, we can achieve effective adaptation to the corresponding sample voice interaction data with fewer incremental model parameters.
[0199] refer to Figure 7B The following is a schematic diagram of the target wakeup recognition model. The target wakeup recognition model includes an input layer, a hidden layer, and an output layer. t represents the input of the target arousal recognition model; y t represents the output of the target arousal recognition model; t represents the index.
[0200] In the middle layer, the fixed model parameters can be expressed as The newly added low-rank parameter can be expressed as Wherein, / represents the number of network layers. In other embodiments, incremental model parameters may be added corresponding to the fixed model parameters of the output layer. This embodiment does not impose any restrictions on the specific model structure of the target wakeup recognition model or the specific location for adding incremental model parameters.
[0201] In an optional embodiment, the incremental model parameters can be matched with corresponding feature tags and packaged as an independent plug-in. Accordingly, the independent plug-in matching the speech feature can be obtained from the server to obtain the baseline model parameters corresponding to the speech feature.
[0202] S730 . Weight the baseline model parameters of the speech features according to the feature weights of the speech features to obtain target model parameters of the speech features.
[0203] The target model parameters can be understood as model parameters obtained by weighting the baseline model parameters of the speech features.
[0204] Exemplarily, the feature weights may be multiplied by the reference model parameters to obtain the target model parameters of the speech features.
[0205] S740: Adjust the model parameters in the target wakeup recognition model according to the target model parameters of the different speech features of each sound source object.
[0206] In some embodiments, for each fixed model parameter, the target model parameters corresponding to the fixed model parameter can be fused to obtain a fused model parameter; the fused model parameter is fused with the fixed model parameter to adjust the model parameters in the target wakeup recognition model.
[0207] Exemplarily, for each fixed model parameter, the target model parameters corresponding to the fixed model parameter may be added together to obtain the fusion model parameter corresponding to the fixed model parameter.
[0208] Exemplarily, for each fixed model parameter, the fusion model parameter may be added to the fixed model parameter to obtain an adjusted fixed model parameter; after each fixed model parameter is adjusted, the target wakeup recognition model is updated.
[0209] In other embodiments, the fusion model parameters may also be added to the target wakeup recognition model to update the target wakeup recognition model. Figure 7C The updated target wakeup recognition model includes fixed model parameters, namely pretrained weights; plugins 1 to n represent different baseline model parameters; W1 to W n , respectively represent the feature weights corresponding to plugin 1 to plugin n.
[0210] In the above steps, the feature weight of each speech feature of the sound source object is determined based on the target proportion of the target wakeup frequency of the sound source object in the total target wakeup frequency. This quantifies the difference in wakeup frequency corresponding to each speech feature, forms a feature importance gradient, and provides a quantitative basis for differential parameter optimization. For each speech feature of each sound source object, the corresponding baseline model parameters are obtained, providing a corresponding benchmark reference for subsequent weighted calculations. Because the baseline model parameters are obtained by training the target wakeup recognition model to be trained using sample speech interaction data with speech features, the matching between the speech features and their corresponding baseline model parameters is improved. By weighting the baseline model parameters of the speech features according to the feature weights of the speech features to obtain the target model parameters of the speech features, and adjusting the model parameters in the target wakeup recognition model based on the target model parameters of the different speech features of each sound source object, the more frequently occurring speech features produce more significant gradient update directions in the parameter space. This allows the updated target wakeup recognition model to adapt to the current application scenario, thereby improving the accuracy of speech wakeup.
[0211] Based on the above embodiment, a method for waking up a display device is described in detail.
[0212] refer to Figure 8 Shown are methods for waking up a display device in some other embodiments, including the following steps:
[0213] S801: The controller determines that the network is accessed for the first time and executes S802.
[0214] S802: The controller sends the voice interaction data collected by the audio input interface to the server within a preset period of time; and
[0215] S803. The controller inputs the voice interaction data into a pre-trained target recognition model to obtain a wake-up recognition result; when the wake-up recognition result is a trigger wake-up result, the controller controls the display to switch from displaying the standby interface to displaying the user interface.
[0216] S804: The server extracts the sound source object and the voice features of the sound source object from the voice interaction data.
[0217] S805: The server sends the sound source object and the voice features of the sound source object to the controller;
[0218] S806: The controller records the sound source object and wake-up time corresponding to the triggering wake-up result in the interaction log record.
[0219] S807. The controller obtains the target wake-up frequency corresponding to each sound source object from the interaction log record; the target wake-up frequency corresponding to the sound source object is the frequency of the sound source object waking up the display within a preset time period.
[0220] S808. The controller determines, for each sound source object, a feature weight of each speech feature of the sound source object based on the target proportion of the target awakening frequency of the sound source object in the total target awakening frequency; the total target awakening frequency is the sum of the target awakening frequencies of each sound source object;
[0221] S809: The controller sends a baseline model parameter acquisition request to the server, where the baseline model parameter acquisition request includes various speech features.
[0222] S810: In response to the request to obtain the reference model parameters, the server feeds back the reference model parameters corresponding to each speech feature to the controller.
[0223] The benchmark model parameters are obtained by training the target wakeup recognition model to be trained using sample voice interaction data; the sample voice interaction data corresponds to voice features;
[0224] S811. The controller weights the reference model parameters of the speech features according to the feature weights of the speech features to obtain target model parameters of the speech features.
[0225] S812: The controller adjusts the model parameters in the target wakeup recognition model according to the target model parameters of the different speech features of each sound source object.
[0226] S813. The controller obtains current voice interaction data collected by the audio input interface at the current moment.
[0227] S814. The controller inputs the current voice interaction data into the updated target wake-up recognition model to obtain a wake-up recognition result corresponding to the current voice interaction data; the wake-up recognition result includes an untriggered wake-up result and a triggered wake-up result.
[0228] S815: When the wake-up identification result is a triggered wake-up result, the controller controls the display to switch from displaying the standby interface to displaying the user interface.
[0229] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0230] Based on the same inventive concept, embodiments of the present application also provide a display device wake-up device for implementing the display device wake-up method described above. The solution provided by this device is similar to the solution described in the above method. Therefore, the specific limitations of the embodiments of the display device wake-up device provided below can be found in the above-mentioned limitations of the display device wake-up method, and will not be repeated here.
[0231] In an exemplary embodiment, Figure 9 As shown, a display device wake-up device is provided, comprising: a first acquisition module 910, a second acquisition module 920, a third acquisition module 930, an update module 940, a fourth acquisition module 950, an input module 960 and a control module 970.
[0232] A first acquisition module 910 is used to acquire a pre-trained target wakeup recognition model;
[0233] The second acquisition module 920 acquires speech attribute data corresponding to target speech interaction data when the model parameter update condition is met; the target speech interaction data is speech interaction data collected by the display device within a preset time period; the speech attribute data includes at least one sound source object and speech features of the corresponding sound source object;
[0234] The third acquisition module 930 acquires the target wake-up frequency corresponding to each sound source object; the target wake-up frequency corresponding to the sound source object is the frequency of the sound source object waking up the display within a preset time period;
[0235] An updating module 940 adjusts model parameters in the target awakening recognition model according to the target awakening frequency of each sound source object and the speech characteristics of the corresponding sound source object to update the target awakening recognition model;
[0236] The fourth acquisition module 950 acquires the current voice interaction data collected by the audio input interface at the current moment;
[0237] Input module 960 inputs the current voice interaction data into the updated target wakeup recognition model to obtain the wakeup recognition result corresponding to the current voice interaction data; the wakeup recognition result includes the untriggered wakeup result and the triggered wakeup result;
[0238] The control module 970 controls the display device to switch from displaying the standby interface to displaying the user interface when the wake-up recognition result is a triggered wake-up result.
[0239] In one embodiment, the update module 940 includes: a first determination unit, which is used to determine, for each sound source object, the feature weight of each speech feature of the sound source object according to the target proportion of the target wake-up frequency of the sound source object in the total target wake-up frequency; the total target wake-up frequency is the sum of the target wake-up frequencies of each sound source object; a first acquisition unit, which is used to obtain, for each speech feature of each sound source object, the baseline model parameters corresponding to the speech feature; the baseline model parameters are obtained by training the target wake-up recognition model to be trained using sample speech interaction data; a processing unit, which is used to weight the baseline model parameters of the speech feature according to the feature weight of the speech feature to obtain the target model parameters of the speech feature; an adjustment unit, which is used to adjust the model parameters in the target wake-up recognition model according to the target model parameters of different speech features of each sound source object.
[0240] In one embodiment, the baseline model parameters are incremental model parameters corresponding to the fixed model parameters in the target wakeup recognition model; during the training of the target wakeup recognition model, the fixed model parameters in the target wakeup recognition model are kept unchanged, and the incremental model parameters are adjusted; accordingly, the adjustment unit includes: a first fusion subunit, which is used to perform parameter fusion on each target model parameter corresponding to the fixed model parameter for each fixed model parameter to obtain a fused model parameter; and a second fusion subunit, which is used to perform parameter fusion on the fused model parameter and the fixed model parameter to achieve adjustment of the model parameters in the target wakeup recognition model.
[0241] In one embodiment, the model parameter update condition includes at least one of the following: detecting that the display device is accessing the network for the first time; triggering a target timing task of the timer; the target timing task is used to indicate that the model parameters are updated according to a preset period or preset frequency; and responding to a model parameter update operation.
[0242] In one embodiment, it also includes: a fifth acquisition module, which is used to obtain, for each sound source object, record items whose wake-up time of the sound source object is within a preset time period from the interaction log record; obtain the number of each record item to obtain the target wake-up frequency; wherein the interaction log record records the sound source object and the corresponding wake-up time corresponding to the triggering wake-up result.
[0243] In one embodiment, the second acquisition module 920 includes:
[0244] The first calling unit is used to call the voiceprint recognition model to extract voiceprint data from the target voice interaction data and identify the sound source object corresponding to the voiceprint data.
[0245] In one embodiment, the speech features of the sound source object include basic attribute features and language and regional features; the second acquisition module 920 includes:
[0246] The second calling unit is configured to call the voiceprint recognition model to extract the voiceprint data from the target voice interaction data and identify the sound source object and basic attribute features of the sound source object in the voiceprint data;
[0247] The third calling unit is used to call the language recognition model to extract the regional characteristic vocabulary of the sound source object in the target voice interaction data and identify the language regional characteristics corresponding to the regional characteristic vocabulary.
[0248] Each module in the aforementioned display device wake-up device may be implemented in whole or in part through software, hardware, or a combination thereof. Each module may be embedded in or independent of a processor in a computer device in hardware form, or may be stored in a computer device memory in software form, so that the processor can call and execute the corresponding operations of each module.
[0249] In an exemplary embodiment, a computer device is provided. The computer device may be a terminal, and its internal structure diagram may be as shown in FIG. Figure 10 As shown. The computer device includes a processor, a memory, an input / output interface, a communication interface, a display unit and an input device. The processor, the memory and the input / output interface are connected via a system bus, and the communication interface, the display unit and the input device are connected to the system bus via the input / output interface. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, near field communication (NFC) or other technologies. When the computer program is executed by the processor, a method for waking up a display device is implemented. The display unit of the computer device is used to form a visually visible picture, which can be a display screen, a projection device or a virtual reality imaging device. The display screen can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covering the display screen, or a button, trackball or touchpad set on the computer device casing, or an external keyboard, touchpad or mouse.
[0250] Those skilled in the art will understand that Figure 10 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0251] In one embodiment, a computer device is further provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.
[0252] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0253] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0254] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile memory and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processor involved in the various embodiments provided herein may be, but are not limited to, a general-purpose processor, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a programmable logic unit (PLC), a data processing logic unit based on quantum computing, an artificial intelligence (AI) processor, and the like.
[0255] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0256] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present application. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.
Claims
1. A display device, characterized in that: include: a display configured to display a user interface or a standby interface; An audio input interface configured to collect voice interaction data; At least one controller, connected to the audio input interface and the display, configured to: Obtain a pre-trained target arousal recognition model; When the model parameter update conditions are met, obtaining voice attribute data corresponding to the target voice interaction data; The target voice interaction data is the voice interaction data collected by the audio input interface within a preset period of time; the voice attribute data includes at least one sound source object and the voice features of the corresponding sound source object; Obtaining a target wake-up frequency corresponding to each of the sound source objects; the target wake-up frequency corresponding to the sound source object is the frequency at which the sound source object wakes up the display within the preset time period; adjusting the model parameters in the target awakening recognition model according to the target awakening frequency of each sound source object and the speech characteristics of the corresponding sound source object to update the target awakening recognition model; Obtaining current voice interaction data collected by the audio input interface at the current moment; Inputting the current voice interaction data into the updated target wakeup recognition model to obtain a wakeup recognition result corresponding to the current voice interaction data; the wakeup recognition result includes an untriggered wakeup result and a triggered wakeup result; When the wake-up recognition result is a triggered wake-up result, the display is controlled to switch from displaying the standby interface to displaying the user interface.
2. The display device according to claim 1, wherein When the controller adjusts the model parameters in the target wakeup recognition model according to the target wakeup frequency of each sound source object and the voice feature of the corresponding sound source object, the controller is configured to: For each sound source object, determining a feature weight of each speech feature of the sound source object according to a target proportion of the target awakening frequency of the sound source object in the total target awakening frequency; The total target awakening frequency is the sum of the target awakening frequencies of all the sound source objects; For each speech feature of each sound source object, obtaining a reference model parameter corresponding to the speech feature; the reference model parameter is obtained by training the target wakeup recognition model to be trained using the sample speech interaction data; weighting the reference model parameters of the speech features according to the feature weights of the speech features to obtain target model parameters of the speech features; The model parameters in the target wakeup recognition model are adjusted according to the target model parameters of the different speech features of each of the sound source objects.
3. The display device according to claim 2, wherein The baseline model parameters are incremental model parameters corresponding to the fixed model parameters in the target arousal recognition model; during the training of the target arousal recognition model, the fixed model parameters in the target arousal recognition model are kept unchanged, and the incremental model parameters are adjusted; Accordingly, when the controller adjusts the model parameters in the target wakeup recognition model according to the target model parameters of the different voice features of each of the sound source objects, it is configured to: For each fixed model parameter, performing parameter fusion on each target model parameter corresponding to the fixed model parameter to obtain a fused model parameter; The fusion model parameters are fused with the fixed model parameters to adjust the model parameters in the target wakeup recognition model.
4. The display device according to any one of claims 1 to 3, characterized in that: The model parameter update condition includes at least one of the following: detecting that the display device is accessing the network for the first time; A target timing task that triggers a timer; the target timing task is used to instruct to update the target wakeup recognition model according to a preset period or preset frequency; Received a model parameter update operation.
5. The display device according to any one of claims 1 to 3, characterized in that: The controller is further configured to: When the model parameter update condition is not met, the pre-trained target wake-up recognition model is used to recognize the wake-up recognition result of the voice interaction data.
6. The display device according to any one of claims 1 to 3, characterized in that: When executing the step of obtaining the target wake-up frequency corresponding to each of the sound source objects, the controller is configured to: For each sound source object, obtaining a record from the interaction log record in which the wake-up time of the sound source object is within the preset time period; wherein the interaction log record records the sound source object corresponding to the triggering wake-up result and the corresponding wake-up time; Obtain the number of each record item to obtain the target wake-up frequency of the sound source object.
7. The display device according to any one of claims 1 to 3, characterized in that: The acquiring of voice attribute data corresponding to the target voice interaction data includes: A voiceprint recognition model is called to extract voiceprint data from the target voice interaction data and identify a sound source object corresponding to the voiceprint data.
8. The display device according to any one of claims 1 to 3, characterized in that: The speech features of the sound source object include basic attribute features and language and regional features; The acquiring of voice attribute data corresponding to the target voice interaction data includes: Invoking a voiceprint recognition model to extract voiceprint data from the target voice interaction data, and identifying a sound source object in the voiceprint data and basic attribute features of the sound source object; A language recognition model is called to extract regional characteristic words of the sound source object in the target voice interaction data, and identify language regional features corresponding to the regional characteristic words.
9. A method for waking up a display device, characterized in that: include: Obtain a pre-trained target arousal recognition model; When the model parameter update conditions are met, obtaining voice attribute data corresponding to the target voice interaction data; The target voice interaction data is the voice interaction data collected by the audio input interface of the display device within a preset time period; the voice attribute data includes at least one sound source object and the voice features of the corresponding sound source object; Obtaining a target wake-up frequency corresponding to each of the sound source objects; the target wake-up frequency corresponding to the sound source object is the frequency at which the sound source object wakes up the display of the display device within the preset time period; adjusting the model parameters in the target awakening recognition model according to the target awakening frequency of each sound source object and the speech characteristics of the corresponding sound source object to update the target awakening recognition model; Obtaining current voice interaction data collected by the audio input interface at the current moment; Inputting the current voice interaction data into the updated target wakeup recognition model to obtain a wakeup recognition result corresponding to the current voice interaction data; the wakeup recognition result includes an untriggered wakeup result and a triggered wakeup result; When the wake-up recognition result is a trigger wake-up result, the display is controlled to switch from displaying a standby interface to displaying a user interface.
10. A device for waking up a display device, characterized in that: include: A first acquisition module is used to acquire a pre-trained target arousal recognition model; A second acquisition module, which acquires speech attribute data corresponding to the target speech interaction data when the model parameter update condition is satisfied; The target voice interaction data is voice interaction data collected by the audio input interface of the display device within a preset period of time; the voice attribute data includes at least one sound source object and the voice features of the corresponding sound source object; A third acquisition module acquires a target wake-up frequency corresponding to each of the sound source objects; the target wake-up frequency corresponding to the sound source object is the frequency at which the sound source object wakes up the display of the display device within the preset time period; an updating module, adjusting model parameters in the target awakening recognition model according to the target awakening frequency of each sound source object and the speech characteristics of the corresponding sound source object, so as to update the target awakening recognition model; A fourth acquisition module acquires current voice interaction data collected by the audio input interface at the current moment; An input module inputs the current voice interaction data into the updated target wakeup recognition model to obtain a wakeup recognition result corresponding to the current voice interaction data; the wakeup recognition result includes an untriggered wakeup result and a triggered wakeup result; The control module controls the display to switch from displaying a standby interface to displaying a user interface when the wake-up recognition result is a trigger wake-up result.