Model training method, voice interaction method and server

By training a multimodal language model that combines speech and image information, the problem of low recognition accuracy in voice interaction of display devices was solved, achieving higher speech recognition and interaction accuracy and improving user experience.

CN121122248AInactive Publication Date: 2025-12-12HISENSE ELECTRONIC TECH (WUHAN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511273225.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-08
Publication Date
2025-12-12
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing display devices suffer from low voice recognition accuracy in voice interaction, which leads to the inability to accurately execute user operations and reduces user experience.

Method used

By training a multimodal language model, combining user voice and page image information from the display device for recognition, training text is generated and training data is constructed. The multimodal language model is then adjusted to improve recognition accuracy.

Benefits of technology

It improves the accuracy of speech recognition and voice interaction, thus enhancing the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121122248A_ABST
    Figure CN121122248A_ABST
Patent Text Reader

Abstract

The invention provides a model training method, a voice interaction method and a server. The model training method comprises the following steps: acquiring a training text for training a multi-modal language model; converting the training text into a corresponding training audio and a control instruction, wherein the control instruction is used for controlling a display device to display a first page corresponding to the training text; triggering a sound playing device to play a training audio, and transmitting the control instruction to display equipment; constructing training data based on the voice when the sound playing device plays the training audio and a first page image when the display device displays the first page according to the control instruction; and training the multi-modal language model based on the training data. The trained multi-modal language model can recognize the user voice in combination with the user voice and the page image of the display device, so that a more accurate recognition result can be obtained, the accuracy of voice interaction is correspondingly improved, and the user experience is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of multimedia playback technology, and in particular to a model training method, a voice interaction method, and a server. Background Technology

[0002] A display device is a terminal device capable of outputting a display image, typically including smart TVs, mobile terminals, smart advertising screens, and projectors. To provide a convenient user experience, some display devices support voice interaction.

[0003] If a display device supports voice interaction, upon receiving a user's voice, the device can recognize the voice using a voice recognition model and then execute the operation corresponding to the recognition result, thus achieving voice interaction with the user. For example, if the voice recognition result instructs the display device to play a certain media file, the device can play the media file without the user needing to operate a remote control, thereby simplifying user operation and improving the user experience.

[0004] However, the same speech can often express different meanings. Therefore, the recognition results obtained by the display device through the speech recognition model are sometimes not the recognition results that the user expects. This can lead to the operation performed by the display device failing to meet the user's needs, resulting in low accuracy of voice interaction. Summary of the Invention

[0005] To address the aforementioned issues, embodiments of this application provide a model training method, a voice interaction method, and a server to improve the accuracy of voice interaction.

[0006] A first aspect of this application provides a model training method applied to a first server, the model training method comprising:

[0007] Obtain training text for training a multimodal language model;

[0008] The training text is converted into corresponding training audio and control instructions, and the control instructions are used to control the display device to display the first page corresponding to the training text.

[0009] The sound playback device is triggered to play the training audio, and the control command is transmitted to the display device;

[0010] Training data is constructed based on the speech when the sound playback device plays the training audio and the first page image when the display device displays the first page according to the control command;

[0011] The multimodal language model is trained based on the training data.

[0012] This approach enables the training of multimodal language models. The trained multimodal language model can then combine user speech with the page images displayed on the screen to jointly recognize user speech, thereby improving the accuracy of speech recognition and consequently enhancing the accuracy of voice interaction.

[0013] In another aspect of the embodiments of this application, obtaining training text for training a multimodal language model includes:

[0014] Based on the first operation that the display device supports can be performed via voice interaction, determine the sentence structure that conforms to the grammatical structure corresponding to the first operation;

[0015] The training text is generated based on the grammatically correct sentence structure corresponding to the first operation.

[0016] Through this scheme, the first server can generate training text based on the first operation, thereby improving the efficiency of obtaining training text and correspondingly improving the efficiency of training multimodal language models.

[0017] In another aspect of this application embodiment, generating the training text based on the grammatically correct sentence structure corresponding to the first operation includes:

[0018] When the first operation is to instruct the display device to play media assets, the media assets that the display device can play are determined based on the account logged into the media asset playback platform on the display device and / or the media assets stored on the display device;

[0019] The training text is generated based on the media assets playable by the display device and the grammatically correct sentence structure corresponding to the first operation.

[0020] With this approach, the first server can generate the corresponding training text when the first operation is to instruct the display device to play media assets, thus eliminating the need for manual input of the training text, simplifying manual operations, and improving the efficiency of obtaining training text.

[0021] In another aspect of this application embodiment, when the training text is used to instruct the display device to play the first media asset, the first page includes media asset information of the first media asset;

[0022] In cases where the training text is used to instruct the display device to jump between pages displaying different images, the first page includes at least one image;

[0023] When the training text is used to instruct the display device to jump between pages displaying different chapters, the first page includes text content.

[0024] In another aspect of the embodiments of this application, the media asset information of the first media asset includes at least one of the following: the name of the first media asset, the introductory information of the first media asset, the poster of the first media asset, at least one frame of the first media asset, and the information of the cast and crew of the first media asset, wherein the information of the cast and crew of the first media asset includes the names and / or photos of the cast and crew.

[0025] In another aspect of the embodiments of this application, training the multimodal language model based on the training data includes:

[0026] The training data is input into the multimodal language model;

[0027] Input training prompts into the multimodal language model. When the training text is used to instruct the display device to play media assets, the training prompts are used to prompt the multimodal language model to pay attention to the display device's browsing history and / or favorites.

[0028] Obtain the output of the multimodal language model after receiving the training data and the training prompts;

[0029] Based on the output results, the multimodal language model is adjusted.

[0030] A second aspect of this application provides a voice interaction method applied to a second server, the second server being configured with a multimodal language model trained using the model training method described in the first aspect, the voice interaction method comprising:

[0031] Receive target voice used to instruct the display device to perform voice interaction;

[0032] The request information is transmitted to the display device, the request information being used to request the acquisition of the second page image of the second page displayed by the display device;

[0033] The target speech and the second page image are recognized by the multimodal language model to obtain the target recognition result output by the multimodal language model;

[0034] The interactive instruction corresponding to the target recognition result is transmitted to the display device, and the interactive instruction is used to instruct the display device to perform the operation corresponding to the target recognition result.

[0035] The solution provided in this application enables speech recognition through a trained multimodal language model. This trained multimodal language model can combine user speech with the page images displayed on the display device to jointly recognize user speech, thereby obtaining more accurate recognition results for user speech. Furthermore, using this trained multimodal language model for voice interaction can also improve the accuracy of voice interaction and enhance the user experience.

[0036] In another aspect of this application embodiment, after transmitting the request information to the display device, the method further includes:

[0037] Input a recognition prompt into the multimodal language model. The recognition prompt is used to prompt the multimodal language model to pay attention to the historical browsing records and / or favorite records of the display device.

[0038] A third aspect of this application provides a server, including:

[0039] The controller is configured as follows:

[0040] Obtain training text for training a multimodal language model;

[0041] The training text is converted into corresponding training audio and control instructions, and the control instructions are used to control the display device to display the first page corresponding to the training text.

[0042] The sound playback device is triggered to play the training audio, and the control command is transmitted to the display device;

[0043] Training data is constructed based on the speech when the sound playback device plays the training audio and the first page image when the display device displays the first page according to the control command;

[0044] The multimodal language model is trained based on the training data.

[0045] A fourth aspect of this application provides a server, including:

[0046] The controller is equipped with a multimodal language model trained using the model training method described in the first aspect, and is configured to:

[0047] Receive target voice used to instruct the display device to perform voice interaction;

[0048] The request information is transmitted to the display device, the request information being used to request the acquisition of the second page image of the second page displayed by the display device;

[0049] The target speech and the second page image are recognized by the multimodal language model to obtain the target recognition result output by the multimodal language model;

[0050] The interactive instruction corresponding to the target recognition result is transmitted to the display device, and the interactive instruction is used to instruct the display device to perform the interactive operation corresponding to the target recognition result.

[0051] The solution provided in this application enables the training of a multimodal language model. The trained multimodal language model can combine user speech with the page images displayed on the display device to jointly recognize user speech. Therefore, compared to solutions that only recognize user speech, the trained multimodal language model can obtain more accurate recognition results for user speech. Furthermore, using this trained multimodal language model for voice interaction can improve the accuracy of voice interaction and enhance the user experience. Attached Figure Description

[0052] Figure 1 This is a schematic diagram of a voice interaction scenario;

[0053] Figure 2 This is a schematic diagram of the structure of a cascaded voice interaction system;

[0054] Figure 3 This is a schematic diagram of another voice interaction scenario;

[0055] Figure 4 This is a schematic diagram illustrating the workflow of a model training method according to some exemplary embodiments;

[0056] Figure 5 This is a schematic diagram illustrating a first page according to some exemplary embodiments;

[0057] Figure 6 This is a schematic diagram illustrating another first page according to some exemplary embodiments;

[0058] Figure 7 This is a schematic diagram illustrating a model training method according to some exemplary embodiments;

[0059] Figure 8 This is a schematic diagram illustrating the structure of a multimodal language model in a model training method according to some exemplary embodiments;

[0060] Figure 9 This is a schematic diagram illustrating the workflow of a model training method according to some other exemplary embodiments;

[0061] Figure 10 This is a schematic diagram illustrating the workflow of a model training method according to some other exemplary embodiments;

[0062] Figure 11 This is a schematic diagram illustrating the workflow of a voice interaction method according to some exemplary embodiments;

[0063] Figure 12 This is a schematic diagram illustrating the interaction between a display device and a second service according to some exemplary embodiments. Detailed Implementation

[0064] The embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described below do not represent all embodiments consistent with this application. They are merely examples of systems and methods consistent with some aspects of this application as detailed in the claims.

[0065] It should be noted that the brief descriptions of terms in this application are only for the convenience of understanding the embodiments described below, and are not intended to limit the embodiments of this application. Unless otherwise stated, these terms should be understood in their ordinary and common meaning.

[0066] The terms "first," "second," "third," etc., used in the specification, claims, and accompanying drawings of this application are used to distinguish similar or related objects or entities, and do not necessarily imply a specific order or sequence, unless otherwise specified. It should be understood that such terms are interchangeable where appropriate.

[0067] The terms “comprising” and “having”, and any variations thereof, are intended to cover but not exclude inclusion, for example, a product or device that includes a range of components is not necessarily limited to all of the components that are clearly listed, but may include other components that are not clearly listed or that are inherent to such product or device.

[0068] Voice interaction is a human-computer interaction method that allows users to speak to a terminal device. The terminal device then performs corresponding operations based on the voice recognition results of the user's voice, thus realizing voice interaction.

[0069] For example, if a terminal device that supports human-computer interaction is a display device, after receiving a user's voice, if the voice recognition result of the user's voice instructs the display device to play a certain media material, then the display device will play the media material according to the voice recognition result, thereby enabling the playback of the media material without the user having to operate a remote control.

[0070] See Figure 1The diagram illustrates a voice interaction scenario where a user can send a voice command to the display device to "play the news." Upon receiving the user's voice command, the display device can play the news based on the voice recognition results.

[0071] In some applications, display devices can achieve voice interaction through cascaded voice interaction systems. See also... Figure 2 The schematic diagram shown illustrates that this cascaded voice interaction system includes modules such as Automatic Speech Recognition (ASR), Natural Language Processing (NLP), and Text-to-Speech (TTS).

[0072] After receiving the user's voice, the display device transmits the voice to the cascaded voice interaction system. The ASR module, which has voice understanding capabilities, converts the user's voice into corresponding text by analyzing sound waves, recognizing phonemes, and matching words. The converted text is then input to the NLP module.

[0073] The NLP module parses the text converted from the user's speech, identifies the user's intent corresponding to the text, extracts key information from the user's intent, and generates corresponding response text. Through this response text, a dialogue with the user is realized through the display device.

[0074] Furthermore, this NLP module can be divided into Natural Language Understanding (NLU), Dialogue Management (DM), and Natural Language Generation (NLG) modules. The NLU module is used to understand the text converted from the user's speech, the DM module is responsible for understanding the context of the text, and the NLG module is used to generate the response text to be sent back to the user.

[0075] The TTS module is used to convert the response text into natural-sounding speech so that the display device can play the speech and respond to the user.

[0076] When the display device plays media resources based on the cascaded voice interaction system, the user can send a user voice of "Play XX" to the display device. Then, the display device transmits the user voice to the cascaded voice interaction system. The cascaded voice interaction system can identify the text corresponding to the user voice through the ASR module (this text is the voice recognition result). The display device can execute the operation corresponding to the text. If the text recognized by the cascaded voice interaction system is "Play XX", the display device plays the corresponding media resources according to the recognized text.

[0077] In some other application scenarios, the display device can implement voice interaction through a language recognition model. Among them, the language recognition model can be a large language model (LLM), etc.

[0078] In the case where the display device implements voice interaction through a language recognition model, refer to Figure 3 As shown in the schematic diagram of the scenario, after the display device 10 receives the user voice, it can transmit the user voice to the server 20 equipped with the language recognition model. The server 20 recognizes the user voice through the language recognition model to obtain the voice recognition result, and then transmits the voice recognition result to the display device 10. The display device 10 executes the operation corresponding to the voice recognition result, thereby realizing voice interaction.

[0079] From the above introduction, it can be seen that the accuracy of the recognition result of the user voice has an important impact on the realization of voice interaction. If the recognition result is inaccurate, it often leads to the display device being unable to perform the operations required by the user, resulting in a low accuracy of voice interaction.

[0080] When performing voice interaction through the above cascaded voice interaction system and language recognition model, only the user voice is used for voice recognition, resulting in the following possible situations during the voice interaction process:

[0081] (Situation 1) If a certain user voice instructs to play media resources, the user voice usually includes the name of the media resource to be played. However, the media resources are updated frequently, and some media resources are updated weekly or even daily. Therefore, there are often names of some new media resources, and some names are not common character combinations. In this case, both the cascaded voice interaction system and the language recognition model only perform voice recognition based on the user voice, and it is difficult to accurately recognize the user voice, resulting in the display device being difficult to determine the media resources to be played.

[0082] For example, the name of a certain movie or TV drama is he(4)yu(3)ye(4). The character combination corresponding to these syllables is relatively rare. It is difficult to determine the correct name of the movie or TV drama, "He Yu Ye", based only on the user voice including this pronunciation.

[0083] Alternatively, the user gives a voice instruction to play a certain song named cheng(1)ying(1)jue(2). Based solely on the user's voice including this pronunciation, it is difficult to determine that the correct song name is 《琤璎诀》.

[0084] (Case 2) The same or similar pronunciations may correspond to different meanings. Therefore, the same user voice may correspond to different operations that require the display device to perform. In this case, after receiving the user voice, both the cascaded voice interaction system and the language recognition model only perform speech recognition based on this user voice, and it is difficult to determine the recognition result that meets the user's needs, resulting in the display device being unable to perform the operations required by the user.

[0085] For example, if the user hopes to view the media resource "嘻兽界", the user's voice may be "我要看嘻兽界". However, during the speech recognition of the user's voice, it may be recognized as "我要看西兽界", resulting in the display device playing the media resource "西兽界".

[0086] For another example, some display devices can display pictures and texts. When displaying pictures, the user sometimes hopes that the display device shows different pictures, so they issue a voice of "下一张" or "上一张", hoping that the display device will jump to display the next picture or the previous picture after receiving this voice. In addition, when the display device is displaying text, the user sometimes hopes that the display device shows different chapters of text, so they issue a voice of "下一章" or "上一章", hoping that the display device will jump to display the next chapter of text or the previous chapter of text after receiving this voice.

[0087] The pronunciations of "上一张" and "上一章" are the same, and the pronunciations of "下一张" and "下一章" are the same. Therefore, when the user issues a voice of "shangyizhang" or "xiayizhang", the display device often cannot determine whether the operation the user hopes for is to jump between pages displaying different pictures or to jump between pages displaying different chapters of text. Among them, if the display device is currently displaying text and the user's intention is to hope that the display device shows the next chapter of text, and the recognition result for the user's voice is "下一张", resulting in the display device thinking that the operation to be performed is to jump to display the next picture. However, since the display device is displaying text, the display device cannot jump to display the next picture, resulting in an operation failure.

[0088] From this, it can be seen that in the case of performing speech recognition only through the user's voice, it will lead to a low accuracy of speech recognition. And the low accuracy of speech recognition will further lead to a low accuracy of voice interaction, resulting in the display device being unable to perform the operations required by the user after the user issues the user voice, reducing the user experience.

[0089] To address the aforementioned issues, embodiments of this application provide a model training method, a voice interaction method, and related apparatus to improve the accuracy of voice recognition of user speech, thereby also improving the accuracy of voice interaction and enhancing the user experience.

[0090] See Figure 4 This application provides a model training method. This model training method can be applied to a first server and includes the following steps:

[0091] Step S100: Obtain training text for training the multimodal language model.

[0092] A multimodal language model is an artificial intelligence model capable of recognizing information from at least two modalities, such as speech and images, to obtain recognition results. Because this multimodal language model combines speech and images simultaneously during speech recognition, it can more comprehensively understand and process human language, thereby improving the accuracy of speech recognition.

[0093] In this step, different types of training text can be set according to the operations performed by the display device. For example, training text can be set to make the display device play media assets, or training text can be set to make the display device jump between different pages.

[0094] Step S200: Convert the training text into corresponding training audio and control instructions. The control instructions are used to control the display device to display the first page corresponding to the training text.

[0095] In the case where the training text is used to instruct the display device to play the first media asset, the first page corresponding to the training text includes the media asset information of the first media asset.

[0096] The media asset information of the first media asset is used to distinguish it from other media assets. In some embodiments, the media asset information of the first media asset may include at least one of the following: the name of the first media asset, the introductory information of the first media asset, the poster of the first media asset, at least one frame of image in the first media asset, and the information of the cast and crew of the first media asset. In addition, the information of the cast and crew may include the name and / or photograph of the cast and crew.

[0097] In some scenarios, the first page can be the details page of the first media asset, which may include one or more of the following: the name of the first media asset, introductory information, poster, and the names of the performers.

[0098] In other scenarios, see Figure 5The first page may be a display page that includes multiple media assets, among which the first media asset is displayed. The display page may display the name of the first media asset, or display the name and poster of the first media asset.

[0099] exist Figure 5 In the display page of the device, there are six media assets: media asset A, media asset B, media asset C, media asset D, media asset E, and media asset F. The first media asset can be any one of these six media assets.

[0100] In other scenarios, the first page may include at least one frame of image from the first media asset. For example, the first page may display one or more stills from the first media asset, or other images other than stills that have been pre-extracted from the first media asset.

[0101] exist Figure 6 The first page is an image captured from the first media asset.

[0102] Additionally, when the training text is used to instruct the display device to jump between pages displaying different images, the first page includes at least one image.

[0103] For example, if the training text includes the character "previous image", the training text instructs the display device to jump from the page displaying the current image to the page displaying the previous image. If the training text includes the character "next image", the training text instructs the display device to jump from the page displaying the current image to the page displaying the next image. In both cases, the training text can be considered as instructing the display device to jump between pages displaying different images.

[0104] When the training text is used to instruct the display device to jump between pages displaying different chapters, the first page includes text content.

[0105] For example, if the training text includes the character "previous chapter", the training text instructs the display device to jump from the page displaying the text content of the current chapter to the page displaying the text content of the previous chapter. If the training text includes the character "next chapter", the training text instructs the display device to jump from the page displaying the text content of the current chapter to the page displaying the text content of the next chapter. In both cases, the training text can be considered as instructing the display device to jump between pages displaying different chapters.

[0106] Step S300: Trigger the sound playback device to play the training audio and transmit control commands to the display device.

[0107] In one feasible implementation, the first server may include a sound playback device, in which case the training audio can be played by the sound playback device in the first server.

[0108] Alternatively, in another feasible implementation, the sound playback device is independent of the first server; in this case, see [link to relevant documentation]. Figure 7 The diagram shows a scenario corresponding to the model training method. The first server can interact with the display device and the sound playback device respectively.

[0109] The first server can select the training text for this training application from a set of multiple training texts, convert the training text into training audio and control instructions, and then transmit the training audio to the sound playback device to trigger the sound playback device to play the training audio.

[0110] Upon receiving the training audio, the sound playback device can play the training audio. For example, the sound playback device may include any one of a speaker, a loudspeaker, and an artificial mouth device. Since the artificial mouth device plays audio at a higher volume, it is preferred as the sound playback device.

[0111] Of course, the sound playback device can also be other devices capable of playing audio, and this application does not limit it.

[0112] In addition, the first server can also transmit control commands to the display device to trigger the display device to display the first page. After receiving the control command, the display device can display the first page corresponding to the control command and obtain the first page image.

[0113] For example, the display device in the embodiments of this application can have various implementation forms. For example, it can be a television, a smart television, a laser projection device, a monitor, an electronic bulletin board, an electronic table, an in-vehicle display, a smart speaker that can display pages (such as a page including song titles or a page including song introduction information, etc.) and other display devices that can perform voice input.

[0114] Of course, this application does not limit the specific form of the actual device, and the display device can also be other display devices that support voice interaction functions.

[0115] If the training text is used to instruct the display device to play a certain media asset, the display device, upon receiving the control command, can display a first page containing the media asset information, such as a details page of the media asset, a display page of the media asset, or a playback page of the media asset.

[0116] If the training text is used to instruct the display device to jump between pages displaying different images, after receiving the control command, the display device can select at least one image from its corresponding media asset library and display it. The media asset library corresponding to the display device may include a local media asset library and / or a media asset library that the display device can access via a network.

[0117] In addition, if the training text is used to instruct the display device to jump between pages displaying different chapters, the display device can select a piece of text content from its corresponding media asset library and display it after receiving the control command.

[0118] Step S400: Construct training data based on the speech when the sound playback device plays the training audio and the first page image when the display device displays the first page according to the control command.

[0119] The first server can collect speech when the audio playback device plays training audio, or the audio can be collected by an audio acquisition device and then transmitted to the first server so that the first server can obtain the speech.

[0120] If the audio is captured by an audio capture device, in one feasible implementation, the audio capture device can be independent of the display device.

[0121] Alternatively, in another feasible implementation, the audio acquisition device may include the display device. In this case, when the sound playback device plays training audio, the display device may acquire the speech and then transmit the speech to the first server.

[0122] This implementation eliminates the need for additional equipment for voice acquisition; instead, it reuses the display device, reducing the cost of training multimodal models.

[0123] In addition, the display device can also transmit the first page image when displaying the first page to the first server so that the first server can obtain the first page image.

[0124] After receiving the voice and the first page image, the first server can construct training data including the voice and the first page image.

[0125] Step S500: Train the multimodal language model based on the training data.

[0126] See Figure 8 The schematic diagram shown indicates that the multimodal language model used in this application embodiment may include: a visual encoder 110, an audio encoder 120, an adapter 130, and a language model 140.

[0127] The visual encoder 110 can convert the features of the received image into corresponding feature blocks after receiving the image, and then stitch the feature blocks together to obtain the feature sequence corresponding to the image. The feature sequence corresponding to the image is then transmitted to the adapter 130.

[0128] In this embodiment of the application, during the training of the multimodal language model, the visual encoder 110 receives the first page image and transmits the feature sequence corresponding to the first page image to the adapter 130.

[0129] The audio encoder 120 can receive audio, extract the features corresponding to the audio, obtain the feature sequence corresponding to the audio, and then transmit the feature sequence corresponding to the audio to the adapter 130.

[0130] In this embodiment of the application, during the training of the multimodal language model, the audio encoder 120 receives the speech collected when the sound playback device plays the training audio, and transmits the feature sequence corresponding to the speech to the adapter 130.

[0131] After receiving the feature sequences corresponding to the image and the feature sequences corresponding to the audio, the adapter 130 converts these two feature sequences into a mode that the language model 140 can process, and then transmits the converted feature sequences to the language model 140.

[0132] After receiving the converted feature sequence transmitted by the adapter 130, the language model 140 can recognize the converted feature sequence and output the recognition result.

[0133] In one feasible implementation, the language model 140 can be a large language model (LLM). Of course, the language model 140 can also be other models, and this application does not limit it.

[0134] Furthermore, the multimodal language model used in the embodiments of this application may also be other architectures, and this application does not limit it.

[0135] By training a multimodal language model, user speech can be recognized in conjunction with the page images displayed on the screen. The page images displayed on the screen often reflect the actions the user needs the screen to perform, thus providing some direction for the actions the user instructs the screen to take. Therefore, compared to solutions that only recognize user speech, the trained multimodal language model can achieve more accurate recognition results. Furthermore, using this trained multimodal language model for voice interaction can improve the accuracy of voice interaction and enhance the user experience.

[0136] For example, if the display device's media asset library includes both "Xi Shou Jie" and "Xi Shou Jie," and the display device's page image shows the media asset information for "Xi Shou Jie," it indicates that the user is interested in this media asset. If the user's voice includes "bofangxishoujie," then through the display device's page image and the user's voice, the trained multimodal language model can determine that the user's voice recognition result is "play Xi Shou Jie," not "play Xi Shou Jie." It can then control the display device to play the "Xi Shou Jie" media asset based on the recognition result, thereby improving the accuracy of user voice recognition and voice interaction, and enabling the playback of media assets that the user is interested in, thus enhancing the user experience.

[0137] For example, if the display device shows a picture on the page and the received user voice includes "xiayizhang", the trained multimodal language model can determine the recognition result of the user voice as "next picture"; if the display device shows a piece of text on the page and the received user voice includes "xiayizhang", the trained multimodal language model can determine the recognition result of the user voice as "next chapter".

[0138] In step S100, an operation is provided to obtain training text for training the multimodal language model, which can be implemented in a variety of ways.

[0139] In one approach, the first server generates corresponding training text based on user input. In another approach, see [link to related documentation]. Figure 9 This operation may include the following steps:

[0140] Step S110: Based on the first operation that the display device supports to be performed through voice interaction, determine the sentence structure that conforms to the grammatical structure corresponding to the first operation.

[0141] In this embodiment, the first operation is an operation that can be performed by the display device through user voice instruction, that is, an operation that the display device can perform through voice interaction. For example, the first operation may include at least one of the following operations: instructing the display device to play media assets, instructing the display device to jump between pages displaying different images, and instructing the display device to jump between pages displaying different chapters.

[0142] Of course, the first operation may include other operations, which are not limited in this application.

[0143] Step S120: Generate training text based on the grammatically correct sentence structure corresponding to the first operation.

[0144] In the case where the first operation is to instruct the display device to play media assets, this operation can be implemented in the following ways:

[0145] First, based on the account logged into the media playback platform on the display device and / or the media assets stored on the display device, determine the media assets that the display device can play.

[0146] The media assets that the display device can play include media assets that the media playback platform allows the display device to play when the display device accesses the media playback platform through the account, and / or media assets stored by the display device itself.

[0147] Then, training text is generated based on the media assets that the display device can play and the grammatically correct sentence structure corresponding to the first operation.

[0148] For example, if the sentence structure corresponding to the first operation is "play + media asset name", and the media assets that the display device can play include news, then the generated training text can be text containing characters such as "play news".

[0149] Through the solution of this application embodiment, the first server can generate corresponding training text based on the media assets that the display device can play and the sentence structure that conforms to the grammatical structure corresponding to the first operation. This allows the first server to automatically generate training text without the need for manual input of the training text, simplifying manual operation, thereby improving the efficiency of the first server in obtaining training text and further improving the efficiency of training multimodal language models.

[0150] In the above embodiments, step S500 discloses the operation of training a multimodal language model based on training data. In some embodiments, see [link to relevant documentation]. Figure 10 This operation can be achieved through the following steps:

[0151] Step S510: Input training data into the multimodal language model.

[0152] The training data includes the audio of the training audio played by the audio playback device, and the image of the first page when the display device displays the first page.

[0153] Step S520: Input training prompts into the multimodal language model. When the training text is used to instruct the display device to play media, the training prompts are used to prompt the multimodal language model to pay attention to the display device's browsing history and / or favorites.

[0154] In this step, if the training text is not used to instruct the display device to play media, the training prompt can be used to prompt the multimodal language model to recognize the training data.

[0155] In addition, if the training text is used to instruct the display device to play media assets, the training prompt is not only used to prompt the multimodal language model to recognize the training data, but also to prompt the multimodal language model to pay attention to the display device's historical browsing records and / or favorite records during the recognition process.

[0156] In some scenarios, the results identified by the multimodal language model based on the training data may contain errors. In such cases, the multimodal language model can further correct the identification results based on the historical browsing records and / or favorite records of the display device.

[0157] For example, in a certain scenario, the multimodal language model initially determines the recognition result based on the training data, which includes "Yumu Village Folk Customs". If the display device has the media asset "Yumu Village Folk Customs" in its collection record, the multimodal language model can correct the initially determined recognition result and obtain the recognition result including "Yumu Village Folk Customs".

[0158] Alternatively, in some scenarios, the media assets that the multimodal language model identifies as needing to be played by the display device may include multiple versions, but the training data does not show which version needs to be played. In such cases, the multimodal language model can determine the version that needs to be played based on browsing history and / or favorites, thereby making the recognition results more in line with user needs, improving the accuracy of voice interaction, and further enhancing the user experience.

[0159] For example, if the asset "Xi Shou Jie" includes multiple seasons, and the audio in the training data is "bofangxishoujie," while the first page image includes asset information for both seasons one and two of "Xi Shou Jie" (such as the name and poster), then based on this training data, the multimodal language model determines the recognition result as "Play Xi Shou Jie." Furthermore, if the display device's browsing history shows that the user viewed both seasons one and two of this asset, and all content from season one has been viewed while some content from season two remains unviewed, the multimodal language recognition model can determine the recognition result as "Play Xi Shou Jie Season 2," thus making the recognition result more closely match the user's needs.

[0160] Step S530: Obtain the output of the multimodal language model after receiving the training data and training prompts.

[0161] The output result is the recognition result of the multimodal language model based on the training data.

[0162] Step S540: Adjust the multimodal language model based on the output results.

[0163] In this step, the deviation between the output result and the training text can be determined. If the deviation is greater than the deviation threshold, the multimodal language model can be adjusted. After adjustment, training data and training prompts are input into the multimodal language model again, and the deviation between the output result of the multimodal language model and the training text is determined again. If the deviation is greater than the deviation threshold, the multimodal language model is adjusted again until the deviation between the output result of the multimodal language model and the training text is less than the deviation threshold, or the number of times the multimodal language model is adjusted reaches a preset adjustment number threshold.

[0164] The solution provided in this application embodiment enables the adjustment of a multimodal language model. Furthermore, in this solution, upon receiving the training prompt, the multimodal language model pays attention to browsing history and / or favorites records during the recognition process. These browsing history and / or favorites records reflect the user's personalized behavior when using the display device, thereby making the recognition results of the multimodal language model more aligned with the user's needs and improving the user experience.

[0165] In another embodiment of this application, a voice interaction method is provided. This voice interaction method is applied to a second server, and the second server is equipped with a multimodal language model trained using the aforementioned model training method.

[0166] The first server and the second server may be the same server, or the first server and the second server may be independent servers. This application does not limit this.

[0167] See Figure 11 The voice interaction method includes the following steps:

[0168] Step S600: Receive target voice for instructing the display device to perform voice interaction.

[0169] In some scenarios, the target language can be captured by a display device and then transmitted from the display device to a second server.

[0170] Step S700: Transmit request information to the display device. The request information is used to request the acquisition of the second page image of the second page displayed by the display device.

[0171] In this step, after the second server transmits the request information to the display device, the display device will obtain the second page image of the second page it is currently displaying and transmit the second page image to the second server.

[0172] Step S800: Recognize the target speech and the second page image using a multimodal language model, and obtain the target recognition result output by the multimodal language model.

[0173] Step S900: Transmit the interactive instruction corresponding to the target recognition result to the display device. The interactive instruction is used to instruct the display device to perform the interactive operation corresponding to the target recognition result.

[0174] In this embodiment, after obtaining the target recognition result, the second server generates a corresponding interaction command based on the target recognition result and then transmits the interaction command to the display device so that the display device can execute the interactive operation corresponding to the target recognition result, thereby realizing voice interaction. Furthermore, this solution utilizes both the display device's page image and the user's voice when determining the target recognition result. Compared to solutions that only recognize the user's voice, the accuracy of the target recognition result is higher, which correspondingly improves the accuracy of voice interaction and enhances the user experience.

[0175] To clarify the interaction process between the second server and the display device in this application, the disclosure... Figure 12 See Figure 12 The voice interaction method provided in this application includes the following steps:

[0176] Step S121: The display device acquires the target voice sent by the user for voice interaction.

[0177] Step S122: The display device transmits the target voice to the second server.

[0178] Step S123: After receiving the target voice, the second server transmits request information to the display device.

[0179] The request information is used to request the image of the second page currently displayed on the display device.

[0180] Step S124: After receiving the request information, the display device takes a screenshot of the second page it is currently displaying and obtains the image of the second page.

[0181] Step S125: The display device transmits the second page image to the second server.

[0182] Step S126: The second server transmits the second page image and the target speech to the multimodal language model and obtains the output of the multimodal language model, which is the target recognition result.

[0183] Step S127: The second server determines the interaction command corresponding to the target recognition result and transmits the interaction command to the display device.

[0184] Step S128: The display device performs the operation corresponding to the target recognition result according to the interactive instruction.

[0185] The solution implemented through steps S121 to S128 of this application embodiment not only enables voice interaction, but also utilizes both the page image of the display device and the user's voice for voice recognition. Consequently, the accuracy of the target recognition result is higher, which improves the accuracy of voice recognition and enhances the user experience.

[0186] In another embodiment of this application, in addition to steps S600 to S900, the following steps are also included:

[0187] After step S700, that is, after transmitting the request information to the display device, an identification prompt is input to the multimodal language model. The identification prompt is used to prompt the multimodal language model to pay attention to the display device's browsing history and / or favorites.

[0188] Through this step, the multimodal language model also considers the historical browsing and / or favorites records of the display device when determining the target recognition result. These historical browsing and / or favorites records can reflect the user's personalized behavior when using the display device, thereby making the recognition result of the multimodal language model more in line with the user's needs and improving the user experience.

[0189] Another embodiment of this application provides a server that includes a controller configured to perform the following operations:

[0190] Obtain training text for training a multimodal language model;

[0191] The training text is converted into corresponding training audio and control instructions, and the control instructions are used to control the display device to display the first page corresponding to the training text.

[0192] The sound playback device is triggered to play the training audio, and the control command is transmitted to the display device;

[0193] Training data is constructed based on the speech when the sound playback device plays the training audio and the first page image when the display device displays the first page according to the control command;

[0194] The multimodal language model is trained based on the training data.

[0195] In some embodiments, the controller performs the action of acquiring training text for training a multimodal language model, specifically configured as follows:

[0196] Based on the first operation that the display device supports can be performed via voice interaction, determine the sentence structure that conforms to the grammatical structure corresponding to the first operation;

[0197] The training text is generated based on the grammatically correct sentence structure corresponding to the first operation.

[0198] In some embodiments, the controller generates the training text by executing a grammatically correct sentence structure corresponding to the first operation, specifically configured as follows:

[0199] When the first operation is to instruct the display device to play media assets, the media assets that the display device can play are determined based on the account logged into the media asset playback platform on the display device and / or the media assets stored on the display device;

[0200] The training text is generated based on the media assets playable by the display device and the grammatically correct sentence structure corresponding to the first operation.

[0201] In some embodiments, when the training text is used to instruct the display device to play the first media asset, the first page includes media asset information of the first media asset;

[0202] In cases where the training text is used to instruct the display device to jump between pages displaying different images, the first page includes at least one image;

[0203] When the training text is used to instruct the display device to jump between pages displaying different chapters, the first page includes text content.

[0204] In some embodiments, the media asset information of the first media asset includes at least one of the following: the name of the first media asset, the introductory information of the first media asset, the poster of the first media asset, at least one frame of the first media asset, and the information of the cast and crew of the first media asset, wherein the information of the cast and crew of the first media asset includes the names and / or photos of the cast and crew.

[0205] In some embodiments, the controller performs the training of the multimodal language model based on the training data, specifically configured as follows:

[0206] The training data is input into the multimodal language model;

[0207] Input training prompts into the multimodal language model. When the training text is used to instruct the display device to play media assets, the training prompts are used to prompt the multimodal language model to pay attention to the display device's browsing history and / or favorites.

[0208] Obtain the output of the multimodal language model after receiving the training data and the training prompts;

[0209] Based on the output results, the multimodal language model is adjusted.

[0210] This server can serve as the first server in the above embodiments. Through this server, a multimodal language model can be trained. The trained multimodal language model can then combine the user's voice with the page images displayed on the display device to jointly recognize the user's voice. Compared to solutions that only recognize the user's voice, the trained multimodal language model can obtain more accurate recognition results for the user's voice. Furthermore, using this trained multimodal language model for voice interaction can improve the accuracy of voice interaction and enhance the user experience.

[0211] Another embodiment of this application provides a server that includes a controller configured to perform the following operations:

[0212] Receive target voice used to instruct the display device to perform voice interaction;

[0213] The request information is transmitted to the display device, the request information being used to request the acquisition of the second page image of the second page displayed by the display device;

[0214] The target speech and the second page image are recognized by the multimodal language model to obtain the target recognition result output by the multimodal language model;

[0215] The interactive instruction corresponding to the target recognition result is transmitted to the display device, and the interactive instruction is used to instruct the display device to perform the operation corresponding to the target recognition result.

[0216] In some embodiments, after transmitting the request information to the display device, the controller is further configured to perform the following operations:

[0217] Input a recognition prompt into the multimodal language model. The recognition prompt is used to prompt the multimodal language model to pay attention to the historical browsing records and / or favorite records of the display device.

[0218] This server can serve as the first server in the above embodiments, enabling voice interaction. Furthermore, when determining the target recognition result, the server utilizes both the display device's page image and the user's voice simultaneously. Compared to a scheme that only recognizes the user's voice, the accuracy of the target recognition result is higher, thereby improving the accuracy of voice interaction and enhancing the user experience.

[0219] For ease of explanation, the above description has been provided in conjunction with specific embodiments. However, the discussion in some embodiments is not intended to be exhaustive or to limit the embodiments to the specific forms disclosed above. Various modifications and variations can be made based on the above teachings. The selection and description of the above embodiments are for the purpose of better explaining the contents of this disclosure, thereby enabling those skilled in the art to better use the embodiments.

Claims

1. A model training method, characterized in that, Applied to the first server, the model training method includes: Obtain training text for training a multimodal language model; The training text is converted into corresponding training audio and control instructions, and the control instructions are used to control the display device to display the first page corresponding to the training text. The sound playback device is triggered to play the training audio, and the control command is transmitted to the display device; Training data is constructed based on the speech when the sound playback device plays the training audio and the first page image when the display device displays the first page according to the control command; The multimodal language model is trained based on the training data.

2. The method according to claim 1, characterized in that, The acquisition of training text for training the multimodal language model includes: Based on the first operation that the display device supports can be performed via voice interaction, determine the sentence structure that conforms to the grammatical structure corresponding to the first operation; The training text is generated based on the grammatically correct sentence structure corresponding to the first operation.

3. The method according to claim 2, characterized in that, The step of generating the training text based on the grammatically correct sentence structure corresponding to the first operation includes: When the first operation is to instruct the display device to play media assets, the media assets that the display device can play are determined based on the account logged into the media asset playback platform on the display device and / or the media assets stored on the display device; The training text is generated based on the media assets playable by the display device and the grammatically correct sentence structure corresponding to the first operation.

4. The method according to claim 1, characterized in that, When the training text is used to instruct the display device to play the first media asset, the first page includes the media asset information of the first media asset; In cases where the training text is used to instruct the display device to jump between pages displaying different images, the first page includes at least one image; When the training text is used to instruct the display device to jump between pages displaying different chapters, the first page includes text content.

5. The method according to claim 4, characterized in that, The media asset information of the first media asset includes at least one of the following: the name of the first media asset, the introductory information of the first media asset, the poster of the first media asset, at least one frame of the first media asset, and the information of the cast and crew of the first media asset, wherein the information of the cast and crew of the first media asset includes the names and / or photos of the cast and crew.

6. The method according to claim 1, characterized in that, The training of the multimodal language model based on the training data includes: The training data is input into the multimodal language model; Input training prompts into the multimodal language model. When the training text is used to instruct the display device to play media assets, the training prompts are used to prompt the multimodal language model to pay attention to the display device's browsing history and / or favorites. Obtain the output of the multimodal language model after receiving the training data and the training prompts; Based on the output results, the multimodal language model is adjusted.

7. A voice interaction method, characterized in that, The method is applied to a second server, which is equipped with a multimodal language model trained by the model training method according to any one of claims 1 to 6, wherein the voice interaction method includes: Receive target voice used to instruct the display device to perform voice interaction; The request information is transmitted to the display device, the request information being used to request the acquisition of the second page image of the second page displayed by the display device; The target speech and the second page image are recognized by the multimodal language model to obtain the target recognition result output by the multimodal language model; The interactive instruction corresponding to the target recognition result is transmitted to the display device, and the interactive instruction is used to instruct the display device to perform the operation corresponding to the target recognition result.

8. The method according to claim 7, characterized in that, After transmitting the request information to the display device, the method further includes: Input a recognition prompt into the multimodal language model. The recognition prompt is used to prompt the multimodal language model to pay attention to the historical browsing records and / or favorite records of the display device.

9. A server, characterized in that, include: The controller is configured as follows: Obtain training text for training a multimodal language model; The training text is converted into corresponding training audio and control instructions, and the control instructions are used to control the display device to display the first page corresponding to the training text. The sound playback device is triggered to play the training audio, and the control command is transmitted to the display device; Training data is constructed based on the speech when the sound playback device plays the training audio and the first page image when the display device displays the first page according to the control command; The multimodal language model is trained based on the training data.

10. A server, characterized in that, include: A controller, equipped with a multimodal language model trained by the model training method according to any one of claims 1 to 6, and configured to: Receive target voice used to instruct the display device to perform voice interaction; The request information is transmitted to the display device, the request information being used to request the acquisition of the second page image of the second page displayed by the display device; The target speech and the second page image are recognized by the multimodal language model to obtain the target recognition result output by the multimodal language model; The interactive instruction corresponding to the target recognition result is transmitted to the display device, and the interactive instruction is used to instruct the display device to perform the operation corresponding to the target recognition result.